Source-linked AI summary
Virtual Testing of Automated Driving Systems through Credible Simulations
Riccardo Dona, Espedito Rusciano, Biagio Ciuffo
TL;DR
ADS complexity makes physical testing alone impractical, while validation-only accreditation does not adequately address heterogeneous, evolving simulation toolchains. The paper proposes a risk-based credibility framework organized around management, analysis, verification, and validation, linking requirements to intended use and decision criticality. It concludes that the framework supports flexible, proportional virtual-testing assessment across ADS and broader road-safety applications, while requiring suitable data, documentation, expertise, and further benchmarking.
Problem
ADS virtual testing increasingly supports safety decisions, but physical testing alone is impractical and validation-only accreditation is limited for complex, multi-tool simulation environments.
Method
The paper adapts NASA STD-7009 into a risk-based framework linking simulation credibility requirements to intended use and decision criticality across four assessment pillars.
Results
The framework accommodates diverse simulation strategies, including multi-fidelity and modular toolchains, and supports subsystem-level and integrated-system validation.
Takeaways & Limitations
The approach provides proportional credibility requirements for virtual testing across ADS safety assessment and broader road-safety simulation uses.
Takeaways & Limitations
Effectiveness depends on suitable validation data, transparent modelling documentation, and expertise, while broad benchmarking across toolchains and learning-based approaches remains incomplete.
Abstract
from arXiv · showhide
Simulation is increasingly used to support safety-related decision-making in road transport, particularly for the assessment and approval of automated driving systems (ADS). The complexity of ADS behavior and size of their operational design domains make exclusive reliance on physical testing impractical, leading to extensive use of virtual testing (VT) during the approval phase. This shift raises critical questions regarding the credibility of modelling and simulation (M&S) results used to support road safety decisions. Current VT accreditation approaches in the ADS domain typically rely on validation-only practices, which have been shown to scale poorly when applied to complex, multi-tool simulation environments. To address this limitation, this paper proposes a risk-based framework for assessing the credibility of simulation toolchains used in ADS safety evaluation, drawing inspiration from established practices in other safety-critical domains, notably NASA's STD-7009 for models and simulations. The framework extends traditional verification and validation (V&V) by explicitly linking credibility requirements to the intended use of simulation outputs and to the safety criticality of the decisions they support within the approval process. It provides a lifecycle-oriented assessment scheme integrating toolchain management, modelling assumptions and limitations, verification, validation, and sensitivity analysis. Credibility acceptance thresholds are defined proportionally, allowing differentiated requirements depending on whether simulation is used for exploratory safety analysis, partial decision support, or as a substitute for physical testing. While demonstrated for ADS, the proposed approach is directly applicable to road safety and simulation studies where VT plays a central role in safety assessment and regulatory decision-making.
1. Introduction
ADS complexity and vast scenario spaces make physical testing alone insufficient, increasing reliance on virtual testing for safety evidence. The paper frames M&S credibility as a risk-based, intended-use-dependent basis for integrating virtual testing into ADS certification.
- ADS complexity, AI, and near-infinite scenario exposure make physical testing alone insufficient for demonstrating safe operation.
- NATM combines organizational processes, safety-case development, multiple testing approaches, and in-service monitoring across the ADS lifecycle.
- Validation-only accreditation is limited because fixed error thresholds can be arbitrary and validated models may be misused outside their domains.
- NASA STD-7009 defines M&S credibility as trust in simulation results for a specific intended use and makes V&V proportional to decision criticality.
- UNECE incorporated M&S credibility into NATM’s test-environment assessment pillar to make virtual testing trustworthy while retaining flexibility.
2. Methodology
The framework was developed through literature review and structured consultations with academic and industrial SMEs. Its methodology addresses evolving toolchains, heterogeneous model fidelity, and the need for use-dependent credibility requirements.
- Literature review and structured consultations with academic and industrial SMEs identified challenges in qualifying ADS virtual-testing toolchains.
- Rapid toolchain evolution conflicts with slower regulatory and policy cycles, complicating stable qualification criteria.
- Sensor models range from abstract object lists to physics-based representations, making fixed-threshold validation unsuitable across heterogeneous toolchains.
- The resulting framework tailors NASA STD-7009 principles to automotive needs while addressing transparency, configuration control, and traceability.
3. Results
The proposed credibility framework organizes ADS simulation assessment around management, analysis, verification, and validation. Requirements are set according to intended use and toolchain criticality, with documented evidence and explicit treatment of assumptions, fidelity, and implementation correctness.
- The framework uses four sub-pillars—management, analysis, verification, and validation—to assess ADS simulation-toolchain credibility.
- Manufacturers document evidence in a credibility handbook, while assessors review it and independently access the integrated toolchain for virtual tests.
- Intended use ranges from exploratory analysis to direct approval support, and criticality links toolchain influence and decision consequences to assessment requirements.
- Criticality classes are red for high, yellow for medium, and green for low, with full assessment required, discretionary, or unnecessary respectively.
- M&S Management: The management pillar covers data pedigree, traceability, uncertainty, personnel competency, and controlled releases for certification evidence.
- M&S Analysis: The analysis pillar documents toolchain construction, modelling rationale, trade-offs across fidelity, and assumptions and limitations spanning black-box, grey-box, and white-box models.
- M&S Verification: Verification checks numerical and implementation correctness, whereas validation evaluates discrepancies between virtual outputs and real-world evidence.
4. Conclusion
The paper presents a risk-based credibility framework for qualifying ADS virtual-testing toolchains, linking assessment rigor to intended use and decision criticality. It offers flexible application across simulation strategies while acknowledging evidence, expertise, and benchmarking limitations.
- The framework links simulation credibility to intended use and decision criticality rather than treating credibility as binary.
- Four pillars—management, analysis, verification, and validation—provide a holistic alternative to validation-only accreditation.
- The approach accommodates multi-fidelity toolchains, modular models of models, and both subsystem-level and integrated-system validation.
- Stakeholder involvement from industry, research, and public authorities supports practical applicability while retaining methodological rigor.
- The framework depends on suitable validation data, transparent modelling assumptions, and expertise in criticality, uncertainty, and sensitivity analyses.
- It has not yet been systematically benchmarked across many independent toolchains or data-driven and learning-based modelling paradigms.