Source-linked AI summary

Criticality Metrics for Automated Driving: A Review and Suitability Analysis of the State of the Art

Lukas Westhofen, Christian Neurohr, Tjark Koopmann, Martin Butz, Barbara Schütt, Fabian Utesch, Birte Kramer, Christian Gutenkunst, Eckard Böde

arXiv:2108.02403v2cs.ROcs.SE

TL;DR

Automated driving needs reliable behavioral-safety assessment, yet selecting criticality metrics for open-context applications remains unresolved. This paper unifies the state of the art and proposes a suitability analysis for matching metrics to application requirements, yielding a structured framework and blueprint for metric selection. The approach also highlights scope limitations, including modeling assumptions and the need to account for evolving metrics and automated-vehicle capabilities.

  • Problem

    Selecting criticality metrics that suit a given automated-driving application remains difficult because metric validity depends on properties, scenario types, and application requirements.

  • Method

    The paper conducts an extensive unified review of criticality metrics, their applications, properties, and interrelations, and uses it to construct a suitability-analysis method.

  • Results

    The paper provides researchers and engineers with structured, visualized information and a blueprint for methodically choosing suitable criticality metrics for automated-vehicle applications.

  • Takeaways & Limitations

    The framework supports deliberate metric selection and statements about scenario coverage when computing criticality within an application.

  • Takeaways & Limitations

    The review's model treatment does not consider motion models with noise, uncertainties, or discretization errors, and metric suitability can change as new metrics and applications emerge.

Abstract

from arXiv · show

The large-scale deployment of automated vehicles on public roads has the potential to vastly change the transportation modalities of today's society. Although this pursuit has been initiated decades ago, there still exist open challenges in reliably ensuring that such vehicles operate safely in open contexts. While functional safety is a well-established concept, the question of measuring the behavioral safety of a vehicle remains subject to research. One way to both objectively and computationally analyze traffic conflicts is the development and utilization of so-called criticality metrics. Contemporary approaches have leveraged the potential of criticality metrics in various applications related to automated driving, e.g. for computationally assessing the dynamic risk or filtering large data sets to build scenario catalogs. As a prerequisite to systematically choose adequate criticality metrics for such applications, we extensively review the state of the art of criticality metrics, their properties, and their applications in the context of automated driving. Based on this review, we propose a suitability analysis as a methodical tool to be used by practitioners. Both the proposed method and the state of the art review can then be harnessed to select well-suited measurement tools that cover an application's requirements, as demonstrated by an exemplary execution of the analysis. Ultimately, efficient, valid, and reliable measurements of an automated vehicle's safety performance are a key requirement for demonstrating its trustworthiness.

1 Introduction

Automated-driving safety requires behavioral-safety assessment beyond functional safety, but choosing criticality metrics is difficult because applications and metric capabilities vary. The paper reviews these metrics and proposes a suitability-analysis blueprint for selecting combinations that match application requirements.

  • Safe deployment of automated vehicles requires assuring and demonstrating safe behavior in open, mixed-traffic contexts beyond functional safety.
  • Criticality metrics provide computational surrogates for assessing traffic criticality in applications including planning, test-case assessment, and scenario-catalog construction.
  • Selecting metrics is challenging because their properties, validity, and scenario applicability vary across applications.
  • 1.3 Overview of the Work: The paper provides an extensive review and a suitability-analysis blueprint, then evaluates the approach through an exemplary application.
  • 1.1 Motivating Example: Adequate safety assessment may require combining several metrics because individual metrics detect only particular aspects of criticality.
  • 1.1 Motivating Example: TTC with a single-point kinematics model is suitable for car-following but has reduced validity or sensitivity in intersection and other nonfollowing scenarios.
  • 1.2 Problem Statement: The proposed selection of metrics enables statements about their coverage of an application's scenario space while accommodating requirements such as output scales and runtime capabilities.
  • 1.3 Overview of the Work: Because new metrics and applications will continue to emerge, the authors invite continued refinement of the catalog and custom suitability analyses.

2 Related Work

Earlier traffic-conflict research established safety-surrogate indicators, but automated driving introduces additional requirements and applications. Existing AV-focused work does not provide a comprehensive comparison and systematic method for selecting suitable metric sets.

  • Traffic-conflict indicators originated in accident-surrogate research and were developed for traffic safety and vehicle-safety analysis.
  • Prior indicator studies did not address automated-vehicle applications, which impose distinct requirements such as formal rigor and runtime capability.
  • In automated driving, safety-surrogate indicators are commonly called criticality metrics, but existing overviews cover only subsets or restricted application settings.
  • Previous work compared neither the advantages and disadvantages of individual approaches nor a systematic procedure for deriving suitable metric sets.
  • Junietz systematically derived requirements for two applications, while newer applications motivate reconsidering and extending the requirement set.

3 Applications

Criticality metrics support computational applications across automated-driving development, testing, analysis, and deployment. The section organizes these applications along the V-model, including safety-oriented implementation, scenario elicitation, testing, and safety argumentation.

  • Criticality metrics quantify aspects of criticality for computational applications such as simulation, planning, and automated-driving analysis.
  • The identified applications are organized along the V-model to derive metric properties from application requirements.
  • Automated-driving functions: Criticality metrics can support safety-oriented function optimization and runtime risk assessment, including evasive reactions.
  • Verification and validation: Testing applications use metrics to define pass/fail criteria, guide search-based testing, and evaluate test-case performance.
  • Scenario elicitation: Scenario elicitation uses criticality metrics to classify scenarios, instantiate safety-relevant cases, and identify relevant data from large driving datasets.
  • Safety argumentation: Safety argumentation uses criticality metrics to quantify hazardous situations and support claims about risk-mitigation effectiveness.

4 Properties of Criticality Metrics

The paper derives metric properties from application needs and uses them as requirements for suitability analysis. These properties cover computational inputs and outputs, subjects and scenarios, measurement quality, classification performance, and prediction models.

  • Metric suitability depends on application-derived requirements for properties such as runtime computation, target values, inputs, outputs, and scenario types.
  • Subject type: Subject type distinguishes metrics applied to human subjects or automated systems, which matters especially in mixed traffic.
  • Inputs and outputs: Inputs may include actor states, road geometry, infrastructure, dynamic objects, and weather, while available inputs depend on the application context.
  • Measurement quality: Reliability concerns consistency under repeated or slightly changed scenes, while validity concerns correspondence with accident probability and severity within a scenario type.
  • Classification performance: Sensitivity measures correctly identified critical situations, and an over-approximating metric has sensitivity one while also identifying uncritical situations.
  • Prediction models: Prediction models differ in prediction horizon and whether they represent one future evolution or multiple branching developments.

5 Models and Metrics

The paper surveys functions that measure traffic-safety aspects and uses this state-of-the-art collection to prepare its suitability analysis. Because the review abstracts away implementation details, several measurement properties cannot be instantiated concretely.

  • The review covers a wide range of functions measuring different aspects of traffic safety.
  • Collecting the state of the art is presented as preparation for executing the suitability-analysis method.
  • The metric review remains highly abstract because implementation assumptions are limited.
  • Reliability, validity, sensitivity, and specificity cannot generally be instantiated at a concrete level without specific metric implementation details.

5.1 Employed Models

The employed models provide alternative ways to predict actor motion and future scene developments for criticality metrics. They range from simple kinematic and vehicle models to potential-based and stochastic reachable-set models.

  • Prediction models generate possible future developments of a scene, and their validity can strongly influence metric measurements.
  • Real-world prediction must account for process and measurement noise, discretization, and limits of relative-motion or coordinate-system representations.
  • Kinematic models: Single-point kinematics approximates motion-variable development with a Taylor polynomial.
  • Vehicle models: The simple-car model maps speed and steering-angle inputs to vehicle position and direction, with axle distance as a model parameter.
  • Vehicle models: Continuous-steering, coordinated-turn, and augmented coordinated-turn models represent vehicle or actor motion with progressively different steering, turn-rate, and acceleration assumptions.
  • Vehicle models: One-track and two-track models represent vehicle dynamics, with the two-track model adding individual-tire dynamics and variable speed.
  • Predictive models: Potential-based models combine object-specific potential functions, while stochastic reachable sets vary control inputs and model multiple possible trajectories.

5.2 Criticality Metrics

The review organizes criticality metrics by their measurement scope, aggregation, prediction-model interfaces, and application-specific properties. It illustrates how metrics such as TTC and its variants support assessment, filtering, and retrospective analysis while remaining sensitive to scenario and modeling assumptions.

  • Review approach: The review presents a wide variety of criticality metrics with descriptions, formulae, constraints, and properties of variables estimated by prediction models.The review uses idealized formulae and clear prediction-model interfaces to abstract suitability analysis toward measurement principles rather than specific implementations.
  • Metric scope: Criticality metrics are distinguished as scene-level functions evaluated at a specific time and scenario-level functions evaluated over a time series.Scene-level metrics can be aggregated over time to derive scenario-level metrics, but the reverse does not generally hold.
  • Aggregation: Metrics defined for one or two actors can be extended to arbitrary actors through aggregation, with the appropriate aggregate depending on whether the assessment is actor-specific or impartial.A designated actor can use a maximum over other actors, whereas an impartial scene assessment can use a mean over actor pairs.
  • Time-based metrics: TTC returns the minimum predicted time until two actors collide, or infinity when their predicted trajectories do not intersect.TTC supports car-following assessment and retrospective aggregation, but its validity is reduced in many intersection scenarios and when multi-actor aggregation is not meaningful.
  • TTC variants: TTC variants extend the basic measure through severity weighting, constrained motion assumptions, distance-dependent encounters, time aggregation, or multiple predicted traces.WTTC is designed for selective data recording and filtering, while PTTC reduces computational cost by imposing scenario and motion constraints at the expense of validity.
  • Time aggregation: TIT aggregates the difference between TTC and a target value over time, thereby reflecting criticality more accurately than TET.TET instead measures the duration for which TTC remains below the target and can be adapted to other metrics.

5.2.12 Post Encroachment Time (PET)

The section describes PET and related measures for temporal separation at conflict areas, alongside metrics that quantify required evasive action, accepted gaps, personal-space conflicts, and maneuver feasibility. These measures differ in whether they use retrospective observations, predictions, spatial requirements, or available control authority.

  • Post Encroachment Time: PET measures the time gap between one actor leaving and another entering a designated conflict area.GT and IAPE are semi-predictive PET variants that predict the second actor’s entry using a constant-velocity model.
  • Predictive encroachment measures: PrET predicts PET with respect to an anticipated intersection point, while SPrET down-weights situations occurring long before the predicted intersection.SPrET thereby incorporates prediction uncertainty, and TA is a constant-velocity special case of PrET.
  • Accepted Gap Size: The AGS measures the predicted spatial gap required for an actor to act, such as entering a pedestrian stream or continuing through a gap.The associated action model predicts whether the actor decides to act given the gap size; larger desired distances generally indicate greater criticality.
  • Required acceleration: Required longitudinal and lateral acceleration quantify the minimum evasive acceleration needed to avoid a future collision.The longitudinal measure can represent braking or positive acceleration, while the lateral measure considers the minimum absolute steering acceleration.
  • Combined metrics: The conditional required acceleration combines required acceleration with SPrET so that dynamic criticality becomes relevant when temporal criticality is present.The paper presents this combination as an example of creating new metrics from existing measures and target values, potentially improving validity by addressing different criticality aspects.
  • Maneuver feasibility: STN divides required lateral acceleration by the maximum lateral acceleration available to the actor, with STN ≥1 indicating that lateral avoidance cannot prevent an impending accident.BTN uses the analogous longitudinal relationship, and DSSM adds braking and reaction-time assumptions for car-following scenarios.
  • Personal-space conflicts: The SOI counts violations of actors’ defined personal spaces over the analyzed period.A violation occurs when another actor’s personal space intersects the designated actor’s personal space.

5.2.24 Pedestrian Risk Index (PRI)

The section covers metrics that estimate pedestrian, vehicle, and scene criticality through temporal conflict conditions, probabilistic collision models, constrained maneuver difficulty, and potential-field functions. It emphasizes that these measures combine predicted interactions with assumptions about motion, uncertainty, and available actions.

  • Pedestrian Risk Index: PRI combines TTZ and impact speed to estimate conflict probability and severity in pedestrian-crossing scenarios.It requires a coherent conflict period in which the pedestrian and vehicle approach the conflict area in a specified temporal order.
  • Probabilistic metrics: CPI estimates the average probability that a vehicle cannot avoid collision by comparing required longitudinal acceleration with a probabilistic minimum available acceleration.The target distribution can depend on road-surface material and vehicle brakes, and the aggregation is normalized over scenario duration.
  • Probabilistic metrics: ACI extends CPI to multiple conditions by constructing a probabilistic collision tree whose leaf nodes represent collision and non-collision outcomes.Each leaf risk combines the probability of satisfying its conditions with a binary collision value before scene-level aggregation.
  • Optimization-based metrics: TCI formulates criticality as the minimum difficulty of satisfying physical and regulatory constraints while avoiding obstacles.Its constraints include required longitudinal and lateral acceleration together with margins for speed and course-angle corrections.
  • Potential functions: Potential-function metrics assign functions to static and dynamic objects, sum them into a scene potential, and can use gradient descent to suggest criticality-reducing vehicle movement.Their properties depend strongly on the specified potential functions, whose definition also raises ethical questions about active decision making.

5.2.33 Safety Potential (SP)

Safety Potential (SP) evaluates collision-avoidance risk between actors by considering whether their safe-control-policy trajectories can intersect. The section also contrasts related metrics, their interpretability, application thresholds, and prediction-model limitations.

  • Safety Potential: SP identifies state combinations in which safe control policies for two actors can lead to a collision, then assigns a numeric safety valuation.It constructs occupied and reachable trajectory sets before defining an unsafe set and potential function.
  • Safety Potential: The SP potential is constrained to positive values for unsafe states and nonnegative values otherwise, while safety procedures cannot increase the potential.The ordering condition compares the current state with states resulting from both actors applying safety procedures.
  • Related Metrics: Accident Metric (AM) records whether any accident occurred, but fails to identify critical scenarios without accidents.It is a binary metric implicitly used in accident databases such as GIDAS.
  • Related Metrics: Collision-severity metrics such as Δv estimate injury or fatality relevance, while Collision Severity (CS) compares evasive-maneuver and predicted-collision severity and incorporates actor mass differences.The Δv literature connects speed change to severe-injury or fatality probability, whereas CS focuses solely on potential-collision severity.
  • Metric Properties: Metric target values are application-dependent: a threshold suitable for pruning minimal-risk maneuvers may be unsuitable for filtering large databases because of low sensitivity.The section gives examples including TTB thresholds of 0.4s, 0.6s, and 1s in different applications.
  • Metric Properties: TTB reliability depends on reliable collision prediction, because small prediction changes can make its value jump to infinity despite only slight changes in actual criticality.For automated vehicles, TTB specificity may also decrease when last-second steering can avoid situations that braking-based reasoning classifies as critical.

5.4 Interrelations of Criticality Metrics

The section maps interrelations among criticality metrics, including dependencies such as BTN relying on along,req. Figure 5 distinguishes scene-level from scenario-level metrics and highlights metrics independent of prediction models.

  • Metric interrelations: BTN depends on the along,req metric, illustrating that the reviewed criticality metrics form an interconnected network rather than isolated measures.These dependencies matter when interpreting metric choices during suitability analysis.
  • Metric interrelations: Figure 5 differentiates scenario-level and scene-level metrics and highlights metrics that do not rely on a prediction model.The visualization summarizes the interrelations developed from the metric descriptions.

6 Suitability Analysis

The suitability analysis is an expert-based process that matches application requirements with criticality metrics and models. An exemplary unprotected-left-turn application demonstrates iterative filtering and returns metrics that can support downstream testing.

  • Generic approach: The analysis seeks adequate metric–model assignments for a described application using available metrics, models, and either requirements or an enlarged candidate set.Its output is a set of suitable metrics assigned suitable models.
  • Generic approach: The five-step process identifies relevant properties, derives and orders requirements, assesses candidate metric–model properties, and iteratively discards failures.Iteration stops when requirements are satisfied with at least one candidate or no candidate remains.
  • Exemplary application: The exemplary application generates critical instances for an unprotected-left-turn scenario by simulating sampled parameters, fitting a regression model, and optimizing the learned model.The resulting parameter combinations are used as representative cases in downstream scenario-based testing.
  • Exemplary application: Requirements are selected over subject type, scenario type, inputs, output scale, reliability, and validity, then ordered according to their importance for the application.Scenario and subject domains are ranked highest, followed by validity and output scale.
  • Exemplary application: The example begins with 43 metrics and removes candidates mismatched to subject types, scenario applicability, or modeling assumptions.The analysis excludes pedestrian-focused metrics and considers restrictions from simple physics-based models and automation-driven behavior.
  • Exemplary application: The resulting metric set, annotated with output scales, can be combined or time-aggregated for fitting a regression model in downstream testing.Metrics removed during filtering may still serve as enhancement factors, such as multiplying a measurement by Δv.

7 Conclusion and Future Work

The paper consolidates criticality-metric knowledge and proposes suitability analysis for selecting metrics in automated-driving applications. It also identifies expert judgment, uncertainty, and the need for more preventative metrics as areas for future work.

  • Conclusion: The review unifies decades of criticality-metric research into a structured knowledge base for automated vehicles.Its sources span traffic conflict research, traffic psychology, and automated-vehicle development and testing.
  • Conclusion: The paper reviews applications, derives metric requirements, evaluates metric properties, and visualizes interrelations before proposing suitability analysis.An exemplary evaluation demonstrates the proposed method for a relevant application.
  • Conclusion: The resulting framework provides structured information and a blueprint for methodically choosing criticality metrics within automated-driving applications.The intended users are researchers and engineers working on automated vehicles.
  • Future work: Many property evaluations rely on expert judgment because traceable evidence is unavailable, motivating quantitative studies using synthetic and real-world data.Future evaluations should also incorporate how different underlying models influence criticality metrics and their properties.
  • Future work: Real-world measurement errors require uncertainty quantification because erroneous inputs propagate through metric computations and affect outputs.Interval arithmetic is identified as one option for tracking numerical error and quantifying its influence on metric outputs.
  • Future work: Most existing metrics measure physics-based symptoms near the end of an accident’s causal network, while formalized causal factors could support earlier detection.Examples of causal factors include environmental conditions, road-network complexity, and traffic-rule violations.

Appendix A: Index of Acronyms of Criticality Metrics

The appendix provides an index of acronyms for criticality metrics, covering measures based on acceleration, distance, time, collision potential, and related safety concepts.

  • Acceleration and braking: The index lists acceleration- and braking-related metrics, including Required Acceleration, Conditional Required Acceleration, and Deceleration Rate to Avoid Crash.
  • Conflict and risk: The index also covers collision, conflict, risk, trajectory, and exposure measures, including Crash Potential Index, Conflict Index, Trajectory Criticality Index, and Time Exposed.
  • Distance and encounter: It includes distance- and encounter-related measures such as Distance of Closest Encounter, Post Encroachment Time, and Time To Closest Encounter.

Appendix B: Requirements on the Properties of Criticality Metrics

The appendix presents a table of minimal application requirements for criticality-metric properties and lists the properties used to characterize such metrics.

  • Requirements table: Table 4 is titled “Minimal requirements of applications on the properties of criticality metrics” and includes an Application and Run-time heading.
  • Requirements table: The appendix associates application requirements with criticality-metric properties rather than presenting a single universal metric-selection criterion.
  • Metric properties: The listed metric properties include capability, target values, subject type, scenario type, inputs, output scale, reliability, validity, sensitivity, specificity, and prediction model.
Loading 2108.02403v2…