Source-linked AI summary
A distribution-free certification framework for trustworthy crash-severity prediction
Amir Rafe, Subasish Das
TL;DR
Crash-severity models lack finite-sample guarantees for noisy ordinal labels, heterogeneous subpopulations, and deployment settings unlike calibration data. The paper wraps any severity model in a distribution-free certification layer that combines ordinal, subgroup, measurement, shift, and risk guarantees. Across 5.2 million Texas records and seven base models, validity was identical, while vulnerable-road-user strata exposed undercoverage and a model-independent floor on set width.
Problem
Crash-severity models support operational decisions but lack finite-sample guarantees for ordinal outcomes, structured report-to-injury mismeasurement, heterogeneous groups, and deployment shift.
Method
The paper wraps any severity model unmodified with distribution-free certification using contiguous ordinal sets, pre-declared partitions, reporting bands, shift certificates, severity-weighted risk control, and an attributable slack budget.
Results
On 5.2 million Texas records across seven base models, validity was identical while six models undercovered unrestrained drivers and motorcyclists under marginal calibration.
Takeaways & Limitations
Certification exposes subgroup failures hidden by marginal coverage and separates a computable model-independent width floor from a bounded but distribution-free unidentifiable remainder.
Takeaways & Limitations
The informativeness floor cannot be certified distribution-free from below; computing a nonvacuous lower bound requires a declared reporting channel, and the remaining width is not identified.
Abstract
from arXiv · showhide
Crash-severity models inform screening, dispatch and site prioritization, yet are deployed without a finite-sample statement of what one prediction means. Off-the-shelf guarantees fail here, because the features that make crash severity distinctive defeat them: the KABCO outcome is ordinal, the recorded label is a field assessment agreeing with medical severity about half the time, erring in a structured way, and deployment crosses jurisdictions and years calibration never saw. We develop a certification layer that wraps any severity model unmodified, with distribution-free guarantees using this structure: contiguous ordinal sets that read as "B or worse"; per-class validity for any pre-declared partition, with an oracle efficiency characterization; transfer of coverage to unobserved true severity through a declared reporting band, with a worst-case sharpness result; a one-sided certificate under deployment shift; and severity-weighted risk control. The guarantees compose with an attributable slack budget. The same analysis bounds what certification can achieve. A certified set's informativeness is governed by a functional of the true law that no base model can evade and that cannot be lower-bounded distribution-free; given a declared misreporting channel identified from record-linkage data, a nonvacuous lower bound on that floor becomes computable. On 5.2 million Texas records across seven base models spanning four decades, the layer attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats, separating it from a remainder that stays bounded but distribution-free unidentifiable. The framework is released as an open-source package with theorem-level tests.
1. Introduction
Crash-severity models support consequential decisions but lack finite-sample guarantees, while ordinal outcomes, structured label noise, heterogeneity, and deployment shift defeat off-the-shelf coverage. The paper introduces a certification layer that uses this structure and characterizes both its guarantees and irreducible limits.
- Motivation: Crash-severity predictions guide funding, dispatch, and screening decisions, yet deployed outputs lack any finite-sample guarantee about what they mean.The gap is between increasingly sophisticated estimators and empty guarantees attached to individual predictions.
- Structural challenges: Ordinal KABCO outcomes require contiguous sets such as “B or worse,” because noncontiguous severity sets are operationally meaningless.The ordering turns bounded reporting error into bounded set expansion and preserves an actionable interpretation.
- Structural challenges: Police-reported severity is a structured, imperfect proxy for injury severity, agreeing with medical categories roughly half the time and tending toward adjacent-category under-reporting.A guarantee on the recorded label therefore concerns the report rather than necessarily the injury.
- Structural challenges: Marginal calibration can conceal failures in severe minority subpopulations, while jurisdictional and temporal shifts invalidate automatic transfer from calibration data.The paper identifies heterogeneity and deployment shift as distinct reasons to avoid relying on average in-distribution coverage.
- Contribution: The certification layer wraps any severity model unmodified and uses ordinal structure, reporting bands, pre-declared partitions, shift certificates, and risk control with an attributable slack budget.Its output is intended to retain finite-sample guarantees while reflecting the recorded-to-true-severity measurement process.
- Scope: The paper explicitly excludes causal claims: certification describes plausible outcomes given recorded data, whereas intervention effects require separate causal analysis.Decoded vehicle fields indicate equipment availability, not whether safety systems were engaged or effective.
2. Related work
Prior work supplies severity modeling, transportation uncertainty methods, and conformal foundations, but not their integration for noisy ordinal crash outcomes under heterogeneity and deployment shift. This paper assembles those ingredients into a framework with contiguous sets, subgroup validity, measurement-aware transfer, shift handling, and risk control.
- Severity modeling: Severity modeling contributes ordered-response, heterogeneity, temporal-instability, and flexible machine-learning traditions, which this paper treats as inputs rather than competitors.The certification layer can condition on latent classes and wrap a random-parameters ordered logit.
- Measurement: Police-reported KABCO labels differ from medically referenced severity in structured, agency-dependent ways, motivating a declared reporting-band approach instead of an assumed transferable noise kernel.The literature reports agreement close to half and disagreement concentrated toward under-reporting and nearby categories.
- Conformal foundations: Generic conformal applications demonstrate demand for uncertainty guarantees but inherit marginal coverage for the calibrated report under exchangeability, without guarantees for true injury severity or subpopulations.The paper positions such work as a predecessor while distinguishing application from foundations.
- Subgroup validity: Class-conditional calibration restores subgroup coverage but can suffer from thin calibration cells, whereas this framework uses minimum cell sizes and rollups to preserve thicker cells.The paper reports a thinnest cell of 4,125 calibration records, compared with 32 in the cited survey study.
- Measurement: Label-noise methods typically require stochastic noise rates or uniformity assumptions, while crash severity offers a more defensible stable ordinal support band.The paper uses the band structure because adjacent misreports are more plausible than large jumps and upward errors from K are rare.
- The gap: Existing literatures separately provide flexible severity models, transportation uncertainty applications, and conformal guarantees, but none addresses the paper’s combined deployment requirements.The missing combination includes true-severity coverage under a reporting band, heterogeneity-class validity, shift certification, severity-weighted risk control, and accountable slack.
3. Framework and guarantees
The framework creates distribution-free crash-severity certificates from calibration quantiles, while explicitly accounting for ordinal outputs, heterogeneity, reporting noise, and limits on informativeness.
- Framework workflow: The workflow treats fitted models as untrusted, creates guarantees through calibration quantiles and declared structure, and exports certificates with attributable slacks.Detection alarms trigger recalibration rather than editing an already issued certificate.
- Ordinal scores, interval sets, marginal validity: The ordinal interval construction preserves operational contiguity and makes reporting-noise expansion additive rather than multiplicative.The framework distinguishes validity from efficiency: contiguity is not required for the guarantee itself, but it keeps the certified expansion actionable and tighter.
- Ordinal scores, interval sets, marginal validity: Exact conditional coverage over individual covariates is impossible in general, so the framework targets finitely many transportation-relevant groups instead.The impossibility result motivates conditioning on pre-declared classes, jurisdictions, or time groups.
- Heterogeneity-conditional validity: Pre-declared groups receive per-class validity, and the oracle characterization identifies the smallest threshold-family predictor that avoids borrowing coverage across classes.Arbitrary partitions preserve validity, while partition quality affects efficiency; guarantees are marginal over each cell rather than simultaneous across cells.
- When certification is informative: The informativeness floor is a functional of the true joint law, independent of the base model, calibration split, and modeling choices.A level-set characterization makes the floor computable in principle, but empirical plug-in estimates remain evidence about the data regime rather than proofs of a high floor.
- When certification is informative: Every valid predictor inherits the floor, and when that floor is sufficiently large, composed interval certificates become vacuous regardless of model, training-data volume, or calibration scheme.The result converts model inadequacy into a question about the data-law functional, while also limiting what agreement among base models can establish.
3.5. The floor cannot be certified from below
The paper shows that certified-set informativeness has a distribution-free floor that cannot be lower-certified without assumptions, and that reporting-band transfer bounds are worst-case sharp.
- 3.5. The floor cannot be certified from below: The informativeness floor cannot be certified from below distribution-free because non-atomic covariates make diffuse and deterministic conditional laws observationally indistinguishable.The impossibility holds for every distribution-free lower certificate and is driven by the covariate law rather than the finite label space.
- 3.6. Coverage on true severity under reporting noise: Under a banded reporting map, coverage transfer expands contiguous sets additively by b−+ b+ rather than multiplicatively, preserving the operational “B or worse” reading.Without ordinal structure, compatibility can encompass the entire label space and the expansion becomes vacuous.
- 3.6. Coverage on true severity under reporting noise: The transfer constant 1−α−δ is sharp, and trimming either side of the declared expansion can reduce true-label coverage below the target.The sharpness result rules out both a larger universal constant and one-sided reductions of the expansion.
- 3.6.1. Numeric grounding of the band: In KABCO linkage evidence, adjacent under-reporting is rare beyond the declared band, while band width two can saturate the five-category scale and create vacuity.The reported beyond-band mass is approximately 0.0062 for one-category upward expansion and approximately 0.0115 after the constant-band downward component.
3.7. The reporting channel imposes an identifiable floor
A declared misreporting channel converts the otherwise non-certifiable informativeness floor into a computable lower bound, while retaining explicit assumptions and partial-identification caveats.
- 3.7. The reporting channel imposes an identifiable floor: A declared channel makes the floor’s lower end computable, changing what is impossible distribution-free into an identifiability statement about the data regime.The channel is treated as a policy input and swept over the set of channels consistent with linkage constraints and observed reported composition.
- 3.7. The reporting channel imposes an identifiable floor: The channel floor is the smallest expected set size needed to cover reported labels at level 1−α when a predictor knows true severity and may randomize.It is computed as a fractional-knapsack problem over channel probabilities, without imposing contiguity.
- 3.7. The reporting channel imposes an identifiable floor: The floor depends only on the stratum channel and true composition, so it is independent of the base model, feature set, and calibration scheme.This makes the quantity a model-independent lower bound on the width of every valid predictor under within-stratum nondifferential reporting.
- 3.7. The reporting channel imposes an identifiable floor: On strata with severity concentrated in smeared interior categories, the channel alone forces wide sets, explaining empirical vacuity independently of the predictor.The middle term remains bounded but not point-identified, while the channel-imposed floor supplies its certifiable lower end.
- 3.7. The reporting channel imposes an identifiable floor: Under differential reporting within a stratum, the floor degrades continuously by the stated total-variation slack rather than remaining exact.The assumption requires reporting to be covariate-invariant conditional on true severity within each stratum.
3.8. Spatiotemporal deployment shift
The framework handles observed strata through group-conditional calibration and new strata through weighted covariate-shift transfer, while exposing explicit diagnostics and scope limits.
- Observed strata: Redistribution of traffic and exposure across already-observed county-year strata cannot break the certificate, but genuinely new strata require a separate transfer theorem.Hierarchical rollup preserves guarantees by calibrating on the coarsest cell actually used.
- New strata: For a new stratum, weighted split conformal transfers calibration using estimated density-ratio weights under covariate shift and exports coverage with a one-sided slack certificate.The method uses unlabeled target covariates, weighted quantiles, and a diagnostic comparing target covariates with weighted calibration covariates.
- New strata: A large classifier-based total-variation lower bound certifies that the transfer certificate is weak, whereas a small value is necessary but not sufficient for reliable transfer.The diagnostic is sample-split and uses exact binomial confidence bounds combined by Bonferroni correction.
- Limitations: The transfer theorem assumes covariate shift, and unlabeled data cannot provide an assumption-free finite-sample upper bound on total variation or detect concept drift retrospectively.Deployment therefore pairs certification with prospective sequential monitoring rather than correction.
- Drift monitoring: Under the online protocol, the conformal test martingale has false-alarm probability at most α_mon, whereas frozen calibration introduces dependence among deployment p-values.The online guarantee follows from independent uniform p-values and Ville’s inequality; frozen-calibration dependence is described as O(1/|I_cal|).
- Severity-cost risk control: Severity-weighted screening costs are bounded by βκ_max, and interval outputs preserve monotone escalation rules for high-severity sets.The result applies under the declared banded-noise assumptions and includes fatal-omission control under an exact-fatality premise.
3.10. The composition theorem: certification calculus
The composition theorem combines class partitions, stratum handling, reporting-error expansion, deployment weighting, and risk control into an auditable certificate with additive, attributable slacks.
- Composition theorem: The composition theorem gives cell-conditional validity on reported labels, transfers it across declared classes and strata, and supports risk-control bounds.Its proof applies the relevant validity, shift, and risk-control results conditionally within each product cell.
- Certification calculus: CHOIR calibrates within product cells, expands sets for reporting incompatibility, and optionally applies risk control before emitting a per-cell certificate.The prescribed order is condition → weight → calibrate → expand → risk-adjust, with rollup for undersized cells and weighted quantiles for new strata.
- Slack budget: All slack terms are additive and attributable to declared reporting-error, covariate-shift, or estimation assumptions.The reporting-error terms are inputs, whereas the transfer penalty is estimated from unlabeled target covariates.
- Interpretation: The exported deployment certificate is an upper bound on the informativeness floor, not a coverage guarantee, because unlabeled data provide only a lower confidence bound on transfer distance.Realized weighted coverage can fall below the certificate, so the artifact is diagnostic rather than probative.
- Scope: The guarantees are architecture-independent: validity holds for arbitrary, even misspecified, base models and class maps.Latent-class identifiability is optional and does not enter the certification guarantees.
4. Data and experimental design
The study analyzes Texas driver-level crash records while preserving crash-level independence, testing random, temporal, and spatial deployment designs, and documenting structured severity heterogeneity.
- Crash-level splitting: The primary analysis samples one driver per crash, yielding 5,213,921 independent records from 9,078,415 valid-severity driver rows.All splits occur at the crash identifier, preventing related occupants from crossing calibration and test sets.
- Experimental design: The experiments use random, temporal, and spatial split designs, including training on 2017–2021 and testing on 2024–2025.The temporal design reflects deployment where future-period covariates are observable but labels are not.
- Features: The feature configurations use pre-investigation crash, person, and vehicle information while excluding outcome-encoding and post-crash fields.Restraint use is retained despite post-crash recording, while injury counts, narratives, blood alcohol results, and identifiers are excluded.
- Sample composition: Fatal crashes comprise 15,774 records, or 0.30% of the sample, and severe outcomes differ structurally from mild outcomes across speed, rurality, motorcycles, and traffic volume.Across the severity ladder, mean posted speed rises from 44.6 to 55.5 mph, while rural and motorcycle shares increase and median traffic volume decreases.
5. Results
Across seven base models, the unchanged certification layer delivers essentially identical validity while model and partition choices determine efficiency and where coverage is repaired. The experiments also show that true-label coverage can be certified under reporting noise, but the resulting guarantees expose unavoidable width floors and residual limitations.
- 5.1. Validity and model agnosticism: Mean set width ranges from 1.51 to 3.41 categories, while coverage remains indistinguishable across models, separating validity from model-dependent efficiency.The balanced gradient-boosting model is widest because class balancing pushes scores deeper into both tails; this affects usefulness, not validity.
- 5.1. Validity and model agnosticism: 20 of 21 model-by-α cells met the pre-registered acceptance criterion, with empirical coverage spanning 0.8989–0.9000 across all seven models and every set contiguous.Gradient boosting at α=0.10 was the sole exception, covering 0.8989 against a 0.8990 criterion; the paper attributes this tiny miss to a criterion tighter than the marginal theory requires.
- 5.2. Heterogeneity, and where marginal validity fails: Marginal coverage falls to 0.374 for unrestrained drivers, 0.583 for motorcyclists, and 0.624 for rural high-speed records, showing that population averages conceal severe-stratum deficits.The paper interprets this as the expected consequence of calibrating one threshold on a population that is 82% uninjured.
- 5.2. Heterogeneity, and where marginal validity fails: Class-conditional calibration restores cells to 0.8982–0.9002 and changes widths for 18.7% of records, while the latent-class partition gives the narrowest tested sets at 3.236 categories.Validity holds for any partition fixed in advance, but efficiency differs: latent classes outperform declared safety strata, KMeans, and the deep-network gate.
- 5.3. The deficit, and what closes it: On the severity ridge, a coordinate-based declared partition raises the deepest cell from 0.338 to 0.847 and pooled coverage to 0.8979, whereas a generic partition reaches only 0.692 at 75 mph.Residual weakness remains in coarse cells, including an hour-wall minimum of 0.825, so the partition guarantees its declared bands rather than exact conditional coverage everywhere.
6. Discussion
The discussion clarifies how the certification framework should be configured and interpreted, while distinguishing its guarantees from generic conformal tooling and identifying evidence boundaries.
- What the guarantees do not mean: The transfer floor is a best-case evidence boundary, not a lower bound on realized target coverage: eight of 29 certified counties fell below their floors.The one-sided construction permits this outcome rather than treating it as estimator failure.
- What the guarantees do not mean: The drift monitor detects but does not correct concept drift; three alarms therefore require recalibration, which cannot be replaced by retroactive certificate edits.Without labels, concept drift remains uncorrectable, and labels arrive with reporting delay.
- Operational guidance: Validity holds for any partition fixed in advance, making partition choice an efficiency and relevance decision rather than a validity requirement.Operationally meaningful partitions protect the coordinates where severity differences actually occur, whereas generic clustering may not.
- Operational guidance: A declared reporting band is a sensitivity input that should be swept, not an estimated confusion kernel assumed to transfer across agencies or eras.The category-dependent map preserves the K boundary when fatality recording can be treated as exact.
- Operational guidance: A budget β determines escalation workload, which saturates at 0.281 for β above roughly 0.01.Figure 7 frames deployment as choosing a staffing-relevant operating point rather than maximizing an accuracy metric.
- Relation to generic conformal tooling: All three methods target 0.90 coverage and attain it with equal set size, while the proposed layer additionally provides contiguity, unobserved-severity coverage, and shift-related guarantees.The comparison uses a small synthetic package dataset and is a sanity check, whereas the research evidence comes from 5.2 million real records.
Conclusion
The paper presents a model-agnostic certification calculus that combines finite-sample guarantees with sharp limits on informativeness. Experiments show identical validity across seven base models, while real-data results expose subgroup undercoverage and the evidence remains bounded by semi-synthetic noise validation and restricted data.
- Contributions: The framework combines contiguous sets, pre-declared per-class validity, reporting-band transfer, shift certification, and severity-weighted risk control under an attributable additive slack budget.These guarantees are distribution-free in the base model but rely on named structural premises where required.
- Identification limits: The informativeness floor is governed by the true severity law and cannot be lower-bounded distribution-free, but a declared channel identified from record-linkage data makes part of that floor computable.On motorcyclists, the reporting channel imposes a model-independent share of certified width that no base model can remove.
- Empirical conclusion: Seven unchanged base models spanning four decades achieved identical validity, while efficiency differed and six of seven models undercovered unrestrained drivers and motorcyclists despite average population coverage.Conditioning repaired every cell at a reported width cost, and certification added seconds beyond model-fitting times of minutes to hours.
- Limitations: The empirical label-noise check is semi-synthetic because true severity is unobserved, so it validates theorem arithmetic rather than field noise magnitudes.Linking crash records to trauma-registry records is identified as the route to measuring jurisdiction-specific noise.
- Implementation: The framework is released as choircert on PyPI under the MIT license, with documentation, a demonstration dataset, and theorem-level tests.Any estimator exporting a conditional CDF can be wrapped without modification.
A. Deferred proofs
The deferred proofs establish classwise efficiency, sharp reporting-band coverage limits, and one-sided transfer validity by combining order-statistic, coupling, and weighted-conformal arguments.
- Theorem 3.13: Classwise validity is achieved by selecting each class threshold at its within-class quantile, which asymptotically minimizes expected set size subject to validity.The oracle threshold is unique up to null sets under strict monotonicity, and marginal calibration can under-cover classes unless class quantiles coincide.
- Theorem 3.33: The reporting-band coverage loss of δ is sharp: without stronger assumptions, no uniform guarantee can exceed 1 − α − δ.The proof constructs mass outside the expanded ordinal band and shows the bound cannot be improved from support information alone.
- Theorem 3.45: The transfer proof first obtains exact validity under the estimated tilted covariate law, then subtracts total variation distance to reach the target-law guarantee.The comparison works because the two joint laws share the same conditional reported-severity distribution given covariates.
- Theorem 3.45: The transfer theorem’s parameter estimation uses a Lipschitz functional and standard M-estimation rate O_p((m ∧ n′)^−1/2) under dominated derivatives.This rate appears under realizability and local regularity assumptions.
Data availability
The Texas crash records are restricted and cannot be redistributed, but the processing pipeline is scripted and a synthetic dataset supports package examples, tests, and demonstrations.
- Access: The CRIS records require a data request and terms-of-use acceptance, and the authors cannot supply the restricted records directly.Personally identifying fields were excluded at ingest and appear in no released artifact.
- Reproducibility: A scripted pipeline lets researchers with an equivalent CRIS extract rebuild the analysis snapshot and reproduce results with one command.A synthetic dataset with the same schema supports examples, tests, and the Section 6.3 comparison without restricted data.
Code availability
The certification framework is publicly released as the MIT-licensed `choircert` package, with source, documentation, tests, and reproducibility materials available online.
- The package is released under the MIT license and distributed on the Python Package Index as `choircert`, imported as `choir`.
- The release includes source code, documentation, theorem-level tests, the data-construction pipeline, experiment harness, and figure- and table-generation scripts.
- The source repository is accompanied by a citable archived Zenodo release with DOI 10.5281/zenodo.21434172.
Reproducibility statement
The reported results are generated from a frozen, deterministic analysis snapshot with safeguards against stale model artifacts and transcribed values.
- Results were computed on a frozen analysis snapshot whose row order is deterministic.
- Cached model artifacts carry split fingerprints, and mismatched artifacts are rejected rather than silently misaligned with the current snapshot.
- Each reported number traces to a stored result file rather than a manually transcribed value.
CRediT authorship contribution statement
The authors’ contributions cover conceptualization and methodology, with Amir Rafe leading software, formal analysis, drafting, and visualization, and Subasish Das providing supervision and review.
- Amir Rafe contributed conceptualization, methodology, software, formal analysis, original drafting, and visualization.
- Subasish Das contributed conceptualization, methodology, supervision, and review and editing.
- Both authors are credited with conceptualization and methodology.