Source-linked AI summary
EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings
Md Thamed Bin Zaman Chowdhury, Moazzem Hossain
TL;DR
In low-resource settings, unreliable crash records and non-scalable expert audits constrain proactive road safety assessment. The paper introduces Expert-Grounded Distillation, an audit-calibrated teacher-to-student pipeline with an open dataset and compact auditor, and reports that grounded supervision improves ordinal risk assessment and can outperform raw model scale.
Problem
Unreliable crash records and non-scalable formal audits constrain proactive road safety management in low-resource settings.
Method
Expert-Grounded Distillation calibrates a teacher against authoritative field audits, gates supervision generation on expert agreement, and distills structured audits into a compact open vision-language model alongside BD-ARSA.
Results
Grounded fine-tuning improves ordinal risk agreement, supporting the conclusion that expert-grounded supervision can outperform raw model scale.
Takeaways & Limitations
The paper supports expert-grounded supervision as a scalable approach to visual road safety auditing in resource-constrained environments.
Takeaways & Limitations
EG-ARSA audits a single street-view image, relies partly on teacher-generated nationwide labels, and concentrates expert ground truth in 18 LGED-audited districts.
Abstract
from arXiv · showhide
Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To address this problem, we propose Expert-Grounded Distillation (EGD), a novel artificial intelligence framework that transfers institutional road safety expertise into a compact vision-language model for scalable visual road safety auditing. The key innovation is a quantified expert-grounding stage in which the teacher vision-language model is calibrated against authoritative field audits. Large-scale annotation is permitted only after the teacher reaches substantial agreement with expert risk assessments (Cohen's kappa = 0.74). The calibrated teacher then generates structured supervision that is distilled into an 8-billion-parameter student vision-language model using Low-Rank Adaptation and a single leakage-free prompt. We also introduce Bangladesh Road Safety Audit (BD-ARSA), the first open, expert-grounded Bangladeshi visual road safety audit dataset containing 21,947 image-audit records with near-national coverage, and Expert-Grounded Road Safety Auditor (EG-ARSA), the first vision-language model developed specifically for this task. Experimental results show that grounded fine-tuning substantially improves ordinal risk assessment over the zero-shot baseline, while blind expert evaluation demonstrates that the compact student outperforms both its 31 billion-parameter teacher and Gemini-2.5-Flash. These findings demonstrate that EGD provides an effective and scalable engineering solution for proactive road safety auditing in resource-constrained environments.
1. Introduction
Bangladesh and other low- and middle-income countries face severe road-safety burdens, but unreliable crash records and the limited scalability of formal audits constrain proactive assessment. The paper addresses this gap with expert-grounded distillation, an open dataset, and a compact visual auditor.
- Motivation: 90% of global road-crash fatalities occur in low- and middle-income countries despite these countries possessing about 60% of the world’s vehicles.Bangladesh experiences around 4,000 deaths and 200,000 injuries annually, with crashes costing up to 5.1% of national GDP.
- Motivation: Bangladesh’s crash records are unreliable, including a greater than 90% discrepancy between government fatalities and WHO estimates and little usable data since 2016.This limits crash-based black-spot identification and treatment.
- Motivation: Formal road safety audits assess infrastructure hazards independently of crash history, but multidisciplinary field audits cannot scale to the full road network.Only about 1.5 million km have been star rated against a UN target near 12 million km by 2030.
- Research gap: Existing automated approaches are either data-hungry and non-interpretable supervised coders or proprietary, zero-shot VLMs that are weak on fine-grained detail and not adapted to Bangladesh’s LGED methodology.Prior approaches also lack open, reusable datasets for this task.
- Approach: EGD calibrates a teacher against authoritative ARI–BUET field audits, gates large-scale supervision on expert agreement, and distills the grounded knowledge into an 8B open student with a leakage-free prompt.The teacher calibration agreement is κ=0.74 before nationwide label generation.
- Artifacts and results: BD-ARSA contains approximately 22,000 open image–audit records, while EG-ARSA produces structured single-image audits for rural and suburban LGED-class roads.The dataset covers 155 road corridors across 63 districts and all eight administrative divisions; national highways and dense city networks remain future work.
- Artifacts and results: Grounded fine-tuning raises ordinal-risk QWK by +0.40 over zero-shot to approximately 0.48, and blind experts rate the 8B student correct on 81% of cases versus 58% for Gemini-2.5-Flash and 42% for the 31B teacher.Automated metrics independently yield 0.74, 0.59, and 0.36 for the same models.
2. Related work
Prior road-safety automation moved from supervised attribute coding toward zero-shot VLM-based visual question answering, while distillation research showed that curated teacher supervision can benefit smaller models. EG-ARSA combines these strands with measured expert grounding and bias-controlled evaluation.
- Road-safety automation: Early road-safety automation used supervised CNN and recurrent architectures to classify many iRAP attributes, but later VLM work reframed auditing as visual question answering.The cited VLM studies used proprietary, zero-shot models on road imagery.
- Road-safety automation: Zero-shot VLMs generalise better than CNNs to unseen classes but underperform on spatial and metric attributes, while specialist models remain stronger on fine-grained sign and object recognition.Frontier VLMs are described as weak on counting and geometry but stronger on qualitative description.
- Knowledge distillation: Prior distillation studies show that teacher rationales, machine-generated multimodal supervision, and active data curation can allow smaller models to match or exceed larger generators.The literature emphasizes careful curation and human verification of teacher-generated data.
- Knowledge distillation: EG-ARSA distinguishes its approach by anchoring synthetic supervision in formal institutional audits rather than generic teacher outputs.Its teacher prompt is validated against expert field audits at κ=0.74 before scale-up.
- Knowledge distillation: EGD uses context distillation: an expert-calibrated teacher prompt supplies context that the 8B student internalises and reproduces through a single leakage-free prompt.The paper frames the student’s performance as a real-domain example of weak-to-strong generalization.
- Efficient adaptation: LoRA reduces adaptation costs by training less than 1% of network parameters through low-rank weight updates.QLoRA further adds 4-bit quantization for more efficient adaptation.
- Evaluation: Because teacher-generated labels can bias same-family judges, the paper anchors headline claims in blind human evaluation and runs the teacher leakage-free.It also reports QWK for ordinal risk and bootstrap confidence intervals.
- Positioning: EG-ARSA occupies an intersection not covered by prior work: an open, fine-tuned VLM distilled from institutional audits for rural/suburban LMIC roads with explicit judge-bias controls.The evaluation also includes bootstrap intervals, a base-model ablation, and blind human review.
3. Data curation: the BD-ARSA dataset
BD-ARSA combines expert-audited records, human-verified alignments, and teacher-audited nationwide imagery into a three-tier dataset. Its taxonomy, provenance hierarchy, location-disjoint splits, and empirical label structure define the dataset’s scope and constraints.
- Expert audit foundation: BD-ARSA derives its ground truth from formal ARI–BUET audits commissioned by LGED across roughly 1,433 km in 18 districts.ARI was responsible for about 1,200 km of the audited network.
- Hazard taxonomy: The dataset uses 12 LGED hazard categories, which were consolidated from 689 auditor descriptions and validated on 100 human-reviewed findings.Two independent zero-shot classifiers proposed the mapping, while disagreements were manually adjudicated.
- Provenance tiers: Gold records retain individually captioned hazard crops, while silver records pair road-scene photos with manually verified expert findings and structured audits.Gold has the highest image-to-hazard certainty; silver is reliable but lower-certainty because source reports lacked explicit photo-to-hazard correspondence.
- Provenance tiers: The street-view tier contains nationwide imagery audited by a 31B teacher using a prompt calibrated against expert audits before label generation.These records have no per-location expert findings, so their image-to-hazard correspondence rests on the validated teacher prompt.
- Coverage and splits: 21,947 records span 155 road corridors, 63 districts, and all 8 administrative divisions, with location-disjoint train, validation, and test splits.The split sizes are 16,082 training, 2,418 validation, and 3,447 test records.
- Dataset analysis: The dataset preserves a long-tailed, correlated hazard structure rather than artificially balancing categories, while drainage and skid_resistance remain constrained by non-visual field checks.Category prevalence is stable across provenance tiers, and single-image assessment is generally reliable only for highly obvious cases in these categories.
4. Methodology
Expert-Grounded Distillation calibrates teacher supervision against expert audits before scaling, then distills that grounded knowledge into a compact vision-language auditor trained for structured, leakage-free road-safety outputs.
- Expert-grounded distillation: κ=0.74 agreement with expert risk judgement gates nationwide teacher-label generation in the measured grounding stage.Calibration uses institutional field audits before large-scale supervision is generated.
- Expert-grounded distillation: The student uses grounded gold, silver, and street-view teacher audits to fine-tune Qwen3-VL-8B-Instruct with LoRA for single-image inference.The vision encoder is frozen while the language backbone is adapted.
- Task formulation and structured output: Each record supervises hazard generation, ordinal overall-risk classification, and recommendation generation through one leakage-free JSON-output prompt.The prompt references neither audit reports nor findings unavailable at inference.
- Training objective and imbalance handling: Task training minimizes weighted, per-task normalized cross-entropy, using equal hazard and risk weights and a lower recommendation weight to prevent longer targets from dominating.The task weights are whazard = 1.0, wrisk = 1.0, and wrec = 0.5.
- Training objective and imbalance handling: Train-only logit adjustment addresses imbalance in the three-class risk head, while hazard generation remains an un rebalanced variable-length multi-label sequence task.The risk adjustment uses temperature τ=1 and class counts (nLow, nMed, nHigh) = (225, 6,340, 9,517).
- Inference and operating-point selection: Validation-optimal post-hoc offsets trade a small amount of ordinal QWK for substantially higher Low recall and are applied once at test time.The offset is selected on validation logits and then frozen for inference.
- Evaluation protocol and metrics: QWK evaluates ordinal risk agreement by penalizing distant errors more heavily, making Low↔High confusion four times as costly as adjacent Medium↔High confusion.At three ordered risk levels, the stated weights are 1 versus 1/4.
5. Results
Grounded fine-tuning substantially improves ordinal risk prediction over the zero-shot base, while the compact student performs strongly against larger models and blind expert evaluation. Results also characterize training behavior, operating-point trade-offs, internal consistency, and an important subset limitation.
- 100% of 3,447 test generations contained a risk token, with parsed text-versus-logit agreement of 0.98.
- All task losses decreased across two epochs while validation QWK rose from 0.385 to 0.524, selecting the early-stop checkpoint.Risk-token accuracy increased from 0.56 to about 0.90 during training.
- +0.40 QWK and +0.17 accuracy separate the fine-tuned student from the zero-shot base, reaching QWK 0.4815 and exact-risk accuracy 0.7171.The zero-shot base achieved QWK 0.0772 and accuracy 0.5440; bootstrap confidence intervals did not overlap.
- 73.7% risk accuracy on 262 expert-grounded entries put the 8B student ahead of Gemini-2.5-Flash at 59.2% and the leakage-free 31B teacher at 35.9%.On silver hazard-category detection, the student traded lower recall for higher precision than Gemini.
- 81% blind-human risk correctness ranked the student above Gemini at 0.58 and the teacher at 0.42, with overall quality of 3.94/5.The independent automated and blind-human evaluations reproduced the same model ranking.
- The raw operating point has Low recall 0.02, while the δ-adjusted variant raises Low recall to 0.42 and F1 to 0.30 at a small ordinal-QWK cost.Residual confusion-matrix errors are predominantly between adjacent risk levels.
6. Discussion
The discussion argues that expert grounding, rather than model scale alone, enables a compact model to produce actionable, affordable audits. EG-ARSA remains constrained by the information available in a single street-view image and by teacher-generated nationwide labels.
- Why Expert-Grounded Distillation works: Expert grounding, rather than scale, is the core mechanism behind the student’s deployment advantage over the unaided teacher.The teacher’s agreement falls from κ=0.74 when grounded to 0.36 under leakage-free evaluation, while the student internalizes the grounded knowledge.
- Practical implications: EG-ARSA outputs structured hazards, severity, ordinal risk, and recommendations, making its audit language actionable for non-expert users.The system is designed to provide more than a single score and can be deployed through a web application.
- Practical implications: A single frozen post-hoc offset allows deployment to select screening recall or ordinal QWK without retraining.This supports different operating points while preserving the same trained model.
- Practical implications: The 8B LoRA student runs on modest hardware and can make audit coverage orders of magnitude cheaper than formal field road-safety audits.The model was fine-tuned on a single A100 in about six hours and is intended for low-resource deployment.
- Generalizability: The EGD recipe is jurisdiction-agnostic and can transfer institutional-audit grounding to settings with small expert-audited corpora.The near-national footprint spans 63 districts and 8 divisions, while future extensions include other LMICs with iRAP or RSA programmes.
- Limitations and future work: Single-image auditing cannot fully recover road geometry or non-visual extremes, and nationwide street-view labels inherit some of the teacher’s ceiling.Expert-grounded gold and silver tiers therefore anchor evaluation, while expert ground truth remains concentrated in 18 LGED-audited districts.
- Limitations and future work: Overhead or satellite imagery could complement street views by supplying geometry and non-visual attributes.Learned fusion of the two modalities is identified as a promising extension.
7. Conclusion
The paper presents EGD and the open BD-ARSA dataset to address scalable road-safety auditing where crash records and expert-led audits are limited. Grounded fine-tuning improves ordinal agreement, and blind expert evaluation favors the compact model over larger alternatives.
- Conclusion: +0.40 quadratic weighted kappa is the student’s ordinal risk improvement over its zero-shot baseline.Blind expert evaluation also shows the compact 8B model outperforming its 31B teacher and a frontier proprietary model.
- Conclusion: EGD grounds teacher supervision in institutional audits, validates it through human review, and distills it into the compact open EG-ARSA model.The pipeline uses a single leakage-free prompt.
- Conclusion: BD-ARSA is released as an open, expert-grounded Bangladeshi road-safety visual-audit dataset.Future work extends EGD to additional road classes and complementary sensing modalities.
CRediT authorship contribution statement
The contribution statement assigns distinct roles across conceptualization, methodology, implementation, validation, analysis, data curation, writing, visualization, supervision, and project administration.
- CRediT authorship contribution statement: Md Thamed Bin Zaman Chowdhury contributed to conceptualization, methodology, software, validation, analysis, investigation, data curation, writing, and visualization.
- CRediT authorship contribution statement: Moazzem Hossain contributed to methodology, validation, writing review and editing, supervision, and project administration.
Data and code availability
The authors release the BD-ARSA annotations, EG-ARSA LoRA adapter, and reproduction code, while providing reconstruction instructions for the non-redistributed street-view imagery. Test results use a frozen location-disjoint split with specified expert-grounded evaluation subsets.
- Released artifacts: The training and evaluation code is released under Apache-2.0 with reproduction steps.
- Released artifacts: The dataset annotations and metadata are released under CC BY 4.0, while street-view imagery is reconstructed locally through the included fetch script.The imagery itself is not redistributed because it is accessed under Google Maps Platform terms.
- Released artifacts: The EG-ARSA model is released as an Apache-2.0 LoRA adapter for Qwen3-VL-8B-Instruct.Use is additionally subject to the Gemma Terms of Use because the distillation supervision was Gemma-generated.
- Evaluation data: n=3,447 is the frozen location-disjoint test split used for all test-set results.Multi-model and human evaluations use n=262 expert-grounded records, comprising 158 expert-gold and 104 expert-silver records.
- Evaluation data: The silver hazard-category overlap is computed over 557 expert findings from 104 silver test records.
Appendix A. NLP pipeline for taxonomy derivation and corpus validation
The appendix describes an NLP pipeline that converts expert audit text into the 12-category BD-ARSA taxonomy and validates the resulting corpus analyses.
- The pipeline operates on the expert finding corpus extracted from ARI–LGED audit reports.It documents taxonomy derivation and validation analyses.
- The validation uses both taxonomy-focused and corpus-level analyses.The table records agreement between the two zero-shot label generators, while the appendix describes broader corpus validation.
- Table A.10 reports inter-classifier agreement between Gemini-2.5-Flash and BART-large-MNLI.
Appendix A.1. Deriving the 12-category schema with dual zero-shot classifiers
The 12-category schema was derived by mapping expert free-text findings with two independent zero-shot classifiers and routing disagreements to human adjudication.
- 689 unique finding texts, covering 208 distinct short labels, were mapped to one of 12 LGED categories.
- Two independent zero-shot classifiers, Gemini-2.5-Flash and BART-large-MNLI, performed the category mapping.
- 82 label-level disagreements were routed to human adjudication after the classifiers showed only moderate agreement.The agreement was not treated as a reliability estimate for the taxonomy.
Appendix A.2. Corpus lexical analysis
Lexical profiling links the detailed expert findings to the hazard taxonomy and uses an independently written auditor-summary layer as a cross-check.
- The dominant detailed-findings vocabulary directly tracks the hazard categories.Frequent terms include route (685), trees (459), traffic (414), intersection (382), and pedestrian (334).
- The corpus analysis compares the detailed-findings layer with an independent text source to cross-check the taxonomy.
Appendix B. Prompts
Appendix B documents the asymmetric Expert-Grounded Distillation prompts: an elaborate LGED-grounded teacher prompt generates structured street-view audits, while standardized student prompts support leakage-free training and inference.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): The teacher prompt uses multiple Street View screenshots and one satellite tile, then applies taxonomy validation, risk-boundary post-processing, and human review.Human review covered a 384-record sample.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): The teacher is instructed to audit Bangladesh’s mixed-traffic road context using the LGED issue categories.The categories include speed management, markings, sight obstruction, drainage, signs, skid resistance, roadside severity, embankments, shoulders, pedestrian facilities, bus stoppages, and intersection quality.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): The teacher checks hazards across 12 categories, including vision obstruction, drainage, traffic signs, skid resistance, roadside severity, and embankment safety.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): Shoulder, pedestrian, bus-stoppage, and intersection rules require context-specific visual evidence and distinguish overlapping hazards.The prompt emphasizes satellite inspection for shoulders, visible stopping activity for bus stoppages, and structural defects for intersection quality.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): The prompt requires examining all images, reporting both hazards and absent safety features, and relying primarily on satellite imagery when Street View is unavailable.
- Appendix B.1. Teacher street-view generation prompt (gemma-4-31b-it): The teacher output is a raw JSON object containing location, road, hazard, severity, overall-risk, imagery-availability, and recommendation fields.
- Appendix B.2. Student single-hazard prompt (expert-gold crops): The student single-hazard prompt identifies one visible hazard and overall risk from a close-up image, optionally adding a remedial recommendation.
- Appendix B.3. Student full-audit, leakage-free prompt (expert-silver and street-view): The full-audit student instruction is standardized and leakage-free, with no grounding context, corrective rules, or report references.