Source-linked AI summary

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban

arXiv:2608.14767v1cs.CVcs.AIcs.CLcs.RO

TL;DR

Automated-driving explanation datasets rarely capture how drivers naturally explain real-world decisions across timing and operating domains. NARRATE addresses this gap with a multimodal Australian dataset and benchmarks, showing that situational-awareness structure is learnable from driver language while fine-grained context recognition and explanation generation remain challenging.

  • Problem

    Existing driving-explanation datasets often rely on observer-written, post-hoc, simulated, or sensor-generated language rather than drivers’ own explanations across real-time and post-drive settings.

  • Method

    NARRATE pairs 2,050 public-road events from 35 Australian drivers with synchronised multimodal streams, timed explanations, action and context labels, and span-level situational-awareness annotations.

  • Results

    Benchmarks show that situational-awareness structure is learnable from driver language, whereas fine-grained context recognition and naturalistic explanation generation remain challenging.

  • Takeaways & Limitations

    NARRATE provides a domain-aware dataset and evaluation testbed for human-centred explanation models in automated driving.

  • Takeaways & Limitations

    NARRATE is modest in scale, limited to a single daytime Brisbane route, and affected by long-tailed action and context distributions with thin rare classes.

Abstract

from arXiv · show

Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.

1 Introduction

NARRATE addresses gaps in human-centred automated-driving explanations by pairing real-world Australian driving events with drivers’ own explanations, multimodal grounding, and cognitively grounded situational-awareness annotations. It supports participant-disjoint benchmarks spanning situational awareness, scenario context, driver actions, and explanation generation.

  • Motivation: Automated vehicles need explanations that help passengers understand both what the vehicle is doing and why the manoeuvre fits the current road situation.The introduction positions explanation as central to safe and trustworthy human–vehicle interaction on public roads.
  • Situational awareness: Driver-produced explanations can express relevant cues, their meaning for the manoeuvre, and anticipated developments through Perception, Comprehension, and Projection.The Situational Awareness framework analyzes which aspects of situational reasoning appear in explanations without requiring every explanation to contain all three levels.
  • Explanation timing: In-vehicle and post-drive explanations capture complementary reasoning under real-time workload and during later video-cued reflection.The introduction argues that both timings are needed rather than treating driver reasoning as a single static text label.
  • Data gap: Australian public-road driving adds underrepresented domain factors, including left-hand traffic, right-hand-drive vehicles, local road rules, signage, lane-use conventions, and road-user expectations.The paper identifies geographic and operational coverage as a gap in existing large driving datasets.
  • NARRATE contribution: 2,050 annotated events from 35 experienced drivers and driving instructors form NARRATE, a multimodal real-world Australian dataset for learning human-centred explanations.The events were collected on public roads in Brisbane, Queensland, and paired with drivers’ own decision explanations.
  • NARRATE contribution: NARRATE provides multimodal event grounding, six high-level and 32 fine-grained scenario-context categories, span-level SA labels, and participant-disjoint benchmarks across four tasks.The benchmarks cover SA classification, context classification, driver-action classification, and explanation generation.

2 Related Work

Prior driving datasets support perception, behaviour, language, and explanation research, but generally lack natural-language explanations produced by the acting driver. NARRATE addresses this gap with span-level Situational Awareness annotations over experienced-driver explanations, covering perception, comprehension, and projection.

  • Driving datasets: Large-scale datasets such as nuScenes and BDD100K advance multimodal perception and visual driving tasks, while HDD adds driver actions, goals, and causes.These datasets provide perceptual or behavioural supervision rather than natural-language explanations from the driver who performed the manoeuvre.
  • Language and explanations: Driving-language datasets target commands, advice, simulation instruction following, or observer-written post-hoc explanations rather than driver-produced explanations.Examples include Talk2Car, HAD, LaMPilot, BDD-X, and BDD-OIA.
  • Situational Awareness: Situational Awareness supports automated-driving interfaces by expressing what the vehicle perceives, how it interprets scenes, and what risks may follow.It is also linked to driver engagement, takeover readiness, and trust calibration.
  • NARRATE’s contribution: NARRATE adds span-level Endsley L1/L2/L3 annotations grounded in experienced-driver language, enabling models to learn both driver actions and explanation content.The labels cover perception, comprehension, and projection.

3 Data Collection

NARRATE collected synchronised multimodal driving data from 35 analysed participants during standardised naturalistic Brisbane drives. Events were paired with explanations elicited through think-aloud comments, safe post-action prompts, or deferred post-drive discussion.

  • Instrumented vehicle: 8,200 camera sequences across 2,050 events were recorded from four lossless views at 10 fps using a roof-mounted sensor rig and ROS 2 edge-compute unit.All streams were time-synchronised via GPS pulse-per-second through the inertial navigation system and stored as ROS 2 bag files.
  • Participants: 35 participants remained after two sessions were excluded: 30 experienced non-instructor drivers and 5 licensed driving instructors.One excluded session lacked recoverable event timestamps, while the other involved non-protocol-compliant driving.
  • Driving protocol: Each session involved approximately 40 minutes of naturalistic public-road driving in Brisbane along a predefined route, preserving variation in traffic, road users, and signal timing.Participants were briefed to drive normally, obey road rules, and prioritise safe vehicle control.
  • Australian road context: A rear-seat data engineer monitored recording and tagged candidate events, while a front-seat researcher elicited explanations only after completed actions and when safe.Three people occupied the vehicle: the participant driver, front-seat researcher, and rear-seat data engineer.
  • Explanation elicitation: Explanations came from participant think-aloud comments, researcher-triggered post-action prompts, or engineer-flagged deferrals when discussing during driving was unsafe or impractical.NARRATE therefore captured both immediate and reflective driver explanations.

4 Annotation Design

NARRATE uses complementary span-level and event-level annotations to capture both drivers’ expressed reasoning and the surrounding driving situation. The design applies a three-level Situational Awareness framework, a 32-category scenario taxonomy grouped into six contexts, and multi-annotator reliability assessment.

  • Annotation layers: NARRATE annotates driver explanations with span-level Situational Awareness labels and events with scenario-context labels, capturing language-based reasoning and driving situations.The two layers are complementary and operate at different annotation granularities.
  • Situational Awareness annotation: Situational Awareness labels follow Endsley’s three levels: L1 Perception, L2 Comprehension, and L3 Projection.Annotators highlighted supporting explanation spans, labelled all expressed levels, and annotated in-vehicle and post-drive explanations independently; one explanation could contain multiple levels.
  • Scenario-context annotation: Scenario context is annotated at the event level using 32 fine-grained categories grouped into six high-level contexts.The contexts are Traffic Compliance; Social Interaction and Traffic Flow; Navigation and Routing; Hazard and Obstacle Management; Special Zones and Stops; and Environmental and Adaptation.
  • Annotators and label resolution: Three experienced annotators with complementary expertise contributed to the annotation process, with Annotator A labelling the full dataset for Situational Awareness and context.A stratified 303-instance subset covering approximately 15% of annotated events and all 35 participants was selected to estimate reliability.

5 Dataset Statistics

NARRATE contains 2,050 annotated events from 35 participants, with long-tailed action and context distributions and explanations collected through multiple trigger sources. Its reliability subset supports robust context annotation and moderate-to-substantial agreement for L1 Perception, while L3 Projection agreement is lower.

  • Scale, splits, and sensor coverage: 2,050 annotated events were retained from 2,138 raw event tags after excluding erroneous, irrelevant, unrecallable, unlocatable, and explanation-less events.The dataset comprises 35 participants.
  • Action, context, and SA distributions: 976 events (47.6%) involved slowing down, followed by 451 (22.0%) lane changes and 187 (9.1%) speed-ups.The four non-action classes collectively represent 9.0% of events.
  • Action, context, and SA distributions: 430 events (21.0%) belonged to the most frequent fine-grained context category, vehicle following, reflecting a similarly long-tailed context distribution.The passage identifies context labels as following a long-tailed pattern.
  • Explanation trigger source: 861 engineer-flagged deferral events produced in-vehicle explanations for 13.2% of rows but post-drive explanations for 95.9%.Participant think-aloud and researcher-triggered prompts yielded near-complete in-vehicle narration rates of 97.4% and 97.7%, respectively.
  • Inter-annotator agreement: 303 reliability-subset instances were used for agreement analysis, with maximum deviation of 3.4 pp from the full-dataset distribution.Context labels were robust across both granularities, while L1 Perception showed moderate-to-substantial three-way agreement and L3 Projection was lower.

6 Benchmarks

NARRATE defines four participant-disjoint benchmarks spanning situational awareness, scenario context, driver actions, and explanation generation. Together, the tasks test whether driver language and structured event information support human-centred understanding and generation, while establishing baseline comparisons rather than proposing a new model.

  • Benchmark overview: Four participant-disjoint tasks evaluate SA classification, scenario-context classification, driver-action classification, and explanation generation.The benchmarks are designed to characterise what NARRATE enables and where it remains challenging, supporting dataset use and comparison.
  • T1: Situational Awareness: T1 tests recovery of three independent SA levels—L1 Perception, L2 Comprehension, and L3 Projection—from in-vehicle and post-drive driver explanations.The task uses binary classifications with BCEWithLogitsLoss and threshold 0.5; macro-F1 across the three levels is primary.
  • T2: Scenario context: T2 infers scenario context from explanations at six high-level and 32 fine-grained multi-label category granularities.Fine-grained macro-F1 is computed over labels with at least five test instances, using primary_text as the in-vehicle explanation when available and otherwise the post-drive explanation.
  • T3: Driver action: T3 infers one of 10 driver-action classes from text, video, motion, or their combinations, primarily using macro-F1 and secondarily accuracy.Compared baselines include text models, frozen CLIP visual features, kinematic features, and simple fusion models combining kinematics with visual or text embeddings.
  • T4: Explanation generation: T4 benchmarks natural-language explanation generation from action labels, context categories, and kinematic traces rather than end-to-end multimodal inputs.Retrieval comparisons use random selection, label matching, frozen visual nearest neighbours, and visual+kinematic nearest neighbours.

7 Limitations

NARRATE is limited by its modest, single-route daytime scope in Brisbane, long-tailed class distributions, and simplified baselines that leave richer multimodal temporal modeling for future work.

  • NARRATE is modest in scale and restricted to a single daytime route in Brisbane.
  • Long-tailed action and context distributions leave rare classes thin.
  • The baselines deliberately use simplified event-level representations as a lower bound, while full temporal, LiDAR, and multi-view exploitation remains future work.

8 Conclusion

NARRATE is a multimodal, real-world Australian driving dataset designed for human-centred explanations in automated driving, combining driver-produced explanations with synchronised multimodal road data and structured annotations.

  • NARRATE contains 2,050 annotated public-road events from 35 experienced drivers and driving instructors.Events pair synchronised visual, LiDAR, localisation, inertial, and kinematic streams with in-vehicle and post-drive explanations produced by the drivers themselves.
  • The dataset provides driver-action labels, scenario-context labels, and span-level Situational Awareness annotations.
Loading 2608.14767v1…