Source-linked AI summary

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

Yuhang Wang, Lingyao Li, Hao Zhou

arXiv:2607.23822v1cs.LG

TL;DR

Naturalistic driving-style research lacks evidence that separates driver-specific behavior from vehicle and context confounds. DriveDNA introduces a large-scale dataset and benchmark with matched evaluations, finding that learned representations identify unseen drivers strongly while apparent style often reflects vehicle, place, or driving conditions.

  • Problem

    Naturalistic driving-style benchmarks lack large-scale evidence that separates driver-specific behavior from vehicle, route, and traffic-context confounds.

  • Method

    DriveDNA constructs a large-scale naturalistic dataset and benchmark that operationalizes style as stable driver-specific motion under comparable contexts and evaluates identification, prediction, and matched comparison.

  • Results

    Learned representations outperform classical descriptors for unseen-driver re-identification, while condition-matched analyses retain driver-specific signal and expose substantial vehicle- and context-related variation.

  • Takeaways & Limitations

    Reliable driving-style claims require measuring leakage alongside utility and matched alongside unmatched performance because driver identity, driving style, and personalized prediction are not interchangeable.

  • Takeaways & Limitations

    The fleet is regionally and demographically selective, and driver–vehicle entanglement limits how completely vehicle and driver effects can be separated.

Abstract

from arXiv · show

Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.

1 Introduction

Driving style reflects systematic, driver-specific control patterns, but naturalistic diversity makes it difficult to separate driver identity from vehicle and situational effects. DriveDNA addresses this challenge with an operationalized behavioral definition and a large-scale, heterogeneous naturalistic benchmark.

  • Motivation: Under comparable conditions, drivers differ systematically in headway, speed choice, acceleration, lane-changing, and other control patterns that can support identification and personalization.These variations motivate personalized driver assistance and autonomous planning.
  • Challenge: Naturalistic driving data captures the diversity of drivers, vehicles, roads, and traffic conditions required for personalization, but this heterogeneity creates a key methodological challenge.Existing controlled or simulated benchmarks provide cleaner comparisons at limited scale and heterogeneity.
  • Operational definition: Driving style is operationalized across driver input, realized vehicle motion, and latent style, with realized motion providing the best-covered and most comparable behavioral space.Driver inputs are inconsistently observable and not directly comparable across vehicles, whereas latent style is inferred rather than directly observed.
  • Dataset contribution: DriveDNA contains 4,121 drives from 465 drivers across 115 vehicle models, totaling 975 hours of human-controlled driving at 10 Hz with forward video and synchronized car-control signals.The dataset is large, open, community-driven, and naturalistic.
  • Dataset contribution: 420 drivers share a vehicle model with at least one other driver, while 20 drivers appear on two or more models, anchoring within-nameplate and cross-vehicle comparisons.DriveDNA is constructed from worldwide community drivers after removing automation-engaged segments and provides a driver/vehicle/drive hierarchy.

2 Related Work

Prior driving datasets largely target scene understanding, trajectories, or controlled human behavior, while driver-identification studies show identity information in telemetry but often leave vehicle, route, and context confounds correlated with identity. Existing personalized-driving benchmarks cover complementary settings, motivating DriveDNA’s cross-drive, cross-behavior evaluation and explicit shortcut-leakage reporting.

  • Real-world driving data: Public autonomous-driving datasets primarily support scene understanding and trajectory research, whereas human-behavior resources include controlled-access and naturalistic driving corpora.Examples include BDD100K, nuScenes, Waymo Open Dataset, highD, SHRP2, and HDD.
  • Driver identification and driving-style modeling: Telemetry studies establish that short vehicle-signal sequences contain driver-identity information, but commonly emphasize closed-set recognition in fixed or weakly controlled configurations.Vehicle dynamics, repeated routes, and context exposure can remain correlated with identity in these settings.
  • Driver identification and driving-style modeling: DriveDNA operationalizes prior insights by using short-term behaviors for annotation while evaluating stable style through cross-drive re-identification, personalized future-motion prediction, and matched-context comparison.The benchmark reports shortcut leakage explicitly and is designed across behaviors and vehicles.
  • Personalized driving benchmarks and planning methods: Existing personalized-driving benchmarks span same-vehicle data, style-conditioned planning, multimodal behavior explanation, and simulated closed-loop evaluation.PDB, StyleDrive, PDB-Eval, and Person2Drive address complementary aspects of personalized driving.
  • Representation learning under confounding: Representation-learning research supplies time-series, multimodal, few-shot, self-supervised, foundation-model, fusion, and domain-adversarial components for evaluating driving-style representations under confounding.The related methods include Transformers, contrastive and margin objectives, prototypical evaluation, frozen image and video encoders, conditional fusion, and domain-adversarial learning.

3 DriveDNA Dataset Development

DriveDNA is a multimodal naturalistic driving dataset built from synchronized windshield-mounted recordings of vehicle signals and forward video during everyday human-controlled driving. Its development pipeline standardizes decoding, extracts benchmark-ready segments and windows, organizes heterogeneous signals, and adds audited maneuver annotations with tiered data release.

  • Data collection: DriveDNA logs synchronize forward-view video with CAN-derived kinematics, driver inputs, and available radar, lane, and pose estimates across 465 drivers and 115 vehicle models.Data were voluntarily contributed from participants’ everyday driving using windshield-mounted comma devices running openpilot logging.
  • Decoding and signal alignment: 0.50 AUROC is the format-prediction performance after decoding every drive from the same log format, reducing exploitable format differences to chance.Lateral motion is represented with realized path curvature rather than steering-wheel angle because steering ratio and wheelbase vary across vehicles.
  • Human-driving extraction: 975 hours of human-controlled driving remain after removing automation-active intervals and discarding segments shorter than 30 seconds.The extraction produces 12,440 segments from 3,989 drives and 452 drivers, including 581 hours above 2 m/s; all benchmark tasks use these segments.
  • Signal organization: 95% coverage is reported for steering signals, compared with 40% for gas and 57% for brake, motivating three variable tiers organized by benchmark role and availability.Tier A contains widely available realized vehicle-motion signals, while Tier B contains vehicle-dependent driver inputs.
  • Windowing and context labels: 62,674 windows from 428 drivers are formed by applying 60-second windows with a 30-second stride, with each window assigned one of six driving contexts.The contexts are car-following, free driving, curve, stop-and-go, high-speed, and urban driving, labeled from speed, curvature, lead-vehicle information, and stopping patterns.
  • Maneuver annotations: 276,248 maneuver-event annotations span six classes and are generated with rule-based methods using vehicle motion, steering, lane, and lead-vehicle signals.Forward-video audits assessed the first five classes using approximately 1,000 sampled events per class, while only 22,322 verified lane changes are released; events are not core benchmark training targets or ground-truth labels.

4 Benchmark Design

DriveDNA evaluates driving style through three core tasks targeting driver identity, predictive utility, and robustness under matched conditions. Its benchmark fixes task protocols, splits, metrics, leakage probes, and evaluation auditing to separate driver-specific behavior from contextual shortcuts.

  • Core tasks: Three core tasks test driver re-identification, personalized future-motion prediction, and same-driver recognition under matched driving conditions.These correspond to identity, utility, and robustness, respectively, with fixed metrics, splits, and procedures.
  • Personalized behavior prediction: From a 5 s history at 10 Hz, models predict acceleration and path curvature over 1, 3, and 5 s horizons, with 3 s primary.Inputs may include front-video context, vehicle parameters, and a few-shot driver support set; personalized models are compared with matched non-personalized counterparts.
  • Condition-matched comparison: 14,868 balanced window pairs are matched on vehicle model, scenario type, speed range, and applicable headway range for controlled driver comparisons.Matching quality is independently validated using vision–language scene attributes.
  • Splits and generalization: The driver-disjoint split uses 212 training, 45 validation, and 45 test drivers, plus a separate 53-driver hold-out for few-shot evaluation.Additional manifests test same-model generalization, cross-vehicle transfer, fixed conditions, and missing-channel robustness.
  • Evaluation protocol: Metrics cover recognition, prediction error, behavior-distribution distances, personalization gain, and leakage of vehicle, drive, and condition information.Video is contextual input only and is evaluated with route- and context-leakage probes because scene information can shortcut re-identification; headline results use three training seeds with fixed evaluation splits.

5 Baseline Settings

The baseline suite uses frozen data splits and thirty configurations across five modeling dimensions aligned with driver identity, predictive utility, and robustness under condition matching. It spans representation, shortcut control, personalization, multimodal context, distributional prediction, and interpretable classical anchors with diagnostic leakage probes.

  • Baseline organization: Thirty baseline configurations are organized across representation learning, shortcut robustness, personalization, multimodal modeling, and distributional prediction on frozen splits.These dimensions provide controlled reference points for driver identity, predictive utility, and robustness under condition matching.
  • Representation: Representation baselines train patch-based Transformer encoders with supervised-contrastive or ArcFace objectives, prototype-based few-shot enrollment, self-supervised learning, and zero-shot MOMENT evaluation on unseen drivers.The encoders include joint-channel and channel-independent PatchTST and iTransformer; self-supervised rows use masked reconstruction or JEPA-style latent prediction.
  • Shortcut control: Shortcut-control baselines isolate driver-specific information by conditioning a population model on state, context, and vehicle, then modeling residual style and adversarial invariance.A utility–leakage Pareto curve completes the shortcut-control stage.
  • Personalization: Personalization baselines form a conditioning ladder from a generic predictor to few-shot support encoding through FiLM, with query injection testing conditioning architecture.The query-injection variant follows the style of TransFuser.
  • Multimodal context: The multimodal predictor fuses CAN and frozen video tokens through cross-attention across six conditioning settings: CAN, +vehicle, +video, +driver, +video+driver, and +residual.Video features compare per-frame DINOv2, DINOv3, and SigLIP2 with temporal V-JEPA 2 clips, alongside CAN↔video contrastive and Qwen3-VL baselines.
  • Distributional prediction: Distributional baselines use mixture-density and conditional-VAE heads to model behavioral uncertainty and multiple future modes, evaluated with likelihood and distribution metrics.Interpretable CAN, radar, and lane descriptors with gradient-boosted classifiers provide a classical lower-bound anchor and interpretation probe, while diagnostic leakage probes run throughout.

6 Results

Results show that driver-specific information survives vehicle and condition controls, but re-identification alone does not establish behavioral or predictive validity. Learned representations outperform descriptors under unseen-driver and matched-context evaluation, while video recognition is dominated by route and vehicle context.

  • Variation and controls: 29–60% of naive behavior-statistic variance is explained by scenario, speed regime, and vehicle, yet residualized drivers remain identifiable at 4× chance.Vehicle-model probes reach 2.3× chance from steering-wheel angle, showing that signal choice also affects apparent driver information.
  • Unseen-driver re-identification: AUROC .935 ± .005 from a supervised contrastive Transformer exceeds classical descriptors at AUROC .707 for unseen-driver re-identification.Top-1 identification is about 33× chance; PatchTST matches the joint encoder (.932 vs. .935), while ArcFace and iTransformer trail.
  • Condition-matched comparison: AUROC .811±.006 is retained on 14,868 matched-context pairs, whereas descriptor verification falls to AUROC .550—chance level—across six scenario types.Trained CAN encoder rankings remain stable under matching, with SupCon and ArcFace highest and iTransformer lowest.
  • Video leakage: AUROC .937 for video-only re-identification is accompanied by route prediction at 347× chance and vehicle prediction at 64×, indicating severe contextual leakage.On within-nameplate splits, CAN drops to .887 while video rises to .962, consistent with place recognition rather than driving behavior.
  • Representation utility: 2.6× chance vehicle leakage remains in the learned realized-motion embedding, while per-frame video features reduce event-forecasting AUROC and add only +0.09% to personalized prediction.DANN cannot reduce the residual leakage without utility cost, and explicit vehicle embeddings reduce unseen-driver prediction accuracy by −0.9 to −1.4%.

7 Release, Privacy, and Ethics

DriveDNA uses a tiered release to balance reproducibility with privacy: public de-identified signals, embeddings, manifests, attributes, harnesses, and code, plus gated blurred raw video. Governance includes informed consent, documented residual risks, takedown support, leakage probes, and restrictions against re-identification.

  • Tiered release: The public tier releases de-identified 10 Hz signal tables, frozen video embeddings, split manifests, VLM scene attributes, evaluation harnesses, and baseline training code.The gated tier contains raw forward video with faces and license plates blurred and requires a research-only data-use agreement.
  • Privacy protections: Driver identifiers are salted hashes with unreleased salt; VINs, device identifiers, GPS coordinates, cabin video, and audio are excluded.Released route identifiers are opaque strings.
  • Consent and governance: Data collection used informed research consent and participant compensation, while the datasheet documents collection terms, de-identification, residual re-identification risks, and takedown support.The collection followed the source platform’s terms of use permitting research use.
  • Intended use and misuse: DriveDNA is intended for research on driving-style representation, personalized prediction, and evaluation methodology, not re-identification.The release includes leakage probes, asks users to report leakage alongside utility, and prohibits re-identification attempts in the data-use agreement.
  • Hosting: The dataset is hosted on Hugging Face with versioned releases and a DOI, while the configuration and code tier is public.This hosting arrangement supports access to the released dataset materials and code.

8 Discussions and Conclusion … F Condition-Matched Pair Construction and Scene-Attribute Validation

The paper frames driving-style evaluation as a controlled measurement problem: identity, style, and personalization must be separated from vehicle, route, and condition confounds. It releases auditable behavioral and maneuver annotations alongside progressively stricter benchmark splits and condition-matching validation.

  • 8.1 Discussion and Broader Impact: Driver identity, driving style, and personalized prediction are related but distinct, requiring utility, leakage, matched performance, and distributional checks.DriveDNA presents these controls as a measurement checklist for interpreting future style claims.
  • 8.2 Limitations: The corpus is regionally skewed, overrepresents aftermarket-system users, and naturally entangles drivers with vehicles because most drivers appear with one car.Behavioral primitives are weak percentile-rule labels rather than ground-truth style or personality measurements.
  • 8.3 Conclusion: DriveDNA contributes a large-scale, multi-vehicle, human-only corpus and benchmark showing that apparent style can reflect vehicle, place, or driving conditions without confound-aware evaluation.The benchmark discriminates representations, personalization mechanisms, and evaluation procedures across thirty baselines.
  • A Decoding and the Log-Tier Confound: Uniform compact-tier decoding removes log-tier and vehicle-model leakage, while scenarios and primitives are assigned by deterministic, scenario-conditioned percentile rules.The pilot tier probe achieved AUROC 0.607; after uniform 10 Hz redecoding, it returned AUROC 0.500.
  • C Annotation Audit and Threshold Sensitivity: 93.0% overall human agreement supports the behavioral labels, and changing percentile thresholds preserves downstream conclusions and every stratified personalization gain’s sign.Threshold changes from Q80/Q20 to Q75/Q85 move re-identification metrics by less than one seed standard deviation; agreement ranges from 84% to 100% across highlighted primitives.
  • D Maneuver-Event Layer and Its Human Audit: The maneuver layer contains six rule-detected event classes, with five sampled classes reaching 94.8–99.6% precision in human audits.Zero-shot VLMs achieve 35.0% accuracy for Qwen3-VL-4B and 25.3% for Qwen2.5-VL-3B, while dynamics-defined events are poorly recognized.
  • E Benchmark Split Definitions: The benchmark’s additional splits progressively control confounds: within-nameplate removes vehicle-model information, cross-vehicle changes cars, and condition matching fixes driving conditions.These evaluations complement rather than replace the main split.
  • F Condition-Matched Pair Construction and Scene-Attribute Validation: Condition-matched pairs balance drivers within scenario, speed, THW, and vehicle-model cells, while visual validation finds higher agreement on matched than unmatched attributes.Qwen3-VL-4B achieves 97.2% parse success; matched-pair agreement is 1.68× random for road type and 1.21× for traffic density, versus 1.07× for weather.

G Vehicle-Instance Sensitivity and Missing-Channel Robustness

The evaluation probes whether driving-style signals survive vehicle-instance changes and missing input channels. Cross-vehicle verification remains above chance, while the frozen encoder degrades gracefully when individual signal groups are removed.

  • Vehicle-instance sensitivity: Condition-matched pairs fix the consolidated vehicle model, not the physical vehicle, leaving vehicle-instance, calibration, and recording-device differences as possible confounds.Within-nameplate splits retain AUROC .887 for CAN embeddings, while vehicle-model probes read 2.3× chance from steering-angle inputs.
  • Vehicle-instance sensitivity: mean AUROC .750 transfers across vehicles, compared with .816 within a single vehicle under the same verification protocol.The cross-vehicle estimate has driver-level bootstrap 95% CI [.64, .86], versus [.74, .90] within a single vehicle.
  • Missing-channel robustness: Removing steering inputs or pedals costs about six AUROC points, whereas removing motion, curvature, lane, or radar costs at most 1.3 points.The frozen SupCon encoder degrades gracefully rather than collapsing under any single missing signal group.

H Distribution Metrics and Per-Driver Personalization Gains … K Baseline Configurations and Metric Implementation

Personalization improves predicted behavior distributions mainly for a subset of drivers, while evaluation-split handling materially changes estimated gains. The paper also provides an auditable baseline index and implementation details for training, representations, multimodal models, personalization, distributional prediction, and metrics.

  • H Distribution Metrics and Per-Driver Personalization Gains: MMD improved from .430 to .422 and Wasserstein-1 from .349 to .339 after personalization across all three seeds.Distances compare predicted and observed future-behavior statistics within scenario buckets using one CVAE predictive sample per anchor.
  • H Distribution Metrics and Per-Driver Personalization Gains: 55% of 80 unseen drivers had positive distributional NLL gains, averaging +2.4 nats, while point-error gains averaged +0.18% and were positive for 51%.The distributional gain was heavy-tailed, whereas personalization primarily changed predicted distributions for some drivers rather than shifting the conditional mean.
  • I Audit of the Evaluation Procedure: +1.16 nats under a moving evaluation split became +0.068 under a frozen split, with frozen-split training-seed gains of +0.07, +0.07, and +0.17 nats.The moving-split reruns produced +1.16, +0.05, and −0.67 nats across three seeds, demonstrating sensitivity to split variation.
  • J Baseline Index: Table 11 makes the evaluated baseline count auditable and organizes every baseline with the location of its reported results.The index covers representations, shortcut control, personalization, multimodal scene information, and distributional heads.
  • K Baseline Configurations and Metric Implementation: All baselines use AdamW with learning rate 3×10−4, weight decay 10−4, batch size 64–256, and a single RTX 5080 with 16 GB.The joint-patch encoder uses a 600×17 window, patch length 10, 60 tokens, dimension 192, depth 4, attentive statistics pooling, and a 128-dimensional output.
  • K Baseline Configurations and Metric Implementation: The baseline suite includes PatchTST, iTransformer, SupCon/ArcFace, Masked-TS SSL, JEPA-style SSL, MOMENT-1, CLIP, video models, and language or vision-language models.It also evaluates ResidualStyle, DANN, FiLM, query injection, re-identification conditioning, vehicle conditioning, Transformer heads, MDN, and CVAE.
  • K Baseline Configurations and Metric Implementation: KL divergence uses 30-bin histograms with 10−6 smoothing, MMD uses an RBF median-distance bandwidth, and Wasserstein-1 uses empirical distributions within scenario buckets.Buckets with fewer than 20 samples are skipped; metrics are averaged over buckets and behavior features, with exact code shipped in the harness.

L Foundation-Model Baselines: LLMs and VLMs · M Conditioning and Recipe Sensitivity · N Component Versions

Zero-shot language and vision–language baselines underperform trained models, while personalization results depend on conditioning and training recipes. Component versions are documented for auditability, and foundation-model results are reported outside the frozen three-seed protocol.

  • L Foundation-Model Baselines: LLMs and VLMs: Zero-shot foundation-model evaluations are single-pass reference results outside the frozen three-seed protocol.This applies to language and vision–language baselines evaluated beyond MOMENT-1.
  • L Foundation-Model Baselines: LLMs and VLMs: AUROC .596 and .598 are achieved by Qwen3-4B and Llama-3.2-3B, respectively, at 5-minute enrollment, below descriptor and trained-encoder anchors.The descriptor anchor is AUROC .707, while trained encoders reach AUROC .935.
  • L Foundation-Model Baselines: LLMs and VLMs: AUROC .759 is the best VLM event-forecasting result, below the trained 64unit GRU’s AUROC .857 on identical anchors.The evaluation uses a fixed 2,897-anchor subsample with one-third positives.
  • L Foundation-Model Baselines: LLMs and VLMs: Qwen2.5-VL-7B scores .607 on identical anchors, below its 3B sibling, indicating scale does not appear to help in these single zero-shot runs.The passage cautions that these gaps may be calibration-sensitive rather than definitive.
  • M Conditioning and Recipe Sensitivity: Replacing the GRU predictor with a Transformer preserves positive gains of +0.3/+0.5/+0.5% at 1/3/5 s, whereas query injection changes the gain to −0.7%.The comparison shows sensitivity to the conditioning mechanism rather than dependence on recurrent prediction heads.
  • M Conditioning and Recipe Sensitivity: With four independent dropout pathways, only about 39% of training steps see the full conditioning stack, shrinking video and driver gains toward zero.The released alternative recipe keeps vehicle conditioning always on and enables the residual pathway after epoch 5.
  • M Conditioning and Recipe Sensitivity: The style-in-query model showed a late-training loss spike in one run, so its reported number carries a caveat.The passage specifically warns that this result should be interpreted with that limitation.
  • N Component Versions: Table 14 lists every pretrained or named component with its release year and role, enabling direct auditing of baseline-suite age.The accompanying table is titled “Component versions used in the baseline suite.”

O Per-Driver Data Distribution · P Datasheet

DriveDNA combines a broad, multimodal naturalistic corpus with safeguards for skewed per-driver coverage, privacy, and confound-aware personalized driving-style research. Its learned embedding space clusters primarily by driver identity rather than maneuver content, while the datasheet specifies composition, preprocessing, access, and permitted uses.

  • O Per-Driver Data Distribution: 449 drivers have moving human-driving data, with a median contribution of about 20 minutes and the ten largest contributors holding 39% of all hours.The benchmark accommodates this heavy-tailed distribution through 1–10-minute few-shot enrollment and evaluation folds requiring at least two drives.
  • O Per-Driver Data Distribution: SupCon window embeddings form compact clusters for individual drivers in a t-SNE projection using at most 60 windows per driver.The qualitative map includes 428 drivers with windows and is trained on the training fold while displaying all folds.
  • O Per-Driver Data Distribution: 1.8× chance maneuver-label agreement versus 130× driver-identity agreement indicates that the embedding organizes primarily by driver identity rather than maneuver content.Maneuver colors mix within driver clusters, showing only mild local structure.
  • P Datasheet: The datasheet follows the referenced framework and is accompanied by a full datasheet and Croissant metadata record.The paper identifies the corpus as supporting rigorous personalized driving-style modeling where vehicle, route, and condition confounding is measurable.
  • P Datasheet: 465 drivers, 115 distinct vehicle models, and 4,121 decoded drives comprise 975 hours of 10 Hz human-controlled signal tables with forward video.The corpus also includes 62,674 tagged windows and over 276,000 audited maneuver events, but it skews toward one region and aftermarket driver-assistance users.
  • P Datasheet: Data collection repurposes logs already produced by consumer vehicles running openpilot, with informed consent and participant compensation through gift cards.Collection is reported between March 2023 and July 2026, excluding a small number of drives with unsynchronized device clocks from this statistic.
  • P Datasheet: Every drive carries a driver identifier shared across that driver’s different vehicles, while released identifiers are replaced with salted hashes.VINs and device identifiers are dropped, GPS is absent from released signals, and faces and license plates are blurred in gated video.
  • P Datasheet: The dataset is intended for driving-style representation learning, personalized behavior prediction, and confound or robustness research, while prohibiting re-identification, surveillance, and identifiable-person insurance or employment decisions.The release is tiered across public signals, embeddings, splits, and code, with raw video governed by a data-use agreement and versioned errata.
Loading 2607.23822v1…