Source-linked AI summary
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
Yang Zhou, Hao Shao, Letian Wang, Zhuofan Zong, Hongsheng Li, Steven L. Waslander
TL;DR
Driving world-model research lacks a rigorous benchmark covering driving-specific visual, trajectory, temporal, controllability, and data-diversity requirements. DrivingGen addresses this gap with a diverse dataset and multifaceted evaluation suite, and its 14-model benchmark reveals trade-offs between visual quality and physical consistency.
Problem
Existing evaluations overlook driving-specific imaging, trajectory plausibility, temporal and agent consistency, controllability, and diverse deployment conditions.
Method
DrivingGen combines diverse driving data with metrics for video and trajectory distribution, quality, temporal consistency, and trajectory alignment.
Results
Benchmarking 14 models reveals that general models can look better while breaking physical consistency, whereas driving-specific models can achieve realistic motion but lower image fidelity.
Takeaways & Limitations
DrivingGen provides a unified framework for evaluating strengths and failure modes in generative driving world models.
Takeaways & Limitations
The current benchmark evaluates open-loop video generation; standardized closed-loop world-generation evaluation remains unavailable.
Abstract
from arXiv · showhide
Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving, this vision gives rise to driving world models: generative simulators that imagine ego and agent futures, enabling scalable simulation, safe testing of corner cases, and rich synthetic data generation. Yet, despite fast-growing research activity, the field lacks a rigorous benchmark to measure progress and guide priorities. Existing evaluations remain limited: generic video metrics overlook safety-critical imaging factors; trajectory plausibility is rarely quantified; temporal and agent-level consistency is neglected; and controllability with respect to ego conditioning is ignored. Moreover, current datasets fail to cover the diversity of conditions required for real-world deployment. To address these gaps, we present DrivingGen, the first comprehensive benchmark for generative driving world models. DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, spanning varied weather, time of day, geographic regions, and complex maneuvers, with a suite of new metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality. DrivingGen offers a unified evaluation framework to foster reliable, controllable, and deployable driving world models, enabling scalable simulation, planning, and data-driven decision-making.
1 University of Toronto 2 CUHK MMLab
The paper provides a project website for DrivingGen.
- DrivingGen’s project website is https://drivinggen-bench.github.io/.
- The benchmark information is available at https://drivinggen-bench.github.io/.
- The authors direct readers to https://drivinggen-bench.github.io/ for the project.
1 INTRODUCTION
DrivingGen addresses the lack of driving-specific evaluation by combining a diverse dataset with metrics for visual quality, trajectories, temporal consistency, and controllability. Benchmarking 14 models exposes trade-offs between visual fidelity and physical consistency.
- Existing evaluations and datasets underrepresent driving-specific imaging constraints and diverse weather, geography, and interactions.
- DrivingGen is proposed as a comprehensive benchmark for generative driving world models.
- The evaluation dataset spans varied weather, times of day, global regions, and complex driving scenarios.
- Its metrics jointly assess distribution realism, visual quality, temporal coherence, and trajectory or control fidelity.
- DrivingGen benchmarks 14 generative world models across general, physics-based, and driving-specialized categories.
- General models can produce visually appealing traffic scenes while violating physical consistency, whereas driving-specific models can prioritize trajectory accuracy over image fidelity.
2 RELATED WORKS
The related-work section focuses on generative world models for autonomous driving and benchmarks for evaluating them. It also points to diverse benchmark scenarios and dataset examples.
- The paper reviews generative world models applied to autonomous driving and benchmarks for evaluating these models.
- DrivingGen’s benchmark gallery includes dense city traffic at night, unusual weather, and complex agent interactions.
- The dataset distribution and representative examples are presented in Figure 2.
3 DRIVINGGEN BENCHMARK
DrivingGen combines diverse open-domain and ego-conditioned data tracks with driving-specific metrics for videos and trajectories. Its evaluation covers distribution, quality, temporal consistency, and trajectory plausibility.
- 3 DRIVINGGEN BENCHMARK: DrivingGen uses a diverse dataset and multifaceted metrics to evaluate generative driving world models from visual and robotics perspectives.
- 3.1 BENCHMARK DATASET: The dataset has open-domain and ego-conditioned tracks for unseen-scenario generalization and trajectory controllability.
- 3.1 BENCHMARK DATASET: Each sample combines a front-view RGB image, scene description, and optional ego trajectory, with 400 samples selected for efficient evaluation.
- 3.1 BENCHMARK DATASET: The benchmark balances weather, time of day, geography, and complex maneuvers to broaden evaluation beyond ordinary daytime driving.
- 3.2.1 DISTRIBUTION: Distribution metrics compare generated videos and trajectories with data using FVD and the proposed Fréchet Trajectory Distance.
- 3.2.2 QUALITY: Quality metrics cover perceptual video quality, driving-specific visual factors, and kinematic trajectory plausibility.
- 3.2.3 TEMPORAL CONSISTENCY: Temporal consistency metrics assess scene and agent continuity in videos and speed and acceleration consistency in trajectories.
4 EXPERIMENTS
DrivingGen evaluates 14 generative world models with comprehensive metrics and reveals consistent trade-offs between visual quality, trajectory fidelity, and robustness. Closed-source models lead overall, while open-source and driving-specialized models are competitive in selected dimensions but remain limited in combined performance and trajectory alignment.
- DrivingGen evaluates 14 generative world models across closed-source, open-source, physical-world, and driving-specific categories.
- Closed-source models lead visual quality and overall ranking while generally preserving agent behavior and scene coherence over time.
- CogVideoX and Wan achieve strong low FVD across both tracks, showing that open-source models can excel on targeted metrics without leading overall.
- No single model combines strong visual realism with trajectory fidelity, as driving-specialized models follow commanded paths more accurately but underperform visually.
- Under ego-trajectory conditioning, significant ADE/DTW errors reflect both video artifacts that impair trajectory recovery and imperfect motion generation.
- Joint evaluation of distribution, perceptual quality, temporal consistency, and trajectory alignment exposes failure modes hidden by single metrics such as FVD.
5 CONCLUSION
DrivingGen introduces a diverse benchmark and multifaceted metrics for evaluating generative driving world models. Benchmarking reveals trade-offs among visual fidelity, physical consistency, and controllability, supporting a unified framework for future development.
- DrivingGen combines diverse driving data with metrics measuring visual realism, trajectory plausibility, temporal coherence, and controllability.
- The benchmark reveals critical trade-offs among visual fidelity, physical consistency, and controllability across current driving world models.
- DrivingGen establishes a unified and reproducible framework for developing reliable and deployment-ready driving world models.
6 FUTURE WORK AND LIMITATIONS
DrivingGen identifies several boundaries of the current benchmark and outlines extensions involving broader data, closed-loop evaluation, richer modalities, scene control, counterfactuals, and score aggregation.
- Expanding More Meaningful Data: 400 samples balance efficiency and practicality but may not fully cover the long tail of driving scenarios.Scaling to thousands of clips is proposed as a future direction as generation and data availability improve.
- Interactive and Closed-Loop Simulation: Closed-loop evaluation is not standardized because current generative video world models target open-loop generation.The paper points to interactive simulators such as CARLA or closed-loop dataset simulation such as Navsim.
- Downstream Tasks Metrics and Enriching data modality: The current single front-view camera limits structural driving generation and prevents fair downstream-task evaluation requiring synchronized multi-camera footage and map knowledge.Future extensions include multi-view video, LiDAR, and HD maps, alongside metrics such as view consistency.
- Evaluation of Scene Controllability and State Transformation: Unified scene-level controllability evaluation is omitted because models differ in control support and datasets pose complexity challenges.Potential controls include pedestrian behavior and lane configuration.
- Counterfactual Reasoning Evaluation: DrivingGen does not explicitly evaluate counterfactual reasoning because it focuses on real driving videos and scenarios that actually occurred.Future metrics could test hypothetical events or modifications.
- Overall Score: Average rank is a quick summary rather than a definitive score, while a composite index would require normalized distributions and aligned metrics.The full metric table is provided transparently.
A.2 BENCHMARKS FOR EVALUATING GENERATIVE WORLD MODELS
Existing video-generation benchmarks evaluate multifaceted outputs, while DrivingGen uses an ego-conditioned track assembled from multiple driving datasets to diversify evaluation conditions.
- Existing Benchmarks: Video-generation benchmarks such as VBench evaluate models with multifaceted metrics based on human-collected prompts.Recent evaluations also extend toward open, dynamic, and complex world-simulation scenarios.
- DrivingGen Ego-Condition Track: DrivingGen’s ego-conditioned track includes distributions and galleries documenting its evaluation data.The supplied figure captions identify these as statistics and gallery views of the track.
- DrivingGen Ego-Condition Track: Data from five open-sourced driving datasets diversify weather, time of day, locations, and driving styles.Videos and ego-trajectories serve as target distributions for metrics including FVD and FTD.
B.2 DETAILS OF OUR SLAM PIPELINE AND COMPARIISION WITH OTHERS
The pipeline reconstructs complete camera trajectories, including failed SLAM cases, and compares trajectory-distribution and trajectory-quality evaluation components. It also defines complementary visual-flicker analysis using luminance spectra.
- SLAM failure handling: Failed SLAM reconstructions are explicitly completed rather than discarded, avoiding artificially inflated scores for unrealistic videos.The method extrapolates trajectories with small random pose-orientation perturbations instead of freezing the camera.
- Trajectory distribution: FTD measures Fréchet distance between generated and reference trajectory distributions using MTR agent-polyline embeddings.Trajectories are windowed to the encoder’s fixed horizon, then pooled by averaging window embeddings before distributional comparison.
- FTD implementation: Trajectory embeddings use window slicing, mean pooling, numerical regularization, and optional covariance shrinkage for stable distance computation.The default recipe uses H=10 steps, stride s=H, and identical slicing for generated and reference trajectories.
- Image-quality analysis: MMP analyzes frame-luminance periodograms to quantify dominant non-DC flicker within a frequency band.The default procedure uses ∆f=0.5 Hz, τ=0.05, fps=10, and one FFT per clip with complexity O(T log T).
B.5 TRAJECTORY QUALITY
Trajectory quality combines kinematic comfort, motion, and curvature, while consistency evaluates smoothness and agent disappearances. Human validation indicates overall metric agreement, but trajectory-related metrics are less accurate.
- Trajectory quality: Trajectory quality aggregates comfort, motion, and curvature through a weighted geometric mean, with higher scores indicating better kinematic quality.Each submetric lies in [0, 1], and dataset means are reported while skipping NaNs.
- Comfort: Comfort scores longitudinal jerk, lateral acceleration, and yaw rate, while non-moving or too-short trajectories receive NaN.Each component uses an inverse transform before geometric aggregation.
- Motion: Motion quality penalizes under-mobility using mean speed, while never-moving trajectories receive 0.The mapping uses vref=6.0 m/s and k=2.5 by default.
- Curvature: Curvature quality maps RMS curvature to Scurv=1/(1+κrms), so larger curvature produces a lower score.Non-moving trajectories return NaN.
- Agent consistency: Disappearance consistency classifies agent tracklets as natural or unnatural and scores videos clean only when all evaluated tracklets are non-abnormal.The final score is the percentage of clean videos, with higher values better.
- Human alignment: Human-alignment evaluation compares distribution, quality, and consistency metrics across videos and trajectories using model win ratios.Trajectory-related metrics are less accurate than human preferences, likely because generated-video artifacts impair SLAM and depth recovery.
B.8 TIME AND RESOURCE FOR DRIVINGGEN
DrivingGen’s main computational bottleneck is video generation, while evaluation is comparatively manageable but still resource-intensive for visual metrics.
- Generation cost: Wan2.2-14B requires about 20–30 minutes to generate one 100-frame video on a single GPU with at least 40 GB memory.Running all metrics for 400 videos with 100 frames takes roughly 1–2 days on a single modern GPU.
- Evaluation cost: Agent and disappearance consistency are the most time-consuming evaluation metrics because they run models for each agent.Trajectory measures such as FTD, quality, consistency, and alignment take only minutes because they use compact embeddings or low-dimensional trajectories.
B.9 HUMAN ALIGNMENT OF DRIVINGGEN
Human alignment is assessed through pairwise model preferences and win ratios across selected video and trajectory metric categories.
- Win-ratio computation: Win ratio equals a model’s total pairwise-comparison score divided by the number of comparisons in which it participated.A selected model scores 1, while tied models each score 0.5.
- Evaluated categories: The fast evaluation covers distribution, quality, and consistency using both videos and trajectories, with results reported in Figure 5.The primary metrics are FVD and FTD, subjective image quality and trajectory quality, and video and trajectory consistency.