Source-linked AI summary
WorldSimBench: Towards Video Generation Models as World Simulators
Yiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang, Enshen Zhou, Lijun Li, Zhenfei Yin, Xihui Liu, Lu Sheng, Jing Shao, Lei Bai, Wanli Ouyang, Ruimao Zhang
TL;DR
Predictive-model research lacks a capability-based categorization and benchmarks for highly embodied models. This paper defines an embodiment hierarchy and introduces WorldSimBench, combining human-centered visual evaluation with action-level testing across three embodied scenarios. Its analyses compare perceptual and manipulative findings while identifying scope beyond robot-focused scenarios.
Problem
Predictive models lack a categorization based on inherent characteristics, while existing benchmarks inadequately evaluate highly embodied models from an embodied perspective.
Method
WorldSimBench classifies predictive models hierarchically and evaluates World Simulators through Explicit Perceptual and Implicit Manipulative Evaluation across three embodied scenarios.
Results
WorldSimBench provides comprehensive evaluation of multiple video generation models, with perceptual and closed-loop manipulative findings generally consistent across most dimensions.
Takeaways & Limitations
The benchmark offers insights intended to guide future research on World Simulators and video generation models.
Takeaways & Limitations
The evaluation covers embodied intelligence from a robot-oriented perspective, while other scenarios and physical representations require further exploration.
Abstract
from arXiv · showhide
Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing benchmarks are unable to effectively evaluate higher-capability, highly embodied predictive models from an embodied perspective. In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench. WorldSimBench includes Explicit Perceptual Evaluation and Implicit Manipulative Evaluation, encompassing human preference assessments from the visual perspective and action-level evaluations in embodied tasks, covering three representative embodied scenarios: Open-Ended Embodied Environment, Autonomous, Driving, and Robot Manipulation. In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulators. In the Implicit Manipulative Evaluation, we assess the video-action consistency of World Simulators by evaluating whether the generated situation-aware video can be accurately translated into the correct control signals in dynamic environments. Our comprehensive evaluation offers key insights that can drive further innovation in video generation models, positioning World Simulators as a pivotal advancement toward embodied artificial intelligence.
1 INTRODUCTION
Predictive models support increasingly embodied predictions, but existing categorization and evaluation methods inadequately assess highly embodied models. WorldSimBench addresses this gap with complementary perceptual and manipulative evaluations across three embodied scenarios.
- Motivation: Predictive models use observations and objectives to predict future states for planning, visual guidance, and action generation in embodied tasks.These models are applied from agent development to robot control.
- Evaluation Gap: Existing evaluations emphasize text planning or aesthetic visual quality, missing physical properties such as perspective consistency and object breakability.This limits assessment of highly embodied predictive models.
- WorldSimBench: WorldSimBench introduces Explicit Perceptual Evaluation and Implicit Manipulative Evaluation to assess visual fidelity and video-to-action transformation.The framework evaluates both visual and action levels.
- WorldSimBench: The benchmark covers Open-Ended Embodied Environment, Autonomous Driving, and Robot Manipulation scenarios.These scenarios test scenario-specific attributes in generated videos and actions.
- Contributions: HF-Embodied contains 35,701 annotated tuples with multidimensional scores and fine-grained human feedback for video assessment.The dataset is constructed from prompted video clips, human feedback, and annotation.
2 RELATED WORK
Related work develops predictive models across text, image, video, and actionable-video modalities, while evaluation methods primarily target planning, task completion, or aesthetic quality. Recent World Simulators improve physical and 3D representations, but predictive videos can still lack physical and logical consistency.
- Predictive Models: Predictive text, image, and video models generate plans, future goal images, and future videos for embodied decision-making and control.These modalities correspond to progressively richer predictive outputs.
- Predictive Models: Predictive video models have advanced embodied control, but limited data or model capacity can leave generated videos without essential physical representations and logical consistency.Such limitations restrict applicability to fixed scenarios and single tasks.
- Predictive Models: Diffusion transformers and large-scale internet video datasets have helped some actionable-video models represent physical laws and 3D environments more precisely.These actionable-video models are also called World Simulators.
- Evaluation of Predictive Models: Existing benchmarks evaluate predictive text models through text-level planning or task completion and predictive image models through aesthetic score-based comparisons.The passage describes evaluation across earlier predictive-model stages.
3 PREDICTIVE MODEL CATEGORY DEFINITION
The paper organizes predictive models into stages S0–S3 according to capability and embodiment. S3 models, called World Simulators, generate videos that follow physical rules and align with executed actions.
- Hierarchy: The hierarchy categorizes predictive models by their capabilities and level of embodiment.The stages progress from textual predictions to actionable video predictions.
- Stage S0: S0 models generate textual predictions and are evaluated through text-level planning and task-completion benchmarks.Their outputs are limited to textual modality.
- Stage S1: S1 models generate visual predictions without temporal information and are evaluated for aesthetic quality.The stage concerns generated images.
- Stage S2: S2 models generate video predictions, but evaluation at this stage focuses solely on aesthetic quality because of limited model capabilities.The passage distinguishes video generation from the evaluation criterion.
- Stage S3: S3 World Simulators generate videos that adhere to physical rules and align with executed actions.WorldSimBench is designed specifically for this stage.
4 WORLDSIMBENCH CONSTRUCTION
WorldSimBench evaluates World Simulators at perceptual and manipulative levels across open-ended environments, autonomous driving, and robot manipulation. Its construction combines human-annotated video assessment with closed-loop video-to-action task evaluation.
- WorldSimBench evaluates embodied capabilities through human-perceived video quality and closed-loop conversion of generated videos into control signals.
- The benchmark covers Open-Ended Embodied Environment, Autonomous Driving, and Robot Manipulation scenarios.
- Explicit Perceptual Evaluation: The HF-Embodied Dataset is built from scenario-specific videos, hierarchical dimensions, generated clips, and fine-grained human annotations with dimension-specific reasoning.
- Explicit Perceptual Evaluation: Hierarchical evaluation dimensions group assessment into Visual Quality, Condition Consistency, and Embodiment, with scenario-specific measures for trajectories, perspective, interaction, and velocity.
- Explicit Perceptual Evaluation: The Human Preference Evaluator takes a generated video and prompt as input and outputs a scenario-specific score aligned with human perception.
- Implicit Manipulative Evaluation: In Implicit Manipulative Evaluation, a video generation model predicts future videos from instructions and observations, while a pre-trained video-to-action model converts them into executable controls.
5 EXPERIMENTS
WorldSimBench evaluates eight video generation models across three embodied scenarios using explicit human-aligned perceptual scoring and implicit video-to-action assessment. Results reveal scenario- and task-dependent weaknesses, while visual quality generally corresponds with closed-loop action performance.
- Eight video generation models are evaluated across Open-Ended Embodied Environment, Autonomous Driving, and Robot Manipulation.
- The Human Preference Evaluator scores selected instruction-generated videos across evaluation dimensions, while video-to-action models assess interactive performance in simulation.
- After fine-tuning on HF-Embodied Dataset, the Human Preference Evaluator consistently surpasses GPT-4o in alignment with human preferences across all scenarios.
- In Open-Ended Embodied Environment, most models struggle with embodied interaction, especially plausible object deformations such as block shattering.
- In Autonomous Driving, models often achieve instruction alignment for simple movements but produce poor 3D depth and unrealistic pedestrians or vehicles.
- Closed-loop results vary substantially by task: first-frame-conditioned models underperform in Open-Ended Embodied Environment, trajectory-capable models perform better in Autonomous Driving, and harder manipulation tasks favor more robust models.
- Visual and closed-loop evaluations are generally consistent, although Dynamicrafter underperforms Open-Sora-Plan in frequent-interaction and long-sequence tasks.
6 CONCLUSION
The paper introduces WorldSimBench as a dual framework for evaluating World Simulators through perceptual and manipulative tests. Its evaluation identifies directions for future World Simulator research, while broader physical scenarios remain underexplored.
- WorldSimBench combines Explicit Perceptual Evaluation and Implicit Manipulative Evaluation to assess video generation models as World Simulators.
- The framework supports comprehensive evaluation and analysis of multiple video generation models as World Simulators.
- The evaluation covers physical rules and 3D content from an embodied-intelligence perspective, but other World Simulator scenarios and physical representations require further exploration.
A TAXONOMY IN EXPLICIT PERCEPTUAL EVALUATION
The explicit perceptual taxonomy organizes evaluation criteria by visual quality, condition consistency, and embodiment across embodied scenarios. These dimensions capture scene stability, instruction fidelity, physical plausibility, and interaction realism.
- Visual Quality: Visual Quality includes background consistency and foreground consistency, which assess whether scene elements remain stable throughout the video.
- Condition Consistency: Condition Consistency includes Instruction Alignment and Scenario Alignment, assessing correspondence between the input instruction, embodied scenario, and generated video.
- Embodiment: Open-Ended Embodied Environment embodiment evaluates velocity, embodied interaction, and trajectory for appropriate and logical object motion.
- Embodiment: Autonomous Driving embodiment evaluates perspectivity, lighting and shadows, trajectory logic, and key elements of the generated scene.
- Visual Quality and Condition Consistency: Robot Manipulation evaluates visual consistency and instruction alignment for the manipulation table, robotic arm, object, and commanded action.
- Embodiment: Robot Manipulation embodiment assesses perspectivity and whether object shape and posture conform to collision rules during interaction.
B DETAILD IMPLEMENTATION OF EXPLICIT PERCEPTUAL EVALUATION
The explicit perceptual evaluation uses scenario-specific instructions, datasets, model conditioning, and a Human Preference Evaluator to score generated videos. The evaluator is trained and compared with GPT-4o across the three scenarios.
- The HF-Embodied Dataset contains scenario-specific video assessment data, while Table 4 identifies positive AD and RM samples as scores higher than 3.
- The study evaluates eight video generation models through Explicit Perceptual Evaluation and Implicit Manipulative Evaluation across OE, AD, and RM.
- Dataset Construction: OE training uses the OpenAI Contractor Gameplay Dataset and supplementary Explore trajectories generated by pretrained Steve-1 agents.
- Dataset Construction: AD training uses nuScenes clips sampled as 25 frames at 10 Hz and maps vehicle actions to textual commands.
- Dataset Construction: RM training uses RH20T-P primitive-level manipulation instructions and videos, excluding coordinate-specific instructions to improve generalization.
- Model Configuration: Open-Sora-Plan(TI2V) is created by adding first-frame conditioning to Open-Sora-Plan(T2V), while other models receive no structural adjustments.
- Evaluator Training: The Human Preference Evaluator receives sampled video frames and prompts specifying the scenario, generation instruction, evaluation dimension, and dimension definition.
- Evaluator Results: After fine-tuning, the Human Preference Evaluator surpasses GPT-4o across all dimensions and scenarios, including challenging RM dimensions where GPT-4o shows negative correlation.
C DETAILED RESULT OF EXPLICIT PERCEPTUAL EVALUATION
The implicit manipulative evaluation tests video-action consistency across open-ended embodied environments, using generated videos to drive low-level control and programmatic metrics. Results reveal strong scene generation in some settings but persistent difficulty with dynamic movement and task-oriented temporal reasoning.
- Open-Ended Embodied Environment: Models generally score highly on Velocity because generated videos contain few moving objects, while dynamic environments remain challenging.The evaluation also finds scenario consistency higher than alignment with task-oriented instructions.
- Open-Ended Embodied Environment: Models generate corresponding scenes more reliably than the temporal actions required for task completion.This gap indicates difficulty reasoning about action sequences over time.
- Robot Manipulation: Robot manipulation models struggle with instruction alignment, often moving without clear objectives rather than producing task-directed actions.Aimless movement can artificially inflate Embodied Interaction and Trajectory scores by reducing object-interaction and penetration errors.
- Open-Ended Embodied Environment: Open-ended embodied evaluation adapts Steve-1’s action space so video generation models function as low-level embodied controllers.The pipeline uses programmatic evaluation to measure control capabilities across collecting, exploration, and digging tasks.
D.3 FULL RESULT
The open-ended embodied and autonomous-driving evaluations measure downstream control under text or multimodal conditions. Results show strong task performance for OpenSora in text-only open-ended tasks, while added image inputs and complex driving scenes expose substantial weaknesses.
- Open-Ended Embodied Environment: The open-ended evaluation reports five tasks and averages scores after dividing travel distance by 10 to equalize task ranges.Models are compared under Text and Text & Image conditions.
- Open-Ended Embodied Environment: OpenSora achieves the highest text-condition average score of 27.80 in open-ended embodied tasks.It performs especially well on collect dirt (70.20) and travel distance (339.87).
- Open-Ended Embodied Environment: Adding image input does not consistently improve open-ended task performance and can reduce success rates.Open-Sora-Plan’s travel distance score drops from 342.91 to 195.14 under the Text & Image condition.
- Autonomous Driving: Autonomous-driving evaluation uses short LangAuto-Tiny routes, with generated videos converted into waypoints and then control signals.The benchmark covers diverse CARLA towns and environmental conditions, while Driving Score combines route completion and infraction score.
E.3 FULL RESULT
The autonomous-driving evaluation connects generated videos to closed-loop control across diverse CARLA conditions and reports both safety and progress metrics. Performance differences emphasize the importance of coherent trajectories, instruction alignment, and accurate environmental representation.
- Results: Open-Sora-Plan produces high-quality videos that support trajectory generation, instruction following, and environment perception in autonomous driving.Its performance is evaluated across eight CARLA metrics, including route completion, driving score, collisions, and traffic violations.
- Results: DynamiCrafter and EasyAnimate struggle to generate detailed, consistent content in complex or dynamic driving scenes.The paper identifies video quality, scene understanding, and task alignment as improvement areas.
- Future Development: Future progress requires more accurate trajectories, stronger instruction following, and better representation of pedestrians, vehicles, and varied terrains.These qualities provide stronger input for real-time decision-making in the control system.
- Future Development: Advancing trajectory accuracy, instruction alignment, and environment representation is identified as crucial for overall autonomous-driving performance.
F.3 FULL RESULT
The robot-manipulation evaluation uses World Simulators with a pretrained video-to-action policy for zero-shot tasks in CALVIN’s held-out environment. Open-Sora-Plan is the most consistent model, while success declines as manipulation sequences become longer and more complex.
- Results: Open-Sora-Plan achieves an average task length of 2.95, indicating consistent completion of manipulation sequences.The reported metrics are success rate and average task length across five-task evaluation sequences.
- Results: DynamiCrafter reaches a 0.95 success rate on the initial task but declines as task complexity increases.This pattern suggests limitations in completing longer manipulation sequences.
- Results: Success rates decrease as task complexity increases, highlighting the need for better video-to-action translation in longer manipulation sequences.Open-Sora-Plan emerges as the most capable model for completing multiple tasks in succession.