Source-linked AI summary

iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework

Jianjie Fang, Yingshan Lei, Qin Wan, Ziyou Wang, Yuchao Huang, Yongyan Xu, Baining Zhao, Weichen Zhang, Chen Gao, Xinlei Chen, Yong Li

arXiv:2605.03941v2cs.CVcs.AI

TL;DR

Existing benchmarks lack diverse scenes, modalities, difficulty levels, and memory tasks for evaluating interactive world models. iWorld-Bench introduces a unified action-based benchmark with six task types and evaluates 14 models, revealing interaction limitations and trade-offs between generation quality and controllability.

  • Problem

    Existing benchmarks lack diverse scenes and perspectives and do not comprehensively test responsiveness to actions or model memory.

  • Method

    iWorld-Bench standardizes diverse video data and uses a modality-agnostic Action Generation Framework to define six interactive task types.

  • Results

    Evaluating 14 interactive world models across six tasks reveals limitations in interaction capability and trade-offs between generation quality and controllability.

  • Takeaways & Limitations

    iWorld-Bench enables systematic comparison of interactive world models across diverse worlds and interaction modalities.

Abstract

from arXiv · show

Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks large-scale datasets and unified benchmarks to evaluate their physical interaction capabilities. To address this, we propose iWorld-Bench, a comprehensive benchmark for training and testing world models on interaction-related abilities such as distance perception and memory. We construct a diverse dataset with 330k video clips and select 2.1k high-quality samples covering varied perspectives, weather, and scenes. As existing world models differ in interaction modalities, we introduce an Action Generation Framework to unify evaluation and design six task types, generating 4.9k test samples. These tasks jointly assess model performance across visual generation, trajectory following, and memory. Evaluating 14 representative world models, we identify key limitations and provide insights for future research. The iWorld-Bench model leaderboard is publicly available at iWorld-Bench.com.

1. Introduction

iWorld-Bench addresses the lack of diverse, unified evaluation for interactive world models by combining a large-scale dataset, a modality-agnostic Action Generation Framework, and tasks assessing visual generation, action following, and memory. It provides 4,900 evaluation tasks and 9 metrics for analyzing 14 interactive world models.

  • Motivation: Interactive world models generate causally consistent environmental responses to external action sequences, enabling bidirectional communication between agents and environments.Actions include camera movements and keyboard inputs.
  • Limitations: Existing benchmarks lack sufficient diversity in scenes and perspectives and omit varying-difficulty and memory tasks for interactive world models.These limitations motivate a dedicated framework for evaluating responsiveness to external action sequences and interactivity.
  • Evaluation: 4,900 evaluation tasks and 9 evaluation metrics assess visual quality, action following, and memory ability across interactive world models.The benchmark evaluates 14 existing interactive world models and analyzes their strengths and limitations.
  • Dataset: 330,000 high-quality video clips cover multiple scenes and perspectives, while 2,100 selected videos form a high-quality evaluation subset.The dataset includes 4 world observation perspectives, 9 outdoor weather variations, 5 indoor lighting differences, day-night transitions, and thousands of diverse scenes.
  • Action Generation Framework: A general Action Generation Framework unifies interaction-capability evaluation across different world-model modalities and supports 6 task types.The framework uses a complete action-space dictionary and modality-agnostic encoding.

2. Related Work

Related work spans camera-parameterized datasets, interactive world models, and their benchmarks. Existing benchmarks and interaction methods leave action-sequence and memory evaluation insufficiently standardized, motivating iWorld-Bench’s Unified Action Generation Framework.

  • Datasets with Camera Parameters: High-quality first-person datasets with precise intrinsic and extrinsic camera parameters support interactive world-model development.The datasets are grouped into autonomous driving, robotics, and drone datasets; 3D reconstruction datasets; and large-scale training datasets.
  • World Model Benchmark: Existing interactive world-model benchmarks primarily evaluate text-to-video generation under text control.They lack assessments for action sequence generation and specifically designed memory-related tasks.
  • Interactive World Model: Interactive world models use text-controlled, one-hot encoding, or intrinsic-and-extrinsic camera-parameter interaction.Text control offers limited degrees of freedom, whereas camera-parameter interaction enables complex camera controls.
  • Interactive World Model: iWorld-Bench introduces a Unified Action Generation Framework to generate standardized action tasks across interactive world models.The framework addresses differing interaction modalities by facilitating unified evaluation.

3. iWorld-Bench

iWorld-Bench provides a unified framework for evaluating world-model interaction capabilities across modalities. It combines a large, processed video corpus with human-verified evaluation samples, a unified action encoding framework, and six task types spanning action following, memory, and camera trajectory following.

  • Dataset Construction: 330,000 video clips were processed into 2,100 selected evaluation videos and 4,900 interactive tasks.The pipeline collected 27.8M multi-image data, standardized and filtered it into 330K video data, then used VLM annotation and human verification.
  • Dataset Construction: 18 high-quality environments across 4 outdoor urban simulators expand coverage beyond predominantly indoor world-model training data.The collection program used 450 manually identified high-quality points within these 18 scenes.
  • Action Generation Framework: The Action Generation Framework accepts world-model inputs from any modality through interactive action encoding and unified encoding mapping.It maps actions to intrinsic and extrinsic camera parameters, one-hot encodings, and text control signals, supporting flexible combinations and extensibility.
  • Action Generation Framework: 81 combined actions form the unified available translational and rotational action spaces after mapping each space to 9 actions.The framework prioritizes these actions and assigns each a unique encoding for compatibility across world models.

4. Experiments

Experiments evaluate 14 representative interactive world models across action control, memory, and camera following, with metrics validated against human preferences. Results show that discrete action conditioning and camera-parameter control improve controllability, while visual quality and trajectory following remain partially trade-offs.

  • Experimental Setup: 14 representative interactive world models are evaluated across text-conditioned, one-hot–conditioned, and camera-parameter control paradigms.The evaluated set includes five text-conditioned models, two one-hot–conditioned models, and seven models using explicit camera intrinsics and extrinsics.
  • Metric Validation: Human preference validation confirms that iWorld-Bench metrics align with subjective world-model performance and remain robust across video resolutions and aspect ratios.The validation study directly tests agreement with human perceptions and metric stability under presentation changes.
  • Action Control and Memory Ability: One-hot models show superior interaction capabilities, text-controlled models lead generation quality, and text-controlled CogVideoX-I2V reaches Brightness Consistency of 0.8988 while sacrificing trajectory accuracy.The results expose a trade-off between visual consistency and action controllability, whereas HY-World 1.5 and videox-fun-Wan provide more balanced profiles.
  • Action Control and Memory Ability: HY-World 1.5 ranks first with an average score of 0.7873, including trajectory accuracy of 0.7472 versus CogVideoX-I2V’s 0.5950.HY-World 1.5 excels in memory ability and trajectory following, supporting the advantage of discrete action signals over continuous text descriptions.
  • Camera Following: Camera-control fine-tuning improves trajectory following but slightly degrades generation-quality metrics relative to base models.This pattern is reported for AC3D versus CogVideoX-I2V and HY-World 1.5 versus HunyuanVideo-1.5.
  • Camera Following: AC3D achieves the best camera-following results, with Trajectory Tolerance of 0.9091, Brightness Consistency of 0.8927, and Motion Smoothness of 0.9919.Across camera-parameter models, Trajectory Tolerance ranges from 0.4286 for ASTRA to 0.9091 for AC3D, while RealCam-I2V leads Image Quality at 0.5889 but has Trajectory Tolerance of 0.7480.

5. Conclusion and Future Work

iWorld-Bench provides a unified, multidimensional benchmark for interactive world models, evaluating six interactive task types across diverse worlds. Extensive evaluations of 14 interactive world models and their base models reveal limitations in current interaction capabilities and differences among interaction paradigms.

  • Benchmark Contributions: iWorld-Bench is a unified benchmark specifically designed for interactive world models.It integrates a multidimensional evaluation framework.
  • Benchmark Contributions: Six types of interactive tasks systematically evaluate model performance across diverse worlds.
  • Findings: Evaluations of 14 interactive world models and their base models expose limitations in current interaction capabilities.The analysis also examines differences and shortcomings across interaction paradigms.

A. Dataset details … A.4. Data Annotation Prompts

The appendix details the benchmark’s heterogeneous source datasets and selected simulator environments, then describes a two-stage refinement pipeline and GPT-4o-based annotation process for producing standardized, high-fidelity data.

  • A.2. Simulator Details: The simulator pool comprises aerial VLN, UAV ON, Openfly, and Embodied City, with 8, 5, 4, and 1 environments selected, respectively.Openfly’s four environments were all retained, while representative high-quality subsets were chosen for aerial VLN and UAV ON.
  • A.3. Fliter: The curation method uses a modular two-stage pipeline that separates visual metric extraction from temporal sequence pruning.This supports incremental filtering updates without re-evaluating the entire dataset through a detect-then-prune workflow.
  • A.3. Fliter: High-fidelity zones are defined by anomaly density below τ = 0.06, then proximal clean segments are merged and segments shorter than Lmin frames are discarded.The procedure yields temporally stable, visually coherent segments with high-fidelity motion trajectories.
  • A.3. Fliter: The broader corpus pipeline consists of multi-source acquisition, structural standardization, spatial rectification, and interactive model synthesis.It calibrates source coordinate systems, maps axes to a global motion reference, and aligns trajectories to a unified right-handed coordinate system.
  • A.4. Data Annotation Prompts: 330,000 videos were annotated with GPT-4o using 119 million input tokens and 21.86 million output tokens at an approximate cost of $518.The process used GPT-4o 2025-01-01-preview and a structured prompt pipeline.
  • A.4. Data Annotation Prompts: The annotation prompt analyzes five frames as a continuous first-person sequence and requires strictly valid JSON output.It determines indoor or outdoor status, describes and categorizes the scene, records weather or lighting, and extracts 15+ visible entities.

A.5. Multi-Model Verification and Human Refinement … B.3. Existing Keyboard and Mouse Support Space

The paper validates annotations through multi-model agreement and targeted human refinement, then defines a unified motion-action space spanning unrestricted combinations, existing controls, and memory mappings.

  • A.5. Multi-Model Verification and Human Refinement: 330,000 annotations were independently verified by Gemini 3.0 Flash, Qwen-VL-Max, and Kimi-K2.5 rather than accepting GPT-4o labels directly.Each verifier made a conservative binary judgment using sampled frames and existing annotation fields.
  • A.5. Multi-Model Verification and Human Refinement: 61,380 non-unanimous clips, or 18.6% of the corpus, underwent human inspection, while 6.35% of flagged clips required edits.Approximately 3,897 clips were edited, representing 1.2% of the full corpus; 268,620 clips were unanimously accepted.
  • A.5. Multi-Model Verification and Human Refinement: 100% of clips received model-based inspection, and 71,380 clips, or 21.6% of the corpus, received manual inspection.Manual review covered all disagreement cases plus 10,000 unanimous-case spot checks.
  • A.5. Multi-Model Verification and Human Refinement: 10 volunteers contributed approximately 1,200 person-hours, while multi-model verification cost approximately $2.8K in API fees.The reported API cost covered all 330,000 clips.
  • A.6. Human Annotation Procedures: Diversity-driven curation uses fine-grained multimodal metadata to support rigorous candidate verification, rejection, grading, and final task definition.The in-house annotation tool combines video playback with a dynamic grading panel and follows a “Verify-Reject-Grade” workflow.
  • B. Framework Details: The framework represents translation and rotation with a quadruple descriptor, using Difficulty, ID, Direction, and Keys alongside three-dimensional coordinates and axis rotations.This representation is designed to make the action space quantifiable and extensible.
  • B.2. Entire Action Space: 729 motions comprise all permutations of 27 translation actions and 27 rotation actions, with difficulty values ranging from 1 to 6.Stationary motion is assigned difficulty 1, and validity depends on action frequency in the collected dataset.
  • B.3. Existing Keyboard and Mouse Support Space: 81 motions form the existing keyboard-and-mouse support space from 9 operable translation and 9 operable rotation actions, with difficulty ranging from 1 to 4.The memory module designs inverses for six translation and rotation motion patterns to generate back-and-forth trajectories for memorability testing.

C. Benchmark Details · C.1. Data presentation of different difficulty

The Difference Verification task presents camera-operation examples across four difficulty levels. Difficulty increases from single-axis movements to complex multi-axis trajectories with view changes, testing detection of subtle pose differences.

  • C.1. Data presentation of different difficulty: The presented Realestate example uses a backward camera movement.
  • C.1. Data presentation of different difficulty: Figure 5 showcases the Difference Verification task across four difficulty levels using camera operation.The levels are organized by rows 1 to 4.
  • C.1. Data presentation of different difficulty: Row 1 contains basic single-axis movements.
  • C.1. Data presentation of different difficulty: Row 2 adds combined translation and rotation.
  • C.1. Data presentation of different difficulty: Row 3 presents sequential composite trajectories.
  • C.1. Data presentation of different difficulty: Row 4 involves complex multi-axis movements with view changes.
  • C.1. Data presentation of different difficulty: The examples demonstrate the model’s ability to detect subtle pose differences across various scenarios.

C.2. Data presentation of memory

The Memory Verification task presents loop-closure scenarios that test whether models can recall an initial state after reversible camera actions. Visual cues support temporal reasoning and long-term consistency checks.

  • Memory Verification: Memory Verification focuses on the difficulty of verifying loop closure.The showcase illustrates memory-dependent interaction scenarios.
  • Memory Verification: Reversible actions such as “up then down” require recalling the initial state.Examples also include turning right and then left.
  • Memory Verification: Red bounding boxes highlight visual cues for temporal reasoning and consistency checks in long-term memory scenarios.

D. Metrics details · E. Expriment details

The metrics suite evaluates generated videos across visual quality, temporal consistency, motion coherence, trajectory control, and memory-related symmetry. It combines frame-level perceptual measures, trajectory comparisons, and long-sequence consistency mechanisms.

  • D. Metrics details: Image Quality (SImage) uses MUSIQ frame scores, averaged across the sequence and linearly normalized, to measure distortions including overexposure, noise, and blur.The metric supports diverse resolutions and aspect ratios through the MUSIQ quality prediction model.
  • D. Metrics details: Brightness Consistency (SBrightness) compares three-level grayscale distributions with the initial frame using modified Softmax enhancement and normalized exponential temporal weighting.The method is designed to monitor brightness drift and style collapse while allowing reasonable brightness changes.
  • D. Metrics details: Color Temperature Constraint (SColor) represents HSV hue across 7 core intervals and applies distant-frame weighting to penalize color drift in long videos.The weighting condition β > α makes the improved weight decay faster and strengthens penalties for scene color-temperature inconsistency.
  • D. Metrics details: Sharpness Retention (SSharpness) compares horizontal and vertical edge-gradient vectors, using BRISQUE-triggered noise penalties to distinguish true details from high-frequency artifacts.Its cosine similarity and penalty logic target systematic visual collapse while retaining texture-direction information.
  • D. Metrics details: Motion Smoothness (SMotion) measures sequence coherence by reconstructing discarded odd frames with a video interpolation model and comparing them using LPIPS-based perceptual deviation.The metric follows a sampling-reconstruction paradigm for quantifying temporal smoothness.
  • D. Metrics details: Trajectory Accuracy (SAccuracy) aligns ViPE-extracted camera trajectories with commanded motion and measures tangent-direction agreement as instruction-level controllability.Trajectory Tolerance (STolerance) instead compares aligned execution against precise system-built ground-truth trajectories, reducing estimator uncertainty and providing an upper-bound reference.
  • D. Metrics details: Memory Symmetry (SMemory) evaluates pixel-wise consistency between symmetric frame pairs with distant-frame weighting, while Trajectory Alignment (SAlignment) measures spatial-topology symmetry in go-and-return camera motion.Both metrics target memory-related failures in long-sequence or cyclic actions, including logical loop closure and route consistency.

E.1. Human Preference Validation Details · E.2. Detailed comparison of world generation models

The benchmark’s automated metrics align strongly with human preferences, detect significant and practically meaningful model differences, and retain fine-grained discrimination among similar models. The paper also documents the inference configurations, variants, release versions, generation times, and camera-control support of evaluated world models.

  • E.1. Human Preference Validation Details: The human-preference study evaluates 14 world models using 16 standardized tasks sampled uniformly across four difficulty levels and 5-point Likert ratings for visual quality and camera-control precision.The analysis pipeline combines Spearman correlation, Kruskal-Wallis testing, effect size with 95% confidence intervals, and Dunn’s post-hoc comparisons.
  • E.1. Human Preference Validation Details: rs = 0.8053 (p = 0.0005 < α = 0.05) shows strong, statistically significant rank-order agreement between objective metrics and human preferences across 14 models.The validation deliberately includes similarly performing models to test metric sensitivity across the full performance spectrum.
  • E.1. Human Preference Validation Details: H = 1496.8994 (p < 0.001, Ntotal = 2,688) rejects identical human-preference distributions across the 14 models.Model identity accounts for approximately 55.5% of preference-score variance, with η2 = 0.5549 (k = 14, N = 2,688).
  • E.1. Human Preference Validation Details: 95% confidence intervals separate broad human-perceived performance tiers, including top-tier mean scores ≥3.48, mid-tier scores 2.49–2.81, and lower-tier scores ≤2.26.The non-overlapping intervals provide visual evidence that benchmark rankings are stable across tiers.
  • E.1. Human Preference Validation Details: 11 out of 15 close-model pairwise comparisons are significant at p < 0.001 under Bonferroni-corrected Dunn’s testing.The four non-significant comparisons are ASTRA–MotionCtrl, ASTRA–WAN 2.2, CameraCtrl–Matrix-game 2.0, and MotionCtrl–WAN 2.2.
  • E.1. Human Preference Validation Details: 1.93 positions is the mean absolute rank difference between objective-metric and human-preference orderings across 14 models.AC3D has the largest discrepancy, with H-Rank 9, O-Rank 4, and |∆Rank| = 5; its objective average is driven by Memory Symmetry = 0.9068 and Motion Smoothness = 0.9919.
  • E.2. Detailed comparison of world generation models: The detailed model comparison reports parameter configurations, model variants, inference time, average generation time per instance on NVIDIA A800 GPUs, official release versions, and explicit camera-pose input support.Version denotes each model’s official release date, while § marks support for explicit camera pose control.
Loading 2605.03941v2…