Source-linked AI summary
The Trinity of Consistency as a Defining Principle for General World Models
Jingxuan Wei, Siyuan Li, Yuhang Xu, Zheng Sun, Junjie Jiang, Hexuan Jin, Caijun Jia, Honghao He, Xinglong Xu, Xi bai, Chang Yu, Yumou Liu, Junnan Zhu, Xuanhe Zhou, Jintao Chen, Xiaobin Hu, Shancheng Pang, Bihui Yu, Ran He, Zhen Lei, Stan Z. Li, Conghui He, Shuicheng Yan, Cheng Tan
TL;DR
The paper addresses the lack of a principled framework for defining general world models that understand and simulate physical reality. It proposes the Trinity of Consistency, reviews the move toward unified architectures, and introduces CoW-Bench; results show strong static priors but persistent consistency failures, especially over long horizons.
Problem
Existing systems approximate physical dynamics, but the field lacks a principled framework defining the essential properties of general world models.
Method
The paper proposes the Trinity of Consistency and reviews specialized-to-unified model evolution, while introducing CoW-Bench for multi-frame consistency evaluation.
Results
CoW-Bench shows closed-source image generation models lead overall, while open-source video generators lag on consistency-sensitive tasks and temporal rule-grounded evolution remains uneven.
Takeaways & Limitations
World-model quality depends on cross-dimensional consistency and constraint satisfaction over time, not visual plausibility or smoothness alone.
Takeaways & Limitations
Existing benchmarks have structural defects because MLLM judges have low precision for fine-grained physical attributes and can reward visually smooth violations.
Abstract
from arXiv · showhide
The construction of World Models capable of learning, simulating, and reasoning about objective physical laws constitutes a foundational challenge in the pursuit of Artificial General Intelligence. Recent advancements represented by video generation models like Sora have demonstrated the potential of data-driven scaling laws to approximate physical dynamics, while the emerging Unified Multimodal Model (UMM) offers a promising architectural paradigm for integrating perception, language, and reasoning. Despite these advances, the field still lacks a principled theoretical framework that defines the essential properties requisite for a General World Model. In this paper, we propose that a World Model must be grounded in the Trinity of Consistency: Modal Consistency as the semantic interface, Spatial Consistency as the geometric basis, and Temporal Consistency as the causal engine. Through this tripartite lens, we systematically review the evolution of multimodal learning, revealing a trajectory from loosely coupled specialized modules toward unified architectures that enable the synergistic emergence of internal world simulators. To complement this conceptual framework, we introduce CoW-Bench, a benchmark centered on multi-frame reasoning and generation scenarios. CoW-Bench evaluates both video generation models and UMMs under a unified evaluation protocol. Our work establishes a principled pathway toward general world models, clarifying both the limitations of current systems and the architectural requirements for future progress.
1 Introduction
The paper frames general world models as internal simulators of physical reality and proposes the Trinity of Consistency as their foundation. It reviews the path from specialized modules toward unified architectures and introduces CoW-Bench for multi-frame evaluation.
- AGI requires agents that learn physical laws, reason about counterfactuals, and predict future states from actions.
- The Trinity of Consistency comprises Modal Consistency, Spatial Consistency, and Temporal Consistency as complementary requirements for robust world models.They respectively address semantic alignment, geometric representation, and physically coherent temporal evolution.
- Modal Consistency aligns heterogeneous inputs in a shared semantic space, Spatial Consistency preserves 3D geometry and object permanence, and Temporal Consistency enforces causal evolution over time.
- The review traces generative models from loosely coupled specialized modules toward end-to-end unified architectures whose dimensions synergize in world simulation.
- CoW-Bench evaluates whether models maintain the Trinity of Consistency in multi-frame reasoning and open-ended generation scenarios.
2.1 The Anatomy of General World Models
The paper presents world-model construction as the integration of modal, spatial, and temporal consistencies. It reviews how specialized models developed these dimensions separately before their prerequisites enabled unified simulators.
- World models integrate modal, spatial, and temporal consistency as an information interface, geometric cornerstone, and dynamic engine.
- Specialized models developed modality alignment, spatial geometry, and temporal dynamics through distinct mechanism shifts.The review covers high-dimensional manifold mapping, explicit 3D primitives, and the progression from frame interpolation to causal dynamics modeling.
2.2 Modal Consistency
Modal consistency addresses the challenge of aligning heterogeneous modalities into a unified, physically complete representation. Its evolution spans geometric alignment, generative manifold design, iterative inference, and increasingly integrated architectures, while persistent modality gaps and error accumulation limit current systems.
- Theoretical Foundations: Modal consistency treats multimodal learning as high-dimensional heterogeneous manifold alignment toward a unified, physically aligned representation space.The target must overcome entropy disparity and topological mismatch across modalities.
- Theoretical Foundations: The Platonic representation hypothesis models images and text as projections of an objective latent physical state, making alignment a joint inverse-projection problem.Visual projections retain high-frequency physical entropy, whereas textual projections abstract discrete symbolic logic.
- Theoretical Foundations: Hypersphere-based alignment can fail because entropy disparity collapses embeddings into narrow cones, producing topological mismatch between visual and textual representations.The cone effect undermines the uniform-distribution assumption of hyperspherical alignment.
- Architectural Evolution: Modal-consistency research progresses from dual-tower geometric isolation through connector and discrete unified interfaces toward integrated architectures that reduce modality-specific gradient conflict.The reviewed paradigm includes Dual-Tower, connector-based, discrete unified, and orthogonally decoupled approaches.
- Generative Manifolds: Continuous Flow Matching avoids quantization by modeling a deterministic transport path in continuous latent space, changing error accumulation from exponential to linear growth.The reported dynamics are ∥δT∥≈T · ϵstep for Flow Matching, compared with ∥δT∥≈LT∥ϵ0∥ for autoregressive generation when L > 1.
- Architectural Evolution: Orthogonal modality decoupling reduces gradient conflict from over 50% in autoregressive paradigms to approximately 30%, with reported gains over LLaVA in typography and long-text tasks.The comparison is reported for Stable Diffusion 3.5 Large under complex typography rendering and long-text comprehension.
- Open Challenges: Current amortized, one-pass mappings remain prone to logical hallucinations on counterfactual tasks requiring multi-step deduction and real-time verification.The limitation is attributed to fitting training-distribution patterns through amortized inference.
2.3 Spatial Consistency
Spatial consistency provides the geometric basis for a world model by progressing from 2D proxies toward explicit 3D representations and enforcing local continuity and global multi-view coherence. The section frames rendering, generation, and motion through physical laws while identifying limitations of 2D dynamics modeling.
- Geometric basis: Spatial consistency requires geometric representations that preserve object structure across viewpoints and support navigation and interaction.The framework emphasizes 3D awareness, geometry, occlusion, object permanence, and multi-view coherence.
- Representation evolution: The reviewed paradigm evolves from 2D proxy manifolds through implicit 3D fields toward explicit Lagrangian primitives.This trajectory is presented as a progression in representations for solving spatial consistency.
- Topological constraints: Local topological consistency constrains microscopic surface continuity, while global consistency enforces unique and coherent object geometry across views.Local continuity is associated with Lipschitz regularity; global coherence is associated with epipolar equivariance and multi-view constraints.
- Physical formulation: NeRF preserves continuous volume-rendered fields, whereas 3DGS uses explicit Gaussian primitives for efficient analytical rasterization and real-time performance.Both are described as discretized solutions to the radiative transfer equation, differing in representation and rendering route.
- Limitations: 2D proxy models fail under occlusion and large rotations because their spatial continuity assumptions and convolutional inductive biases do not capture strict 3D consistency.The cited limitations include non-differentiable optical flow after depth mutations, inability to model object permanence, and missing SE(3) equivariance.
2.4 Temporal Consistency
Temporal consistency treats world models as systems that preserve identity, obey physical constraints, and maintain causal evolution over time. The section reviews frequency-aware evaluation and the progression from temporal inflation to causal 3D tokenizers and autoregressive modeling, alongside their limitations.
- Temporal dynamics: Temporal consistency requires stable subject identity, smooth motion trajectories, and causally ordered events across frames.The framework distinguishes physical constraints that suppress flicker from causal constraints such as object permanence.
- Evaluation: FVD mainly measures spatial feature-distribution similarity and can miss high-frequency temporal flicker and non-physical deformations.This motivates evaluation methods that explicitly model temporal frequency behavior.
- Evaluation: VCD measures generated-versus-natural video feature differences in the temporal frequency spectrum, where flicker appears as high-frequency energy fluctuations.The method uses feature extraction and a short-time Fourier transform along the time axis.
- Evaluation: Veo 3 achieves a task success rate exceeding 70% on zero-shot physical interaction tasks such as predicting domino toppling.The evaluation combines temporal artifact measures, Physics-IQ, and causal reasoning tasks.
- Model evolution: Temporal inflation extends frozen 2D image models with inserted temporal modules, but its 2D anchor causes severe texture stretching under large viewpoint changes or newly appearing content.The paper characterizes this paradigm as a transitional solution rather than true temporal dynamics modeling.
- Model evolution: Discrete autoregressive models use causal 3D tokenizers, while long-sequence generation remains constrained by exposure bias and exponentially amplifying prediction errors.Next-scale prediction, context packing, and anti-drift sampling are described as responses to sequence variance and memory decay.
2.5 Outlook of the Consistencies
The three consistencies have emerged as distinct computational capabilities, but a general world model requires their integration into a shared cognitive substrate. The outlook therefore shifts from optimizing isolated modules toward architectural unification across modalities, space, and time.
- Evolution of consistencies: Modal, spatial, and temporal consistency have respectively advanced semantic translation, explicit 3D representation, and causal world simulation.The section presents these as three computational engines developed through specialized-model evolution.
- Architectural outlook: Independent optimization of specialized modules cannot form a coherent world simulator without a shared cognitive substrate.The stated bottleneck is architectural rather than merely component-level.
3.1 The Rise of Large Multimodal Models
Large Multimodal Models unify heterogeneous inputs by mapping them into an LLM-centered representation and increasingly coordinate specialized modules through planning and tool use. Their architecture combines semantic alignment with programmatic reasoning and closed-loop verification.
- Representation bridging: LMMs map heterogeneous modality data into the LLM’s word-embedding space through visual connectors or adapters.This mapping provides semantic alignment and conversion across modalities.
- Representation bridging: Visual encoders and projection layers convert image feature maps into visual tokens that the LLM can process alongside text tokens.The visual tokens are projected into the language model’s dimensional space and participate in self-attention.
- Representation bridging: Q-Former and Perceiver Resampler architectures compress massive pixel features into a fixed number of learned queries, reducing sequence redundancy.The paper interprets this as semantic pooling and an information bottleneck favoring language-relevant features.
- System coordination: LMMs increasingly use the LLM as a kernel that schedules resources, coordinates logic, and invokes specialized modules on demand.This extends multimodal systems beyond static end-to-end mappings.
- System coordination: Visual programmatic reasoning decomposes ambiguous visual instructions into executable sub-tasks and combines low-level operators for logical self-consistency.VisProg and ViperGPT exemplify recursive task decomposition through executable code flows.
- System coordination: ReAct-based systems perform closed-loop verification by pausing generation to call external expert models that check and correct outputs.The mechanism targets physical hallucination through test-time computation.
3.2 Integration of Modal and Spatial Consistency
Modal and spatial consistency are integrated to align language with geometric structure, enabling precise spatial control while preserving object layouts and physical common sense. Research progresses from implicit pixel-space mappings toward explicit geometric conditions and structured editing controls.
- Modal-spatial integration enables precise responses to spatial instructions involving occlusion, surrounding, and perspective.
- Four technical paths explore modal-spatial integration: implicit emergence, explicit synergy, structured isomorphism, and reinforcement learning.
- View-space mapping introduces camera poses and depth maps as geometric conditions, allowing semantic and geometric flows to coexist through cross-attention.
- Instruction-Driven Image Editing: Instruction-driven editing combines gradient decoupling with attention injection to change semantics while preserving spatial structure.
- General Image Generation: Joint image-text modeling captures pixel-level spatial constraints implicit in descriptions such as “a cat on a table.”
- 3D-consistent noise initialization and temporal attention enable geometrically controllable generation and multi-view sequences with continuity.
3.3 Integration of Modal and Temporal Consistency
Modal and temporal consistency connect semantic instructions to coherent time evolution, moving generative systems from static images toward probabilistic simulation of spatiotemporal causality. End-to-end video architectures pursue this through scalable modeling, causal compression, native 3D attention, and progressive alignment.
- Modal-temporal integration makes video content adhere to text or image instructions while modeling spatiotemporal causality.
- End-to-End Scalable Modeling: End-to-end scalable modeling fits the joint distribution p(ximg|c) from multimodal inputs to video outputs by expanding model and data scales.
- End-to-End Scalable Modeling: Flow Matching replaces Gaussian-noise prediction with a deterministic ODE trajectory between prior noise and data distributions.
- End-to-End Scalable Modeling: Rectified Flow reduces transport curvature by forcing latent variables along linear trajectories, supporting dynamic textures with physical conservation in few steps.
- Causal Spatiotemporal Compression: Causal 3D VAEs prevent future-information leakage through asymmetric temporal padding and provide a foundation for streaming inference.
- Native 3D Attention Modeling: Native 3D DiT models joint spatiotemporal attention, correcting the decoupling of space and time in earlier factorized architectures.
3.4 Integration of Spatial and Temporal Consistency
Spatial and temporal consistency jointly require objects to preserve geometry, identity, and physically coherent motion across time and occlusion. The field advances from implicit video-prior learning toward explicit geometric anchoring, trajectory constraints, and manifold alignment.
- Dynamic object permanence requires objects to preserve geometric form, physical motion trajectories, and intrinsic properties during spatiotemporal evolution.
- Four evolutionary stages span implicit spatiotemporal learning, explicit geometric anchoring, and later transitions toward physical realism in 4D world models.
- Implicit Spatiotemporal Learning: Video-prior distillation offers a training-free alternative to full-parameter fine-tuning by modulating posterior scores from heterogeneous priors.
- Implicit Spatiotemporal Learning: Multi-view and video diffusion priors are combined with time-dependent weights, while systematic solutions address gradient conflicts and distribution mismatches.
- Explicit Geometric Anchoring: Mapping camera extrinsics Pt ∈ SE(3) to video timestamps forces geometric parallax to be interpreted as optical-flow motion.
- Explicit Geometric Anchoring: Acceleration regularization minimizes second-order position differences, suppressing high-frequency jitter and supporting smooth motion.
3.5 Preliminary Emergence of World Models
Unified multimodal architectures increasingly combine semantic interfaces, geometric structure, and causal evolution into internal world simulators. This emerging direction extends from video generation toward interactive physical environments and embodied decision-making.
- Modal consistency provides interaction interfaces, spatial consistency constructs geometric structure, and temporal consistency supplies causal evolution.
- Video World Simulators: Sora maintains perspective constancy during camera movements and exhibits collisions and occlusions that adhere to temporal causality.
- Video World Simulators: Open-Sora verifies Video DiT principles through a transparent open-source testbed for exploring the three consistencies.
- Interactive World Simulators: Interactive world models introduce action operators, extending trinity consistency from passive video projection to active simulation.
- Interactive World Simulators: Genie discretizes action spaces, while Matrix-Game 2.0 models multi-agent game-theoretic logic for causal arbitration in simulated environments.
- Embodied AI: Explicit state reconstruction enforces rigid-body dynamics through st+1 = fphysics(st, at), truncating or penalizing trajectories that violate physical laws.
- The paper concludes that trinity synergy supports General World Simulators with internalized physical laws and counterfactual reasoning capabilities.
4 Challenges, Benchmarks, and Outlook
Current world models remain limited by gaps in physical authenticity, long-range causal stability, controllability, and evaluation validity. The paper frames CoW-Bench as a benchmark designed to test modality, spatial, and temporal consistency through multi-frame reasoning, long-range evolution, and causal intervention.
- Open Challenges: Current models optimize pixel- or token-level likelihood, producing visually plausible videos that can violate mechanics such as support, momentum conservation, and material behavior.The paper identifies physical authenticity as a differentiable gap and calls for embedding physical structures such as Hamiltonians and conservation laws.
- Open Challenges: Long-term causal chains remain brittle because short-range attention cannot prevent identity and event-logic failures from accumulating over hour- to day-scale simulations.A proposed direction is hierarchical implicit dynamics spanning symbolic narratives or scene graphs, sparse 4D event representations, and high-dimensional local modeling.
- Open Challenges: Future world models must support active intervention, allowing users to modify forces, materials, and boundary conditions while receiving physically consistent feedback through an end-to-end action-to-observation loop.The paper also extends the target from physical simulation toward agentic and digital ecosystems involving social causality and operating-system state transitions.
- Benchmark Gaps: Existing benchmarks disconnect scores from physical capability because visual judges can reward smooth outputs that violate physical laws, while in-distribution datasets mask out-of-distribution failures.The paper specifically criticizes model-as-judge evaluation and memorization of real-world or standard game recordings.
- Benchmark Gaps: Short-sequence evaluations conceal state drift: without online process verification or physical-constraint correction, small errors can amplify over time and collapse long simulations.Unlike frame-by-frame physics engines, generative models primarily rely on probabilistic sampling and lack deep process-verifiability probes.
- Benchmark Outlook: CoW-Bench addresses these gaps with hardcore physical standards, dynamic long-range evolution, causal intervention, and separate evaluation of modality, spatial, temporal, and pairwise consistency.Its broader motivation is to replace perceptual metrics such as FID and FVD when evaluating world models as physical simulators.
5 CoW-Bench
CoW-Bench is a balanced, multi-frame benchmark designed to evaluate world-model consistency across modal, spatial, temporal, and cross-consistency tasks. Its results show that current models often achieve local or static plausibility but struggle with long-horizon constraint enforcement, persistent semantics, and instruction fidelity.
- Benchmark Design: CoW-Bench contains 1,485 samples across 18 balanced fine-grained sub-tasks spanning modal, spatial, temporal, and cross-consistency regimes.Each sub-task contains 69–91 samples, including 50 hard Maze cases.
- Benchmark Design: The benchmark spans instruction complexity from 7.1-word atomic tasks to 74.8-word compositional tasks, creating a broad semantic difficulty gradient.
- Benchmark Design: CoW-Bench evaluates generative simulation rather than passive perception, using dynamic constraint satisfaction instead of static question-answer accuracy.
- Main Results: Closed-source image generation models lead overall, while open-source video generators lag on most consistency-sensitive tasks.GPT-image-1.5 achieves the best overall performance, followed by Nano Banana Pro and GPT-image-1.
- Main Results: Temporal control is harder than visual continuity: models can generate smooth footage while violating rule-grounded evolution or causal constraints.Sora reaches 9.32 on T-WL, whereas T-Rule and T-Stage-Order vary sharply across models.
- Main Results: Cross-consistency tasks expose failures in globally anchored spatial structure and persistent world-state maintenance despite strong local or per-frame plausibility.Nano Banana Pro reports 4.46 on TS-Maze-2D, while leading models approach ceiling performance on some property-persistence tasks.
- Main Results: Overall, UMMs perform well in static or single-view settings but drop sharply when stable world states must persist across changes in time and space.
- Modal Consistency: Modal consistency remains limited by unstable identity-and-attribute binding and widespread constraint backoff.GPT-image-1.5 scores 1.19 on Id+Attr, Nano Banana Pro 1.25, and HunyuanVideo 0.01; closed-source image models typically show Backoff around 1.6–1.8.
5.6 Cross-Axis Consistency
Cross-axis evaluation shows that current models often preserve local visual plausibility while failing to bind instructions, geometry, and temporal evolution into a coherent world state. CoW-Bench operationalizes these failures through multi-frame, constraint-focused evaluation.
- Semantic role binding is a dominant bottleneck, with models often attaching instructed actions or relations to the wrong entity despite coherent geometry.
- Avoiding forbidden configurations is generally easier than constructing exact required spatial relations, as Neg-rel and Excl. outperform Pos-rel across many mid-tier models.
- Models may preserve scene stability while failing to execute instructed temporal direction, scheduling, or pacing, producing smooth sequences that violate semantic commitments.
- Triggered events reveal timing and persistence failures: Trigger and Post-comp degrade even when Pre-hold is relatively strong.
- Maze-2D sharply discriminates models because plausible maze motion can still fail to maintain a single trajectory from the correct start to the correct goal.
- CoW-Bench uses multi-frame reasoning and constraint satisfaction to distinguish visual generation from physical simulation.
6 Conclusion
The conclusion frames consistency across modality, space, and time as the defining criterion for a World Model. It introduces CoW-Bench as a unified diagnostic and identifies interaction representation as a structural limitation of current systems.
- The Trinity of Consistency defines World Models through cross-dimensional robustness rather than single-axis visual performance.
- World-model paradigms progress from opaque Vector-as-Action and predefined Key-as-Action spaces toward semantically expressive Prompt-as-Action interaction.
- CoW-Bench unifies video generation models and UMMs under a multi-frame protocol that treats consistency as strict constraint satisfaction.
- The analysis identifies constraint backoff as a structural artifact of interaction representations, not merely a consequence of insufficient data or scale.
- A system that produces compelling pixels without cross-dimensional consistency remains a texture synthesizer rather than a world simulator.
7 Contributions
The listed contributors span multiple institutions, including Shanghai Artificial Intelligence Laboratory, the University of Chinese Academy of Sciences, and several universities.
- The paper lists Jingxuan Wei, Siyuan Li, Cheng Tan, Yuhang Xu, Zheng Sun, Junjie Jiang, Hexuan Jin, and Caijun Jia among its contributors.
- Additional contributors include Honghao He, Xinglong Xu, Xi Bai, Chang Yu, Yumou Liu, Junnan Zhu, Xuanhe Zhou, and Jintao Chen.
- The author list also includes Xiaobin Hu, Shancheng Pang, Bihui Yu, Ran He, Zhen Lei, Stan Z. Li, Conghui He, and Shuicheng Yan.
- Affiliations include Shanghai Artificial Intelligence Laboratory and the University of Chinese Academy of Sciences.
- Other listed affiliations are National University of Singapore, Shanghai Jiaotong University, and China University of Petroleum (East China).