Source-linked AI summary

RISE-Video: Can Video Generators Decode Implicit World Rules?

Mingxin Liu, Shuran Ma, Shibei Meng, Xiangyu Zhao, Zicheng Zhang, Shaofeng Zhang, Zhihang Zhong, Peixian Chen, Haoyu Cao, Xing Sun, Haodong Duan, Xue Yang

arXiv:2602.05986v1cs.CVcs.AI

TL;DR

TI2V models have achieved strong perceptual quality, but their ability to reason over implicit world rules remains insufficiently evaluated. RISE-Video introduces a human-annotated, reasoning-centric benchmark with multidimensional metrics and an automated LMM judging pipeline. Evaluation of 11 representative models reveals persistent difficulties with higher-level and implicit reasoning despite strong perceptual quality.

  • Problem

    Existing TI2V evaluation largely emphasizes perceptual quality and temporal coherence, leaving implicit reasoning over world rules insufficiently assessed.

  • Method

    RISE-Video provides 467 human-annotated samples across eight reasoning domains, evaluates four dimensions, and adds an automated LMM-based judging pipeline aligned with human judgments.

  • Results

    11 representative TI2V models show systematic reasoning limitations, with all evaluated models achieving relatively low reasoning accuracy and Hailuo 2.3 reaching 22.5%.

  • Takeaways & Limitations

    The benchmark supports holistic, scalable assessment of TI2V reasoning beyond perceptual fidelity and exposes persistent challenges in higher-level implicit reasoning.

Abstract

from arXiv · show

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a pioneering reasoning-oriented benchmark for Text-Image-to-Video (TI2V) synthesis that shifts the evaluative focus from surface-level aesthetics to deep cognitive reasoning. RISE-Video comprises 467 meticulously human-annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi-dimensional evaluation protocol consisting of four metrics: \textit{Reasoning Alignment}, \textit{Temporal Consistency}, \textit{Physical Rationality}, and \textit{Visual Quality}. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human-centric assessment. Extensive experiments on 11 state-of-the-art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world-simulating generative models.

1 Introduction

RISE-Video addresses the lack of reasoning-focused evaluation for TI2V models by benchmarking implicit world-rule understanding across diverse dimensions. It combines a human-annotated benchmark, four evaluation metrics, and scalable LMM-based assessment to examine current systems.

  • Existing TI2V evaluations emphasize visual fidelity and temporal coherence, leaving implicit reasoning over world rules comparatively under-evaluated.
  • RISE-Video contains 467 human-annotated samples organized across eight reasoning domains covering low-level perception through high-level logical inference.
  • The benchmark evaluates generated videos with Reasoning Alignment, Temporal Consistency, Physical Rationality, and Visual Quality.
  • An automated LMM judging pipeline uses manually designed reasoning-aware questions and prompts to support scalable assessment aligned with human judgments.
  • The study evaluates 11 representative TI2V models and reports systematic limitations in their ability to execute implicit world rules.

2 Related Work

Video-generation research has progressed from basic synthesis methods toward architectures that improve temporal coherence, resolution, and text alignment. Benchmarking has likewise moved from coarse perceptual measures toward structured and semantically grounded evaluation.

  • Diffusion-based video generation evolved from adding temporal modules to image generators toward architectures designed for longer-range temporal coherence.
  • Lumiere generates entire clips in one pass using a space–time U-Net, while CogVideoX scales diffusion transformers with a 3D VAE.
  • Early video benchmarks measured overall realism with frame- or video-level metrics but failed to capture motion coherence.
  • VBench introduced a unified evaluation framework with 8 data categories and 16 dimensions to assess diverse video-generation capabilities.

3 Method

RISE-Video organizes 467 human-annotated samples across eight reasoning dimensions and evaluates TI2V models using four complementary metrics. Its evaluation combines targeted LMM judging with specialized strategies for structured visual reasoning tasks.

  • Data Construction: RISE-Video partitions reasoning into eight dimensions covering experiential, commonsense, temporal, societal, perceptual, spatial, subject-specific, and logical knowledge.The taxonomy spans low-level perceptual cues and high-level abstract inferences.
  • Data Construction: Logical capability tests rule-following in game actions, puzzle solving, and geometric reasoning under explicit structural constraints.Examples include mazes, board games, word-linking puzzles, and symmetry-like geometric tasks.
  • Data Construction: The benchmark evaluates experiential, commonsense, subject-specific, perceptual, and societal reasoning through knowledge-grounded video-generation scenarios.These dimensions include intentions and procedures, everyday physics and health, discipline-specific knowledge, visual attributes, and social or cultural contexts.
  • Data Construction: Temporal and spatial knowledge assess reasoning across multiple time spans, viewpoint transformations, object arrangements, and three-dimensional manipulation.Temporal categories range from short-term events to longer-duration changes, while spatial evaluation follows specified camera trajectories and object relationships.
  • Data Construction: 467 human-annotated samples provide diverse and representative coverage of reasoning scenarios.Each sample is curated and annotated by human experts.
  • Evaluation Metrics: Four complementary metrics assess model performance beyond perceptual fidelity, with Reasoning Alignment using manually designed, knowledge-aware questions tailored to each reasoning type.The evaluation pipeline also includes Temporal Consistency, Visual Quality, and Physical Rationality, with dimension-specific frame extraction and LMM judging.
  • Evaluation Metrics: Schematic Puzzles require specialized evaluation: trajectory constraints verify maze navigation, grid alignment measures symmetry, and reference images support board-game comparison.These strategies address rigid geometry and the difficulty of expressing ground-truth states linguistically.

4 Experiments

Experiments on 11 TI2V models show that reasoning remains difficult despite progress in visual generation. Results compare reasoning categories, dynamic instruction following, and LMM-based evaluation against human judgments.

  • Experimental Setup: 11 representative TI2V models are evaluated across four metrics and eight reasoning categories.The study includes both closed-source and open-source systems.
  • Main Results: 22.5% accuracy is achieved by Hailuo 2.3, the best-performing model, followed by Veo 3.1 at 22.3% and Sora 2 at 21.3%.Hailuo 2.3 also exceeds Wan 2.6 in Reasoning Alignment by 6.6%.
  • Reasoning Categories: Perceptual Knowledge is relatively strong, whereas Logical Capability remains consistently low across models.Logical tasks require combining perceptual evidence with abstract, rule-based reasoning.
  • Dynamic Behavior: Kling 2.6 often produces minimal motion and fails to perform required commonsense transformations such as camouflage and capillary action.Wan 2.6 and Hailuo 2.3 show stronger instruction following and more dynamic generation behavior.
  • Temporal and Visual Quality: Veo 3.1 and Sora 2 exhibit temporal discontinuities, while Veo 3.1 also fails to adapt a chameleon’s color sufficiently to its surroundings.These abrupt frame changes disrupt temporal smoothness and overall video quality.
  • Automated Evaluation: GPT-5 shows the most robust alignment with human preference across most metrics, while Qwen3-VL-235B has lower Temporal Consistency error but a high-score bias.The bias is accurate for perfect samples but weakens discrimination between high-quality outputs and severely defective ones.

5 Conclusion

RISE-Video evaluates whether TI2V models generate videos that satisfy diverse reasoning requirements rather than only perceptual criteria. Across 11 models, strong visual quality coexists with difficulty in higher-level implicit reasoning.

  • Benchmark and Evaluation: RISE-Video organizes evaluation into eight reasoning categories and four dimensions for holistic assessment beyond perceptual fidelity.An automated LMM-based judging pipeline supports scalable evaluation while aligning closely with human judgments.
  • Conclusion: 11 representative TI2V models retain strong perceptual quality but struggle with higher-level and implicit reasoning.The findings underscore a gap between visual realism and rule-consistent reasoning.

A.1 Data Source

The benchmark’s input images come from generated images, permissively licensed websites, and manually curated RISEBench samples. These sources support visual fidelity, diversity, and transfer to TI2V reasoning tasks.

  • Image Sources: High-quality image-generation models provide images selected for visual fidelity and diversity in downstream TI2V tasks.
  • Image Sources: Permissively licensed website images are collected according to their respective usage terms.
  • Image Sources: Suitable images are manually curated from RISEBench and adapted for TI2V reasoning tasks.

A.1.1 Privacy-Preserving Image Stylization

The dataset uses image stylization for tasks involving real individuals when preserving original appearance is unnecessary. This reduces identifiable details while retaining information needed for reasoning evaluation.

  • Privacy Preservation: Image stylization removes identifiable visual details from selected real-person images while preserving structural and semantic information required for evaluation.The processing is intended to mitigate privacy concerns without invalidating the reasoning tasks.

A.2 Prompt for Judgement

The judging prompts specify structured evaluation procedures for reasoning alignment, temporal consistency, physical rationality, and visual quality. They use rubric-based scoring and strict JSON outputs, with reference-image comparison for targeted alignment judgments.

  • Reasoning Alignment: Reasoning Alignment compares the generated video’s final frame with a reference image for the attribute or region specified by a focused question.The judge assigns a 1–5 similarity score based only on visible agreement in the specified aspect.
  • Temporal Consistency: Temporal Consistency scores whether instruction-required changes occur while unrelated object attributes, identities, backgrounds, and layouts remain stable.The rubric ranges from perfect consistency to severe continuity breaks, using a 1–5 scale.
  • Output Format: All evaluation prompts require machine-readable JSON responses containing scores or answers and concise reasons.The formats prohibit markdown wrappers and specify the required output fields.
  • Physical Rationality: Physical Rationality evaluates motion, contact, fluidity, object permanence, material transitions, and physical artifacts on a 1–5 scale.Higher scores indicate seamless physical behavior, while lower scores reflect increasingly severe continuity or physics violations.

A.3 Analysis on the Judge Models

This section presents qualitative judge-model comparisons for temporal consistency, physical rationality, and visual quality, alongside example evaluation instructions and a perfect physical-rationality judgment.

  • Evaluation instructions: The examples include continuation instructions for generating the most natural or reasonable next action after an initial scene.These instructions frame temporal generation as a plausible continuation of the captured moment.
  • Physical Rationality: 5 is assigned to a cube-stacking example because colors, shapes, identities, and the instructed stacking sequence remain consistent.The judgment specifically describes the absence of unintended visual or spatial inconsistencies.

A.4 More Vsualizations

The visualizations show TI2V examples across the benchmark’s reasoning categories, including perceptual, commonsense, temporal, experiential, logical, societal, spatial, and subject knowledge.

  • Perceptual Knowledge: Perceptual Knowledge examples depict a vehicle continuing forward until it becomes fully visible.The example illustrates a visibility change during motion.
  • Commonsense Knowledge: Commonsense Knowledge examples ask models to generate long-term changes in toe joints under high uric acid conditions.The prompt begins from a normal person’s toe bone.
  • Temporal Knowledge: Temporal Knowledge examples require showing time reversing over five seconds.The visualizations include outputs from HunyuanVideo-1.5-720P-I2V variants.
  • Experiential Knowledge: Experiential Knowledge examples show the process of hands taking a letter out of an object.This visualization represents a procedural action sequence.
  • Logical, Societal, Spatial, and Subject Knowledge: Logical, societal, spatial, and subject examples respectively depict correct water-container filling, Thanksgiving food placement, ordered cookie shapes, and a chemical reaction.The prompts specify the intended outcome or arrangement for each scenario.
Loading 2602.05986v1…