Source-linked AI summary

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark

Ziyu Guo, Xinyan Chen, Renrui Zhang, Ruichuan An, Yu Qi, Dongzhi Jiang, Xiangtai Li, Manyuan Zhang, Hongsheng Li, Pheng-Ann Heng

arXiv:2510.26802v1cs.CVcs.AIcs.CL

TL;DR

The paper asks whether high-fidelity video models can reason zero-shot in challenging visual tasks. It evaluates Veo-3 across 12 reasoning dimensions using the MME-COF Chain-of-Frame benchmark and finds strong local abilities but unreliable complex reasoning, motivating complementary use with dedicated reasoning models.

  • Problem

    It remains unclear whether coherent video generation reflects robust zero-shot reasoning or surface-level pattern learning.

  • Method

    The study evaluates Veo-3 across 12 reasoning dimensions and curates MME-COF to standardize Chain-of-Frame assessment.

  • Results

    Video models show promising short-horizon spatial coherence, fine-grained grounding, and local dynamics, but remain limited in long-horizon causal reasoning, geometric constraints, and abstract logic.

  • Takeaways & Limitations

    Video models are not yet reliable standalone zero-shot reasoners but show potential as complementary visual engines alongside dedicated reasoning models.

  • Takeaways & Limitations

    The models struggle with causal and physical logic, long-horizon rule adherence, and consistent constraint-aware geometric understanding.

Abstract

from arXiv · show

Recent video generation models can produce high-fidelity, temporally coherent videos, indicating that they may encode substantial world knowledge. Beyond realistic synthesis, they also exhibit emerging behaviors indicative of visual perception, modeling, and manipulation. Yet, an important question still remains: Are video models ready to serve as zero-shot reasoners in challenging visual reasoning scenarios? In this work, we conduct an empirical study to comprehensively investigate this question, focusing on the leading and popular Veo-3. We evaluate its reasoning behavior across 12 dimensions, including spatial, geometric, physical, temporal, and embodied logic, systematically characterizing both its strengths and failure modes. To standardize this study, we curate the evaluation data into MME-CoF, a compact benchmark that enables in-depth and thorough assessment of Chain-of-Frame (CoF) reasoning. Our findings reveal that while current video models demonstrate promising reasoning patterns on short-horizon spatial coherence, fine-grained grounding, and locally consistent dynamics, they remain limited in long-horizon causal reasoning, strict geometric constraints, and abstract logic. Overall, they are not yet reliable as standalone zero-shot reasoners, but exhibit encouraging signs as complementary visual engines alongside dedicated reasoning models. Project page: https://video-cof.github.io

1 Introduction

The study asks whether video models can reason zero-shot through Chain-of-Frame generation rather than merely produce coherent content. It introduces a standardized 12-dimension evaluation and finds promising local reasoning but unreliable complex reasoning.

  • Study scope: The study investigates whether video models can serve as zero-shot visual reasoners across 12 spatial, geometric, physical, temporal, and embodied dimensions.The analysis focuses on Veo-3 while comparing several state-of-the-art models through the MME-COF benchmark.
  • Findings: Video models show promising short-horizon spatial coherence, fine-grained grounding, and locally consistent dynamics, but struggle with long-horizon causal consistency, geometry, and abstract logic.The study distinguishes common success patterns from failure patterns to assess when behavior reflects reasoning versus pattern replay.
  • Conclusion: Current video models are not yet reliable as standalone zero-shot reasoners, despite encouraging signs of emergent reasoning.The authors position them as potentially useful alongside specialized reasoning systems rather than as independent reasoners.
  • Benchmark: MME-COF standardizes Chain-of-Frame evaluation with a compact taxonomy and protocol for category-wise assessment beyond surface-level visual fidelity.The benchmark is designed to quantify reasoning potential and enable directly comparable scores and qualitative behaviors across categories.

2 Deep-Dive Analysis on Veo-3

The deep-dive evaluates video reasoning through a structured taxonomy, expert-curated cases, standardized prompts, and zero-shot generations. Veo-3 can ground salient visual details but remains vulnerable to instruction divergence, artifacts, and weak grounding for difficult targets.

  • Task taxonomy: The study organizes reasoning into 12 categories, including visual detail, visual trace, spatial, physics-based, table-and-chart, and object-counting reasoning.Each category contains representative cases targeting specific reasoning abilities.
  • Data and prompts: Five experts curate representative cases and construct explicit prompts, while a unified style controls visual constraints, motion, framing, and linguistic ambiguity.The design aims to make output differences reflect reasoning potential rather than prompt variability.
  • Evaluation setup: Each case generates six 1280×720, 24 FPS, 8-second samples in a unified zero-shot setting without fine-tuning, supervision, or auxiliary tools.Outputs are judged as Good, Moderate, or Bad using correctness, clarity, and temporal stability.
  • Evaluation criteria: Moderate and bad outcomes include blur, incomplete framing, instability, unintended motion, ambiguous targets, severe artifacts, and extraneous objects.These criteria limit confident interpretation when the target or attribute cannot be reliably inferred.
  • Visual detail reasoning: Veo-3 performs well on fine-grained attributes and spatial relations for salient, well-grounded targets, but fails when objects are small, occluded, or cluttered.It can also produce plausible yet instruction-divergent outcomes reflecting stylistic generation preferences.

2.3 Visual Trace Reasoning

Veo-3 can produce locally coherent reasoning traces in simple settings, but its visual reasoning degrades with long horizons, complex spatial transformations, and strict object or rule grounding.

  • Visual Trace Reasoning: The spatial evaluations assess temporal coherence and goal consistency, while using qualitative Good, Moderate, and Bad ratings for reasoning outputs.
  • Visual Trace Reasoning: Figure 4 highlights short-horizon pathfollowing successes, object-grounding failures, and step omissions or mistakes in multi-step traces.
  • Visual Trace Reasoning: Figure 5 shows long-horizon planning breakdowns, inconsistent arrow or trajectory rendering, and failures to preserve comparative or sequential information across frames.
  • Visual Trace Reasoning: Veo-3 can produce coherent step-by-step paths in simple, low-branching settings, but this behavior is not robust across trials.
  • Visual Trace Reasoning: The model systematically fails at long-horizon planning, rule-grounded execution, and object-persistent manipulations.
  • Real-World Spatial Reasoning: Veo-3 handles basic spatial layouts but struggles with complex viewpoints, orientation changes, depth, and precise perspective transformations.

2.5 3D Geometry Reasoning

Veo-3 shows potential on basic 3D transformations, but its geometric reasoning becomes fragile when transformations are complex, multi-step, or require consistent relations among multiple objects.

  • 3D Geometry Reasoning: The evaluation measures geometric accuracy, structural completeness throughout transformations, and visual continuity across frames.
  • 3D Geometry Reasoning: Veo-3 performs reasonably on simple, single-step geometric transformations but degrades noticeably on multi-step or compositionally complex transformations.
  • 3D Geometry Reasoning: The model frequently produces misaligned or self-intersecting structures, leading to loss of geometric consistency.
  • 3D Geometry Reasoning: Veo-3 can partially understand individual object shapes but lacks coherent coordinate systems and spatial relationships among multiple objects.

2.6 2D Geometry Reasoning

Veo-3 shows basic success on simple 2D geometric constructions but struggles to preserve geometric constraints, drawing order, and structural consistency in complex tasks.

  • The evaluation covers planar construction and movement tasks, including connecting points, adding lines, and moving shapes while preserving described geometric relations.
  • Geometric movements can roughly follow instructions but may introduce jitter, misalignment, disconnected shapes, or distorted paths.
  • Complex connection tasks often produce incorrect sequences, altered figures, and drawings that continue beyond the required construction.
  • Veo-3 demonstrates foundational ability on simple geometric connection tasks but remains far from a robust geometric reasoner.
  • A moderate success rate of 83% is reported for one dot-connection example.

2.7 Physics-based Reasoning

Veo-3 can render locally plausible dynamics and simple reflections, but it does not reliably preserve quantitative physical constraints or causal fidelity in complex interactions.

  • The task evaluates gravity, collisions, reflection, momentum, and energy conservation through physically plausible and temporally coherent motion.
  • Moderate outputs remain visually plausible despite irregular acceleration, timing mismatch, or slight conservation violations.
  • Reflection and general trajectory shape can be rendered plausibly, although angular and timing offsets remain common.
  • Frictional, force-driven, and mechanically constrained cases frequently violate quantitative physical constraints and causal fidelity.
  • Veo-3 produces locally plausible short-term dynamics but fails on energy, momentum, causal ordering, and contact mechanics in constrained scenarios.

2.9 Table and Chart Reasoning

Veo-3 can focus approximately on relevant chart regions and handle basic counting, but imprecise localization, altered visual elements, and scene instability limit reliable analysis.

  • The chart and table task evaluates relevant-region identification, smooth transitions, visual continuity, clarity, and proper scaling.
  • Charts are often zoomed into approximately correct regions, while tables may receive random selections and altered or distorted elements.
  • Veo-3 falls short of functioning as a precise and reliable chart-table reasoner.
  • A valid counting output requires accurate labels or boxes with controlled motion and a largely static scene.
  • Veo-3 demonstrates basic counting but lacks the spatial control and robustness needed for reliable enumeration in dynamic or complex scenes.

2.11 GUI Reasoning

Veo-3 shows limited GUI interaction ability, but inaccurate clicks, inconsistent effects, and altered interface elements reveal shallow functional understanding.

  • The task evaluates click accuracy and temporal coherence while requiring the scene and irrelevant interface elements to remain consistent.
  • A bad GUI output has imprecise or erratic clicks and substantially altered original data or icons.
  • Across GUI cases, click positions and resulting on-screen effects are often inconsistent, while icons and text may be altered or newly generated.
  • The model shows partial GUI responsiveness and some visual feedback in one Web-system case.
  • Veo-3 imitates GUI click behaviors without fully grasping the underlying functional logic.

2.12 Embodied Reasoning

The embodied-reasoning evaluation tests whether video models can recognize affordances and maintain faithful, stable visual sequences during manipulation tasks. Veo-3 can generate plausible affordances for clearly defined objects, but lacks reliable planning and stability for dynamic or spatially constrained interactions.

  • 2.12 Embodied Reasoning: The evaluation assesses static and dynamic affordances, manipulation-relevant attributes, visual stability, and reasoning fidelity without implausible shortcuts.
  • 2.12 Embodied Reasoning: Good outputs fairly frame manipulation-relevant geometry with stable scale and no cropping, while bad outputs bias or distort evidence enough to prevent fair decisions.
  • 2.12 Embodied Reasoning: Robobench samples support evaluation of both static attributes and direct reasoning about generated static and dynamic affordances.
  • 2.12 Embodied Reasoning: Veo-3 can generate plausible manipulation affordances for a clearly defined object, but dynamic affordances expose workarounds and insufficient stability.
  • 2.12 Embodied Reasoning: Veo-3 remains limited to basic object recognition rather than reliable embodied reasoning in real-world interactions.The model lacks the planning and stability needed to interpret and act on dynamic or spatially constrained instructions.

2.13 Medical Reasoning

The medical-reasoning evaluation tests localization, anatomical attributes, pathological patterns, and binary decisions while requiring stable, undistorted views. Veo-3 can manipulate medical images, but struggles with terminology, organ modeling, and detail-preserving zooms.

  • 2.13 Medical Reasoning: The evaluation covers lesion or structure localization, anatomical attributes, pathological patterns, binary decisions, manipulation correctness, and surrounding-region stability.
  • 2.13 Medical Reasoning: Good outputs settle on the correct anatomical level or lesion without distortion, whereas bad outputs miss targets, hide cues, or confuse medical terminology.
  • 2.13 Medical Reasoning: Medical evaluation data are sampled from ViTAR and include prompts for anatomically controlled views such as chest, lung, and lumbar imaging.
  • 2.13 Medical Reasoning: Veo-3 retains some ability to manipulate medical images, but its lack of medical knowledge limits the quality of the resulting reasoning videos.
  • 2.13 Medical Reasoning: Veo-3 struggles to manipulate correct objects when prompts use medical terminology, and medical-image zooms introduce substantial distortion and detail loss.The reported failures occur across the medical cases and reflect difficulty modeling medical organs precisely.

3 MME-COF

MME-COF is a benchmark for systematically measuring zero-shot reasoning in generative video models across diverse categories and standardized dimensions. Results show generally limited reasoning, with visual stability strongest, model-specific specialization, and mean scores below 2.0 out of 4.

  • 3 MME-COF: MME-COF is introduced as a benchmark designed specifically to reveal and quantify video models’ reasoning potential.
  • 3 MME-COF: MME-COF contains 59 curated entries and instruction prompts spanning 12 reasoning categories.The benchmark’s composition is summarized in Table 1, Figure 2b, and Figure 19.
  • 3 MME-COF: The benchmark evaluates leading video models in a zero-shot setting, using six generated samples per prompt and mean scores across samples.
  • 3 MME-COF: Gemini-2.5-Pro scores each generated video from 0 to 4 on instruction alignment, temporal consistency, visual stability, content fidelity, and focus relevance.
  • 3 MME-COF: Most models show limited reasoning across MME-COF, while Visual Stability has the highest average and outputs remain largely pattern replay rather than genuine reasoning.
  • 3 MME-COF: Mean scores remain below 2.0 out of 4 across evaluated models, despite relative specialization in different reasoning aspects.Sora-2 performs relatively better in physics-based, embodied, and medical reasoning; Veo-3 in real-world spatial reasoning; and Seedance in rotation and 3D geometry.

4 Related Work

Related work spans video representation and generation, video reasoning with specialized multimodal models, and zero-shot evaluation of video generation models. Prior studies motivate examining whether generative video models can perform reasoning beyond content synthesis.

  • 4 Related Work: Video understanding methods learn representations for downstream tasks, while newer approaches encode videos as tokens for captioning, event localization, and high-level reasoning.
  • 4 Related Work: Video-reasoning research increasingly uses large reasoning models and specialized multimodal systems for temporal and spatio-temporal reasoning.
  • 4 Related Work: Recent zero-shot studies evaluate video generation models in general-purpose vision, medical imaging, and world-model settings, including tasks such as segmentation and image editing.

5 Conclusions and Insights

Video models show intuitive, local reasoning in simple visual worlds but remain unreliable on complex, long-horizon reasoning. Their emerging Chain-of-Frame abilities nevertheless support use as complementary visual engines alongside dedicated reasoning models.

  • Video models exhibit intuitive yet local reasoning, including short-horizon coherence, realistic dynamics, and coherent embodied manipulation paths.These strengths are supported by qualitative and quantitative results, but remain foundational rather than comprehensive reasoning abilities.
  • They fail on causal and physical logic, long-horizon rule adherence, complex geometry, and functional interaction planning.Observed failures include implausible motion, broken causal order, state loss, violated geometric constraints, and imitation without functional understanding.
  • Current video models are not reliable standalone zero-shot reasoners because plausible outputs often follow surface patterns rather than general principles.They prioritize visual plausibility or symmetry over precise spatial and geometric instructions, producing instructionally flawed results.
  • Despite these limitations, their emergent visual reasoning abilities suggest a role as complementary visual engines alongside dedicated reasoning models.Carefully designed prompts may guide these models through visual problems step by step within collaborative reasoning systems.
Loading 2510.26802v1…