Source-linked AI summary

VGI-Bench: Probing Visual Intelligence in Video Generation Models

Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai

arXiv:2608.19583v3cs.CVcs.AI

TL;DR

Reliable evaluation of zero-shot visual reasoning in video generation remains difficult because benchmarks may mismatch models’ visual priors, omit evolving processes, or miscalibrate feasibility. VGI-BENCH addresses these gaps with realistic, process-sensitive, difficulty-calibrated tasks and evaluates current systems broadly. Results show emerging but unreliable reasoning: Seedance 2.0 achieves 51.0 under the benchmark criteria, while analyses identify bounded transfer, input sensitivity, failure modes, and limited denoising self-correction.

  • Problem

    Existing benchmarks may use visually mismatched inputs, fail to require visual rollouts, and include tasks that are too difficult or insufficiently calibrated for current video models.

  • Method

    VGI-BENCH uses photorealistic-style inputs, process-sensitive downstream tasks, calibrated difficulty, and a two-level taxonomy of domains and skill tags.

  • Results

    Current generative systems solve a subset of visually grounded reasoning tasks but remain far from reliable, with Seedance 2.0 achieving 51.0 under the evaluation criteria.

  • Takeaways & Limitations

    VGI-BENCH provides a diagnostic testbed showing emerging visual reasoning alongside sensitivity to inputs, bounded synthetic-transfer gains, and limited self-correction during denoising.

  • Takeaways & Limitations

    The benchmark is limited to roughly 5–10-second image-to-video tasks at fixed 16:9 aspect ratio, with English prompts and a representative rather than exhaustive task suite.

Abstract

from arXiv · show

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/

1 Introduction

VGI-BENCH addresses gaps in visual-reasoning evaluation by using realistic inputs, process-sensitive tasks, calibrated difficulty, and a two-level taxonomy. Broad evaluation finds emerging reasoning ability but substantial limitations, with Seedance 2.0 reaching only 51.0 under the criteria.

  • Motivation: Existing benchmarks may use abstract inputs that mismatch video models’ natural-image priors, weakening whether failures measure reasoning limitations.Controlled comparisons found more collapse and constraint ignorance for abstract inputs than realistic counterparts.
  • Benchmark design: The taxonomy combines mutually exclusive task domains with non-exclusive skill tags for fine-grained capability analysis.Representative panels show real-scene inputs and icons encoding skill tags.
  • Motivation: Existing visual reasoning tasks often omit explicit scene rollouts, while long-horizon and knowledge-heavy tasks can exceed practical feasibility.The benchmark therefore targets visually grounded procedures with calibrated difficulty near current capability boundaries.
  • Benchmark design: VGI-BENCH uses photorealistic-style inputs, requires valid intermediate trajectories, and filters tasks to remain challenging yet partially feasible.Tasks are organized by one mutually exclusive domain and one or more skill tags.
  • Findings: 51.0 is the strongest model score reported for Seedance 2.0, while current models still struggle with coherent multi-step execution.Common failures include physical collapse, rule violation, and object/state inconsistency; analyses also examine input sensitivity, synthetic-transfer limits, and denoising behavior.
  • Contributions: The benchmark contributes broad evaluation and multifaceted analyses of failure modes, input sensitivity, synthetic fine-tuning transfer, and denoising dynamics.These analyses identify factors that shape and limit visual reasoning performance.

2 Related Works

Video generation models have progressed from content creation and world simulation toward possible zero-shot visual reasoning through generated frame sequences. This motivates evaluating their ability to express solutions for multimodal reasoning tasks.

  • Background: Video generation models were initially evaluated mainly for visual quality, motion realism, and condition alignment.As temporal coherence and physical plausibility improved, they were increasingly viewed as visual world simulators.
  • Background: Generated videos can explicitly predict how scenes, objects, and interactions evolve over time.This world-simulation perspective treats temporal evolution as a model capability.
  • Visual reasoning: Recent studies suggest video models may act as zero-shot visual reasoners by expressing solutions through generated frame sequences.Follow-up benchmarks evaluate spatial, physical, logical, and other generative reasoning tasks.
  • Related benchmarks: Related efforts include TiVi-Bench, V-ReasonBench, MMGR, and VBVR, which extend evaluation or adaptation of generative visual reasoning.These works span multiple reasoning domains and, in VBVR, connect task collections with benchmark-specific adaptation.

3 VGI-BENCH

VGI-BENCH is a process-sensitive benchmark organized by domains and skill tags, with calibrated tasks and complementary evaluation metrics. It measures both progress toward a visual goal and validity of the evolving procedure, while also supporting image-output adaptations.

  • Taxonomy: The taxonomy defines four mutually exclusive domains: Visual Organization, Physical Manipulation, Structured Puzzles, and Spatiotemporal Dynamics.These domains distinguish arranging objects, physical actions, rule-governed transformations, and temporal state evolution.
  • Taxonomy: Non-exclusive skill tags provide a capability-level view complementary to domain organization and enable coarse- and fine-grained diagnosis.Seven tags are used, including Affordance.
  • Task design: Each task gives a text prompt and input image as the first frame, requiring a generated video to complete a visually grounded, process-sensitive procedure.Tasks emphasize reasoning-intensive objectives and avoid low-level recognition or localization.
  • Task design: Tasks use three difficulty levels and are selected near current model capability boundaries rather than as an exhaustive task suite.References specify acceptable solutions or desired final states and guide task proposals and evaluation criteria.
  • Difficulty calibration: Pre-generation testing accepts a task only when at least one tested model solves a sampled instance and at least one fails it.This filters out tasks that are trivially solvable or entirely infeasible for current models.
  • Evaluation: Completeness measures global progress, Rubric Score measures local process validity, and Final Score multiplies them so either failure penalizes the result.Rubric items use 1/(x + 1), while adaptive sampling inspects transient violations more closely.
  • Image adaptation: Image-output adaptations preserve task goals while changing the evaluation from temporal procedure expression to target-state inference and rendering.This extends selected VGI-BENCH tasks to image generative models.

4 Experiments

VGI-BENCH evaluates contemporary video and image generation models, finding that commercial video models lead but remain far from reliable, while evaluator ablations support the full VLM-based evaluation design.

  • Evaluation protocol: The evaluation samples half of each task’s instances using a fixed random seed, with stability against full-set evaluation analyzed separately.Gemini-3-Flash is used as the automatic evaluator.
  • Video-model results: Seedance-2.0 achieves the best overall video score at 51.0, while commercial models consistently outperform open-source models.Structured Puzzles is the most challenging domain, and Topology and Temporal are the weakest skill dimensions.
  • Image-model diagnostic: Image models are evaluated on an output-adapted subset that tests static goal-state realization without valid intermediate processes.Nano-Banana-Pro leads with Avg. 55.0, followed by Seedream-5.0-Pro with Avg. 52.1.
  • Evaluation reliability: The full VLM-based evaluator agrees more strongly with human judgments than either ablation removing adaptive frame sampling or the sliding focus window.Table 3 reports AUC and pairwise accuracy against human preferences.

5 Discussion

The discussion diagnoses failures in physical coherence, rule following, state tracking, input sensitivity, synthetic-transfer boundaries, and denoising-based reasoning. Across these analyses, later generation steps rarely correct early reasoning errors, and improvements remain structurally bounded.

  • 5.1 Failure Modes: Generated videos exhibit physical collapse, rule violations, and object/state inconsistency during goal-directed process generation.Failures include unrealistic deformation, object penetration, prohibited actions, skipped steps, and unstable identities or intermediate states.
  • 5.2 Input Condition Sensitivity: Oracle prompts still leave models struggling with fine-grained spatiotemporal transitions, instruction following, and physical simulation under strong rule constraints.The performance gain is particularly weak for HunyuanVideo 1.5.
  • 5.3 Synthetic Fine-tuning Transfer: Synthetic fine-tuning transfers more effectively across structurally aligned tasks, but overlap does not guarantee improvement and some non-overlap tasks still benefit.Task-level gains are uneven, with UNTIE_KNOT showing little overall success-rate change despite a substantial rubric-score rise.
  • 5.3 Synthetic Fine-tuning Transfer: Improvements concentrate on planning and spatial reasoning, whereas physical interaction and strong temporal dependency remain difficult and may degrade.The transfer boundary follows the structural coverage of the synthetic training distribution.
  • 5.4 When Is Visual Reasoning Decided?: The denoising analysis decodes intermediate states and labels transitions as stable, changed, correct→wrong, wrong→correct, or wrong→wrong′.This protocol examines whether later denoising steps revise incorrect intermediate states.
  • 5.4 When Is Visual Reasoning Decided?: Self-correction remains below 1% throughout denoising, while wrong-to-wrong′ transitions reach 23.1% at 4 →10 and 24.8% at 10 →20.Once readable, states become increasingly stable, reaching 90.6% stability in the second half; later steps mainly lock in and refine early hypotheses.

6 Conclusion

VGI-BENCH evaluates visual intelligence in video generation models using realistic-style inputs, process-sensitive tasks, and calibrated difficulty levels. Results show emerging reasoning ability but substantial limitations, motivating systematic diagnosis and future model development.

  • VGI-BENCH provides a benchmark for evaluating visual intelligence in video generation models.
  • Realistic-style inputs, process-sensitive task design, and calibrated difficulty levels test whether models can solve tasks through valid visual rollouts.
  • Current models exhibit emerging reasoning ability but remain far from general visual intelligence.
  • The benchmark’s analyses cover input sensitivity, bounded synthetic-fine-tuning transfer, and limited self-correction during denoising.
  • VGI-BENCH is intended to support systematic diagnosis and development of future video and multimodal foundation models.

Limitations

The benchmark has explicit scope boundaries around generation length, conditioning format, language, and task coverage.

  • Tasks target current models’ roughly 5–10s generation length, excluding longer-horizon procedural reasoning.Examples outside scope include multi-minute assembly and long-trajectory planning.
  • The benchmark covers only image-to-video generation with a fixed 16:9 aspect ratio.Text-to-video, multi-image conditioning, and audio-conditioned generation remain future work.
  • English prompts and rubrics leave multilingual and cross-lingual evaluation unaddressed.
  • The task suite represents a focused, non-exhaustive slice of visual reasoning domains and difficulty levels.It may need extension as model capabilities improve.

A.1 Task Material Collection

VGI-BENCH constructs task materials from matched web or dataset images and generated images, then refines them for visual correctness and standardized evaluation. Prompts specify goals, constraints, actions, and required state changes, while representative tasks demand valid evolving procedures.

  • Input Image Construction: Input images come from web images, existing visual datasets, or image-generation models.Generated inputs use direct textual scene descriptions or construction pipelines when appropriate.
  • Input Image Construction: Human reviewers check generated images for object correctness, attributes, layout, and artifacts before accepting or regenerating them.All input images are standardized to a 16:9 aspect ratio.
  • Text Prompt Construction: Text prompts state task goals and specify object rules, visual attributes, allowed and prohibited actions, required state changes, and generation controls.Background, spatial layout, and camera-motion instructions reduce irrelevant variation.
  • Representative Tasks: MAZE requires continuous corridor-following motion from a start cell to a red goal without clipping, jumping, teleportation, or sudden cuts.
  • Representative Tasks: RECOVER_2D_NET requires rigid faces to rotate continuously about hinge edges into a closed 3D solid while preserving geometry and adjacency.
  • Image Output Subset: An auxiliary image-output subset contains 16 tasks whose goals can be meaningfully represented by a single image, excluding inherently continuous temporal interactions.

B.1 Model Generation Configuration

The primary evaluation uses a broad set of video generation models under a unified text-and-image input format, with auxiliary single-image evaluation and detailed metric reporting. Difficulty trends and strict success rates require cautious interpretation because current models remain unstable and end-to-end success is rare.

  • Video Evaluation (Primary): The primary evaluation covers nine representative video generation models using a unified text prompt and input image format.For each task, half of its instances are evaluated due to cost.
  • Image Evaluation (Auxiliary): The auxiliary evaluation adapts benchmark tasks for single-image generation and unified multimodal models at 1280 × 720 resolution.
  • Access and Budget: Commercial models are accessed through APIs, while open-source models run locally on NVIDIA H20 GPUs.
  • Per-metric Full Results: Final is the mean of per-instance Completeness × Rubric products, so it generally differs from Comp. × Rub. computed from separate means.The difference reflects instance-level alignment between completeness and rubric performance.
  • Difficulty-level Fluctuations: Difficulty scores need not increase monotonically within every model–domain cell because layout, object configuration, and action-pattern changes interact with model priors.Difficulty trends are interpreted mainly at the aggregate level.
  • Strict Success Rate: Strict success rates are sharply lower than aggregated scores because success requires perfect completeness and no rule violations.Human ceiling performance is included for reference.

B.3 Human Ceiling Evaluation

The human ceiling evaluation tests whether VGI-BENCH tasks are solvable and compares human performance with model outputs. It uses task-adapted response formats and strict binary success judgments, while supplementary analyses document human failures and evaluation stability.

  • Human Study Protocol: Human participants completed randomly sampled instances across task difficulty levels under response-time limits calibrated from pilot studies.The study targets task solvability and clarity rather than response speed.
  • Human Study Protocol: Human performance is measured by strict success rate, because checklist violations common in generated videos rarely occur for participants.Each instance receives a binary judgment of successful task completion or substantial achievement of the intended objective.
  • Human Failure Cases: Human failures remain non-trivial at the highest difficulty level, including miscounted moves, missed unique cells, and geometric tiling mis-fits.These examples indicate that the benchmark includes challenging instances even for human participants.
  • Stability Analysis: The main evaluation uses half of the instances at each difficulty level, while a representative full-instance comparison tracks the half-subset within 5 percentage points on average and preserves model rankings.The comparison covers 5 video models and 6 tasks spanning four domains.

C.3 Human Correlation

The evaluator is validated against human rankings of generated videos using filtered pairwise preferences. Agreement remains stable across domains and skills, although sampled-frame evaluation is weaker for fine-grained physical and topological changes.

  • Human Preference Protocol: Human annotators rank triplets of model videos by task progress, rule violations, and overall correctness, producing more than 1,100 filtered preference pairs.Contradictory preferences and cycles are removed before correlation analysis.
  • Correlation Metrics: Evaluator quality is assessed through pairwise agreement and ROC-AUC against strict human preferences.Score differences below 0.05 are treated as ties for pairwise agreement.
  • Correlation Results: Agreement stays within roughly 0.73–0.85 AUC and 69%–80% pairwise accuracy across every domain and skill tag.No single taxonomy slice drives the aggregate reliability.
  • Correlation Results: Structured Puzzles show the strongest human agreement, whereas Physical Manipulation, Topology, and Physics are weakest.The weaker slices involve subtle interactions or connectivity changes distributed across frames, which sampled-frame evaluation can miss.
  • Interpretive Context: The benchmark’s process-sensitive design makes rule violations and intermediate-state validity central to judging visual reasoning.Examples include physical collapse, prohibited actions, incomplete Eulerian paths, and simultaneous disk moves in Tower of Hanoi.

D.3 Scaling Fine-tuning on Synthetic Data

Synthetic fine-tuning transfers most effectively to tasks structurally aligned with its training distribution, but task-level gains are uneven. Denoising analysis further indicates that models rarely correct wrong reasoning after early structural decisions are formed.

  • Scaling Fine-tuning on Synthetic Data: Average fine-tuning gains decrease from overlap to semi-overlap and then non-overlap tasks.Structurally aligned training examples better support following task rules, reaching goals, and avoiding rubric violations.
  • Scaling Fine-tuning on Synthetic Data: Task-level transfer is non-uniform: some overlap tasks degrade, while some non-overlap tasks improve.The paper attributes these outcomes to domain gaps, execution requirements, controlled generation, and transferable spatial or logical capabilities.
  • Reasoning along Denoising Trajectory: Self-correction is evaluated by comparing consecutive decoded intermediate videos across selected denoising-step pairs.The protocol labels transitions between readable solution states rather than only endpoint correctness.
  • Reasoning along Denoising Trajectory: The denoising study analyzes 702 annotated video pairs from 117 task instances across six non-uniform step pairs.Four open-source models are evaluated under default 40-step generation settings.
  • Reasoning along Denoising Trajectory: Wrong → correct transitions remain at or below 0.9% and stop after step 4, while wrong → wrong′ changes persist into middle denoising.The answer is largely settled early, with later steps mainly refining the initial solution state.

E.1 Detailed Comparison with Related Works

The comparison evaluates related video benchmarks by input appearance, process sensitivity, and difficulty control. VGI-BENCH emphasizes realistic inputs, explicit state transitions, and feasibility-calibrated task difficulty.

  • Comparison criteria: Table 20 compares video benchmarks across input appearance, process sensitivity, and difficulty control.Each pie shows the fraction of tasks satisfying or not satisfying a desideratum; Real. indicates photorealistic-style inputs and Abs. non-photorealistic inputs.
  • Input appearance: Reasoning-heavy benchmarks vary substantially in their use of realistic inputs, with some relying on line-art, schematic, or script-generated visuals.PhysGenBench and WorldSimBench are fully realistic, whereas VBVR-Bench is entirely script-generated for controllability and scalability.
  • Process sensitivity: VGI-BENCH designs tasks around explicit state transitions and evaluates intermediate states and transition validity.This targets process-sensitive reasoning rather than answers that can be produced directly from a static input.
  • Difficulty control: VGI-BENCH combines pre-generation and manual feasibility review with explicit multi-level difficulty construction.In contrast, several reasoning-heavy benchmarks lack model-facing feasibility calibration, and only TiVi-Bench among them provides multi-level difficulty design.
  • Use and licensing: The benchmark artifact is intended for research evaluation and diagnostic analysis rather than unrestricted redistribution of third-party outputs.Commercial APIs and open-source models are used under applicable terms, while proprietary outputs and third-party weights are not redistributed beyond permission.

F.5 Human Annotation Recruitment, Payment, and Data Consent

The human annotations were produced by the paper’s co-authors for benchmark validation, without external recruitment or separate compensation. The study reports anonymization and aggregate reporting, but no formal ethics review approval or exemption.

  • Recruitment: The paper’s co-authors provided the human-ceiling and preference annotations as part of the research process.No external participants were recruited through crowdsourcing platforms or student pools.
  • Payment: No separate compensation was provided for the annotations.
  • Consent and data handling: Annotators were informed that their responses and preference annotations would support benchmark validation, including human-ceiling estimation and VLM-as-Judge validation.The annotations were reported in aggregate and did not include names, contact information, demographic attributes, or other sensitive personal information.
  • Ethics review: The study did not obtain formal ethics review board approval or exemption.The paper states that no external participants were recruited and that annotations were used only for benchmark validation.
Loading 2608.19583v3…