Source-linked AI summary

Trimming the Long-Tail of Visual World Modeling Evaluation

Bingxuan Li, Yining Hong, Cheng Qian, Hyeonjeong Ha, Jiateng Liu, Zhenhailong Wang, Yue Guo, Yunzhu Li, Heng Ji

arXiv:2606.24256v1cs.CV

TL;DR

Current visual world models are rarely tested on irregular physical interactions, leaving their generalization of physical principles unclear. TailOR benchmarks regular, unconventional, and impossible tool-use scenarios, revealing performance degradation as interactions depart from familiar patterns.

  • Problem

    It remains unclear whether realistic visual world models internalize physical principles or rely on statistical regularities from familiar tool–task pairings.

  • Method

    TailOR evaluates image and video world models across progressively challenging Regular, Unconventional, and Impossible tool-use scenarios.

  • Results

    Performance consistently declines from Regular to Unconventional and Impossible scenarios, with video models scoring lower than image models across evaluation metrics.

  • Takeaways & Limitations

    Current world models rely heavily on memorized interaction templates rather than compositional physical reasoning beyond common tool–task co-occurrences.

  • Takeaways & Limitations

    TailOR’s tool-use scope leaves richer embodied settings, multi-step manipulation, and complex environment dynamics for future benchmarks.

Abstract

from arXiv · show

Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irregular interactions remains underrepresented. Although recent visual world models, including image and video generation models, achieve impressive realism on existing benchmarks, they primarily focus on simulating common physical interactions. This raises a central question: Do current visual world models internalize and generalize physical principles? In this work, we introduce Tailor-Bench, a benchmark that challenges world models to simulate irregular physical interactions. To enable systematic evaluation, we design three scenario modes that progressively challenge model reasoning: Regular scenarios reflect common tool-task pairs, Unconventional scenarios replace conventional tools with attribute-compatible substitutes to test affordance generalization, and Impossible scenarios introduce attribute-violating tools to probe constraint awareness. Additionally, we design two complementary settings under a unified evaluation protocol: predictive generation requires inferring outcomes without guidance, while descriptive generation specifies the target outcome for faithful realization. Our experimental results reveal a clear long-tail gap in physical world modeling: performance degrades from Regular to Unconventional and Impossible scenarios, indicating limited generalization beyond common interactions. Failure analysis further shows that models rely on superficial visual patterns: image models fail to realize correct state changes, while video models further suffer from temporal inconsistencies.

1. Introduction

The paper argues that current visual world-model evaluations emphasize common physical interactions, leaving their ability to generalize physical principles under irregular, long-tail interactions unresolved. It introduces TailOR to evaluate this gap through progressively unfamiliar scenarios and reports consistent degradation from Regular to Unconventional and Impossible interactions.

  • Motivation: Physical interactions follow a long-tailed distribution in which familiar head scenarios dominate human experience and large-scale visual data.Examples include cutting bread with a knife, shattering glass with a ball under gravity, and hammering a nail with a hammer.
  • Motivation: Existing evaluations largely test standard world assumptions, including familiar object behavior, intended tool use, and expected physical outcomes.These evaluations primarily cover head scenarios rather than the broader range of irregular interactions.
  • Benchmark: TailOR is introduced as a benchmark for evaluating world models under irregular physical interactions through scenarios that progressively depart from familiar interactions.The framework begins with Regular scenarios that mirror common tool–task pairs and includes multiple evaluation axes.
  • Results: The benchmark evaluates state-of-the-art image and video models and reveals a pronounced long-tail performance gap across both generation modalities.Performance consistently degrades from Regular to Unconventional and Impossible interactions.

2. Related Work

Recent multimodal generative models have improved image and video synthesis quality, temporal consistency, controllability, physical realism, and audio-visual coherence. Related work also examines these models as implicit world simulators and develops benchmarks for compositionality, object correctness, prompt alignment, and faithfulness.

  • World Modeling with Multimodal Generative Models: Large-scale image models demonstrate high-fidelity image generation aligned with prompts, while video models improve temporal consistency and controllability.Examples include Qwen-Image, Gemini Image, GPT-Image-1, Wan, MovieGen, Kling, and Pika.
  • World Modeling with Multimodal Generative Models: Veo 3, Sora 2, and Seedance extend generation to longer sequences with improved physical realism and audio-visual coherence.These systems represent more recent progress in multimodal generation.
  • World Modeling with Multimodal Generative Models: Several studies investigate whether large video generation models learn internal representations resembling world simulation.This work frames generative models as potential implicit world simulators beyond visual synthesis.
  • Evaluation of Multimodal Generative Models: Evaluation benchmarks assess text-to-image and text-to-video generation through compositionality, object-level correctness, prompt alignment, and interpretable faithfulness.T2I-CompBench, GenEval, and GenAI-Bench target the first three properties, while TIFA uses question-answering-based evaluation for faithfulness.

3. Problem Formulation

The problem formulation tests whether visually realistic world models internalize physical principles or rely on statistical regularities by evaluating tool-use interactions across regular, unconventional, and impossible scenarios. It further separates implicit outcome prediction from faithful realization of specified outcomes for image and video models.

  • Scenario Formulation: Each evaluation instance pairs a task goal and canonical tool with required functional attributes, attribute-compatible unconventional substitutes, and physically incompatible tools.This formulation supports testing whether substitute-object success follows functional attributes rather than memorized tool–task associations.
  • Scenario Modes: Regular scenarios use frequent canonical tool–task pairings, unconventional scenarios use atypical tools with matching attributes, and impossible scenarios use tools violating critical attributes.The three modes progressively probe in-distribution reproduction, attribute-level generalization, and respect for physical constraints.
  • Evaluated Model Types: Image models must produce spatially coherent final states, whereas video models must additionally represent temporal dynamics such as motion trajectories and contact events.The two modalities therefore test different levels of physical modeling complexity.
  • Evaluation Settings: Predictive generation requires inferring whether an interaction succeeds or fails and what state change occurs without revealing the final result.It evaluates implicit physical reasoning based on internal representations of properties such as rigidity, sharpness, friction, and structural compatibility.
  • Evaluation Settings: Descriptive generation provides the desired outcome in advance and tests whether models faithfully realize it while maintaining physical consistency and plausible scene dynamics.This setting emphasizes controllability and adherence to stated physical constraints.

4. Benchmark Design and Curation

TailOR is a structured benchmark for long-tail physical interactions, built by generating and verifying regular, unconventional, and impossible tool-use scenarios. It evaluates image and video generation through predictive and descriptive prompts using checklist-based measures of instruction adherence, interaction accuracy, and physical realism.

  • Evaluation preparation: The curation pipeline adds predictive and descriptive prompts, checklist-based rubrics, and a final human quality-control pass to each benchmark item.Rubrics separately assess instruction adherence and interaction accuracy using task specifications and prompts.
  • Data resources: TailOR curates scenarios from a common action ontology and an object–affordance graph, combining semantic task descriptions with tool-property constraints.The ontology is initialized from HICO-DET and manually filtered to 18 actions; the graph is built from ConceptNet properties.
  • Scenario construction: Each task generates one regular tool, two attribute-compatible unconventional substitutes, and two attribute-violating impossible tools through staged LLM generation and affordance inversion.Human annotators verify physical plausibility, impossibility, ambiguity, and visual clarity before retaining instances.
  • Dataset scale: 80 tool-use tasks span diverse action categories and indoor and outdoor environments, yielding 400 evaluation tasks and 1,600 prompt-based evaluation instances.Each task is evaluated for image and video generation under predictive and descriptive prompt types.

5. Evaluation

The evaluation spans leading image and video generation models across Regular, Unconventional, and Impossible scenarios, revealing a systematic long-tail performance gap. Video models underperform image models, while automatic and human evaluations largely agree on model rankings.

  • Models evaluated: The benchmark evaluates four image models and four frontier video models across its scenario settings.Image models include Z-Image, Qwen-Image, GPT-Image-1, and Nano-Banana-2; video models include HunyuanVideo-1.5, Wan-2.2, Sora-2, and Veo-3.1.
  • Long-tail scenarios: Performance consistently declines from Regular to Unconventional and Impossible scenarios across all four evaluation dimensions.The results indicate a systematic long-tail gap: models perform better on common head-distribution interactions than on attribute-level generalization.
  • Video-model difficulty: Video models score lower than image models across every evaluation metric, reflecting the added challenge of temporally coherent state transitions.Video models must depict plausible interactions while maintaining coherent changes across frames.
  • Video-model difficulty: Video models often revert to conventional training-distribution behaviors instead of executing intended dynamics, while early-frame errors amplify into cascading causal failures.These failures break temporal consistency, causal consistency, and physical plausibility.
  • Evaluation consistency: Automatic and human evaluations show strong agreement in model rankings across scenarios and metrics, with both identifying Sora-2 as the best video model.Most models ranked highly automatically also rank among the best in human assessment.

6. Discussion

TailOR reveals distinct failure patterns across scenario types and generation settings. Models often rely on familiar visual interaction templates rather than physically grounded reasoning, with video generation additionally exposing temporal and causal inconsistencies.

  • Image generation: Image-model failures shift from incorrect outcomes in Regular scenarios to affordance misgeneralization in Unconventional and instruction-adherence failures in Impossible scenarios.Regular failures also involve inaccurate attributes and physical violations; Unconventional failures include physical violations, while Impossible failures include incorrect outcomes and violations.
  • Video generation: Video-model failures shift from implausible dynamics in Regular scenarios to affordance misgeneralization in Unconventional and temporal inconsistency in Impossible scenarios.Interaction misexecution follows the leading failure in Regular and Unconventional scenarios, while Impossible scenarios also show physical violations and interaction misexecution.
  • Generation-setting comparison: Descriptive prompts generally improve image-model entity completeness and scene validity, whereas video models benefit less from descriptive guidance.Across most models, predictive performance is comparable to or higher than descriptive performance, especially for video models; descriptive performance often falls further behind in Unconventional scenarios.
  • Predictive generation: Predictive generation failures indicate insufficient physical knowledge, as models infer interactions through visual patterns rather than attribute-level physical abstractions.Regular scenarios can be handled by familiar interaction templates, whereas Unconventional and Impossible scenarios require compositional and counterfactual reasoning about properties such as rigidity, sharpness, leverage, and force direction.
  • Predictive generation: Video generation amplifies physical-reasoning limitations because multi-frame sequences expose violations of force propagation, state continuity, and physical constraints.These violations compound across frames during temporal generation.
  • Descriptive generation: Descriptive generation still fails because models often default to high-frequency object-function pairings instead of following unconventional outcomes and physically simulating the requested interaction.In video generation, forward propagation preserves an object’s familiar role across subsequent frames, hindering intended causal dynamics even when the outcome is explicitly specified.

7. Conclusion and Future Works

TailOR exposes a clear long-tail gap in visual world models: performance degrades on irregular interactions, with video models additionally challenged by temporal coherence and valid state transitions. The analysis attributes failures to weak physical inference and reliance on familiar visual patterns, motivating stronger physical inductive biases and improved temporal modeling.

  • Conclusion: TailOR evaluates whether visual generative models internalize physical principles rather than mainly relying on statistical regularities from training data.Experiments reveal a clear long-tail gap across both image and video generation.
  • Conclusion: Video models are more brittle than image models because they must maintain temporally coherent dynamics and physically valid state transitions over time.This adds temporal requirements beyond generating a visually plausible outcome.
  • Failure Analysis: Predictive generation fails when models cannot infer outcomes from transferable properties, while descriptive generation often defaults to familiar visual patterns despite explicit outcome instructions.The relevant physical properties include rigidity, geometry, and force transmission.
  • Future Work: Future work should add inductive biases for object attributes, affordances, and causal dynamics so models learn reusable physical primitives rather than holistic templates.The benchmark is intended as a testbed for physically grounded world modeling.
  • Future Work: Video generation could improve through long-horizon state tracking, force-consistent motion modeling, and constraint-aware temporal planning.These mechanisms target the temporal and physical consistency problems identified in the conclusion.

A. Benchmark Details … A.3. Human Verification

The benchmark’s data curation pipeline uses GPT-5 to generate task scenarios, tool variants, evaluation prompts, and rubrics grounded in physical attributes and affordances. Two rounds of human verification then filter examples and align prompts and criteria with realistic, visually checkable interactions.

  • A.1. Use of LLMs for Data Curation: GPT-5 instantiates tasks, proposes attribute-compatible unconventional tools, generates impossible tools by reversing affordances, and creates evaluation prompts and rubric questions.The pipeline also provides full prompt templates for each curation step.
  • A.2.1. Step 1.Action-to-Task Generation: Each action definition becomes 20 diverse, realistic, visually demonstrable tasks containing a conventional tool, expected outcome, and physics-derived required attributes.Tasks are designed for clear completion within a 5 second video clip and must show an obvious visible change.
  • A.2.2. Step 2: Unconventional Tool Generation: Unconventional tools are physically plausible substitutes that are less canonical but share relevant functional properties with the original tool.The prompt asks for everyday objects that match some required attributes and may still accomplish the task.
  • A.2.3. Step 3: Opposite Attribute and Impossible Tool Generation: Impossible tools oppose or lack critical attributes, making them physically incompatible with the action and useful for testing constraint awareness and failure recognition.The procedure generates objects that clearly cannot accomplish the task because of conflicting affordances or physics.
  • A.2.4. Step 4: Evaluation Instance and Prompt Construction: Each task expands into five evaluation instances—one regular, two unconventional, and two impossible—with predictive and descriptive prompts for both image and video generation.Predictive prompts withhold the final outcome, whereas descriptive prompts specify successful or failed final states; video prompts additionally cover the process.
  • A.2.5. Step 5: Rubric Generation: Each instance receives four rubrics measuring instruction adherence and interaction accuracy through checklist items tailored to entities, attributes, scene validity, state changes, affordances, and motion.The dimensions are scored from 0-100%, with motion plausibility applying specifically to video generation.
  • A.3. Human Verification: Four human volunteers conduct two verification rounds, filtering ambiguous or unrealistic tasks and checking tool plausibility, constraint violations, prompt clarity, and rubric alignment.The annotators include three master’s computer science students and one undergraduate physics student.

B. Evaluation Details

This section provides additional details about the evaluation protocol of TailOR.

  • The section elaborates on TailOR’s evaluation protocol.

B.1. Evaluation Data · B.2. Human Annotation

The study evaluates a resource-constrained subset of Tailor-Bench through human annotation, covering all scenario modes and generation settings. Annotators use structured checklists and open-ended ratings, with substantial agreement across both evaluation types.

  • B.1. Evaluation Data: B.1. Evaluation Data: 80 high-quality prompts span Regular, Unconventional, and Impossible scenarios plus Predictive and Descriptive settings, producing 1080 generated data points.The prompts were sampled from the full benchmark for human evaluation because annotation and computational resources were limited.
  • B.1. Evaluation Data: B.1. Evaluation Data: Blocked prompts, including cases involving Sora-2 and Veo3, are replaced with alternatives from the full benchmark to preserve the target number of evaluation instances.The passage attributes these blocks likely to safety filtering policies.
  • B.2. Human Annotation: B.2. Human Annotation: Evaluation is conducted by annotators with prior experience in visual content evaluation.This experience is intended to support reliable assessment of generated outputs.
  • B.2. Human Annotation: B.2. Human Annotation: Nine computer science students with computer vision and multimodal generation backgrounds receive detailed guidelines on benchmark objectives and evaluation criteria.The group includes four PhD students, three undergraduate students, and two master’s students.
  • B.2. Human Annotation: B.2. Human Annotation: Each sample is presented with its prompt and evaluation questions, then assessed using checklist-based structured metrics and open-ended quality ratings.Annotators read the scenario prompt, inspect the generated image or video, answer checklist questions, and assign open-ended quality scores.
  • B.2. Human Annotation: B.2. Human Annotation: Each generated sample is independently evaluated by at least three annotators, with final human scores computed by averaging their ratings.Annotators may optionally add comments about the sample or evaluation questions.
  • B.2. Human Annotation: B.2. Human Annotation: Agreement reaches an average Krippendorff’s α of 0.72 for checklist metrics and an average pairwise Spearman correlation of 0.68 for open-ended scores.The reported values indicate substantial agreement and consistent scoring behavior, respectively.

B.3. Automatic Evaluation with VLM Judge

The paper supplements human evaluation with a VLM judge for scalable and reproducible assessment. The judge evaluates generated media using the human annotators’ information and rubric, producing structured metric scores after reasoning through visible evidence and physical plausibility.

  • A VLM judge is used alongside human evaluation to improve scalability and reproducibility.
  • Judge Model: The automatic evaluation model is gemini-2.5-pro.
  • For each sample, the judge receives the prompt, generated media, and evaluation rubric available to human annotators.
  • Evaluation Prompt: The judge answers rubric questions, produces structured scores for each metric, and explains its reasoning before assigning final scores.
  • Evaluation focuses on object presence, physical attributes, interaction behavior, and resulting state change, using only visible evidence and strict physical-plausibility criteria.
Loading 2606.24256v1…