Source-linked AI summary

VideoPhy: Evaluating Physical Commonsense for Video Generation

Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, Aditya Grover

arXiv:2406.03520v2cs.CVcs.AIcs.LG

TL;DR

Existing text-to-video models may become physical-world simulators, but their adherence to text and physical commonsense remains unclear. VideoPhy benchmarks these capabilities with diverse interaction prompts and human evaluation, then introduces VideoCon-Physics for scalable assessment. CogVideoX-5B achieves joint caption-and-physics correctness on 39.6% of instances, while the benchmark concludes that current models remain far from accurate physical simulation.

  • Problem

    It remains unclear whether text-to-video models generate videos that both follow prompts and obey physical commonsense, limiting evidence about their suitability as physical-world simulators.

  • Method

    VideoPhy curates diverse captions covering solid-solid, solid-fluid, and fluid-fluid interactions, evaluates twelve models with human judgments, and fine-tunes VIDEOCON into VideoCon-Physics.

  • Results

    39.6% of CogVideoX-5B instances satisfy both semantic adherence and physical commonsense, while other evaluated models score below 20%.

  • Takeaways & Limitations

    VIDEOPHY finds that existing video models significantly lack physical commonsense and semantic adherence, and are far from general-purpose world simulators.

  • Takeaways & Limitations

    The evaluation excludes some closed models, including Sora, Kling AI, and Genmo, because their videos were unavailable through unsupported APIs.

Abstract

from arXiv · show

Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic motions and render complex objects. Hence, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate diverse prompts that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., CogVideoX) and closed models (e.g., Lumiere, Dream Machine). Our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we propose an auto-evaluator, VideoCon-Physics, to assess the performance reliably for the newly released models.

1 Introduction

VideoPhy addresses whether modern text-to-video models can serve as physical-world simulators by evaluating caption adherence and intuitive physical behavior. It introduces a benchmark of human-verified real-world dynamics and finds substantial shortcomings.

  • Internet-scale video pretraining has produced models capable of generating realistic motions, complex scenes, and videos conditioned on text prompts.
  • The benchmark targets an unresolved question: whether generated videos from text-to-video models adhere to physical laws.
  • VideoPhy curates human-verified captions and evaluates generated videos for both semantic adherence and physical commonsense.
  • 39.6% of instances satisfy both caption adherence and physical accuracy for the best-performing model, CogVideoX-5B.

2 VIDEOPHY Dataset

VIDEOPHY is a benchmark built from diverse, human-verified captions describing physical interactions and annotated for simulation difficulty. Its design spans material interactions, activities, objects, and the complexity of rendering motions.

  • VIDEOPHY covers daily activities, physical objects, interactions between material types, and perceived rendering complexity.
  • The caption-generation pipeline classifies real-world dynamics into solid-solid, solid-fluid, and fluid-fluid interactions, including inviscid and viscous fluids.
  • Human verification filters captions for clarity, manageable complexity, and accurate interaction-category assignments.
  • 688 captions remain after verification, comprising 289 solid-solid, 291 solid-fluid, and 108 fluid-fluid interactions.
  • Two physics-based simulation researchers label captions easy or hard according to perceived modeling and simulation complexity.
  • Difficulty is category-relative and depends on material modeling and numerical factors such as deformability, velocity, and high-order PDE terms.
  • The dataset includes 11,330 generated videos and 138 unique actions, while captions average 8.5 words.

3 Evaluation

VideoPhy evaluates generated videos through human judgments of semantic adherence and physical commonsense, while VideoCon-Physics provides a scalable automatic alternative. The metrics use binary judgments to isolate these capabilities despite their entanglement with other video properties.

  • The evaluation focuses on semantic adherence and physical commonsense because video quality, motion, temporal consistency, and text alignment are intertwined.
  • Semantic adherence, SA, is a binary measure of whether caption actions, events, entities, and relationships are grounded in generated frames.
  • Human evaluation uses qualified workers who judge each caption-video pair for semantic adherence and physical commonsense without seeing the generating model.
  • Annotators rely on intuitive physical knowledge rather than listing individual physics violations, keeping judgments less costly and time-consuming.
  • Human evaluation is more accurate for benchmarking but expensive and time-consuming at scale.
  • VideoCon-Physics fine-tunes VIDEOCON on human semantic-adherence and physical-commonsense annotations to enable cheaper, scalable evaluation.

4 Setup

The setup evaluates twelve open and closed text-to-video models on VIDEOPHY using train/test splits and human scoring. It also documents annotation volume, agreement, and access limitations affecting the model set.

  • The study evaluates twelve diverse open and closed text-to-video models, including CogVideoX, Lumiere, Dream Machine, Pika, and Gen-2.
  • VIDEOPHY prompts are split equally into 344 training and 344 test prompts for automatic-evaluator training and benchmarking.
  • Each test prompt produces one video per model, which three human annotators judge for semantic adherence and physical commonsense.
  • The study could not evaluate some closed models, including Sora, Kling AI, and Genmo, because their videos were inaccessible without API support.
  • Human judgments show 75% agreement for semantic adherence and 70% for physical commonsense across 24,500 test annotations.
  • VIDEOCON-Physics training uses two videos per training prompt from nine models and 12,000 human annotations.

5 Results

Across VIDEOPHY evaluations, current text-to-video models frequently fail to satisfy both caption semantics and physical commonsense, with performance varying by model, material interaction, and caption complexity.

  • Performance on VIDEOPHY Dataset: 39.6% of cases generated by CogVideoX-5B achieved both semantic adherence and physical commonsense, while every other model scored below 20%.This result comes from human evaluation using SA = 1 and PC = 1 as the success condition.
  • Performance on VIDEOPHY Dataset: 53% physical-commonsense performance made CogVideoX-5B the best model, followed by CogVideoX-2B at 34.1%.The reported scaling pattern suggests larger network capacity improves capture of physical constraints in internet-scale video data.
  • Performance on VIDEOPHY Dataset: Dream Machine achieved 61.9% semantic adherence but only 21.8% physical commonsense, showing that following the prompt does not necessarily imply physical plausibility.Among closed models, Pika achieved positive judgments on both criteria in 19.7% of cases.
  • Variation with the States of Matter: All models performed worst on solid-solid interactions, where CogVideoX-5B reached 24.4% for accurate semantic adherence and physical commonsense.Performance was strongly influenced by the states of matter involved, with Pika performing best on fluid-fluid interactions.
  • Variation with Complexity: Semantic adherence and physical commonsense decreased as caption complexity increased, making physically harder captions more difficult to control through conditioning.The benchmark therefore exposes a gap between easy and hard captions for video generation.
  • Qualitative Analysis: Qualitative failures included unequal water-stream speeds, sudden bread-shape changes, airborne droplets remaining still, and material-inconsistent deformations.The analysis also identifies conservation-of-mass and Newton’s First Law violations among common failure modes.
  • Automatic Evaluation: VIDEOCON-PHYSICS outperformed GPT-4Vision and Gemini-1.5-Pro on ROC-AUC for both semantic-adherence and physical-commonsense judgments.The comparison evaluates automatic methods on testing prompts from VIDEOPHY.

6 VIDEOCON-PHYSICS: Automatic Evaluator for VIDEOPHY Dataset

VIDEOCON-PHYSICS is an automatic evaluator designed for scalable assessment of semantic adherence and physical commonsense in generated videos. It generalizes to unseen prompts and generative models, and its automatic rankings largely track human rankings.

  • VIDEOCON-PHYSICS provides scalable and reliable automatic evaluation of semantic adherence and physical commonsense.
  • 17 points and 19 points: VIDEOCON-PHYSICS outperforms zero-shot VIDEOCON on semantic adherence and physical commonsense judgments, respectively, for unseen prompts.
  • 15 points and 15 points: an ablated VIDEOCON-PHYSICS outperforms VIDEOCON on semantic adherence and physical commonsense judgments for unseen generative models.
  • Automatic and human leaderboards show similar relative model rankings, supporting VIDEOCON-PHYSICS as a reliable evaluation tool.
  • Finetuning video models on VIDEOPHY data decreases semantic adherence while leaving physical commonsense unchanged.

7 Related Work

Prior video-generation evaluation methods assess distributional similarity, text alignment, or broad video quality, but largely overlook physical commonsense. Video-generation research spans diffusion and autoregressive architectures, while physics modeling provides established simulation frameworks for solids and fluids.

  • Video Generation Models: Diffusion-based and autoregressive architectures form the two primary approaches to video generation discussed in prior work.
  • Evaluating Video Generation Models: FVD requires reference videos, favors video quality, and can miss unrealistic motions, limiting its suitability for physical-commonsense evaluation.
  • Evaluating Video Generation Models: CLIPScore measures video-text similarity but is unsuitable for evaluating physical commonsense in generated videos.
  • Evaluating Video Generation Models: Existing comprehensive evaluation frameworks introduce multiple quality metrics, yet largely overlook physical commonsense.
  • Physics Modeling: Physics modeling includes rigid-body simulation for nondeforming solids and deformable-solid models that account for strain and stress.

8 Conclusion

The paper introduces VIDEOPHY as a benchmark for physical commonsense in generated videos and VIDEOCON-PHYSICS as a scalable automatic evaluator. It constructs diverse material-interaction captions and evaluates video-generation models with multimodal judgments.

  • 8 Conclusion: VIDEOPHY is introduced as a dataset for assessing physical commonsense in generated videos, while VIDEOCON-PHYSICS enables cheap, scalable evaluation.
  • 8 Conclusion: The study benchmarks diverse open and closed video models, including Zeroscope, LaVIE, SVD-T2I2V, Gen-2, Pika, Lumiere, and CogVideoX.
  • 8 Conclusion: The benchmark covers daily activities, solid-solid, solid-fluid, and fluid-fluid interactions, plus the perceived complexity of rendering objects and motions.
  • 8 Conclusion: GPT-4 prompts generate concise captions for solid-solid, solid-fluid, and fluid-fluid material interactions using common materials and excluding nonphysical social actions.
  • 8 Conclusion: The evaluation pipeline prompts multimodal evaluators for yes/no judgments of semantic adherence and physical commonsense using the video and conditioning caption.

G Fine-Grained Diversity Analysis

The benchmark analyzes fine-grained statistics across physical-interaction categories and compares video-generation models across task complexity.

  • Fine-grained collection statistics are visualized across different physical-interaction categories.
  • Video-generation models are compared across different levels of task complexity.

I Fine-Grained Results

The evaluation reports fine-grained semantic-adherence and physical-commonsense scores across physical interaction categories and caption-difficulty levels.

  • Interaction categories: Scores are computed across solid-solid, solid-fluid, and fluid-fluid interaction categories.
  • Difficulty levels: Scores are also compared across difficulty levels 0 and 1.
  • Evaluation dimensions: The evaluated dimensions are semantic adherence and physical commonsense.

J Automatic Evaluation Baselines

The automatic evaluation uses vision-language models to judge whether generated videos follow their captions and physical commonsense, with inputs adapted to each model's video-processing capabilities.

  • Automatic evaluation: GPT-4Vision receives the caption and eight uniformly sampled video frames for zero-shot evaluation.
  • Scoring: The evaluator assigns binary semantic-adherence and physical-commonsense scores.
  • Probability interpretation: The token probabilities for Yes and No need not sum to one because VIDEOCON predicts over the full vocabulary.
  • Video input: Gemini-Pro-Vision-1.5 evaluates the caption and entire generated video because GPT-4V does not process videos natively.

L Training Details for VIDEOCON-PHYSICS

VIDEOCON-PHYSICS is trained as an automatic evaluator and used to rank video generators by semantic adherence and physical commonsense, while comparisons examine quality, motion, and human agreement.

  • Training: VIDEOCON-PHYSICS is fine-tuned with LoRA on attention-block projections for five epochs using Adam optimization.The configuration uses r = 32, α = 32, dropout = 0.05, 50-step warmup, linear decay, and a peak learning rate of 1e-4.
  • Video processing: Videos are processed as 32 segments with one middle frame sampled from each segment after resizing frames to 224 × 224.
  • Leaderboard construction: Model rankings average automatic semantic-adherence and physical-commonsense scores and are compared with human joint performance rankings.
  • Leaderboard agreement: The automatic leaderboard reliably tracks the human leaderboard for open and closed video generators.
  • Quality comparison: Gen-2 combines a high video-quality score of 5.8 with a poor semantic-adherence and physical-commonsense score of 7.6.

P Finetuning video model with VideoPhy data

The paper tests whether VideoPhy data can improve a video generator, but fine-tuning Lumiere-T2I2V reduces semantic adherence while leaving physical commonsense unchanged; qualitative examples document varied violations.

  • Training data: The fine-tuning experiment uses 1000 training video-caption pairs with joint physical-commonsense and semantic-adherence scores of 1.
  • Fine-tuning outcome: Semantic adherence decreases by a large margin after fine-tuning, while physical commonsense remains unchanged.
  • Interpretation: The paper attributes the result to possible sample scarcity, optimization difficulties from mixed on-policy and off-policy videos, and unsuitable vanilla fine-tuning.
  • Qualitative violations: Qualitative examples span physical-law violations across multiple generators, including deformation, penetration, mass conservation, and Newtonian-law failures.
  • Model coverage: Examples also include violations from SVD-T2I2V, Pika, Lumiere-T2V, Lumiere-T2I2V, CogVideoX, and Dream Machine.
Loading 2406.03520v2…