Source-linked AI summary
Principia: Relational Physics Tests for Video Models
Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad
TL;DR
Video generators can look realistic while violating physical laws, and absolute motion measurements are often unavailable or ambiguous in generated video. Principia addresses this with calibration-independent relational tests over paired objects, finding weak physical consistency in generators and limited violation detection by VLMs.
Problem
Visual realism does not imply physical consistency, while absolute motion measurements depend on frame rate, scale, and camera calibration that may be ambiguous or unavailable.
Method
Principia evaluates eight Newtonian phenomena using relational consistency between paired objects, image-space scoring, controlled real-world scenes, and scalable Isaac Sim scenarios.
Results
Across six state-of-the-art video generators, no model exceeds 0.45 on Principia despite around 0.8 on visual-quality benchmarks, while the best VLM achieves only 67% accuracy.
Takeaways & Limitations
Principia indicates that current video generators’ visual realism is not accompanied by reliable relational physical consistency, and VLM detection remains weak.
Takeaways & Limitations
Principia is limited to macroscopic Newtonian mechanics and does not address fluid dynamics, soft-body deformation, or thermodynamic phenomena.
Abstract
from arXiv · showhide
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.
1 Introduction
Principia addresses the gap between visual realism and physical consistency in video generation by testing calibration-independent relationships between paired objects. It combines controlled real-world scenes, relational scoring, and scalable synthetic scenarios to evaluate generators and VLMs.
- Motivation: Existing physical-consistency benchmarks rely on plausibility judgments, reference trajectories, or explicit laws, each with important evaluation limitations.Plausibility is subjective, trajectory matching can penalize valid alternative futures, and law-based evaluation may require recovered physical quantities.
- Benchmark: Principia evaluates eight physical phenomena by measuring whether paired objects’ relative motions satisfy the expected physical relation.The benchmark covers gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulums, and mass-spring systems.
- Calibration-independent evaluation: Relational invariants enable image-space evaluation without estimating absolute mass, scale, velocity, acceleration, or camera parameters.For example, matched blocks sliding under friction should reach the bottom simultaneously despite different masses.
- Dataset construction: About 500 scenes cover eight Newtonian phenomena across translational, rotational, collisional, and oscillatory dynamics, using synchronized and controlled real-world experiments.Automated and manual validation filter videos affected by geometric, temporal, surface, or release asymmetries.
- Scalable evaluation: Principia also provides an Isaac Sim pipeline for controlled and counterfactual scenarios that evaluates both video generators and VLMs.The synthetic testbed supports scalable testing and can be used to develop methods for improving physical consistency.
- Results: No evaluated generator exceeds 0.42 on Principia, despite scores around 0.8 on standard visual benchmarks.The result indicates that visual realism does not imply physical consistency.
2 Related Work
Prior physics benchmarks use calibrated environments, VQA, human or VLM judgments, or real-video comparisons, often requiring absolute measurements. Principia extends calibration-independent relational testing from gravity to eight Newtonian phenomena and evaluates both generators and VLMs.
- Existing benchmarks: Existing physics benchmarks span simulation reasoning, VQA, commonsense scoring, and real-video evaluation, but all require calibration, scale, or recorded ground-truth parameters.These quantities may be ambiguous in generated video.
- Prior relational evaluation: The closest precursor tests ratios of fall times between two objects to isolate Galileo’s principle without calibration.That protocol is unit-free, relational, and quantitative.
- Principia: Principia generalizes relational evaluation to eight Newtonian phenomena across translational, rotational, collisional, and oscillatory dynamics.It also evaluates both video generators and vision-language models.
3 The Principia Dataset
The Principia dataset uses paired-object scenes whose relational invariants remain valid across camera, scale, and frame rate. Controlled real-world recordings, filtering, synthetic counterfactuals, and physics-specific invariants support evaluation across eight phenomena.
- Dataset overview: Principia contains 500+ real-world paired-object scenes spanning eight Newtonian phenomena.Its benchmark combines real video data with relational, quantitative physics evaluation for generators and VLMs.
- Experimental design: Each scene holds selected parameters fixed while varying one factor, so paired motions isolate the physical law under test.The invariant must hold independently of camera, scale, or frame rate.
- Construction protocol: Approximately 750+ videos were recorded and heavily filtered because small asymmetries can mimic genuine relational physics violations.Filtering targets geometric, timing, release, motion, spin, and surface anomalies.
- Measurement: Object tracking and all subsequent measurements occur in pixel space, making the consistency score independent of camera intrinsics, frame rate, and metric scale.The score uses ratios and equalities of pixel-space quantities.
- Translational and collisional phenomena: Gravity and restitution compare height and time or rebound ratios, while friction tests equal arrival times for different-mass blocks.These relations instantiate calibration-independent invariants derived from the corresponding laws.
- Rotational inertia: Rotational inertia compares arrival-time ratios for matched solid and hollow cylinders whose moments of inertia differ.The higher-inertia cylinder arrives later under no-slip rolling.
- Momentum and projectile motion: Momentum and projectile scenarios test distance or range relations, including greater displacement from greater launch height and shorter travel for heavier blocks.The projectile invariant is an ordering: the higher-launched sphere lands farther.
- Oscillatory phenomena: Pendulum and mass-spring scenarios test period ratios based on length and extension ratios based on suspended mass.The pendulum relation uses T1/T2 = l1/l2, while spring extension scales as x1/x2 = m1/m2.
4 Experiments
The experiments evaluate video generators and vision-language models with Principia’s relational consistency measures across physical phenomena. Visual quality remains high while physical fidelity is substantially lower, and performance varies across phenomena and model families.
- Evaluation Setup: Principia reports continuous relational consistency Sϕ across physical phenomena, while Table 3 summarizes mean ± standard deviation across scenario scores.Higher Sϕ indicates better adherence to the underlying physical invariant, and the Overall score summarizes performance across phenomena.
- Visual Quality and Physical Fidelity: Six generators score around 0.8 on VBench but between 0.14 and 0.42 on Principia, showing visual quality and physical fidelity are nearly orthogonal.Models at the visual-quality frontier are no more likely to satisfy physical invariants than lower-quality alternatives.
- Generator Results: Wan2.2-14B achieves the best overall generator score at 0.419, narrowly ahead of Omni at 0.409 and Veo-3.1 at 0.38.No generator exceeds an average consistency score of 0.5 across the evaluated phenomena.
- Generator Results: Scaling effects are uneven: Wan2.2-14B improves substantially on several phenomena but decreases momentum consistency by −0.05, while Cosmos scaling also produces mixed changes.For Wan2.2, gains include restitution (+0.35), gravity (+0.32), friction (+0.40), inertia (+0.26), and projectile (+0.22).
- Per-Phenomenon Results: Generator failure profiles differ by phenomenon, with Omni strong on inertia and friction, Wan2.2-14B strongest overall, and all models performing poorly on momentum.No model dominates across all physical phenomena, and none fully covers the radar chart.
- Vision-Language Models: Vision-language models are evaluated on detecting relational physics violations using real-world and synthetic samples, with no model exceeding an average agreement of 0.7.The synthetic Principia-Synth dataset includes both physically valid videos and videos that explicitly violate each phenomenon’s relational invariant.
5 Discussion and Conclusion
Principia reveals a large gap between visual realism and physical consistency in current video generators, while VLMs also struggle to detect relational violations. The benchmark targets interpretable videos governed by macroscopic physical laws and identifies scope boundaries for evaluation.
- No evaluated generator exceeds 0.45 on Principia’s continuous metric, despite scores around 0.8 on visual-quality benchmarks.The result spans six state-of-the-art generators.
- Current generators produce realistic renderings but fail to enforce the relational invariants required by physical laws.
- Most VLMs perform near chance on violation detection, with the best model reaching only 67% accuracy.They distinguish physically consistent videos from videos containing physics violations.
- Principia evaluates physical-law adherence in otherwise interpretable videos, rather than automatically separating hallucinated or severely deformed objects from physics violations.
- Principia covers macroscopic Newtonian mechanics but excludes fluid dynamics, soft-body deformation, and thermodynamic phenomena.
- The evaluation protocol scores generated videos according to whether their motion satisfies the tested relational constraints.
A.1 Dataset Statistics:
The dataset is built from 401 in-house real-world scenes that are edited and augmented into 529 conditioning scenes. Figure 6 summarizes their distribution across physical phenomena and real-world versus augmented scenarios.
- 401 real-world scenes were captured in-house and used as the dataset’s source recordings.
- 529 scenes result after removing experimenters and apparatus from conditioning frames and augmenting them for visual diversity.
- Figure 6 separates the available samples for each physical phenomenon into real-world and augmented scenarios.
A.2 Evaluation protocol
The evaluation extracts object trajectories and phenomenon-specific events from generated videos, then compares measurements against relational physics expectations. Generations are conditioned on edited first frames and prompts under default model configurations.
- Object trajectories are automatically extracted with SAM3, and detected impacts, arrivals, turning points, and related events provide measurements for each invariant.
- Restitution and gravity use ball trajectories to detect ground impact, rebound apex, drop height, rebound height, and flight time.
- Friction and rotational inertia are evaluated from the first frame when each object’s bottom pixel crosses the annotated incline endpoint.
- Projectile evaluation tracks the projectile centroid, detects ground impact, and measures horizontal range from the annotated launch point.
- Pendulum periods are estimated by detecting horizontal-velocity reversals after midpoint crossings, while spring tests compare first-oscillation extension ratios against mass ratios estimated by m ∝ l^3.
- Momentum evaluation compares maximum post-collision target displacement across two interactions to assess the expected ordering of momentum transfer.
A.5 Evaluation Compute.
Directional Consistency Score filters videos by qualitative motion direction before Principia scoring, using image-space displacement conventions that differ for pendulum motion. Generating the corpus required substantial compute.
- 2,800 A100-hours were required across four open-weight models, equivalent to 120 days of continuous computation on one A100.Wan2.2-14B and Cosmos-2.5-14B required approximately 1,100 and 1,300 A100-hours, respectively.
- DCS compares each displacement direction with the expected direction and ranges from −1 for consistently opposite motion to 1 for perfect agreement.
- For non-pendulum scenarios, the tracked coordinate is vertical and downward image motion is positive; pendulum motion instead uses the horizontal coordinate toward equilibrium.
- Videos with DCS below 0.8 receive a Principia Score of 0, filtering failures to generate the expected qualitative motion.
A.7 Scaling Within Architecture
Scaling within a fixed architecture improves some physical phenomena but degrades others, so larger video generators do not uniformly achieve higher physical fidelity.
- A.7 Scaling Within Architecture: Wan improves on friction (+0.40), inertia (+0.26), and pendulum (+0.17) but regresses on momentum (−0.05) when scaling from 5B to 14B parameters.
- A.7 Scaling Within Architecture: Both model families degrade on at least one phenomenon as parameter count increases, indicating that scaling does not uniformly improve physical fidelity.
- A.7 Scaling Within Architecture: Physical fidelity is unlikely to be solved by scale alone within current architectural choices.
A.8 Camera Sensitivity Analysis
Principia’s image-space measurements require a static camera because camera motion can introduce apparent object motion and confound attribution of deviations to physics.
- A.8 Camera Sensitivity Analysis: Camera motion can introduce apparent object motion and interfere with SAM-mask measurements of relational invariants.
- A.8 Camera Sensitivity Analysis: Uncontrolled scenes can produce deviations from expected trajectories through several confounding factors, making physics-specific attribution difficult.
- A.8 Camera Sensitivity Analysis: Table 6 reports Principia scores on real-world videos under simulated camera motion, varying pan and zoom magnitudes as percentage changes from the original video.
B Vision-Language Model Evaluation
Principia evaluates vision-language models on detecting relational physics violations in real and synthetic videos using phenomenon-specific PASS/FAIL judgments. Performance does not improve with model scaling, and the best reported model reaches only 67% accuracy.
- B Vision-Language Model Evaluation: VLMs classify whether videos satisfy relational invariants, unlike video generators, which are scored on preserving those invariants in generated motion.
- B Vision-Language Model Evaluation: VLMs receive real-world or anti-physics videos and phenomenon-specific prompts requesting PASS/FAIL judgments with brief explanations.
- B Vision-Language Model Evaluation: Phenomenon-specific prompts target restitution, gravity, friction, rotational inertia, projectile motion, momentum, mass-spring extension, and pendulum period.
- B Vision-Language Model Evaluation: Agreement score A_ϕ is the fraction of scenes where a VLM’s PASS/FAIL judgment matches the Principia-synth ground-truth binary score.
- B Vision-Language Model Evaluation: Scaling does not improve VLM performance: Gemini and Qwen show overall reductions of 0.12 and 0.04, respectively.
C Additional Qualitative Results of Video Generators on Principia
Qualitative examples show that video generators often violate relational physics through implausible motion, object hallucination, deformation, or failure to preserve expected comparisons. Some larger models produce plausible behavior for selected phenomena but still fail elsewhere.
- C Additional Qualitative Results of Video Generators on Principia: The project webpage is recommended for viewing model-wise qualitative results from Figures 10–14.
- C Additional Qualitative Results of Video Generators on Principia: Wan2.2-5B produces implausible restitution, friction, and projectile motion, including a hallucinated third ball and deforming blocks.
- C Additional Qualitative Results of Video Generators on Principia: Wan2.2-14B violates expected restitution and projectile relationships while also producing implausible sliding motion with hallucinating and deforming blocks.
- C Additional Qualitative Results of Video Generators on Principia: Cosmos2.5-2B fails qualitatively by hovering instead of falling, remaining stationary instead of sliding or rolling, and changing object identities during projectile motion.
- C Additional Qualitative Results of Video Generators on Principia: Cosmos2.5-14B produces plausible restitution and friction behavior but fails on projectile motion by landing the right ball on the cardboard box.
- C Additional Qualitative Results of Video Generators on Principia: Veo-3.1 and Omni preserve plausible friction timing but violate restitution and projectile relationships.