Source-linked AI summary

Interpreting Physics in Video World Models

Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal, Thomas Fel, Shahab Bakhtiari, Blake Richards, Mike Rabbat

arXiv:2602.07050v1cs.CVcs.AI

TL;DR

Video world models can perform physical reasoning, but their internal representational regime remains unclear. The paper probes encoder representations across layers, subspaces, patches, and attention circuits, finding a sharp Physics Emergence Zone where physical variables become accessible. Physical signals peak at intermediate depth, scalar motion variables appear earlier than direction, and direction uses a distributed circular population code rather than a compact factorized state.

  • Problem

    It remains unclear whether video world models internally represent physical variables as factorized, reusable states or as task-specific distributed representations.

  • Method

    The study uses layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations across encoder-based video transformers.

  • Results

    Across tested models, physical information emerges at a sharp intermediate-depth Physics Emergence Zone, with scalar motion variables accessible earlier than direction and physics representations degrading toward output layers.

  • Takeaways & Limitations

    The findings favor distributed, task-specific physical representations over compact reusable latent variables, while showing that these representations support physical predictions.

  • Takeaways & Limitations

    The study is limited to masked-objective encoder transformers, coarse physical reasoning and controlled motion variables, and synthetic toy-ball data that may not reflect natural video.

Abstract

from arXiv · show

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such variables in a task-specific, distributed manner. While modern video world models achieve strong performance on intuitive physics benchmarks, it remains unclear which of these representational regimes they implement internally. Here, we present the first interpretability study to directly examine physical representations inside large-scale video encoders. Using layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations, we characterize where physical information becomes accessible and how it is organized within encoder-based video transformers. Across architectures, we identify a sharp intermediate-depth transition -- which we call the Physics Emergence Zone -- at which physical variables become accessible. Physics-related representations peak shortly after this transition and degrade toward the output layers. Decomposing motion into explicit variables, we find that scalar quantities such as speed and acceleration are available from early layers onwards, whereas motion direction becomes accessible only at the Physics Emergence Zone. Notably, we find that direction is encoded through a high-dimensional population structure with circular geometry, requiring coordinated multi-feature intervention to control. These findings suggest that modern video models do not use factorized representations of physical variables like a classical physics engine. Instead, they use a distributed representation that is nonetheless sufficient for making physical predictions.

1. Introduction

The paper asks how video world models internally represent physical information and examines whether these representations resemble compact physics-engine variables or distributed, task-specific structure. It introduces interpretability analyses that locate and characterize these representations across physical reasoning tasks.

  • Prior work largely evaluated physical reasoning behaviorally, leaving where physical information is constructed and how it is organized across layers and patches unresolved.
  • The study uses layerwise probing, subspace analysis, and targeted ablations to map physical information’s accessibility, structure, and computational substrate.
  • The analysis combines possible-versus-impossible video judgments with a synthetic toy-ball dataset controlling velocity and acceleration across V-JEPA 2 and VideoMAE-v2 G encoders.
  • Physical representations emerge sharply at approximately one-third depth, peak in middle layers, and degrade toward the output; speed and acceleration appear earlier than direction.
  • Direction and possible-versus-impossible judgments occupy nearly orthogonal subspaces, while both rely on localized attention and direction uses a high-dimensional circular population code.

2. Related Work

Related work frames video world models as learned systems for reusable environmental structure and places this study within debates about physical reasoning and video interpretability. The paper focuses on encoder-based models because their persistent intermediate representations support representation-level analysis.

  • 2.1. Video world models: World models are learned systems whose internal representations capture reusable environmental structure for prediction, imagination, or planning.
  • 2.1. Video world models: Diffusion generators distribute computation across denoising steps, complicating representation-level analysis and motivating a focus on encoder-based video world models.
  • 2.2. Physical reasoning: Video world models can perform physical reasoning but show failures on causal, counterfactual, and violation-of-expectation benchmarks.
  • 2.3. Representations of physical reasoning: Prior cognitive-science work debates whether physical judgments rely on compact reusable latent states or heuristic, domain-specific reasoning.
  • 2.4. Interpretability of video models: Video interpretability remains comparatively limited because temporal complexity and dimensionality make analysis harder than in text and images.

3. Models and Probing Methodology

The study compares frozen encoders from V-JEPA 2 and VideoMAE-v2 using layerwise probes to determine where physical information is linearly accessible and how spatial-temporal structure is retained.

  • The study evaluates two state-of-the-art video transformer architectures with layerwise probing for physical reasoning.
  • V-JEPA 2 uses a latent prediction objective that maps spatiotemporal patches to representations and forecasts masked or future patches; analysis uses its frozen encoder.
  • VideoMAE-v2 uses masked autoencoding to reconstruct missing pixels, and the study analyzes its retained frozen encoder alongside V-JEPA 2.
  • Linear probes on mean-pooled space-time patches measure what is directly linearly available, while patch-preserving attentive-MLP probes retain spatial and temporal structure.

4. The Physics Emergence Zone

Physical reasoning becomes detectable at a consistent intermediate-depth transition called the Physics Emergence Zone. Across tested models, possible-versus-impossible discrimination rises sharply there, while physical representations are strongest at intermediate depth and weaken toward the output.

  • The Physics Emergence Zone is the layer range where probes begin performing well on possible-versus-impossible physical reasoning.
  • Across V-JEPA 2 scales, probe accuracy rises from approximately ∼50% near chance to ∼85–95% at approximately one-third encoder depth.
  • The transition is consistent across V-JEPA 2 model sizes and also appears in VideoMAE-v2-G, whereas smaller VideoMAE-v2 variants lack reliable emergence.
  • Object permanence, shape constancy, and spatiotemporal continuity violations show a similar emergence pattern, rather than one restricted to a single violation type.
  • Possible-versus-impossible representations peak in the middle third of the encoder and degrade toward the output layers.

5. Velocity and Acceleration in the Physics Emergence Zone

Layerwise probing separates early-accessible scalar motion quantities from direction, which emerges sharply at the Physics Emergence Zone. This pattern generalizes beyond single-object motion.

  • Method: Ground-truth velocity and acceleration are probed in Cartesian and polar representations using controlled synthetic ball videos.The dataset holds other factors fixed while measuring motion variables in pixels per frame.
  • Cartesian representations: Acceleration is decodable from early layers without requiring an explicit intermediate velocity representation.Both Cartesian acceleration components show high-R2 early-layer decoding.
  • Polar representations: Speed and acceleration magnitude are available early, whereas motion direction becomes reliably decodable only at the Physics Emergence Zone.The transition is therefore more closely associated with direction than scalar motion quantities.
  • Generalization: The same Physics Emergence Zone signature appears for object-level direction probes across CLEVRER object types and model scales.This extends the direction-emergence pattern beyond single-object motion.

6. What is the Relationship between Possible-vs-Impossible Physics and Direction?

Possible–impossible physics judgments and motion direction emerge at the same intermediate depth but use nearly separate representations. They instead share localized spatiotemporal processing in the Physics Emergence Zone.

  • Task specificity: The Physics Emergence Zone selectively appears for tasks requiring global spatiotemporal coherence, including intuitive physics and shuffled-video detection.CLEVRER counting and SSv2 classification lack the characteristic one-third emergence signature.
  • Representational geometry: Direction and possible–impossible judgments occupy nearly orthogonal subspaces, with principal angles averaging 69°–83° and projection overlap of only 7–13%.Speed has 81° average proximity to IntPhys, while direction has 69°; less than 3% of the IntPhys subspace projects onto speed.
  • Shared substrate: Local and long-range attention heads coexist uniquely at the Physics Emergence Zone, producing a sharp increase in attention-distance diversity.Other layers show relatively homogeneous attention profiles.
  • Functional ablation: Suppressing local attention there severely degrades direction decoding and possible–impossible discrimination while largely sparing ImageNet classification.The intervention identifies localized spatiotemporal processing as a shared circuit-level substrate rather than a shared representational subspace.
  • Scope: The analysis is limited to coarse spatiotemporal reasoning and does not examine contact dynamics or force-based inference.Those compositional computations may rely on additional mechanisms.

7. Steering the Direction Variable

Direction is represented by a distributed circular population code rather than a single controllable feature. Effective steering requires coordinated intervention across many dimensions.

  • Population geometry: Direction-selective MLP units at the end of the Physics Emergence Zone tile 360° and organize into a unit-circle geometry.Individual units show smooth direction tuning, whereas analogous circular organization is not observed for speed.
  • Population geometry: Manipulating only the unit-circle subspace does not effectively steer decoded direction, indicating a higher-dimensional representation.The circular geometry is therefore not sufficient by itself for causal control.
  • Dimensionality: Direction decoding requires roughly 40–50 independent features at the Physics Emergence Zone, increasing to up to 80 near output layers.Possible–impossible discrimination requires approximately 20 independent features at the emergence zone.
  • Structured redundancy: The sawtooth probe pattern is consistent with structured redundancy from approximately sinusoidal feature pairs such as sine–cosine components.This supports a distributed, paired-feature encoding of direction.
  • Causal steering: Coordinated interventions across many orthogonal probe directions reduce angular error smoothly, reaching < 0.5° at layer 8 versus > 80° for single-probe interventions.Effective steering requires manipulating a large fraction of the representational subspace.

8. Discussion

The findings support distributed, task-specific physical representations rather than compact shared latent variables. Direction nevertheless forms a structured circular geometry, paralleling biological motion coding.

  • Representational regime: Motion direction and possible–impossible judgments occupy nearly orthogonal representational subspaces, providing no evidence for shared low-dimensional latent variables.The result supports distributed, task-specific computation at the representational level.
  • Biological parallel: Direction in video world models emerges as a distributed circular geometry that parallels direction coding in biological vision.The paper connects this organization to neuroscience findings in which direction-selective neurons tile angular space.
  • Domain dependence: The observed distributed high-dimensional geometry contrasts with low-dimensional physical steering reported for a transformer trained on PDE simulations.Whether learned models expose compact physical state variables depends on the training domain and objective.

9. Limitations

The analysis is scoped to masked-objective encoder-based video transformers and two diagnostic settings, while leaving richer physical reasoning and complete circuit mechanisms unexamined.

  • The study covers encoder-based video transformers trained with masked objectives, so autoregressive and diffusion models may have different representational structures.
  • Its diagnostics target possible–impossible discrimination and controlled motion variables rather than contact dynamics, force inference, or long-horizon interaction.
  • The methods characterize representational accessibility and coarse causal influence, not a complete circuit-level mechanism.
  • The synthetic toy-ball dataset may not reflect how physical structure is represented in natural video.

10. Conclusion

Across V-JEPA 2 and VideoMAE-v2 G, physics-relevant information emerges at a sharp mid-depth transition, peaks afterward, and weakens toward the output layers. Scalar motion magnitudes are accessible earlier than direction, which becomes linearly accessible only at this transition.

  • Across V-JEPA 2 and VideoMAE-v2 G, possible–impossible discrimination and motion direction emerge at the Physics Emergence Zone.
  • Physics signals peak after the transition and then weaken toward the output layers.
  • Scalar motion magnitudes are available early, whereas direction becomes linearly accessible only at the Physics Emergence Zone.

Impact Statement

The impact statement reports no societal consequences requiring specific emphasis. The supplied material also describes controlled synthetic motion datasets, probe methodology, and consistent subtask emergence patterns.

  • Impact Statement: The authors state that their work has potential societal consequences but that none must be specifically highlighted.
  • Dataset construction: The controlled datasets use single spheres moving along straight-line trajectories with known ground-truth dynamics.
  • Dataset construction: The velocity dataset contains 392 videos spanning 8 directions, 7 speeds, and 7 start positions.
  • Dataset construction: The acceleration dataset contains 280 videos spanning 8 directions, 5 accelerations, and 7 start positions.
  • Subtask analysis: Probe performance shows the same one-third emergence pattern for object permanence, shape constancy, and spatiotemporal continuity.

C.1.4. THE MIDDLE LAYER POSSIBLE-VS-IMPOSSIBLE PHYSICS REPRESENTATIONS GENERALIZE TO BETTER PERFORMANCE ON A DOWNSTREAM INTUITIVE PHYSICS TASK

Intermediate encoder representations provide stronger possible-versus-impossible physics signals for downstream prediction than final-layer representations. The middle-layer advantage accompanies a broader transition toward distributed, spatially generalizable physical representations.

  • Middle-layer representations yield the best downstream predictivity on the violation-of-expectation intuitive physics task, outperforming final-layer representations.The task measures next-frame plausibility and is more real-world than the earlier setup.
  • Early encoder layers can show poor standalone performance, while a trained predictor can still extract useful physical information from them.The predictor itself learns from the representation, so encoder-only probe performance need not determine downstream utility.
  • Control tasks such as object counting, image classification, and standard video classification lack the one-third emergence signature, indicating that it is not a generic depth effect.Shuffled video detection is the exception, consistent with its reliance on temporal order.
  • Around the Physics Emergence Zone, direction information spreads across patches and becomes individually decodable, unlike earlier location-bound signals.This shift explains the abrupt rise in per-patch performance while mean-pooled performance improves more gradually.
  • After the emergence zone, direction decoding generalizes across spatial regions, indicating a globally accessible, position-invariant representation.The transition reflects a shift from local, retinotopic signals to globally distributed encoding.
  • Combined spatial and temporal attention ablation destroys direction encoding, while temporal ablation strongly affects both direction and intuitive-physics reasoning.Spatial ablation minimally affects direction R² but degrades per-patch localization.
Loading 2602.07050v1…