Source-linked AI summary

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

Finn Rasmus Schäfer, Korbinian Moller, Yuan Gao, Christian Oefinger, Sebastian Schmidt, Johannes Betz

arXiv:2607.05966v1cs.RO

TL;DR

Long-horizon world-model failure is commonly framed as compounding error without specifying which error compounds. This paper reframes the issue as kinematic versus dynamic imagination, operationalizes it with iKCE and conditioning perturbations, and finds that DreamerV3 rollouts remain insensitive to a friction-driven regime change that collapses policy reward.

  • Problem

    Compounding-error accounts do not distinguish whether long-horizon world-model deterioration reflects kinematic or dynamic deficiencies.

  • Method

    The paper introduces imagined Kinematic-Consistency Error and a conditioning-perturbation protocol to test imagined rollouts across physical-regime changes.

  • Results

    Across a friction sweep crossing the gait-collapse boundary, WM iKCE is statistically flat while trained-policy reward falls from ∼650 at µ=0.5 to ∼200 at µ=0.10.

  • Takeaways & Limitations

    The results support a kinematic-not-dynamic signature in which imagined rollouts are insensitive to a physical regime change affecting real behavior.

  • Takeaways & Limitations

    Evidence is limited to a single 2D 9-DOF DMC walker-walk embodiment, a restricted (z, vz) state slice, and one open-weight world-model family.

Abstract

from arXiv · show

Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does not distinguish what kind of error compounds. We propose a kinematic-vs-dynamic reframing: world models tend to imagine kinematically rather than dynamically. We operationalize this as the imagined Kinematic-Consistency Error, a per-step diagnostic that measures how far a rollout departs from a closed-form kinematic null, paired with a perturbation protocol that tests whether iKCE responds when physical conditions cross a regime boundary. We instantiate the diagnostic on a released DreamerV3 checkpoint trained on DMC walker-walk, where imagined iKCE runs roughly two orders of magnitude above that of matched real-physics rollouts. Across a friction sweep that crosses the gait-collapse boundary, the model's iKCE stays statistically flat even as the trained policy's reward collapses through the same range, providing the kinematic-not-dynamic signature. The diagnostic distinguishes kinematic from dynamic imagination at horizons longer than the embodiment's gait period.

I. Introduction

The paper reframes long-horizon world-model failure as a distinction between kinematic continuation and dynamic fidelity. It introduces iKCE and a perturbation protocol to test this distinction, finding two signatures in DreamerV3.

  • Motivation: World models support planning and self-supervised learning, but long-horizon rollout deterioration is usually described only as compounding error.That framing does not identify which error type or feature dimensions deteriorate.
  • Kinematic versus dynamic motion: Kinematic motion uses position, velocity, and acceleration trajectories without reproducing the forces or constraints that produce them.Dynamic motion requires constraints such as mass, friction, and contact to be reproduced correctly.
  • Core diagnosis: The paper argues that current world models extrapolate internally consistent position-velocity-acceleration trajectories while violating the physical constraints underlying real motion.
  • Related accounts: Kinematic fallback is presented as a third account of long-horizon failure alongside predictable-representation engineering and error-compounding bounds.A representation can be stable and accurate within training conditions yet remain biased toward kinematic continuation across physical-regime boundaries.
  • Contributions: The paper introduces iKCE and a conditioning-perturbation protocol as a falsifiable diagnostic for kinematic-versus-dynamic imagination.
  • Contributions: At T=16, the kinematic-null residual is ∼180× above matched physics, while imagined rollouts remain invariant across a friction sweep crossing gait collapse.The paper identifies regime invariance, rather than absolute iKCE magnitude, as the diagnostic signature.

II. Evidence

Four heterogeneous observations across driving and VLA settings converge on a structural deficit: learned representations emphasize kinematic features while under-representing dynamic features needed for regime-conditional behavior.

  • Evidence: Four existing observations are individually inconclusive but jointly form a coherent structural signature motivating the paper’s diagnosis.
  • Representational diagnostic: WPCR rises from ∼20 without visual input to ∼97 with one frame, then remains essentially flat as frames are added or temporally shuffled.Text-only input already achieves 59.6% BAcc, while reintroducing video recovers only ∼2.6pp under the best encoding.
  • Sensor-degraded behavior: Under heavy Gaussian noise (σ = 70), Alpamayo R1’s trajectory decoder collapses kinematic priors while its language branch remains coherent but safety-irrelevant.
  • Trajectory-prediction baselines: Kinematic-only inputs match perception-based planners on nuScenes open-loop L2, achieving 0.29 m versus 0.37 m for VAD-Base.The cited authors attribute this to the trajectory distribution and coarse collision-evaluation grid, calling for rethinking open-loop evaluation.
  • Physics-consistency scoring: KCE shows no monotonic relationship to model scale or modality, with Gemini-3-Pro at 0.06–0.11 m overlapping smaller fine-tuned models at 0.08–0.12 m.
  • Synthesis: Across these observations, learned representations are dominated by kinematic features and under-represent dynamic features required for physical-regime-conditional behavior.

A. iKCE: Imagined Kinematic Consistency Error

iKCE measures how much imagined next states depart from a closed-form kinematic continuation. Its interpretation is asymmetric: low iKCE indicates kinematic consistency, whereas dynamic imagination requires positive, horizon-growing, regime-responsive error.

  • Definition: iKCE is defined over imagined world-model rollouts using a chosen kinematic state vector such as [x, y, v, a, θ]⊤.
  • Definition: The kinematic predictor is any closed-form model, such as constant-velocity or constant-acceleration prediction, matched to the embodiment and output space.
  • Diagnostic role: At test time, iKCE measures each predicted next state’s departure from the kinematic extrapolation of its predecessor rather than supervising the model to reduce that error.The prefixed “i” denotes imagined rollouts.
  • Interpretation: Near-zero iKCE does not certify dynamic imagination because it can result from exact kinematic continuation.
  • Interpretation: Dynamic imagination requires iKCE that is positive, grows with horizon, and responds to physical-regime conditioning such as friction transients or contact events.Accordingly, iKCE is necessary but not sufficient to certify dynamic imagination.

B. Conditioning perturbations

The perturbation protocol evaluates iKCE across controlled changes to physically meaningful conditioning parameters. A kinematic imaginer should respond smoothly to perturbation magnitude rather than to physical-regime changes.

  • Protocol: For each base rollout, the protocol generates K imagined rollouts under controlled perturbations of initial velocity, friction, or lateral-acceleration limits.
  • Protocol: The protocol converts static single-rollout iKCE into a dose-response curve over the conditioning state.
  • Predicted signature: A kinematic imaginer produces iKCE that scales smoothly with perturbation magnitude regardless of physical regime.This follows because it extrapolates the same linear update structure in every regime.
  • Diagnostic interpretation: The protocol is analogous to weighted pairwise consistency testing, replacing question-pair checks with rollout comparisons under controlled conditioning perturbations.

IV. Experiments and Results

The DreamerV3 checkpoint shows substantial imagined iKCE relative to matched physics and remains friction-invariant across an empirically defined gait-collapse boundary, while controls support the kinematic-not-dynamic signature.

  • H1: Non-degenerate iKCE: ∼180×: imagined iKCE exceeds matched real-physics rollouts at T=16, while remaining at least an order of magnitude higher at both measured horizons.The ratio narrows at T=64 because the world model’s smooth post-transient tail dilutes the integrated metric.
  • H2: Friction invariance: The friction sweep spans 13 magnitudes from 0.1 to 1.7 and crosses the empirically determined boundary at µ=0.20, where trained gait reward first falls below 50% of baseline.The boundary is defined from the trained policy’s reward collapse rather than selected a priori.
  • H2: Friction invariance: At T=64, WM iKCE is statistically flat across friction, with a 1.32× max/min spread and overlapping 95% CIs at every friction value.The WM regression slope is βWM = −0.009 with 95% CI [−0.096, +0.082], which contains zero.
  • H2: Friction invariance: The same sweep produces physics-side sensitivity: βphys = −0.220 with 95% CI [−0.301, −0.142], while trained-policy reward falls from ∼650 at µ=0.5 to ∼200 at µ=0.10.Real-physics iKCE is also elevated at low friction, with means 1.04–1.30 × 10−4 versus 0.70–0.92 × 10−4 at higher friction.
  • Controls: Controls show that WM friction invariance persists after actor-training-horizon ablation and under domain-randomized friction, while kinematic-axis perturbations still affect the WM.The retrained actor retains the 1.32× WM friction spread, and the domain-randomization control retains a zero-containing WM slope.
  • Controls: The physics-side contrast survives with the policy in-distribution at every swept friction, supporting a kinematic-not-dynamic distinction rather than mere policy out-of-distribution behavior.The domain-randomization result bounds, but does not eliminate, the policy-OOD confound.

V. Discussion & open directions

The diagnostic is structural rather than predictive: it identifies kinematic imagination through regime insensitivity, while leaving accuracy and downstream policy quality separate questions. The discussion reports scope limits and proposes embodiment, conditioning, and control extensions.

  • Discussion: iKCE identifies kinematic imagination through absent regime sensitivity, not dynamic prediction quality or rollout accuracy.Low iKCE can indicate a kinematic predictor, whereas high iKCE with friction sensitivity indicates dynamic structure but not accuracy.
  • Discussion: Long-horizon actor training achieved ∼400 episodic reward versus ∼955 for the default h=15 checkpoint and remained unstable, without a claimed causal link.The observation is presented as consistent with the kinematic-not-dynamic hypothesis, while acknowledging multiple factors affect Dreamer training.
  • Limitations: The evidence is limited to one 2D 9-DOF walker, one open-weight world-model family, and the (z, vz) sub-slice.The authors propose testing quadruped, humanoid, driving, and additional world-model embodiments and families.
  • Limitations: The domain-randomization control bounds but does not eliminate policy-distribution confounds, while five-observation conditioning limits the regime evidence available to imagined rollouts.Closed-loop gait adaptation remains on the physics side, and the short prefix may not convey all relevant regime information.
  • Limitations: Per-step displacement decays over the horizon, so low long-horizon iKCE partly reflects reduced motion magnitude rather than purely kinematic structure.This measurement limitation qualifies interpretation of the long-horizon metric.
  • Open directions: Future tests include richer-contact embodiments, fixed open-loop actions, explicit friction or contact conditioning, and conditioning-prefix sweeps from 5–64 observed steps.These experiments would separate representational absence from architectural insensitivity and insufficient regime evidence.

A. Methodological Details

The friction-sweep diagnostic tests whether imagined iKCE responds to changing physical conditions across horizons and seeds. Results show flat WM slopes, horizon-emergent physics sensitivity, and a nontrivial WM distance from the kinematic baseline.

  • Flatness test for H2: βphys = −0.220, CI [−0.301, −0.142], while |βWM| is bounded above by ∼0.13 across one decade of friction changes.The physics interval excludes zero, whereas WM iKCE changes by at most approximately 13% per decade.
  • Horizon-emergence test: The dynamic physics signature emerges with horizon, changing from βphys = +0.012 at T=8 to −0.221 at T=64, while WM slopes remain near zero.The physics confidence interval exits the CI-contains-zero region between T=32 and T=64; the WM intervals straddle zero at every tested horizon.
  • Horizon-emergence test: The horizon test re-integrates saved T=64 per-step traces at T ∈ {8, 16, 32, 64}, requiring no new rollouts.The slopes are computed from the same friction-sweep traces, with rollout-level bootstrap confidence intervals.
  • Trivial-WM scale anchor: The measured WM lies farther from the trivial-kinematic baseline than real physics, so low absolute iKCE is not the diagnostic signature.A trivial kinematic predictor has iKCE = 0 by construction; the paper instead identifies friction invariance under perturbation as the signature.
  • Actor-training-horizon control: Retraining the checkpoint at imagination horizon 64 tests whether T=64 flatness is caused by evaluating beyond the default imagination horizon of 15.The control matches the measurement horizon while keeping seed, hyperparameters, and total training budget fixed.

1) Actor-training-horizon control.:

A domain-randomization control separates policy out-of-distribution effects from the friction-sweep evaluation. The randomized policy avoids reward collapse across the tested friction range, so the original regime boundary is specific to the fixed-friction policy.

  • Actor-training-horizon control: Domain-randomizing friction during training keeps the policy in-distribution across the sweep and prevents reward collapse anywhere in the tested range.Friction is sampled per episode from U(0.1, 1.7), and evaluation reward is 930 ± 36 across the full range.
  • Actor-training-horizon control: The µ=0.20 regime boundary does not transfer to the domain-randomized control, showing it is a property of the fixed-µ policy.The original boundary was defined from reward collapse under a policy trained only at µ=1.0.

2) Domain-randomization control.:

Domain randomization preserves the central contrast: imagined iKCE remains friction-invariant despite full-range training, while physics retains a friction response. Additional controls show this pattern is not explained by actor horizon, representation choice, or insensitivity to kinematic perturbations.

  • Domain-randomization control: The DR-trained world model remains statistically flat across the full friction range despite observing friction variation during training.Its slope is βDR_WM = −0.026, with CI [−0.123, +0.076].
  • Domain-randomization control: Under domain randomization, physics retains a strictly negative friction slope, with a point estimate roughly half the default-policy magnitude.The DR physics slope is βDR_phys = −0.114, CI [−0.201, −0.024], versus −0.220 for the default policy.
  • Per-step structure decomposition: Per-step physics residuals show friction-dependent contact spikes, whereas WM residuals show a short transient followed by a smooth, friction-invariant tail.The differing temporal structure indicates more than a magnitude difference between imagined and physical residuals.
  • Actor-training-horizon ablation: The h=64 actor reproduces the default actor’s 1.32× WM friction spread, ruling out the h=15 training horizon as the explanation.Confidence intervals overlap at every friction value.
  • Representation and actor controls: The transient-plus-smooth-tail WM structure is unchanged across actor horizons, and the friction-insensitivity reappears in the gait-DOF representation.These controls reduce concerns that the result depends on actor horizon or the identity (z, vz) view.
  • Perturbation controls: Both physics and WM respond to joint-noise perturbations, while only physics responds to the dynamic friction sweep.The joint-noise control therefore distinguishes dynamic insensitivity from general perturbation insensitivity.

C. Reproducibility

The paper releases its diagnostic pipeline, trained checkpoints, perturbation data, and figure sources, while specifying the kinematic predictors and perturbation procedures used for reproducibility.

  • Released artifacts: The released artifacts include the diagnostic pipeline, trained checkpoints, perturbation-sweep CSVs, and PGFPlots figure sources.The implementation uses the NM512 PyTorch DreamerV3 port at commit 6ef8646.
  • Kinematic predictors: The identity-view predictor uses constant-velocity extrapolation of root vertical position and velocity with a 25 ms timestep.The gait view instead embeds each joint angle on the unit circle and extrapolates each joint one step.
  • Perturbation protocol: Friction perturbations scale every geom’s MuJoCo friction at reset, while joint noise is added to observed joint positions before WM encoding.The physics state remains unperturbed by joint noise, and WM rollouts condition on the first five perturbed observations.
Loading 2607.05966v1…