Source-linked AI summary
The Intervention Gap in Latent World Models
Donna Vakalis
TL;DR
Learned world models may represent current task variables without correctly predicting how interventions change them. The paper develops capture-gated matched-intervention audits and finds that intervention fidelity is distinct from reward fit, conditional across candidates and supports, and must be evaluated directly on the native model interface.
Problem
Standard training signals do not separately certify task-variable capture or correct intervention consequences in learned world models.
Method
The paper evaluates coarse operator error and a capture-gated audit relative to declared queries, horizons, supports, and native interfaces.
Results
Intervention fidelity is neither revealed by reward fit nor ensured by task-anchored training, and failures vary across candidates, tasks, seeds, and supports.
Takeaways & Limitations
Intervention fidelity should be audited directly, capture-first, on the model surface that carries the task query, with uncertainty trusted only near training support.
Takeaways & Limitations
The evidence covers Cheetah locomotion and Finger Spin with one primary scalar query and a five-step horizon, and supports neither cross-family inference nor population-level architecture ranking.
Abstract
from arXiv · showhide
Planning-time intervention fidelity is a distinct, measurable property of a learned world model: whether the model's own open-loop transitions move task variables the way matched environment interventions do. In the settings we test, it is neither revealed by reward fit nor ensured by task-anchored training. Across released TD-MPC2 checkpoint sizes, episode return falls as an operator-error diagnostic on task observables grows, while reward-prediction error stays small and nearly flat, and a self-supervised world model trained without task signal preserves the same operator substantially better than a task-anchored model on the shared task. A capture-gated matched-intervention audit then localizes what fails. On Cheetah, three LeWorldModel checkpoints capture the current task query and support decodable real intervention effects; however, their imagined five-step effects are worse than predicting no effect and worse than an environment-endpoint oracle. The failure is task-direction rotation with excess gain, not feature collapse. This severe pattern is conditional: five PreJEPA seeds retain an oracle-relative deficit without it, Finger Spin experiments extend the deficit beyond locomotion with heterogeneous severity across seeds, and shared-bank effect geometry is both candidate- and support-dependent. We also test practice-side questions. In DreamerV3 the posterior distribution, not its sample, carries the current query; ensemble disagreement ranks error only near training support; and a frozen support-aware score degrades held-out error ranking in both tested transfer directions while native disagreement remains informative in both. We conclude that intervention fidelity must be audited directly, capture-first, on the model's native interface.
1 Introduction
The paper argues that planning-time intervention fidelity is distinct from reward fit and task-anchored training, and must be audited directly. Across several tests, fidelity failures can coexist with good capture and reward-related diagnostics, with severity depending on candidate, task, seed, and support.
- Core claim: Planning-time intervention fidelity measures whether open-loop model transitions move task variables like matched environment interventions.Standard training signals certify neither task-variable capture nor correct intervention consequences separately.
- Value-based signals: TD-MPC2 episode return falls as operator error grows, while reward-prediction error remains small and nearly flat.The operator diagnostic also ranks checkpoints differently from Bellman-residual and value-slice metrics.
- Value-based signals: A self-supervised world model preserves the shared-task operator substantially better than a task-anchored model.The comparison supports insufficiency of task anchoring for securing operator fidelity.
- Matched-intervention findings: On Cheetah, LeWorldModel checkpoints pass current-query capture and real-effect resolvability but produce poor five-step imagined effects.Their effects are worse than zero-effect prediction and the environment-endpoint oracle; the distortion is task-direction rotation with excess gain rather than feature collapse.
- Scope of findings: The severe Cheetah pattern is conditional: PreJEPA retains an oracle-relative deficit without the same conjunction, while Finger Spin shows heterogeneous severity across seeds.Shared-bank effect geometry is candidate- and support-dependent rather than an architecture ranking.
- Auditing in practice: DreamerV3’s posterior distribution carries the current query, while disagreement is informative near training support and native disagreement remains informative across tested transfers.A frozen support-aware score degrades held-out error ranking in both transfer directions.
2 Two Audits of Planning-Time Fidelity
The paper presents a coarse operator-error instrument and a finer capture-gated audit for evaluating planning-time intervention fidelity. The protocol fixes the task query, horizon, support, and native interface, then separates state capture, real-effect resolvability, and model propagation.
- Shared evaluation setup: The audit fixes a query family, physical horizon, and intervention-support distribution before evaluating any candidate model.All diagnostics remain relative to these choices and the model’s native interface.
- Operator-error audit: The coarse instrument decodes task queries from latent states, rolls the model forward natively, and compares predicted outcomes with real environment outcomes.It reports probe quality separately from operator error, avoiding silent attribution of poor decodability to transitions.
- Capture-gated audit: Gate 1 tests whether the task target is readable from held-out real planning states, and Gate 2 tests whether matched real effects are resolvable at endpoints.Propagation is interpreted only for cells passing both prerequisites.
- Two audit resolutions: The coarse score can conflate missing task information with incorrect latent transitions, motivating the finer instrument’s separate tests.The finer audit decomposes fidelity into capture, real-effect resolvability, and propagation.
- Capture-gated audit: Gate 3 compares native imagined effects with real effects, persistence or zero effect, and the environment-endpoint oracle.The model and environment execute matched action sequences from the same real roots.
- Interpretation: The protocol’s finite-horizon error decomposition makes propagation interpretable only when root capture error is small.The decomposition distinguishes capture error, closure-restricted transition error, and approximation residuals without guaranteeing realized return.
- Calibration and inference: Calibration uses prespecified synthetic acceptance cases, while grouped disjoint splits and nested cross-validation isolate environment-side targets, readouts, and evaluation.The controls include shuffled labels and features, horizon-zero identity, persistence, executed-environment endpoints, wrong actions, and action-null conditions.
3 The Value Channel Neither Reveals Nor Ensures Operator Fidelity
Value-based signals neither reveal nor ensure intervention fidelity: operator error diverges from reward and Bellman diagnostics, and task anchoring does not prevent it.
- Return and reward fit: −0.90 rank correlation links rising operator error with falling episode return across five dependent TD-MPC2 sizes.The association remains −0.89 after controlling for reward-prediction accuracy.
- Return and reward fit: [0.03, 0.09] is the flat reward-prediction-error range despite the same sweep’s operator failures.The largest checkpoint combines near-best reward fit with worst operator error and planning return.
- Comparison with value metrics: +0.30 rank correlation shows weak agreement between full-observable operator error and Bellman-residual rankings, while the value-restricted variant reaches +1.00.The operator score adds information from task-observable directions not constrained by the value channel.
- Limits of the coarse diagnostic: The coarse diagnostic cannot distinguish absent latent information from information moved incorrectly and aggregates over support without matched interventions.The finer capture-gated audit is therefore needed to separate these failure sources.
- Task anchoring: 0.457 ± 0.011 is the operator-error ratio for self-supervised LeWorldModel runs versus task-anchored TD-MPC2 on shared observables.The comparison remains separated under single- versus multi-task anchoring, restricted observables, and an MLP probe.
4 Capture Without Propagation
The capture-gated audit shows that Cheetah LeWorldModel checkpoints can capture current task information and real effects while severely mispredicting their own imagined five-step effects.
- Prerequisites: Three frozen LeWorldModel checkpoints pass current-query capture and real-effect resolution on matched five-step Cheetah interventions.Each evaluation family uses three roots and six matched actuator contrasts, with thresholds fixed on 48 calibration families before scoring 96 evaluation families.
- Propagation failure: Predicted-effect R2 is far below zero for every checkpoint, making imagined effects worse than predicting no effect.Every whole-family 95% interval lies below zero.
- Propagation failure: Every checkpoint’s imagined endpoint is worse than the environment-endpoint oracle under the same readout.Thus, current-query capture and real-effect resolvability coexist with severe five-step transition distortion.
- Rotation with excess gain: 0.85–0.89 contrast cosine shows strong alignment of complete predicted feature changes with real changes, while task projections are nearly orthogonal and 1.7–2.6 times too large.The descriptive characterization points to task-direction misalignment with excess gain rather than wholesale feature collapse.
- Scope: The primary result does not establish nonlinear-decoder absence of task information, other horizons or supports, planner failure, return effects, or generalization beyond the tested roster.These are explicit scope boundaries for the dissociation.
5 The Deficit Replicates; Severe Distortion Does Not
The oracle-relative propagation deficit recurs across tested rosters and tasks, but severe below-zero distortion is conditional on candidate, task, and seed rather than universal.
- Boundary results: The oracle-relative deficit reappears in every tested roster, while severe distortion remains conditional on model family, task, and seed.The paper distinguishes the recurring deficit from the sharper severe conjunction.
- PreJEPA non-replication: Five PreJEPA predictors pass capture, real-effect resolution, and oracle-relative comparison, but none meets the severe below-zero condition seen in LeWorldModel.This is a non-replication of severe distortion, not evidence of accurate propagation.
- Finger Spin: Three Finger Spin LeWorldModel runs all show positive predicted-minus-oracle error, while only one satisfies the severe below-zero certificate.The task uses endpoint negative hinge velocity, and raw scores are not comparable to Cheetah.
- Shared-bank geometry: Shared evaluation on identical roots, five-step actions, and endpoints yields no unanimous architecture signature.The figure treats candidate labels as row-specific and does not establish an architecture ranking or size law.
- Conclusion: Capture and propagation remain separable, while failure severity and geometry depend on candidate, task, seed, and support.The sharp Cheetah classification is valid on its frozen support but is not support-invariant.
6 Auditing in Practice: Surfaces, Uncertainty, and Support
The auditing results identify which representations and uncertainty signals are useful, and show that support-aware corrections require transfer validation rather than automatic trust.
- Surfaces: The posterior distribution is the only tested DreamerV3 surface that carries the current query across checkpoints.Centered posterior logits pass both capture gates; sampled latents, their mode, the deterministic state, and their concatenation do not.
- Uncertainty: Within PreJEPA, disagreement ranks future-feature error under task-policy support (ρ = 0.678) but is inconclusive under environment-random support (ρ = 0.027).The support-linked relation deteriorates along both root-context and future-action departure axes.
- Support geometry: Shared action-effect geometry uses matched-effect cosine across supports, but every architecture-by-support summary remains inconclusive.Hatched cells are uncertified when capture or real-effect resolvability fails, and TD-MPC2 is descriptive only.
- Support transfer: A frozen equal-weight support augmentation improves within-family ranking but degrades held-out error ranking in both tested transfer directions.Native disagreement remains informative, whereas support distance alone is much weaker.
- Gate order: In a matched-shell TD-MPC2 audit, task-policy roots passed capture but real-effect resolution failed, while environment-random roots failed absolute capture.No propagation cell was interpretable in either direction.
- Audit protocol: The proposed protocol declares query, horizon, and support first; qualifies real roots and endpoints; then compares imagined propagation with persistence, zero effect, and the environment-endpoint oracle.It reads the model's native interface and uses the distribution rather than a sample when the state is distributional.
7 Related Work
Related work frames the paper at the intersection of latent world-model planning, value-equivalence, rollout error, structural evaluation, operator methods, probing, uncertainty, and self-supervised representations.
- World models for planning: Latent world models expose encoder states and open-loop transitions directly to planners or imagined-rollout learners.The paper evaluates this planner-exposed interface rather than only training loss or realized return.
- Value equivalence: Value-aware model learning restricts accuracy to planning-relevant classes, but one-step likelihood can correlate poorly with control performance.The paper's distinction is operational: it tests capture and intervention propagation against externally constructed targets.
- Rollout error: Prior rollout-error work studies compounding error, unstable sample rollouts, or model exploitation, whereas this protocol measures matched intervention propagation after capture and real-effect gates.The comparison concerns evaluation target and gating, not a claim that prior work measures the same quantity.
- Evaluation beyond return: Structural evaluation work motivates testing transition coherence and interventional reasoning beyond return, while this paper adds continuous-control matched interventions with explicit gating and oracle comparators.Its interventions are matched in the environment and scored at executed endpoints rather than compared only across model rollouts.
- Operator theory: Koopman-style language is used to estimate finite-horizon environment-side closure, not to claim global linearity or spectral recovery.This limits the operator-theoretic interpretation to the measured observable family and horizon.
- Probing: Probe methodology motivates separating representational readability from behavioral use, so capture is paired with an independent propagation test.Readability alone is not treated as evidence that the model uses the representation correctly.
- Uncertainty and support: Ensemble disagreement is established as an uncertainty and off-support rollout signal, while its reliability can degrade under distribution shift.The paper tests this support dependence directly rather than assuming disagreement is uniformly calibrated.
- Self-supervised world models: LeWorldModel and PreJEPA instantiate self-supervised latent-dynamics candidates evaluated as frozen models rather than proposed representation or training methods.Their role here is comparative auditing within the self-supervised world-model family.
8 Scope and Limitations
The evidence is limited to specific tasks, model families, queries, supports, and correlational analyses. Capture and support-aware conclusions also depend on fixed readouts, calibrated intervention shells, and untested decision value.
- Scope: The learned-model evidence covers only Cheetah and Finger Spin, each with one scalar query and a five-step horizon.The eligibility survey found five-run PreJEPA as the only family passing every original primary criterion, so the evidence does not support cross-family inference or population-level architecture ranking.
- Scope: The TD-MPC2 size-sweep association is correlational, based on five dependent checkpoints from one training recipe, and is not a selection rule.The anchoring comparison also uses two families, one anchored seed, different modalities, and different native prediction modes.
- Measurement boundaries: Capture is relative to a fixed readout family and does not establish availability under every nonlinear decoder or planner use of decoded information.Propagation fidelity likewise need not imply high return, and the experiments provide no checkpoint-selection result, return improvement, or guarantee.
- Measurement boundaries: The environment-side closure is synthetically calibrated, qualifies only on matched action shells, and has untested candidate-model decision value.A first model-side shell audit terminated at the capture gates.
- Interpretation: Support-reference analyses remain correlational, so observed support movement and disagreement–error associations do not identify a causal mechanism.The geometric descriptions constrain candidate mechanisms without selecting one.
9 Conclusion
The paper defines intervention fidelity as a distinct property of learned world models and shows why it requires direct, capture-first auditing on the native interface. Its evidence supports a qualified conclusion rather than a universal architecture claim or return guarantee.
- Conclusion: Planning-time intervention fidelity measures whether open-loop model transitions move task variables like matched environment interventions.The property is treated as a relation among the model, query, horizon, intervention support, and native interface.
- Conclusion: Fidelity can fail despite passing current-state capture, real-effect resolvability, and overall feature-alignment checks, so it must be audited directly.The audit is capture-first and reads the surface that actually carries the query.
- Conclusion: The propagation decomposition separates root-capture error, accumulated transition error, and target-approximation error, but does not bound policy discontinuity or realized return.The bound therefore diagnoses measured error sources rather than guaranteeing control performance.
- Conclusion: The synthetic suite passes ten designed cases, validating the procedure and stopping rules rather than establishing learned-model adequacy or empirical-closure value.
C Per-Candidate Results
Per-candidate results show heterogeneous, task- and seed-dependent intervention behavior: severe distortion is not universal, oracle-relative deficits recur, and candidate/support summaries remain separate.
- PreJEPA: None of five PreJEPA candidates meets the strict below-zero distortion condition, although accurate propagation is not established.
- Finger Spin: Run 01 meets the severe below-zero certificate, while Runs 02 and 03 retain positive predicted-effect signal; all three show an oracle-relative deficit.
- LeWorldModel: Every held-out LeWorldModel run has a strictly negative combined-minus-native interval, so the reciprocal complete-role label is not driven by one seed.
- Support analysis: The support-departure evaluation uses 384 independent families over a 4 × 4 context–action grid, with technical replicate runs treated as paired repeats.
- Evaluation design: Candidate models remain separate during evaluation, with disjoint identities for target construction, probing, calibration, and testing plus whole-family bootstrap resampling.
E Dreamer Interface Boundaries
Dreamer’s current-query information is localized to the categorical posterior’s centered logits, while environment-side qualification and model-side shell audits expose important boundaries in the evaluation protocol.
- Current-query localization: Centered posterior logits are the sole tested Dreamer surface passing both current-query coordinate gates at every checkpoint.Posterior probabilities pass only at scales 16 and 64; sampled states, modes, deterministic recurrent states, and their tested concatenation pass none.
- Current-query localization: Exact-logit instruments pass across four free_nats runs, but no sampled budget passes both coordinates for all runs despite monotone approximation improvement.
- Closure qualification: The original Cheetah closure qualification returns rank zero on both supports, leaving later target, probe, candidate-score, and evaluation assignments unused.
- Closure qualification: The rank-zero result cannot be read as absent action-conditioned structure because action effects and their H = 1 through H = 5 ordering persist despite weak common linear predictability.The independently sampled action slots functioned as exchangeable identifiers, and physical dimensionality was not identifiable from their SVD.
- Matched-shell audit: The frozen full-quadratic estimator qualifies on matched fixed-radius shells but fails over the all-radii action ball, while paired-radius testing does not confirm task-policy radius conditioning.
- Matched-shell audit: All first model-side matched-shell audit cells terminate at capture gates, with endpoint-effect capture failing for task-policy roots and absolute capture failing for environment-random roots.
G Inferential Breadth
The evidence supports only narrow, conditional inferences: one PreJEPA family meets the original primary criteria, while TD-MPC2 findings remain exploratory and support-sensitive. Environment-side shell analyses add qualified localization results but do not confirm the primary radius-conditioning contrast.
- Eligibility survey: Only the five-run PreJEPA family satisfies every original primary criterion, so the survey supports no cross-family inference, monotonic size law, or model ranking.Four PreJEPA runs are primary and one is an extra run; five TD-MPC2 sizes remain descriptive examples rather than independent training runs.
- Environment-side localization: The frozen full-quadratic estimator fails the all-radii action-ball gate but qualifies on matched shell tests across both supports.The paired inner/outer shell test qualifies both shells, yet the primary task-policy radius-conditioning contrast is not confirmed.
- Exploratory TD-MPC2 evidence: The exploratory TD-MPC2 panel contains 80 complete scores, with performance varying by candidate, model size, and starting-state support.The 317M model combines accurate capture with extreme predicted action-effect error, while a 19M model beats persistence on policy support but not environment-random starts.
- Exploratory TD-MPC2 evidence: The TD-MPC2 contrasts establish support sensitivity within a dependent released-size panel, not confirmatory architecture evidence.The same candidate changes relative to persistence across policy-support and environment-random starting states.