Source-linked AI summary

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

Malo de Pastor

arXiv:2608.14650v1cs.LGcs.RO

TL;DR

The paper asks whether a Medium-derived prediction interface can identify when switching to a frozen Full predictor improves decisions enough to justify sequential overhead. Using paired exact-reset evaluation and audits, it finds useful selection gains over fixed and prediction-withheld references, including a stronger current-DINO control.

  • Problem

    World-model predictions can change which candidate actions matter across states and tasks, motivating evidence on when additional predictive computation improves decisions.

  • Method

    The paper uses a paired exact-reset physical evaluation and audit protocol with frozen Medium/Full predictors, shared candidate actions, and thresholded prediction-interface routing.

  • Results

    The tested prediction interface beats fixed and prediction-withheld references on PushT and a dimension-matched current-DINO router, with favourable effects for all three checkpoint pairs.

  • Takeaways & Limitations

    Under paired exact-reset outcomes, the tested Medium-derived interface supports useful post-hoc selection between two frozen predictors.

  • Takeaways & Limitations

    The study does not establish compute saving, closed-loop value, prediction causality, sufficiency, or generality across world-model families.

Abstract

from arXiv · show

Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. We study a narrower question: under paired exact-reset physical outcomes, can a Medium-derived interface predict when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential overhead? Our contribution is a paired evaluation and audit protocol, not a new generic routing rule: all candidate actions are executed from the same reset state, Medium and Full act on the same candidate set and task, and their paired physical-loss difference defines the routing target. On a fresh PushT bank (V106; 1,600 states, 39 tasks, three checkpoint pairs), a frozen prediction-interface router lowers overhead-inclusive decision cost relative to standalone Medium, standalone Full, and a latency-advantaged task-only router. We then prospectively seal a second 1,600-state PushT confirmation (V107) against a stronger current-state control using the task, a dimension-matched projection of current DINO features, and all five candidate actions, with no DINO encoder latency charged. The prediction interface lowers priced physical decision cost by 0.002549 (state-clustered 95% interval [-0.002867, -0.002238]; one-sided 95% upper bound -0.002286), with negative effects for all three checkpoint pairs. A controlled-PyBullet audit independently supports a composite task-prediction-regime router. The sequential router remains slower than fixed policies, and its advantage is restricted to low compute prices. The evidence supports incremental routing information in the tested prediction interface beyond one deliberately favoured current-DINO control, but not causal sufficiency, compute saving, closed-loop value, or cross-family generality.

1 Introduction

The section frames whether Medium’s predicted consequences can identify when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential overhead. It proposes a paired exact-reset evaluation that distinguishes decision-value gains from compute saving.

  • Motivation and question: The study asks whether a Medium-derived interface predicts when switching to Full reduces task-specific decision loss enough to offset sequential overhead.The question is evaluated under paired exact-reset physical outcomes.
  • Information interface: The prediction-interface router uses the task and Medium’s predicted consequence, while excluding true futures, oracle regret, Full prediction, and privileged simulator state.Exact physical outcomes supervise and evaluate the router offline but are unavailable at routing time.
  • Evaluation protocol: Exact resets evaluate Medium, Full, and all candidate actions on identical physical queries, making realized escalation benefit observable offline.At test time, the sequential policy evaluates Medium before deciding whether to pay for Full; the task-only control routes before Medium and is latency-advantaged.
  • Baselines: 99.1% of the latency-advantaged task-only control’s cases route to Full, motivating comparison against a control that is not artificially weakened.The passage identifies this routing behavior for the V106 control.
  • Interpretation: The learned router reduces priced physical regret enough to beat both fixed policies, yet remains slower in raw milliseconds.The objective includes Medium, Full, task-evaluation, and router overhead, so the result supports selective decision value rather than compute saving.

2 Decision ledger for Medium-to-Full escalation

The decision ledger defines escalation as a paired, exact-reset comparison of Medium and Full candidate-action outcomes, with positive benefit meaning that priced escalation reduces decision cost. It separates conditional optionality from sequential duplication cost and learned-selection error, while treating the task-only comparator as latency-advantaged rather than an estimator of the population identity.

  • Paired decision ledger: Exact reset reveals every candidate’s physical consequence, so Medium and Full losses are observed on the same query and paired across PushT and PyBullet states.The protocol makes no claim about actions outside the candidate set or real-robot causal effects.
  • Bayes escalation accounting: Positive B means escalation reduces the priced decision objective, and the Bayes rule escalates when its conditional benefit exceeds the incremental Full cost.The resulting value measures conditional escalation improvement before comparing always-paid sequential latency with standalone policies.
  • Bayes escalation accounting: Sequential routing beats standalone Full only when optionality exceeds the signed latency gap κG; in both protocols κG > 0, creating a duplication penalty.This accounts for Medium’s always-paid interface cost plus additional Full evaluation relative to standalone Full.
  • Nested interfaces: Adding Medium’s predicted consequence can strictly increase population information value when coarse outcomes contain both positive- and negative-benefit regions.This is an information statement, not a guarantee that a finite-sample learner will exploit the interface.
  • Deployment interpretation: Learned-router gains decompose into conditional optionality, acquisition or duplication cost, and imperfect selection, while the sealed task-only control remains a latency-advantaged deployment comparator.The task-only control routes before Medium and is not a plug-in estimator of the nested-interface identity.

3 Paired physical routing

The section defines paired physical routing by evaluating Medium and Full on identical candidate sets from exact-reset states, making their realized loss difference the routing target. It also specifies the sequential overhead rule and limits interpretation to the audited simulator estimand rather than unrestricted causal generalization.

  • Paired evaluation: Exact resets and shared candidate action sets make Medium–Full physical-loss differences paired rather than confounded by state sampling.Each task scores predicted consequences, selects an action, and evaluates physical regret using the realized consequence.
  • Inference limits: The audit’s finite candidate-set simulator intervention does not establish unrestricted robotics or real-world causal-generalization results.Population confidence statements require a declared physical-state sampling law or valid design, and three fixed checkpoints define only a finite equally weighted checkpoint estimand.
  • Routing rule: The V106 router regresses the paired gross regret difference ∆r = LM,r −LF,r and converts it to a net-benefit rule by thresholding at λdF,r when calibrated.The derivation relies on constant measured incremental Full latency within each checkpoint pair; the fitted score is not assumed perfectly calibrated.
  • Routing rule: At audit time, Medium is always evaluated, while escalation evaluates Full and uses its action, with sequential latency sG + πdF.The reported objective uses λ = 0.002, and candidate-minus-reference differences are negative when the candidate is better.

4 Frozen audit evaluations

The frozen audits evaluate routing interfaces on sealed PushT and controlled-PyBullet banks under exact-reset, multi-candidate decision protocols. They preserve fixed models, tasks, interfaces, thresholds, and latency-measurement procedures while testing finite-state capacity reversals across physical settings.

  • PushT: 1,600 fresh PushT states form a sealed audit set with five candidate actions, horizon 30, 39 tasks, and three checkpoint pairs.Shards 8–9 were never opened during router or threshold selection.
  • PushT: V107 adds two non-overlapping 800-state shards with the same 39 tasks, five candidates, horizon, and three frozen checkpoint pairs.The prediction and current-DINO/action routers were selected during development, then copied and hashed before V107 generation.
  • Controlled PyBullet: Controlled PyBullet freezes six 2,000-state banks spanning identity replication and prespecified changes in horizon, velocity, obstacle count, and their composition.The audit also crosses nine spatial targets and nine loss geometries, including safety-penalized losses.
  • Capacity reversals: Medium has lower decision cost on 51.9% of controlled PyBullet states and 45.1% of PushT states, with Full winning on the complements.These realized reversals do not establish Blackwell incomparability or value available to the deployed interface.
  • Latency protocol: All costs use one physical state, one downstream task, and five candidate actions as the decision unit, with world-model forwards measured on NVIDIA H100 GPUs.PushT latency uses 200 warmups, seven rounds of 1,000 repeats, 512 sampled queries, and the median of round medians; p90 accounting is a sealed sensitivity.

5 Results

Across fresh PushT and controlled PyBullet evaluations, prediction-derived routing reduces priced decision cost and improves physical outcomes relative to fixed and task-only references under the tested low compute price. However, the router remains slower than fixed policies, is not near-oracle or compute-saving, and loses its advantage at higher prices.

  • V107 PushT confirmation: 0.002549 lower priced physical decision cost than the dimension-matched current-DINO/action control on fresh V107 PushT.The state-clustered 95% interval was [−0.002867, −0.002238], with negative effects for all three checkpoint pairs.
  • Mechanism of the priced gain: 0.002617 lower regret offset the router’s 0.000068 priced latency disadvantage versus the current-DINO/action control.Prediction routing also escalated less often: 47.2% versus 66.2%.
  • Controlled PyBullet audit: 0.002800 and 0.003668 lower candidate-minus-margin and candidate-minus-random costs, respectively, in controlled PyBullet.The audit supports selective routing for the composite prediction–task–regime interface but does not independently isolate prediction from regime information.
  • Fixed-policy comparisons: 0.004308 and 0.002577 improvements over Fixed Medium and Fixed Full, respectively, on PushT.Controlled PyBullet improvements over the corresponding fixed references were 0.003677 and 0.003746.
  • Operational limits: 0.651 ms router latency exceeded 0.281 ms for Medium and 0.364 ms for Full on PushT, despite lower mean regret and decision cost.The router’s decision cost was 0.03037 versus 0.03411 for Medium and 0.03222 for Full, while escalating 42.0% of query–task pairs.
  • Compute-price sensitivity: 0.01 was the highest compute price through which the candidate remained favorable; it lost to every listed reference at 0.025 and 0.05.At λ = 0.01, its difference versus Fixed Full was −0.000385, so the advantage is restricted to low compute prices.

6 Related work

Prior work studies predictive planning, adaptive computation, deferral, decision calibration, and informativeness, whereas this paper uses prediction interfaces to select between frozen computations. Its experiments do not claim causal sufficiency or formal statistical-experiment comparisons.

  • World models and predictive planning: DINO-WM predicts future DINOv2 patch features for zero-shot planning, while this work uses frozen semantic features to select between frozen computations.Decision-focused learning instead trains predictions for downstream optimization loss.
  • Adaptive WAM inference: Adaptive world-action-model systems condition extra computation on predicted visual futures or intermediate states and evaluate task success or latency.The cited examples include gated geometric Best-of-N and SANTS.
  • Adaptive cascades, deferral, and information value: Adaptive networks and deferral methods use cheap-stage signals or conditional risk differences to decide continuation or selection among fixed experts.Confidence-based cascade analysis and two-stage acquisition further study when additional information is worthwhile.
  • Calibration and statistical experiments: Decision calibration evaluates predictions through induced decision losses, while Blackwell, Le Cam, and deficiency frameworks formalize informativeness.The current experiments do not estimate a garbling kernel, deficiency, or a separating class of all decision problems.

7 Limitations and conclusion

Under paired exact-reset outcomes, the tested Medium-derived prediction interface supports post-hoc selection between two frozen predictors and beats the study’s fixed, prediction-withheld, and current-DINO references. The evidence remains limited: it does not establish prediction causality, compute saving, closed-loop value, or generality across world-model families.

  • Limitations: The evidence does not establish prediction causality, sufficiency, or superiority to all current-state encodings, architectures, or compute-matched ensembles.The study also does not establish compute saving, closed-loop value, or generality across world-model families.
  • Limitations: The sequential prediction router is slower than fixed policies, and the original PushT sensitivity supports only a low-price operating region rather than a continuous frontier.Benefit prediction is imperfect, the clairvoyant lower-bound gap remains substantial, and PyBullet validates a composite task–prediction–regime interface.
  • Conclusion: Under paired exact-reset outcomes, the Medium-derived prediction interface supports useful post-hoc selection between two frozen predictors.The predictors are related DINO-style calibrated interfaces rather than genuinely different world-model families.
  • Conclusion: The interface beats fixed and prediction-withheld references on the first fresh PushT bank and a dimension-matched, action-conditioned current-DINO router on a second prospectively sealed bank.The current-DINO comparison charged no encoder latency, and all three fixed checkpoint pairs showed the reported result.
  • Scope: The contribution is a paired physical evaluation and audit protocol, not a new generic routing theory.The study’s action sets and downstream loss families are controlled, and none of the evaluations is a closed-loop robotics deployment.

A Proofs · A.1 Proof of the standard Bayes escalation result (Proposition 1)

The proof establishes the standard Bayes escalation result by conditioning on G, identifying the pointwise escalation rule, substituting it into the objective, and taking expectations.

  • A.1 Proof of the standard Bayes escalation result (Proposition 1): Conditioning Eq. (2) on G provides the proof’s starting point.
  • A.1 Proof of the standard Bayes escalation result (Proposition 1): Thus, the proof proceeds from conditional analysis to pointwise minimization and finally to an expectation-level conclusion.
  • A.1 Proof of the standard Bayes escalation result (Proposition 1): Because π ∈{0, 1}, the pointwise minimizer escalates exactly when bG > 0.
  • A.1 Proof of the standard Bayes escalation result (Proposition 1): The proof then substitutes the pointwise minimizing escalation rule into the conditioned expression.
  • A.1 Proof of the standard Bayes escalation result (Proposition 1): This substitution converts the pointwise decision characterization into the corresponding minimized expression.
  • A.1 Proof of the standard Bayes escalation result (Proposition 1): Taking expectations completes the argument for the stated Bayes escalation result.

A.2 Proof of the accounting identity (Proposition 2) … E V107 fixed-threshold price sensitivity

The paper formalizes sequential-routing cost identities, regret bounds, representation cautions, and auditability conditions, then reports sealed V107 and prior-family comparisons under state-clustered inference. V107’s fixed-threshold result favors prediction over the current-DINO/action control at λ = 0.002, without establishing dominance over fixed policies.

  • A.2 Proof of the accounting identity (Proposition 2): Sequential-cost identities derive the excess costs of Fixed Medium, always-escalating routing, and standalone Full relative to the Bayes router.Always escalating costs LF + λ(sG + dF), while standalone Full removes λκG from that cost.
  • A.3 Proof of the nested-interface Jensen identity (Theorem 1): Nested-interface Jensen analysis shows that conditioning on a richer prediction interface cannot reduce the expected positive-part term, with strictness when both conditional positive and negative parts occur.The conditional Jensen gap is min(P0, N0), and deployment comparison subtracts interface-overhead terms.
  • A.4–A.6 Sign-regret results: Sign-regret results identify excess cost as |bG| on sign mismatches and bound it by the prediction error |bb − bG|, with corresponding deployment-cost corollaries.When prediction error is at most ε, sign errors imply |bG| ≤ ε; zero error yields zero weighted excess cost.
  • A.7–A.8 Threshold and fixed-rate results: The scalar-ordering obstruction shows that a single threshold must incur positive excess cost when Bayes-escalation regions are separated in the wrong order, while fixed-rate routing is optimized by thresholding bG.For rate q, the threshold rule maximizes E[bGg], exceeding independent routing by Cov(bG, g⋆) ≥ 0 unless bG is almost surely constant.
  • A.9–B Auditability and representation limits: Complete executed records make frozen-model decisions and realized policy costs deterministic and finite-audit means exactly computable, whereas the public package releases only policy-facing arrays.A measurable recoding can preserve all unrestricted Bayes risks while changing coordinate-dependent geometry, so latent distances are not a universal comparison currency.
  • C.1 Information restrictions: Information restrictions distinguish prediction-based routing from the deliberately favored V107 current-DINO/action control, while PyBullet additionally uses predicted physical consequence and regime.The PushT prediction candidate routes only after mandatory Medium computation; V107’s control routes before Medium and uses a 43-dimensional PCA-compressed current-DINO block.
  • C.2–C.3 Sealing and executed estimands: V107 seals its protocol, router hashes, projection, thresholds, latency components, implementation, primary comparison, and bootstrap seed before physical outcomes, using state-clustered estimands over 1,600 states, 3 checkpoints, and 39 tasks.Inference averages tasks and checkpoint effects within state before resampling states; the one-sided quantile is 0.95 for the single primary comparison.
  • D.1–D.2 Complete primary comparisons; E V107 fixed-threshold price sensitivity: -0.0025 [−0.0029, −0.0022] is the PushT V107 current-DINO + actions result, while prior comparisons report -0.0043 for Fixed Medium, -0.0026 for Fixed Full, and -0.0026 for the task-only input router.At λ = 0.002, negative values favor prediction over the current-DINO/action control, but the comparison does not assert dominance over fixed policies.

F Claim guardrails

The evidence supports a priced-decision improvement, not compute savings or a universally superior expert. Claims are limited to the tested aligned interfaces, evaluated decision families, and related calibrated physical settings.

  • The router lowers a priced decision objective but remains slower than both fixed policies, so it does not establish compute savings.The router’s overhead is part of the priced objective, while its sequential execution remains slower than fixed policies.
  • Full is not uniformly better: its value is state- and task-dependent, both fixed capacities win on substantial subsets, and the models are Blackwell-incomparable.Capacity ordering reverses empirically for the evaluated decision families.
  • The router is not near oracle because a substantial gap remains to the realized-loss clairvoyant lower bound, which uses unavailable audit outcomes and zero router overhead.This lower bound is therefore not an attainable routing policy.
  • The tested aligned prediction interface beats one dimension-matched current-DINO/action router on one prospectively sealed PushT bank, but this does not establish arbitrary world-model generality.Support comes from two controlled physical mechanisms and two fresh PushT banks under related calibrated DINO-style interfaces.

G Reproducibility and archival note

The release candidate archives hash-verified frozen sources, sealed protocols, and reusable audit implementations while preserving the executed run code. It supports independent auditing, but full physical-outcome regeneration still depends on external assets.

  • Archived layers: The release candidate separates hash-verified frozen sources, compact sealed protocols and summaries, and a reusable implementation of decision-value, routing, bootstrap, and lineage checks.The reusable implementation was not retroactively substituted for the executed code.
  • Archived layers: The archive includes the executed V104D and V104D2 scripts, reports, router checkpoints, and threshold material.
  • Reproducibility limits: The package makes statistics, frozen router lineage, and source semantics independently auditable, but full physical-outcome regeneration remains an external-asset task.Missing assets include heavy raw observations, all-candidate physical outcome tensors, DINO feature banks, upstream world-model checkpoints, and the original HPC environment.
Loading 2608.14650v1…