Source-linked AI summary

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

Javier Aguilar Martín

arXiv:2608.17956v1cs.LGcs.AIeess.SY

TL;DR

The paper asks whether sampling verification can certify continuous world models when rare, localized modes are omitted. It analyzes the danger law and finds catastrophic exploitation of accepted mode-blind models, with repair succeeding for 1D clamps but failing for 2D region modes.

  • Problem

    The paper examines whether a sampling gate can accept a model that is exact outside a small continuous-state mode region yet catastrophically exploited by its planner.

  • Method

    The paper combines an expected-risk danger law, localization bounds, planner exploitation tests, LLM repair experiments, and version-space identification certificates.

  • Results

    105 of 111 1D clamp draws were repaired, while 2D region recovery was 0/156 and accepted mode-blind models were exploited at nearly the whole attainable return.

  • Takeaways & Limitations

    Sampling acceptance certifies sample consistency but not safety against localized discontinuous modes, while repair is strongly dependent on the mode geometry.

  • Takeaways & Limitations

    The empirical repair failure is scoped to the tested 2D region treatments, while other prompts, longer contexts, refinement methods, and fitting tools remain untested.

Abstract

from arXiv · show

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1-r)^N; an independent acceptance sample adds its budget to the exponent. On three hybrid instruments the accepted mode-blind model is exploited: the planner is pinned at the mode boundary at a regret of nearly the whole attainable return. We prove a localization budget, valid at boundary points: models with Lipschitz constant at most L differing by eta at a point disagree above tolerance eps on a region of volume at least kappa((eta-eps)/L)^(d+m); the discontinuous reset modes studied pay no such budget. With real LLM synthesis, GPT-5.x repairs an omitted 1D clamp in 105 of 111 mode-containing draws -- every attempt exact on 50 of 56 instrument-stream blocks (95% CI [0.781, 0.960]). On 2D regions no artifact recovers the rule (0/156); eight targeted interventions leave the failure in place, and positive controls locate it: a located rule is not induced, while given form and location the constants follow exactly. A version-space certificate proves identification is class-relative: at the widest dose the declared fit succeeds in 20/20 blocks and every sample-consistent circle is within tolerance in 18/20. We prove a class of entry rules exactly consistent with every sample yet harmless at play, so identifiability is a measurable property of the instrument. Re-scoring all 1034 artifacts on independent samples confirms acceptance certifies sample consistency and no more: where the gate is provably informative it covers about two percent of the exploited planner's queries.

1 Introduction · 2 The instruments · 3 Theory: what transfers, what changes, what is new

The paper shows that sampling verification can accept continuous world models that are exact off rare mode regions yet catastrophically exploited by planners. It transfers the exact gate-miss and identifiability results, adds localization bounds for Lipschitz models, and tests how instrument geometry changes synthesis and certification.

  • 1 Introduction: Accepted mode-blind continuous models can pass sampling gates while planners exploit their omitted rare modes catastrophically.The continuous question is whether a model can pass the gate, remain exact outside a small mode region, and still be exploited by its trusting planner.
  • 2 The instruments: The three instruments are cart-with-wall, pendulum-with-stop, and a 4D bi-modal PatchField2D, evaluated with random-shooting MPC and CEM.The study uses one fixed configuration per base planner family and a common harness without per-instrument planner recalibration.
  • 1 Introduction: 105 of 111 mode-containing draws were repaired exactly by GPT-5.x on 1D clamps, with 50/56 instrument–stream blocks exact and 95% CI [0.781, 0.960].On 2D regions, no artifact recovered the rule in any treatment, so the synthesis result is scoped to the measured geometry.
  • 2.1 Cart-with-wall: The cart’s omitted inelastic wall resets the next state to (xwall, 0), while the blind model removes that branch and is bit-exact off-mode.Its rarity knob sweeps mode probability from 0.317 to 0.0024 as xwall ranges from 2 to 10, and the blind planner drives toward the wall while the truth planner goes left.
  • 2.2 Planner and play cost: Random-shooting MPC replans from sampled action sequences, and its true-environment returns compare the truth planner, model planner, and uniform-random policy on paired seeds.The truth planner is a benchmark rather than a proved optimal policy.
  • 2.3 Three design conditions on a continuous danger instrument: The instruments require rare mode firing, exploitable omission, qualitatively different truth-planner behavior, plateau rewards, and a drag time-constant inside the planning horizon.I.i.d. candidate sampling can make truth and blind models indistinguishable in imagination, while piecewise-constant blocks and constant candidates prevent that failure.
  • 3 Theory: what transfers, what changes, what is new: For any measurable critical event of probability r, N i.i.d. gate rollouts miss it with exact probability (1-r)^N; two-mode misses instead use the union event and sharp Fréchet–Hoeffding bounds.The product form cannot be restored by a fixed correction because measured dependence changes sign across the knob grid.
  • 3.1 The estimand: danger as an expected risk, and which factor of it is exact: Lipschitz models differing by η at one state-action point must disagree above ε on a region with volume scaling as ((η−ε)/L)^(d+m), including boundary constants, whereas discontinuous reset modes pay no such localization budget.The bound applies in joint state-action space and shows that exact localization at fixed amplitude and vanishing volume requires unbounded L.

4 The mechanism and the threshold law

The danger law follows the exact risk factor play cost × (1−r)^N: it is negligible inside the random envelope, then rises through a threshold and reaches full play cost. Across cart, pendulum, and PatchField2D instruments, mode-blind planners exploit omitted boundary modes through high query reach, while comparisons against random depend on reward-tail behavior.

  • Threshold law: Danger equals play cost × (1−r)^N, with simultaneous 95% rarity intervals propagated monotonically through the risk estimate.The paired cells additionally combine rarity intervals with paired play-cost bootstraps.
  • Threshold law: Danger is ≈0 inside the random-rollout envelope, rises through the elbow, and plateaus at full play cost as N shifts the threshold.On the wall-position sweep, the elbow lies entirely inside the tested range, while play cost is approximately 1 and knob-invariant except for far-plateau leakage at the largest knob.
  • Robustness and limitations: 100 paired episodes preserve total exploitation, but the below-random comparison is baseline-sensitive: on the cart, Jrand − Jblind = 0.249 with interval [0.024, 0.559], yet the blind planner wins 86 of 100 seeds.The random baseline has a heavy-tailed distribution, with median 4.7×10−4, maximum 10.58, and skewness 6.1.
  • Cross-instrument mechanism: PatchField2D has per-mode rarities r1 ∈[0.085, 0.245] and r2 ∈[0.0067, 0.0100], Jblind = 0, Jrand = 0.11, and play cost [1.005, 1.006].At 600 rarity rollouts, six of nine cells are censored zeros; 50,000-rollout measurements find dependence negative at (2, 6) and (3, 7), positive at (4, 6), and undecided at (4, 7).
  • Cross-instrument mechanism: A second base planner family remains low-reach across both 1D instruments and PatchField2D, while Proposition 8 links low query reach on disagreement regions to low play cost.High query reach permits high play cost but does not force it.

5 Axis separation: which errors the gate catches at ε = 10−2, and which it misses

At ε = 10−2, the gate catches pervasive errors but misses localized hard modes according to the (1 −r)^N danger law. Across tolerances, tightening ε remains a pervasive-error control rather than a mode-detection mechanism.

  • Axis separation: Danger occurs only when a mode is both rare and hard, with d@N = play cost · (1 −r)^N.Figure 4 places only the rare hard-mode quadrant at high gate-miss probability and play cost.
  • Axis separation: 0.0001 reveal-rarity predicts a 0.9940 pass rate, within the measured [0.9814, 0.9994].Reveal-rarity was measured on 20,000 rollouts after increasing the sample because the sub-ε arm was censored at zero.
  • Gate behavior: 0.003 vs 0.0030 and 0.667 vs 0.6622 are the measured-versus-predicted pass@40 values for the two wall rows.Both predictions fall inside their 300-gate Wilson intervals, supporting pass@40 = (1 −r)^40 at gate scale.
  • Gate behavior: A global drag bias above tolerance is revealed on every rollout and never accepted, whereas a sub-tolerance bias is accepted at 0.997.The gate therefore polices pervasive error, with the acceptance rate for sub-tolerance bias predicted by the danger law.
  • Tolerance sweep: ε ∈{10−9, . . . , 0.3} leaves mode-arm reveal-rarity flat while pervasive-bias arms switch at their own error scale.For mode arms, pass@40 ≈(1 −r)^40 continues across the full tolerance grid.

6 The exploitation is planner-mediated: a distrust-region fence collapses it on the 1D instruments

The exploitation is planner-mediated, not model-mediated: distrust-region replanning fences refuted predictions and truncates trajectories crossing those fences, collapsing exploitation without changing the model or gate. On both 1D instruments, one contact fences the mode on all eleven cases.

  • Planner-mediated exploitation: The exploitation is planner-mediated rather than model-mediated, so a planner-side fix collapses it without changing the model or gate.The gate can still accept a wrong model, leaving the danger law and Proposition 5 untouched.
  • Planner-mediated exploitation: Distrust-region replanning fences positions associated with refuted model predictions and truncates imagined trajectories that cross a fence.The intervention acts on planning rather than model synthesis or gate acceptance.
  • 1D instrument fence: 1 contact suffices to fence the mode on all eleven cases across the two 1D instruments.A single planner-side contact is enough to establish the distrust-region fence in every reported 1D case.

7 LLM synthesis: on the 1D clamps the danger reduces to the identifiability event

On 1D clamps, GPT-5.x usually repairs a revealed mode exactly, while samples that miss the mode produce accepted, wall-blind models exploited at play. The results separate synthesis failure from sample identifiability: 2D region induction fails despite successful rule translation and targeted interventions.

  • 1D clamps: GPT-5.x produced correct synthesis that was float-exact through the sandbox, with gate 1.000, zero refinement iterations, and play at truth parity across 40 seeds.Thus ε = 10−9 costs nothing when the rule is given or correctly synthesized.
  • 1D clamps: 20/20 wall-absent seeds passed the gate at 1.000, remained wall-blind, and were exploited at play with cost 0.999.The wall-absent event matched (1 −0.0114)40 ≈0.63, and the 20 seeds covered 20 distinct gate-sample blocks.
  • 1D clamps: 20/20 wall-present seeds were repaired: large 10/10 in 0–1 iterations and mini 10/10 in 0–5 iterations.The repaired artifacts wrote the true global clamp rather than a curve fit.
  • Evidence and identifiability: 111 draws occupied 36 distinct gate-sample blocks, so the instrument–stream block—not each draw—was the primary inference unit.Shared random streams mean changing instrument, knob, patch shape, or prompt variant does not create an independent sample.
  • 2D regions: 0/76 incomplete-arm seeds recovered the circular disc rule, while eight ablations found the 2D failure persisted across tested prompting, budget, curvature, and related causes.Given the rule, artifacts wrote discs, squares, slabs, and boundary projections at gate 1.000 in zero refinement iterations.

8 An independent acceptance sample

An independent acceptance sample changes the protocol from training-sample consistency to prospective validation, and its evidence adds only where the tested hypotheses hold. Re-scoring shows improved off-sample exactness outside mode regions, but accepted artifacts can remain globally wrong when the evaluation distribution is uninformative.

  • Independent gate: 102 of 102 spot-checks reproduce the originally stored accuracy, validating that reconstructed training blocks are the original samples.The three blocks are verified disjoint at the individual rollout-seed level.
  • Hypotheses and exponents: 60 mode-blind artifacts show exact agreement between held-out acceptance and mode-free acceptance: 25 are accepted with the mode absent from Dg, and 35 are rejected with it present.This supports the tested hypothesis that a mode-blind artifact fails independent acceptance exactly when the sample contains a mode contact.
  • Independent gate: Independent acceptance rejects 40 of 650 draws that scored 1.000 on their own sample, or 6.2% of draws across 25 rollout-seed blocks.The counts share rollout-seed blocks, so no binomial interval is attached.
  • Off-sample evaluation: 613 independently accepted draws are exact outside the mode region on a further independent 100-rollout evaluation sample.Every failing evaluation transition, where present, is a mode contact; the 613/613 fraction receives no binomial bound because draws share seed blocks.
  • Identifiability limit: A class of entry rules remains accepted at ε = 10^-9 by the own sample, independent acceptance sample, and 100-rollout evaluation sample despite being wrong at play.Freeze semantics remove the evidence separating membership rules from entry rules, so acceptance does not remove this identifiability limit.

9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable · 10 Localization is representational: smooth learners and a mode-capable baseline on the same samples

The 2D collapse reflects limits of instrument identifiability and a synthesis prior, not missing evidence or inability to fit known geometric parameters. On the same samples, smooth learners remain mode-blind or diffuse error, whereas mode-capable classes localize and exactly repair identified events.

  • 9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable: 2000 rollouts per instrument show that identifiability depends on whether samples reach a region’s far side, relative to a stated class and tolerance.Population support constrains the region from more than one side but does not itself guarantee rule identification.
  • 9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable: Eight interventions do not overcome the 2D failure, while positive controls show that constants follow from form and location but nothing follows from form alone.The evidence therefore separates a failure to perform the fit from failures of evidence, constant fitting, or representation.
  • 9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable: 20 of 20 synthesizers infer the withheld radius exactly when given the region’s form and centres, achieving IoU = 1.000 and agreement on all 9020 state–action-grid points.All 20 pass independent gate and 100-rollout evaluation at ε = 10−9; given form alone, recovery is 0 of 20.
  • 9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable: The reachable coverage range is 111◦–185◦, beyond which freeze semantics prevent full-circle coverage from every bearing and leave the instrument unable to answer the question.Ring starts widen angular coverage while holding contacts at the baseline median of 14.5, so the manipulation changes coverage rather than quantity.
  • 9 What the 2D collapse is about: eight interventions, two controls, and a target that is sometimes not identifiable: 8 of 8 incomplete-arm artifacts accepted on their own samples are rejected by independent acceptance samples, with held-out accuracies 0.9944 to 0.9997.The artifacts span 6 distinct rollout-seed blocks, showing what independent acceptance adds when the evidence is not identifying.
  • 10 Localization is representational: smooth learners and a mode-capable baseline on the same samples: 10−15 off-mode error lets the wall-free linear model pass the ε = 10−9 gate while remaining wall-blind, with probe error 4.18 like synthesized blind code.This instantiates learner-independent non-identifiability: on the miss event, acceptance provides no information distinguishing mode-blind from true models.
  • 10 Localization is representational: smooth learners and a mode-capable baseline on the same samples: Four contact rows out of 3200 tilt the linear fit’s off-mode maximum error from 1.7e−14 to 1.2e−02, failing both gates while retaining probe error 4.17.The smooth hypothesis leaks error throughout the domain, whereas synthesized code localizes the event; a mode-capable learned class can match code’s float-exactness.

11 Related work · 12 Limitations · 13 Conclusion

The paper places its sampling-gated synthesis pipeline within CEGIS, statistical model checking, active identification, hybrid-system learning, and decision-aware planning, while showing that its guarantees are limited by rare modes, geometry, planner coverage, and experimental scope. Its conclusion is that continuous state spaces preserve the measure-theoretic failure law, but localized discontinuous boundaries make mode omission especially dangerous and difficult to repair from data.

  • 11 Related work: The pipeline is a CEGIS loop whose counterexamples come from random transition sampling rather than a verifier or equivalence oracle.This substitution explains why additional refinement iterations cannot guarantee completion when random samples miss rare modes.
  • 11 Related work: Rare-property verification requires sample sizes scaling inversely with event probability, while directed falsification and adaptive generators target inputs that uniform sampling misses.The paper characterizes its gate as passive random testing with blind spot probability (1 −r)^N.
  • 11 Related work: Recoverability depends on excitation and hypothesis class: active queries can outperform passive sampling, while mode-capable classes and bounded-Lipschitz classes support different forms of identification.The paper uses this distinction to motivate experiment design, explicit mode structure, and its localization results.
  • 12 Limitations: The study’s empirical scope is limited to three designed deterministic instruments, modest shared-sample synthesis cells, two fixed planner configurations, and incomplete model-family coverage.External dynamics, process noise, broader hyperparameter sweeps, and prevalence beyond the designed instruments remain unestablished.
  • 12 Limitations: 0/156 2D-region recovery attempts failed across GPT-5.x sizes, including guided treatment and targeted interventions, reversing the 1D repair result.The failure includes curved and flat boundaries, while positive controls show that located rules and constants can be recovered.
  • 12 Limitations: The mode-blindness probe detects whether the true active mode is encoded but misses invented modes outside the sampled region, such as a phantom stop at θ = −1.4.Detecting invented modes requires code inspection or probes seeded outside the sampled region.
  • 12 Limitations: 1.9% of the exploited planner’s queries lie in the certified smooth region, while discontinuous reset modes have no finite local Lipschitz constant and evade that certificate.The geometric localization bound is unconditional at boundary points, but converting geometry into visitation probability is instrument-specific.
  • 13 Conclusion: Continuous state spaces preserve the gate-miss law, identifiability argument, and play-cost bound, but discontinuous reset boundaries localize the dangerous omitted rule.On 1D instruments GPT-5.x repairs mode-containing draws in 105 of 111 cases, whereas on the 4D bi-modal instrument it recovers no 2D region rule in 0/156 attempts.

Supplementary Material

The supplementary results show that reveal-rarity can be exactly ε-invariant only over a sample-defined range, while population behavior has no positive threshold and instead follows a rate. They also establish that minimum-based thresholds are unstable and sample-dependent.

  • ε-invariance: pass@N = (1 −r)^N is exactly ε-invariant on the range above the smallest positive contact disagreement.This identity is defined through the rollout statistic D and applies when the model differs from truth only at the hard mode.
  • Rate replacing the threshold: 0.420, 0.123, and 0.041 are the running minima after 25, 200, and 3200 wall@4 firing rollouts, respectively, with no positive limiting threshold.The population threshold is zero; the meaningful population statement is a rate holding at every ε.
  • Rate replacing the threshold: 0 ≤ rfire − reveal-rarity(ε) ≤ C ε^2 for every ε > 0 in clamped semi-implicit plants.The bound separates faint mode contacts from mode firing and depends on the density constant C.
  • Rate replacing the threshold: 222 is the bound C for the cart and pendulum, while 2.10 is the pendulum’s near-equality confirmation of the proved exponent.The bound is loose, but the two-constraint mechanism is close to tight for the nonlinear pendulum.
  • Threshold limitations: 1.8 and 1.4 are the observed cross-stream variation factors for ε∗ on wall@4 and wall@8, showing that the minimum is unstable and sample-dependent.At ten times the sample, ε∗ falls to 0.065 on wall@4 and 0.272 on wall@8.

B The play-cost normalizers, and why knob-invariance is arithmetic

The derived normalizers make play cost an explicit, nearly attained bound when the blind planner repeatedly targets the wall. Knob-invariance is arithmetic: if truth and random returns are knob-independent, only the exploited planner’s residual reward can vary.

  • Play-cost normalizers: Jrand = 0.53 and Jtruth = 17.77 on the cart, while Jrand = 0.06 and Jtruth = 20.08 on the pendulum; qhit(E) ≈1 saturates normalized play cost ≈1.The blind planner queries the wall region in every episode, with Jblind ≈0 on the cart.
  • Play-cost normalizers: 18.0359/17.238 = 1.0463 bounds the cart’s play cost, versus 1.0299 measured, which reaches 98.4% of the derived ceiling.The derived normalizers use Jmax ≤ 18.0359 and Jmin ≥1.33 × 10−6 at xwall = 8, with qhit = 1.
  • Play-cost normalizers: 2 × 10−7 is the push-left trajectory’s maximum gap from the normalizer at xwall ≤6, so Jmax is known there rather than merely bounded.The same verification searches several policy families and 4000 random block policies as a redundancy check.
  • Phantom targeting: 2 × 10−4 phantom-targeting probability from rest is negligible, 5 × 10−3 is estimated under the planner’s sampling law, and the probability becomes 1.000 from (2, 3) onward.Thus phantom-targeting is self-reinforcing but not self-starting; the lure’s beginning and contact = 1.00 remain measured.
  • Knob-invariance: 17.757356407381 and 17.772246981024 are the cart truth-planner returns at all seven knobs for the sharp and default variants, while knob dependence reduces to exploited residual reward.The sharp and default random returns vary by 3.9×10−9 and 1.2 × 10−4, respectively; a pinned planner gives play cost exactly Jtruth/(Jtruth−Jrand).

C What the gate does certify: coverage certificates for Lipschitz pairs

Coverage certificates can bound model error over gate-covered regions, but their resolution is constrained by sampling geometry and density assumptions. The strongest certificate is partition-based, yet certified coverage may barely overlap the planner’s queried region.

  • Packing certificate: A packing proof certifies sup_U ∥f − f̂∥∞ ≤ 2.97 from N = 40 rollouts, using ρ = 1.165 under c = 5/6, L = 1.27, ε = 0.01, and δ = 0.05.The corner condition is barely satisfied: U’s narrowest extent is 0.6 versus ρ/2 = 0.583; the wall disagreement is 4.2.
  • Partition certificate: The exact partition argument improves the certificate to ρ = 0.600 with K = 8, roughly halving the packing-route bound without new measurements or assumptions.At M = N = 40, K = 8 is admissible because 8 · (7/8)^40 = 0.038 ≤ δ, while K = 9 exceeds δ.
  • Scope conditions: The density-based certificate is hypothesis-specific: it requires additive action gains, excludes PatchField2D’s directional force map, and must verify its region assumptions directly.The two-action conditional density is exactly c only when the relevant clamp-free history and target-image conditions hold.
  • Later-step coverage: Later steps expand certified level-set volume but reduce density, so at fixed N = 40 they provide no certificate under a shape-free ball-mass bound.Certified volumes are 3.02 at step 20 for {p ≥ 0.05}, 5.77 at step 40 for {p ≥ 0.02}, and 5.27 at step 80 for {p ≥ 0.01}.
  • Certificate relevance: Only 1.9% of the exploited planner’s 1.3 million queries per episode pair fall inside the dependence-exact certified box, despite the partition certificate being empirically calibrated.For K = 8, independent validation covered the exact partition in 384/400 trials, with measured failure rate 0.0400 and 95% interval [0.0248, 0.0640].

D The detectability rate, and the gate’s visitation density

For smooth model errors, a lower-bounded gate visitation density converts spatial disagreement into a quantitative detection rate, with miss probability at most (1−q)^N. The guarantee is geometric and instrument-dependent: it can constrain hiding in supported regions but is vacuous where the gate has zero visitation density, including the hard mode wall.

  • Detection rate: A density lower bound c on an interior disagreement ball gives per-rollout reveal probability q = c(2ρ)^(d+m), so N rollouts miss with probability at most (1−q)^N.Here ρ = (η−ε)/(2L); without interiority, use q = c vol(B ∩ (S × A)).
  • Detection rate: The hiding threshold scales as N^(1/(d+m)); when cN/ln(1/δ) = 48, increasing dimension lowers the required Lipschitz constant for hiding.The dimensional comparison is stated for cN/ln(1/δ) > 1; below 1, the exponent reverses the comparison.
  • Instrument-specific density: For the cart at step 1, c = 5/6 and the guaranteed disagreement ball occupies the corresponding fraction of the one-step reachable volume.The step-1 law is uniform on a reachable set of volume 1.2, and Monte Carlo agrees within 1.3%.
  • Instrument-specific density: At ε = 0.01 and N = 40, hiding η = 0.5 and η = 1 requires L ≥1.78 and L ≥3.60, respectively, versus the plant constant 1.27.At η = 0.5, a pair no rougher than 1.4× the plant cannot hide an error of that size here.
  • Scope conditions: The result is vacuous for the hard mode wall: c = 5/6 applies on |x1| < 0.53, while the wall lies at xwall ∈[2, 10] where step-1 visitation density is exactly zero.The guaranteed ball does not bridge this gap under the stated plant Lipschitz constant.

E A second planner family: play cost is planner-dependent

Play cost depends on planner query reach: Proposition 8 gives low query reach as sufficient for low play cost, while high reach only permits, rather than forces, exploitation. Across the measured CEM settings, blind-model play cost stays near zero despite lower imagined crossing than MPC, with important planner-specific caveats.

  • Planner-dependent play cost: Low query reach forces low play cost, whereas high query reach permits but does not force high play cost.Proposition 8 bounds play cost by the planner’s probability of querying the disagreement region; the reported “two branches” are measured regimes, not a predicted dichotomy.
  • CEM results: [−0.0213, 0.0248] is CEM’s blind-model play-cost range on every row, with seed-paired 95% t-intervals including zero across all 11 rows.On the five cart rows, every per-seed difference is exactly 0.0, yielding degenerate [0, 0] intervals; imagined crossing remains strictly below MPC’s throughout.
  • Planner-specific qualifications: 70% and 25% are CEM’s pendulum contact rates at θstop = 0.8 and 1.0, while contact is zero from θstop ≥1.2.These contacts do not enter MPC’s pinned, below-random regime, so contact alone is not exploitation; moreover, crossing is only a proxy for qhit(E).
  • Planner-specific qualifications: 15.36 to 16.46 is CEM’s pendulum truth-return range, versus MPC’s 20.08.The comparison is blind-CEM against truth-CEM, not a claim that CEM is globally optimal; limited reach can also miss real rewards, and only one fixed CEM configuration is tested.
  • PatchField2D: 0.070 < 0.208, 0.027 < 0.149, and 0.009 < 0.094 are CEM-versus-MPC crossing fractions at knobs (2, 6), (3, 7), and (4, 8).Blind CEM is not exploited on the 4D bi-modal instrument, and no 2D competence gap appears, although per-seed variance remains.

F Planner-side mitigation: distrust-region replanning

Distrust-region replanning removes the planner-mediated exploitation without changing the model or acceptance gate, but its fencing cost is governed by boundary geometry and does not guarantee completion. In one dimension a single fence suffices, whereas in two dimensions distinct fences scale with packing complexity and duplicate violations can remain large.

  • Method and effect: Planner-side distrust-region replanning collapses exploitation while leaving the blind model and acceptance gate unchanged.The gate can still accept a wrong model; the mitigation acts only during planning.
  • Method and effect: pc blind remains ≈0.94–1.03, while pc mit never exceeds 0.81 across all 11 mitigation-sweep rows.Exactly one violation fences the mode on every row; the residual mitigated cost reflects unavoidable first contact and increases with lure distance.
  • Packing bound: The number of new-coverage violations is bounded by the packing number Npack(F, ε), and once fences ε-cover the entry barrier B, imagined entries are truncated.The bound applies without additional hypotheses; the barrier-cover conclusion requires ε-coverage of B.
  • 1D behavior: 1D instruments average exactly 1.00 violations on all 11 rows, and pendulum returns and violation counts remain bit-identical as ε sweeps from 0.1 to 0.005.A single point beyond the boundary disconnects the agent from the phantom, so the mechanism is insensitive to ε.
  • 2D behavior: 2D distinct-fence counts are at most 2/5/6 at knobs (2, 6)/(3, 7)/(4, 8), while the two worst episodes record 28 violations each.The distinct counts are at most a quarter of the 24-fence budget, but raw violations are dominated by duplicates.
  • Dimensional cost and limits: The deployment fencing cost grows exponentially with boundary dimension p, but this caps only pairwise new-coverage fences, not completion time, total violations, or phantom closure.Duplicate fences are unlimited by the proposition, and the 2D campaign observes an episode with 24 duplicate violations.

F.1 The 2D mitigation: a partial collapse, and lock-in at the far knob … K Cross-family arms, and a fourth artifact class

The paper shows that continuous-world-model mitigation is fragile at multidimensional boundaries, while multi-mode gate guarantees require dependence-aware bounds. Across sharper rewards, tolerance sweeps, bounded noise, and model families, acceptance remains fundamentally limited by what samples expose.

  • F.1 The 2D mitigation: a partial collapse, and lock-in at the far knob: 0/20, 2/20, and 7/20 episodes end pinned at blind-level return as patch distance grows, despite the 2D fence’s soundness.The failure is lock-in at a reachable boundary: unsigned distance-to-fence tie-breaking cannot distinguish rounding away from plunging through the patch.
  • G Multi-mode gates: the sharp bracket, and the measured dependence: The exact multi-mode miss-probability range is the sharp Fréchet–Hoeffding bracket, so independence cannot replace dependence information.At six of nine PatchField2D knobs, no rollout out of 600 contacts both patches, attaining the disjointness endpoint.
  • G Multi-mode gates: the sharp bracket, and the measured dependence: Three of four high-budget knobs resolve dependence, with opposite signs: P(both) = 8.6 × 10−4 versus r1r2 = 19.0 × 10−4, and 12.8×10−4 versus 6.2×10−4.The remaining knob’s interval straddles the product, confirming that dependence is empirically variable rather than uniformly signed.
  • H The residual reward leak, and the sharp-plateau variant: 6/6 knobs fall below random with the asymmetric sharp-plateau variant, while play cost spans only [1.0028, 1.0029] and Jtruth remains 20.08.Removing the phantom plateau’s sigmoid tail strengthens exploitation and makes knob-invariant play cost exact to the reported scale.
  • I The ε-sweep: the axis separation is tolerance-invariant: Mode-arm reveal-rarity is flat across ε ∈{10−9, 10−6, 10−4, 10−3, 10−2, 3×10−2, 0.1, 0.3}.On PatchField2D, patches-omitted is 0.147, patch-1-only is 0.142, and patch-2-only is 0.005 at every ε.
  • J Bounded observation noise: the gate law survives to a measured masking boundary: At η = 0.1, the analytic blind-pass increase is exactly 0, while η = 1 is the first pre-specified level exceeding the frozen +0.05 boundary.The correct model passes every block at every tested noise level; bounded noise preserves the gate’s content through the primary level.
  • K Cross-family arms, and a fourth artifact class: The gate certifies sample-covered exactness, not global correctness: a phantom symmetric stop at θ = −1.4 was accepted at gate 1.000 despite never being reached.This artifact was exact on every sampled transition but encoded a hard stop that does not exist.

L The 2D artifacts: a code and behavioural audit, and three ablations … O.1 The two messages

The 2D audit finds that translation is reliable but induction of the circular boundary collapses into a stable half-plane prior across prompting, model-family, arity, evidence, and budget interventions. Coverage certificates transfer only to smooth settings, while the protocol’s sampled-transition messages and refinement checks define what artifacts are actually tested.

  • L The 2D artifacts: a code and behavioural audit, and three ablations: ≈74/76 artifacts reproduce the exact 4D integrator and reward, while 38/76 replace the 2D disc with a half-plane clamp.The audit localizes the failure to induction of the circular boundary rather than plant translation.
  • L The 2D artifacts: a code and behavioural audit, and three ablations: 0/40 repairs result from richer prompting and a 3× budget, although the intervention changes the failure class.The treatment uses 120 examples, 40 failure lines, region-first guidance, an explicit anti-1D-threshold instruction, and 15 refinement iterations.
  • M The eight ablations on the 2D mode, campaign by campaign: 19/20 mode-containing seeds are repaired by gpt-5.4, but none recover the rule: all accepted artifacts implement the same half-plane.Each passes its own gate, an independent acceptance sample, and an independent 100-rollout evaluation at ε = 10^-9, yet scores 0 on the mode-blindness probe.
  • M The eight ablations on the 2D mode, campaign by campaign: 0/40 artifacts recover the region in each interior-witnessing campaign, while the dominant half-plane class remains 24/40 and 25/40.Landing and clamp achieve best agreement of 0.10 and 0.26, respectively, versus 0.50 for the disc.
  • M The eight ablations on the 2D mode, campaign by campaign: 7.25% mode share in landing provides eleven times more mode evidence than 0.66% in freeze, but both landing and clamp still fail.Clamp holds mode share at 0.75%, matching freeze, while increasing rule complexity instead.
  • N The coverage certificate’s scope, stated in full: Proposition 14 certifies supU ∥f − f̂∥ ≤ ε + 2Lρ for L-Lipschitz pairs once the gate ρ-covers U, but fixed sample size trades extent for resolution.The certificate transfers to smooth cases; no step-t level set certifies below the step-1 figure, because the gate’s size is limiting.
  • O The LLM protocol, in full: Each campaign cell is one (arm, seed) pair issuing one synthesis call followed by zero or more refinement calls, with counts regenerated from the campaign code.The appendix states that its protocol facts and per-campaign counts are read directly from the code revision producing the results.
  • O.1 The two messages: N = 40 rollouts yield 3200 transitions, but the default prompt shows 30 lines and guided variants show 120; refinement supplies truncated failure lines.The refine text requests matching every transition within ε in x, v, and reward, while the enforced check uses the sup-norm over all state components and reward.

O.2 The refine loop … O.9 Prompt variants

The pipeline accepts the first artifact matching every sampled transition to ε or stops at its iteration budget, while extraction, retries, sampling, and classification follow fixed operational rules. Across campaigns, call costs varied sharply by instrument, and prompt guidance addressed a landing-state identification failure without changing the contract.

  • O.2 The refine loop: The loop stops at the first artifact with accuracy 1.0 on the whole sample or when max iters is exhausted; ε = 10^-9 makes matching effectively exact.The default budget is 5 iterations per campaign, while the two region-guidance campaigns use 15.
  • O.3 Code extraction, and an unparseable reply: Replies without code fences are passed as code, and sandbox failure, empty output, or non-JSON output yields accuracy 0.0 so refinement continues.The separate audit path classifies such an artifact as invalid.
  • O.4 Retry and error policy: Every cell is the first and only draw for its (arm, seed), because provider errors are not retried, resampled, or dropped at the application level.Checkpoints support restart without recomputing completed cells.
  • O.5 Sampling parameters, deployments, API version: Sampling leaves temperature, top p, and request-level seed at provider defaults; deployments are gpt-5.4 for large, gpt-5.4-mini for mini, and never gpt-5-nano in paper-2 campaigns.The cross-family arm uses Qwen/Qwen3-Coder-30B-A3B-Instruct, and Azure API version is 2025-04-01-preview.
  • O.6 LLM calls per seed: 1998 calls across 20 API campaigns, 625 cells, and two agent-relayed Claude campaigns averaged 3.13 calls per cell, ranging from 1–16.One-dimensional instruments typically succeeded on the first try, whereas PatchField2D at k = (3, 7) exhausted its budget in 20/40 cells per campaign.
  • O.7 What the gate is: Each cell collects 40 i.i.d. uniform-random truth rollouts, uses them both as examples and as the gate, then logs gate passage and whether the sample contains a mode.PatchField2D records mode blindness as a per-mode dictionary and places its mean in wall blindness.
  • O.8 Operational classification rules: Behavioural classification probes artifacts on an 81 × 81 grid at action 0.3, with deviation threshold 10^-6, two velocity slices, and mask-based coverage measures.The one-dimensional repaired predicate requires gate passage plus zero wall blindness, while PatchField2D repair is gate passage alone among mode-present cells.
  • O.9 Prompt variants: 36/40 region-guidance artifacts conditioned freezing on the current position rather than the landing position (x2, y2), motivating a landing-state prompt intervention.Prompt variants change only max examples, max failures, and guidance; the contract text is unchanged.

O.10 The agent-relayed Claude arms … Declarations

The paper documents relayed Claude campaigns, reproducibility infrastructure, and a clear separation between pre-specified findings, post-result additions, and post-hoc theory. It also declares funding, data/code availability, AI use, and anonymization status.

  • O.10 The agent-relayed Claude arms: 20 relayed calls covered 8 one-dimensional cells, while 19 covered 4 two-dimensional cells; duplicate relay files were excluded from the protocol count.The two-dimensional directory would contain 21 files only if the discarded duplicate replies were included.
  • P Reproducibility: CPython 3.12.8 ran on macOS 15.7.7 with seeded pure-Python or numpy float64 CPU results; printed twelve-digit values and 10^-15 residuals are most sensitive to libm differences.The reported truth values are Jtruth = 17.757356407381 and 17.772246981024.
  • P Reproducibility: 9 of 15 tables and figures were cell-parsed, 5 were claims-only, and 1 was unaudited; static and dynamic manifest checks agree after an initial wrapper-and-glob failure.The manifest records the unaudited status of Table 5 rather than leaving the gap implicit.
  • P Reproducibility: 1959 API calls across 625 cells and 20 campaigns are exact, while named CPU and LLM campaigns total 10.79 h as a lower-bound runtime.The repository’s measured companion-paper cost is $3.73, presented only as an order-of-magnitude anchor.
  • Q Pre-specification, and what was added after the result: 14 items were pre-specified and confirmatory, whereas 15 were added after results, including diagnostic repairs, exploratory ablations, post-hoc analyses, and mixed items.The paper identifies ten theory results as post-hoc and treats the one-dimensional synthesis result as the pre-specified core.
  • Q Pre-specification, and what was added after the result: 10 process claims were audited: 7 were supported, 1 was supported statistically rather than temporally, and 2 unsupported dated-record claims were corrected.The corrected claims concerned the sharp-plateau competence risk and the statement that partial repair had been predicted.
  • Declarations: The work received no external funding, and the author declares no competing interests.LLM API costs were borne by the author.
  • Declarations: All code, result artifacts, and quoted synthesized programs are available under Apache-2.0 and CC-BY-4.0, but no archived DOI-bearing release exists; AI systems supported both the study and manuscript preparation.The author states that theorem statements, proofs, experimental designs, and conclusions are their own, and the version is not anonymized.
Loading 2608.17956v1…