Source-linked AI summary
An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models
Javier Aguilar Martín
TL;DR
Sampling gates certify only a code world model’s behavior on reachable queries, leaving unreachable structure unconstrained. This paper formalizes that quotient on an annular freeze-mode instrument, proves the closed-ring filled-disc artifact is certified yet harmless, and measures how reach, sensing, repair, and mitigation change its consequences. The results show that identical topology can produce opposite danger when reach differs, while sensor resolution and parameter identifiability limit repair.
Problem
Sampling-gate certification can leave the behavior and topology of unreachable regions undetermined, raising the question of what certified models know and what their errors cost.
Method
The paper combines a gate-quotient analysis, a minimal annular freeze-mode instrument, theorem-backed channel-width sweeps, and LLM synthesis and mitigation studies.
Results
A facing channel collapses exploitation over a knee at γ ≈0.1, while a hidden channel preserves play cost 1.116 despite identical topology; dimension-matched mitigation reduces cost from 0.999 to 0.058.
Takeaways & Limitations
Danger is topology relative to planner reach, and repair is constrained by both reachable evidence and the sensor’s finite resolution rather than topology alone.
Takeaways & Limitations
The sensor conclusion is detector-specific: at n = 6 only a 2-plane projection diagnostic was computed, which is not evidence about β5.
Abstract
from arXiv · showhide
A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior. The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument we prove the extreme case (a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play) and measure, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly, and instantly falsified. Three principles organize the empirics. First, danger is topology relative to reach: a channel the planner can use collapses the blind model's exploitation (play cost 1.09 to ~0 over a knee at gamma ~ 0.1), while a hidden channel with the same first Betti number keeps it at full strength (1.12). Second, repair is parameter-bound and sensor-bound: no family recovers the region from outside evidence; from inside, models pose the right topology but cannot pin its parameters, and the posed topology tracks the guiding persistent-homology summary's wrong beta_1 (a sensor with a measured geometric resolution limit), not the truth. Third, mitigation must match the error's dimension and direction: point fences fail against the one-dimensional boundary, a dimension-matched persisted fence collapses exploitation to a two-lesson transient (0.999 to 0.058), and the dual freedom certificate collapses the invented-mode failure symmetrically (1.769 to 0.029). In n dimensions the shell makes misidentification near-certain while the danger stays fully exploitable: the two axes are independent.
1 Introduction
The paper reframes model correctness as topology relative to reachable queries: sampling gates certify only what they can reach, while errors beyond reach remain gauge. A ring instrument and LLM studies show how reach, sensing, repair, mitigation, and dimension determine whether wrong topology is harmless or dangerous.
- Motivation: An annular freeze mode is the minimal setting where an unreachable interior makes certification and topology inseparable.Outside trajectories stop at the outer wall, so the existence of a hole is not determined by outside data.
- Motivation: Acceptance-with-certainty fixes the model only on the reachable query set; beyond reach, model behavior is gauge.The gate quotient makes certifiability relative to the gate policy and its reachable queries.
- Theory and mechanism: At γ = 0, the filled-disc artifact is unfalsifiable by every sampling gate and bitwise harmless during play.The artifact is wrong-topology but produces identical planner-induced real trajectories.
- LLM synthesis: Across GPT-5.x, Qwen, and Claude, every mode-absent sample produced a certified blind exploited model at play cost ≈1.12.No tested family repaired the ring from outside evidence; inside, topology could be posed but parameters remained difficult to identify.
- Empirical regimes: At a facing channel, exploitation collapses over a knee near γ ≈0.1, whereas a hidden channel preserves full danger at 1.116 despite identical topology.The difference is whether the planner can reach the channel, not the channel’s Betti-number change.
- Sensing and mitigation: A dimension-matched persisted fence reduced play cost from 0.999 to 0.058, while point fences failed against the one-dimensional boundary.The successful fence restored gate-equivalence on the operative side.
- Sensing and mitigation: The evidence sensor reported ˆβ1 = 1 below its geometric resolution limit although the true β1 = 0, and the posed artifact followed the sensor report.The pre-registered flipped-summary crossover was directionally consistent but not significant (p = 0.065).
2 The instrument: a mode with an inside
RingField2D is a minimal two-dimensional thrust-and-drag plant whose annular freeze band encloses a phantom reward and separates the plane. Three fixed knobs alter reachability without changing the mode’s local physics.
- Plant: RingField2D uses state (x, y, vx, vy), bounded scalar thrust, gain 3.0, drag 0.3, and dt = 0.1.The plant retains the companion paper’s semi-implicit integrator and planner constants.
- Mode geometry: The annular freeze band has inner radius 3.5 and outer radius 5.0, enclosing a high-reward phantom lode at c = (12, 0).Landing in the annulus freezes the state at the previous position with zero velocity.
- Reachability knobs: Channel width γ changes the ring from β1 = 1 at γ = 0 to a contractible β1 = 0 C-shape for any γ > 0.The channel is an angular gap cut from the ring.
- Reachability knobs: Channel orientation leaves topology unchanged but determines whether the gap faces the start or is hidden on the far side.Its effect is through the planner’s path relative to the channel.
- Reachability knobs: The start side selects outside evidence, where the interior is unreachable at γ = 0, or inside evidence, where the hole can be sampled.The three knobs were pre-registered and alter reachability rather than local mode physics.
- Calibration: The default freeze-band thickness w = 1.5 exceeds the maximum per-step displacement Δ ≤ 1.0, preventing trajectories from jumping over the band.The outside contact rarity at calibration was r = 0.0312 per rollout.
3 Theory: the gate quotient
The gate quotient formalizes certification as agreement on reachable queries, making beyond-reach disagreement invisible. The ring theory then shows how channel width continuously reopens identifiability and how contact and entry probabilities behave.
- Gate quotient: A model matching the truth on the reachable query set is accepted with probability 1 for every gate size, and policies staying within reach behave identically.Conversely, certainty acceptance forces disagreement to have zero query-occupation measure.
- Ring specialization: Outside trajectories generate pathwise-identical contact processes for disc and annulus modes, so no outside gate can distinguish their topology.This equivalence also applies to repair loops and topological summaries.
- Ring specialization: At γ = 0, the filled-disc artifact is accepted by every gate and yields play cost = 0 exactly under deterministic planners starting outside.The theorem separates wrongness from both certifiability and behavioral consequence.
- γ-curves: For larger channels, contact events are nested pathwise, making r(γ) nonincreasing in γ.This nesting was observed without violations in 44,000 common-random-number checks.
- γ-curves: Direct interior entries are pathwise monotone in γ, while deviations from full monotonicity are bounded by funnel-assisted entry mass.The funnel mass peaked at 22/50,000 in the reported common-random-rollout measurement.
- γ-curves: The γ-curves are continuous, and interior-entry probability is positive for every γ > 0 while rint(0) = 0 exactly.Thus the channel reopens identifiability continuously from an exact zero.
- Query lower bound: Under the query-lower-bound conditions, a planner with an imagined phantom-entering candidate must query the disagreement region, with measured blind contact rate 1.0 at every γ.The result connects imagined high reward to certain contact with the omitted mode.
4 Mechanism without LLMs: three regimes, one knob
A fixed wrong-topology artifact moves among three certification and gameplay regimes as channel width and reach change: unfalsifiable-and-harmless, falsifiable-and-costly, and instantly falsified.
- Experimental setup: The mechanism grid evaluates the filled-disc and mode-blind artifacts across channel width, channel placement, and start side using gate rollouts and paired MPC episodes.Each cell uses 400 uniform-random rollouts for gate-side falsifiability and 16 paired-seed MPC episodes for play cost.
- Three regimes: At γ = 0, the filled-disc artifact is unfalsifiable and costless, with disagreement rate 0 and pcfill = 0.This is the observed bitwise equivalence predicted by Proposition 2.
- Three regimes: With a facing channel, the same artifact becomes falsifiable at ∼10^-3 per transition and costly, with pcfill values 0.343/0.220.The open channel places the artifact’s wrong region within the planner’s operative reach.
- Three regimes: From an inside start, the artifact is instantly falsified, with disagreement rates of 0.77–0.97.The interior is directly reachable from the start, exposing the invented freeze region.
- Invented-mode exploitation: For an inside start, the invented mode is exploited below random: pcfill = 1.769 because imagined freezes repel the planner from the reward.This is the dual of omitted-mode lure, where the planner is pulled into danger rather than away from value.
- Topology relative to reach: Hidden and facing channels have identical topology but opposite consequences because only the facing channel intersects the planner’s operative reach.The hidden channel remains observationally identical to the closed ring, while the facing channel changes falsifiability and play.
5 LLM synthesis on the closed ring: three families
The closed-ring synthesis study tests three model families with behavioral audits and held-out gates, showing family-independent blind artifacts, failed outside repair, and parameter-bound inside repair.
- Protocol: The synthesis protocol uses four pre-registered closed-ring cells, 20 seeds per condition, N = 40 evidence rollouts, ε = 10^-9, and at most five refinement rounds.Cells vary prompting, topological guidance, per-seed summaries, and inside versus outside starts.
- Mechanism grid: The full mechanism grid preserves the three regimes and reports r, rint, disagreefill, and pc across 400-rollout and 16-episode evaluations.Its qualitative separations exceed the stated Wilson confidence bounds by orders of magnitude.
- Behavioral audit: The behavioral audit probes synthesized step() functions on an 81 × 81 state grid at two velocity slices to classify effective freeze or mode sets.This audit distinguishes behavioral geometry from misleading comments or plausible-looking but misparameterized clauses.
- Outside evidence: Across GPT-5.x, Qwen, and Claude, every mode-absent sample produced a certified blind artifact exploited at play cost ≈1.12.The same blind program explains the bit-identical play result across families; identifiability depends on the sample, not the model.
- Held-out audit: Held-out acceptance coincided with the independent gate’s mode miss in 156/156 artifacts.The replication tested 903 synthesized artifacts on disjoint 40-rollout gate blocks across 39 synthesis conditions.
- Outside evidence: None of the three families repaired the ring from outside evidence; gate-passing mode-present artifacts were behaviorally blind textual point fits.Their freeze masks were empty and wall blindness was 1.0, despite comments hypothesizing localized traps.
- Inside repair: From inside, models posed geometric structures but rarely passed: large achieved 1/20, while median inner-radius error stayed 0.5 across 40–320 evidence rollouts.The posed structures commonly used a round inner radius r = 4.0, indicating parameter guessing rather than evidence-based estimation.
- Inside repair: The sole gate-certified inside recovery used a disc-complement form with gate 1.0, blindness 0.0, and play cost 0.0 despite a wrong outer boundary.The outer boundary remained gauge because inside probes reached only the inner boundary.
6 The open-ring arm: danger is topology relative to reach
The open-ring experiments show that danger depends on whether the planner can reach the omitted region, not solely on its topology. Channel orientation changes exploitation while topology remains fixed, and certified inside-start passes remain essentially non-repairs.
- Reach-relative danger: A facing channel collapses blind-model exploitation over a knee at γ ≈0.1–0.15, whereas an equally wide hidden channel preserves play cost = 1.116.The channel becomes usable when its arc width admits the planner’s step, making the blind plan executable in truth.
- Reach-relative danger: The facing and hidden channels have the same β1 = 0 but opposite danger because only the facing channel intersects the planner’s operative reach.The hidden channel remains observationally identical to the closed ring, while the facing channel makes the phantom directly accessible.
- Certification versus repair: Inside-start gate-pass remains essentially zero across γ, so certified passes do not establish repairs.The strongest repairer reconstructed “ring r = 3.5, gap 2.4 rad” and passed in 2/3 seeds, yet remained mode-blind; it was harmless because the wide opening removed the obstruction.
- Independent axes: The synthesis and play axes vary independently: increasing γ raises mode-absent samples from 6 to 16 of 20 while reach-dependent danger can collapse or persist.The same instrument therefore separates whether evidence contains the mode from whether its omission obstructs planning.
- Metric converse: Thin-neck controls keep topology fixed while changing metric access, with play cost collapsing only at the thinnest neck and recovering to 0.96–0.99 for necks 0.2 and wider.At neck 0.1, pcblind = 0.451 and blind contact rate = 0.56; thicker necks are pinned by a landing in the freeze band.
7 Mitigation against an enclosed mode: a covering law
Mitigation succeeds only when it matches both the geometry of the error and its directional sign. Persistent tangential fences restore truth-equivalent planning for omitted modes, while freedom certificates address invented modes.
- Covering law: Point fences fail against the ring because a zero-dimensional fence cannot cover its one-dimensional reachable boundary.The planner can re-cross through unfenced imagined arcs; the required number of violations scales with boundary measure over fence radius.
- Dimension-matched defense: A persistent nerve fence links violations into tangentially extended segments and can make the mitigated planner exactly truth-equivalent when coverage and return-gap conditions hold.Proposition 8 states that every crossing candidate truncates before crossing and that the argmax then coincides with the truth planner.
- Direction-matched defense: Freedom patching handles the opposite error direction: it locally substitutes the pinned integrator where a model falsely predicts freezing.Under the freedom-certificate condition, patched imagination equals truth imagination for every candidate.
- Direction-matched defense: The omitted-mode lure and invented-mode repulsion require opposite defenses, and each defense’s cost is governed by its failure’s lie rate.Distrust is inert against an invented mode because hallucinated freezes flatten all imagined values before fences can affect selection.
8 The evidence sensor has finite resolution
The evidence sensor has a geometric resolution limit: persistent-homology reports can confidently preserve a false loop, and synthesized topology follows that report rather than truth. Relative homology improves specificity when certified-free trajectories are represented as paths, while recovery remains start- and dimension-dependent.
- Sensor resolution: The pre-registered sensor reports ˆβ1 = 1 for channels narrower than approximately 2 arc-units although the true β1 = 0 for every γ > 0.The detector flips near γ ≈1.8, and increasing evidence dose can strengthen the false report rather than correct it.
- Sensor resolution: The measured resolution flip does not move across detector budgets and evidence doses: ˆβ1 = 1 for γ ≤1.2, 0 at γ = 2.4, with γ = 1.8 wobbling.This supports a geometric rather than budgetary resolution limit.
- Topological criterion: Persistent-homology guarantees follow a gap- and thickness-dependent sandwich, with the pointwise law reproducing 78/80 factorial rows and the two misses lying in the undecided band.The criterion uses mean radius, largest angular gap, and persistence threshold; the displayed predictor is measured rather than exact.
- Relative-homology repair: Relative homology replaces edge deletion and achieves 20/20 open-ring correctness when contact evidence is paired with certified-free trajectory polylines.Plain Rips scored 9/20 in the cited comparison, while relative homology avoids infinite bars and uses the free paths to certify passage.
- Nested structures: Nested boundaries expose a deeper resolution failure: the detector recovered the true β1 = 2 only once in 20 samples, and no synthesized artifact posed a two-band structure.The detector returned 0 on 9/20 and 1 on 10/20; synthesis instead produced single arcs hugging one boundary.
- Dimensional scaling: In ShellField-n, outside evidence never recovers βn−1, whereas inside evidence recovers it through n = 5; the n = 6 result is only a 2-plane projection diagnostic.Outside density falls from 0.18× to 0.002× the NSW covering density, making recovery start-governed and dimension-limited.
9 Dimension as the rarity knob
Increasing dimension makes shell contact exponentially rarer, so misidentification becomes near-certain, while competent planners retain full exploitation; these axes remain independent.
- Independent axes: By n ≥3, the identifiability event is near-certain, although truth-MPC reaches the real lode at every tested n ≤6.The paper separates gate-side rarity from play-side reachability.
- Rarity collapse: The shell-contact rarity decays geometrically with dimension, at measured factor 0.411 per dimension and spherical-cap prediction 0.4167.The measured log-linear fit has R2 = 0.999; the spherical-cap prediction is rout/L = 5/12 = 0.4167.
- Mechanism: The exponential decay follows the mode’s angular size from the start rather than variance concentration.The cone condition yields an explicit exchangeability bound, while isotropic actions give the spherical-cap rate exactly.
- Mechanism: The cube-uniform instrument also retains exponential decay through conditional sign independence and symmetric action coordinates.The spherical argument does not directly apply because the cube-uniform interface has only finite symmetry.
- Planner competence: Planner competence depends on the action interface: omitting deterministic axial candidates can mask exploitation, while adding the 2n axial candidates restores it.Direction-uniform random candidates alone do not fix the weak candidate set.
- Topological boundary: A non-separating solid torus preserves path-relative danger but lacks an exact gauge region because it does not separate space.With the hole on or off the optimal path, exploitation changes from 0.02 to 0.90 despite the same rarity and topology.
10 Related work
The paper places sampling-certified code world models at the intersection of verification, model-based reinforcement learning, hybrid reachability, and topological data analysis.
- Code world models: Code world models synthesize executable models searched by classical planners, with sampling-based acceptance functioning as property-based testing of dynamics.The gate quotient formalizes exactly what this acceptance can pin down.
- Model-based reinforcement learning: The paper contrasts structurally unreachable localized error with pervasive-error analyses and visited-state value-loss bounds in model-based reinforcement learning.Those bounds are silent where the gauge region lacks coverage.
- Hybrid systems and reachability: Falsification and reachability methods inherit the same reachability limit: policies querying only R cannot exit the unobserved region.The paper connects this limitation to hybrid-system mode and guard analysis.
- Topological data analysis: Persistent homology supplies the evidence sensor, while stability and sampling theory calibrate when topology inference can work.The paper uses TDA as a component of a running system rather than introducing a new TDA method.
11 Boundary results and limitations
Boundary analyses delimit the paper’s claims: several quantitative conclusions depend on model–planner randomness, detector resolution, measured defects, and a narrow instrument family.
- Empirical scope: The paper’s core instruments are mostly round, separating shells with freeze semantics; non-round separators and moving boundaries remain future work.The square ring and solid torus provide limited ablations rather than general results.
- Empirical scope: The LLM synthesis cells are modest, cross-family comparisons are spot-checks, and one Qwen cell is missing, so cross-family evidence fixes mechanisms rather than rates.The within-gap contrast remains directional-only even after 60 paired seeds.
- Boundary analyses: The heavy play-cost tail is a delay cost from wrong aim-point ranking, not a freeze transient or route-commitment effect.The loss occurs without execution contact; truth arrives at steps 34–36 while the blind planner arrives at 37–38.
- Boundary analyses: A tight play-cost bound must condition on future planner randomness, making cost a property of the model–planner pair rather than the model alone.Holding one dirty-step imagination fixed while varying future seeds moves realized cost over [−1.71, +0.60].
- Boundary analyses: The direct-entry monotonicity results are limited by a measured funnel defect: unconditional pathwise monotonicity is false, including under a velocity-updating variant.The paper reports 91/91 separated violations in the original setting and 13/30,764 entering-pair violations in the variant.
- Boundary analyses: The sensor flip is detector-specific: its γ ≈1.8 resolution threshold is established for one Rips detector and one registered density, while broader propagation is unproved.The flipped-summary intervention was tested in one cell and left the causal share of the claim line unearned.
- Boundary analyses: At wide γ, truth baselines change because the phantom lode becomes reachable, so danger conclusions rely on paired facing-versus-hidden contrasts.Each play-cost value uses its own cell’s baseline, with the normalization issue shared symmetrically.
12 Conclusion
The paper frames certifiability, correctness, and consequence as distinct properties determined by reach. Practically, unreachable fences are both uncertifiable and harmless until a planner-accessible channel makes the omission exploitable.
- A sampling gate certifies the model only on its reachable query set; errors beyond that set remain unconstrained.
- The same artifact can move from unfalsifiable and harmless to falsifiable and costly as a channel changes its reach relationship.
- An enclosed fence cannot harm play while it remains unreachable, but danger appears when the omission intersects a competent planner’s optimal path.
- Repair requires evidence from inside the boundary and a sensor with sufficient topological resolution at the operative sampling density.
Reproducibility
The paper provides a reproducibility package spanning formal verification, mechanism and synthesis drivers, behavioral audits, calibration, held-out evaluation, and committed execution metadata.
- Experimental drivers: The mechanism grid, γ-probes, synthesis harness, agent-relay protocol, and open-ring driver expose the main experimental components.
- Audits and aggregation: The aggregator, ShellField-n scripts, and behavioral audit support result aggregation and artifact checking with an oracle self-test.
- Evaluation integrity: Rarity calibration uses 30,000 rollouts per synthesis configuration, while held-out re-scoring uses disjoint gate and evaluation blocks.
- Formal verification: The Lean formalization records the deterministic results and the continuity modulus’s measure steps, with the library built by continuous integration.
- Execution provenance: Every seed, checkpoint, classified artifact, invocation, serving identity, result JSON, and claims audit is committed or recorded for exact reruns.
- Archival snapshot: The public repository’s paper3-v1 snapshot includes the Python environment lock and SHA-256 manifest for committed results and paper sources.