Source-linked AI summary

Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution

Yijun Liao, Fanwei Liang

arXiv:2608.24460v1cs.CL

TL;DR

The paper asks when model preferences in knowledge conflict can be attributed to data when natural cues agree. It constructs an exact coextensive synthetic setting and uses minimal causal edits to separate the cues. Mechanism identity varies across runs, while escape timing and shortcut ceilings replicate; attribution can also reverse before circuit formation is safely established.

  • Problem

    Natural data rarely makes competing conflict-resolution cues disagree, so behavioral accuracy cannot reveal which cue a model uses or attribute that preference to the corpus.

  • Method

    The paper trains transformers on a synthetic language with coextensive recency and rarity, separating them through minimal edits that invert multiplicity while preserving truth and document structure.

  • Results

    Mechanism identity is not reliably reproducible across seeds, whereas escape timing is monotone in redundancy; probing before escape reverses attribution in 32 of 75 runs at unchanged accuracy.

  • Takeaways & Limitations

    A corpus can fix when a mechanism appears without fixing which mechanism it is; mechanistic attribution requires that the objective not be indifferent between alternatives.

  • Takeaways & Limitations

    The reported map uses one budget and schedule because doubling schedule length moves terminal sign fraction by up to 0.55, while two corpus collinearities remain only partly broken.

Abstract

from arXiv · show

When a context asserts two values for one fact, a model commits to a cue -- recency, repetition, position -- but natural data rarely makes these disagree, so behavior cannot reveal which. We train 26M-parameter transformers on a synthetic language where recency and rarity are exactly coextensive, and separate them with a minimal causal edit that inverts one cue while holding the truth, token count and answer position fixed. All 75 runs reach accuracy >= 0.999, including where the trivial heuristic fails, so no held-in evaluation distinguishes them. Under intervention the per-cell readout does not replicate: 13 of 25 cells differ by more than 0.3 in sign fraction across three seeds, the largest by 0.879 against a standard error of 0.025. The construction predicts this -- coextensive rules leave the objective indifferent between them -- and the variance is ordered by how much of the optimization each comparison releases. What replicates is timing: escape from a positional shortcut with a closed-form ceiling, monotone in redundancy. Probed before that escape, attribution reverses sign in 32 of 75 runs at unchanged accuracy, and gating on circuit formation is necessary but not sufficient. The corpus fixes when a mechanism appears, not which one -- a criterion for when mechanistic attribution to data is available at all, and our construction makes the unavailable case exact.

1 INTRODUCTION

Natural documents make recency and rarity agree, so accuracy cannot reveal which cue a model uses. The paper separates them with causal edits and finds that mechanism identity is not reliably determined, while escape timing is.

  • Natural documents usually make later values also rarer, leaving cue preferences observationally indistinguishable.The commitment becomes visible only when recency and repetition diverge, as in stale or inconsistently updated data.
  • The synthetic task makes recency and rarity exactly coextensive, then inverts their multiplicities while preserving truth, token count, and answer position.This minimal edit makes rarity predict the opposite value without changing other document features.
  • All analyzed cells reach accuracy ≥0.999, including where the trivial last-update heuristic fails, so held-in evaluation cannot distinguish the mechanisms.The same score is compatible with opposite dependence on repetition count.
  • Escape timing remains reproducible and monotone in redundancy, whereas the corpus fixes when a mechanism appears, not which mechanism it is.The paper therefore offers a criterion for when mechanistic attribution to data is available.

2 RELATED WORK

Prior behavioral studies identify systematic conflict-resolution preferences but cannot attribute them to corpus statistics when candidate cues agree. Related mechanistic and phase-diagram work motivates the paper’s controlled separation of observationally aliased rules.

  • Behavioral studies find preferences tracking popularity, coherence, update plausibility, agreeing-source count, repetition, authority, and superseded state.
  • Because natural documents rarely make candidate cues disagree, observed preferences cannot be attributed to the fixed corpus rather than correlated cues.
  • Mechanistic studies localize components and demonstrate circuits, but completeness remains difficult to certify.
  • Prior work documents non-unique explanations, including multiple faithful circuits and conclusions sensitive to analysis-pipeline variation.
  • Unlike prior phase diagrams based on task-distribution properties or output matching, this paper varies cue statistics while holding the task and Bayes-optimal policy fixed.

3 SETUP

The setup uses an infinite synthetic assignment language whose answer is the latest value for a queried slot. It varies redundancy and update-to-query distance while preserving the target function and controlling positional shortcuts.

  • Documents contain assignment statements followed by a query, with the answer equal to the most recent value assigned to the queried slot.
  • The queried slot repeats vold Rold times, rebinds to vnew, and places the query ∆D statements later; other slots fill remaining positions.
  • The generator makes the current answer value occur exactly once before the query, removing a baseline noise term from measurements.
  • Both varied axes preserve the target function while changing the implementation cost of competing mechanisms.
  • Rold controls superseded-value repetition, while ∆D measures statements between final rebinding and query.
  • A positional copy rule has ceiling 1/|supp(∆D)|, falling from 0.333 to 0.059 as support widens, and models saturate that ceiling.

4 READING OUT THE RULE

The paper reads out aliased rules with a paired intervention that changes one cue while preserving the ground truth. It summarizes the intervention using the sign of a log-odds change across held-out documents.

  • The input intervention edits one rule’s prediction while leaving the ground truth unchanged.
  • The readout is a paired difference in log-odds for a common contrast value between edited and base documents.
  • Multiplicity inversion retains one vold copy and rewrites the others as vnew, making only RARITY flip while RECENCY remains unchanged.PRIMACY and first-occurrence positional accounts cancel between the two terms.
  • The primary measure is the fraction of 400 held-out documents whose ∆ has the expected sign, with an exact binomial test.Median and interquartile range are reported, but sign fraction is primary because per-document ∆ is heavy-tailed.

5 RESULTS

Held-in behavior cannot distinguish the competing mechanisms: all analyzed cells reach near-perfect accuracy, while intervention reveals redundancy-ordered but non-replicating readouts across seeds. The objective's indifference explains the variance, whereas row-level ordering remains interpretable.

  • All 75 analyzed runs reach accuracy ≥0.999, including the stratum where the trivial last-update heuristic fails.
  • Pooled sign fraction rises monotonically with redundancy through Rold = 12, but the row ordering breaks at Rold = 16 because one seed drives most of the decrease.Row means are 0.479, 0.768, 0.858, 0.962, and 0.879 for Rold = 3, 5, 8, 12, and 16.
  • 13 of 25 cells span more than 0.3 in sign fraction across seeds, with the widest range reaching 0.879.The widest cell has sign fractions 0.098, 0.477, and 0.977 across seeds.
  • The low-redundancy row is least determined: seed 0 averages 0.270, whereas seeds 1 and 2 average 0.538 and 0.630.The fifteen runs in this row span 0.098 to 0.977, and the ∆D axis has no interpreted ordering.
  • The multiplicity inversion predicts negative ∆ for occurrence-count mechanisms, yet pooled gated medians are positive at every Rold ≥5 in all three seeds.Four cell-level exceptions lie within 0.031 nats of zero.
  • The objective is indifferent between recency and rarity because both select the same value on every training document, so seed, checkpoint, and schedule variation can determine the readout.Seed spread reaches 0.879, within-run checkpoint variation 0.303, and schedule changes move cells by up to 0.55.

6 SHORTCUT DYNAMICS

Models first exploit a positional shortcut whose accuracy has a closed-form ceiling, then may abruptly acquire retrieval as redundancy increases. Probing before or immediately after circuit formation can therefore reverse the apparent mechanism without any accuracy change.

  • 6.1 THE PLATEAU HAS A CLOSED-FORM CEILING AND MODELS SATURATE IT: A fixed backward token offset is correct only when ∆D takes one particular value, giving a positional ceiling of 1/|supp(∆D)|.The answer offset is 4∆D + 6 because each statement occupies four tokens.
  • 6.1 THE PLATEAU HAS A CLOSED-FORM CEILING AND MODELS SATURATE IT: Plateau models saturate the positional ceiling within 2% for narrow support, while wider-support cells reach 70–73% of their falling ceilings.
  • 6.2 ESCAPE IS A CIRCUIT FORMATION EVENT: During the plateau, the copy diagnostic remains 0.001–0.046 while sequence accuracy is 0.077–0.326, indicating that retrieval has not been acquired.
  • 6.2 ESCAPE IS A CIRCUIT FORMATION EVENT: Escape is marked by a near-chance-to-above-0.95 jump in the copy diagnostic and sequence accuracy reaching 1.000; loss derivatives independently show one interior peak.
  • 6.2 ESCAPE IS A CIRCUIT FORMATION EVENT: Escape timing decreases monotonically with Rold across four separable levels in both seeds, with peak positions falling from thousands of steps at Rold = 3 to 400 steps at Rold = 16.Peak height correlates with escape step, with average-rank Spearman ρ = 0.933.
  • 6.3 ATTRIBUTION ON THE PLATEAU REVERSES THE CONCLUSION: Gating on circuit formation is necessary but insufficient: one run remains frequency-type for 7000 steps after copy accuracy reaches 1.000, then crosses while accuracy stays 1.000.Its sign fraction moves from 0.23–0.36 to 0.86–0.94 without a corresponding loss peak near the crossing.

7 DISCUSSION

When RECENCY and RARITY are coextensive, the objective cannot select a mechanism, so mechanism identity is not reproducible at one cell. Escape timing and shortcut behavior remain measurable, but attribution requires circuit-formation gating and still needs replication across runs.

  • Mechanism identity: Coextensive rules leave mechanism choice underdetermined by the loss, making the run—not the single cell—the unit of replication.The construction makes this indifference exact because RECENCY and RARITY select the same value on every training document.
  • Measurement validity: 32 of 75 runs yield a confident reversed attribution before the retrieval circuit exists, despite no behavioral signal of failure.Task accuracy therefore cannot substitute for a circuit-formation gate.
  • Measurement validity: Gating on circuit formation is necessary but insufficient because even a gated single-checkpoint readout can report either direction.Mechanism identity remains unresolved when the objective is indifferent between alternatives.
  • What survives: Escape timing is monotone in redundancy across four levels in seeds 0 and 1 and three in seed 2, while mechanism identity at one cell is not.The timing result is recovered from an unthresholded quantity, alongside the positional shortcut’s ceiling and a within-cell dose-response.
  • Implications: Aliased constructions expose underdetermination, while natural data can conceal the same indifference behind confident single-seed measurements.The paper leaves comparison with a nonaliased control open because every cell is aliased by Equation 1.

8 LIMITATIONS

The study’s limitations concern optimizer dependence, collinear covariates, excluded causal comparisons, residual slot-identification cues, and restricted model and task coverage. These boundaries prevent separating several causes of the observed spread or generalizing beyond the tested setting.

  • Optimization dependence: Doubling the cosine-schedule length moves terminal sign fraction by up to 0.55, so the endpoint depends jointly on corpus and optimizer.The study reports one budget and schedule and cannot determine how much cross-seed spread reflects incomplete convergence.
  • Collinear covariates: The length–sparsity complex remains unresolved: matching slot count recovers 85–97% of the sign-fraction gap but only 17–74% of the median gap.Matching at fixed length by moving pupdate removes almost none of the dispersion difference, and different pooling statistics disagree.
  • Causal scope: Two data knobs admit no causal readout, so the corresponding settings are excluded rather than reported as mechanisms.Explicit update markers change sequence length, while history-indexed queries make the target rule no longer the last value.
  • Generalization: The experiments use one architecture family and one synthetic task family, with depth and capacity covarying at fixed width.Escape ordering and within-cell gradients survive tested depths, but per-cell readouts vary with depth as much as with seed.
  • Reproducibility: The appendix measures generator properties before training and provides code reproducing the tables, generator, edits, and per-cell readout logs.These implementation materials do not broaden the tested architecture or task scope.

A.1 INVARIANTS AND FAILURE MODES

The generator was designed to prevent memorization and trivial positional or token-copy solutions, while documenting residual imperfections and loader or probe checks. Collision-rate measurements verify which candidate rules are aliases, complements, or partial shortcuts.

  • Generator invariants: bindMax is bounded by 4, supporting that bindings are not memorizable across documents in the infinite-stream construction.Raw occurrence counts would be misleading because repeated queried-slot values occur within a single document by design.
  • Generator invariants: The realized ∆D–length correlation stays at |r| ≤0.043 in every cell, preventing distance from being confounded with document length.The ∆D axis therefore varies update-to-query distance without an associated token-count signal.
  • Generator invariants: ansLast is zero in all 25 cells, so the answer is never the last value token before the query by construction.Values are sampled without replacement within each document, removing a trivial copy baseline.
  • Generator invariants: updDens remains within [0.946, 1.102] over filler slots, while the queried slot is excluded because its forced rebinding would otherwise violate the diagnostic.This preserves approximately uniform update density in the measured filler region.
  • Residual confounds: Queried-slot antecedent distance exceeds filler-slot distance by 2–22%, so the fifth generator invariant holds only partially.The largest ratio is 13.32/10.92 = 1.22 at Rold = 3, ∆D = 2, matching the construction’s 40/3 spacing.
  • Failure modes: A reused dataloader seed can create an apparent phase transition through memorization, detectable when training loss falls below the analytic entropy floor while held-out loss rises.Comparing training loss with the closed-form floor is required before interpreting a transition in that curve.
  • Probe checks: Bit-identical reruns show that changing probe density does not alter training loss, while training and terminal probes agree within 0.012 on average and 0.060 at worst in sign fraction.The two probes therefore report the same scale without contributing shared-stream numerical nondeterminism.
  • Collision rates: RECENCY and RARITY collide at 1.000 in every cell, whereas FREQUENCY and PRIMACY are complements with collision rates of 0.001–0.309 and 0.000.The table separates exact aliases from candidate rules that agree with the target only under degeneracies.

B.1 THE QK-NORMALIZATION GAIN

QK normalization places attention-logit scale under learnable per-dimension gains, and the γ = 1 arm tests whether this scale affects retrieval formation. The results treat the arm as an observation without identifying a mechanism, while depth preserves escape-time ordering but not readout ordering.

  • Normalization setup: Per-head QK normalization uses learnable query and key gains, initialized uniformly at γ = 2.0.The gains apply across 1024 query/key components.
  • Logit requirement: At sequence length n = 224, achieving attention probability p = 0.9 requires a pairwise logit gap of 7.6 under equal distractors.The requirement is sufficient rather than necessary when particular positions can already be suppressed.
  • Logit requirement: At initialization, the target–distractor cosine-gap requirement is 0.95 for γ = 1 versus 0.24 for γ = 2.The paper sets γ = 2 so Equation 5 has margin for every n ≤224.
  • Gain arm: At γ = 1, terminal logit bounds of 20.5–36.9 do not explain why the Rold = 3 circuit fails to form while Rold = 16 forms quickly.The Rold = 3, γ = 1, seed-1 run reaches 36.9 without forming retrieval, whereas Rold = 16 crosses the gate at step 1000.
  • Limitations: The γ = 1 arm is reported as an observation without a mechanism because endpoint gains cannot distinguish early binding, cold-start difficulty, and gain–formation interaction.The paper states that separating these accounts requires gain trajectories, which the arm does not record.
  • Interpretation and depth: At γ = 2, optimization lowers terminal mean gains below initialization, while depth preserves coarse escape ordering but produces no stable readout ordering.The depth arm reports Rold = 3 as the slowest cell, but fixed-seed readouts change sign across depths.

C EDIT VALIDITY IN DETAIL

The edit-validation analysis checks that the causal contrast changes only the intended multiplicity cue and preserves key document invariants. It also derives the positional shortcut’s ceiling and identifies conditions under which that bound or the plateau analysis becomes unreliable.

  • Edit invariants: The edit preserves token count, answer position, ground-truth answer, statement count, and verified antecedent distance on retained documents.Violations are discarded, contributing to edit coverage below 1.00.
  • Edit checks: Idealized predictors produce ∆ = 2κ for the target-rule predictor and ∆ = 0 for the ground-truth predictor, matching +16.0 and |∆| < 10^-9 at κ = 8.These values validate the construction rather than constitute a model result.
  • Edit checks: A trained model before circuit formation yields approximately zero readout, supporting that the edit itself does not perturb logits without a mechanism to move.At Rold = 3, ∆ is +0.000 with 0.49 positive documents at step 2000 and −0.002 with 0.40 at step 4000.
  • Controls: The filler-slot control reads −0.009 to +0.015 across the grid, generally below 1% of targeted readouts exceeding 1 nat.Some negative-readout cells have matching control signs and are therefore treated as edit noise or uninterpretable.
  • Controls: The off-distribution mass gate requires contrast-pair mass of at least 0.5 per document, with analyzed-grid mass ranging from 0.59 to 1.00.The gate prevents cell means from concealing documents where the model has left the task.
  • Positional ceiling: The positional answer offset is 4∆D + 6, independent of document length, so a fixed-offset rule is correct for only one ∆D value and is bounded by 1/|supp(∆D)|.The bound requires ∆D to be unobservable from other features; length correlation or slot-marking could invalidate the positional interpretation.
  • Scope boundaries: Four Rold = 2 cells remain outside the analyzed grid because retrieval fails within budget, while the ∆D = 5 cell reaches neither positional nor retrieval behavior.Its terminal accuracy is 0.0018 with copy diagnostic 0.003, so escape time is not monotone in band width.
  • Observed plateau heights: The two |supp| = 3 cells saturate the positional ceiling within 2% and agree to 0.0002, while wider-support cells reach only 70–73% or less.The ratio is reported because both observed occupancy and the ceiling fall with support width.

D.4 THE Rold = 2 CELLS

The excluded Rold = 2 cells expose a two-phase failure mode: several models remain positional, and probing them can produce frequency-like readings despite absent retrieval. These cells are excluded because circuit formation, rather than accuracy, determines whether attribution is valid.

  • Exclusion decision: Four of five Rold = 2 cells never acquire retrieval within 12000 steps: three terminate POSITIONAL and one terminates NEITHER.The cells were excluded from the main grid before reporting rule readouts.
  • Probe artifact: The three positional runs return strongly frequency-type probe readings, including −0.47, −1.30, and a pilot −0.94.Their copy diagnostics are near chance, so the readings cannot represent a contest between recency and rarity.
  • Probe artifact: A positional explanation is plausible because the edit rewrites copies at fixed statement positions and can flip the value selected at offset 4∆D + 6.The explanation cannot be tested from data-side quantities because the model’s chosen offset is not identified by a distinct ∆D mode.
  • Interpretation: The plateau readings are therefore reported as artifacts with unidentified mechanisms, and identifying the mechanism would require reading the model’s offset directly.The exclusion does not affect the main results, which depend on cells where the retrieval circuit forms.

E PER-CELL STATE CLASSIFICATION

All analyzed runs form retrieval circuits, but per-cell attribution is stable only in the interior; boundary readouts vary across seeds, checkpoints, and budgets despite unchanged task accuracy.

  • State classification: All 75 runs pass the circuit gate and terminate in the RETRIEVAL state, excluding nothing within the analyzed grid.The gate primarily excludes the Rold = 2 column and timestamps readouts rather than selecting among analyzed runs.
  • Circuit timing: Escape time is monotone in Rold, with gate row means reaching 1000 steps by Rold ≥8 in seeds 0 and 2.Seed-specific row means remain later at lower redundancy, while the gate saturates at higher Rold.
  • Budget dependence: Interior sign fractions reproduce across an 8× budget range, whereas the boundary at Rold = 3, ∆D = 2 shifts by 0.55.Interior medians still vary without budget ordering, while boundary location is budget-dependent.
  • Checkpoint stability: Post-formation sign fractions remain unstable: within-run standard deviation reaches 0.303, and one trajectory crosses zero while accuracy stays 1.000.Across 69 runs, the median within-run standard deviation is 0.069; the largest occurs in an Rold = 3, ∆D = 3 run.
  • Replication: Doubling the schedule narrows ranges insignificantly overall, while the widest cells remain broad and non-convergent over available checkpoints.The mean range falls from 0.335 to 0.305, but 13 of 20 cells narrow at p = 0.077 and the widest cells do not consistently narrow.
  • Replication: Slot-matched results move toward the Rold = 16 readout, but only the sign fraction clears the replication criterion across seeds.The median spans a fourfold spread across seeds, while the sign fraction covers 85%, 97%, and 95% of the reference gap.

G.2 SLOT COUNT AT FIXED LENGTH

The fixed-length pupdate arm matches slot count while changing filler-side update density, but it cannot isolate slot count and does not reproduce the shortened arm.

  • Design: The pupdate arm matches slot count at 11.51 versus 11.55 while keeping both legs at 205 tokens.Training and evaluation otherwise match the main grid, and pupdate is the only changed parameter.
  • Results: Both pupdate legs increase the readout, contrary to the prediction that matched slot counts should make them converge toward each other.At Rold = 8, sign fractions are 0.154, 1.000, and 0.997; at Rold = 16 they are 0.997, 1.000, and 1.000.
  • Interpretation: The pupdate arm cannot isolate slot count because its parameter also changes filler-side update density while leaving dispersion nearly fixed.The product of dispersion and realized redundancy is constant within 4%, dispersion changes under 6%, and slot count changes by factors of 1.64 and 1.71.
  • Interpretation: Across both legs, slot count and filler update density move in opposite directions while the readout rises in both, ruling out monotone dependence on either covariate alone.This excludes simple monotone accounts but does not identify what the readout follows.
  • Caveat: One Rold = 8 seed-0 readout is uninterpretable because its control shares the negative sign and reaches 36% of the median.The run’s 0.154 sign fraction is therefore not treated as a reversal.
  • Replication: All six arm runs form retrieval circuits, and the within-cell redundancy gradient matches each cell’s effect direction in every run.The negative case also follows its own direction, at −0.063 versus −0.047.
  • Fixed-width control: Holding the positional ceiling at 1/9 preserves cross-seed spread, including a 0.766 range at d = 1.Because the ceiling is identical, the spread cannot be mediated by differences in positional-shortcut payoff.
  • Fixed-width control: The fixed-width arm does not order the row by either gradient or escape time, so it distinguishes neither account.The d = 16 loss-derivative peak varies sevenfold within cell, and row means have no ordering.

H SUMMARY STATISTICS

The analysis prioritizes gated sign fractions because means are heavy-tailed and medians can be capped by contrast-pair mass; distributional checks clarify low- and high-redundancy behavior.

  • Readout definition: The primary dependent variable is the expected-sign fraction over 400 held-out documents, supplemented by medians and interquartile ranges.An exact binomial test accompanies the sign fraction; the sign fraction is preferred because it is insensitive to magnitude and scale.
  • Summary statistics: At Rold = 5, ∆D = 5, the mean is +0.176 while the sign fraction is 0.52, because four documents exceed +8 nats.The binomial test gives p = 0.48, illustrating mean sensitivity to heavy tails.
  • Distribution shape: At Rold = 3, distributions concentrate near zero with a thin negative tail, whereas Rold = 16 spans 0 to 12 nats.The shape difference is consistent with low-redundancy instability rather than establishing an opposite mechanism.
  • Distribution shape: All five Rold = 3 cells are unimodal, so their negative median is not explained by a mixture of two rule-using populations.Two adjacent secondary local maxima remain within counting noise.

I ATTEMPTING TO MEASURE THE GEOMETRY

A gradient-based geometry test reproduces the readout but does not provide independent evidence about its mechanism, because the constructed directions and loss geometry are not diagnostic.

  • Measurement: The geometry measurement uses loss gradients from fresh documents and probe gradients, removes the loss-gradient component, and evaluates both losses along the normalized orthogonal direction.Evaluations use ε = 3 × 10^-4∥θ∥ across twelve checkpoints in two cells.
  • Validation: The readout reproduces at 16000 steps, matching medians and sign fractions from the main tables in both cells.The values agree to two decimals: −0.0874 / 0.126 and +1.8218 / 0.998.
  • Validity window: The comparison is restricted to ten checkpoints where contrast-pair mass is at least 0.5 and the loss gradient is estimable.The estimability criterion requires half-batch cosine at least 0.3; one pre-gate checkpoint is discarded.
  • Geometry result: Across the valid window, the aliased direction has |cos(gL, g∆)| ≤ 0.018 against chance 2.0 × 10^-4, and its curvature is one to two orders lower than along the loss-gradient direction.Removing the loss-gradient component changes ∥g∆∥ by less than 1.7 × 10^-4.
  • Geometry result: Collision rate does not order the seven readouts: within-checkpoint rank correlations average −0.023, with four positive and six negative values.The constrained readouts exceed the aliased one in 8 and 7 of 10 checkpoints, but neither comparison is significant.
  • Limitation: The measurement is not independent evidence because rejected-gradient curvature reflects the bulk loss surface, while saturated token loss makes the readout nearly orthogonal to its gradient.The cosine between readout and loss gradient falls from 0.82 and 0.64 at step 400 to 0.03–0.13 at steps 14000–16000.
Loading 2608.24460v1…