Source-linked AI summary

Lagged Coupling: Internal Representations Become Readable Before They Become Causal

Xining Xun

arXiv:2609.01048v1cs.CLcs.AI

TL;DR

The paper asks whether readable internal representations become effective behavioral control surfaces, and tests that question across training and model scale with a pre-registered three-track design. It finds that internal readability reliably precedes behavioral readability and causal efficacy, while the single-onset hypotheses remain indeterminate and the observed lag does not shrink with scale.

  • Problem

    Prior work often treats readability from a representation as evidence that the representation is a useful behavioral lever, but when and at what scale this link develops has not received a pre-registered test.

  • Method

    The study tracks internal readability, behavioral readability, and causal efficacy across the Pythia training grid using frozen estimands, matched nulls, and a pre-registered decision tree.

  • Results

    Internal readability leads behavioral readability and causal efficacy: steering is inert in 43/48 cells, while headroom grows 2.0–56.8× and causal write-in remains ≤ 0.11% of headroom.

  • Takeaways & Limitations

    Probe accuracy should not be treated as evidence of steerability; representation formation can outpace causal readout consolidation.

  • Takeaways & Limitations

    The main claims rely on one Pythia suite, eight checkpoints leave sub-grid events unresolved, and both pre-registered onset axes resolve INDETERMINATE.

Abstract

from arXiv · show

Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as early as step 1,000 at every scale -- yet steering along that same reading direction remains null-equivalent in 43 of 48 model-checkpoint cells. Internal readability systematically outruns causal efficacy, and the lag does not shrink with scale. We call this structure lagged coupling and decompose it into three dissociable tracks: (i) internal readability, saturated (AUROC >= 0.990) from the first checkpoint everywhere; (ii) behavioral readability, which develops gradually and progressively later at larger scales (12B reaches 0.909 only at the final checkpoint); (iii) causal efficacy, almost always null-equivalent, occasionally counterproductive early, with one isolated positive pulse (12B, step 8,000, z = +2.49) our grid cannot resolve. The ordering is dominantly read-before-write (11/11 units, no inversion). Representation headroom along the probe direction grows up to 57x with training and scale while causal write-in stays below 0.11% of headroom -- the variable is increasingly written into the representation and increasingly ignored by the readout. Under a fully pre-registered protocol, both single-onset hypotheses resolve INDETERMINATE (scale slope +0.24, 95% CI [-0.60, +0.87]; time vote 3:3) -- a disciplined negative explained by the three-track decomposition. A pre-registered OLMo-2 replication preserves the direction at attenuated magnitude. Our results caution against inferring steerability from probe accuracy and establish a developmental bottleneck: representation formation reliably outpaces causal readout consolidation.

1 Introduction

The paper tests whether readable internal representations are also effective behavioral control surfaces, tracking readability and causal efficacy across training and scale. It finds a stable lagged-coupling structure in which internal readability leads behavioral readability and causal write-in.

  • Research question: Readable representations do not reliably function as causal levers, motivating a pre-registered test of when the two properties emerge.The study tracks internal readability, behavioral readability, and intervention efficacy simultaneously.
  • Protocol: A fully audited pre-registration protocol freezes estimands, nulls, decision rules, and data-integrity checks before evaluation.The protocol included a content-hash gate that caught a corrupted public checkpoint lineage.
  • Findings: Three tracks develop on different clocks: internal readability appears earliest, behavioral readability develops later, and causal efficacy remains mostly null-equivalent.This decomposition yields a stable cross-scale structure rather than a verdict for either single-onset fork.
  • Findings: 43/48 cells show inert steering, while four early cells show significantly harmful effects.The reported ordering is dominantly read-before-write, with no observed inversion.
  • Mechanism: 2.0–56.8× representation headroom growth contrasts with causal write-in remaining ≤ 0.11% of headroom.This provides a mechanistic account of why increasingly strong representations can remain behaviorally ignored.

2 Related Work

Prior work establishes that information can be probed and internal structures can be localized or edited, but developmental measurement of causal leverage remains limited. This paper extends developmental analysis from circuit formation to whether readable directions become behavioral levers.

  • Development and scaling: Scaling research characterizes changes with model size and data, while the timing and sharpness of capability emergence remain contested.Developmental analyses of internal mechanisms are comparatively rarer.
  • Development and scaling: Grokking and induction-head studies show that internal progress and circuit formation can occur at identifiable points during training.The paper extends this timing perspective beyond circuit formation.
  • Probing: Probes recover structure from hidden states, but probe accuracy can confound representation with probe capacity and task priors.The paper treats probe AUROC as one track and pairs it with a matched causal readout.
  • Causal intervention: Circuit-level methods localize behaviorally relevant structures and support causal mediation or weight-level intervention.These approaches motivate studying whether readable internal directions become effective intervention sites.
  • Developmental gap: Steering and editing research generally studies fully trained models, leaving developmental timing of causal leverage unmeasured.The paper reports that readable directions can precede causal efficacy and may not become causal across observed times and scales.

3 Pre-registered Setup and Decision Protocol

The study uses a frozen, pre-registered grid and decision protocol to compare internal and behavioral readability with causal efficacy across model sizes and checkpoints. Its intervention and measurement choices are designed to make developmental and cross-scale comparisons consistent and auditable.

  • Grid and instrumentation: 192 units span six Pythia sizes, eight training checkpoints, and four task families under a common frozen evaluation pipeline.The pipeline includes probe, output-level, intervention, and per-unit resolution tracks.
  • Estimands: Track 1 trains a logistic probe at a pre-registered site and evaluates inducibility with dev- and test-split AUROC.The pre-registered dev↔test discrepancy flag fired in only 3 of 192 units, all at 12B and all below 0.12.
  • Estimands: Track 2 measures the AUROC of the model’s own answer score against the gold label, independently of any probe.Internal and behavioral readability are distinct quantities by construction.
  • Estimands: Track 3 applies an additive, unit-norm edit along the frozen probe direction at one site and compares its paired effect with matched random-direction nulls.Causal write-in is normalized by representation headroom, while each unit’s noise floor determines whether an effect is resolved.
  • Decision protocol: The decision procedure tests scale and time axes with weighted regressions and takes an explicit INDETERMINATE branch when a fork fails.The protocol prohibits threshold lowering, estimand switching, seed completion, and post-hoc re-gridding; the OLMo-2 replication verdict is capped at “weak” with fewer than three sizes.
  • Decision protocol: All estimands, thresholds, nulls, decision rules, and degradation branches were frozen before criterion data were produced.Logged execution changes did not alter thresholds, nulls, or the decision tree.

4 Results

The preregistered onset analysis resolves INDETERMINATE on both axes, while the three-track results reveal a stable lagged-coupling structure: internal readability precedes behavioral readability and causal efficacy across scales.

  • 4.1 The pre-registered verdict: INDETERMINATE on both axes: The scale-axis slope is +0.24 with 95% CI [−0.60, +0.87], while the time-axis vote splits 3:3, so both preregistered axes resolve INDETERMINATE.No scaling law is claimed under the frozen decision rules.
  • 4.2 Three tracks, three clocks: Internal readability reaches ≥ 0.990 from step 1,000 at every size, whereas behavioral readability develops gradually and later at larger scales.The three-track decomposition separates probe readability from output-level behavioral readability.
  • 4.3 The ordering is read-before-write, and scale does not close the lag: Read-before-write ordering occurs in 11 of 11 pattern-1 units, with no inverted case, and the catch-up magnitude does not increase systematically with scale.The catch-up regression has 95% CI [−0.722, +0.438], crossing zero.
  • 4.4 Interventions are largely inert, occasionally counterproductive: Causal efficacy lies within the null band in 43 of 48 cells, with four significant early backfires and one positive 12B step-8k pulse of +2.49.The intervention curves remain flat across tested strengths in null cells, supporting an interpretation about the direction rather than a single dose.
  • 4.5 Mechanism: headroom outruns write-in by orders of magnitude: Representation headroom grows up to 56.8×, while causal write-in remains at or below 0.11% of headroom and absolute effects never exceed 0.016.The probe direction increasingly expresses the target variable, while downstream computation uses that component weakly in relative terms.
  • 4.6 The prior-bias arm: a mechanical positive result: The prior-bias effect grows with scale, with WLS slope 95% CI [+0.00062, +0.291], indicating larger models are harder to steer off their priors.The lower confidence bound is close to zero, and the per-unit effect is described as small in magnitude.

5 Discussion

The study separates readability from causal control, showing that probe-derived directions can remain poor behavioral levers despite strong internal signal. Its three-track account interprets this gap as a developmental bottleneck in which representation formation outpaces causal readout consolidation.

  • Implications for AI safety: Probe-derived directions are handles on behavior only asymptotically and weakly in the study’s grid.The discussion distinguishes activation-level interpretability’s central practical claim from the observed developmental pattern.
  • Implications for AI safety: 43/48 cells showed inert steering, while four early cells were significantly harmful.These results concern runtime activation control along reading directions.
  • Implications for AI safety: Monitoring can rely on Track 1, but intervention-based control relies on Track 3, which the data fail to support across nearly all training.The paper explicitly separates early-warning sensing from runtime activation control.
  • Developmental interpretation: The three-track decomposition rejects single-onset emergence claims because internal readability, behavioral readability, and causal efficacy follow different clocks.Internal readability is present from the beginning, behavioral readability develops later at larger scales, and causal efficacy does not consolidate monotonically.
  • Developmental interpretation: Representation headroom grows 2.0–56.8× while causal write-in remains ≤ 0.11% of headroom.The proposed bottleneck is that the probe direction increasingly occupies the representation while readout reliance does not keep pace.
  • Developmental interpretation: The isolated 12B step-8k positive pulse occurs during a behavioral transition and disappears after stabilization, but the coarse grid cannot establish a timing window.The rewiring-window explanation is explicitly presented as a hypothesis for future work.
  • Methodological contribution: The study’s methodological contribution combines preregistered decision rules, instrument checks, content-hash gates, audit ledgers, and restart-safe execution.Two audit gates fired during execution, including detection of a corrupted checkpoint lineage and recovery from a cloud restart without data loss.

6 Limitations and Threats to Validity

The paper limits its claims to a controlled runtime-activation setting and explicitly reports uncertainty from intervention strength, task breadth, statistical power, pulse interpretation, resolution, and checkpoint coverage.

  • Intervention scope: The intervention tests minimum-energy, signal-aligned single-site runtime edits, not weight-space editing or trained distributed interventions.Dose–response curves are flat across the strength grid, but broader intervention classes remain open.
  • Task scope: The four task families use controlled answer-format contrasts, leaving open-ended generation and multi-turn control for future study.The controlled design was chosen to make per-track estimands comparable across 192 units.
  • Statistical scope: The negative onset verdict is based on n = 6 sizes, although the dissociation, ordering, and headroom contrasts use n = 48, 191, and 24.INDETERMINATE means the preregistered onset hypotheses failed, not that the study measured nothing.
  • Pulse interpretation: The isolated pulse is one cell in 48, lacks a timing-window inference, and is excluded from the headline claims.Four same-signed, time-clustered early-backfire cells are treated separately from the pulse.
  • Measurement resolution: 0.00398 (160M) → 0.0328 (12B), so null-equivalence at 12B is evaluated against a coarser resolution floor.The paper publishes this calibration and defines null statements relative to each size’s floor.
  • External validity: The main claims rely on Pythia, while the two-size OLMo-2 replication is capped at a weak verdict and eight checkpoints leave sub-grid events unresolved.The reported structure is explicitly descriptive rather than a law.

7 Reproducibility and Audit

The reproducibility pipeline was deterministic, configuration-driven, and fully audited, with atomic recovery and content-hash verification protecting the developmental analysis from execution and lineage failures.

  • Deterministic execution: The pipeline derives unit seeds from md5 digests, disables bytecode caching, and writes results atomically with completion markers.Interrupted runs resume by skipping verified-complete units.
  • Deterministic execution: 174/192 test units resumed bit-identically after a cloud restart with zero data loss.The restart-safe design preserved the completed computation.
  • Auditability: The append-only audit ledger records the corruption quarantine, aggregation hotfix, manifest refresh, restart recovery, and final seal.The hotfix was logged with before-and-after hashes for touched files.
  • Lineage verification: A content-hash gate caught silently corrupted Pythia-2.8B checkpoint lineage before criterion data were produced.Independent community and lineage audits later corroborated the detection.

Appendix A The weight-corruption incident: a mandatory content-hash gate

Appendix A documents a silently corrupted Pythia-2.8B lineage and the mandatory hash-based detection, quarantine, and revision-pinning process used to protect the measurements.

  • Incident: Distinct Pythia-2.8B checkpoint labels contained byte-identical weight content, invalidating their apparent temporal distinction.The affected generation-5 standard revision was distinguished from the intact v0 lineage through lineage auditing.
  • Corroboration: Independent reports corroborated byte-for-byte identity across early steps 0–6,300 and reconstructed five generations of mirror copies.The external account matched the study’s detection chain point for point.
  • Response: The detection chain was content-hash gate → lineage quarantine → revision pinning.All reported 2.8B measurements use the verified-healthy pinned revision, while contaminated manifests were voided and refreshed.
  • Implication: Per-checkpoint content hashing is recommended as a hard gate because file-size and load-success checks cannot detect this silent failure mode.The authors characterize the failure as silent and common enough to affect a flagship checkpoint suite.

Appendix B Full causal-efficacy grid

The full 48-cell test grid finds causal efficacy overwhelmingly null-equivalent, with significant effects concentrated in early negative cells and one positive pulse. The corresponding 12B step-8k pulse also appears in the development phase.

  • 48 model×checkpoint cells were evaluated, with five reaching |z| ≥ 2: one positive and four negative, all negative cells early in training.The grid-wide result is summarized in Table B1.
  • +2.04 was observed for the development-phase counterpart of the 12B step-8k cell, alongside the test-phase pulse.The dev↔test probe-discrepancy flag fired in 3 of 192 units, all at 12B and below 0.12.

Appendix C Resolution, behavioral readability, and configuration summary

Behavioral readability develops gradually and later at larger scales, while internal probe readability remains saturated across sizes and checkpoints. The appendix also documents the pre-registered probe, intervention, null, and decision configuration.

  • Behavioral readability: Internal probe readability is saturated at approximately 1.000 across sizes and checkpoints, with a dev minimum of 0.990 at 12B.Behavioral readability is reported separately at the output level.
  • Behavioral readability: Behavioral readability’s developmental gradient steepens with scale across the complete four-size checkpoint series.The table reports output-level answer-score AUROC on the dev split.
  • Configuration summary: The protocol uses frozen-dev logistic probes, additive unit-norm edits on a frozen λ grid, matched random-direction nulls, and weighted 95% CI plus majority-vote decisions.Aggregates are computed mechanically under the stated pre-registered configuration.
Loading 2609.01048v1…