Source-linked AI summary
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
Simon Lam-Muir
TL;DR
Behavioural thresholds, internal accessibility, and training-time development may capture different events, leaving the acquired computation unidentified. This paper separates these observables in a controlled recurrent-depth reasoner across two training surfaces and finds divergent acquisition trajectories, pre-arrival answer accessibility, and a non-identifiable developmental accessibility curve.
Problem
Behavioural thresholds, final checkpoints, and hidden-state readouts can disagree without identifying the computation the model acquired.
Method
The study uses an oracle-defined closed relational world and a recurrent-depth transformer, combining held-out behavioural trajectories with preregistered probes and controls.
Results
Three-hop competence required 70 versus 13,055 logical epochs across symbolic and verbal surfaces, while weak answer information was accessible before behavioural arrival.
Takeaways & Limitations
Behavioural competence, internal accessibility, and training-time development are distinct observables, and neither behaviour nor decoder accessibility identifies the acquired computation.
Takeaways & Limitations
Results concern one closed synthetic world and recurrent-depth architecture class, with no automatic generalisation to frontier language models.
Abstract
from arXiv · showhide
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. These quantities need not identify the same event. We study a 30M-parameter recurrent-depth relational reasoner in a closed, oracle-defined world, using dense behavioural trajectories, two training surfaces, preregistered pre-arrival hidden-state probes, prospectively checked evaluability, and explicit untrained and negative controls, holding the training-time and inference-time axes separate throughout. Behaviour first: under one frozen acquisition criterion, three-hop competence cost 70 logical epochs on the symbolic surface and 13,055 on the verbal surface, a 186.5-fold contrast, after which verbal four-hop competence cleared in 8 logical epochs. Across the 13,055-epoch grind, four-hop held-out behaviour never exceeded 3/40 and ended at 0/40. Internal measurement next: on the verbal surface a linear probe recovered future-answer identity before behavioural arrival at 0.056159 against uniform chance 0.025, an untrained control of 0.024758 and a population frequency baseline of 0.048309 (p = 0.012987; 16/40 answer classes contributing). Analogous pre-arrival accessibility survived the surface change, reaching 0.1020 against a zero-step control of 0.0460 (p = 0.000999) at the upstream structural position and 0.0618 at the readout comparator (p = 0.004), with 21/40 classes contributing. Finally, the natural attempt to track that accessibility across training was not cleanly evaluable: probe eligibility is defined by behavioural arrival, so the measured population changes with the measurand. Behavioural competence, internal accessibility, and training-time development are distinct observables, and neither behaviour nor decoder accessibility identifies the computation training acquired; causal intervention is the necessary next step.
1 Introduction
The paper argues that behavioural thresholds, hidden-state accessibility, and retrospective checkpoint analysis are distinct measurements that need not identify the computation a model acquired. Using a mechanically checkable relational setting, it shows why these measurements are insufficient without causal intervention.
- Measurement problem: Behavioural thresholds, internal probes, and retrospective checkpoint analyses are three common but potentially discordant measures of capability acquisition.Their disagreement does not reveal which measurement corresponds to the computation the model acquired.
- Setting: A closed oracle-defined relational world enables exact checks of item novelty, path overlap, and answer-class coverage across two training surfaces.The two surfaces express the same underlying relational task family differently.
- Behavioural development: Changing only the bundled training surface produces radically different acquisition costs and developmental shapes that a single competence threshold collapses to one point.The reported shapes include long grinds on one surface and rapid races on the other.
- Internal accessibility: Pre-arrival future-answer information is linearly accessible from hidden states before behavioural production on the verbal surface and at two preregistered symbolic positions after a surface change.The accessibility is weak and heterogeneous rather than confined to one surface.
- Identification limits: Tracking pre-arrival accessibility across training is not validly estimable because probe eligibility is defined by the behavioural event under study, so causal intervention is the necessary next step.The paper also reports a quarantined execution and an earlier validation branch that correctly returned no result.
2 Experimental setting and measurement framework
The study uses a closed, mechanically verifiable relational world and a 30M-parameter recurrent-depth transformer whose fixed acquisition criterion and matched training surfaces support controlled measurement. Its framework separates training time from inference time and requires prospectively frozen, population-supported hidden-state probes with explicit reference controls.
- Closed-world task: Every answer and intermediate is mechanically known in a fixed-seed closed world, enabling exact novelty checks and verified zero-overlap held-out batteries.Evaluation items can be checked against training exactly rather than by estimation.
- Model and acquisition: 30M parameters, a 768-dimensional state, and a repeatedly applied 4-layer block decouple inference computational depth from parameter count.Curriculum stage k trains on depths at most k and clears only when held-out accuracy strictly exceeds 0.95.
- Matched surfaces: Two separate trained realizations use symbolic and verbal surfaces while holding the world, curriculum, architecture, optimiser configuration, and acquisition criterion fixed.Comparisons therefore concern matched realizations, not one model evaluated twice.
- Two clocks: Training time indexes checkpointed parameter updates, whereas inference time indexes recurrent iterations within one forward pass; claims about one clock do not identify the other.This separation is required for interpreting internal measurements and developmental analyses.
- Internal measurement: Linear probes are fitted independently by layer at registered event-relative positions under a frozen fold map and scored with uniform chance, population-frequency, and untrained zero-step references.The zero-step control is reconstructed deterministically from recovered historical provenance, but byte-identity with the original initialisation is unverified.
- Prospective evaluability: Analyses run only after the population is mechanically shown in advance to support the intended minimum sufficient statistic and estimand.This design principle prevents running planned analyses merely because activations exist.
3 Behaviour has hidden structure
Behavioural acquisition has distinct trajectories that a single threshold obscures: the verbal surface required a prolonged three-hop grind, then cleared four-hop competence rapidly, while held-out later-depth behaviour remained at floor. These curves also expose limits of threshold-based and extreme-value developmental inference, motivating internal measurement without establishing what computation was acquired.
- Acquisition trajectories: 13,055 versus 70 logical epochs: three-hop competence differed sharply between verbal and symbolic surfaces under one frozen criterion, a 186.5× contrast.The matched task realizations produced developmental trajectories of different shape, not merely different ratios.
- Acquisition trajectories: 8 logical epochs after 13,055 epochs at k = 3, the verbal realization cleared k = 4.This sequence does not establish whether the earlier grind purchased reusable machinery or whether four-hop capability was already present but unexpressed.
- Held-out behaviour: 3/40 peak four-hop accuracy: across 1,643 frames of the verbal k = 3 grind, held-out four-hop behaviour ended at 0/40 against a 0.95 criterion.Four-hop accuracy was exactly zero in 71.3% of frames, while five-hop accuracy was zero in 93.3% and also ended at 0.
- Inference limits: A single 3/40 frame among 1,643 caused the registered classifier to label k = 4 as distributed-climb, showing that maximum-based statistics misdescribe long developmental series.The k = 5 label likewise depended on an exact threshold equality, underscoring the fragility of transporting classifiers validated on short series.
- Inference limits: Behaviour measures expression, not availability: prolonged absence of later-depth competence-scale behaviour does not establish absence of internal development.This motivates measuring inside the model, while eventual criterion crossing says little about the route taken to it.
4 Internal accessibility is a different observable
Pre-arrival answer-relevant information was weakly but reproducibly accessible in hidden states on both surfaces, yet accessibility was distinct from behavioural arrival and did not identify the acquired computation. Registered probes therefore measured a different observable from local interior-read ordering and decisive readout commitment.
- Observable distinction: 93.25 content-sensitive differential exceeded a parse-control baseline of 88.5 and a preregistered confound boundary of 89.5 while novel two-hop accuracy was 0.025.The 3.75 clearance was a local descriptive ordering observation, not an estimate of mechanism acquisition or an accessibility curve.
- Verbal-surface accessibility: 0.056159 held-out accuracy exceeded uniform chance 0.025, an untrained zero-step control of 0.024758, and a population frequency baseline of 0.048309 before behavioural arrival.Restricted permutation yielded p = 0.012987, with 16 of 40 answer classes contributing.
- Interpretive limits: Accessibility is not use: linear probes can exploit distributed partial information without establishing a completed answer representation, causal use, or identification of the acquired computation.The probes concern final-answer identity in hidden states, whereas earlier readout work concerned composed intermediates exposed through the tied vocabulary readout.
- Cross-surface accessibility: 0.10201 at the symbolic primary position exceeded its zero-step control of 0.04598, with p = 0.000999, and analogous accessibility survived at both registered locations.The upstream structural position carried the stronger signal; the two locations had separate preregistered thresholds.
- Cross-surface accessibility: 4.1× uniform chance and 2.2× its control describe the symbolic figure, while the verbal result sits +0.785 above its frequency baseline; no registered cross-population comparison was performed.The symbolic primary was +5.60 points above its frequency baseline, but the paper does not claim a statistically larger effect than the verbal result.
5 Measurement has hard limits
Longitudinal pre-arrival probing is not cleanly evaluable because behavioural eligibility changes with the behaviour being measured, confounding representational change with population migration. This design limitation is distinct from execution-integrity and support failures, and does not establish that accessibility fails to develop.
- Longitudinal identifiability: Probe eligibility changes as items achieve behavioural correctness, so checkpoint-wise probe curves confound representational change with population migration.The failure mode is selection coupled to the measurand: the measured population changes because it improves.
- Longitudinal identifiability: 0 of 136 windows at width 2 met the fixed-panel requirement of forty classes with at least two items each; wider windows likewise yielded none.At width 10, zero of 128 windows qualified; at width 50, zero of 88; across the whole stage, zero of 1.
- Interpretation: The result is not cleanly evaluable, rather than a null result or evidence that accessibility fails to develop.The design cannot measure the developmental question under longitudinal eligibility.
- Interpretation: Three failure modes remain distinct: longitudinal eligibility makes the estimand non-identifiable, quarantine reflects execution-integrity failure, and insufficient holdout support prevents a defined statistic.Only the first is a claim about identifiability.
- General lesson: Observation, including refined decoder access, does not identify the computation that training acquired.Behaviour may miss internal accessibility, accessibility may precede behaviour, longitudinal probing may be non-identifiable, and decoder success does not establish causal use.
6 Related work
Prior work shows that behavioural emergence can mask gradual internal change, probe accessibility need not reflect use, and representation can affect compositional performance. This paper separates behavioural competence, inference-time accessibility, and training-time development, identifying a specific evaluability failure in longitudinal measurement.
- Behaviour and internal change: Prior reverse-engineering and circuit-level studies show continuous mechanism formation and stabilisation beneath discontinuous behavioural emergence.These studies track how circuits appear and stabilise across training and scale.
- Longitudinal probing: Longitudinal checkpoint probing has measured when linguistic, factual, commonsense, and reasoning information becomes accessible, but this work targets changing probe eligibility.Liu et al. report linguistic knowledge is acquired quickly and stably, whereas reasoning abilities are not stably acquired.
- Training-stage analysis: Counterfactual and activation-patching studies show latent-reasoning faithfulness depends on training stage, supporting the insufficiency of final checkpoints.The cited study uses a different latent-reasoning architecture and task, without the recurrent-depth surface comparison studied here.
- Present contribution: The paper’s contribution is separating three observables and showing that longitudinal accessibility is unidentifiable when the measured population is defined by behavioural arrival.Under matched conditions, Act 1 measures 13,055 against 70 logical epochs at one depth, a 186.5× contrast, while Section 5 establishes the evaluability failure.
- Interpretability limits: Probe-decodable information need not be used, while decoder choice and hidden-state location can affect interpretability in recurrent-depth models.Prior work reports recoverable properties that are unnecessary for the trained task and strong dependence on layer index and decoding method.
7 Discussion
The discussion separates behavioural competence, internal linear accessibility, and training-time development as related but non-interchangeable observables. It concludes that causal computation remains unmeasured and requires intervention, not improved decoding.
- Behavioural competence, internal linear accessibility, and training-time development are distinct quantities that can come apart.Behaviour is what the model produces, accessibility is what a decoder recovers, and development tracks changes across checkpoints.
- Matched relational tasks showed radically different acquisition trajectories across two surfaces, but the experiment does not isolate which surface property caused the difference.Vocabulary, grammatical structure, relation order, and sequence length varied together, so decomposing their effects requires additional arms.
- Final checkpoints cannot distinguish a long grind from a short race, while behaviour-dependent probe populations can invalidate probe curves unless evaluability is established in advance.Behavioural thresholds conceal route structure, and probe curves may reveal information behaviour does not express.
- The causal computation is not measured; the next test is to manipulate oracle-defined intermediate semantics and assess whether predicted recipient computations result.The proposed intervention transplants an intermediate state while preserving the recipient’s remainder and tests consistency across entities and depths.
8 Limitations
The study’s conclusions are scoped to one closed synthetic world, one recurrent-depth architecture class, and a frozen forty-class answer universe. Its weak, heterogeneous probe effects and separate surface realizations support descriptive accessibility findings, not broad generalization, completed representations, causal use, or shared cross-surface mechanisms.
- Scope: The results concern one closed synthetic world and one recurrent-depth architecture class, without automatic generalization to frontier language models.The answer universe is frozen at forty classes, with neither unseen-answer nor unseen-head generalization tested.
- Probe heterogeneity: 21/40 and 16/40 classes contribute to the probe effects, while the thinnest classes contribute nothing, making pooled figures non-uniform across the class space.The reported effects are weak and strongly heterogeneous across classes.
- Surface comparison: The two surfaces are separate trained realizations rather than one model observed twice, so their comparison is descriptive only.No registered cross-population effect-size comparison between surfaces was performed.
- Inference limits: Linear accessibility does not establish a completed representation, causal use, shared representation across surfaces, or a shared transition law.These limitations constrain what can be inferred from decoder accessibility and cross-surface comparisons.
9 Conclusion
Behavioural thresholds conceal developmental structure: matched realizations differed sharply in acquisition time, while weak answer information was accessible before behavioural arrival across surfaces and positions. Yet observing internal information still does not identify the computation training acquired.
- Conclusion: 70 and 13,055 logical epochs were required for the same competence depth across training surfaces, followed by an 8-epoch race.The realizations matched in world, architecture, curriculum, and acquisition criterion.
- Conclusion: Weak answer information was linearly accessible before behavioural arrival, surviving a surface change and appearing at two preregistered positions.This shows that internal accessibility can precede behavioural competence across measurement locations and training surfaces.
- Conclusion: Internal observation reveals more than behaviour, but it does not identify what computation training acquired.The conclusion distinguishes richer internal observation from identifying the learned computation.
Data and code availability
The companion repository provides the paper’s instruments, registered specifications, cryptographic-digest receipts, proofs, and figure-generation scripts, while executed registrations and receipts are released with the paper.
- Repository contents: The companion repository contains instrument sources, registered specifications, per-result receipts with cryptographic digests, population and support proofs, and figure-generation scripts for every panel.The repository is available at https://github.com/primeca libre-research/ltg-replication-receipts.
- Released materials: Executed registrations and receipts supporting the paper’s reported claims are released with the paper.
- Sealed experiments: Prospectively registered but unexecuted experiments remain sealed until execution.