Source-linked AI summary

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

Dekun Yang

arXiv:2608.13626v1cs.AIcs.CLcs.LG

TL;DR

The paper asks whether decodable and causally usable hidden state signals also support stable action maps that transfer to held-out sources and compose. Using an evidence lattice, positive-control carriers, and Qwen3-4B interventions, it finds that state availability, causal use, local geometry, and reusable closure separate, with composition remaining unsatisfied.

  • Problem

    Existing evidence shows state signals can be encoded and causally relevant, but does not establish stable action functions that transfer to held-out sources and compose.

  • Method

    The paper independently tests state availability, causal local use, and reusable closure using an evidence lattice, calibrated affine carriers, structured stresses, and model interventions.

  • Results

    Across tested carriers, state availability, causal use, local geometry, and reusable closure separate; early-layer and refit maps improve local fits but fail unchanged composition.

  • Takeaways & Limitations

    Calling an internal map a causally faithful reusable operator requires state availability, causal local use, action-structured geometry, and reusable affine closure together in one eligible carrier.

  • Takeaways & Limitations

    The evidence is restricted to one post-trained model family, an absorbing setter domain, four sampled final-token layers, and three causal datasets from one checkpoint.

Abstract

from arXiv · show

A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus .398 for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to .474 (.469 with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes.

1. Introduction

The study tests internal action maps with an evidence lattice that separates state availability, causal local use, geometry, and reusable affine closure. Across a Qwen/Qwen3-4B task and two finite worlds, the results support separable capabilities rather than a general reusable operator claim.

  • Measurement framework: The evidence lattice tests state availability, causal local use, and reusable closure/algebraic laws independently, requiring both branches for a “causally faithful reusable operator.”Geometry without an output effect and causal use without held-source closure remain informative but insufficient.
  • Calibration: A known contextual S5 carrier passes the held-source tests, while parity-support splits, smooth observation curvature, and held-domain conjugacy stress behavior under misspecification.The positive control checks implementation recovery before structured stresses probe robustness.
  • Grounded results: Within-domain cross-fitting improves h28 reconstruction, but randomized entity splits do not support entity-dominated maps; one-step reconstruction favors h4/h16, whereas paired interventions affect h28/h36.Action identity dominates the tested parameter geometry, while lexical and state-probe controls do not resolve whether h4 is a surface carrier.
  • Contributions: The study distinguishes reconstruction, causal use, action-dominant geometry, one-step specification sensitivity, relative algebraic discrimination, and global affine closure.Its method uses natural endpoints, held sources, distinct negative and blocked outcomes, and independent audits of formal follow-ups.
  • Scope: The conclusions apply only to the tested carriers and hypotheses and do not constitute a general denial of internal operators.The experiments cover one controlled language-model setting and two exact finite worlds.

2. Related Work · 3. Evidence Framework

The paper situates its contribution at the intersection of state tracking, affine and compositional representation methods, and algebraic world-model diagnostics. Its evidence framework separates reusable action-map closure from decoding, causal local use, relative laws, and preregistered evidence-status gates.

  • 2.1 State tracking and causal state use: Prior work shows that hidden activations can track dynamic entities and support interventions, while designed memory states update through learned action operators.These studies motivate testing both state availability and causal state use rather than treating decodability alone as sufficient.
  • 2.2 Linear transformations and composed functions: Affine relation decoders, task vectors, function vectors, and continuous compositional latents provide precedents for representing relations and multi-step computations.The paper distinguishes these broader representation mechanisms from action-conditioned maps evaluated on natural held endpoints.
  • 2.3 Algebraic structure and world models: Algebraic-structure studies examine associativity, homomorphism error, held-out transitions, drift, commutators, and structured recurrent interfaces, while behavioral work shows local prediction can coexist with incoherent world models.Probe controls likewise warn that measurements may reward fitting capacity rather than the intended construct.
  • 2.3 Algebraic structure and world models: The documented search through 12 August 2026 found no prior study combining known-algebra calibration, held-source natural endpoints, matched interventions, law tests, behavioral eligibility, and independent artifact audits.This is explicitly a search-bounded description of the integrated protocol.
  • 3.1 From a state signal to an action map: The framework defines h(s, c) and evaluates whether action maps transfer beyond observed pairs to withheld entities, contexts, histories, or source states using natural post-action activations.Alchemy uses row-relative endpoint error, while finite worlds use NRMSE against a frozen target-variance reference; identity has error one when displacement is nonzero.
  • 3.1 From a state signal to an action map: Decoding, action order, and absolute closure are separate tests: probes and order effects may succeed even when composed maps miss the held source’s natural final activation.Causal local use is separately tested with fixed probes, random-label controls, and matched wrong-state or translation interventions; success licenses a usable direction, not a naturally traversed transition.
  • 3.2 Absolute closure and relative laws: H1–H5 combine one-step controls, two-step endpoint closure, direct-map discrepancy, setter baselines, commutativity, inverses, and behavioral covariance rather than relying on order contrast alone.Thresholds, partitions, and inferential units were frozen in advance; NOT SUPPORTED, NOT TESTED, UNTESTABLE, and RIGHT-CENSORED distinguish failed gates, blocked experiments, missing antecedents, and budget-limited events.

4. Experimental Settings

Experiments used a frozen post-trained Qwen/Qwen3-4B checkpoint with final-token representations and tested action maps across grounded, synthetic, calibration, lexical, causal, and learned-world settings. The protocol emphasized held-source evaluation, prospective freezing, matched replication, and explicit separation of state availability, causal use, geometry, and closure.

  • Grounded model and protocol: The grounded setup froze post-trained Qwen/Qwen3-4B, using 36 layers, width 2,560, five container states, and final-prompt-token representations.Weights were never updated; actions emptied containers or filled them with one color.
  • Grounded model and protocol: h28 was frozen as the earliest sampled layer passing direct state decoding and a paired state-specific intervention, with splits holding out entities and generated instances.State categories and the two primary templates were not held out.
  • Grounded model and protocol: Three regenerated dataset seeds used 500 train, 100 validation, and 200 test pairs per action, with H2 covering 20 ordered distinct-action sequences.H3 matched same-entity and disjoint-entity pairs; the planned natural-language inverse test stopped after failed behavior manipulation.
  • Calibration and controls: Structured calibration varied nuisance ratios 0, .05, .10, .20, .40, and .80 across 20 seeds per ratio while scanning h4, h16, h28, and h36 with full and low-rank affine maps.Phase 8 tested support splits, observation warps, and held-domain orthogonal conjugacy at the same strengths.
  • Calibration and controls: Entity/context conditioning used seven cyclic outer splits, four training entities, one selection entity, and two held-out entities, alongside a five-fold within-entity reference.Stable-hash sampling retained 80 rows per entity-action cell; covarying entity, token, and context features prevent a pure entity mechanism interpretation.
  • Learned-world evaluation: The learned-world experiments used three decoder-only Transformer scales, three seeds, and 224 = 16.78 million training examples, evaluating longer trajectories, inverse loops, equivalent histories, and ordered action pairs.Only six medium/large trajectories entered the preregistered cross-world denominator; small models were capacity controls.

5. Results

Results separate state decoding and causal use from reusable affine closure. The known affine carrier passes all gates, whereas model representations show partial geometry, causal effects only at later layers, and no composition rescue from refitting or learned worlds.

  • Known affine carrier: 1.000 held-state, one-step, two-step, and inverse decoding, approximately 3.31×10−8/5.02×10−8/6.55×10−8 NRMSEs, and 1.000 commuting-pair AUROC passed every affine-carrier gate.The independent embedding retained every gate direction within the frozen 5% reproduction tolerance.
  • Structured perturbations: 23/30 maximum-strength seed-fold cells crossed at least one closure gate, while strength correlated perfectly with median one/two-step error in curved and held-domain families.At strength .80, curved median one/two-step NRMSE was .0118/.0127, versus .814/1.050 under held-domain conjugacy, with decoding remaining 1.000.
  • Model representation: 82.87% ± 1.81% h28 probe accuracy and a 2.334 ± .231-logit intervention effect coexisted with .5189 ± .0286 held-source affine error, versus .3979 ± .0130 within-domain cross-fit error.Only 4/15 original cells were below the frozen .50 gate, while within-test-domain cross-fit passed every cell.
  • Layer comparison: .299 h4 and .347 h16 mean one-step errors beat .519 h28 and .549 h36, but 0/6 early-layer cells passed the joint H2 gate and h4 conflict-state decoding stayed below the .80 criterion.Causal effects replicated only at h28/h36: h4 and h16 intervals contained zero, whereas all six later-layer intervals excluded zero.
  • Refitting and composition: .4686 weighted train-plus-validation refitting and .4744 unweighted refitting improved h28 one-step error, but every refit failed composition, with direct-map gaps of .8059 and .8199.Original composition endpoint error/direct-map discrepancy was .8110/.7945; the decisive gap remained far above .50 in all three datasets.
  • Learned finite worlds: Learned worlds reached shared-chart and task/state eligibility diagnostics, but K was not observed within 33.55 million examples; held-source one-step NRMSE was .786 ± .036 versus .260 ± .006 seensource.Held composition was .835 ± .022, teacher-forced composition .540 ± .015, and a separately fitted depth-one map .497 ± .030.

6. Discussion and Limitations

The discussion treats layer findings as separable partial outcomes: early layers support geometry but not causal use, whereas h28 supports causal use but not frozen reusable closure. It limits the claims to tested models, domains, layers, datasets, and function classes while proposing a reporting standard that keeps these branches distinct.

  • Interpretation of layer results: h4/h16 are geometry-positive and causal-use-negative, whereas h28 is causal-use-positive and frozen-closure-negative; no layer supports their conjunction.Choosing a layer for decoding or patching alone would overstate a local result as a model-wide mechanism.
  • Calibration limits: Structured held-domain conjugacy reaches grounded error ranges, but the curved family remains below the gates, so universal statistical power is not established.The limitation is part of the calibration result rather than evidence that the stress suite is universally decisive.
  • Alternative explanations: The h28 transfer gap does not establish pure entity binding: local within-domain fit, modest within-entity gains, and action-dominant frozen-map comparisons leave multiple effects entangled.The supported account is a distribution-dependent map with action-structured parameters, while entity, token, template, and contextual effects remain unresolved.
  • Protocol and closure: Outcome-aware refitting changes h28 H1 estimates under a different protocol, but its failure on H2 shows that better one-step interpolation does not imply reusable composition.The refit is a legitimate generalization estimate because it does not use test data, not a post hoc revision of the original decision.
  • Scope and limitations: The evidence is limited to one post-trained model family, an absorbing setter domain, four sampled final-token layers, correlated entity/context factors, and three causal datasets from one checkpoint.The conflict probe is weak, the learned-world K is right-censored, and nonlinear, attention-mediated, cross-token, and related alternatives remain unexcluded.
  • Reporting standard: The proposed reporting standard separates state availability, causal use, and held-source closure, reports natural endpoints and continuous error, and includes composition, commutativity, and inverse-cycle tests.It also distinguishes within-domain reference fits from held-domain mechanisms and requires behavioral eligibility, sufficient statistics, manifests, and independent audits.

7. Conclusion

The study recovers a known affine action algebra while showing that state availability, causal local use, action-structured geometry, and reusable affine closure can diverge across tested carriers and depths. Its positive claim is deliberately bounded: causal faithfulness requires their conjunction in the same eligible carrier, while richer operators remain possible.

  • Held-source measurement recovers a known affine action algebra, but structured domain mismatch can push its gates into the empirical failure range.
  • In Qwen3-4B, action-conditioned geometry, state decoding, and dataset-seed causal-use results appear at different sampled depths.Direct entity tests do not support a purely entity-specific account, lexical controls remain unresolved, and h28 one-step results vary with final fitting scope.
  • Within the tested carriers, state availability, causal local use, action-structured geometry, and reusable affine closure are distinct empirical constructs.An internal map should be called a causally faithful reusable operator only when these constructs are jointly supported in the same eligible carrier.

Data Availability

Source rows and reproducibility metadata are available with the arXiv version, alongside a compact bundle. Human or personal data were not collected, while large model artifacts are not redistributed.

  • Data availability: Source rows for Figures 2, 3, and S1–S4 accompany the arXiv version under anc/source_data/.Associated metadata JSON files include result, manifest, audit, and source-table SHA256 values.
  • Data availability: A compact reproducibility snapshot is provided as anc/reproducibility_bundle.zip.
  • Data availability: No human-participant or personal data were collected.
  • Data availability: Large model weights, raw activation tensors, intermediate checkpoints, and multi-gigabyte fitted artifacts are not redistributed.

Code Availability

The reproducibility bundle provides the materials needed to inspect and regenerate the experiments, while excluding the third-party state-probes code and large artifacts from redistribution.

  • Code Availability: The ancillary bundle includes preregistrations, frozen configurations, experiment code, wrappers, auditors, plotting scripts, tests, environment specifications, audit reports, and source tables.The state-probes submodule is identified by its upstream URL and pinned commit but is not redistributed.
  • Code Availability: Excluded large artifacts can be regenerated from recorded model identifiers and frozen configurations, with hashes and passports preserved in reports and metadata.This supports reproducibility without distributing the large artifacts themselves.

Ethics Declaration

The study used computational models and synthetic systems only, with no human participants, personal or clinical data, or animal research.

  • Ethics Declaration: The study involved pretrained and from-scratch computational models, procedurally generated prompts, and finite synthetic transition systems, without human participants, personal data, clinical data, or animal research.

Funding

This work received no external funding.

  • No external funding was received for this work.

AI-Assistance Disclosure · 8. Supplementary Results

OpenAI Codex supported implementation and drafting, while the supplementary results document threshold sensitivity, failed composite validation, composition-gate failure, and layer-scale diagnostics. Experimental claims and numerical values were checked against versioned artifacts and independent audits.

  • AI-Assistance Disclosure: OpenAI Codex assisted with code, experiments, validation, figures, evidence organization, and drafting; the authors retained responsibility for design, interpretation, and final text.Claims and numerical values were checked against versioned machine-readable artifacts and independent audit outputs.
  • 8.1 Threshold sensitivity without verdict movement: All 15 h28 action-by-seed cells first pass the frozen H1 threshold at .58, above the red .50 gate.The figure also varies H2 thresholds and reports Phase 7 h4 order-sub-gate and full-H2 counts.
  • 8.2 Failed H5 composite and collapse diagnostic: ρ= .091 with trajectory-cluster 95% CI [−.035, .249] links algebra violation and world incoherence across 99 nonzero checkpoints, while the preregistered H5 task gate fails.H3 remains diagnostic and cannot override absolute gates.
  • 8.2 Failed H5 composite and collapse diagnostic: A trajectory shows H1 error rising as collapsed state representations differentiate, reinforcing the diagnostic separation between state differentiation and algebraic validity.Signed margins are defined for H1, H2, AUROC, and H4, while H3 remains diagnostic.
  • 8.3 Layer-local composition gates: No layer-seed cell passes joint H2, so early-layer one-step geometry does not satisfy the frozen composition construct.Raw cells report endpoint error, direct-map gap, order d_z, probe gain, and joint H2 status; h28 rows are compatibility comparators.
  • 8.4 Layer scale and effective dimension: Scalar RMS normalization preserves the sampled-depth ordering while the early representations remain low dimensional.The figure reports residual norm, action displacement norm, covariance participation ratio, and train-fitted normalized one-step error across three activation datasets.
Loading 2608.13626v1…