Source-linked AI summary
Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit
Anqi Peter Li
TL;DR
Order-dependent activation statistics can reflect intervention-site asymmetry rather than interaction. The paper introduces a no-fit audit using single-intervention baselines and an antisymmetrized second difference; across eighteen cells, the baseline explains 84.3–97.7% of bracket norm, while corrected residuals isolate mixed interaction only to second order.
Problem
Order-dependent activation statistics are often read as interaction, although where interventions enter the network can confound that interpretation.
Method
The audit predicts the open-path order-swap from four single-injection responses and uses the antisymmetrized second difference to remove first-order and pure self-curvature terms to second order.
Results
84.3–97.7% of the bracket norm is explained by the single-intervention baseline across eighteen cells, while the self-curvature term is 1.8–5.2× larger than the corrected residual at the primary configuration.
Takeaways & Limitations
Run the single-intervention prediction before interpreting an order-swap as interaction or geometric structure; if it explains the bracket, form the second difference instead.
Takeaways & Limitations
Residuals for the six language models remain finite-scale mixtures because empirical bounds on third derivatives were not supplied, and the claims are scoped to activation-space interventions at distinct sites.
Abstract
from arXiv · showhide
Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network. We introduce a no-fit site-asymmetry audit. For a twice-differentiable readout, the open-path order-swap decomposes into a canonical additive response measured by single interventions and an antisymmetrized second difference free of first-order and pure self-curvature terms to second order. Across six open-weight language-model families, the single-intervention baseline explains 84.3-97.7 percent of the bracket norm (mean 93.7 percent), while the no-interaction self-curvature term is 1.8-5.2 times larger than the corrected residual in the two families with the plus/minus injection split. The corrected residual clears a generic-interaction null in three of six families under a confound-free prompt split and two of six after configuration robustness. A known-positive surrogate recovers planted mixed interaction, while a matched site-separation test changes the baseline share and a random architecture reproduces the first-order regime. The same estimator transfers to released non-language references: trained residual fractions fall below a fixed Gaussian-direction null in 11/12 contrasts (5/6 ViT-B/16, 6/6 ResNet-50), a portability check rather than pooled evidence. The contribution is a reusable measurement criterion: run the single-intervention baseline before reading an order-swap vector as interaction or geometric structure; if it explains the vector, form the second difference instead. All claims are scoped to activation-space interventions at distinct sites; we do not claim that representation geometry is globally Abelian.
1. Introduction
The paper introduces a no-fit site-asymmetry audit showing that activation-space order swaps can be dominated by a canonical additive response rather than interaction. It recommends forming an antisymmetrized second difference after measuring the single-intervention baseline.
- Contribution: The audit separates construction-forced site asymmetry from model-dependent interaction without fitting parameters.It uses single-intervention responses to predict the order-swap bracket before interpreting residual structure.
- Mechanism: The first-order artifact arises as (J1 −J2)(vi −vj), which is linear in each intervention and generically non-zero under site mismatch.This mechanism also persists in closed loops and does not require interaction between the directions.
- Results: 84.3–97.7% of the bracket norm is explained by the first-order baseline across eighteen family-by-configuration cells.The mean baseline share is 93.7%, while the second difference contributes 0% of that bracket decomposition by definition.
- Second-order correction: The self-curvature term is interaction-free yet 1.8–5.2× larger than the corrected residual at the primary configuration.A two-term fit therefore combines self-curvature and mixed interaction rather than measuring interaction alone.
- Validation: The corrected residual clears a generic-interaction null in 3 of 6 families under the confound-free design.Supporting controls include a known-positive surrogate and a randomly initialized network reproducing the first-order regime.
2. Related Work
Related work places this audit alongside studies showing that order-dependent or scale-dependent statistics can reflect first-order structure or uncalibrated measurement pipelines. The cited weight-space analogues remain outside the paper’s activation-space criterion.
- Interaction theory: Prior work proves that second-difference interaction equals a Hessian bilinear form and vanishes for locally affine maps.The paper adopts this theory for its corrected residual Q.
- Methodological context: Studies across transformers and other model analyses report first-order predictability, unstable pairwise composition, or artifacts from uncalibrated pipelines.These works motivate randomized or calibration-based controls related to the paper’s audit.
- Weight-space comparisons: Weight-space Lie-bracket and ordering statistics are methodological neighbours but remain outside the paper’s activation-space criterion.A first-order correspondence between spaces makes agreement closer to entailment than independent replication under the analyzed conditions.
3. Method: the estimator and its first-order term
The estimator separates site-dependent first-order effects and pure self-curvature from mixed interaction in activation-space order swaps. It measures single-site responses to form a parameter-free baseline, then uses the residual as a finite-scale interaction diagnostic.
- First-order artifact: Distinct downstream maps make order-swap brackets carry a first-order term even without interaction.The term is (J1 − J2)(vi − vj), so site mismatch acting on direction differences can generate a nonzero bracket.
- No-fit estimator: The residual equals the antisymmetrized mixed second derivative plus an uncontrolled O(α3) remainder at finite injection scale.Known-Hessian surrogates recover planted interaction, but model-side residuals remain finite-scale mixtures rather than exact interaction estimates.
- Second-order decomposition: The bracket’s second-order coefficient also contains self-curvature that is quadratic in each direction without coupling them.Only the mixed term couples the two directions; a two-term fit therefore combines self-curvature and interaction.
- Controls: Linear antisymmetric structure has the form W(vi − vj), but finite-data fitting leaves slack that must be tested rather than assumed away.The paper therefore measures J1 and J2 directly instead of relying only on a high-capacity fitted correction.
- No-fit estimator: The canonical baseline bLij uses matched single-site responses and requires no fitted parameters.It predicts the bracket from s1 and s2, while the antisymmetrized second difference removes site-Jacobian and pure self-curvature terms to second order.
- Controls: Closed-loop statistics are not automatically protected: distinct-site four-leg loops can retain first-order residuals without interaction.Cancellation requires each return leg to re-enter at its outgoing site with an identity composition.
4. Results: validating the audit
The audit validates a strong additive, first-order explanation for order-swap brackets while isolating a smaller corrected residual and a larger no-interaction self-curvature term. Robustness tests support only a limited interaction claim.
- Baseline validation: The additive baseline reaches cos(bL, Br) ∈ [0.9955, 0.9986] across six families, leaving 5.2–7.6% of the bracket norm unexplained.Across eighteen family-by-configuration cells, the cosine never falls below 0.9869.
- Curvature control: The no-interaction second-order term is 1.8–5.2× larger than the operator extracted from the residual at the primary configuration.This comparison covers Llama and OLMo across three injection configurations.
- Interpretation: The fitted W(vi −vj) correction removes linear structure but cannot establish that the remaining nonlinear residual represents interaction.The paper states that a residual can be mostly a term with no interaction in it.
- Null validation: A random-init residual network reproduces the high baseline cosine across ε = 0.01–4, supporting an architecture-level first-order null rather than a training-specific explanation.The reported cosine range is [0.9956, 0.9988].
- External reference: The estimator transfers to trained vision references, falling below a fixed Gaussian-direction null in 11/12 contrasts: 5/6 ViT-B/16 and 6/6 ResNet-50.The paper treats this as portability evidence rather than pooled evidence.
- Interaction evidence: The corrected residual clears the generic-interaction null in 3 of 6 families under prompt splitting, but the behavioural direction survives multiplicity correction in only 2 of 6.The surviving candidate interaction is therefore reported without overstating its stability.
Appendix A. Protocol details
The appendixed protocol specifies released checkpoints, direction extraction, held-out readouts, site choices, and reproducibility boundaries across language and vision references.
- Language models: The six language-model checkpoints are base, non-instruct models from DeepSeek, Gemma, Llama, Mistral, OLMo, and Qwen.The exact checkpoint identifiers are listed in the protocol.
- Reproducibility: Most analysis runs CPU-only from committed per-pair tensors, but Table 2 requires rerunning forward passes because single-injection responses are released only as scalar summaries.The external CPU reference is fully rerunnable from released raw held-out outputs.
- Vision protocol: The external protocol uses a CPU residual MLP, torchvision ViT-B/16, and torchvision ResNet-50 with fixed public checkpoints and a fixed Gaussian-direction null.It also specifies class-balanced direction extraction and disjoint held-out readout images.
- Design dimensions: The six families use residual widths Dout ∈ {3584, 4096}, while 16 trait contrasts produce 120 unordered pairs per family.The pair count follows from the stated contrast construction.
- Seeds and alignment: Disjoint prompt samples make cross-seed agreement measure direction re-extraction rather than rereading a single fit, with mean direction reliability 0.784.The paper cautions that downstream statistics can transform or normalize direction errors.
- Numeric provenance: A provenance script binds every printed number to a named result-record field and rejects perturbed decoy values.This is presented as protection against value-proximity certification errors.
Appendix B. Delimiting the claim: full treatment
The full treatment evaluates corrected residuals with overlap-stratified cosine statistics and progressively stronger controls. The strongest supported claim is confined to three families under the confound-free design, and two after intersecting robustness controls.
- Diagnostic choice: The signed corrected statistic is near zero, but that collapse is entailed by antisymmetry and is not evidence.Its reported mean is −0.0100 with coherence 0.070.
- Shared-argument statistic: Rsh compares mean absolute cosine for pairs sharing one trait against pairs sharing none; the observed value is 2.132.Reference values are 0.999 for constant-Jacobian slack, 1.324 for generic antisymmetric interaction, and 6.058 for trait-varying first-order Jacobian.
- Confound control: Disjoint prompt halves reduce the mean from 2.132 to 1.602 to 1.391, leaving only Llama, OLMo, and Qwen above the recalibrated null of 1.325.The paper therefore claims the effect in 3/6 families under the confound-free design.
- Robustness boundary: Intersecting configuration, prompt, trait, and seed controls leaves Llama and OLMo as the families surviving every control.The full records and per-family strata are reported in the associated materials.
Appendix C. The zero-parameter first-order measurement
The appendix measures first- and second-order contributions without fitting, showing that self-curvature can exceed the corrected residual and that finite-scale residuals are not automatically interaction estimates.
- Finite-scale certificate: Third-order Taylor remainders make the finite-scale residual a mixture unless empirical bounds on M, M1, and M2 are available.The certificate follows from applying third-order remainders to both orderings and four single-site paths.
- Second-order split: The exact identity bL = L + S is an arithmetic self-check, not evidence that the second-order split is interaction-free or clean.The reported closure is 1.9 × 10−7 relative, consistent with float32 rounding.
- First-order measurement: The zero-parameter prediction leaves 5.2–7.6% of the bracket norm at the primary configuration across families and configurations.The residual is the antisymmetrized mixed second derivative plus an unbounded O(α3) remainder, so it is not an interaction estimate without remainder control.
- First-order measurement: At one weaker configuration, DeepSeek reaches cosine 0.9869 with a 15.7% residual, the largest reported residual in the table.The passage notes the co-location with failure to exclude the generic null without identifying them as the same effect.
Appendix D. Nulls, robustness, and residual geometry
The appendix separates exact identities from empirical evidence, compares residuals with calibrated nulls, and tests whether conclusions survive prompt, trait, configuration, and site controls.
- Residual geometry: The identity Brij = (Dij − Dji) + bLij separates the raw bracket’s first-order artifact from the second difference’s corrected residual.Readers seeking interaction should measure D directly rather than interpret the raw bracket.
- Null construction: The generic-interaction and trait-varying-Jacobian nulls match the observed residual fraction, whereas the constant-Jacobian arm is an exact floor rather than a competitor.The constant-Jacobian residual is only 0.007–0.023 of the bracket versus the observed 0.029–0.099.
- Null comparison: The observation exceeds the generic-interaction null in 5 of 6 families and remains far below the trait-varying-Jacobian model in 6 of 6.The null arms use brackets from true directions and fits on independently estimated directions.
- Prompt-split replication: The cross-half mean is 1.391 against a recalibrated null of 1.325, with Llama, OLMo, and Qwen clearing it.DeepSeek, Gemma, and Mistral remain within 0.02 of pure linear slack.
- Trait robustness: A disjoint 14-trait inventory preserves null clearance for Llama, Mistral, and OLMo, with mean retention 1.04±0.07.Size-matched subsampling slightly deflates rather than inflates the result, with design effect 0.969.
- Site and configuration robustness: The effect collapses to approximately 1 when compared layer pairs share no site, while shared deeper-site comparisons retain 2.733, 2.101, and 1.946.After all controls, Llama and OLMo remain; the effect is therefore scoped to specific site configurations.