Source-linked AI summary
A Pre-Specified Construction-Confirmation Test of Operation-Level Causal Transfer Across Finite Isomorphic Symbolic Domains
Xinyi Shan
TL;DR
The paper asks whether activation interventions move an operation-level structure between isomorphic symbolic domains, rather than merely changing answers. It estimates per-input operation contrasts and applies them to mapped recipient inputs with controls and independent confirmation. One pre-specified candidate passed both PyVene splits and was numerically replicated by NNsight, with confirmation Holm-adjusted p-values of 0.006943 and 0.007141.
Problem
The study asks whether an activation intervention transports a reusable operation between symbolic domains, beyond demonstrating task performance, decodability, or answer changes.
Method
The design applies per-source-state operation contrasts to mapped recipient inputs, compares operation-specific and no-op controls, and isolates construction from confirmation.
Results
One pre-specified candidate passed both PyVene splits, and a frozen NNsight replication reproduced the selected route across construction and confirmation.
Takeaways & Limitations
The evidence supports an existence result for one route replicated across two intervention implementations on one model revision and layers 20–21.
Takeaways & Limitations
The claim is limited to one route-bound candidate replicated across two intervention implementations on one frozen model revision and one frozen layer interval.
Abstract
from arXiv · showhide
Behavioral accuracy, linear decodability, and successful activation interventions do not by themselves show that a model carries an operation-level structure from one symbolic domain to another. We ask a narrower question in finite isomorphic state spaces: if the hidden-state difference between two operations is estimated separately for each source input, does adding that difference to a mapped recipient input move the model toward the corresponding recipient answer? The design compares this input-specific intervention with wrong-operation, norm-matched random, and no-op controls, and separates candidate construction from an independently isolated confirmation split. On a frozen Qwen2.5-7B-Instruct model at layers 20--21, one route--domain--operation candidate from a family pre-specified and frozen before confirmation access, transparent | integer_mod16--letters16 | successor->predecessor, passed both PyVene splits; its confirmation intersection--union p-value was 0.000198 and its 36-family Holm-adjusted p-value was 0.006943. A subsequent NNsight 0.7.0 experiment, pre-specified and frozen before its confirmation access, tested only this selected prompt route, without candidate or layer reselection. It reproduced all 12 confirmation effect estimates, confidence intervals, and exact sign-flip p-values numerically; its 36-family Holm-adjusted p-value was 0.007141. The result is therefore limited to one prompt route and one candidate, replicated across two intervention implementations on one model revision and one layer interval. It does not establish cross-model generalization, full-family backend independence, domain-general transfer, or algebraic invariance.
1 Introduction · 2 Related Work
The paper distinguishes causal operation transfer from behavioral success, decodability, and generic activation effects, using input-specific operation contrasts and stringent controls. It situates this test among prior work on steering, symbolic mechanisms, causal abstraction, isomorphic transfer, cross-format arithmetic, and preregistered confirmation.
- 1 Introduction: The study asks whether an operation-specific hidden-state difference, estimated per source input, transfers to a mapped recipient input while surviving alternative-explanation controls.The intervention transports a difference between operation-conditioned source representations rather than copying a successful donor activation.
- 1 Introduction: The narrow contribution combines per-source operation contrasts, recipient-per-input interventions, wrong-operation and norm-matched random controls, and isolated construction-confirmation adjudication.Exactly one pre-specified candidate passed both PyVene splits, and the selected route was separately confirmed with NNsight on the same model revision and layer interval.
- 2.1 Function vectors, task vectors, and activation steering: Prior vector-based methods show that internal directions can encode tasks or behaviors and causally alter outputs, but they do not by themselves distinguish fixed task directions from operation transfer.Function Vectors, Task Vectors, In-context Vectors, and activation engineering motivate the distinction while including related causal or contrastive interventions.
- 2.2 Successor mechanisms and algorithmic primitives: Related studies analyze successor heads, symbolic successor/predecessor transformations, and reusable algorithmic primitives composed through activation operations.These works provide precedent for studying ordered-token mechanisms and primitive representations.
- 2.3 Causal abstraction and invariance: Causal abstraction emphasizes alignments and interchange interventions, while prior work shows causal effects need not remain invariant under input-format changes.This motivates testing transfer claims beyond probe accuracy or isolated intervention success.
- 2.4 Othello and isomorphic-domain transfer: Othello studies use interventions and token-remapped or isomorphic games to examine representational alignment, corresponding actions, and causal steering across domains.These studies connect causal representation analysis with cross-language and isomorphic-domain transfer.
- 2.5 Isomorphic procedures and algebra transport: Research on analogical reasoning, isomorphic procedures, and algebra transport examines geometric or functor-like relational transfer and preservation of specified operational laws.The present framing belongs to this broader literature on transferring structure across representations.
- 2.6 Cross-format arithmetic mechanisms: Naganna et al. identify arithmetic heuristic neurons across symbolic arithmetic, natural-language problems, and Python code, using causal keep-only and knockout interventions.Their cross-format transfer comparison copies donor activation states into failed target executions, differing from the present operation-difference transport.
3 Preliminary Design Stages and Pre-Specification · 4 Methods
The study separated preliminary capability and design gates from a frozen causal test across finite isomorphic domains. Methods specified input-dependent contrasts, matched controls, independent confirmation, and multiplicity-adjusted adjudication, followed by a selected-route replication.
- 3 Preliminary Design Stages and Pre-Specification: P3A checked task and measurement feasibility, while P3B defined domains and operations but stopped before causal testing when behavioral screening failed.Neither preliminary stage supplied causal evidence for P3C; P3C fixed candidates, mappings, controls, eligibility, split separation, and the 36-candidate Holm procedure before formal causal data were examined.
- 4.1 State spaces, operations, routes, and candidates: The formal design used three 16-state domains, three frozen operations, two formal prompt routes, and 36 candidates spanning domain pairs and ordered operation transfers.The abstract states were binary4, integer_mod16, and letters16; the diagnostic abstract route could not promote candidates.
- 4.2 Input-dependent operation contrasts: Each source input supplied its own operation difference, which was added to the mapped recipient representation at the corresponding state without reusing a donor activation or broadcasting one direction.Patching used the single residual-stream vector at assistant_output_boundary, with source–recipient directions measured separately.
- 4.3 Control Interventions and Norm Matching: Each bucket contained 38 records: one real contrast, one shuffled-operation control, 32 matched-random controls, and four no-op paths.Matched-random and shuffled controls were accepted only when FP16 L2 norms matched the real contrast exactly at the bit-pattern level.
- 4.5 Construction and Independent Confirmation: Confirmation data remained inaccessible until construction fixed the candidate set and layers, after which eligible formal candidates were measured under the same statistical family.Diagnostic results remained separate and could not promote a candidate; split isolation constrained adaptation to confirmation data.
- 4.6 Effects and Statistical Adjudication: The analysis prespecified 12 components per candidate, clustered-bootstrap confidence intervals, exact state-level sign-flip tests, and a 36-candidate Holm correction.A candidate passed only if all 12 lower confidence bounds were positive, all four no-op checks passed, the Holm-adjusted p-value was at most 0.05, and it was measured.
- 4.7 NNsight Selected-Route Replication Protocol: The NNsight 0.7.0 chain tested only transparent| integer_mod16--letters16|successor->predecessor with frozen layers, controls, estimands, tests, and the retained 36-member multiplicity family.It was not a second 36-candidate discovery or screening pass; unmeasured family members received p = 1 before Holm correction.
5 Results
Independent construction and confirmation splits produced one intersecting candidate: transparent|integer_mod16--letters16| successor->predecessor. This selected route passed causal and control tests in both PyVene and NNsight, with matching 12-component confirmation effect objects.
- Confirmation dataset: 243/288 formal exact accuracy, 95/144 diagnostic exact accuracy, and 338/432 overall exact accuracy were obtained in confirmation.Confirmation contained 432 unique clean-prediction records, and three of six eligible formal candidates passed Holm correction.
- Split-specific candidate selection: |Pconstruction ∩Pconfirmation| = 1, with the unique intersecting candidate transparent|integer_mod16--letters16| successor->predecessor.Construction had five passes and confirmation had three; split-specific pass sets were intersected only after independent adjudication.
- Causal and control records: 1,344 no-op effects were exactly zero across 12,768 causal/control records.The dataset included 336 real, 336 shuffled-operation, 10,752 matched-random, and 336 records for each of four no-op classes.
- PyVene causal effects: 0.0001983642578125 intersection–union p-value and 0.0069427490234375 36-family Holm-adjusted p-value accompanied the selected route.Aggregate normalized recovery was 0.55358 [0.45874, 0.65043], while real-minus-matched-random recovery was 0.51961 [0.42152, 0.62057].
- NNsight replication: All 12 NNsight confirmation estimates, 95% confidence bounds, and exact one-sided sign-flip p-values matched the PyVene confirmation object.The equality applies only to the pre-specified 12-component confirmation effect object, not backend-internal raw tensors or independently sampled data.
6 Pre-Specification, Split Isolation, and Reproducibility
The study froze key design choices before confirmation access and isolated construction from confirmation to reduce outcome-dependent decisions. A standard-library-only audit reproduced the reported numerical traces, but it was an internal reproducibility check rather than external replication.
- Design motivation: Preliminary stages shaped the formal experiment by motivating exact-output eligibility, state-level causal measurement, and controls for propagation, magnitude, operation identity, and pipeline effects.These stages changed the design rather than contributing evidence to the formal result.
- Pre-specification and split isolation: The 36-candidate family, source–recipient map, layer rule, controls, component tests, multiplicity correction, and diagnostic boundary were specified before confirmation access.The choices were preserved in a private content-addressed audit trail, not a public preregistration registry; construction and confirmation used separately computed eligibility and independent adjudication.
- Reproducibility audit: The standard-library-only audit reproduced 12 component statistics, IUT and Holm values, figures, tables, and manuscript numeric traces without rerunning model inference.It reconstructed the candidate family and split-specific pass sets and verified the clean and causal/control record structure.
- Reproducibility audit: The audit was an internal reproducibility check, not external replication.It recomputed the paper’s numerical claims without importing the original adjudicator or rerunning model inference.
7 Discussion
A single pre-specified candidate passed both PyVene construction and confirmation splits and was replicated on the same route across PyVene and NNsight. Narrowly, the evidence supports an input-dependent, operation-level causal-transfer candidate, but remains route-, candidate-, model-, and layer-bound.
- Validated candidate: One formal candidate passed both the construction and independently isolated confirmation PyVene splits under pre-specified controls, a 12-component IUT, and 36-family Holm correction.The candidate was pre-specified and frozen before confirmation access.
- Validated candidate: The same frozen candidate, model revision, and layer interval passed a pre-specified NNsight selected-route replication, demonstrating replication across two intervention implementations.This was not a second full-family candidate screen and does not establish backend independence.
- Interpretation: The result is consistent with an input-dependent, operation-level, recipient-per-input causal-transfer candidate constrained by matched-random, shuffled-operation, zero/identity, and self-replacement controls.These controls constrain random-direction, wrong-operation, and no-op explanations but do not exclude every alternative mechanism.
- Interpretation: 0.37498 was the real-minus-shuffled recovery for D1→D2, versus 0.15868 [0.09283, 0.23190] for D2→D1; shuffled-operation control retained approximately 0.287 aggregate normalized recovery.The asymmetric margin means the evidence does not attribute all observed recovery to an operation-specific component.
8 Responsible-Use Statement
The study frames mechanistic interventions as useful for auditing but potentially misleading without validation, and limits its conclusions to finite synthetic symbolic tasks. Future applications require revalidation for the target model, backend, task distribution, and operational context.
- Risk Mitigation: Pre-specification, construction-confirmation isolation, explicit controls, and multiplicity correction mitigate risks of overstating causal explanations or steering model behavior.The safeguards include no-op and wrong-operation controls, with candidate and control families frozen before confirmation access.
- Scope Limits: The claim boundary excludes cross-model, domain-general, full-family backend-independent, and algebraic conclusions.These limits apply to the study’s causal-transfer claims.
- Scope Limits: The experiments concern finite synthetic symbolic tasks and do not evaluate deployment decisions or human-facing applications.The paper therefore does not establish effects in operational or user-facing settings.
- Future Use: Future use should revalidate effects for the target model, backend, task distribution, and operational context.Revalidation is required before extending the findings beyond the studied setup.
9 Use of AI Systems
The sole human author retained responsibility for the study and manuscript while AI systems assisted with ideation, implementation, validation, execution, statistical, and literature-search tasks.
- Human oversight: SHAN XINYI was the sole human author and accepted responsibility for the study and manuscript.She set the research question, scope, and decision criteria.
- Human oversight: The author authorized executions, approved evidence and interpretation boundaries, froze versions, and reviewed verification reports.She also approved the final interpretations and claims.
- AI assistance: AI systems assisted with research ideation, implementation and validation code, human-gated cloud execution, statistical pipelines, and literature search.The passage also identifies manuscript and figure dra… as assisted activities.
10 Code and Data Availability · 11 Conclusion
The manuscript provides public-safe materials for inspecting methods, component-level results, and deterministic numerical checks, but not for an independent end-to-end rerun. Its conclusion is limited to one route-bound candidate reproduced across PyVene and NNsight, without establishing universal or cross-model structure.
- 10 Code and Data Availability: Public materials include the manuscript source, embedded component-level tables, ancillary CSV files, and public-safe verification materials.These materials support inspection of reported methods, component-level results, and deterministic numerical checks.
- 10 Code and Data Availability: The private confirmation workspace, raw tensors, credentials, host identity, private paths, and governance records are not public.The available materials are not claimed to support an independent end-to-end rerun of all experiments.
- 11 Conclusion: One route-bound candidate passed construction and independently isolated confirmation in PyVene under a frozen 36-candidate family.Eligibility was split-specific, controls were operation-specific, and adjudication was multiplicity-corrected.
- 11 Conclusion: A separate NNsight 0.7.0 chain reproduced the same selected route across construction and independently isolated confirmation.The chain was pre-specified and frozen before confirmation access and reproduced the complete 12-component result set.
- 11 Conclusion: The result does not establish a universal internal language, cross-model generalization, backend independence, or full-family cross-backend replication.These limitations constrain the scope of the conclusion beyond the selected route and implementations.
- 11 Conclusion: The result does not establish preserved algebraic structure.The conclusion therefore remains restricted to the reported route-bound candidate and replication setting.
A Supplementary Tables … A.5 Confirmation record composition
The supplementary tables are mechanically generated from the frozen derived-data baseline and document stage roles, preliminary metrics, frozen execution identities, controls, and confirmation-record composition. Preliminary gates established feasibility but did not support the later 7B causal result, while the strict causal stage remained closed in P3B.
- A Supplementary Tables: The supplementary tables are mechanically generated from the frozen derived-data baseline, with machine-readable supplements S1–S4 provided as ancillary files rather than new analyses.This covers the stated role of the supplementary tables and ancillary files.
- A.2 Preliminary-stage process metrics: P3A achieved 0.9175 semantic accuracy, 0.9150 strict generation accuracy, and nine of ten task families passed on Qwen2.5-1.5B-Instruct.These results served only as a capability and measurement gate.
- A.2 Preliminary-stage process metrics: P3B recorded 0.6056 canonical exact accuracy, with 11/36 behavior cells and 33/144 atomic-pair cells passing during its 720-sample nonblind clean screen.The strict gate remained closed, so P3B did not enter its planned blinded causal stage.
- A.3 Frozen execution contract: The frozen execution contract identifies Qwen/Qwen2.5-7B-Instruct, model revision a09a35458c702b33eeacc393d103063234e8bc28, and a 14-file snapshot.The passage also records the snapshot manifest hash.
- A.3 Frozen execution contract: The execution contract fixes the PyVene backend at version 0.1.8, alongside Python 3.11.9, torch 2.11.0+cu128, and transformers 4.57.6.The backend commit and manifest hash are also recorded in the frozen contract.
- A.3 Frozen execution contract: The contract records tokenizers 0.22.2, an NVIDIA GeForce RTX 3090, the assistant-output-boundary patch site, and transformer-block model-instance details.The supplied execution identity also includes GPU driver information.
- A.4 Control roster: Table 5 provides the per-bucket and total control roster for the supplementary record.The supplied passage identifies the table’s scope but does not enumerate its rows.
- A.5 Confirmation record composition: Table 6 provides the clean and causal/control record composition for the confirmation record.The supplied passage identifies the table’s scope but does not enumerate its rows.
A.6 Full 36-candidate split adjudication · A.7 Twelve pre-specified components · A.8 Claim boundary
The paper adjudicates 36 formal candidates, reports 12 frozen component analyses, and restricts its supported conclusion to a single selected route replicated across two intervention implementations. It does not support broader claims about models, domains, operations, abstract routes, algebraic relations, or universal language.
- A.6 Full 36-candidate split adjudication: The full adjudication covered all 36 formal candidates in separate construction and confirmation status tracking.
- A.7 Twelve pre-specified components: The component analysis froze 12 estimates with 95% intervals and exact one-sided sign-flip p-values.
- A.8 Claim boundary: The supported candidate was transparent|integer_mod16–letters16|successor->predecessor.
- A.8 Claim boundary: The evidence is limited to route-bound, single-candidate replication across PyVene and NNsight on one frozen model revision and one frozen layer interval.
- A.8 Claim boundary: NNsight did not repeat full 36-candidate screening, so the result does not establish backend independence or full-family cross-backend replication.
- A.8 Claim boundary: The paper does not support cross-model generalization or claims covering all domains and all operations.
- A.8 Claim boundary: The diagnostic route had zero eligible candidates, so abstract-route causal confirmation is not supported.
- A.8 Claim boundary: Composition, inverse, involution, conjugacy, and other algebraic relation preservation remain untested, while “universal language” and “first word” are prohibited scientific claims.
A.9 Registered FP16 norm contract
The registered FP16 norm contract defines the current scientific claim, while confirmation-only positives and retry-development history are excluded from interpretation.
- Confirmation-only positives must not be interpreted or promoted into the current claim.
- Retry-development history is not part of the scientific claim under the final registered-norm contract.