Source-linked AI summary
Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
Nura Aljaafari, Andre Freitas
TL;DR
Mechanistic interpretability lacks a shared vocabulary for how transformer circuits compose across tasks and architectures. The paper introduces CPC, which describes computation through four operator roles and evaluates their signatures across 15 models from five architecture families. The results support CPC as a comparative vocabulary while showing that operator depth and geometric expression remain architecture-specific.
Problem
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures.
Method
CPC models transformer computation as iterative coherence refinement through alignment, unification, suppression, and routing, and evaluates predictions using weight-space signatures, activation measures, and causal tests.
Results
Across 15 models from five architecture families, operator signatures correlate with held-out activation-level role measures above random baselines, while suppression is more stable across tasks than unification.
Takeaways & Limitations
CPC provides a shared vocabulary for comparing transformer mechanisms and post-training effects, with predictions that should remain conditional on architecture.
Takeaways & Limitations
CPC is not a mechanistic proof, covers decoder-only English models, and uses a coherence proxy that conflates semantic coherence with residual-stream geometry.
Abstract
from arXiv · showhide
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
1. Introduction
Mechanistic interpretability has identified individual transformer circuits, but lacks a shared account of how their functions compose. CPC addresses this gap with four operator roles and evaluates their signatures across models and tasks.
- Motivation: Existing circuit analyses characterize mechanisms individually, but their broader functional relationships remain underspecified.The gap concerns how identified mechanisms fit within a broader framework of transformer computation.
- Framework: CPC models transformer computation as iterative coherence refinement through alignment, unification, suppression, and routing.Alignment proposes candidate relations, unification integrates compatible assignments, suppression prunes incompatible alternatives, and routing carries the selected interpretation to readout.
- Framework: The framework treats operator roles as a higher-level vocabulary for components whose correspondence is not one-to-one.One component may instantiate several roles, while one role may be distributed across multiple components.
- Evaluation: The evaluation tests signature validity, suppression’s dependence on alignment, contradiction sensitivity, and post-training changes in operator structure.These dimensions correspond to four empirical questions posed in the introduction.
2. Coherence in Language and Cognition
Coherence concerns how linguistic and cognitive elements fit into a structured whole through supportive and conflicting relations. Constraint-based formulations represent this as selecting a mutually compatible configuration under weighted positive and negative constraints.
- Linguistic coherence: Linguistic coherence links clauses, sentences, or discourse spans through relations such as elaboration, cause, and contrast.These relations explain how one discourse unit is interpreted with respect to another.
- Linguistic coherence: Local coherence includes neighboring-span relations, discourse-entity continuity, and lexical or semantic continuity.These sources describe how nearby text remains connected and interpretable.
- Linguistic coherence: Semantic opposition does not necessarily imply incoherence when the relation between opposing contents is explicitly represented.Contrast, correction, disagreement, and concession can remain coherent under this view.
- Constraint satisfaction: Constraint-based coherence assigns weighted positive and negative relations among elements and seeks a partition maximizing total satisfied weight.Positive constraints reward mutually supported elements, while negative constraints represent conflict.
- Constraint satisfaction: Exact maximization of the coherence objective is computationally difficult, motivating connectionist approximations with excitatory and inhibitory interactions.The approximation progressively separates incompatible alternatives while reinforcing mutually supportive elements.
- Constraint satisfaction: In discourse, elements may be propositions, entities, events, or spans, with coherence depending on the configuration of their supportive and incompatible relations.This formulation bridges epistemological and linguistic perspectives.
3. Coherentist Probabilistic Compositionalism
CPC frames transformer computation as iterative coherence refinement in the residual stream, using four functional roles: alignment, unification, suppression, and routing. Its global functional rewards coverage and local compatibility while penalising incompatible assignments, but operator implementations may overlap across components and layers.
- Coherence refinement: CPC models transformer computation as layerwise construction and refinement of an interpretation in the residual stream.The framework treats depth as approximate coherence-refinement steps rather than an explicitly solved constraint problem.
- Coherence state: The coherence functional rewards coverage and local fit while penalising grounded fragments that cannot be jointly maintained.Scoh(g | m, G) measures grounding fit, while the contradiction kernel vanishes for compatible pairs.
- The four operators: Alignment identifies candidate relations, unification incorporates compatible content, suppression reduces incompatible alternatives, and routing carries selected content to the readout.These roles correspond to QK interactions, additive writes, subtractive writes, and mover-style routing, respectively.
- The four operators: Routing changes where information is available without directly changing the coherence functional.It controls whether resulting content can affect prediction rather than necessarily adding representational content.
- Mechanistic correspondence: Operator labels are higher-level functional abstractions: one mechanism may combine roles, and one role may be implemented by several components.The present weight-space signatures are defined for attention heads, while feedforward unification and position-dependent routing are evaluated through functional or activation-level measures.
- Mechanistic correspondence: Representative circuits map non-uniquely onto CPC roles, including induction mechanisms and the four-role IOI circuit.The IOI mapping associates duplicate-token heads with alignment, S-inhibition with suppression, and name-mover heads with routing and supportive OV writes.
4. Methodology
The methodology evaluates CPC through weight-space signatures, held-out activation measures, causal ablations, contradiction-sensitive coherence analyses, and base-versus-instruction-tuned comparisons. Experiments span 15 models from five architecture families and use preregistered-style decision rules with matched controls and corrected significance tests.
- Evaluation design: The study evaluates four predictions: operator-signature validity, suppression-alignment coupling, coherence under contradiction, and post-training operator reweighting.Layerwise composition and cross-task stability are secondary analyses.
- Models and data: 15 models from five families are evaluated, with P1 using model weights without inference and P4 comparing available instruction-tuned variants.The families include GPT-2, Pythia, Qwen 2.5, Gemma 2, and LLaMA 3.2.
- Operational measures: Operator signatures use four head-level weight-space coordinates: salign, ssup, sunify, and scopy.Activation-level measures use DLA for unification and suppression, while routing is measured separately at activation level.
- P1: operator signatures: P1 correlates suppression, unification, and routing weight-space scores with held-out activation measures against random projections, magnitude rankings, and shuffled task relations.Direct causal validation ablates the top five heads per signature dimension against layer-matched random and magnitude-matched controls.
- P2: suppression-alignment coupling: Suppression-alignment coupling is tested by comparing alignment-head ablation with random-head ablation and matched no-conflict prompts.The decision rule requires a larger suppressive reduction after alignment ablation and a weaker reduction on no-conflict prompts for full support.
- P3: contradiction sensitivity: Contradiction sensitivity compares consistent and contradictory prompt pairs across depth using raw and whitened coherence proxies plus lexical and shuffled-position controls.The whitened variant reduces the contribution of global residual-stream covariance.
- P4: post-training reweighting: Base and instruction-tuned variants are compared through early- and late-half per-head signatures and corresponding-head induction-score correlations.Preservation is assessed with Fisher’s z-transformation under H0: ρ≤0.9 and H1: ρ>0.9.
- Decision rules: Decision rules use Holm-corrected tests, with P1, P2, and P3 distinguishing SUPPORT, PARTIAL, MIXED, or AGAINST outcomes.P3 separately scores the directional prediction on the whitened proxy.
5. Results and Discussion
Across 15 models, CPC operator signatures capture structured functional information, but causal specificity, depth organization, and post-training reweighting vary by architecture. Suppression is more task-stable than unification, alignment affects downstream suppression broadly rather than selectively, and whitened coherence gaps follow the predicted direction.
- Operator signatures: Suppression, unification, and routing scores correlate positively with held-out activation-level role measures, while alignment scores do not track attention-based alignment.Mean per-model Spearman correlations range from 0.21 to 0.58; the alignment weight score has r≈−0.04 with the attention-based measure.
- Operator signatures: Structured signatures occur across all architecture families, but preferred cluster counts and score geometry remain architecture-dependent.Best k ranges from 2 to 6, and clusters retain substantial structure after regressing out layer position.
- Layerwise composition: Suppression peaks early or midway depending on architecture, whereas routing tends to peak late; the three-stage depth organization is a tendency, not a strict sequence.The full timing rule holds in 8 models, while the remaining 7 satisfy two of three conditions.
- Causal ablation: Routing ablation lowers the correct-token logit in 14 models and exceeds both controls in 10, while unification ablation lowers it in 10 models and exceeds layer-matched controls in 9.GPT-2 Small is a counterexample for unification: ablation raises the logit by +1.27 versus −0.80 for its control.
- Causal ablation: Suppression ablation is task-specific only weakly: it exceeds the layer-matched control in 5 models, while activation-ranked suppression restores the predicted direction across all models.Weight scores measure total subtractive capacity, whereas task-relevant targets depend on the residual state.
- Suppression–alignment coupling: Alignment ablation reduces suppressive activity beyond random-head controls in 10 models, but comparable effects on no-conflict prompts indicate generic upstream dependence rather than conflict-specific coupling.The pre-specified outcome is 1 SUPPORT, 9 PARTIAL, and 5 AGAINST.
- Coherence under contradiction: The raw coherence gap is significant in 14 models but architecture-dependent in sign; after whitening, the directional prediction is positive in every model and significant in 14.The shuffled-pair control reproduces the raw gap in 14 models, attributing the sign flip to residual-stream covariance.
- Post-training: Induction-head score structure is preserved after post-training with r=0.98–1.00, while operator-specific late reweighting receives mixed support across families.The operator-specific rule yields 2 SUPPORT and 2 PARTIAL verdicts.
6. Related Work
CPC builds on mechanistic circuit analysis while differing from abstraction-level frameworks by naming component-level functions. It connects weight-space and activation-level analyses to a shared vocabulary for transformer computation.
- Mechanistic interpretability: Circuit analysis supplies CPC’s empirical basis through QK/OV decomposition, induction heads, IOI circuits, copy suppression, feedforward memories, and related component studies.The framework draws on prior analyses of greater-than, factual recall, and component reuse.
- Mechanistic interpretability: Sparse autoencoders, feature circuits, attribution graphs, path patching, automated circuit discovery, and edge-attribution methods connect interpretable features to causal subgraphs.These methods locate components and causal relationships used by mechanistic analyses.
- Theoretical frameworks: Unlike Bayesian and algorithmic in-context-learning frameworks, CPC names the component-level functions those approaches abstract away: alignment, unification, suppression, and routing.Post-training objectives are part of the related theoretical context.
7. Conclusion
CPC offers a shared vocabulary for transformer mechanisms, while its empirical expression remains dependent on architecture.
- CPC describes transformer computation through four operator roles involved in coherence construction and readout.
- Across 15 models from five architecture families, CPC identifies cross-architecture axes of head variation.The number of separable signature profiles, suppression depth, and contradiction responses depend on the architecture family.
- CPC predictions should be stated conditionally on architecture because signature profiles, suppression depth, and contradiction responses vary by family.
- Verifying operator assignments through causal abstraction and formalising suppression–unification task asymmetry are proposed next steps.
Limitations
The paper’s claims are bounded by the coherence proxy’s geometric confounding and by limited, task-dependent validation of suppression and coupling.
- The coherence proxy conflates semantic coherence with residual-stream geometry and functions as a global contradiction marker rather than within-prompt relational structure.The raw gap reverses sign between families and is reproduced by cross-prompt position pairs.
- A relational-coherence proxy that isolates pair-specific structure remains future work, while the ascent conjecture is weak or absent in GPT-2 and Pythia-410M.
- Suppression weight scores do not reliably isolate task-specific suppression, although task-targeted DLA rankings restore causal reliability at the cost of requiring task labels.A layer-matched control exceeded the weight score in 5/15 models.
- The evaluation uses synthetic tasks, one naturalistic negation set, four base–tuned pairs, and limited coverage for greater-than tasks and coupling controls.The coupling test uses the validated IOI circuit only for GPT-2 Small; elsewhere it uses automatically detected heads.
Ethics Statement
CPC is presented as an analytical framework for internal representations rather than system deployment, with possible implications for safety mechanisms.
- CPC concerns internal representations and does not directly address system deployment.
- Improved understanding of suppression and instruction-following circuits may inform both the design and circumvention of safety mechanisms.
Appendix A: The CPC Functional and Coherentism
CPC models transformer computation as incremental coherence refinement, with graded compatibility and four functional operator roles.
- CPC functional: CPC treats transformer computation as iterative refinement of grounded representational fragments under a global coherence functional.Grounded fragments are linked through relational structure, and local scores reward compatible interpretations while the contradiction kernel penalises incompatibility.
- CPC functional: The formulation uses graded support, token-by-token updates, and implicit optimisation rather than explicit constraint solving.
- Four operators: Alignment identifies candidate relations, unification integrates compatible information, suppression reduces incompatible alternatives, and routing carries selected information toward readout.
- Four operators: The correspondence between operators and transformer mechanisms is functional rather than one-to-one across components and roles.A mechanism may combine several roles, while one role may be implemented by several components.
- Measurement: The empirical coherence proxy measures one possible consequence of compatible representations becoming more similar, not the complete coherence functional.It does not directly identify discourse relations, entity continuity, or incompatibility between specific groundings.
- Formal elements: A fragment specifies positions, semantic or syntactic content, and compositional constraints, while a grounding maps positions partially into a relational graph.Compatibility requires agreement on shared positions and joint satisfaction of both constraint sets.
Appendix C: Task Definitions
The appendix defines three next-token tasks and their correctness criteria: indirect-object identification, greater-than year comparison, and factual recall.
- Shared evaluation format: All three tasks use token-sequence contexts and admit examples only when the model predicts correctly under the task-specific criterion.Multitoken targets are scored on their first token.
- Indirect Object Identification: Indirect Object Identification requires producing the non-repeated name after a transfer clause introduces two names and repeats one as subject.The repeated name supplies the wrong-token contrast for ablation logit margins.
- Indirect Object Identification: Prompts for indirect-object identification use 15 template frames over roughly 100 names, varying places and objects.
- Greater-Than Comparison: Greater-Than Comparison requires completing a two-digit year continuation that forms a consistent interval relative to the stated start year.Correctness requires a single-token two-digit year satisfying the inequality; this excludes nine models.
- Factual Recall: Factual Recall requires predicting the gold object of a subject-relation prompt, with correctness determined by matching the object's first token.Examples come from CounterFact.
Appendix D: Parameter and Threshold Choices
The appendix records preregistered experimental parameters, sampling and clustering choices, activation-measurement conventions, head-set sizes, statistical thresholds, and contradiction-prompt construction.
- Parameter precommitment: All constants were fixed before confirmatory runs, with stochastic procedures repeated at seeds {0, . . . , 4} and means and ranges reported across seeds.Values were preregistered, matched to validated reference circuits, or selected as conventional defaults.
- Prompt counts and splits: The design targets 500 model-correct examples per task and uses a fixed 50/50 discovery/held-out split.Candidate prompts are oversampled fourfold before correctness filtering, and hashing fixes split assignments.
- Activation measures: Activation measures use the model’s top-10 predicted tokens, while layer thirds provide the summary and testing resolution across architectures.All layerwise quantities are still computed at every layer; whitening uses fixed shrinkage α=0.1.
- Head-set sizes: P2 detection uses five alignment heads and six suppressors, whereas causal ablations remove the top five heads per operator.The detection sizes mirror validated GPT-2 Small IOI reference sets.
- Statistical thresholds: Holm correction uses α=0.05, preservation uses a preregistered threshold r=0.9, and cross-task stability uses r>0.5.The preservation rule requires shared variance above 0.81 rather than relying on a nonsignificant difference.
- Contradiction prompts: The contradiction dataset contains 3,854 globally unique consistent/contradictory pairs generated once with seed 42.Pairs share surface structure up to a key substitution introducing or removing contradiction, across two pools and seven contradiction types.