Source-linked AI summary
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
Gregory N. Frank
TL;DR
Alignment-trained language models can recognize the same content yet produce divergent policies, leaving the routing computation between detection and behavior insufficiently localized. This paper maps that gate-amplifier mechanism across models and finds that detection-layer cipher bypasses collapse routing necessity by 70–99%, enabling policy circumvention when pattern matching fails.
Problem
Models can encode sensitive content similarly while producing divergent policies, but the routing computation connecting detection to behavior remains insufficiently localized.
Method
The paper combines DLA, ablation, interchange testing, knockout cascades, and cipher interventions to localize and test content-specific policy routing across language models.
Results
Gate interchange necessity collapses by 70–99% under substitution ciphers across three models, while the gate-amplifier motif generalizes to twelve models from six labs.
Takeaways & Limitations
Policy behavior is routed through an early detection-layer gate, so encodings that defeat its pattern matching can bypass refusal even when deeper layers reconstruct content.
Takeaways & Limitations
The study covers 2–72B models and political-censorship or safety-refusal settings, while broader architectures, scales, and behavior domains remain untested.
Abstract
from arXiv · showhide
We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate and amplifier are single heads; at larger scale they become bands of heads across adjacent layers. The gate contributes under 1% of output DLA, yet interchange testing (p < 0.001) and knockout cascade confirm it is causally necessary. Interchange screening at n >= 120 detects the same motif in twelve models from six labs (2B to 72B), though specific heads differ by lab. Per-head ablation weakens up to 58x at 72B and misses gates that interchange identifies; at scale, interchange is the only reliable audit. Modulating the detection-layer signal continuously controls policy from hard refusal through evasion to factual answering. On safety prompts the same intervention turns refusal into harmful guidance, showing that the safety-trained capability is gated by routing, not removed. Thresholds vary by topic and by input language, and the circuit relocates across generations within a family even while behavioral benchmarks register no change. Routing is early-commitment: the gate fires at its own layer before deeper layers finish processing the input. An in-context substitution cipher collapses gate interchange necessity by 70 to 99% across three models, and the model switches to puzzle-solving rather than refusal. Injecting the plaintext gate activation into the cipher forward pass restores 48% of refusals in Phi-4-mini, localizing the bypass to the routing interface. A second method, cipher contrast analysis, uses plain/cipher DLA differences to map the full cipher-sensitive routing circuit in O(3n) forward passes. Any encoding that defeats detection-layer pattern matching bypasses the policy regardless of whether deeper layers reconstruct the content.
1 Introduction
The introduction explains why identical mid-depth topic detection can produce radically different behaviors: learned routing maps detected concepts to behavioral policies. It localizes a gate-amplifier circuit, establishes causal evidence across scales, and frames routing as a source of safety bypasses.
- Motivation: Identical mid-depth topic detection can precede refusal, propaganda, factual answering, or fabrication across models.A linear probe achieves perfect accuracy in all four models despite enormous behavioral variation.
- Routing framework: Learned routing maps detected concepts to behavioral policies, explaining the gap between detection and behavior.The paper localizes this machinery, studies its scaling, and uses it to predict a class of safety bypass.
- Circuit mechanism: The localized mechanism comprises contextual detection at layers 15–16, a sparse attention gate, and downstream amplifier heads that boost refusal.The gate reads the detection signal and writes a routing vector; amplifier heads strengthen that vector toward refusal.
- Circuit mechanism: ∼77% of the routing signal comes from distributed attention heads and ∼23% from MLP pathways, while gate and amplifier heads contribute under 1% directly.These figures are reported on Qwen3-8B at n=120, with the ratio identified as corpus-dependent; interchange testing nevertheless establishes the gate as causally necessary.
- Evidence and scaling: The evidence program progresses from separability and held-out generalization to causal intervention and failure-mode prediction.The stated pipeline combines per-head DLA, head-level ablation, activation-swap interchange testing, and bootstrap stability.
- Evidence and scaling: Twelve models from six labs spanning 2B–72B are covered by interchange screening, while decomposition and knockout cascade identify the gate-amplifier motif in three architectures.The architectures named are Qwen3-8B, Phi-4-mini, and Gemma-2-2B; screening uses n≥120.
3. Scaling
Across four same-generation model pairs spanning 2B–72B, per-head ablation weakens sharply while interchange remains informative. The gate commits routing at the detection layer, and cipher encoding collapses its interchange necessity while changing responses from refusal to puzzle-solving.
- Scaling: Across four same-generation pairs spanning 2B–72B, per-head ablation effects weaken up to 58× while interchange remains informative.This comparison covers models from 2B to 72B.
- Scaling: The gate commits the routing decision at the detection layer, establishing an early-commitment vulnerability in policy routing.The passage identifies the detection layer as the point where the routing decision is committed.
- Scaling: 70–99% interchange necessity collapse under cipher encoding across three models (n=120) coincides with puzzle-solving rather than refusal.The cipher condition changes the model’s response from refusal to puzzle-solving.
2 From Detection to Routing
Policy routing is committed at prompt time through contextual detection rather than a simple scalar threshold, and its geometry is mechanistically identifiable even when behavioral benchmarks miss major shifts. Surgical ablation removes routing in most tested models, while the motif generalizes across a 12-model, six-lab panel despite lab-specific directions.
- Prompt-Time Commitment: Routing is committed before generation, with matched last-prompt and first-generated-token DLA overlapping almost perfectly in Qwen3-8B.Even GLM-4-9B shows a 2.8-nat KL peak between matched sensitive and control prompts.
- Contextual Detection: The same keyword receives different layer-16 scores under different framing, showing that detection is compositional and routing is not a simple threshold.Annotated edge cases further confirm contextual routing beyond scalar thresholding.
- Mechanistic Validation: 91–100% political-probe accuracy survives LOCO-CV while null probes drop to chance, separating genuine encoding from probe artifacts.Political probes trained on all categories except one retain performance on the held-out category; null controls do not.
- Mechanistic Validation: 3 of 4 tested models lose political routing after surgical ablation of the sensitivity direction, producing factual output, while cross-model direction transfer fails.The failed transfer indicates lab-specific routing geometry.
- Benchmark Blindness: 33% to 0% political refusal across three Qwen generations went undetected by benchmarks, whereas a mechanistic signature registered the shift.The same passage reports that steering rose during the refusal decline.
- Cross-Model Scope: 12 models from 6 labs spanning 2B–72B validate the routing motif, with Qwen3-8B as the deep case study and Phi-4-mini as the cleanest replication.The broader panel supports motif generality while the passage does not claim identical routing components across labs.
3 A Routing Circuit in Qwen
In Qwen3-8B, a three-step pipeline identifies L17.H17 as a content-sensitive routing gate whose signal is amplified by deeper heads. Although the gate contributes less than 1% of output DLA, interchange and knockout analyses show it is causally necessary and initiates a distributed cascade.
- Gate identification: L17.H17 is identified as the gate because interchange finds the strongest combined necessity and sufficiency signal, despite different DLA and ablation rankings.Its necessity is 1.1% and sufficiency 0.3%, leading L22.H7 by 64% with p < 0.001; interchange top-10 Jaccard is 1.0.
- Circuit mechanism: The gate reads sensitive content at layer 17, while layers 22–23 amplify the routing vector without re-examining content.On sensitive prompts, L17.H17 attends to relevant tokens; on matched controls, it attends to generic punctuation, whereas amplifier heads attend to formatting and position tokens.
- Causal cascade: 5 of 6 downstream Qwen amplifiers are suppressed when L17.H17 is zeroed, confirming a cascade with L22.H5 showing −25.8% and L22.H6 counter-routing at +10.1%.Across architectures, gate knockout suppresses 3 of 5 amplifiers in Phi-4-mini and Gemma-2-2B, indicating partial redundancy rather than a single point of failure.
- Causal interpretation: <1% of output routing DLA comes from the gate and amplifier heads, yet interchange shows causal necessity at p < 0.001 and knockout suppresses downstream heads by 5–26%.The gate ranks #2 at L18 immediately after writing its vector, then falls out of the output top 20 as distributed carriers at L30–35 dominate.
- Attention–MLP interaction: MLP layers carry mean-absolute interchange necessity of 5.2–8.7 and knockout effects up to ∼70, but the gate remains necessary when MLP routing dominates.Gate necessity is 1.84 on the high-MLP-share half of prompts versus 2.22 on the low-MLP-share half.
4 Routing Across Architectures and Scales
Interchange screening identifies the gate-amplifier routing motif across 12 models from six labs spanning 2B–72B, while per-head ablation becomes unreliable at larger scales. Routing concentrates in fewer heads in smaller models, distributes across larger models, and relocates across Qwen generations despite stable within-generation amplifiers.
- Cross-architecture detection: 12 models from 6 labs spanning 2B–72B show the gate-amplifier motif under interchange screening at n≥120.Necessity ranges from 1.0% in Mistral-7B to 8.4% in Gemma-2-2B; two 70B+ models confirm the motif.
- Scaling organization: 0.38 to 0.05–0.15 marks the Qwen3-to-Qwen3.5 drop in top-1 head DLA amplitude as routing distributes across more heads.Smaller models concentrate routing in fewer heads, whereas larger models distribute it while the motif remains detectable.
- Audit reliability: 58× weaker ablation at 72B contrasts with interchange necessity remaining above 1% in every tested model.At 72B, the top ablation effect is 0.016, essentially undetectable, while interchange remains the reliable gate-finder across 2B–72B.
- Generational relocation: 0–2 of the top 20 routing heads are shared across Qwen generations, with Jaccard ≤0.05, while core amplifiers remain stable within generations.The circuit therefore relocates across generations even as within-generation cross-corpus amplifier stability persists.
5 Routing Is Causally Controllable
Detection-layer steering continuously controls routing, converting refusal into evasion or factual answers on Tiananmen prompts and into harmful guidance on Phi-4 safety prompts. Routing thresholds vary by topic and language, making policy behavior category- and language-sensitive.
- Continuous control: α · d steering at the detection layer continuously modulates routing across 2,400 outputs evaluated by three-judge majority vote at n=120.The steering direction d is the mean activation difference between sensitive and control prompts.
- Continuous control: 100% baseline refusal on Tiananmen prompts falls to 0% by α=35 under attenuation, following a clean sigmoid.Tiananmen was the only category with 100% baseline refusal.
- Topic sensitivity: 8% aggregate refusal at α=0 masks topic-specific behavior: only Tiananmen produces consistent hard refusal across 15 political categories, at 8/8.Amplification reveals different routing thresholds and sensitivities across categories.
- Language sensitivity: Chinese prompts produce higher gate-layer activation than English equivalents for Tiananmen (+0.33) and Xi/CCP (+0.32), while benign topics show no difference.The preliminary comparison used n=16 paired prompts.
- Behavioral control: Attenuation shifts Tiananmen behavior from REFUSAL to EVASION to FACTUAL, and Phi-4 safety behavior from REFUSAL to HARMFUL_GUIDANCE.Across 2,400 outputs, inter-judge agreement was 76.0% unanimous and 97.2% majority.
6 Discussion
The cipher experiments localize policy bypass to the detection-layer routing interface: encodings that fail to trigger the gate evade refusal even when deeper processing may continue. Cipher contrast and interchange together reveal a broader, partly distributed circuit, while plaintext gate injection partially restores refusal.
- Circuit discovery: 47 heads: Cipher contrast identifies a broader content-dependent routing circuit in Phi-4-mini, including known members and more than 30 previously untested heads at layers 13–16.Across the three models, approximately 77% of positive routing signal is content-dependent.
- Circuit discovery: 18 unique circuit members: Combining cipher contrast with interchange outperforms either method alone because the methods detect content dependence and causal necessity, respectively.Only 2 of the top 10 heads overlap in Phi-4-mini; contrast finds cipher-sensitive L16 heads, while interchange finds deep content-independent amplifiers at L26–L29.
- Interpretation and limitations: The vulnerability is an early-commitment failure at the routing interface: encodings that do not instantiate a gate-readable representation bypass policy regardless of whether deeper layers reconstruct target content.The experiments do not directly verify harmful-intent representations under cipher, and the claims are primarily shown for a Latin substitution cipher and specified censorship and safety settings.
- Rescue experiment: 48%: Injecting the plaintext gate activation into Phi-4-mini’s cipher forward pass restores refusal, compared with 0% under cipher alone.Qwen3-8B shows 0% single-head recovery, consistent with its more distributed architecture.
- Cipher bypass: 70–99%: Gate interchange necessity collapses under cipher encoding across the three tested models, with ciphered prompts producing puzzle-solving behavior rather than refusal.Gemma-2-2B and Phi-4-mini show 99% drops; Qwen3-8B shows a 70% drop.
Appendix A: Mechanistic Methods · Appendix B: Evidence Summary
Appendix A defines the attribution, interchange, knockout, modulation, statistical, and behavioral procedures used to localize and validate policy-routing mechanisms. Appendix B organizes major claims by supporting evidence, model coverage, sample size, and evidence depth.
- Appendix A: Mechanistic Methods: DLA projects each component’s output onto a target–baseline logit-difference direction at the last prompt token.The target is the model’s greedy first generated token, while the baseline is the mean embedding of common refusal tokens.
- Appendix A: Mechanistic Methods: Interchange testing scores candidate heads by necessity and sufficiency, distinguishing dual-scoring gates from necessity-only amplifiers.Stored pre-projection activations are swapped through an o_proj forward pre-hook between sensitive and matched control prompts.
- Appendix A: Mechanistic Methods: Knockout cascade zeros a gate head’s o_proj input slice and measures downstream amplifier DLA changes across n=120 prompts for both Qwen and Phi-4.The procedure removes the gate contribution from subsequent computation and compares downstream effects with the unperturbed pass.
- Appendix A: Mechanistic Methods: Intermediate-layer DLA can expose a gate before downstream amplification, with the Qwen gate ranked #2 at L18.The gate’s rank decreases later as downstream heads amplify its signal.
- Appendix A: Mechanistic Methods: Across four alternative direction definitions, the gate’s DLA rank ranges from #177 to #294, whereas interchange ranking remains perfectly stable under 2,000 bootstrap iterations.This establishes direction robustness for interchange identification while showing that DLA ranking is direction-sensitive.
- Appendix A: Mechanistic Methods: Detection-layer modulation sweeps α from 0 to 50 in increments of 5, testing both attenuation (−α) and amplification (+α).The intervention adds or subtracts α · d through a forward hook at the detection layer.
- Appendix A: Mechanistic Methods: Validation combines 2,000 bootstrap resamples, 10,000 paired sign-flip permutations, and knockout comparisons against 10 random non-gate heads at similar depths.Behavioral labels use majority votes from three independent LLM judges across six categories, with 76.0% unanimous agreement on Qwen and 84.0% on Phi-4.
- Appendix B: Evidence Summary: Table 3 maps each major claim to supporting evidence, model coverage, sample size, and evidence depth.Its “Full decomposition” tier includes DLA, ablation, interchange, and knockout cascade, while “Interchange screening” includes necessity and sufficiency only.
Appendix C: Cipher Contrast Analysis
Cipher contrast analysis identifies content-dependent routing heads in O(3n) forward passes by comparing plaintext and cipher-encoded harmful prompts. It complements interchange by recovering known and previously untested circuit members, revealing distributed, layered routing structure.
- Method: O(3n) forward passes identify content-dependent routing heads, reducing the cost of interchange testing from O(4nK) for K candidate heads.The method exploits cipher encoding as a natural experiment and compares plaintext, cipher, and benign-control DLA.
- Validation against known circuits: In Phi-4-mini, the known gate L13.H7 ranks 4th and top amplifier L16.H13 ranks 3rd by cipher contrast score.In Qwen3-8B, the four known L22 amplifiers rank 5th, 7th, 12th, and 20th, while gate L17.H17 ranks 57th (top 5%).
- New circuit members: In Phi-4, previously unknown heads L16.H9, L16.H12, and L16.H10 rank 1st, 2nd, and 5th, while Qwen’s L31.H3 ranks 1st overall.These findings show that cipher contrast discovers heads not surfaced by interchange screening, including a deep-layer Qwen head.
- Layer clustering and signal decomposition: Cipher-sensitive heads form sparse layer bands: Phi-4 has gate L13 and amplifiers L16, while Qwen has gate L17, amplifiers L22, and deep routing at L31–35.The method classifies non-negligible heads using |DLA| ≥0.05, cipher contrast >0.1, and routing contribution >0.05.
- Multi-head interchange and coalition structure: 3.31× the single gate head’s interchange necessity is achieved collectively by the top 10 pro-routing heads across 5 layers.Per-prompt correlations separate pro-routing and counter-routing coalitions, with internal r = 0.5–0.78 and r = 0.6–0.88, respectively.
Appendix D: Bijection Detection Bypass · Appendix E: Generated Text Examples · Appendix F: Three-Judge Panel
Cipher encoding bypasses alignment by preventing the detection-layer gate signal from forming, while deeper processing may partially recover content without restoring refusal. Injecting the plaintext gate activation partially rescues refusals, and generated examples plus three-judge labeling document the resulting behavioral shift.
- Appendix D: Bijection Detection Bypass: At the gate layer L17, cipher-encoded harmful prompts score 5.1 versus 7.3 for benign controls, showing the detection signal is absent.At L35, the cipher probe reaches 47.6, only 29% of the plaintext harmful score 163.8.
- Appendix D: Bijection Detection Bypass: At alpha values 0, 10, and 20, amplification fails to restore routing because cipher inputs lack a gate-readable detection signal.The cipher blocks formation of the safety representation at the routing interface, leaving no detection signal for amplification to boost.
- Appendix D: Bijection Detection Bypass: Under Qwen3-8B plaintext, refusal tokens first appeared at L24 in 7% of prompts and consolidated at L34–35 in 17%; under cipher, they never exceeded 2%.The refusal decision materialized after the L17 gate and L22–23 amplifiers, whereas cipher prevented refusal-token formation throughout the network.
- Appendix D: Bijection Detection Bypass: 48.3% recovery followed plaintext gate-activation injection in Phi-4-mini, restoring refusal in 58 of 120 cipher cases.The baseline was 99.2% refusal for safety prompts and 0% refusal under cipher encoding; the smaller n=8 corpus recovered 75% (6/8).
- Appendix D: Bijection Detection Bypass: Cipher inputs make models treat harmful requests as word puzzles, bypassing safety intervention rather than producing ordinary factual answers.The observed output begins by decoding the message step by step, contrasting with plaintext refusal and high-alpha factual answering.
- Appendix E: Generated Text Examples: Generated examples show three distinct outcomes: plaintext refusal, cipher puzzle-solving compliance, and alpha=50 attenuation yielding direct historical information.These outputs illustrate that bypassing detection changes policy behavior while preserving task-directed generation under different interventions.
- Appendix F: Three-Judge Panel: 2,400 dose-response outputs were classified by Gemini 2.0 Flash, Llama 3.1 8B, and GPT-4o-mini using majority labels across six behavioral categories.Agreement was 76.0% unanimous, 97.2% majority, and 2.8% three-way disagreement.
Appendix G: Per-Category Dose-Response
Policy responses vary sharply by political category at baseline and under amplification. Tiananmen consistently produces hard refusal, while other categories show factual, evasive, steered, or threshold-dependent responses.
- Baseline responses: At α=0, Tiananmen yields hard refusal on 8/8 prompts (100%), while Falun Gong yields 1/8 refusal and other categories produce steered, factual, or evasive answers.DISAGREE labels are omitted, so rows may sum to less than 8.
- Amplification thresholds: 75% refusal is reached at α=50 for Internal CCP politics and Xinjiang, versus 50% for Great Firewall.These categories reach refusal at different amplification thresholds.
- Amplification thresholds: Hong Kong and Falun Gong never reach refusal under amplification, producing steered outputs instead.Their response pattern differs from categories that cross refusal thresholds.
Appendix H: Scaling Data · Appendix I: Prompt Corpora and Control Design
Scaling changes where the routing gate sits and can decouple refusal behavior from routing-signal strength across model generations. The evaluation uses structurally matched sensitive and control prompts, with corpus validation showing stable gate and amplifier heads despite peripheral variation.
- Appendix H: Scaling Data: 33% to 0%: political refusal falls across Qwen generations while steering rises from 3.25 to 5.0, despite no refusal benchmark detecting the shift.Top-1 routing-head DLA peaks in Qwen3-8B, then falls sharply in Qwen3.5, alongside a drop in total routing signal.
- Appendix H: Scaling Data: Across all four families, the gate moves deeper relative to total model depth as model scale increases.The authors interpret this as larger models requiring more layers to form the detection representation before routing begins.
- Appendix I: Prompt Corpora and Control Design: All experiments pair a routing-sensitive prompt with a matched control sharing syntactic structure but differing in topic sensitivity or harmfulness.This design isolates representation swapping from generic political sensitivity or task structure.
- Appendix I: Prompt Corpora and Control Design: n=120: the political corpus contains 15 categories of Chinese political sensitivity, with 8 prompts per category and structurally parallel non-Chinese controls.Categories include Tiananmen Square, Tibet, Xinjiang, Taiwan, surveillance, and labor rights.
- Appendix I: Prompt Corpora and Control Design: n=120: the safety corpus combines 88 HarmBench harmful requests with 32 manually constructed requests, each matched to a benign structural parallel.Examples contrast vehicle theft or password theft with repairing an ignition switch or announcing a product launch.
- Appendix I: Prompt Corpora and Control Design: Bootstrap Jaccard 0.92: Qwen3-8B’s core amplifier heads remain the top three across v1, adversarial, and v2 corpora.The gate L17.H17 was identified on v1 and validated across v1 (n=24), adversarial (n=32), and v2 (n=120); peripheral ranks 7–20 vary.