Source-linked AI summary
From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge
Wenkang Wei, Yuan Fang, Renhe Jiang, Hong Cheng, Xingtong Yu
TL;DR
The paper asks how language models progressively retrieve and use internal knowledge, separating query-routing information from answer-supporting content through layerwise hidden-state interventions. Across Qwen, Llama, and Gemma, it finds that routing and content have distinct, representation-specific trajectories, with global routing dependence decreasing later while fitted content remains causally relevant.
Problem
How language models progressively retrieve and use internal knowledge across layers, including how query-routing information differs from answer-supporting knowledge, remains unclear.
Method
The study uses layerwise interventions on question-end hidden states across Qwen, Llama, and Gemma, comparing fitted request directions, content representations, and selection candidates across answer formats.
Results
Routing and content emerge in an ordered but model- and representation-specific trajectory: global routing dependence decreases later while fitted content persists, whereas Qwen's pair-conditioned direction retains a late effect.
Takeaways & Limitations
The results distinguish early readability, natural routing strength, causal steering, and later content dependence rather than treating them as a single measure of knowledge use.
Takeaways & Limitations
The handoff result concerns deletion of a fitted global direction, with unequal intervention lengths, and does not establish that fitted content is its unique mediator.
Abstract
from arXiv · showhide
How does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on the hidden state at the end of the question. Across Qwen, Llama, and Gemma, we compare country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct. A pair-conditioned request direction describes which country is queried in natural single-country questions; a global request direction describes first- versus second-country requests in paired questions; separate selection candidates test control among contents already available in the hidden state. A diagnostic reanalysis of frozen Qwen natural-question states shows that the pair-conditioned direction grows stronger before interventions on it begin to alter later fitted knowledge, with this causal window opening while answer-supporting content is still forming. The paired three-model trajectories are not uniform: Gemma shows a partially overlapping mid-layer routing-content profile, whereas Llama has no sustained routing-effect window under the same gates. In the paired protocol, dependence on the global request direction decreases from fixed earlier to later layer sets while dependence on fitted content persists. A matched Qwen comparison shows that the pair-conditioned direction retains a late effect, so this operational handoff concerns the global fitted direction rather than all request information. These results separate early readability, natural strength, causal steering, and later content dependence.
1 Introduction
The paper separates hidden-state content from routing information and traces how their roles change as language models form answers across layers. It finds ordered routing and content emergence, followed by a causal handoff from routing toward answer-supporting content.
- The analysis causally separates answer-supporting content from routing, distinguishing parameter routing from hidden-state routing.Parameter routing points toward knowledge in model parameters, whereas hidden-state routing directs computation toward content already represented in the hidden state.
- Early layers integrate query-relevant knowledge into answer-supporting content before either routing form has a detectable answer effect.
- Parameter routing strengthens while content is still forming, directs computation toward relevant parameter knowledge, and is followed by progressive content consolidation.
- Hidden-state routing becomes effective later, directing computation toward newly formed content for further processing toward the answer.
- After answer-supporting content forms, causal dependence shifts from routing toward the content itself: routing interventions weaken later, while content interventions remain consequential.
2 Related Work
Prior research has studied where pretrained knowledge is stored, how it is retrieved, and which hidden states influence predictions. These lines of work leave unresolved how query information and answer-supporting knowledge interact across layers.
- Parameter studies locate and edit pretrained knowledge associations but do not explain how an ordinary forward pass accesses that knowledge.
- Retrieval studies decompose knowledge use into functional components and stages, but do not establish the full cross-layer relation between query information and formed answers.
- Causal hidden-state interventions identify information that influences knowledge-based predictions, yet an answer effect alone does not reveal whether that information carries content or guides subsequent retrieval.
3 Parameter-Retrieval Routing in Natural Questions
The pair-conditioned request direction is readable across layers before interventions measurably affect answers, while later interventions can steer computation toward the paired country’s knowledge. Its natural strength rises before later-knowledge effects emerge, during ongoing fact formation, but deletion does not establish necessity for later knowledge.
- 3.5 Early Readability Precedes Detectable Answer Dependence: The country-request direction identifies the queried country on every held-out question across all 36 layers, but detectable answer effects occur only at layers 28–36 and peak at layer 32.Readability is therefore present early, whereas causal answer dependence is concentrated late.
- 3.5 Early Readability Precedes Detectable Answer Dependence: At the fixed layer-32 to layer-36 pair, reversal changes both the answer margin and later fitted knowledge, whereas deletion changes the answer margin without confirmed later-knowledge loss.The registered criterion supports reversible steering but not deletion-based necessity for later knowledge formation.
- 3.6 Natural Route Strength Precedes Its Later-Knowledge Effect: The natural route coefficient has a sustained state-relative increase from layer 21, before route interventions affect later fitted knowledge and while the fitted fact is still developing.The state-relative mean rises from 0.00905 early to 0.01412 at layer 21, 0.04963 at layer 27, and 0.10128 at layer 28.
- 3.6 Natural Route Strength Precedes Its Later-Knowledge Effect: Fitted-content support grows from 69.8% at layer 21 to 96.9% at layer 28 and reaches 100% at layer 33, while route deletion or reversal affects the fixed layer-36 fact from layers 28–31 or 28–32.These are pointwise diagnostic intervals from an existing-data reanalysis, not simultaneous confidence bands or universal onset claims.
- 3.7 Reversal Steers Later Knowledge without Fixed-Pair Deletion Necessity: The routing candidate preserves the country distinction across wordings and can steer later computation toward the paired country’s knowledge, but its naturally occurring coefficient is not shown to be necessary for that fact.This separates early request representation, natural coefficient strength, causal steering, and later content dependence.
4 Object Selection within Hidden States
The paper tests whether hidden-state selection can control which of two country-related contents determines an answer, distinguishing supplied-answer choice from retrieval of stored facts. Evidence is hierarchical: stable transfer appears only for Qwen adjective answers, while a stronger two-continuation test succeeds only in a separate Qwen base-model capital task.
- 4.1 Calibrating Selection when Both Answers Are Supplied: The calibration task isolates choice between supplied markers by changing the requested record without requiring either marker to be recalled from model weights.The fitted contrast averages first-record states minus second-record states across paired prompts.
- 4.2 Transferred Selection Is Stable in One Model–Format Setting: Transferred selection passes every calibration and split-level gate only for Qwen adjective answers, at layers 28–30 and 32, with a largest effect of 0.11 answer-margin units.Other model–format curves remain measurements but do not qualify as stable transfer.
- 4.3 Direct Selection in a Separate Base-Model Task: The strongest two-continuation test succeeds in the Qwen base-model capital task, with directed selection changes exceeding controls on 37/40 and 44/44 validation questions.The intervention changes downstream selection and answers while preserving non-selection readouts; joint pass rates meet the threshold at receiving layers 27–29.
- 4.3 Direct Selection in a Separate Base-Model Task: Early edits can alter the answer path without demonstrating two usable candidates, while the measured failures constrain this linear interface rather than proving selection is absent.The paper therefore treats object selection as conditional on model, task, and representation rather than as a universal processing stage.
- 4.5 Capital Questions across the Three Instruction Models: Gemma yields a reproducible internal selection candidate, but its answer effects fail the prespecified three-layer threshold, so testing stops before two-candidate restoration.Selection changes exceed controls on all 52 questions, whereas answer changes exceed the strongest pointwise control on 32/52 and 31/52 questions.
- 4.6 Selection Evidence Is Hierarchical and Task-Dependent: The evidence forms a hierarchy: marker choice, cross-task transfer, and same-state two-continuation control establish progressively stronger claims.No instruction model reaches two-candidate restoration in the reported continent and capital experiments.
5 The Development and Use of Knowledge Content
Fitted continent content remains causally important through late layers across Qwen, Llama, and Gemma, and a noun-derived intervention transfers to adjective answers but rarely to arbitrary codes.
- 5.2 Fitted Content Becomes Consequential and Persists Late: The fitted content component is tested by deleting it and comparing answer-margin damage against matched random changes, separating readability from causal dependence.Pointing measures whether the state is closer to the correct continent reference; deletion tests whether later computation depends on that component.
- 5.2 Fitted Content Becomes Consequential and Persists Late: Content deletion damages continent answers through late layers in all three models, with positive effects at layers 27–36 for Qwen, 14–28 for Llama, and 18–34 for Gemma.The layer ranges are model-specific rather than aligned by a common depth.
- 5.2 Fitted Content Becomes Consequential and Persists Late: A separate natural single-country Qwen experiment also finds positive late content-deletion effects, but its float32 results are not pooled with paired-task bfloat16 values.Every pointwise 95% interval for layers 27–36 lies above zero.
- 5.3 Fitted Content Transfers to Adjectives but Rarely to Codes: A noun-derived content change transfers to adjective answers in all three models without refitting adjective references, supporting shared factual content across natural expressions.The corrected mean answer-margin shifts are 15.72 for Qwen, 19.78 for Llama, and 46.09 for Gemma; these are not accuracies or proportions.
- 5.3 Fitted Content Transfers to Adjectives but Rarely to Codes: Transfer to arbitrary code answers is limited: only 1/5 baseline-qualified mappings pass in Qwen, while 0/4 pass in Llama and 0/2 in Gemma.A map must first produce correct codes on all baseline questions in both validation splits.
6 Layerwise Changes in Routing and Content Dependence
Across models, routing and content effects occupy overlapping but nonuniform layer profiles. In the paired protocol, late answers become less sensitive to the fitted global request direction while remaining sensitive to fitted content, whereas a matched Qwen comparison preserves a late pair-conditioned routing effect.
- 6.1 Overlapping Routing and Content Profiles: The three models show overlapping rather than shared routing-content profiles: Gemma has a partially overlapping mid-layer pattern, while Llama lacks a sustained route-effect window under the same checks.Qwen shows a mid-layer coordinate rise before sustained content dependence.
- 6.2 Global-Route Dependence Decreases on New Countries: In later fixed layer sets, deleting the fitted global request direction causes less answer damage while deleting fitted content remains consequential across the paired-country evaluation.The result concerns a global fitted direction, not all request representations or a shared processing boundary.
- 6.2 Global-Route Dependence Decreases on New Countries: Late states remain separable along the global request direction, so reduced deletion damage does not show that request information has vanished.The paper uses “routing–content handoff” only for the operational contrast between global-route and content deletion effects.
- 6.3 Controls for Request Specificity and Repeated Deletion: Qwen controls show that the true request direction exceeds balanced wrong-label directions, while repeated late content deletion exceeds repeated late routing deletion.The controls use the same 24 country pairs and matched random paths.
- 6.3 Controls for Request Specificity and Repeated Deletion: Repeatedly deleting the fitted global component from layers 32–36 adds only 0.21% normalized margin damage, below the prespecified 10% practical threshold.The one-sided 95% upper bound is 0.45%, indicating no substantial recovered late dependence in this test.
- 6.4 Pair-Conditioned Routes Retain Late Dependence: Changing from a global to a pair-conditioned direction retains positive late effects, so the handoff is specific to the global fitted direction rather than all request information.The matched comparison rules out task format alone as the explanation for the different late effects.
7 Limitations
The paper limits its conclusions to specific country–continent tasks, fitted measurements, intervention designs, and model/checkpoint settings. These constraints primarily bound generalization and interpretation rather than the reported within-task findings.
- The evidence is centered on country–continent associations, with paired experiments across three instruction-tuned models but complete natural single-country intervention evidence only for Qwen.Noun, adjective, and code formats vary expression of one association rather than representing separate knowledge domains.
- The natural-route onset is diagnostic rather than a prospectively registered change point because it reuses frozen Qwen validation states.The three-model paired analysis uses a different global direction and normalization, so it is not a direct Llama/Gemma replication of the Qwen timing result.
- Because prompts omit the target association, the experiments study how released weights use the associations without assigning each fact exclusively to pretraining.Later instruction tuning or another training stage may also contribute to the evaluated behavior.
- The fitted linear directions and content spaces capture only selected aspects of representation, and projecting them out does not remove every possible encoding.The pair-conditioned natural-question direction is also constructed separately for each country pair, limiting transfer claims.
- The reported handoff concerns deletion of the fitted global direction, not all request representations, because pair-conditioned and global interventions also differ in deletion length.The study did not compare the two orientations at one common targeted length, and fitted content is not shown to be the unique mediator.
- Object-selection evidence is task- and checkpoint-dependent: instruction-model tests do not reach the full restoration criterion, while noun-to-adjective transfer generally fails for arbitrary codes.These failures constrain the tested directions and finite readouts rather than showing that models lack selection.
8 Conclusion
The conclusion distinguishes request information from answer-supporting content and tracks how their contributions vary across depth. Results support a representation- and model-specific routing–content handoff rather than one universal processing order.
- Request information and answer-supporting content make distinct, changing contributions across depth, requiring separate measures of readability, strength, causal effect, and later knowledge.Natural single-country Qwen evidence shows readability from the first layer, a later sustained rise in request strength, and causal effects while content is still forming.
- Across three instruction models, dependence on the global request candidate decreases from earlier to later layer sets while dependence on fitted content persists.New-country tests, request-specificity controls, and repeated deletion support this operational handoff.
- A matched Qwen comparison retains a late effect for pair-conditioned request deletion, showing that the handoff concerns the fitted global direction rather than all request information.The inferred depth profile therefore depends on the fitted representation, center, intervention size, model, and task.
- Noun-derived content changes transfer to adjective answers without adjective refitting but rarely transfer to arbitrary codes.Direct selection evidence is also setting-dependent: one Qwen base-model capital task succeeds, while instruction-model tests stop earlier.
- The paper reports reproducibility support through archived plotting data, hashes, frozen records, and fresh-load comparisons.These materials accompany the source and preserve model, tokenizer, prompt, intervention, output, and comparison records.
A.1 Task and Data
The paired task varies which country is requested while withholding the country–continent association, then tests the same association in continent-name, adjective, and code formats. Layerwise hidden-state measurements separate target content from parameter-routing and object-selection candidates before controlled interventions assess answer and knowledge effects.
- Task and data: Each paired question names two countries from different continents and asks about one, so changing the request changes the target knowledge without supplying either association.The two requests are evaluated as separate model inputs.
- Task and data: The same country–continent association is expressed as a continent name, adjective, or arbitrary code, testing format dependence rather than three knowledge domains.The code format supplies a mapping such as Africa = dax and Asia = wug, but not the queried association.
- Task and data: The original data use four country-disjoint partitions of 24 countries, with fit, selection, and two validation roles separating construction from recurrence testing.Each partition contains six countries arranged into three pairs, and continent position is balanced across requests.
- Measured components: At the question-end position, each layer’s hidden state is decomposed into a fitted target-knowledge space and two request contrasts.The target space is a two-dimensional continent basis; parameter routing uses first-minus-second main-task requests, while object selection uses an auxiliary marker task.
- Measured components: The object-selection candidate measures control over information already supplied by the hidden state, unlike parameter routing, which targets stored associations outside the fitted knowledge space.Its calibration requires directional marker switching to outperform matched random changes before transfer to the continent task.
- Interventions and scoring: Interventions add or delete one fitted component at the question-end position while leaving model parameters and other token positions unchanged.Random Gaussian controls are scaled to match deletion length, and effects are evaluated using controlled two-candidate answer margins and later knowledge scores.
B.1 Model and Task Records
The model-record protocol fixes extraction, fitting, intervention, scoring, and validation conventions for paired-country experiments. It distinguishes fitted measurement directions and reports stable effects using cross-partition criteria without treating plotted ranges as significance intervals.
- Model and task records: The frozen records use Qwen, Llama, and Gemma instruction checkpoints, with interventions applied at the question-end position while model parameters remain fixed.Exact model, tokenizer, prompt, weight, and intervention records are archived for fresh-load checks.
- Model and task records: Hidden states are extracted at the final input-token position across layers, with answer candidates scored token by token using their own preceding tokens.Layerwise knowledge scores are read at the same position after subsequent layers.
- Fitted measurements: Each model, layer, and answer format receives its own fitted knowledge space rather than a shared coordinate system across formats.The space is built from centered continent means and its first two singular vectors.
- Results reporting: Table 4 reports stable 1-based layers and peak answer-margin effects for parameter routing, transferred object selection, and target content within each model–format row.Raw magnitudes are not pooled or ranked across models, and peak values can include non-stable layers.
- Fitted measurements: Global parameter routing and object-selection directions are residualized against the knowledge space and the other candidate before normalization.The resulting directions are separate measurement vectors, not an additive orthogonal decomposition of the hidden state.
- Results reporting: Stability requires positive effects in the selection split and both validation partitions, while plotted validation ranges show their minimum and maximum rather than statistical significance.A non-stable layer is not evidence of exactly zero effect.
C Expanded Evaluation and Repeated Interventions
The expanded evaluation fixes country-pair layer sets and uses pair-level bootstrap controls to test routing and target-knowledge interventions on held-out questions. Repeated interventions and label controls support the robustness of the measured effects, while a language extension provides no handoff evidence.
- Expanded evaluation: All 192 held-out questions preferred the correct candidate, using 24 non-overlapping country pairs and eight variants per pair.The evaluation covered 48 countries excluded from the original fitting and evaluation partitions.
- Layer sets: The analysis fixed earlier and later comparison layers separately for Qwen, Gemma, and Llama rather than treating them as continuous processing stages.Qwen used layers 29 and 31 versus 32 and 34; Gemma used layer 24 versus 25 and 30; Llama used layers 15 and 24 versus layer 28.
- Uncertainty: Bootstrap intervals resampled country pairs with all eight variants and compared conditions within each resample.The procedure used 2,000 pair-level resamples and produced percentile intervals for specified contrasts, not simultaneous bands across layers.
- Reproducibility: The three-model experiment was computationally reproducible: fresh-load primary and repeated answer margins agreed for every recorded question–layer condition.The agreement covered 6,912 Qwen, 6,528 Gemma, and 5,376 Llama records, without pooling repeats as independent observations.
- Routing controls: The true routing direction passed both fixed-earlier comparisons, although the separate layer-29 interval spanned zero while layer 31 had positive intervals.The reported comparisons retained these single-layer results rather than collapsing them into one continuous window.
- Intervention controls: Repeated routing and target-knowledge deletions were matched to random perturbations, with rounding-aware length controls enforcing relative length error within 5% at nearly all active positions.Only 12 of 12,288 earlier-routing positions and 5 of 7,680 later-routing positions exceeded the threshold after correction.
- Scope boundary: The original-language extension stopped before intervention fitting because only 52.3% of 768 screening questions were correct and no complete disjoint pairing was available.It therefore supplies no intervention evidence about the routing–content handoff.
D Cross-Format and Fitting-Robustness Checks
Cross-format and robustness checks preserve the distinction between fitted content and routing while showing that their layerwise effects depend on model and comparison protocol. Expanded profiles and raw-damage decompositions clarify where corrected effects change.
- D.1 Fixed Knowledge Changes across Answer Forms: The fixed knowledge direction was transferred from noun answers to adjective and code formats under a registered positive-shift criterion.Each model tested six code mappings, with validation requiring positive random-corrected shifts across both partitions and most country pairs.
- D.2 Fit and Random-Seed Sensitivity: Content effects passed all listed comparison layers for Qwen, Gemma, and Llama, whereas routing robustness varied across models and layer sets.Qwen passed 3/4 fits earlier, Gemma 4/4, and Llama only 2/4; later routing passed 0/4 while content passed 4/4.
- D.2 Fit and Random-Seed Sensitivity: The expanded profiles report paired routing and target-knowledge pointing fractions and answer effects with separate model-specific scales and bootstrap shading.Figure 12 covers 24 new country pairs using each model’s full decoder depth and separately fitted components.
- D.3 Raw Damage, Random Damage, and Deletion Length: Qwen’s mean true route-deletion damage drops from 0.8021 at layer 31 to 0.0104 at layer 32, while mean strongest-random damage rises from 0.2500 to 0.4167.Deletion length increases from 11.2052 to 12.0985 across the same transition.
- D.3 Raw Damage, Random Damage, and Deletion Length: Gemma shows the same directional decomposition at its transition: true damage falls from 3.2083 to 0.2292 between layers 24 and 25, while random damage falls from 1.2500 to 0.6875.These raw values support a reduction in true deletion damage rather than attributing the corrected-score change solely to the random baseline.
- E.1 Why the Paired Three-Model Records Are Not a Direct Replication: The paired protocol uses a global first-versus-second request direction, unlike the natural analysis’s pair-conditioned direction and normalization.Its normalized coordinate is within-model and cannot be numerically compared with the natural route-strength measure.
- E.1 Why the Paired Three-Model Records Are Not a Direct Replication: The expanded study retains both validation partitions and marks layers only when both partitions and at least 80% of individual questions meet the targeted-effect gate.Lines average validation A and B, while shading spans their two values.
E.2 Model-Specific Trajectories
The paired three-model trajectories do not form one universal route-first schedule. Qwen, Llama, and Gemma differ in when routing coordinates are readable, when routing deletions affect answers, and when fitted content becomes consequential.
- Qwen: Qwen’s global request coordinate enters a higher mid-layer range before fitted-content deletion becomes stably consequential, but global-route deletion never passes the strict answer or later-fact gates.The paired protocol therefore complements rather than reproduces the natural-question route-leading result.
- Llama: Llama’s request coordinate is high early and through roughly layer 15, while routing has no sustained answer-effect window and content becomes stable earlier.Only code answers show two isolated routing passes, at layers 15 and 24; no source passes the fixed layer-28 fact gate.
- Gemma: Gemma provides a partial overlap: the request coordinate is strong before a continent-name routing pass at layer 24, while content passes at layer 21 and continuously from layers 25–33.Adjective and code formats retain sustained content effects without a sustained routing window.
- Cross-model comparison: Across models, the comparison separates coordinate strength, answer dependence, and later-fact dependence rather than establishing a universal ordering.Qwen is the clearest natural route-leading case, Gemma a partial paired analogue, and Llama a boundary case with dispersed routing effects.
- Protocol: The paired analysis holds Qwen weights, computation, scoring, country pairs, and answer task fixed while changing paired templates and request position.Template A constructs pair-conditioned directions, whereas templates B and C evaluate them.
- Direction construction: The fitted request constructions exclude both paired-task content and an auxiliary marker contrast before normalization and distinguish shared-fit from country-pair directions.The natural comparison uses its own pair-conditioned direction and pair center with task-specific content exclusion.
F.2 Aggregation, Uncertainty, and Actual Length
The common-protocol analyses aggregate effects at the country-pair level, compare interventions with matched controls, and expose a substantial dependence on deletion definition and length. Selection readouts provide a finite linear proxy whose numerical checks pass but whose scope remains limited.
- F.2 Aggregation and uncertainty: All 24 country pairs qualify before intervention, so qualified-pair and all-pair estimates coincide under the fixed bootstrap aggregation.The maximum random-control mean is recomputed after pair averaging within each bootstrap sample.
- F.2 Actual length: Table 6’s late effects depend strongly on deletion definition and projected length: mean deletion lengths are 16.38 for the global condition versus 80.47 for the pair-conditioned condition.Matched random controls do not equalize the true interventions, so the contrast is not orientation alone.
- F.2 Actual length: At layer 32, corrected layer-36 knowledge loss is −82.7 for the global direction versus 2648.4 for the pair-conditioned direction, a direct difference of 2731.1.The intervals are −330.1 to 72.2, 1223.3 to 3611.0, and 1382.9 to 3691.5, respectively; these are not percentages.
- F.2 Uncertainty sensitivity: Pooling random controls leaves the deletion knowledge effect inconclusive, while the pooled reversal effect remains positive at 6270.8 [3261.0, 8747.7].The sensitivity analysis does not upgrade deletion necessity, although directional reversal remains supported.
- Selection readouts: Selection readouts separately target country identities, each candidate continent, and the requested first/second label using ridge-fitted linear components.The candidate continent targets are three-dimensional one-hot vectors, while selection is a scalar +1 or −1 target.
- Selection controls: The direct selection direction is unit-normalized, and controls include orthogonal Gaussian perturbations plus a wrong-position swap matched to the true swap length.No answer association is supplied in these inputs.
- Selection qualification: Sources qualify only with source-content preservation of at least 0.70, selection change of at least 0.25, positive Equation 46, and at least three receiving layers meeting F_l→l′,joint ≥0.60.Numerical qualification additionally requires 95% of active cases to satisfy 5% target-write and random-length tolerances.
- Selection limitation: All 1,584 actual writes per model load passed the 5% numerical tolerance, but finite linear readouts incompletely measure identity and candidate knowledge.The results do not justify inferring that selection is absent from the model.
F.6 Base-Model Capital Selection and Two-Continuation Restoration
The base-model capital task isolates which of two named countries is requested and tests whether that selection can be causally transmitted across layers. The validated intervention changes the selected object while preserving other fitted readouts, with restoration confined to the question-end source and layers 27–29.
- The task keeps the country pair and capital relation fixed while changing whether the first or second listed country is requested.This isolates object selection from the named countries and relation.
- The fitted protocol uses separate country, capital, and two-entry selection targets, with disjoint screening and validation pairs fixing the source at layer 26.Eight disjoint pairs screen all layers and five input positions before held-out validation.
- The intervention passes the two-stage validation criterion on 37/40 questions in validation A and 44/44 in validation B while preserving all three other readouts.The downstream selected-object change exceeds all controls at layers 27–30.
- Two continuation branches restore either the original-request or opposite-request selection coordinate without copying the full clean hidden state.The branches add only fitted-basis projections at the receiving layer.
- The confirmed result is limited to the layer-26 question-end source and layers 27–29; an earlier positional edit failed the same two-usable-candidate definition.Layer 30 is also excluded because validation A reached only 23/40 on the two-answer criterion.
F.7 Capital Replication in Three Instruction Models
The three-instruction-model replication applies the same capital-selection framework across Qwen, Llama, and Gemma, with model-specific screening and held-out gates. Qwen and Llama fail to open the full later stage for different reasons, while Gemma shows a near-threshold candidate but fails held-out validation.
- F.7 Capital Replication in Three Instruction Models: The replication keeps the capital relation, phrasings, disjoint splits, scoring rule, and functional gates fixed while replacing base weights with three instruction models.Each model uses its own chat template and decoder depth, with float32 hidden-state calculations; the task contains 196 questions.
- F.7 Capital Replication in Three Instruction Models: Ridge regression fits four readouts from fit questions: both country identities, both candidate capitals, and a first-versus-second selection label.Country and capital targets use fit-only principal-component coordinates, and held-out results do not choose rank or coefficient.
- F.7 Capital Replication in Three Instruction Models: Screening tests source preservation, selection movement, answer effects, and at least three qualifying receivers before held-out validation can open.Two covariance-based random directions and a wrong-position application provide controls.
- F.7 Capital Replication in Three Instruction Models: Qwen has no qualifying source because its best candidate’s mean answer change is 0.008 answer-margin units below the strongest control and its best joint count is 5/14.The candidate nevertheless preserves non-selection readouts on 10/14 questions and moves selection by 0.995 of the full target difference.
- F.7 Capital Replication in Three Instruction Models: Llama’s source preserves the other readouts and strongly changes selection, but only two receivers qualify, so held-out validation is not opened.Its joint counts are 14/16, 11/16, and 9/16 at layers 15–17; ten passes are required.
- F.7 Capital Replication in Three Instruction Models: Gemma is the only model with a screening-passing candidate, but both 52-question held-out groups fail the fixed-receiver gate and restoration remains locked.Validation A produces 32, 32, and 31 joint passes; validation B produces 31, 31, and 29, while every receiver requires 32/52.
- G Natural-Question Wording and Complete Country Lists: The natural-question appendix uses separate single-country inputs, three question templates, and disjoint fitting, screening, and validation country sets.Template A supplies fitting and request-direction states; templates B and C supply intervention-evaluation questions, with each validation pair contributing four evaluated questions.
- G.3 Screening and Validation Pairs: Validation pairs are independent resampling units, and fixed source and receiver layers are selected before new countries are evaluated.The appendix lists the country-pair structure and keeps each pair’s wordings and intervention conditions together.