Source-linked AI summary
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
Markus J. Buehler
TL;DR
The paper asks whether correct scientific answers reveal mechanism representations or merely shortcuts. It tests readability, relational state transformations, and causal control, finding that physical relationships are more visible in controlled state changes than absolute states alone.
Problem
Correct scientific answers do not establish that language models represent or use governing mechanisms rather than relying on lexical cues, memorized associations, or numerical shortcuts.
Method
The paper separates readability, representation, and causal use using vocabulary readouts, option-free hidden-state geometry, controlled state transformations, and a neutral-anchored benchmark.
Results
Across readout, geometry, relational, and intervention tests, mechanism information was readable and causally steerable, while absolute geometry was confounded by lexical or numerical comparison.
Takeaways & Limitations
Physical relationships were more visible in controlled state changes than in absolute states alone, with selected directions controlling constrained scientific decisions.
Takeaways & Limitations
The study uses one model, and its neutral-anchored relational benchmark is narrow, development-selected, explicitly staged, and based on one frozen state direction.
Abstract
from arXiv · showhide
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.
1 Introduction
The study asks whether an open-weight language model merely produces correct materials-science answers or represents and causally uses the governing mechanisms. It separates concept readability, relational representation, and causal use through constrained readouts, state comparisons, and interventions.
- Problem: Correct scientific behavior alone does not establish that the model represents or uses the mechanism relevant to its answer.The introduction frames this as the central interpretability problem for increasingly capable AI systems in scientific research.
- Motivation: Materials science is a demanding test because reasoning must connect indirect evidence and mechanisms across atomic, mesoscale, and structural-performance scales.Examples include loading history, fracture surfaces, heat treatment, and lattice morphology.
- Measurement strategy: Gemma’s intermediate 2,560-dimensional states across 42 layers are measured with constrained lenses that reuse the fixed decoder, producing vocabulary rankings rather than generated explanations.Layers are numbered 0 through 41; direct, tuned, and Jacobian lenses make different assumptions about intermediate representations.
- Study aim: The investigation tests whether Jacobian-lens measurements transfer from Claude and broad language tasks to an open Gemma checkpoint and indirect materials-science descriptions.It also examines whether internal directions can steer model behavior and inform future training objectives.
- Conceptual framework: The study distinguishes readability, representation, and causal use rather than treating recovery of a known concept as evidence of mechanistic reasoning.Controlled readouts recover declared physical words, open discovery searches for recurring word sets, and relational tests compare complete hidden vectors without selecting vocabulary targets.
- Interpretive limitation: Readable materials-science information need not uniquely encode constitutive physics, so causal intervention is required to determine whether internal directions change scientific decisions.Apparent geometric organization may instead reflect lexical identity or numerical comparison.
2 Results
Jacobian vocabulary readouts were highly reproducible and sometimes improved mechanism-term recovery, but gains were selective across families and prompts rather than universal. Target-free vocabulary exposed meaningful engineering neighborhoods, while filtering and exact-word criteria revealed substantial noise and limitations.
- Jacobian readout recovery: Mean Jacobian recovery AUC was 0.025765 versus 0.012832 for direct unembedding, a +0.012933 difference, but family-level evidence supported selective rather than universal improvement.The relative increase was 100.8%, with a 95% bootstrap interval of −0.018674 to +0.048609 and one-sided family sign-flip p = 0.2344.
- Jacobian readout recovery: Family effects varied substantially: boundary attack improved most at ∆AUC = +0.1167, while cyclic damage favored direct unembedding at −0.0550.Notch resistance, rapid transformation, ductile failure, and line-defect motion also favored Jacobian readouts; particle strengthening favored direct unembedding, and three families tied.
- Prompt-level examples: Individual prompts showed both strong recoveries and counterexamples, including corrosion ranked 1/1/1 under Jacobian lenses versus 111 directly, but fatigue ranked 1,279–1,345 versus 7.Other examples included coalescence at 24/25/18 versus 17,418, and strengthening at 351–402 versus 11.
- Jacobian readout recovery: The three independently fitted Jacobian lenses agreed almost perfectly, with Spearman correlations of 0.9987, 0.9980, and 0.9978 across 150 controlled pairs.Their family-clustered bootstrap lower bounds were 0.9975, 0.9959, and 0.9945, whereas the Jacobian mean versus direct unembedding correlated only 0.212.
- Open vocabulary discovery: Exact declared-word overlap favored direct lists, appearing in 2 of 10 Jacobian lists versus 6 of 10 direct lists, although semantic neighborhoods captured related terms.Examples included boundaries for boundary and martens for a martensitic neighborhood; weak families also produced generic or indirect vocabulary.
- Open vocabulary discovery: Decontaminated lists identified 29 of 50 prompts (58%; macro-F1 0.568) versus 32 of 50 (64%; macro-F1 0.615) for direct lists, while both exceeded the 10% balanced-class null.A prompt-word TF–IDF classifier reached 56%, and direct-only correctness exceeded Jacobian-only correctness by 8 to 5 prompts; filtered displays removed connective-word noise but did not imply causal word sequences.
3 Discussion
The discussion concludes that materials-science information is reproducibly readable and can causally influence selected answers, but absolute representations do not uniquely establish constitutive physics. Controlled relational comparisons and interventions provide stronger evidence while exposing benchmark, transfer, and readout limitations.
- Readouts: The Jacobian lens is reproducible but selective: 36 of 50 controlled prompts are zero under both readouts, with no family-level mean advantage.Independent WikiText fits yield nearly identical term ranks, but reproducibility does not imply universal superiority.
- Readouts: Both predetermined and target-free word sets enable blinded identification of 9 of 10 mechanism families, while zero of 523 target-free words survives enrichment correction.Prompt TF–IDF remains competitive for family classification, so the readouts are complementary rather than uniquely physical.
- Absolute versus relational structure: Full-state geometry, direct unembedding, prompt embeddings, and graph neighborhoods are descriptive evidence because physical structure can collapse to lexical or numerical comparison.Within each family, the graph target is algebraically identical to numerical-direction agreement when y = sx.
- Absolute versus relational structure: Controlled numerical reversals isolate relational organization: the neutral-anchored benchmark removes fixed context and distinguishes direct, neutral, and inverse physical laws.This experiment changes the measured object from absolute state placement to endpoint transformation, avoiding the earlier graph’s global sign ambiguity.
- Late-stage convergence: Relational signals and physical-equivalence geometry converge late in the network, with geometry rising after approximately 60% depth and grain-size patching reaching 10% and 50% of peak at 58.5% and 63.4%.The disjoint 12-law relational sweep is continuously strong across the latter half of the network.
- Causality and limitations: Interventions establish selected causal effects rather than universality: the grain direction reverses its output effect between refinement and coarsening, while transfers generalize across answer formats and orientations.The discussion cautions that readable structure is not automatically constitutive physics, and direct optimization risks readout hacking.
4 Conclusion
Materials-science information is readable in Gemma, but absolute representation geometry is partly confounded by lexical and numerical comparison. Controlled state transformations and selected causal directions reveal a narrower physical abstraction that controls constrained scientific decisions.
- Conclusion: Materials-science information is reproducibly readable, while much absolute geometry remains compatible with lexical or numerical comparison.A narrower physical abstraction appears when measurement follows controlled transformations between states.
- Conclusion: The framework combines calibrated measurement, representation analysis, intervention, and identifiability to locate readable structure, causal control, and limits of physical interpretation.It distinguishes what is present without Jacobian transport, what survives option-free controls, and where causal effects fail under transfer.
- Future work: Follow-up evaluations should preserve neutral anchors and matched reversals while testing graded responses, standardized option-free prompts, unseen governing laws, and additional model families.The proposed mechanism, constitutive-orientation, counterfactual, graph-persistence, intervention-specificity, and uncertainty scores could serve as separate fine-tuning rewards under readout-independent evaluation.
5 Materials and Methods … 5.3 Mechanism-direction construction
The study fixed a specific Gemma model, inference protocol, and layer grid, then compared direct decoding with Jacobian-transported readouts. Mechanism directions were fitted from semantic concept contrasts, independently of tested answer words, with layers selected in a preliminary study.
- 5.1 Model, revisions, and registered layer grid: Gemma-4-E4B-it used immutable revision a4c2d58be94dda072b918d9db64ee85c8ed34e3f, 42 transformer layers, and hidden width 2,560.Evaluations used bfloat16 inference, a fixed final-prompt-token position, one greedy continuation token, 25 registered layers, and a 38–92% source-depth band.
- 5.1 Model, revisions, and registered layer grid: The vocabulary contained 262,144 tokens, and held-out evaluations followed a fixed registered layer grid and source-depth band.The protocol also fixed the final-prompt-token score position and used one greedy continuation token for output-leakage checks.
- 5.2 Jacobian-lens estimator and direct baseline: Jacobian lenses estimated layer-to-final transport from source states to target-layer positions, with causality restricting contributions to later or equal target positions.The implementation excluded the first 16 positions and the final position without a next-token target, then averaged valid source positions within records and records across the dataset.
- 5.2 Jacobian-lens estimator and direct baseline: Both Jacobian and direct readouts used Gemma’s own vocabulary decoder, differing only in whether layer-to-final transport was applied.No separately trained classifier converted states to words; direct unembedding used the matched score.
- 5.2 Jacobian-lens estimator and direct baseline: Three Jacobian lenses were fitted on 1,000 unique 128-token wikitext-103-raw-v1 training records using corpus seeds 0, 1, and 2.All lenses shared the same model revision, penultimate target, 25 source layers, and released estimator.
- 5.3 Mechanism-direction construction: Mechanism directions used four prespecified positive and four negative concept tokens per family rather than the answer words being tested.For each frozen direction-fitting prompt, the concept score gradient was pulled back through the fitted Jacobian map and normalized.
- 5.3 Mechanism-direction construction: Three prompt-specific unit vectors were averaged and renormalized into one direction per family, lens fit, and layer, with a matched direct control through final_norm without J.Answer outcomes such as higher, lower, grooves, clean, hard, and soft were excluded from direction construction.
- 5.3 Mechanism-direction construction: Candidate layers 10, 16, 20, 24, and 28 were evaluated once in a disjoint preliminary study, selecting layers 16, 24, and 16 for corrosion, transformation, and grain size.Selection used the mean signed −4% to +4% endpoint minus one population standard deviation, with the lower layer breaking ties; the broad screen reused those layers.
5.4 Localized intervention and scientific-answer scoring
The evaluation localized interventions to the final prompt token while leaving other positions and layers unchanged, then scored complete answer strings by teacher-forced log probability. Conditions used symmetric intervention doses, reversed answer orders, matched Jacobian and direct directions, and seeded random-vector controls.
- Localized intervention: Intervention was restricted to the final prompt token at the frozen source layer, with all other positions and layers unchanged.The intervention used ℓ,t = hℓ,t + a 100∥hℓ,t∥2d with a ∈ {−4, −2, 0, 2, 4}.
- Scientific-answer scoring: Complete lowercase answer strings were scored with teacher forcing using exact no-leading-space continuation tokenization.For multi-token answers, intervention remained at the original final prompt position while later answer pieces used ordinary network attention.
- Scientific-answer scoring: The measured contrast was positive-answer minus negative-answer log probability, with the endpoint subtracting the −4% contrast from its +4% value.This defined the intervention-dependent answer preference contrast.
- Control directions: Each condition was tested twice with answer words reversed, using three matched Jacobian directions, a matched direct direction, and ten seeded random unit-vector controls.Random vectors were orthogonalized against the seed-0 matched Jacobian direction and the remaining independent direct-direction component.
5.5 Steering cohorts, endpoints, and inference · 5.6 Exploratory full-state activation patching
The study used large, preregistered steering cohorts with explicit endpoints and independent-unit inference, then explored full-state activation patching using matched donor states and counterfactual-aligned log-odds shifts. Additional controls addressed perturbation magnitude and compared causal sensitivity with representation readability.
- 5.5 Steering cohorts, endpoints, and inference: 6,000 intervention rows covered 60 exact prompts across mechanism directions, random directions, direct directions, and five perturbation doses.The broad screen used ten conditions per family and two answer orders, averaging answer orders and applicable lens fits at the physical-condition level.
- 5.5 Steering cohorts, endpoints, and inference: 2,400 intervention rows came from six new matched material pairs, each contributing refinement and coarsening conditions in both answer orders.The primary endpoint was defined so positive values always indicated movement toward the frozen correct answer.
- 5.5 Steering cohorts, endpoints, and inference: The dose-shape audit reused all five collected doses, while inference retained six independent material pairs with 30,000 bootstraps and exact 26 sign flips.Strict monotonicity was descriptive, and the analysis was explicitly a robustness analysis rather than a second prospective confirmation.
- 5.5 Steering cohorts, endpoints, and inference: Two later transfer protocols froze the unchanged grain direction, layer, doses, scoring, and answer orders before model outputs were generated.One cohort used six new polycrystals with complete-string increase/decrease answers, while the inverse Hall–Petch cohort used six nanocrystalline materials below their material-specific crossover.
- 5.6 Exploratory full-state activation patching: Activation patching reused 24 exact prospective-study prompts and replaced final-prompt-token residuals at each of 25 registered layers with matched donor residuals.The protocol was frozen before patching output but after the original steering output had been inspected.
- 5.6 Exploratory full-state activation patching: Counterfactual-aligned shifts were defined from exact next-token higher-minus-lower log odds, with positive values indicating movement toward the answer appropriate to the opposite relation.The primary scalar averaged over the fixed 38–92% band, using matched material pairs as independent units and 30,000 pair-cluster bootstrap resamples.
- 5.6 Exploratory full-state activation patching: A separate falsification control rescaled same-relation and order-only directions to match the Euclidean norm of reverse displacements before repeating patching and inference.This post hoc control tested perturbation magnitude but could not make the inspected cohort an independent confirmation.
- 5.6 Exploratory full-state activation patching: Readability and causal sensitivity were compared by correlating 25-layer relation-separation curves from Jacobian and direct readouts with matched patching curves against a 49-member circular-shift null.Relation separation was computed within material from refinement and coarsening contrasts averaged over answer orders and, for Jacobian readout, lens fits.
5.7 Frozen cross-mechanism activation patching · 5.8 Practical command-line workflow for a new experiment
Frozen cross-mechanism patching uses a prespecified factorial design to test whether donor states transfer mechanism-specific decision signals. The command-line workflow makes lens fitting, controlled prompts, steering studies, graph audits, patching analyses, and new-mechanism confirmations reproducible from frozen manifests and records.
- 5.7 Frozen cross-mechanism activation patching: 1,920 patches replaced each receiver’s final-question-token residual with donor residuals across 24 receivers, 480 ordered pairs, and four frozen layers.The factorial design covered four material cases in each of six mechanisms and layers 16, 24, 32, and 37.
- 5.7 Frozen cross-mechanism activation patching: Patching effects were computed from positive-minus-negative answer margins and averaged across receivers, frozen layers, and both directions of each mechanism pair.The scoring alternatives were higher/lower or greater/smaller, with prespecified subsets including all 15 pairings.
- 5.7 Frozen cross-mechanism activation patching: Strong evidence required both exact pair-sign and structured donor-label tests, with at least five of six donor-family physical-outcome contrasts above zero.The structured test enumerated 46,656 balanced donor-family label assignments, and 30,000 fixed-seed pair bootstrap resamples supplied 95% intervals.
- 5.8 Practical command-line workflow for a new experiment: The repository exposes operational commands for fitting, prompt execution, raw-data storage, and visualization, including --fit-only lens saving and optional Hugging Face upload.The fit-only workflow writes the weight file and metadata sidecar without evaluating prompts.
- 5.8 Practical command-line workflow for a new experiment: A controlled prompt is represented as a JSON record specifying its mechanism shape, domain, text, final-token readout, tracked terms, and absence constraints.The example fatigue prompt tracks fatigue, crack, and propagation while requiring those terms to be absent from input and output.
- 5.8 Practical command-line workflow for a new experiment: One exploratory prompt uses the demo wrapper because the quantitative wrapper requires at least 50 independent items per evaluated shape, while retaining the 1,000-record paper-recipe checkpoint.The demo command writes machine-readable results and token-by-layer and semantic-stream visualizations, plus tracked-term diagnostics when applicable.
- 5.8 Practical command-line workflow for a new experiment: Versioned manifests and analysis scripts regenerate steering studies, representation and graph audits, option-free analyses, and cross-mechanism patching outputs, with analysis-only rebuilding available without another Gemma forward pass.The workflow also supports Apple-silicon execution via --device mps instead of cuda and stores committed machine-readable products for inspection.
- 5.8 Practical command-line workflow for a new experiment: For a new mechanism, users freeze the layer and controls after development, define concept ensembles and direction prompts, and keep scored outcome words out of direction ensembles.The protocol specifies two four-word concept ensembles, three direction-fitting prompts, two scientific answer strings, a frozen layer, and confirmation conditions; raw JSON records rendered prompts.
5.9 Held-out evaluation design and leakage controls
The held-out evaluation used a preregistered-style, duplicate-free 50-prompt manifest finalized before execution, with low lexical overlap with the earlier suite and explicit tokenizer-based leakage exclusions.
- Held-out evaluation design: 50 prompts were finalized before held-out execution, with five prompts per Table 2 family and no duplicates of development items.The exact manifest was finalized on 14 July 2026.
- Leakage controls: 0.0061 mean word-5-gram Jaccard overlap and 0.0952 maximum overlap limited lexical similarity with the earlier association suite.These overlap values were measured against the earlier association suite.
- Leakage controls: No items failed the rule requiring every tokenizer-resolved declared term to be absent from the input and generated one-token continuation.Failures would have been excluded identically for all methods and seeds; multi-token terms were excluded from the one-token rank endpoint as registered.
5.10 Lexical-adversarial physical-equivalence cohorts
The section describes two frozen lexical-adversarial protocols testing whether physical equivalence is distinguished from wording changes across mechanism families and material systems. After the primary analysis failed, a disjoint second cohort was introduced with unchanged triplet construction and lexical preflight.
- First protocol: The first protocol covered six mechanism families, each with four material systems, using anchors, physically equivalent paraphrases, and near-verbatim lexical counterfactuals.The families included grain size/yield strength, crack size/fracture stress, temperature/diffusivity, fiber angle/axial stiffness, fiber fraction/elastic modulus, and martensite fraction/hardness.
- Scoring: The registered margin compared anchor similarity with physical-paraphrase and lexical-counterfactual similarity, with positive values indicating that physics outranked wording.The primary scalar averaged this margin over the fixed 38–92% layer band for the three-fit Jacobian ensemble.
- Second protocol: After the primary analysis failed and its retained layer curve showed an unregistered late rise, a second protocol used six new mechanisms and 24 new material systems.The exact triplet construction and lexical preflight were unchanged, and the cohort was disjoint from the first protocol.
5.11 Graph construction, falsification, and statistical tests
This section defines the graph constructions, frozen comparisons, null models, and statistical tests used to assess mechanism-specific organization and its robustness. It also records that several audits were post hoc and that the natural-question comparison does not isolate a single linguistic cause.
- Cross-phrasing mechanism graph: The cross-phrasing graph used 50 descriptions, ten mechanism families, five phrasing folds, 25 layers, and three fitted lenses, with 200 directed edges per layer and 10% balanced chance.Edges linked each source to its most similar target in each other phrasing fold, using cosine similarity averaged across the fixed 38–92% band.
- Falsification and statistical tests: Statistical controls included a word- and character-overlap regression, blocked and exact nulls, pairwise ROC–AUC, mean reciprocal rank, leave-one-mechanism-out AUC, and family-bootstrap uncertainty.The standard null used 50,000 permutations, while the exact base-case null enumerated 46,656 balanced assignments; family-bootstrap uncertainty used 50,000 resamples.
- Signed-relation graph across different material cases: The signed-relation graph used 72 prompts and 144 directed edges, restricting each source to nearest neighbors from other material cases within its mechanism family.The prompts comprised six mechanism families, four material cases per family, and three surface variants per case: anchor, physical paraphrase, and lexical counterfactual.
- Natural-question position audit: A natural-question robustness run reused the identical 72 scientific stems but stopped at the natural final token without answer choices, answer words, A/B codes, response-format instructions, or checkpoint markers.The three-position comparison also included the contextual pre-mapping checkpoint and post-mapping final state, with every method using the same 38–92% layer band.
- Graph-identifiability theorem and audit chronology: The broader graph-identifiability, continuous cross-law, label-blind matching, graph-learning, community, sparsification, and partition audits reused existing states and were explicitly post hoc where stated.The graph-identifiability audit introduced no new prompts or Gemma forward passes, while the three-position comparison was designed to map realistic analysis boundaries rather than isolate one linguistic cause.
5.12 Neutral-anchored relational constitutive benchmark
The neutral-anchored relational benchmark tests whether frozen hidden-state changes follow supplied constitutive laws rather than relying on absolute-state geometry. It uses balanced direct, inverse, and neutral relations with matched prompts and preregistered contrast-based endpoints.
- Task construction: The benchmark asks the model to classify supplied equations as direct, inverse, or unchanged and infer whether a numerical control rises or falls, without intermediate answers or rationales.Method development used 16 earlier laws.
- Task construction: The final cohort contains 20 direct, 20 inverse, and 20 physically neutral relations across 13 domains, yielding 960 exact prompts and 480 matched comparisons.Each law varies algebraic surface, material case, numerical direction, and answer order; matched prompts differ only in the physical input direction.
- Frozen state direction: The frozen readout uses unit-normalized layer-34 residual states and development-set centroids, with no refitting on the 60-law cohort.The contrast is raw-state based rather than Jacobian-transported or vocabulary-decoded.
- Neutral anchoring: Direct laws predict positive contrast, inverse laws negative contrast, and neutral relations no systematic change after calibration to the validation-neutral median and robust scale.Direct, inverse, and validation-neutral labels do not affect the calibration parameters.
- Endpoints and inference: Primary endpoints include pairwise ROC–AUC comparisons, calibration-neutral direct/inverse accuracy, equation-surface AUC, and Spearman correlation with registered law sign.AUC intervals use 50,000 class-stratified law bootstraps, while the ordinal test uses 100,000 fixed-seed labels.
5.13 Answer-scaffold and arbitrary-code audits
This section audits answer-scaffold effects and arbitrary-code falsification, emphasizing that the scaffold comparison was post hoc and that pre-choice states were collected before answer mappings were available. The arbitrary-code design separated scientific content from assigned answer labels, with an initial tokenization mismatch preventing marker detection.
- Answer-scaffold comparison: The answer-scaffold comparison was post hoc, combining pre-choice layer-39 states with ordinary final-output states from disjoint scientific stems.The pre-choice checkpoint followed the complete scientific question but preceded answer mapping, whereas the ordinary state followed a supplied semantic answer pair.
- Arbitrary-code falsification: The arbitrary-code falsification assigned different A/B codes to anchor and physical paraphrase despite shared answers, while counterfactuals had opposite answers but both required A.Its contextual checkpoint preceded the mapping, preventing access to future assignments; the first execution found no marker because isolated and in-context tokenization differed.
5.14 Controlled recovery and population inference
Controlled recovery was evaluated with rank-based pass@k and normalized AUC measures, while population-level uncertainty and lens agreement were assessed using hierarchical resampling, exact sign-flip testing, and family-level correlations. A leave-one-mechanism-family-out audit tested whether these summaries depended on any single family.
- Controlled recovery: Recovery used each term’s best full-vocabulary rank across registered layers, evaluating pass@k at k ∈{1, 2, 5, 10, 20, 50, 100} and normalized AUC versus log k.Jacobian AUC was averaged across three fitted lenses within each prompt; lens fits were treated as repeated measurements rather than populations.
- Population inference: Population inference contrasted prompt-level Jacobian mean AUC with direct AUC using a 20,000-resample hierarchical bootstrap and an exact sign-flip test over 210 = 1, 024 family-level assignments.The bootstrap sampled mechanism families with replacement, then phrasings within sampled families.
- Population inference: Lens agreement was measured with pairwise Spearman rank correlations over all 150 resolved prompt–concept rows, with 5,000 bootstrap intervals obtained by resampling whole families.The correlations compared the fitted lens results across resolved prompt–concept cases.
- Population inference: A post hoc influence audit removed each mechanism family in turn and recomputed the remaining nine-family mean AUC difference and all three pairwise lens-fit Spearman correlations.The audit introduced no new threshold or hypothesis test.
5.15 Complete-sequence scoring for multi-token technical terms · 5.16 Target-free candidate generation · 5.17 Target-free cross-phrasing classification
The analyses extend mechanism readouts to complete multi-token terms, generate candidates without declared targets, and classify mechanism families from filtered consensus lists under preregistered controls and baselines.
- 5.15 Complete-sequence scoring for multi-token technical terms: Multi-token terms were excluded from controlled rank endpoints because first-piece rank does not measure complete-word probability.Robustness scoring separately examined transgranular versus intergranular and martensite versus bainite across their held-out prompts.
- 5.15 Complete-sequence scoring for multi-token technical terms: The robustness study scored two exact multi-token contrasts across all five held-out prompts in their respective families.The contrasts were transgranular versus intergranular, both split into two pieces, and martensite versus bainite, both split into three.
- 5.16 Target-free candidate generation: Candidate generation scanned unrestricted top-1 decoded tokens at every prompt position and retained only tokens from the fixed source-layer band.Tokens present in the input or constituting one-token continuations were removed before retaining candidates.
- 5.16 Target-free candidate generation: 64 candidates per stored run were retained, and a prompt-level candidate survived only when present in all three stored lens records.The same three-record consensus rule was applied to matched direct-unembedding lists; direct scores were identical across records.
- 5.16 Target-free candidate generation: Declared terms were excluded from candidate generation, filtering, retention, and ranking, with exact overlap added only afterward as descriptive annotation.Family candidates were ranked by mean three-record consensus score multiplied by log(50/fw), then averaged over five family phrasings.
- 5.17 Target-free cross-phrasing classification: The secondary classifier used complete filtered three-fit consensus candidate lists, with candidate weights incorporating inverse document frequency computed only from training prompts.The feature vocabulary was also restricted to training prompts in each fold.
- 5.17 Target-free cross-phrasing classification: A cosine nearest-centroid classifier was compared with a training-fold TF–IDF bag of input words and balanced label permutations as baselines.The supplied passage introduces these baselines but does not provide their resulting accuracy values.
- 5.17 Target-free cross-phrasing classification: Jacobian and direct candidates were evaluated under three frozen filters, including target-agnostic function-word removal and two stricter lexical-overlap controls.The most conservative rule removed candidate/input pairs of at least five letters when either word began with the other.
5.18 Semantic-stream construction · 5.19 Exploratory latent geometry · 5.20 Automated blinded secondary rating
These sections define retrospective semantic-stream displays, a post hoc exploratory latent-geometry protocol, and an opaque automated rating procedure. The geometry analysis used normalized hidden-state transports and controlled classification and visualization choices, while the rating withheld prompts, methods, and answer keys from the evaluator.
- 5.18 Semantic-stream construction: Semantic streams were retrospective displays rather than inferential endpoints, pairing two deterministic renderings of five selected frozen prompts.The unrestricted rendering used fit 0 and displayed seven words with the largest depth-integrated scores.
- 5.18 Semantic-stream construction: The unrestricted stream assigned each normalized leading token at rank r a layer score (W −r)/W within a stored list of width W.The passage specifies the score and ranking rule for the displayed words.
- 5.19 Exploratory latent geometry: The inferential geometry protocol was frozen before hidden-vector extraction but after held-out vocabulary results, making it post hoc exploratory.Raw final-prompt states were extracted for 50 prompts and 25 layers.
- 5.19 Exploratory latent geometry: The analysis averaged three normalized transports and renormalized them, producing 3,750 transported 2,560-dimensional vectors alongside raw, target-layer, lexical, and discovered-word vectors.The array combined transported and comparison representations.
- 5.19 Exploratory latent geometry: At every layer, nearest-centroid classification used cosine similarity with five phrasing folds, while a corrected null permuted ten balanced labels and retained the maximum across 25 layer accuracies.The protocol used 5,000 permutations, although the supplied passage truncates before completing that description.
- 5.19 Exploratory latent geometry: For Figure 6C, 50 best-layer vectors were reduced to 25 principal components explaining 85.1% variance, then embedded with UMAP using cosine distance, 10 neighbors, minimum distance 0.20, and a fixed seed.Five additional UMAP seeds and a two-dimensional PCA sensitivity plot were also recorded.
- 5.19 Exploratory latent geometry: The best-layer visualization was chosen post hoc for legibility after the joint all-layer display was dominated by layer progression, without affecting classification or permutation results.The passage explicitly distinguishes the visual choice from inferential analyses.
Data and code availability
The study’s code, data, protocols, outputs, analysis materials, and manuscript sources are retained in public repositories, with versioned manifests linking results to experimental inputs and outputs.
- The accompanying Substrates repository retains code, exact prompt manifests, versioned protocols, raw model outputs, analysis scripts, figure-generation code, and manuscript sources.Versioned machine-readable manifests link each reported result to its corresponding experimental inputs and outputs.
- The three fitted lens bundles are stored in the Hugging Face repository lamm-mit/gemma4-jacobian-lenses.