Source-linked AI summary

The information geometry of large language models is shared, learned, and controllable

Dario Picozzi

arXiv:2609.11063v1cs.LGcs.CL

TL;DR

The paper asks what structure large language models share and how behaviour can be changed without unnecessary disturbance. It develops Fisher–Rao output geometry as a coordinate-invariant framework and tests it across architectures, language statistics, acquisition, and interventions. The geometry predicts shared semantic and spectral structure, acquisition timing, and lower-disturbance reusable control across operations.

  • Problem

    It remains unclear what structure large language models share and how to change one behaviour without disturbing others.

  • Method

    The paper uses Fisher–Rao geometry of next-token probabilities to analyze shared structure, language statistics, acquisition, and minimum-disturbance interventions.

  • Results

    Output geometry is shared across models, tracks human predictions, follows the language law, predicts acquisition timing, and improves matched control across steering, editing, attribution, dictionary learning, and fine-tuning.

  • Takeaways & Limitations

    A single pullback-metric correction provides a reusable, behavior-centered basis for comparing models and changing outputs while preserving other behaviour.

  • Takeaways & Limitations

    The simple margin equation does not hold for arbitrary structures under RMS population blur; valid statements must specify the observable, margins, Lipschitz constants, and error type.

Abstract

from arXiv · show

Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.

Results

The paper shows that next-token predictive behaviour defines a shared, coordinate-invariant output geometry whose structure reflects language statistics and human completion patterns. This geometry predicts knowledge acquisition and intervention costs, enabling reusable, lower-disturbance control.

  • Predictive behaviour fixes a canonical, identifiable output geometry: The pullback Fisher geometry predicts local output change and yields coordinate-covariant minimum-disturbance updates.Measured-to-predicted KL ratios had median 1.000 across 99 model–depth–objective cells, while reparameterisation tests preserved the natural-gradient direction.
  • Predictive behaviour fixes a canonical, identifiable output geometry: Behaviour identifies the language-defined read-out subspace, with held-out profiled loss growing quadratically with subspace distance.Across four model sizes, every path had R2 ≥0.9999, and controlled language assignments recovered more than 99% of imposed geometric separation.
  • Shared geometry and human completion structure: Human completion geometry improves with model scale and training, while model-only calibration further improves prediction of human word choices.Semantic distance fell from about 0.34 at 70M to 0.30 at 2.8B parameters, and the scale trend replicated on independent narratives.
  • The geometry inherits the statistics of language: Token statistics and read-out directions jointly predict spectral structure and effective dimension without fitted parameters.Held-out spectral-exponent errors were 0.0623–0.0856, while effective-dimension RMSEs were 0.0270–0.0521; probability concentration predicted mode counts and read-out geometry predicted decay.
  • Corpus n-gram statistics predict acquisition, and evidence depth shifts its timing: Corpus n-gram margins predict fact-acquisition trajectories, whereas randomized deeper evidence delays acquisition across tested settings.Trajectory R2 was 0.775–0.792 with timing errors of 0.77–0.96 log2 training-step units; deep evidence shifted acquisition by about 2.10 log2 steps.

Discussion

The paper argues that Fisher–Rao geometry provides a shared, behaviorally determined structure across models and a principled basis for low-disturbance control. Its evidence links predictive behavior to geometry, language statistics to acquisition and spectra, and the same correction to several model-analysis and intervention tasks.

  • Spectral measurements and effective-dimension prediction: Token statistics predict the spectrum and effective dimension without fitted parameters, while the weighted read-out profile improves spectral-shape prediction across all seven families.Probability concentration accounts for effective dimension, and the read-out profile captures additional spectral structure.
  • Knowledge acquisition: Corpus n-gram evidence predicts individual-fact acquisition with median absolute errors of 0.77–0.96 log2 training steps, while deeper randomized evidence causally delays acquisition.The prediction holds at all three tested model sizes, and the controlled language experiments establish the acquisition delay across architectures and evidence constructions.
  • Minimum-disturbance interventions: Fisher geometry predicts minimum-disturbance control costs and enables donor-prompt updates to transfer to unseen prompts while reducing reference-sequence change three- to sixfold versus Euclidean control.The paper connects this reusable control construction to steering, editing, feature analysis and training applications.
  • Applications and scope: The geometry can support model auditing by comparing checkpoints, post-training variants and independently developed models without shared hidden coordinates, architectures or tokenizers.Consensus–residual decomposition separates shared structure from model-specific deviations.

Data availability

The reported analyses use public checkpoints and benchmark datasets spanning language models, vision models, semantic encoders and human prediction datasets.

  • Data availability: Public resources include Pythia, GPT-2, GPT-Neo, Qwen, Qwen1.5, Mistral-7B, Mamba, RWKV, StarCoder2, BLOOM, OLMo, Sentence-BERT and DINOv2, plus WikiText, LAMBADA, SST-2, TruthfulQA, CounterFact, CIFAR-100 and human prediction datasets.Derived data supporting the findings are available from the corresponding author on reasonable request.

Extended Data

Extended-data analyses test coordinate covariance, shared output structure, effective dimension, acquisition timing, and geometric intervention costs. Together they show that output geometry is stable under valid coordinate mappings, captures cross-model structure, and supports calibrated prediction and control.

  • Coordinate covariance: Mapped reference metrics reproduce native damped directions with relative error at most 7.4 × 10^-13, whereas resetting the reference to identity changes the regularised problem.The same figure reports output-geometry agreement of 0.899 at token level and 0.910 after first-byte push-forward.
  • Shared output structure: Shared eigenvector structure beyond displacement magnitude concentrates in leading modes, with residual overlap falling from 0.080 at ranks 0–5 to 0.013 at ranks 90–120.About 89% of leading-band raw overlap is explained by shared displacement magnitudes; the residual carries rank-dependent cross-family structure.
  • Effective dimension: Bounded effective-dimension curves outperform unbroken power laws on all 80 held-out curves, reducing median error from 0.0814 to 0.00493.The median bounded-to-power error ratio is 0.0749.
  • Evidence depth and acquisition: Pretraining corpus statistics predict held-out fact-acquisition trajectories without refitting, while deeper evidence delays acquisition across controlled randomized assignments.The supplied extended-data passage establishes the delay experiment, while the quantitative trajectory metrics appear elsewhere in the paper.
  • Intervention geometry: The output metric predicts sparse-feature causal cost and supports lower-preservation-cost natural-gradient adaptation, while intervention benefits depend on layer and magnitude.The supplied figure passages specify sparse-feature selectivity, natural-gradient LoRA preservation, and layer-dependent correction, but do not provide complete numerical comparisons.
  • Concept geometry: Cyclic concepts are near-isometries of the metric, whereas analogy relations move probability mass along shared directions with mean alignment 0.49 versus 0.33 for permuted controls.The cyclic paths return toward their starting sectors after a full circuit, and natural-gradient steering accumulates less off-target change.

Identification of the read-out geometry

This section identifies the read-out Fisher geometry while separating it from non-identifiable hidden activation geometry. Local and global arguments show when read-out subspaces are recoverable and how predictive error transfers into recovery guarantees.

  • Local identification: The profiled Fisher curvature is positive normal to the containment fibre, enabling local identification of the target read-out subspace after profiling intercept and context coordinates.The proof map combines the local Schur-complement curvature with an inverse-softmax near/far argument.
  • Global identification: Global identification is exact up to the overcomplete containment fibre: K_D,d(S)=0 if and only if T⊂S, while rank-matched read-outs identify the unique subspace T.The global theorem separates uniformly near candidates from candidates with a fixed Hellinger penalty.
  • Approximate recovery: Approximate recovery converts predictive root-probability error into subspace-distance guarantees and inherited partition-capture bounds.The supplied corollaries formalize recovery and capture rates, while the proof uses the global margin from the identification theorem.
  • Geometric identifiability: The output Fisher metric is the first Hessian of the read-out’s Gibbs law, while hidden activation geometry remains non-identifiable under output-preserving invertible reparameterisations.The Fisher metric is tied to the affine-softmax read-out, whereas identical model behaviour can coexist with changed Euclidean activation geometry.
  • Controlled validation: Across the controlled synthetic-language experiment, all 96 labelled test runs met the common excess-risk requirement ϵ<Δ2/32.The protocol spans three architectures, two capacities, and eight independent pairs of 64-context, 64-outcome languages.

Supplementary Note 4

The supplementary results formalize why relational geometry converges with predictive fit while hidden geometry does not. They connect per-context probability disagreement, margins, coherence, tokenizer coarse-graining, and attenuation floors to observed cross-model agreement.

  • Excess-risk convergence: Theorem 6 bounds relational discrepancies by excess predictive risk, and under vanishing risk the pairwise agreement statistic converges to one.The rate depends on the reference-law regularity and the models’ excess risks.
  • Rank stability: Relational rank stability depends on entrywise margins and perturbation errors, so discordance is concentrated among pairs with small margins.The Kendall discordance rate is bounded by the probability that the margin does not exceed the perturbation magnitude.
  • Scope limitation: The simple margin equation is not valid for arbitrary observables under RMS population blur; valid guarantees must specify the observable, Lipschitz constant, margin, and error type.The remark distinguishes hard threshold counts from smooth effective-dimension responses.
  • Coherence mechanism: In trained models, large per-context disagreement does not destroy relational agreement because disagreement directions have low coherence with the relational structure.Median coherence is 0.024–0.034 for well-trained pairs, with observed stability substantially better than worst-case bounds.
  • Attenuation floors: The attenuation identity yields a positive agreement floor under bounded idiosyncratic deviations and nonnegative cross-covariance, but no corresponding floor exists for hidden-layer geometries.Identical output behaviour can accompany hidden-geometry Pearson agreement −0.2053387 and Spearman agreement −4/17.

Supplementary Note 5

The shared component of output-geometric representations carries semantic structure that transfers across models, while model-specific residuals do not. Independent activation-space comparisons and cross-model probing reinforce the distinction between shared and idiosyncratic structure.

  • Semantic structure: Shared-component alignment with external semantic geometry exceeds residual alignment, reaching 0.440 on factual relations and 0.231 after controlling surface form.Residual alignments are negligible on natural, templated, and factual-relation batteries.
  • Transferable geometry: A shared Fisher–Rao component supports eight-way cross-model probe accuracy of 0.659 and consensus-only accuracy of 0.700, versus chance 0.125.The model-specific residual reaches 0.407 within its own model but only 0.091 across models.
  • Cross-modal comparison: Last-token hidden-state geometry from Pythia-1.4B and GPT-2 Large agrees with vision-only DINOv2 geometry at 0.385, with 84% remaining after controlling textual co-occurrence.The supplied passage presents this as an independent activation-space extension beyond textual co-occurrence.

Part III: Learned

The output Fisher spectrum is constrained by token probabilities and read-out geometry, with effective dimension and spectral shape predicted across models. The learned structure includes a high-probability singleton head and interchangeable clusters, but the partition is not universal across contexts.

  • Spectral constraints: Finite-width majorization forces profile mass beyond rank d into the first d spectral modes but does not determine where individual eigenvalues leave the profile.The inheritance bounds depend on measurable read-out frame and tail constants, without requiring randomness or independence.
  • Spectral prediction: Profile predictions reproduce held-out spectra and effective-dimension curves across five external model families, outperforming flat-spectrum and rank-shuffled alternatives in every family.The profile was fixed before exact spectra were revealed and generalized across nine models from five external families.
  • Spectral prediction: Probability concentration predicts effective dimension more accurately, whereas read-out geometry predicts spectral decay more accurately across all seven tested families.The equal-family mean log-error ratios favor the probability-only profile for effective dimension and the weighted profile for spectral shape.
  • Singletons and clusters: Weighted K-means achieves mean capture 0.718 and worst-direction capture 0.252 on the fixed language-mode battery, with high-probability tokens concentrated in singleton or near-singleton cells.Twenty-six of 128 cells contain fewer than eight tokens, and those cells carry probability mass 0.3315.
  • Singletons and clusters: The resolved geometry separates a probability-declared singleton head from a support-rich bulk, while fresh-context tests reject a universal fixed partition.The proposed within-cell attenuation failed in all six model-halves, although the fixed-battery anatomy and interchangeability theorem were supported.

Supplementary Note 7

The paper predicts fact acquisition from pretraining corpus statistics and tests evidence depth causally in controlled languages. Deeper evidence delays persistent acquisition across architectures, while held-out loss provides a better cross-scale alignment than training step.

  • Statistical acquisition law: Pretraining unigram, bigram and trigram margins predict held-out fact-acquisition trajectories without refitting, with trajectory R2=0.775–0.792 and acquisition-status agreement 0.792–0.854.The predictor also outperforms constant acquisition-time and label-permutation baselines.
  • Evidence depth: Deepest evidence delays acquisition by a factor 2^Δ≃4.3 in training steps, with Δ=2.103 [1.79, 2.40] across the fixed lower-tail quantile band.The effect is a near-uniform translation in logarithmic acquisition time and concerns persistent signed commitment.
  • Architecture generalization: Evidence depth delays acquisition across GPT-2, GPT-NeoX and Llama, while evidence construction modulates the delay on average.All six architecture-by-construction intervals are positive, whereas fixed-treatment and randomized-label controls contain zero.
  • Cross-scale alignment: Held-out loss aligns cumulative acquisition curves across scale more accurately than training step, with a quantified alignment gap of 0.130 versus 0.413.The real-checkpoint schedule remains observational; causal depth effects come from the randomized synthetic system.
  • Acquisition law: The acquisition law combines corpus composition, evidence depth, cumulative acquisition curves and held-out-loss re-indexing, with per-fact additive fits reaching R2=0.72–0.83.The direct causal result is the near-uniform lower-tail quantile shift.

Supplementary Note 8

This note formalizes weighted margin trajectories and distinguishes exact path accounting from stronger acquisition-time claims. It derives first-passage, scaling, and evidence-depth results under explicit assumptions.

  • Weighted margin projection: A weighted margin projection fixes the fact population, weights, and response-independent coordinates, then tracks its trajectory along parameter paths.The construction supports instantaneous and finite-step accounting identities.
  • Assumptions: Fixed evidence coordinates are essential because checkpoint-dependent resolvers can create apparent projection movement even when margins remain unchanged.Allowing z_f to vary adds a product-rule term to the trajectory identity.
  • Exact path accounting: The endpoint telescope is exact without smoothness, whereas tangent-based identities require retained remainders or stated differentiability assumptions.A zero pooled remainder can hide nonzero fact-level remainders.
  • First-passage ordering: Exact cumulative-change dominance orders ordinary first entry: if the shallow trajectory is never below the deep trajectory through a deep crossing, shallow reaches the threshold by then.A uniform-error version gives the corresponding approximate ordering.
  • Scaling relations: A constant trajectory scaling produces a logarithmic-time translation, but a single quantile-band shift does not establish full distributional scaling.Population scaling need not mean every fact shares the same multiplier.
  • Depth and acquisition: Nominal evidence depth alone cannot universally order acquisition time because learner-relative effective rates can reverse shallow and deep acquisition orders.Randomized evidence-depth experiments identify an empirical causal effect, while projection identities provide conditional path accounting.

Supplementary Note 9

This note links output-margin changes, read-out subspace motion, and language residuals through geometric identities. The results clarify when smooth sensitivity predicts decisions and which updates rotate the resolved subspace.

  • Decision transitions: The discrete margin theorem had zero violations over 45,315 observed cross-seed flips, and LAMBADA flip frequency fell from 0.519 to zero when normalized margin z crossed one.This connects a smooth geometric trajectory to an abrupt decision transition.
  • Objective-rate factorization: On 80 held-out Pythia-70M examples, mean alignment squared is 0.6811 (95% interval 0.6503–0.7117), while sensitivity alone cannot determine objective rate.The factorization held exactly, but the correlation between log sensitivity and log squared objective rate was 0.541.
  • Read-out kinematics: Only the normal component of read-out motion rotates its subspace; updates confined to the current span, including scalar decoupled weight decay, produce no rotation.The result follows from differentiating the orthogonal projector.
  • Residual-driven rotation: The same approximation residual governing static recovery also bounds unresolved-language-driven read-out rotation, while scalar weight decay contributes exactly zero.A controlled smooth-path experiment supported the projector identity and zero-decay result in all 24 held-out seeds.
  • Semantic geometry: Relations showed raw probability-displacement alignment of 0.49 versus near zero for cyclic concepts, while the permuted-geometry control preserved marginals at 0.33.Cyclic concepts instead showed a low metric-isometry defect of 0.29 compared with 1.02 for relations.

Part IV: Controllable

The controllable framework uses Fisher-derived local costs to find minimum-disturbance interventions and compare them with Euclidean control. Experiments show geometric predictions remain useful for finite paths, editing, and reusable multi-objective control.

  • Minimum disturbance: The unique minimum-cost intervention satisfies the behavior constraint through the local metric, and its excess cost over the natural step equals a quadratic off-target disturbance term.The decomposition also extends to a positive-semidefinite metric on its resolved range under stated conditions.
  • Damping: Damping reduces the natural-gradient advantage by shrinking anisotropy: the Euclidean-to-natural cost ratio and its worst-case ceiling decrease toward one.The result applies to B + αI for positive semidefinite B and α > 0.
  • Coordinate covariance: Natural control is chart-covariant, whereas coordinate-Euclidean control depends on the chosen chart and can change when a fresh identity reference is imposed.Reference-metric transport is required to preserve covariance under invertible reparameterizations.
  • Full cubic correction: Adding the network-Hessian term reduced mean absolute error from 0.345 to 0.0176 against a directional finite-difference estimate in the 410M full-curvature analysis.The map-Hessian term vanished at affine final-layer anchors.
  • Finite intervention paths: Across 216 prompts, geometric-mean Euclidean/Fisher path-cost ratios ranged from 11.48 to 127.12, with all 72 paired bootstrap intervals positive in log space.The paths were relinearized after every accepted step under a 0.02-nat prompt-KL trust region.
  • Reference disturbance: Across 24 model–objective–checkpoint settings, the Euclidean/reference-Fisher ratio was 2.8986–6.2882, while donor-Fisher/reference-Fisher was at least 2.129.These comparisons measure mean sequence KL on reference continuations.
  • Reusable control: Reusable within-model activation updates produced positive truth-preference and anti-sycophancy effects on new prompts while lowering reference-sequence KL than corresponding Euclidean compositions.The experiment did not require aligned hidden coordinates between models.

Matrix-free computation

The implementation avoids full Fisher and Jacobian matrices by combining operator-based solves, spectral methods, and low-rank approximations. Relative damping gives a width-independent conjugate-gradient guarantee.

  • Operator implementations: Matrix-free solves apply Jacobian–vector and vector–Jacobian products without materializing H, J, or G, while dense spectral computation is faster at evaluated widths d ≤ 1024.Matrix-free storage grows linearly with activation dimension, whereas dense-operator storage grows quadratically.
  • Conjugate gradients: At relative damping c = 10^-2, the standard conjugate-gradient bound permits at most 85 iterations for Euclidean relative residual 10^-6.The regularized condition number is bounded by (1 + c)/c, independent of width.
  • Low-rank Woodbury: A top-mass support truncation factors the approximate Fisher as B B^T and reduces the damped solve to an |S| × |S| Woodbury system.The construction uses |S| vector–Jacobian products to form the reduced operator.
  • Damping selection: Power iteration estimates the largest eigenvalue so damping can be set relative to the spectrum, typically with c approximately 10^-2.This supports the relative-damping condition used in the convergence bound.
  • Numerical stability: Range restriction removes output-inert null-space components, while severe ill-conditioning motivates float32 solves and makes bfloat16 directions only weakly to moderately aligned in evaluated cases.The reported absolute cosine alignments ranged from 0.056 to 0.542.

Part V: Evidence, translation and reproducibility

The paper validates output-geometry comparisons and shows that Fisher geometry supports interpretable cross-model analysis, predictive calibration, and lower-disturbance interventions. Controls distinguish shared output structure from coordinate-dependent activation effects and clarify where predictive relationships remain unresolved.

  • Knowledge editing: Knowledge-editing comparisons distinguish a fixed routed-edit test from separate query-time benchmarks, preventing direct conflation of their scores.The fixed GPT-2 comparison reuses one solved delta and shared gating, whereas the remaining table sections cover GPT-2-XL query-time protocols with different edit counts and routing conditions.
  • Intervention validation: The Fisher cost predicts intervention disturbance substantially better than activation magnitude and supports matched-effect gains over Euclidean updates.Across 303 interventions, Fisher prediction correlates 0.997 with measured KL versus 0.896 for activation magnitude; matched-effect off-target KL reductions reach factors of 8–74, 30–190, and 4–550 across objectives.
  • Cross-model convergence: Output-distribution geometry outperforms activation geometry across similarity measures and remains invariant under random invertible reparameterisations.Agreement is 0.778 versus 0.710 for mutual k-nearest-neighbour alignment, 0.975 versus 0.935 for centred-kernel alignment, and 0.935 versus 0.755 for distance-matrix agreement; activation measures degrade as condition number increases while output agreement remains constant.
  • Cross-model convergence: The shared output geometry is driven mainly by displacement magnitudes, with only a small residual eigenvector-overlap component after magnitude-matched null correction.Top-six overlap is 0.698 against a shuffled null of 0.025, but the magnitude-matched null is 0.636, leaving residual overlap of +0.063.
  • Semantic structure: Output geometry separates consensus semantic structure from model-specific residuals, retaining semantic alignment after surface-text partialling.On the semantic battery, consensus alignment is 0.440 and remains +0.231 after partialling out surface geometry, while residual alignment is negligible.
  • Human alignment: Human semantic agreement improves with model scale, while a risk-based relation predicts agreement across architectures without refitting.Mean squared model–human semantic distance declines from 0.33885 at 70M to 0.29576 at 2.8B; the transferred relation correlates at 0.758320 with RMSE 0.029465.

Supplementary Note 20

The supplementary material documents the evidence map and computational reproducibility conditions. Experiments use multiple public model families, controlled synthetic systems for causal acquisition tests, and specified hardware and numerical-precision settings.

  • Evidence map: Supplementary Table 4 maps each central claim strand to its evidence and role in the manuscript.The table organises the paper’s analytical, observational, and interventional evidence chain.
  • Reproducibility: The computational analyses use public checkpoints, controlled synthetic systems, float32 execution by default, and specified GPU, CPU, and parameter-scale limits.Most GPU runs use one RTX 5000 Ada; full GPU-resident pullback solves reach 2.8B parameters, while 6.9B analyses use graph-free spectra and CPU-based calibration or offloading.
Loading 2609.11063v1…