Source-linked AI summary
Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens
Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling
TL;DR
The paper addresses the limited theoretical account of how the J-lens reads verbalizable information from intermediate language-model states. It formalizes the J-lens as an averaged first-order transfer operator, analyzes its approximation and bias, and studies its sparse Jacobian-energy structure. The analysis identifies short-horizon and sparse-concept readouts and motivates filtering and decoupling methods that improve the practical lens.
Problem
The J-lens lacks a detailed mathematical interpretation of its causal readout, approximation behavior, bias, and sparse positional structure.
Method
The paper combines local Jacobian linearization, global least-squares projection, a Stein-based identity, and Jacobian-energy analysis to study the J-lens.
Results
The J-lens is characterized as an expectation over anticipated future outputs whose energy concentrates in diagonal short horizons and sparse concept positions.
Takeaways & Limitations
Energy filtering and decoupling methods further validate the interpretation and improve the J-lens's ability to read out correct intermediate concepts.
Takeaways & Limitations
The Stein bridge requires a smooth positive input density, differentiability, controlled growth, and vanishing boundary terms.
Abstract
from arXiv · showhide
The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal structure. Besides treating the J-lens as a heuristic probe, we further regard it as a first-order causal transfer operator from intermediate activations to expected future readouts. We study the Jacobian matrix as the optimal local linear approximation of the downstream mapping, analyze its global approximation behavior and bias, and identify its mathematical meaning as an expectation over anticipated future readouts. Further analysis of the Jacobian energy distribution reveals that its causal geometry is highly sparse. The energy decays with depth, concentrates in an extremely small proportion, and decomposes into diagonal pathways and specific critical positions. This decomposition further resolves the expectation of the J-lens over future outputs into short-horizon and sparse concept predictions, providing a more intuitive attribution and explanation for the ability of the J-lens to visualize concepts during the thinking process. Based on the theory, we propose a simple but effective improvement strategy and decoupling method for the J-lens, which significantly enhances the ability of the J-lens to read out correct intermediate concepts.
1 INTRODUCTION
The paper gives the J-lens a mathematical and causal interpretation, then analyzes sparse Jacobian-energy structure to explain its short-horizon and sparse-concept readouts. It also proposes filtering and decoupling methods to improve intermediate-concept readout.
- Mathematical interpretation: The J-lens is interpreted as an averaged first-order causal transfer operator from intermediate activations to expected future readouts.The paper connects local linearization and global fitting to the downstream latent mapping.
- Sparse causal geometry: Jacobian-energy analysis finds larger energy in earlier layers, concentration in few source-target pairs, and diagonal or critical-position patterns.These patterns correspond to short-horizon token prediction and highly sensitive tokens at critical positions.
- Sparse causal geometry: The paper names the two resulting readout modes short horizons and sparse concepts, resolving anticipated future outputs into these positional modes.This provides an attribution-oriented interpretation of J-lens readouts.
- Improving the J-lens: A top-energy position filtering strategy and a method for decoupling the two readout modes are proposed to improve the practical J-lens.The filtering strategy uses only a subset of position pairs with the highest Jacobian energy.
2 BACKGROUND: J-LENS AND GLOBAL WORKSPACE
The J-lens reads intermediate activations through their average first-order effects on present and future outputs, producing verbalizable directions in a sparse J-space. The paper frames this as a response to the limited direct interpretability of intermediate residual-stream states.
- Background: Intermediate residual-stream states are computational high-dimensional vectors that are not directly verbal, while final states map to vocabulary scores through the unembedding matrix.The logit lens often fails early because representations have not yet rotated into output coordinates.
- J-lens: The J-lens averages Jacobians from intermediate activations to later final-layer states across source positions, target positions, and prompts.It therefore reads an activation by its average first-order effect on present and future outputs.
- J-space: J-lens readout tokens describe what an activation is disposed to say across contexts, and sparse nonnegative combinations of token-associated directions form the J-space.The rows of WU Jℓ define the token-associated readout directions.
- Open problem: The paper identifies a missing theoretical account of why averaged Jacobians approximate optimal readouts and how nonlinear and non-Gaussian effects influence that approximation.It also investigates the sparse causal structure over positions underlying the interpretation.
3 MATHEMATICAL EXPLANATION OF J-LENS: EXPECTATION OF THE FUTURE READOUTS
The J-lens averages local Jacobians to approximate the expected future readout of an intermediate activation, linking local sensitivity to a global affine readout. Its accuracy is limited by nonlinear and non-Gaussian bias, which grows with source–target distance and helps explain early-layer failure.
- Future-readout interpretation: The J-lens maps an intermediate hidden state to averaged future final-layer responses through the generally nonlinear function f_l.The target averages final-layer responses at positions t′≥t.
- Local approximation: A Jacobian gives the optimal local first-order sensitivity map, but its usable approximation radius shrinks with curvature and it remains pointwise rather than global.For a small intervention δ, the predicted response change is J_f_l(h_l,t)·δ.
- Global least-squares view: The global affine readout is the L2 projection of the regression function onto affine functions, requiring neither Gaussianity nor differentiability.Its slope is a statistical compromise over the data distribution.
- Stein bridge: Under Gaussian inputs, the averaged Jacobian equals the population least-squares slope through the Stein bridge; non-Gaussian inputs introduce systematic bias.The equality depends on the stated regularity and growth conditions.
- Bias analysis: The J-lens combines first-order linearization with averaging local Jacobians, so nonlinear and non-Gaussian biases amplify as source–target distance increases.The nonlinear residual cannot be removed by any linear slope, while distributional bias reflects deviation from Gaussianity.
- Empirical validation: Empirical layer-wise comparisons show nonlinear and non-Gaussian errors decrease near the target layer, while large early-layer biases can move readouts outside the confidence interval.The paper uses this agreement between theoretical and empirical error to validate its interpretation.
4 JACOBIAN ENERGY STRUCTURE IN J-LENS: SHORT HORIZON AND SPARSE CONCEPTS
The J-lens’s averaged Jacobian hides a highly concentrated positional structure: energy decays with depth and separates into short-horizon diagonal pathways and sparse concept positions. This decomposition interprets future-output averaging as two positional modes.
- 4.2 LAYER-WISE DECAY AND CONCENTRATION: The energy analysis identifies which source-target pairs influence the J-lens most strongly, rather than treating all positions as equally informative.Higher-energy positions contribute more strongly to the J-lens readout.
- 4.2 LAYER-WISE DECAY AND CONCENTRATION: Jacobian energy decays with depth and is heavy-tailed across position pairs, so a small minority of pairs dominates the averaged operator.Early layers have larger influences, while concentration increases with depth.
- 4.3 JACOBIAN STRUCTURE: SHORT HORIZONS AND SPARSE CONCEPT: High-energy pairs form diagonal pathways for short-term token expressions and horizontal or vertical patterns associated with sparse concept positions.Diagonal energy is especially prominent in later layers; broadcast origins and integration points define the other pattern.
- 4.3 JACOBIAN STRUCTURE: SHORT HORIZONS AND SPARSE CONCEPT: Stable high-energy positions attached to interpretable tokens are grouped as sparse concept positions, while diagonal and adjacent positions define short horizons.These two patterns reduce the aggregate future-output readout to short horizons and sparse concepts.
- 4.3 JACOBIAN STRUCTURE: SHORT HORIZONS AND SPARSE CONCEPT: The formalized J-lens readout is approximated by separating the aggregate operator into short-horizon and sparse-concept positional modes.The approximation uses the positions denoted by t* for sparse concepts.
5 EXPERIMENTS
The experiments evaluate energy-based filtering and mode decoupling on Qwen3-8B across six tasks using defined metrics for next-token prediction and intermediate-concept recall. Filtering generally improves over the vanilla J-lens, while diagonal removal reveals coupling between the two readout modes.
- 5.1 EXPERIMENTS SETTINGS: Qwen3-8B is evaluated on association, multihop reasoning, multilingual, ordering operations, poetry, and typos tasks using Wikitext for J-lens construction.All J-lenses use the same 200 training examples and configuration.
- 5.2 METRICS: The short-horizon layers metric measures how early lens readout converges to the model’s next-token prediction through trailing layer agreement.For each position, it counts the trailing layers whose top-1 prediction matches the true next token.
- 5.2 METRICS: Intermediate-concept recall rate measures how broadly top-10 lens predictions contain annotated intermediate-concept tokens across the position-layer readout.The metric is computed on prompts annotated with concept tokens.
- 5.3 RESULTS: Top-j energy filtering strengthens concentration and outperforms the vanilla J-lens on most tasks by focusing readout on highlighted sparse concepts.The filtering strategy uses only the highest-energy fraction of position pairs.
- 5.3 RESULTS: Diagonal filtering weakens next-token prediction without eliminating it and also reduces sparse-concept performance, indicating deeper coupling between the two patterns.The result points to decoupling as an area for future work.
6 CONCLUSION
The paper interprets the J-lens as an averaged first-order transfer operator and explains its readout through sparse Jacobian-energy structure. Energy filtering preserves or improves next-token consistency and intermediate-concept recall, while the authors acknowledge that the work remains in progress.
- 6 CONCLUSION: The J-lens is interpreted as an averaged first-order transfer operator from intermediate activations to expected future readouts.Local linearization, global least squares, and a Stein bridge connect the operator to optimal linear readout under Gaussian inputs.
- 6 CONCLUSION: Nonlinear curvature and non-Gaussian score residuals are identified as sources of systematic bias and early J-lens failure.These deviations explain why the approximation can be inaccurate in relevant regimes.
- 6 CONCLUSION: Jacobian energy decays with depth, concentrates in few source-target pairs, and separates into diagonal short horizons and sparse concept positions.These structures provide the paper’s interpretation of J-lens readout.
- 6 CONCLUSION: Energy filtering preserves or improves next-token consistency and intermediate-concept recall, supporting the proposed interpretation of J-lens readout.The conclusion presents these improvements as validation of the interpretation.
- 6 CONCLUSION: The authors acknowledge that the work has limitations and remains in progress rather than representing final results.The conclusion does not specify a narrower limitation.