Source-linked AI summary
Learning reshapes power-law anisotropy in internal representations
Asahi Nakamuta, Jun-nosuke Teramae
TL;DR
Power-law anisotropy in neural representations changes during learning, but its formation from input and task structure is unclear. The paper exactly analyzes a wide two-layer linear network and finds that feature learning reshapes local exponents across modes and time, unlike lazy learning.
Problem
Existing theories often treat internal-representation spectra as fixed, leaving unclear how learning reshapes input power laws and whether exponents vary across modes.
Method
The authors exactly solve population-gradient-flow dynamics for a wide two-layer linear teacher–student network with aligned power-law input covariance and teacher structure.
Results
Feature learning produces up to four asymptotic local power-law exponents across modes and training times, whereas lazy learning leaves the exponent nearly unchanged.
Takeaways & Limitations
Power-law internal representations are reshaped over time and across modes through interactions between input statistics and task learning.
Takeaways & Limitations
The analysis assumes simultaneously diagonalizable input and teacher structure, infinite width, and population-limit gradient flow.
Abstract
from arXiv · showhide
Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models to the mouse cerebral cortex. This anisotropy is a key geometric property of high-dimensional information processing and underlies a variety of theoretical analyses. However, the mechanism by which it emerges from input structure and task-driven learning has remained unclear. Here, we characterize this formation process by exactly solving the learning dynamics of a wide two-layer linear neural network in a teacher--student setting with power-law input and teacher structures. We show that, in the feature-learning regime, the local power-law exponent of the internal-representation spectrum evolves nonmonotonically over the course of training and exhibits up to four distinct asymptotic regimes across modes and training times. By contrast, in the lazy regime, the exponent remains essentially unchanged. We further demonstrate numerically that similar exponent dynamics arise in more realistic nonlinear networks. Together, these results suggest a general mechanism by which the dynamic interaction between input statistics and task structure gives rise to power-law internal representations.
1 Introduction
Power-law anisotropy is widely observed in neural internal representations, but existing theories largely treat their spectra as fixed rather than explaining how learning forms and reshapes them. This study analyzes that process using a wide two-layer linear teacher–student model and shows that input structure, teacher anisotropy, and mode-dependent learning jointly shape representations.
- Motivation: Internal-representation covariance eigenspectra are reported to decay as power laws across biological and artificial neural systems.Examples include mouse visual cortex, computer-vision models, and large language models.
- Motivation: Assuming a fixed power-law exponent cannot explain whether internal representations inherit input structure or are reshaped by task learning.The input covariance of natural data can itself exhibit power-law decay, making this distinction necessary.
- Approach: The study analyzes hidden-layer representation dynamics in a wide two-layer linear teacher–student network with aligned power-law input covariance and teacher-map structures.Both layers are jointly trained under population gradient flow, inducing nonlinear learning dynamics despite a linear input–output map.
- Main findings: Power-law internal representations are dynamically reshaped by teacher-signal anisotropy and mode-dependent learning timescales, rather than merely reflecting input statistics.The results provide an analytical starting point for explaining observed representations through interaction between input structure and learning.
Related Works
Prior work has documented power-law spectra in biological neural populations, natural-image covariances, and learned representations, while showing that input covariance alone does not fully explain internal spectra. Deep linear networks provide analytically tractable models for studying nonlinear gradient dynamics and evolving hidden representations.
- Power-law spectra of internal representations: Power-law covariance eigenspectra have been observed across biological neural populations, including mouse V1, with eigenvalues decaying approximately as k^-1.Related scaling has been reported across species and brain regions.
- Power-law spectra of internal representations: A broken power law fits mouse V1 data better than a single power law, supporting a local, mode-dependent exponent description.
- Input structure and spectral theories of fixed representations: Power-law covariance spectra occur in natural images, but persistent power-law spectra in V1 responses after spatial whitening show that input covariance alone is insufficient.
- Learning dynamics of deep linear networks: Deep linear networks offer analytically tractable feature-learning models because factorized-weight gradient dynamics are nonlinear despite linear input–output maps.Exact analyses have derived aligned-mode learning curves with plateaus and stage-like learning, as well as temporal evolution of hidden representations and neural tangent kernels.
2 Setting
The paper studies a mean-field, two-layer linear teacher–student network with infinite hidden width, focusing on the simplest equal input–output dimension setting, P = D. Inputs and teacher structure are aligned, and the network is trained by gradient flow to analyze internal-representation spectra over time.
- Model setting: The model is a mean-field two-layer linear network with input dimension D, hidden width N, and output dimension P in a teacher–student setting.The teacher is represented by a matrix Θ, with inputs and hidden weights in R^D and outputs in R^P.
- Model setting: The analysis takes the hidden-layer width to infinity and sets the equal finite input–output dimensions to P = D.The choice P ≥ D is motivated by large language models, diffusion models, and autoencoders, and is required for spectral evolution across a broad range of modes.
- Structural assumptions: Inputs and the teacher matrix have aligned structure: the input covariance and teacher Gram matrix are diagonal in the same orthonormal basis.The input distribution is Gaussian with covariance Λ = diag(λ1, . . . , λD).
- Training and analysis: Weights are independently Gaussian-initialized and trained by gradient flow on mean-squared error with learning rate ν.The paper derives the temporal evolution of internal-representation covariance eigenvalues in the infinite-data limit and characterizes their power-law behavior.
3 Results
Learning adds a teacher-induced component to the input-shaped representation spectrum, producing up to four asymptotic local power-law exponents and nonmonotonic exponent evolution in the feature-learning regime. In the lazy regime, the exponent changes little, while nonlinear networks show the same α-to-α + β evolution predicted by the linear theory.
- Spectrum evolution: The spectrum combines an initialization component inherited from the input with a teacher-induced component shaped by learning dynamics.Thus, representation anisotropy depends on both input and teacher structure.
- Spectrum evolution: At initialization, ρk(0) has the input exponent α; during training, teacher-derived structure is acquired from low to high mode index.The analytical eigenspectrum dynamics closely match finite-width numerical results.
- Asymptotic regimes: Up to four asymptotic exponents coexist across bulk regions of the (k, t) plane, with non-power-law interpolation near domain boundaries.Near boundaries, the effective exponent varies continuously between adjacent asymptotic values.
- Asymptotic regimes: The local exponent first steepens and then decreases because mode-dependent learning rates incorporate teacher anisotropy into the representation.This nonmonotonic evolution is specific to feature learning, where the representation changes substantially from initialization.
- Nonlinear networks: In nonlinear networks, the feature-learning spectrum evolves from α to α + β over a finite mode range, whereas the lazy-regime exponent remains at α.These results agree quantitatively with the linear-theory predictions.
4 Conclusion · Appendix · A The anisotropy of the internal representation corresponds to the eigenvalues of M(t)
The analysis shows that power-law anisotropy is dynamically reshaped across modes and training through interactions between input statistics and task learning, rather than governed by a single fixed exponent. The appendix further characterizes this anisotropy through the eigenvalues of M(t), which coincide with the nonzero eigenvalues of the internal-representation operator.
- 4 Conclusion: The hidden-representation spectrum decomposes into a baseline component inherited from input covariance at initialization and a learning-dependent component.
- 4 Conclusion: In the feature-learning regime, the local effective power-law exponent can take four asymptotic values over suitable finite mode ranges away from boundaries.
- 4 Conclusion: These regimes reflect initialization dominance, early mode-dependent changes, post-learning first-layer deformation, and deformation distributed across both layers.
- 4 Conclusion: Early learning transiently steepens the spectrum through mode-dependent learning rates set by input eigenvalues before later exponent transitions.
- 4 Conclusion: Nonlinear-network experiments preserved the predicted exponent transitions and the distinction between feature-learning and lazy regimes.
- 4 Conclusion: The analysis assumes simultaneously diagonalizable input covariance and teacher matrices, infinite width, and population-limit gradient flow.
- A The anisotropy of the internal representation corresponds to the eigenvalues of M(t): In the population limit, the internal-representation kernel’s nonzero eigenvalues are obtained from the matrix M(t).The kernel eigenvalue problem becomes an integral-operator problem, whose nonzero eigenvalues coincide with those of M(t).
B Warm-up: Derivation of the eigenvalues when the second layer is fixed
With the second-layer weights fixed, training only the first layer yields an analytical time evolution for M(t). Assuming ΘΘ⊤ is diagonal, M(t) remains diagonal and its eigenvalue dynamics decouple across modes.
- Derivation: Fixing the second layer reduces the warm-up analysis to first-layer training and provides an analytical solution for M(t).The dynamics are studied under gradient-flow training in the infinite-data, population-limit setting.
- Eigenvalue dynamics: Because ΘΘ⊤ is diagonal, M(t) remains diagonal throughout training, so its eigenvalues ρ_k evolve independently across k.The resulting eigenvalue-spectrum dynamics are illustrated in Fig. 5.
- Derivation: The learning rate is set to ν = N so that parameter dynamics remain O(1) as N →∞.This scaling is part of the infinite-width training setup.
C Warm-up: Derivation of the power-law exponent boundaries when the second layer is fixed … D Analytical eigenvalue solution when both layers are trained
The fixed-second-layer analysis derives the local exponent’s time evolution and partitions the (t, k) plane into learned and unlearned regimes with distinct asymptotic exponents. When both layers train, diagonal Gram-matrix dynamics decouple across modes and yield an analytical eigenvalue solution through conserved quantities and variable transformations.
- C Warm-up: Derivation of the power-law exponent boundaries when the second layer is fixed: The local power-law exponent αeff(k, t) is obtained analytically from the time-dependent spectrum ρk(t) and the derivative of Q(k, t).This expression provides the analytical time evolution used to derive boundaries in the (t, k) plane.
- C Warm-up: Derivation of the power-law exponent boundaries when the second layer is fixed: The learning boundary is klearn(t) = (cλ0t)1/α, with modes unlearned for k ≫ klearn(t) and learned for k ≪ klearn(t).The boundary separates modes according to whether zk(t) is much smaller or larger than 1.
- C.1 Learned modes: For learned modes, input-induced anisotropy and learning-induced anisotropy combine to produce exponent α + β at finite k, while sufficiently large k approaches α.The learning-induced term decays more rapidly with k, so the input exponent α eventually dominates.
- C.2 Unlearned modes: For unlearned modes, finite-k spectra can decay with exponent 3α + β, whereas sufficiently large k necessarily has asymptotic exponent α.The faster-decaying second term explains the eventual large-k exponent.
- C.2 Unlearned modes: Across the three asymptotic regions, the exponent is α for k > kcross(t), 3α + β for klearn(t) < k < kcross(t), and α + β otherwise.These regions summarize the learned and unlearned mode behavior and are illustrated in Fig. 6.
- D Analytical eigenvalue solution when both layers are trained: When both layers train, diagonal Λ, ΘΘ⊤, Gwa, Gw, and Ga remain diagonal because their time derivatives are diagonal from diagonal initial conditions.This invariance reduces the dynamics to independent diagonal entries.
- D Analytical eigenvalue solution when both layers are trained: Two conserved quantities eliminate two variables, allowing ηk(t) to be expressed as a function of ξk(t) and yielding the sequence ξk(t) →ηk(t) →ρk(t).The conserved quantities reduce the coupled mode dynamics before solving for ξk(t).
- D Analytical eigenvalue solution when both layers are trained: A change of variables, separation of variables, partial fractions, and inversion produce the analytical eigenvalue solution ρk(t) = λkηk(t).The full analytical solution is not required to evaluate the power-law exponents.
E Derivation of the power-law exponent boundaries when both layers are trained … E.3 Boundary tlayer(k) governing the allocation of learning
The derivation identifies three power-law exponent boundaries from relations governing mode learning, dominance of learning over initialization, and allocation between layers. These boundaries are defined through ξ_k(t), η_k(t), and u_k(t), with larger-k behavior causing divergence for the latter two.
- E Derivation of the power-law exponent boundaries when both layers are trained: Three relations suffice to derive the power-law exponents when both network layers are trained.The derivation uses the stated relations as its basis.
- E.1 Learning boundary t_learn(k): Because ξ_k(t) rises monotonically from 0 to γ_0θ_k, t_learn(k) is defined when ξ_k(t) reaches 0.9γ_0θ_k.This marks the boundary between unlearned and learned modes.
- E.1 Learning boundary t_learn(k): k_learn(t) is defined as the inverse of t_learn(k).The inverse boundary expresses the learned-mode cutoff as a function of training time.
- E.2 Boundary t_cross(k) at which learning becomes dominant: t_cross(k) is obtained by equating the learning and input contributions in η_k(t).It marks when the second term overtakes the first term already present at initialization.
- E.2 Boundary t_cross(k) at which learning becomes dominant: For larger k, t_cross(k) diverges.Thus, sufficiently high-index modes require increasingly long times for learning to become dominant.
- E.3 Boundary t_layer(k) governing the allocation of learning: The boundary t_layer(k) is defined by u_k(t) = 1, equivalently by ξ_k(t) = S.This boundary governs how learning is allocated between the two layers.
- E.3 Boundary t_layer(k) governing the allocation of learning: For larger k, t_layer(k) diverges, and its condition can equivalently be expressed as a constraint on k.The appendix states the equivalent k-condition without providing it in the supplied passage.
E.4 Conditions for the existence of the asymptotic regions
The asymptotic regions have distinct existence conditions determined by initialization scales and γ0. When the feature-learning conditions fail, the entire spectrum lies in Region I and its power-law exponent remains nearly unchanged, defining the lazy regime.
- Region I: Region I always exists because R_k(0) = 0 < 1 at initialization.This condition holds at t = 0 for every mode.
- Region II: Region II exists over a wide range when the intersection of t_cross(k) and t_learn(k) occurs at a k-coordinate much larger than 1.Its upper boundary is set by that intersection.
- Region IIIa: Region IIIa requires the second-layer initialization scale to be sufficiently larger than the first-layer scale.Its broad-range condition is determined by the upper and lower boundaries k_u and k_l.
- All regions: The four regions coexist only when the combined conditions for Regions II, IIIa, and IIIb are satisfied.The paper states these conditions collectively without reproducing their equations in the supplied passage.
- Feature-learning and lazy regimes: Feature learning generally corresponds to a larger second-layer initialization scale than first-layer scale and large γ0, whereas the lazy regime places all modes in Region I.In the lazy regime, the spectrum’s power-law exponent remains nearly unchanged.
F Proof that the ordering of the eigenvalues ρk(t) is preserved
The proof shows that eigenvalue ordering is preserved throughout training: for k > l, ρ_k(t) ≤ ρ_l(t) at every time t. Thus, the input-dimension index k remains aligned with eigenvalue rank.
- Proof of ordering preservation: For k > l, ρ_k(t) ≤ ρ_l(t) for every t, so eigenvalue ordering is preserved throughout training.This establishes that eigenvalues cannot cross during training.
- Proof of ordering preservation: The scalar dynamics ξ(t; λ, θ) are monotonically increasing in time, input eigenvalue λ, and teacher eigenvalue θ.The proof derives nonnegative time, λ, and θ derivatives for ξ.
- Proof of ordering preservation: Because λ_k and θ_k are ordered descending, k > l implies ξ_k(t) ≤ ξ_l(t).This follows by applying monotonicity first in λ and then in θ.
- Proof of ordering preservation: Since ρ_k increases monotonically with ξ_k and λ, the index k coincides with eigenvalue rank.The ordering of the internal eigenvalues therefore remains tied to input dimension.
G Numerical settings
The numerical experiments used power-law input and teacher spectra with matched exponents, fixed initialization, and dimensions (D, P, N, M) = (512, 512, 8192, 4096).
- G Numerical settings: For all plots, the input and teacher spectra followed λ_k = k−1 and µ_k = ζ(2)−1k−1, with α = β = λ_0 = 1 and µ_0 = 1/ζ(2).Inputs and targets were generated as x ∼N(0, Λ) and y = Θ⊤x, with diagonal Λ and Θ matrices.
- G Numerical settings: The second-layer initialization scale was fixed at σ_a = 1.0.
- G Numerical settings: The numerical settings were (D, P, N, M) = (512, 512, 8192, 4096).