Source-linked AI summary
Dimension-Corrected Hitting Times for Heavy-Tailed Spectral Emergence in Neural Optimizer Dynamics
Zongmin Liu
TL;DR
The paper addresses the poorly understood step complexity of heavy-tail emergence by formulating it as a right-censored hitting-time problem. It finds a reproducible dimension-corrected spectral-gap law and develops conditional theory clarifying why Adam recurrences alone do not establish redistribution. The evidence is strongest in controlled teacher–student dynamics and does not claim a first-principles Adam theory or optimizer universality.
Problem
The step complexity of neural-network spectral heavy-tail emergence is poorly understood, so final-time tail exponents do not adequately characterize onset dynamics.
Method
The paper models heavy-tail emergence as a right-censored hitting time and combines finite-onset regression, censored survival analysis, and conditional redistribution theory.
Results
The dimension-corrected model is favored over gap-only alternatives, with R^2 = 0.683, γ = 0.626, ρ = 0.772, and AIC improving from 706.62 to 628.70.
Takeaways & Limitations
The results provide a quantitative spectral hitting-time framework while limiting claims about first-principles Adam redistribution and optimizer universality.
Takeaways & Limitations
The main evidence is controlled synthetic teacher–student dynamics, and the full derivation of redistribution from teacher–student Adam covariance is not claimed.
Abstract
from arXiv · showhide
Heavy-tailed empirical spectral densities of neural-network weight matrices are widely used as diagnostics of implicit self-regularization, but the step complexity of heavy-tail emergence remains poorly understood. We formulate spectral heavy-tail formation as a right-censored hitting-time problem: a run that does not reach a heavy-tail diagnostic within the observation horizon is treated as censored rather than discarded. In controlled full-batch teacher--student dynamics, we find that the first-step spike--bulk gap alone does not explain onset time. Instead, finite-onset regression supports a dimension-corrected spectral-gap law, (τ_{\mathrm{HT}}\approx CΔ_1^{-γ}d^ρ), with (R^2=0.683), (γ=0.626), and (ρ=0.772) across 330 completed runs. Right-censored lognormal accelerated-failure-time models further favor the dimension-corrected model over a gap-only model, improving AIC from 706.62 to 628.70. Theoretically, we prove that exact early loss dynamics in linear networks do not determine factor spectral tails, that Adam recurrences alone do not imply spectral redistribution, and that projected singular-basis spreading implies contraction of a spectral-tail potential and hence a dimension-corrected hitting-time bound. Empirically, projected-kernel profiles support the sufficient spreading mechanism, Adam and AdamW agree under tested grids, GD and signGD do not reach onset in the same regimes, and real pretrained Qwen2.5-0.5B and Pythia-70M transformer weights show non-Gaussian spectral-tail structure relative to matched Gaussian nulls. The result is a reproducible spectral hitting-time law with rigorous conditional theory, not a claim that Adam necessarily generates heavy tails from first principles.
1 Introduction
The paper reframes heavy-tail emergence as a right-censored hitting-time problem and finds that onset is better explained by a dimension-corrected spectral-gap law than by the first-step gap alone. It combines empirical timing models with conditional theory separating loss dynamics, optimizer recurrences, and spectral redistribution.
- Main finding: The first-step spike–bulk gap alone is insufficient to explain onset, motivating a dimension-corrected spectral-gap coordinate.Finite-onset regression and censored likelihood comparisons favor including dimension.
- Hitting-time formulation: Heavy-tail emergence is defined by its first diagnostic hit, while runs without a hit by the horizon contribute right-censored observations rather than being discarded.This distinguishes onset dynamics from final-time tail estimates.
- Censored analysis: A censored lognormal accelerated-failure-time model including log d improves AIC from 706.62 to 628.70 relative to a gap-only model.Right-censoring preserves information from runs that do not hit the diagnostic within the observation horizon.
- Conditional theorem and mechanism evidence: Adam recurrences alone do not imply spectral redistribution, so the theory proves a conditional spreading-to-hitting implication and tests projected-kernel evidence empirically.Rank-one locked gradients provide the no-free-redistribution boundary.
- Relation to exact early-dynamics theory: Exact early loss dynamics in linear networks do not identify spectral-tail hitting because loss depends on the realized product map while factor spectra depend on the factorization.The paper therefore treats predictive-loss dynamics as complementary rather than determinative for spectral emergence.
2 Problem Setup
The setup uses a two-layer teacher–student model with Gaussian inputs and fixed readout, updating the first-layer weights by full-batch optimization. Heavy-tail onset is diagnosed from multiple spectral statistics and treated as right-censored when no checkpoint qualifies by the horizon.
- Teacher–student model: The teacher–student setup samples Gaussian inputs, chooses a normalized teacher direction, and uses a two-layer student network.The main dynamics update W_t while keeping the readout fixed.
- Spectral initialization: The covariance spectrum is compared with the Marchenko–Pastur upper edge to define the first-step spike–bulk gap Δ_1.The gap is the positive excess of the top eigenvalue over λ+.
- Heavy-tail diagnostic: The primary onset requires Hill index 1.4–3.0, KS score at most 0.35, top-tail mass at most 0.85, and increased effective rank.These criteria are evaluated at checkpoints.
- Observation rule: Runs that satisfy no onset criterion by the maximum horizon T are recorded as right-censored with τ_HT > T.Censoring retains information about delayed or unobserved onset.
3 Theory: From Adam Spikes to Conditional Hitting Times
The theory separates spectral-tail hitting from exact loss dynamics and shows that Adam recurrences alone do not guarantee redistribution. Under a measurable projected spreading condition, a spectral-tail potential contracts and yields a conditional dimension-corrected hitting-time bound.
- Theory boundary: Exact two-step loss dynamics determine residual maps and layer Gram matrices but not the empirical spectral tail of an individual factor.Rescaling factors by S and S^-1 preserves the product map and initial loss while changing a factor’s spectrum arbitrarily.
- Theory boundary: Adam recurrences alone do not imply spectral redistribution because full-batch gradient sequences can keep updates rank-one locked.The paper therefore treats singular-basis spreading as an additional mathematical condition rather than an automatic consequence of Adam.
- Adam mechanism: A first-step sign-Adam spike can produce a large outlier, but spike formation is distinct from subsequent spike-to-bulk redistribution.The leading outlier is lower bounded by a term of order η^2c^2t_h up to cross and bulk terms.
- Conditional mechanism: The projected-kernel profile measures singular-basis injection, while the spectral-tail potential combines deviation from a power-law envelope with a penalty against isolated outliers.The construction fixes a tail window, exponent interval, and spike threshold before testing contraction.
- Conditional mechanism: Under projected spreading and bounded cross-term perturbations, the theorem establishes contraction of the spectral-tail potential and a dimension-corrected hitting-time law.The result is conditional by design; deriving the spreading assumption directly from teacher–student Adam covariance is explicitly left open.
- Censored inference: Right-censored likelihoods retain non-onset trajectories as survival information rather than discarding them, avoiding conditioning on only early events.Discarding non-onset runs biases comparisons toward high-gap, fast-onset regimes.
4 Experiments
The experiments evaluate controlled optimizer sweeps and robustness extensions for finite-onset and right-censored spectral hitting times. The reported figures show that dimension correction organizes onset better than gap alone and remains preferred when non-onset runs are censored.
- Experimental design: The robustness program included 255 original-design runs and 330 runs including extensions across dimensions, learning-rate regimes, seeds, optimizers, and projected-kernel profiles.The extensions test boundary behavior and diagnostic stability rather than estimating the scaling exponents.
- Right-censored analysis: The censored dimension-corrected model improves AIC from 706.62 to 628.70 relative to the gap-only model.Low-gap trajectories often remain below the consensus criterion by the observation horizon and are treated as right-censored rather than omitted.
5 Results
The results favor a dimension-corrected onset model over a gap-only explanation, while mechanism tests support projected spreading and guard against broad optimizer or generalization claims.
- 5 Results: The dimension-corrected model is preferred over a gap-only model in both finite-onset regression and right-censored comparisons.Low-gap trajectories often remain censored at the observation horizon, and excluding high learning rates preserves the qualitative conclusion.
- 5 Results: Projected-kernel profiles outperform random-basis baselines and show increasing fitted mixing weight before onset, supporting projected singular-basis spreading.Raw projected correlation remains competitive, so the evidence does not establish uniqueness of the arcsine transform.
- 5 Results: Adam and weak-decay AdamW agree closely, whereas GD and signGD do not meet the primary onset criterion under the tested grids.The comparison supports only partial optimizer universality, even when GD and signGD produce nonzero gaps at high step sizes.
- 5 Results: Ridge-readout diagnostics show non-monotone test MSE across Hill-alpha bins, so heavier tails are not claimed to improve generalization universally.The main contribution remains the spectral-dynamical hitting-time law rather than a generalization guarantee.
6 Limitations
The study’s strongest theoretical gap is that the first-principles derivation of projected spectral spreading from teacher–student Adam covariance remains open.
- 6 Limitations: A full first-principles theorem deriving projected spectral spreading from teacher–student Adam gradient covariance remains open.The experiments are controlled synthetic teacher–student dynamics, with only an appendix sanity check beyond the main setting.
7 Conclusion
The paper reframes heavy-tail emergence as a censored hitting-time phenomenon and finds a dimension-corrected spectral-gap law in controlled full-batch Adam dynamics.
- 7 Conclusion: Heavy-tail emergence is best viewed as a censored hitting-time phenomenon governed by a dimension-corrected spectral-gap law rather than the first-step gap alone.The law is supported by finite-onset regression, censored survival modeling, eta-bridge robustness, and projected-kernel mechanism tests.
A.1 Proof of Proposition 1
The proofs establish boundaries on what early loss dynamics and Adam recurrences identify, then derive tail-potential contraction and censored hitting-time modeling under explicit spreading and residual controls.
- A.1 Proof of Proposition 1: Exact product-preserving reparameterizations can arbitrarily change factor conditioning while keeping predictive loss fixed, so early loss dynamics cannot identify factor spectral tails.For invertible S, the transformed factors satisfy eB eA = BA while the singular values of eA can become arbitrarily ill-conditioned.
- A.1 Proof of Proposition 1: Right-censored likelihoods retain non-onset runs through survival terms, whereas dropping them overweights high-gap trajectories and changes the target estimand.Observed hits contribute densities, while runs with τ_i > T_i contribute survival probabilities.
- A.1 Proof of Proposition 1: Under rank-one locked gradients, Adam moments remain separable and the limiting update is rank one, so Adam recurrences alone do not imply redistribution across singular directions.The argument motivates a separate nondegenerate spreading condition.
- A.1 Proof of Proposition 1: Projected kernels measure injected energy along current singular directions, with the nontrivial quantity being the projected diagonal rather than the literal correlation-matrix diagonal.The projected profile is formed by evaluating the kernel in the right singular basis of W_t.
- A.1 Proof of Proposition 1: Weyl and Davis–Kahan controls reduce normalized tail-mass evolution to the injected profile plus an O(ω_tΨ_t) residual.The result assumes bounded cross terms and tail-energy errors.
- A.1 Proof of Proposition 1: Convexity of KL divergence makes mixing toward the reference profile contract the KL component of the spectral-tail potential.The spike penalty also contracts when top mass is mixed away from a single spike, with residual terms absorbed under smallness conditions.
C Complete Experiment and Robustness Summary
The experiments combine controlled optimizer sweeps with projected-kernel, optimizer, onset-metric, generalization, and tiny-image robustness checks. The appendix image study supports qualitative transfer beyond the teacher–student simulator but is excluded from the main scaling-law estimates.
- Experiment coverage: 4860 projected-kernel rows across 180 runs, 9 checkpoints, and 3 profile types provide the main robustness dataset.The full package contains 330 runs including robustness extensions, while the optimizer comparison covers 150 runs at d = 500 and d = 1000.
- Tiny-image sanity check: The tiny digits MLP and ConvNet train normally while their monitored spectral diagnostics change, supporting a qualitative external sanity check.The study uses full-batch Adam on 8-by-8 digit images and is deliberately not used to estimate γ or ρ.
- Tiny-image sanity check: The tiny-image table reports heavier monitored tail diagnostics at larger Adam learning rates, but remains appendix-only evidence rather than part of the main scaling law.The check broadens the codepath without replacing the controlled teacher–student experiments.
F Real pretrained transformer sanity checks
Real pretrained transformer matrices exhibit heavier sampled spectral tails than matched Gaussian nulls, while optimizer, generalization, onset, and failure-mode checks qualify interpretation. These results support diagnostic applicability outside simulation but do not establish a transformer training-trajectory hitting-time law.
- Optimizer comparison: Adam and weak-decay AdamW agree closely, whereas GD and signGD do not reach consensus onset under the tested grids.This optimizer comparison is limited to the tested regimes and does not establish universal optimizer behavior.
- Real pretrained weights: Qwen2.5-0.5B scans found a mean null-relative Hill-alpha shift of −1.559, with 94.05% of matrices lower and 98.21% showing larger top-tail mass than matched nulls.The scan covered 168 internal attention and MLP projection matrices.
- Real pretrained weights: Pythia-70M matrices likewise show heavier sampled spectral tails than matched Gaussian nulls, with attention and MLP projection families separated by the diagnostics.A selected-matrix few-step Adam run verified that spectral monitoring can execute on real pretrained matrices.
- Cautions and sensitivity: Test MSE is non-monotone across Hill-alpha bins, and low-gap trajectories remain censored while extreme spike regimes do not imply monotone generalization improvement.Consensus, Hill-only, and related diagnostics are compared to separate robust onset from single-metric artifacts.
H Reproducibility statement
The reproducibility package centers on fixed-seed synthetic teacher–student experiments and includes the simulation, theory, likelihood, optimizer, robustness, onset-sensitivity, and real-transformer materials. These artifacts document the controlled evidence and supplementary checks supporting the paper.
- Artifact scope: All main experiments use fixed-seed synthetic teacher–student data, with simulation tables, figures, and audit summaries included in the artifact package.The supplementary code package includes the sweep runner and configuration files.
- Supplementary materials: The appendix includes theorem proofs, the right-censored likelihood derivation, optimizer comparisons, projected-kernel robustness, onset-metric sensitivity, and real-transformer sanity checks.These materials complement the main controlled experiments rather than replacing them.