Source-linked AI summary
A Full Adam Theorem for Spectral Heavy-Tail Onset
Zongmin Liu
TL;DR
The paper addresses the missing Adam-specific derivation of spectral heavy-tail onset in a closed Gaussian Stein-Hermite state-evolution model. It derives the full chain from Adam recurrences through gradients, momentum kernels, denominator homogenization, Hermite edge transfer, and Gram dynamics. The theorem yields τ_ε=Θ(Δ_1^-γd^ρ log(Ψ_0/ε)) while restricting the claim to the specified model.
Problem
The missing result is a full Adam-specific theorem connecting actual Adam dynamics to spectral heavy-tail onset, rather than assuming the final spectral drift.
Method
The paper closes the chain using Stein-Hermite gradients, finite-width covariance concentration, a non-centered Gaussian sign kernel, denominator homogenization, Hermite edge transfer, and exact Gram updates.
Results
τ_ε=Θ(Δ_1^-γd^ρ log(Ψ_0/ε)) after the first spike-bulk gap, with matching upper and lower hitting bounds in the theorem's state-evolution setting.
Takeaways & Limitations
Within the closed state-evolution model, Adam's momentum and normalization can be connected formally to a spectral heavy-tail hitting law.
Takeaways & Limitations
The theorem does not extend to arbitrary-gradient Adam recurrences, and exact two-step linear-network loss dynamics do not identify factor spectra or heavy-tail hitting times.
Abstract
from arXiv · showhide
We prove a full Adam theorem for spectral heavy-tail onset in a closed Gaussian Stein-Hermite teacher-student state-evolution model. The theorem begins with the actual full-batch Adam recurrences, derives the population gradient by Stein-Hermite calculus, proves finite-width covariance concentration, converts multi-step Adam momentum into an exact non-centered Gaussian sign kernel, controls the diagonal Adam denominator by a basis-homogenization theorem, derives a regularly varying projected update response from a Hermite edge-transfer theorem, pushes the response through the exact Gram update, and proves approximate-target KL contraction with matching upper and lower hitting bounds. The final law is (τ_\varepsilon=Θ(Δ_1^{-γ}d^ρ\log(Ψ_0/\varepsilon))), where (Δ_1) is the first spike-bulk spectral gap. The result is full in the following precise sense: every step from Adam's momentum and denominator to the spectral hitting law is formalized inside the closed state-evolution model. We also prove that a stronger arbitrary-gradient Adam theorem is impossible, and that exact two-step linear-network loss dynamics do not identify factor spectra or heavy-tail hitting times.
1 The full-theorem target
The paper targets a full Adam-to-hitting theorem inside a closed Gaussian Stein-Hermite state-evolution model, deriving spectral heavy-tail onset after the first spike-bulk gap. Its scope includes explicit Adam dynamics, gradient and covariance closures, and a defined boundary against arbitrary-gradient claims.
- 1 The full-theorem target: The proof derives the trajectory from actual Adam recurrences, Stein-Hermite gradients, finite-width covariance concentration, and the exact Gram update.
- 1 The full-theorem target: The theorem proves how a full-batch Adam trajectory reaches a spectral heavy-tail window after the first spike-bulk gap Δ_1 appears.
- 1 The full-theorem target: New closures include a non-centered Gaussian sign kernel, diagonal-denominator homogenization, Hermite edge transfer, scalarized Gram cross terms, and approximate-target KL contraction.
- 1 The full-theorem target: The result is full only within the specified state-evolution model: arbitrary Adam recurrences alone cannot force onset because vanishing future gradients stop Adam's movement.
- 1 The full-theorem target: The spectral-tail potential measures deviation from a tail envelope while penalizing isolated spikes, and the hitting time is defined when it falls below ε.
4 Adam momentum, non-centered sign kernels, and denominator homogenization
The paper converts bias-corrected multi-step Adam momentum into an exact Gaussian sign kernel rather than relying on the centered arcsine special case. This kernel is derived from the jointly Gaussian gradient decomposition and retains nonzero mean effects.
- 4 Adam momentum, non-centered sign kernels, and denominator homogenization: Bias-corrected Adam momentum is treated through the jointly Gaussian decomposition of vectorized gradients, enabling an exact multi-step sign-kernel calculation.
- 4 Adam momentum, non-centered sign kernels, and denominator homogenization: The theorem uses the full non-centered Gaussian sign kernel; the centered arcsine law appears only as a corollary when the mean is zero.
Thus the correlation-dependent spectral exponent is the same as that of R(m)
The denominator closure shows that positive coordinatewise Adam normalization preserves the leading projected tail exponent under basis balance and delocalization. Thus the correlation-dependent spectral exponent remains unchanged.
- 4 Adam momentum, non-centered sign kernels, and denominator homogenization: A positive coordinatewise Adam denominator rescales projected sign energy but does not change the leading tail exponent in a delocalized singular basis.
- 4 Adam momentum, non-centered sign kernels, and denominator homogenization: The homogenization theorem requires a positive diagonal preconditioner that is balanced in the top singular window, with χ_t=o(1).
5 Hermite edge-transfer theorem
The Hermite edge-transfer theorem derives the regularly varying projected update response from a spectral-edge profile and a nonzero Hermite transfer coefficient. The resulting response is then propagated through a scalarized Gram update whose radial cross term cancels from the normalized profile.
- 5 Hermite edge-transfer theorem: The projected response is derived from a spectral edge and a nonzero Stein-Hermite transfer coefficient rather than assumed as the final drift.
- 5 Hermite edge-transfer theorem: The transfer function begins at the first nonzero Hermite-edge order q, with b_q≠0; it is polynomial for finite expansions and analytic for summable expansions.
- 5 Hermite edge-transfer theorem: The Gram theorem handles a signed cross term by decomposing it into a scalar radial component and a small trace-free perturbation.
- 5 Hermite edge-transfer theorem: The scalar radial cross term changes total top-window mass but cancels from the normalized spectral profile.
7 Approximate-target contraction and hitting
The theorem handles approximate target mismatch through KL contraction and establishes a two-sided spectral heavy-tail hitting law under explicit state-evolution assumptions.
- Approximate-target contraction: The approximate target profile requires contraction lemmas because b(t) is only approximately q, not exactly equal to it.
- Theorem assumptions: Theorem 8.1 assumes Stein-Hermite gradients, covariance concentration, Gaussian multi-step gradients, bounded momentum thresholds, denominator balance, edge transfer, scalarized spectral work, and controlled target mismatch.
- Approximate-target contraction: If target mismatch remains bounded relative to the current KL divergence, approximate contraction proceeds at rate 1 − cκ before hitting.
- Hitting law: The theorem supplies matching upper and lower hitting bounds uniformly before τε for both the deterministic closed flow and its high-probability finite-width approximation.
- Hitting law: The hitting exponents are derived from post-spike edge exponents and the first nonzero Stein-Hermite transfer order, rather than fitted in the proof.
9 Separation from exact two-step linear loss dynamics
Exact two-step loss dynamics in two-layer linear networks do not identify individual factor spectra, so they cannot determine factor-weight heavy-tail hitting times.
- Scope: Figure 2 provides algebraic and asymptotic theorem audits rather than empirical substitutes for Theorem 8.1.
- Non-identifiability: For every invertible S, the reparameterization (BS−1)(SA) preserves the predictive map, residuals, and loss while changing factor spectra.
- Implication: Therefore, exact two-step linear loss formulas cannot determine factor-weight heavy-tail hitting times.
10 Finite-size theorem audits and external spectral sanity checks
Finite-size checks audit the theorem links, while external transformer spectral checks provide only plausibility evidence and do not establish the dynamic Adam theorem or its exponents.
- Finite-size theorem audits: Numerical checks audit covariance concentration, momentum sign kernels, regular-variation preservation, Gram perturbation, and two-sided hitting recovery rather than replacing the analytical proof.
- External spectral checks: Static transformer spectra for Qwen2.5-0.5B and Pythia-70M show stable upper-tail geometry beyond an entry-shuffle null, but do not prove the dynamic theorem or estimate its exponents.
- Theorem mechanism: Adam’s coordinatewise normalization and momentum induce a non-centered sign kernel; edge transfer yields a regularly varying update profile, and Gram mixing contracts the spectral-tail potential.
- Proof audits: The underlying proof chain uses Stein-Hermite Gaussian identities, covariance concentration, and Gram perturbation within the closed state-evolution model.
C Proof of the non-centered momentum sign kernel
The proof derives an exact non-centered Gaussian sign kernel for Adam momentum, with the centered arcsine identity appearing only as a special case.
- Non-centered kernel: For Gaussian variables with nonzero means, sign-product expectations are expressed through univariate and bivariate Gaussian threshold probabilities.
- Centered special case: The centered case reduces to the classical arcsine law, but the theorem uses the full non-centered sign kernel.
- Denominator homogenization: Basis-delocalized directions make quadratic forms involving the diagonal preconditioner concentrate around normalized traces, enabling denominator homogenization.
- Regular-variation transfer: Bounded threshold factors and scalar denominator normalization preserve the regular-variation exponent of the correlation response.
F Proof of Gram profile mixing
The proof inserts the normalized top-window profile into the exact Gram identity, controls perturbations spectrally, and establishes matching upper and lower hitting bounds.
- F Proof of Gram profile mixing: The normalized profile is formed from diagonal top-window entries r_i(t) through R_t and b_i(t)=r_i(t)/R_t.
- F Proof of Gram profile mixing: Perturbations shift top-window eigenvalues by at most their operator norm, yielding a bound on the profile error via Weyl’s inequality.
- F Proof of Gram profile mixing: The update-energy profile supplies the d^-ρ scale used in the Gram-profile analysis.
- F Proof of Gram profile mixing: Convexity of KL in its first argument and a local Lipschitz bound control the perturbation contribution under the stated lower-coordinate condition.The local bound applies on x ≥ q_i/2, with the combination constrained by b ≤ q_1 + o(1) ≤ θ + o(1).
- F Proof of Gram profile mixing: Matching upper and lower hitting bounds follow by summing tail probabilities and iterating a no-teleportation inequality while the deterministic potential remains above ε.The upper bound comes from tail-probability summation; the lower bound preserves the potential above ε.
H Proof of the main theorem
The main theorem composes exact gradient, covariance, Adam-kernel, denominator, edge-transfer, Gram, and KL steps within the closed state-evolution model, while identifying scope limits for external checks and arbitrary gradients.
- H Proof of the main theorem: The proof derives the result through exact population gradients, finite-width covariance closure, a non-centered momentum sign kernel, denominator homogenization, Hermite edge transfer, Gram updates, and KL contraction.
- H Proof of the main theorem: Full-batch Adam in the closed Stein-Hermite Gaussian state evolution satisfies a two-sided spectral heavy-tail hitting law.
- H Proof of the main theorem: A centered arcsine identity is only a corollary because the theorem uses the full non-centered Gaussian sign kernel.
- H Proof of the main theorem: The Adam denominator preserves the leading projected spectral exponent only under basis balance or delocalization conditions.
- H Proof of the main theorem: The regularly varying projected response follows from an edge-regular profile and a nonzero Hermite transfer order, rather than being assumed directly.
- H Proof of the main theorem: Exact two-step linear-loss formulas do not identify factor spectra, while arbitrary adapted gradients cannot guarantee heavy-tail onset; static transformer checks provide plausibility rather than theorem proof.The gauge transformation preserves the map, residuals, and loss while changing factor spectra; vanishing future gradients give a concrete arbitrary-gradient counterexample.