Source-linked AI summary
The Principles of Deep Learning Theory
Daniel A. Roberts, Sho Yaida, Boris Hanin
TL;DR
Deep neural networks require an effective theory that captures finite-width effects, representation learning, and training-algorithm dependence beyond the infinite-width description. The book develops recursive and RG-flow-based analyses of neural networks and derives conclusions about criticality, kernel behavior, and representation learning.
Problem
Infinite-width networks can mismatch practical observations, depend inadequately on training-algorithm properties, and fail to learn representations of their inputs.
Method
The book uses layer-to-layer iteration equations, nonlinear learning dynamics, and recursive marginalization to construct effective descriptions of neural networks.
Results
The framework analyzes deep learning through representation group flow and representation learning induced by nonlinear dynamical interactions at finite width.
Takeaways & Limitations
The principles developed support analyzing the deep representation flow and finite-width learning dynamics of nonlinear models.
Abstract
from arXiv · showhide
This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.
6 Bayesian Learning
Section 6 develops Bayesian learning, covering Bayesian probability, inference, model fitting, model comparison, and inference at finite and infinite width.
- Bayesian learning is introduced through probability, inference, model fitting, and model comparison.
- The section analyzes Bayesian inference at infinite width, including criticality and the absence of representation learning.
- It also examines Bayesian inference at finite width, including a treatment of Hebbian learning.
7 Gradient-Based Learning
Section 7 introduces supervised learning and connects gradient descent with function approximation.
- The section begins with supervised learning.
- It then examines gradient descent and function approximation.
8 RG Flow of the Neural Tangent Kernel
Section 8 develops the RG flow of the neural tangent kernel, beginning with its forward equation and analyzing fluctuations across layers.
- 8 RG Flow of the Neural Tangent Kernel: The section derives a forward equation for the neural tangent kernel from the MLP preactivation iteration equation.
- 8 RG Flow of the Neural Tangent Kernel: It analyzes the NTK across the first, second, and deeper layers, including deterministic behavior, fluctuations, and their accumulation.
- 8 RG Flow of the Neural Tangent Kernel: The section also covers interlayer correlations, the NTK mean, NTK–preactivation cross-correlations, and NTK variance.
9 Effective Theory of the NTK at Initialization
Section 9 studies the effective theory of the NTK at initialization through criticality analysis and universality classes.
- 9 Effective Theory of the NTK at Initialization: The section analyzes criticality of the neural tangent kernel.
- 9 Effective Theory of the NTK at Initialization: It characterizes a scale-invariant universality class and a K⋆= 0 universality class.
- 9 Effective Theory of the NTK at Initialization: It relates criticality to exploding and vanishing problems.
10 Kernel Learning
Section 10 covers kernel learning through incremental and large-step analyses, including Newton’s method and algorithm independence.
- Section 10 begins with a small-step analysis of kernel learning.
- The section then examines a giant leap, including Newton’s method.
- Algorithm independence is treated as a topic within the giant-leap analysis.
11 Representation Learning
Section 11 develops the differential neural tangent kernel and follows its representation-group-flow behavior across network layers.
- Section 11 introduces the differential of the neural tangent kernel.
- The section analyzes representation-group flow of the differential neural tangent kernel.
- The forward equations show zero differential neural tangent kernel in the first layer, nonzero values in the second, and growth in deeper layers.
335 ∞.1 Two More Differentials
The supplementary material extends the analysis to additional differentials, finite-width training, information, residual networks, and optimal aspect ratios.
- 335 ∞.1 Two More Differentials: The appendix studies two more differentials and finite-width training, including prediction at finite width.
- 335 ∞.1 Two More Differentials: Information-theoretic appendices examine criticality at infinite width and the optimal aspect ratio at finite width.
- 335 ∞.1 Two More Differentials: Residual-network appendices analyze criticality, finite-width optimal aspect ratios, and residual building blocks.
Initialization
The book frames deep learning as an effective-theory problem: realistic deep networks are difficult to understand macroscopically, but their known microscopic rules motivate tractable principles for analyzing learned functions.
- Modern AI systems combine billions of components whose training enables tasks once considered extraordinarily complex.
- Deep learning uses multilayer neural networks to learn increasingly refined representations helpful for underlying tasks.
- Deep learning theory remains disconnected from practice because many analyses rely on unrealistic assumptions and insufficiently address depth.
- The book aims to provide principles for theoretically analyzing deep neural networks of actual relevance.
- Statistical mechanics motivates deriving macroscopic regularities from the deterministic dynamics of many microscopic constituents.
- The central analytical goal is to compute the distribution of network functions simply across learning algorithms and training datasets.
A Principle of Sparsity
The infinite-width description simplifies neural-network analysis but misses finite-width interactions, algorithm dependence, and representation learning. A 1/n effective theory restores these effects, with the depth-to-width ratio controlling its validity and networks' effective behavior.
- Higher derivative terms vanish, function distributions become independent, and training dynamics become linear and algorithm-independent in the infinite-width limit.
- The infinite-width trained distribution is Gaussian and analytically tractable, but it mismatches practical multilayer networks whose outputs depend on the learning algorithm.
- At order 1/n, the trained-network distribution is nearly Gaussian and captures neuron interactions, learning-algorithm dependence, and nontrivial representation learning.
- The depth-to-width ratio L/n controls the relative size of expansion terms and the strength of interactions, providing the perturbative scale for the effective theory.
- Residual connections can shift the optimal aspect ratio to larger values, extending practical trainability and quantitative description toward deeper networks.
Wick Wick Wick: combinatorial derivation
The derivation models finite-width preactivation statistics with effective actions whose couplings reproduce two- and four-point correlators. Wick contractions, marginalization, and recursion relations show how nearly-Gaussian behavior and representation flow emerge across layers.
- Correlator derivation: Wick contractions express second-layer two- and four-point correlators through expectations of the stochastic metric.The connected four-point correlator is obtained after subtracting two-point contributions.
- Finite-width effects: The quartic vertex generates connected four-point correlations, linking finite width to correlations between activation pairs and feature learning.The quartic coupling is absent in the strict infinite-width limit.
- Effective action: A quartic action reproduces second-layer preactivation correlators to leading order in 1/n1.Matching the action’s correlators determines its quadratic and quartic couplings.
- Layer recursion: Recursive two- and four-point correlator equations extend the effective description through deeper layers while retaining leading finite-width corrections.Replacing metric quantities by their leading couplings gives a recursion for the quartic coupling.
- Marginalization and RG flow: Marginalization reduces integrals over all dataset samples to compact integrals over only a few samples and defines representation group flow.RG flow tracks the transformation of layer representations through repeated marginalization of fine-grained features.
- RG relevance: Finite-width couplings remain relevant under RG flow, so deep finite-width networks do not simply become infinite-width Gaussian models.The effective theory therefore retains qualitative differences from the quadratic Gaussian description.
Effective Theory of Preactivations at Initialization
At initialization, the effective theory analyzes deep networks through preactivation correlators and criticality conditions. The depth-to-width ratio controls finite-width corrections, while initialization and activation choice determine whether stable critical behavior is possible.
- Effective preactivation theory: Explicit two-point and four-point preactivation correlators are derived in asymptotic limits for deep networks.These correlators provide the effective description used to analyze initialization behavior.
- Finite-width scale: The depth-to-width ratio L/n controls the validity and strength of finite-width corrections to the infinite-width description.It acts as a cutoff scale for the effective theory.
- Critical initialization: Criticality requires tuning initialization hyperparameters so parallel and perpendicular susceptibilities satisfy the criticality conditions.The two conditions jointly determine the critical initialization parameters for a given activation.
- Critical behavior: At criticality, perturbations avoid exponential growth or decay, whereas χ < 1 produces exponentially vanishing perturbations around the fixed point.For scale-invariant activations, a critical line of nontrivial fixed points can occur.
- Universality classes: For scale-invariant activations, the critical prescription recovers Kaiming initialization for ReLU and identifies a shared limiting behavior across activation functions.This shared behavior motivates activation-function universality classes under representation group flow.
- No criticality: Some activations cannot be tuned to criticality: softplus fails the condition, while sigmoid requires an unphysical negative bias variance.Nonlinear monomials also fail the stated criticality requirement.
Deep asymptotic analysis for the midpoint kernel
The asymptotic analysis characterizes how critical midpoint kernels approach fixed points and how perturbations decay with depth. For the K⋆=0 universality class, the leading exponents are universal, while amplitudes retain activation- and data-dependent information.
- Midpoint-kernel asymptotics: At criticality, the K⋆=0 midpoint kernel approaches its fixed point through a mild power-law decay with universal exponent p0 = 1.This exponent is independent of the particular activation function within the universality class.
- Stability: The fixed-point stability condition requires (−a1) > 0; otherwise the asymptotic kernel becomes invalid and the fixed point repels exponentially.SWISH and GELU exhibit the stated instability near K⋆=0.
- Asymptotic expansion: The asymptotic expansion requires subleading 1/ℓ2 and log(ℓ)/ℓ2 terms to cancel higher-order contributions consistently.The expansion can be refined to arbitrary degree by adding higher-order corrections.
- Crossover behavior: For large inputs, the kernel first decays faster than exponentially bounded intermediate behavior before entering the power-law regime near the fixed point.The crossover leaves an undetermined constant ℓ0 that captures leading input-norm dependence.
- Universality: Deep RG flow makes activation details increasingly irrelevant: functions in the K⋆=0 class with the same first three Taylor coefficients become indistinguishable asymptotically.The critical exponent remains generic even though coefficients depend on activation details.
- Parallel perturbations: Parallel perturbations decay with universal exponent p∥ = 2, reflecting differences between input-dependent diagonal kernels.Their leading difference is governed by distinct input-dependent constants.
Deep asymptotic analysis for perpendicular perturbations
The analysis characterizes how perpendicular perturbations and finite-width corrections scale under RG flow across activation-function universality classes. Odd activations yield a definite perpendicular critical exponent, while fluctuation strength and correction relevance vary across classes.
- K⋆= 0 Universality Class: For odd activations such as tanh and sin, perpendicular perturbations decay as 1/ℓ, matching the midpoint kernel's approach to its fixed point.This gives the critical exponent p⊥ = 1 for the odd-activation universality class.
- Deep asymptotic behavior: Perpendicular perturbations can be more important than parallel perturbations because they exhibit a milder asymptotic falloff with depth.Their relative importance follows from comparing their depth-dependent scaling under the RG flow.
- Half-Stable Universality Classes: SWISH and GELU belong to half-stable classes that break scale invariance, and the analysis concludes that both are inferior to ReLU; tanh is preferred among smooth activations.Their critical initializations are close to ReLU's because these activations are small perturbations of ReLU.
- Fluctuations: The leading finite-width fluctuations scale with the depth-to-width ratio L/n, so deeper networks increasingly depart from the infinite-width limit.The effective theory is therefore most reliable in a regime where depth is not too large relative to width.
- Finite-width corrections: With an appropriate O(1/n) correction to CW, the NLO metric is subdominant to the kernel and vanishes in the interpolating limit.Such corrections can be neglected for most wide networks of reasonable depths, whereas four-point vertex corrections remain relevant.
- Universality classes: The scale-invariant universality class has pV = −1, whereas the K⋆= 0 class has pV = 1, reflecting opposite four-point-vertex scaling with depth.The scale-invariant class has V(ℓ) ∼ℓ, while the K⋆= 0 class has V(ℓ) ∼1/ℓ.
- Universality and fluctuations: The dimensionless four-point-vertex correction remains relevant under RG flow across different universality classes, even though its coefficient depends on the activation function.For example, the coefficient is 5 for ReLU and 2/3 for tanh, while their same-depth, same-width sensitivity is mostly similar.
Criticality analysis of the polar angle
The analysis characterizes how scale-invariant activation functions propagate polar-angle and perpendicular perturbations, identifying universal critical behavior and a caution around singular Gaussian expectations. It also connects criticality with Bayesian evidence and stable deep-network learning.
- ρ measures activation-function kinkiness at the origin, with ρ=0 for linear activation and ρ=1/π for ReLU.
- ρ=0 is the only case that exactly preserves the polar angle and the full two-input kernel matrix.
- pψ = 1 is universal except in the degenerate linear limit ρ=0, where pψ = 0.
- p⊥= 2 because the relevant quantity crosses from near-constant behavior at small depth to ∼1/ℓ2 decay at large depth.
- Singular Gaussian expectations require care when analyzing multiple inputs for nonlinear scale-invariant activations.
- For sufficiently deep networks, Bayesian model comparison prefers critical initialization hyperparameters, while infinite-width Bayesian posteriors lack representation learning.
Kernel Learning
At infinite width, kernel-based learning becomes analytically tractable: training dynamics simplify, fully trained solutions become independent of optimization paths, and representation learning disappears. The resulting predictions and generalization behavior can be characterized through the frozen NTK, bias–variance analysis, and nonlinear interpolation around training examples.
- Small-step learning: Infinite-width gradient-descent training is described by the frozen NTK, with network-output changes consistently truncated to linear order in the global learning rate.The output components move independently at the initial training step.
- Algorithm independence: The fully trained solution is determined by the initial network output, the frozen NTK, and the training set, while the mean prediction and covariance describe ensemble behavior on unseen inputs.At infinite width, the solution is algorithm-independent across the analyzed optimization procedures.
- Generalization: The generalized bias–variance decomposition separates test generalization error into bias from mean-prediction deviation and variance from realization-to-realization uncertainty.The decomposition is a good proxy when predictions remain close to the mean prediction.
- Interpolation and extrapolation: Nonlinear activation functions allow fully trained networks to nonlinearly interpolate or extrapolate around pairs of training examples, exposing activation-induced inductive bias.The analysis extends from test inputs near one training sample to test inputs near two training samples.
- Representation learning: Representation learning vanishes in the infinite-width limit because update means and covariances in the penultimate layer are suppressed by 1/n.The distributions before and after the learning update become equal in the strict infinite-width limit.
- A giant leap: A single theoretical gradient-descent step can fully train an infinite-width network, reaching the same minimum as many gradient-descent or stochastic-gradient-descent steps.This equivalence is used as a theoretical tool for analyzing fully trained networks.
Nonlinear ∗-Polation by Smooth Nonlinear Deep Networks
Smooth nonlinear deep networks produce nonlinear ∗-polations whose local curvature reflects activation- and architecture-dependent inductive bias, while infinite-width predictions reduce to fixed-feature linear models. Finite-width corrections make the neural tangent kernel dynamic, providing a mechanism for representation learning.
- Nonlinear networks can nonlinearly ∗-polate, and fully trained predictions depend on the initialized network output through the final-layer nonlinearity.
- The ratio δδΘ[0]/Θ00 measures local ∗-polation curvature, encoding the activation function and architecture’s non-universal inductive bias for generalization.
- Deep linear networks compute only linear functions, whereas nonlinear activation functions require solving a δδΘ[0] recursion to characterize their function classes.
- As s →0 or 1, ∗-polation approaches the corresponding training output with absolute certainty, and the nearest training point contributes most to a test prediction.
- At infinite width, Bayesian inference, neural-tangent-kernel regression, and kernel methods share equivalent mean-prediction forms based on fixed random features.
- Finite-width effects make the NTK evolve during training; the dNTK captures this evolution and connects deep networks to nearly-kernel methods with data-dependent feature learning.
Q-recursion
Finite-width corrections are governed by dNTK-related recursions whose leading effects scale with depth-to-width ratio ℓ/n. These corrections make the NTK evolve, yielding representation learning and nearly-kernel descriptions with feature evolution.
- Q(ℓ) remains order one across layers because its recursion starts from Q(1)=0 and preserves that scale.
- Leading dNTK-preactivation cross correlations are 1/n-suppressed and therefore become visible only at finite width.
- The dNTK recursions introduce critical exponents pP and pQ that describe the asymptotic depth scaling of P(ℓ) and Q(ℓ).
- All leading finite-width effects in the dNTK-preactivation joint distribution are controlled by the same ℓ/n perturbative cutoff across two universality classes.
- Finite-width networks are representation learners because nonzero dNTK-preactivation cross correlations grow as ℓ/n, while the dNTK enables the NTK to evolve during training.
- Nearly-kernel methods incorporate nontrivial feature evolution into kernel-like predictions, with effective kernels shifted by meta-kernel and training-output contributions.
B.2 Residual Infinite Width: Criticality Analysis
Residual connections modify the kernel recursions and critical initialization conditions by combining identity and MLP branches. The resulting finite-width analysis links residual strength to optimal aspect ratio and can extend effectively-deep networks to arbitrarily large aspect ratios.
- B.2 Residual Infinite Width: Criticality Analysis: Residual MLPs add the previous layer’s metric to the next-layer metric through the rescaled identity term.
- B.2 Residual Infinite Width: Criticality Analysis: Residual connections shift the critical initialization hyperparameters because the identity branch changes the kernel recursion.
- B.2 Residual Infinite Width: Criticality Analysis: The customary residual choice γ=1 cannot satisfy the criticality conditions, whereas each 0<γ^2<1 has an associated critical rescaled weight variance.
- B.2 Residual Infinite Width: Criticality Analysis: Critical residual solutions form a one-parameter family CW(γ), trading off identity-branch and MLP-branch contributions while preserving the infinite-width kernel for 0<γ^2<1.
- B.2 Residual Infinite Width: Criticality Analysis: Mutual information at initialization provides an estimate of the optimal finite-width aspect ratio r⋆.
- B.2 Residual Infinite Width: Criticality Analysis: As γ increases, ν(γ) decreases monotonically and the optimal aspect ratio r⋆(γ) increases for both universality classes.
- B.2 Residual Infinite Width: Criticality Analysis: The optimal residual hyperparameter can keep the relevant regime finite even as aspect ratio r tends to infinity, extending effectively-deep networks to arbitrary aspect ratios.