Source-linked AI summary

Mean Field Analysis of Neural Networks: A Central Limit Theorem

Justin Sirignano, Konstantinos Spiliopoulos

arXiv:1808.09372v2math.PRmath.STstat.ML

TL;DR

The paper addresses limited mathematical understanding of single-hidden-layer neural networks in a regime with large network size and many stochastic gradient steps. It proves a central limit theorem using linearization and weak-convergence methods in a Sobolev-space setting. The limiting fluctuations around the mean-field limit are Gaussian and satisfy a stochastic partial differential equation.

  • Problem

    Limited mathematical understanding remains for neural networks, including their empirical parameter distributions in large-size, many-iteration regimes.

  • Method

    The proof linearizes the discrete-time nonlinear evolution and uses weak convergence, relative compactness, and uniqueness in a suitable dual Sobolev space.

  • Results

    The fluctuation process converges in distribution to a Gaussian limit satisfying a stochastic partial differential equation.

  • Takeaways & Limitations

    The central limit theorem quantifies finite-network fluctuations around the mean-field limit and supports a same-order scaling between hidden units and stochastic gradient steps.

Abstract

from arXiv · show

We rigorously prove a central limit theorem for neural network models with a single hidden layer. The central limit theorem is proven in the asymptotic regime of simultaneously (A) large numbers of hidden units and (B) large numbers of stochastic gradient descent training iterations. Our result describes the neural network's fluctuations around its mean-field limit. The fluctuations have a Gaussian distribution and satisfy a stochastic partial differential equation. The proof relies upon weak convergence methods from stochastic analysis. In particular, we prove relative compactness for the sequence of processes and uniqueness of the limiting process in a suitable Sobolev space.

1 Introduction

The paper proves a central limit theorem for single-hidden-layer neural networks trained by stochastic gradient descent, jointly in large-width and large-iteration regimes. The Gaussian fluctuations around the mean-field limit satisfy a linear SPDE, with convergence established through weak-convergence arguments.

  • The paper rigorously proves a central limit theorem for the empirical parameter distribution of single-hidden-layer neural networks.
  • The asymptotic regime simultaneously sends the number of hidden units and stochastic gradient descent iterations to large values.
  • The mean-field limit is a deterministic law-of-large-numbers description, while the central limit theorem gives the first-order finite-width correction.
  • The proof linearizes the discrete-time empirical evolution, controls vanishing remainder terms, and proves relative compactness and uniqueness in a dual Sobolev space.
  • The number of hidden units and stochastic gradient steps should be of the same order for convergence and statistically good behavior under the paper’s scaling.
  • The fluctuation limit is Gaussian and satisfies a linear stochastic partial differential equation coupled to the nonlinear mean-field PDE.

2 Sobolev Spaces

The analysis studies convergence in a dual Sobolev space on a fixed bounded domain. Compact support and sufficiently high regularity provide the setting for the fluctuation-process arguments.

  • The parameter measures are uniformly compactly supported with respect to network width and time.
  • W^J,2_0(Θ) is a Hilbert space, and W^-J,2(Θ) is its dual.
  • The convergence analysis uses the dual space W^-J,2(Θ) of the Sobolev space W^J,2_0(Θ).
  • The study takes J ≥3 for convergence in the relevant Sobolev space.
  • The domain Θ is bounded and fixed across network widths and times, although it may depend on fixed problem parameters.

3 Preliminary Calculations

This section develops the scaled empirical-measure evolution, separating drift, martingale, and remainder terms to prepare the fluctuation analysis. The remainder is shown to be O(N^-1).

  • The decomposition separates drift terms, martingale terms, and remainder terms in the fluctuation dynamics.
  • The drift components D1,N(t) and D2,N(t) are approximated by time integrals.
  • The process is càdlàg, with jumps occurring at discrete stochastic-gradient-descent update times.
  • O(N^-1) bounds control the remainder term under uniformly bounded parameters and compactly supported data.
  • The scaled empirical measure is expanded as a telescoping sum over the discrete-time particle evolution.

4 Relative Compactness

The section establishes relative compactness of the pre-limit fluctuation processes in negative Sobolev path spaces. Uniform bounds, compact containment, and short-time regularity provide the needed tightness conditions.

  • The fluctuation processes are shown to be relatively compact in D_W^-J,2([0,T]).
  • Uniform bounds in N and t provide the main estimate for the fluctuation process.
  • The proof uses Sobolev estimates, Parseval’s identity, martingale inequalities, and compactness arguments.
  • Compact support and bounded parameters ensure that the fluctuation measures vanish outside a fixed compact set K.
  • Short-time increment estimates establish the regularity condition required by the relative-compactness theorem.
  • Compact containment follows in W^-J,2(Θ), with J ≥ J2 + 1 = 3.

5 Continuity properties and identification of the limiting equation

This section identifies the fluctuation limit through convergence of martingale terms and pre-limit dynamics. The limiting process satisfies the linear stochastic evolution equation driven by a Gaussian martingale.

  • The pre-limit fluctuation processes take values in C_W^-J,2([0,T]).
  • The martingale terms converge to distribution-valued mean-zero Gaussian martingales with specified variance and covariance structures.
  • The Gaussian martingale has covariance structure determined by the limiting mean-field measure and the model’s data distribution.
  • Any limit point satisfies the stochastic evolution equation obtained by passing to the limit in the pre-limit equation.

6 Uniqueness of the stochastic evolution equation

This section proves uniqueness of solutions to the limiting stochastic evolution equation in the negative Sobolev space. The argument reduces the difference of two solutions to a deterministic equation and shows it vanishes.

  • The proof compares two candidate solutions and studies their difference Φ_t.
  • The difference satisfies a deterministic equation because the stochastic driving terms cancel.
  • A zero initial difference, Sobolev estimates, and Grönwall-type bounds imply that Φ_t remains zero.
  • The limiting stochastic evolution equation has a unique solution in W^-J,2.

7 Proof of the Main Result

The proof establishes convergence of the neural-network fluctuation processes by combining relative compactness, identification of limit points, and uniqueness of the limiting SPDE solution.

  • Relative compactness is established for the sequence of empirical-measure processes in the stated product space.
  • Every limit point satisfies the stochastic partial differential equation governing the fluctuations.
  • Uniqueness of the limit point, followed by Prokhorov’s theorem, yields convergence of the fluctuation process to the Gaussian correction.
  • The fluctuation process ηN converges in distribution to ¯η in the Skorokhod space DW −J,2([0, T ]).

8 Conclusion

The paper rigorously proves a central limit theorem for single-hidden-layer neural networks trained by stochastic gradient descent in a joint large-network and large-iteration regime.

  • The paper studies single-hidden-layer neural networks as both network sizes and stochastic gradient descent iterations become large.
  • It proves a central limit theorem for the empirical distribution of neural-network parameters.
  • The limiting fluctuation process satisfies a stochastic partial differential equation and has a Gaussian distribution.

A Proof of Lemma 4.2

The lemma’s proof handles discrete-time fluctuation dynamics by controlling remainder terms and working with compactness and Sobolev-embedding estimates.

  • The proof identifies and estimates a remainder term within the fluctuation analysis.
  • A compact set K ⊂ R1+d and the Sobolev embedding theorem provide the stated analytical setting and regularity control.
  • Compactness of X × Y and Young’s inequality are used together with an established bound to control the relevant expression.
  • Intermediate estimates are combined to obtain the target bound and complete the lemma.
  • The empirical-measure process changes at discrete times because stochastic gradient descent produces jumps.
  • The argument uses conditional expectations and the initial values of the particle parameters in its discrete-time estimates.

B Auxiliary lemmas

The auxiliary lemmas derive analytic bounds using smooth test functions, integration by parts, Young’s inequality, and dimension-wise estimates.

  • For smooth compactly supported g, the proof establishes a finite constant C controlling the relevant integral expression.
  • The auxiliary estimate is first proved in one dimension, with higher-dimensional algebra described as similar but more tedious.
  • Integration by parts and the smoothness of g supply one of the intermediate inequality steps.
  • Young’s inequality and smoothness assumptions are reused to bound the final term in the auxiliary estimate.
  • Rearrangement produces a finite constant C for the resulting bound.
Loading 1808.09372v2…