Source-linked AI summary

Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach

Ryo Karakida, Shotaro Akaho, Shun-ichi Amari

arXiv:1806.01316v3stat.MLcond-mat.dis-nncs.LG

TL;DR

The paper asks how parameter-space geometry can be characterized universally across deep neural networks, where FIM analysis has been limited. Using mean-field theory with random weights and large-width limits, it derives asymptotic eigenvalue statistics showing widespread flatness and concentrated distortion, then connects these statistics to generalization measures and learning-rate selection.

  • Problem

    The geometric structure of parameter spaces shared across diverse deep neural networks lacks a general theoretical framework, despite the FIM's central role in statistics, machine learning, and parameter-space geometry.

  • Method

    The study uses mean-field analysis of FIMs for networks with random weights and biases in the large-width limit, deriving eigenvalue statistics through recurrence relations.

  • Results

    The FIM has mean eigenvalue O(1/M), variance O(1), and maximum eigenvalue O(M), implying most eigenvalues are near zero while the spectrum has a huge edge.

  • Takeaways & Limitations

    The statistics imply parameter spaces that are locally flat in most dimensions but strongly distorted in others, and support Fisher-Rao norm and learning-rate applications.

  • Takeaways & Limitations

    The paper studies random-weight and random-bias settings, and conditions used for its global-minimum learning-rate result do not generally hold after training.

Abstract

from arXiv · show

The Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large width limits, which enables us to utilize mean field theories. We investigate the asymptotic statistics of the FIM's eigenvalues and reveal that most of them are close to zero while the maximum eigenvalue takes a huge value. Because the landscape of the parameter space is defined by the FIM, it is locally flat in most dimensions, but strongly distorted in others. Moreover, we demonstrate the potential usage of the derived statistics in learning strategies. First, small eigenvalues that induce flatness can be connected to a norm-based capacity measure of generalization ability. Second, the maximum eigenvalue that induces the distortion enables us to quantitatively estimate an appropriately sized learning rate for gradient methods to converge.

1 Introduction

This paper develops a universal theoretical account of parameter-space geometry in deep neural networks by analytically studying the Fisher information matrix under random weights and large-width limits. Its eigenvalue statistics imply local flatness in most directions and strong distortion in others, with applications to generalization measures and learning-rate selection.

  • Motivation: Random connectivity and large-width limits provide a framework for deriving architecture-independent properties of deep networks.Prior work used this perspective to study expressivity and universal gradient behavior across layer counts and activation functions.
  • Motivation: The paper addresses the limited theoretical understanding of geometric structure shared across diverse deep neural networks.Existing flat-landscape analysis has mainly concerned shallow ReLU networks or special wide-network regimes.
  • Main results: The study analytically derives universal FIM eigenvalue statistics for a wide class of shallow and deep networks with varied activation functions.The mean, variance, and maximum eigenvalue are obtained through recurrence relations for macroscopic variables.
  • Main results: O(1/M) mean, O(1) variance, and O(M) maximum eigenvalue imply that most eigenvalues are near zero while the distribution edge is huge.Because the FIM defines the parameter-space metric, these statistics correspond to flatness in most dimensions and strong distortion in others.
  • Learning-strategy applications: The derived statistics connect small eigenvalues to the Fisher-Rao norm as a capacity measure and the largest eigenvalue to learning-rate estimation.The paper presents these as potential uses for assessing generalization ability and selecting convergent steepest-descent rates.
  • Related work: Mean-field analysis supplies theoretical evidence for the long-observed spectrum that peaks near zero and has a very long tail.Earlier theoretical evidence and evaluation of this skewed distribution were described as unresolved.

2 Preliminaries

The preliminaries define the FIM and explain its geometric interpretation, then specify the fully connected random network and mean-field framework used for analysis.

  • Fisher information matrix (FIM): The FIM is defined as the expectation of the outer product of score gradients for the network’s statistical model.For the Gaussian output model, it can also be expressed through network-output gradients and estimated empirically from training samples.
  • Scope: The framework studies the empirical FIM for arbitrary sample size T, which converges to the expected FIM as T →∞ and extends straightforwardly to related output models.The analysis focuses on Gaussian inputs, while finite-variance input distributions can also be considered.
  • Geometric interpretations: The FIM acts as a Riemannian metric, with dθ^T Fdθ representing local KL divergence and output robustness under parameter perturbations.It also determines the local loss landscape and coincides with the Hessian at the global minimum under the stated teacher-signal conditions.
  • Network setting: The study analyzes fully connected feedforward networks with linear outputs, Gaussian inputs, random weights, and hidden-layer widths proportional to a large width M.Activation functions and layer widths may vary, while the main setting keeps the output dimension C constant.
  • Mean-field approach: Mean-field analysis uses recursively computed feedforward and backpropagated macroscopic variables under an independence assumption for the two backpropagated chains.The large-width limit makes pre-activations amenable to Gaussian integration and central-limit arguments.

3 Fundamental FIM statistics

The paper derives universal large-width statistics for FIM eigenvalues across a broad class of networks: the mean is small, the variance remains finite, and the maximum grows with width.

  • 3.1 Mean of eigenvalues: O(1/M) is the asymptotic mean of the FIM eigenvalues as network width M becomes large.The mean follows from mλ = Trace(F)/P, with the recurrence coefficient κ1 remaining O(1).
  • 3.1 Mean of eigenvalues: Most FIM eigenvalues approach zero, making the parameter space locally flat in most dimensions as M grows.Because the FIM is positive semidefinite, the vanishing mean directly implies concentration near zero for most eigenvalues.
  • 3.1 Mean of eigenvalues: N(λ ≥k) ≤ min{ακ1CM/k, CT} bounds the number of eigenvalues above any fixed k > 0.Thus, at most O(M) eigenvalues can remain O(1), which is much smaller than the total parameter count P.
  • 3.2 Variance of eigenvalues: O(1) is the variance of the FIM eigenvalue distribution, while its mean is O(1/M), implying a very large distribution edge.The coexistence of a near-zero bulk and finite variance produces a strongly skewed spectrum.
  • 3.3 Maximum eigenvalue: O(M) is the order of the maximum FIM eigenvalue, which produces strong distortion in certain parameter directions.The maximum is proportional to α, so increasing depth strengthens this directional distortion.
  • 3.4 Numerical verification: Numerical experiments with tanh, ReLU, and linear networks agreed very well with the theoretical values of mλ, sλ, and λmax in the large-width limit.The experiments used random connectivity and 100 random ensembles with different seeds.

4 Connections to learning strategies

The paper connects universal FIM statistics to two learning strategies: Fisher-Rao norm estimates for generalization capacity and maximum-eigenvalue estimates for stable learning rates.

  • 4.1 The Fisher-Rao norm: The Fisher-Rao norm provides a capacity measure connected to generalization ability, with a theorem evaluating its typical value through κ1.The analysis averages the norm over parameters with fixed variances.
  • 4.1 The Fisher-Rao norm: In the large-width limit, the Fisher-Rao norm is independent of network width, matching an earlier empirical conjecture.The result concerns the average norm and is quantified by κ1.
  • 4.2 Learning rate for convergence: A learning rate η < 2(1 + µ)/λmax is necessary for steepest gradient descent to converge to a zero-error global minimum.The critical rate is defined as ηc := 2(1 + µ)/λmax; rates above it do not converge.
  • 4.2 Learning rate for convergence: Theorem 7 predicts that wider networks require smaller learning rates, while deeper networks require finer tuning, creating a trainability–expressive-power trade-off.The paper contrasts increasing optimization difficulty with exponentially growing expressive power as depth increases.
  • 4.2 Learning rate for convergence: Experiments on artificial data and MNIST found that theoretical learning-rate estimates aligned well with observed convergence and divergence regions.The experiments varied width and learning rate; CIFAR-10 results were reported as almost the same as MNIST.
  • 4.3 Multi-label classification with high dimensionality: For outputs with C = O(M), the mean eigenvalue remains O(1), while the maximum is between O(M) and O(M^2), broadening the distribution.This extends the main C = O(1) setting to classification problems with many labels.

5 Conclusion and discussion

The study presents universal asymptotic FIM statistics for broad classes of deep networks and discusses their implications, extensions, and practical boundaries.

  • Conclusion: Across networks with varying depths and activation functions, the FIM has a small eigenvalue mean and a huge maximum eigenvalue obtained through recurrence relations.These statistics imply local flatness in many directions and strong distortion in others.
  • Conclusion: The derived statistics support applications involving the Fisher-Rao norm and learning-rate selection for steepest gradient descent.The paper presents these as examples of connecting FIM statistics to learning strategies.
  • Discussion: Agreement between Gaussian-prior experiments and theory suggests broader applicability to i.i.d. finite-variance parameter distributions, pending further experiments.The applicable scope of the mean field approach remains to be clarified.
  • Discussion: Natural-gradient methods face instability because many FIM eigenvalues are near zero, making damping crucial in practice.The paper identifies the damping term ε in (F + εI)^−1∇θE(θ) as important for performance.
  • Discussion: Extending the framework to normalization methods, Hessian spectra, global parameter-space structure, and extremely deep residual networks remains future work.Extremely deep residual networks require careful treatment of width order and diverging macroscopic variables.

A.1 Theorem 1

Theorem 1 derives the asymptotic mean of FIM eigenvalues using macroscopic variables and large-width mean-field calculations.

  • Mean eigenvalue: The FIM eigenvalue mean is computed from the matrix trace divided by the parameter count.The analysis separately considers diagonal blocks and network outputs.
  • Mean eigenvalue: Large-width central-limit calculations express the relevant gradient sums through recursively defined macroscopic variables.The transformation applies across input samples and uses recurrence relations for the forward and derivative variables.
  • Mean eigenvalue: Weight entries dominate the mean eigenvalue asymptotically, while bias contributions are smaller in the M ≫ 1 limit.The scaling follows from the relative number and magnitude of the corresponding matrix entries.
  • Eigenvalue concentration: Nonnegative FIM eigenvalues allow Markov’s inequality to bound the number whose values exceed a fixed positive threshold.The resulting bound follows by combining positivity with the asymptotic mean.

A.3 Theorem 3

Theorem 3 analyzes the squared sum of FIM eigenvalues through a dual sample-space matrix, whose off-diagonal activity correlations determine the leading behavior.

  • Dual-matrix representation: The FIM is represented as F = B B^T/T, and its dual F* = B^T B/T has the same nonzero eigenvalues.This transfers the spectral calculation from parameter space to the T × T dual matrix.
  • Squared eigenvalue sum: The squared eigenvalue sum is obtained from the Frobenius norm and the entries of the dual matrix.For T = 1 and C = 1, the dual matrix reduces to the squared norm of the network-output derivative.
  • Asymptotic calculation: Large-width central-limit analysis makes the dual matrix asymptotically describable by macroscopic variables and Gaussian fluctuation terms.The recurrence relation for one macroscopic variable requires the stated assumption.
  • Exceptional case: If κ2 = 0, lower-order terms may become non-negligible, and their evaluation depends on the T/M ratio.This exceptional case lies outside the study’s scope.

(ii) C > 1 of O(1)

For a fixed, O(1) number of outputs, the dual FIM is analyzed through diagonal and non-diagonal blocks. The diagonal blocks dominate asymptotically, yielding an O(M) maximum eigenvalue while cross-output terms are lower order.

  • Matrix reduction: The dual matrix F* is used because it shares the FIM’s non-zero eigenvalues and is smaller when output dimension is fixed.F* is a CT × CT matrix, while F is P × P.
  • Block structure: The diagonal blocks have leading order O(M), whereas non-diagonal blocks become lower-order because different outputs do not share the final-layer weights.The backpropagated cross-output terms vanish through hidden layers, producing entries of lower order.
  • Asymptotic decomposition: The leading part of F* is block diagonal, with each diagonal block given by αMK/T; the residual term represents cross-block contributions.The residual can become non-negligible when C = O(M), but this section considers fixed C = O(1).
  • Maximum eigenvalue: λmax = O(M) because the upper and lower bounds asymptotically coincide at this order.The upper bound uses the spectral norm and the lower bound sums entries of a diagonal block.
  • Other statistics: The eigenvalue mean and second moment are evaluated from the corresponding block contributions, with cross-output terms incorporated through their asymptotic bounds.The second-moment analysis reduces the upper bound to a sum of the C = 1 results.

A.6 Theorem 5

The paper derives a Fisher-Rao norm result and analyzes gradient convergence near a zero-error global minimum. It also gives activation-specific recurrence relations for computing the macroscopic quantities used in the analysis.

  • Fisher-Rao norm: The Fisher-Rao norm is introduced through entries of the FIM and expanded around infinitesimal random weights.A Taylor expansion is substituted into the norm and averaged over the weight parameters to obtain its leading term.
  • Fisher-Rao norm: Terms involving distinct parameter entries vanish after averaging, simplifying the leading-order Fisher-Rao norm expression.The cancellation follows from averaging products containing independent zero-mean weight variables.
  • Gradient convergence: Near a zero-error global minimum, diagonalizing the FIM reduces the momentum update to independent scalar recurrences for each eigen-direction.The coordinate transformation preserves gradient stability.
  • Gradient convergence: η < 2(1 + µ)/λmax is necessary for the steepest gradient to converge to the global minimum.The condition follows from requiring ηλi < 2(1 + µ) for every FIM eigenvalue.
  • Activation-specific recurrences: For erf, ReLU, and linear activations, the macroscopic recurrence relations can be evaluated analytically without numerical integrations.The erf case uses Gaussian-integral identities, while ReLU and linear activations permit explicit evaluations of the required recurrences.

C.2 Training on CIFAR-10

Figure C.2 presents training losses after one epoch of SGD for Tanh, ReLU, and linear networks trained on CIFAR-10.

  • The color map compares training losses after one epoch of SGD across Tanh, ReLU, and linear networks on CIFAR-10.
Loading 1806.01316v3…