Source-linked AI summary
Loss-Parameterized Fisher Width Along Learning Trajectories
Vu Khac Ky
TL;DR
The paper asks when training loss can serve as an effective coordinate for Fisher width along learning trajectories, given that equal-loss parameters need not share Fisher geometry. It combines controlled-model theory with dynamical analysis and experiments, finding branchwise loss parametrization: learning dynamics can select a distinguished branch, while optimizer dependence remains in the nonlinear setting.
Problem
Equal-loss parameters need not have the same Fisher width, so the paper studies whether loss can nevertheless parametrize this observable along learning trajectories.
Method
The paper derives exact Fisher-width identities and stability bounds, analyzes population gradient flow in a Gaussian-teacher logistic model, and tests matched-loss behavior in controlled full-Fisher and nonlinear MLP experiments.
Results
The teacher-aligned state maximizes Fisher trace and Euclidean-ball Fisher width on each loss level below log 2, gradient flow selects this branch, and nonlinear experiments find GD and SGD close while Adam is substantially displaced.
Takeaways & Limitations
Loss parametrization is supported branchwise rather than universally: a distinguished branch may be selected by problem geometry and learning dynamics, but loss alone does not determine Fisher geometry throughout parameter space.
Takeaways & Limitations
The exact envelope and dynamical results are limited to population Gaussian-teacher logistic regression, while the nonlinear experiment uses a diagonal model-Fisher approximation and does not establish behavior for full Fisher matrices or neural networks generally.
Abstract
from arXiv · showhide
Fisher width measures the Gaussian width of a probe set after deformation by the local Fisher geometry. We study its evolution along learning trajectories and ask when training loss can serve as an effective coordinate for this quantity. We first derive an exact trace--shape factorization and a deterministic stability bound for fixed compact probes. In a population Gaussian-teacher logistic model, the teacher-aligned state is extremal on every loss level below $\log 2$: it has minimal parameter norm and maximizes both Fisher trace and Euclidean-ball Fisher width. We then show that population gradient flow asymptotically selects this branch, with explicit rates for the aligned and orthogonal coordinates. This yields, for $d\geq2$, \[ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}π\mathbb E[χ_{d-1}]. \] Controlled full-Fisher experiments support the matched-loss branch and the population predictions. In a nonlinear MLP with a diagonal model-Fisher approximation, GD and SGD remain close at matched loss, whereas Adam follows a substantially displaced branch; the fixed probes tested retain highly similar temporal shapes. These results support a branchwise, rather than universal, loss parametrization of Fisher width.
1 Introduction
The paper studies when training loss can serve as an effective coordinate for Fisher width along selected learning trajectories, rather than as a universal parameter-space law. Exact factorization, population geometry, gradient-flow analysis, and experiments support a branchwise interpretation.
- 1.1 Fisher Width Along Learning Trajectories: Fisher width is studied as a dynamical observable whose local Fisher geometry changes along learning trajectories.The paper asks whether loss can approximately coordinate Fisher width on selected trajectories, not whether iteration time or loss universally determines it.
- 1.3 Contributions and Scope: Controlled logistic experiments find closely agreeing full-Fisher-width curves at matched loss across tested optimizers and small initializations, while nonlinear MLP results retain optimizer dependence.In the MLP, GD and SGD remain close under diagonal model-Fisher geometry, whereas Adam follows a substantially displaced branch.
- 1.2 Matched-Loss Fisher-Width Branches: Below log 2, the teacher-aligned state has minimal parameter norm and maximizes Fisher trace and Euclidean-ball Fisher width on each loss level.The aligned relation is an upper envelope, not an identity: equal-loss parameters can have different Fisher widths.
- 1.2 Matched-Loss Fisher-Width Branches: Population gradient flow asymptotically selects the aligned branch, with near-branch residual O(|β|) for general probes and O(β2) for rotation-invariant probes.Along the actual population trajectory, this selection combines with low-loss asymptotics to produce the stated Fisher-width law for d ≥2.
- 1.2 Matched-Loss Fisher-Width Branches: A trace–shape factorization and deterministic perturbation bound separate Fisher scale from normalized Fisher shape for fixed compact probes.The bound includes probe-dependent sensitivity, so stable normalized geometry alone does not imply a loss-only Fisher-width law.
- 1.3 Contributions and Scope: The supported conclusion is branchwise loss parametrization: loss does not globally determine Fisher width, and the theory applies specifically to the controlled population setting.Finite-sample results provide pointwise or finite-checkpoint concentration, while the nonlinear experiment uses a diagonal Fisher approximation.
2 Dynamic Fisher Width and Matched-Loss Agreement
The controlled logistic experiments compare Fisher-width trajectories after reparametrizing them by shared observed loss. Matched-loss curves agree closely across tested optimizers and small initializations, but the evidence is exploratory and does not establish a global law.
- 2 Dynamic Fisher Width and Matched-Loss Agreement: The experiments compare optimization trajectories at matched loss rather than matched iteration to test whether they follow a common observed Fisher-width branch.Comparisons use only the shared observed loss interval and piecewise-linear interpolation on a common grid.
- 2 Dynamic Fisher Width and Matched-Loss Agreement: The controlled setup uses binary logistic regression with fixed teacher, dataset, initial parameters shared across optimizers, and the full model-Fisher matrix.The experiment evaluates GD, damped NGD, Noisy GD, and Adam across six initializations; no diagonal or trace approximation is used.
- 2.2 Matched-Loss Agreement Across Optimizers: The four tested optimizers produce closely agreeing full-Fisher-width curves over their shared observed loss ranges.NGD lies slightly above the GD reference on average, but the displacement is small relative to the width scale.
- 2.3 Stability Across Initializations: Within the tested small-initialization regime, full-Fisher width is highly reproducible after conditioning on loss.This statement concerns the tested initialization scale rather than arbitrary initial conditions.
- 2.3 Stability Across Initializations: The matched-loss reference curve is defined by piecewise-linear interpolation only within the observed loss range, without fitting a global functional form.The experiment does not infer a power law, log-linear law, or other extrapolated relation.
- 2.3 Stability Across Initializations: The controlled evidence is exploratory because the same fixed dataset generates trajectories and evaluates the sample Fisher matrix, so it is not a finite-sample trajectory theorem.Independent evaluation data and concentration guarantees are treated separately.
3 Trace–Shape Factorization and Probe Stability
The section factors Fisher width into a common trace scale and a probe-dependent normalized-shape term, then bounds shape-induced variation for fixed compact probes. Controlled diagnostics find small residuals, while emphasizing that this does not establish a loss-only law.
- Trace–shape factorization: Fisher width separates into a probe-independent trace-scale factor and a normalized Fisher-shape factor.Similar Fisher traces alone do not guarantee similar widths because normalized geometry also matters.
- Probe stability: The deterministic stability theorem applies to any fixed nonempty compact probe and depends on a probe-sensitivity factor.The square-root Fisher matrix is used directly, with a positive spectral lower bound needed to transfer matrix perturbations to square-root perturbations.
- Probe stability: Adaptive probes are outside the theorem because both the Fisher metric and the probe vary along the trajectory.Unbounded probes must first be restricted to compact, norm-constrained sets.
- Controlled diagnostic: 1.3 × 10^-2 is the largest observed normalized residual for the ellipsoid; ball, subspace, and sparse probes are smaller.The residuals remain below the empirical stability benchmarks, whose worst-case values are one to two orders of magnitude larger.
- Controlled diagnostic: The controlled GD diagnostic separates most tested width variation into a common trace scale and a slowly varying fixed-probe shape factor.It does not explain why Fisher trace should be organized by loss; that requires the population theory.
4 Controlled Gaussian–Logistic Theory
The Gaussian–logistic population model identifies the teacher-aligned ray as an extremal low-loss branch and shows that population gradient flow asymptotically selects it. Combining this geometry and dynamics yields a late-time Fisher-width law, while ruling out a global loss-only identity.
- Controlled model: The population analysis identifies a distinguished loss-parametrized branch, proves its low-loss extremality, and derives its late-time Fisher-width law.The model is intended to expose geometric and dynamical structure rather than establish a general loss-only description.
- Fisher geometry: The Fisher matrix has one longitudinal eigenvalue and a repeated transverse eigenvalue, with both scalar coefficients strictly decreasing along the aligned ray.Trace-normalized geometry is near-isotropic, although unnormalized longitudinal and transverse scales may differ asymptotically.
- Aligned envelope: For every loss below log 2, the aligned state has minimal parameter norm and maximizes Fisher trace and Euclidean-ball Fisher width.Both inequalities are strict away from the aligned branch.
- Aligned envelope: Equal-loss aligned and non-aligned parameters have different Euclidean-ball Fisher widths, so no global loss-only Fisher-width function exists.The aligned relation is therefore an envelope rather than a universal identity.
- Dynamical selection: Population gradient flow sends the aligned coordinate to infinity while the orthogonal coordinate decays, asymptotically approaching the aligned branch.The teacher-aligned ray is invariant, and arbitrary population gradient-flow trajectories receive the dynamical link to the branch.
5 Finite-Sample Control and Controlled Validation
The section establishes pointwise and finite-checkpoint empirical Fisher-trace control using independent evaluation data, then numerically checks the population envelope, dynamical selection, and late-loss law. Held-out ratios remain descriptive rather than trajectory-uniform guarantees.
- Finite-sample control: Pointwise empirical Fisher-trace concentration holds when the evaluated parameter is independent of the evaluation data.The proof uses a population decomposition with scalar and conditional Bernstein bounds.
- Finite-sample control: Finite-checkpoint trace control follows by applying the pointwise result with failure probability δ/N and a union bound.The same evaluation sample may be reused across checkpoints, although the estimates then remain dependent.
- Finite-sample control: A fixed evaluation sample loses resolution as the population Fisher scale becomes small, and the result is not trajectory-uniform.The effective sample size is governed by n m(r) in the low-variance regime.
- Population validation: Equal-loss quadrature agrees with the global aligned envelope and its near-aligned expansion.All quantities in this check are population quantities, so finite-sample approximation is not involved.
- Population validation: At t = 5000, the dynamical-selection diagnostics are within a few percent of limiting constants, while the normalized Fisher-width ratio is already very close to one.These values check the theoretical limits but do not establish experimental convergence rates.
- Held-out validation: Held-out Fisher-width ratios include finite-sample error in both loss and width and are used only as descriptive diagnostics.The controls apply to empirical Fisher trace, not the full Gaussian-width ratio, and do not strengthen the population asymptotic law.
6 Beyond the Controlled Regime
The nonlinear MLP experiment tests which matched-loss Fisher-width patterns persist beyond the controlled full-Fisher setting. Under a diagonal model-Fisher approximation, GD and SGD remain close, Adam separates, while fixed probes retain similar temporal shapes.
- Setup: The experiment uses a two-hidden-layer ReLU network on binary MNIST, comparing GD, mini-batch SGD, and Adam across six initializations.The network has hidden widths 256 and 128; training runs for 600 iterations with gradient clipping.
- Fisher geometry: The model-Fisher geometry is evaluated through a diagonal approximation on a fixed subset of 500 training examples, rather than the full neural-network Fisher matrix.The approximation is a conditional model Fisher, distinct from an empirical Fisher based on observed-label gradients.
- Optimizer-dependent matched-loss behaviour: GD and SGD remain close at matched loss, with average diagonal model-Fisher ball-width displacement below two percent, whereas Adam is substantially farther from the GD branch.Comparisons use only the loss interval shared by the trajectories, with no extrapolation; the displacement pattern is consistent across six initializations.
- Optimizer-dependent matched-loss behaviour: The MLP differs from the controlled logistic regime because loss reparametrization does not remove optimizer dependence: GD and SGD agree, but Adam follows a displaced branch.The logistic comparison uses the full 30 × 30 Fisher matrix, while the MLP uses diagonal model-Fisher geometry.
- Fixed-probe temporal stability: The fixed ball, random-subspace, sparse, and ellipsoidal probes retain highly similar temporal shapes along GD, with the sparse probe’s mean correlation above 0.99.Matched-loss coefficient of variation across initializations ranges from approximately 2.6% to 4.7% across the four probes.
- Scope: These findings support a branchwise rather than universal interpretation: some loss-parametrized organization persists, but the branch depends on the optimizer in this nonlinear experiment.The conclusion is restricted to the tested network, training protocol, diagonal model-Fisher approximation, and fixed probe family.
7 Discussion and Conclusion
The paper supports a branchwise account of loss-parametrized Fisher width: an aligned branch is extremal and dynamically selected in the controlled model, while experiments show optimizer- and representation-dependent deviations beyond it.
- Below log 2, the teacher-aligned state minimizes parameter norm and maximizes Fisher trace and Euclidean-ball Fisher width at each loss level.
- Population gradient flow asymptotically selects the aligned envelope, so the resulting Fisher-width relation holds along the actual trajectory.
- For d ≥2, combining dynamical selection with aligned low-loss asymptotics yields the stated limiting Fisher-width ratio along population learning trajectories.
- Near alignment, matched-loss residuals scale as O(|β|) for general compact probes and O(β2) for rotation-invariant probes.The linear bound for general probes is not claimed to be sharp.
- Controlled full-Fisher experiments show closely agreeing matched-loss curves across several optimizers, whereas the MLP’s diagonal model-Fisher experiment shows Adam on a displaced branch while GD and SGD remain close.
- The conclusions are limited by the controlled model’s isotropic Gaussian, noiseless linear-teacher assumptions and by experiments using one MLP, protocol, and diagonal Fisher approximation.Finite-sample theory covers fixed parameters or finitely many checkpoints, not continuous adaptive trajectories.
- The overall mechanism is level-set envelope, dynamical branch selection, then loss-parametrized Fisher-width asymptotics rather than a universal loss-only law.
A Technical Proofs
The appendix collects deferred technical proofs: the Gaussian-width probe-stability bound and finite-sample Fisher-trace concentration result.
- The appendix proves the Gaussian-width perturbation bound used in Section 3 and the finite-sample Fisher-trace concentration result used in Section 5.
A.1 Proof of the Probe-Stability Bound
The probe-stability proof applies the trace–shape identity to a nonempty compact probe and divides by its scale when that scale is positive.
- For a nonempty compact probe T, the proof invokes the trace–shape identity to establish the probe-stability bound.
- When the probe scale aT is positive, division by aT yields the normalized form of the bound.
A.2 Details for the Finite-Sample Trace Bound
The finite-sample trace proof combines rotational reduction, sub-exponential concentration, conditional control of transverse fluctuations, and a union bound over recorded checkpoints.
- Rotational invariance reduces the fixed-parameter analysis to a convenient coordinate representation.
- The proof controls scalar and transverse contributions using uniformly bounded sub-exponential norms and Bernstein inequalities.
- Opposite monotonicity of Z4 and v(rZ) supplies a covariance inequality used in the concentration argument.
- The resulting absolute and relative bounds hold with explicitly controlled failure probabilities.
- Conditioning on the training sample makes recorded parameters fixed and independent of evaluation data, enabling simultaneous control over N checkpoints by a union bound.