Source-linked AI summary

Critical initialization destabilizes higher input derivatives in wide scalar-input networks

Prashant Singh, Pranav Singh

arXiv:2609.09244v1stat.MLcs.AIcs.LG

TL;DR

The paper asks how initialization controls higher input derivatives needed by derivative-dependent objectives, beyond the first-order behavior captured by edge-of-chaos theory. It derives mean-field jet-covariance recursions and shows that critical fully connected networks grow unstable at higher orders, while L^-1/2 residual scaling bounds each fixed finite order. These are initialization results, not claims about trained-network performance.

  • Problem

    Derivative-dependent objectives require higher input derivatives, but the edge-of-chaos criterion specifies only first-order perturbation propagation.

  • Method

    The paper derives third-order mean-field recursions for the input-derivative covariance using infinite-width joint Gaussianity and analyzes depth-scaled residual networks.

  • Results

    At criticality, first-derivative variance stays constant, second-derivative variance grows linearly with nonzero activation curvature, and third-order variance grows quadratically; L^-1/2 residual branches uniformly bound every fixed finite order.

  • Takeaways & Limitations

    Independent layerwise initialization cannot preserve a nonzero second derivative and stabilize its variance at criticality simultaneously, whereas depth-scaled residual branches avoid this growth.

  • Takeaways & Limitations

    The results concern initialization, assume smooth activations, and use the infinite-width Gaussian-jet theorem at fixed depth rather than a depth-uniform limit.

Abstract

from arXiv · show

The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher input derivatives. For smooth scalar-input fully connected networks, using a joint Gaussianity of the finite derivative jet that holds in the infinite-width limit at each fixed depth, we derive mean-field recursions through third order that are exact at the variance fixed point, with finite-depth corrections that decay geometrically. At criticality, the first-derivative variance is depth-invariant, whereas the second-derivative variance grows linearly whenever the activation has nonzero curvature. The resulting third-order system closes on mean-field susceptibilities. For residual networks with branch scale L^{-1/2}, we prove that every fixed finite derivative order has uniformly bounded variance under explicit regularity assumptions. Simulations verify the critical growth laws, the residual bound, and the closed recursion. The results concern initialization, not trained-network performance.

1. Introduction

Higher-derivative objectives expose an initialization problem not captured by the edge-of-chaos criterion for first-order perturbations. The paper derives the resulting derivative-variance behavior and contrasts ordinary fully connected networks with depth-scaled residual networks.

  • Motivation: Physics-informed, gradient-regularized, and score-matching objectives depend on input derivatives beyond output variance alone.Physics-informed losses use differential-equation residual derivatives, while explicit score matching depends on second input derivatives.
  • Motivation: The edge-of-chaos condition preserves first-order perturbations but does not determine propagation of second and higher input derivatives.This gap motivates analyzing derivative-dependent objectives at initialization.
  • Theory: At criticality, second-derivative variance grows linearly and third-derivative variance grows quadratically when activation curvature is nonzero.The first-derivative variance remains constant under the same critical condition.
  • Residual networks: For residual layers scaled by L^-1/2, every fixed finite derivative order has uniformly bounded variance under bounded-derivative assumptions.The analysis uses a joint-Gaussian derivative jet in the infinite-width limit and derives a closed recursion through third order.
  • Scope: The conclusions concern initialization rather than trained-network performance, and the analysis assumes smooth activations.The connection to trained accuracy is explicitly not extended by the paper.
  • Conclusion: Independent layerwise initialization cannot both preserve a nonzero second input derivative and stabilize its variance at criticality, whereas depth-scaled residual branches avoid this growth.Finite-width simulations validate the recursions and residual prediction.

2. Related work

The paper situates its contribution between classical mean-field signal propagation, initialization remedies for physics-informed networks, and prior correlator hierarchies. Its distinctive focus is the depth-wise variance of higher input derivatives and the reach of residual scaling across derivative orders.

  • Mean-field initialization: Classical mean-field theory uses variance and correlation maps to characterize signal propagation and identifies criticality through the correlation-map slope.The edge-of-chaos setting maximizes trainable depth in that framework.
  • Physics-informed networks: Prior physics-informed-network studies analyze initialization-related training pathologies, kernels, or input encodings rather than the higher input-derivative variances studied here.The paper distinguishes its object from neural-tangent-kernel imbalance and first-derivative input encoding remedies.
  • Residual networks: The residual remedy of Wang et al. uses a trainable scalar initialized to zero, whereas this paper analyzes fixed attenuation that remains unchanged during training.Both approaches attenuate residual branches at initialization, but they are mechanistically distinct.
  • Residual networks: The L^-1/2 branch exponent was known to control forward variance, while this paper shows that the same choice bounds every input-derivative order simultaneously.The new result concerns the scaling’s reach, not the exponent itself.
  • Higher-derivative theory: Earlier higher-derivative work addresses computational cost, kernel-argument derivatives, or correlator hierarchies rather than depth-wise input-derivative variance behavior.The paper describes these questions as related but distinct.
  • Novelty: The paper claims that prior work does not isolate the criticality growth law for input-derivative correlators, while qualifying the literature search as non-exhaustive.The novelty claim is therefore stated cautiously.

3. Setup

The setup studies scalar-input fully connected networks at random initialization in the infinite-width, fixed-depth regime. It tracks a derivative jet and its covariance under mean-field Gaussianity, with explicit qualifications on the assumptions and limits.

  • Network model: The model is a width-n, depth-L fully connected network mapping a scalar input through componentwise activation, independent weights, and biases.The analysis takes n to infinity at fixed finite depth and derivative order, typically using tanh.
  • Mean-field map: In the infinite-width limit, preactivation variance follows a deterministic map with a fixed point, while χ1 controls first-order perturbation propagation.χ1 greater than, less than, or equal to one yields growth, decay, or preservation of first-order variance.
  • Propagation: The chain rule and Faà di Bruno’s formula propagate jet components through the layers, while the output uses a standard independent linear readout.The resulting covariance is the object entering the derivative recursions.
  • Derivative jet: The paper tracks the K-jet of layer preactivations and its covariance, including derivative variances and cross-order correlations.The zeroth-order preactivation is retained because it couples to higher jet moments.
  • Assumptions: The finite derivative jet is jointly Gaussian in the infinite-width limit at each fixed depth under smoothness and polynomial-growth conditions.This theorem supplies the jet analogue of ordinary mean-field Gaussianity, but does not give a depth-uniform width limit.
  • Assumptions: The derivation does not require independence between activation factors and jet factors; a separate factorization condition is explicitly not used by the results.Finite-width simulations examine the Gaussianity question at the widths used.

4. Theory

Mean-field Gaussian integration by parts and fixed-point identities produce closed recursions for derivative covariances. At criticality, resonance drives linear second-order and quadratic third-order growth, while residual scaling avoids this growth.

  • Recursion derivation: Gaussian integration by parts and fixed-point identities evaluate expectations coupling activation derivatives with products of jet components.The jet covariance includes zeroth-order couplings needed to prevent incorrect vanishing terms.
  • Finite-depth corrections: Finite-depth corrections to fixed-point identities decay geometrically at a rate governed by the variance map.For an attracting fixed point, the variance and its first three input derivatives approach their limits geometrically.
  • Mean-field coefficients: The correlation-map slope χ1 and variance-map slope λ play distinct roles: χ1 controls first-order propagation, while λ controls variance-map transients.This distinction separates critical signal propagation from convergence toward the variance fixed point.
  • Order two: At criticality, first-derivative variance is depth-independent, while second-derivative variance grows linearly when χ2 is nonzero.The order-two recursion is exact at constant preactivation variance, with finite-depth corrections otherwise.
  • Order three: At criticality, v1 and c12 are depth-independent, v2 and c13 have equal and opposite linear slopes, and v3 grows quadratically.The third-order system closes on χ1, χ2, and χ3, and κ13 cancels from the relevant recursion.
  • Order three: The closed system’s remaining forcing terms become positive after c13 turns negative, yielding the two-sided bound v3 = Θ(L^2).The sign change makes the quadratic growth statement stronger than an upper bound alone.

4.5. The recursion closes on the susceptibilities

At the variance fixed point, the derivative-variance recursion closes on susceptibilities: activation moments outside that family cancel through third order. Criticality preserves first-derivative variance but produces linear second-order and quadratic third-order growth, while residual scaling bounds fixed finite orders.

  • Closure on susceptibilities: Through third order, non-susceptibility activation moments cancel, leaving recursion coefficients expressed only through mean-field susceptibilities.The observed cancellations occur in paired squared and adjacent cross terms and rely on fixed-point identities.
  • Closure on susceptibilities: The general susceptibility-only closure is conjectured for every derivative order, but it remains unproved beyond orders one through three.A proof would need to show that fixed-point derivative identities cancel the Stein correction against adjacent cross terms at every order.
  • Growth at criticality: At criticality, the first-derivative variance is depth-invariant, the second-derivative variance grows linearly, and the third-derivative variance grows quadratically.The growth arises from collision of the two solution roots at χ1 = 1, producing a resonant linear term for second order.
  • Output and validation: The conclusions transfer from preactivation derivatives to network-output derivatives through the independent readout, and simulations validate the predicted growth and residual boundedness.The transfer holds for every derivative order covered by the analysis.
  • The dichotomy: For critical networks with nonzero curvature, independent layerwise initialization cannot preserve a nonzero second input derivative while keeping its variance stable with depth.The dichotomy is that χ2 > 0 yields unbounded second-derivative variance, whereas χ2 = 0 makes the second derivative vanish almost everywhere.
  • Residual scaling: With residual branch scale L^-1/2, every fixed derivative order up to K has variance uniformly bounded in depth under bounded activation derivatives and finite initial jet moments.The bound may depend on K, the activation, and the initial jet; the scaling exponent is shared across fixed finite orders.

5. Numerical method

The numerical method propagates derivative jets directly while exploiting Gaussian weight-product sampling and stable estimators. It controls initialization transients and tests polynomial growth using finite differences rather than asymptotic log-log slopes.

  • Jet propagation: Derivative jets are propagated with the Faà di Bruno expansion, using closed-form derivatives through order four for tanh.The implementation avoids repeated automatic differentiation and uses double precision.
  • Gaussian sampling: Conditional on the current jet, weight products are jointly Gaussian, enabling exact distributional sampling at cost O(n(K + 1)^2) per layer.The covariance is computed from the realized previous-layer jet, so finite-width fluctuations are retained rather than replaced by their mean-field limit.
  • Growth estimation: Polynomial degree is tested with constant finite differences because finite-depth log-log slopes underestimate the true growth degree.Measured slopes of 0.85, 1.80, and 2.63 for k = 2, 3, and 4 fall below the true degrees 1, 2, and 3.
  • Stable estimators: Covariances are evaluated in centered ratio form rather than as differences of separately estimated means to avoid cancellation when ratios are near unity.The alternative loses significant digits because the factorized and unfactorized expectations are close.
  • Experimental setup: Measurements use width n = 2^16, depth L = 256, and 16 independently sampled networks, with fitting over depths 64 through 256.The first layers are excluded because they contain the largest initialization transient, not because the theory fails there.

6. Parity arguments

Parity alone does not eliminate the cross terms in the order-three recursion because activation factors and derivative jets share dependence on the central preactivation. The resulting term is quantitatively important and can evade a natural ratio diagnostic.

  • Parity does not eliminate cross terms: For odd activations and symmetric central preactivations, adjacent activation-derivative products have zero expectation, but this does not make their jet-weighted cross terms vanish.Factorization fails because both the activation factor and the jet components depend on the same central preactivation.
  • Magnitude of the error: The omitted order-three cross term contributes -39.65 at depth 256, equal to 29% of the total forcing of 134.53.This demonstrates that the parity-based shortcut is not a negligible approximation.
  • Order dependence: Order one is immune because its recursion contains only the squared susceptibility χ1, so adjacent derivative products first arise at order two.Consequently, standard mean-field analyses of variance and input-output Jacobians do not encounter this failure.
  • Diagnostic failure: The usual factorization-ratio diagnostic is uninformative here because the factorized value is exactly zero, producing a 0/0 ratio.Other terms pass the diagnostic within 0.1% even though these cross terms are wrong by construction.

7. Numerical results

The simulations support the predicted critical growth laws, residual boundedness, and third-order closure, while showing that finite-width closure residuals fall sharply before plateauing near 3%.

  • Gaussianity and closure: Skewness and excess kurtosis of h(1), h(2), and h(3) remained below 0.01 in magnitude throughout the measurement window.The joint Gaussianity assumption for the finite derivative jet was numerically supported at the tested widths.
  • Critical growth laws: At χ1 = 1, the order-one slope was 3.6×10^-5 ± 5.1 × 10^-5, consistent with depth-invariant first-order variance.The order-two slope was 0.1719 ± 0.0089 versus a prediction of 0.1700 ± 0.0066, while the order-three quadratic coefficient was 0.2475 ± 0.0352 versus 0.2288 ± 0.0134.
  • Critical growth laws: The fitted third-order amplitude ratio stayed within 5% of unity across every measured χ1.The comparison used theoretical χ2 and an amplitude derived from v1, without fitting the theory-side quantities.
  • Residual networks: Residual scaling reduced variance growth from factors of 23, 590, and 1.8×10^4 to 1.39, 1.70, and 2.22 for orders two, three, and four.The comparison covered depths 8 to 512 at γ = 1/2; order four was used only as a boundedness check.
  • Gaussianity and closure: The closed covariance system matched all five measured ensemble ratios within 4% at the tested width and depth range.The integration used only initial χ1, χ2, and χ3 values after starting at ℓ = 64.
  • Width dependence: The largest closure deviation fell from 14.79% at n = 214 to 3.94% at n = 216 and 2.90% at n = 218, then appeared to plateau near 3%.The authors do not treat the fitted width exponents as a convergence rate because the resolved width changes are limited.

8. Discussion

The discussion argues that independent layerwise tuning cannot stabilize higher derivative variances at criticality, whereas attenuated residual branches provide the architectural route to boundedness.

  • Discussion: For losses depending on second or higher derivatives, no setting of the weight and bias variances avoids both first-derivative instability and second-derivative growth.Criticality stabilizes the first derivative but induces linear second-order growth; moving off criticality produces exponential first-derivative behavior.
  • Discussion: Branch attenuation by L^-1/2 keeps accumulated layer contributions bounded at every fixed derivative order.The paper identifies attenuation, rather than the skip connection alone, as the relevant architectural feature.
  • Scope: The analysis concerns initialization only and does not establish consequences for training dynamics or trained accuracy.Following those consequences would require tracking the derivative quantities under gradient descent.

9. Conclusions

The paper closes a third-order mean-field system for derivative covariances and shows that criticality preserves first derivatives while resonantly amplifying higher derivatives; residual scaling prevents this growth at initialization.

  • Conclusions: The derivative covariance recursions through third order are exact at the variance fixed point, with geometrically decaying finite-depth corrections.The joint Gaussianity of the finite derivative jet holds in the infinite-width limit at each fixed depth, and the closed system depends only on χ1, χ2, and χ3.
  • Conclusions: At χ1 = 1, first-derivative variance remains constant, second-derivative variance grows linearly when activation curvature is nonzero, and third-order variance grows quadratically.The linear growth arises from resonance between the recursion modes at criticality.
  • Conclusions: Simulations confirm the predicted growth exponents, parameter-free resonance amplitude, residual boundedness, and closed recursion.These validations concern the initialization analysis.
  • Open questions: Closure at every derivative order remains conjectural, and no order-four mean-field growth law is claimed.The evidence currently covers orders one through three; order four is reported only as a boundedness check.

CRediT authorship contribution statement

Prashant Singh and Pranav Singh contributed equally across conceptualization, formal analysis, investigation, methodology, software, validation, visualization, and writing.

  • CRediT authorship: Both authors contributed equally to the paper's conceptualization, analysis, investigation, methodology, software, validation, visualization, and writing.The statement assigns the listed roles to both Prashant Singh and Pranav Singh.

Funding

The paper reports no specific grant funding.

  • No specific grant supported the research.
Loading 2609.09244v1…