Source-linked AI summary

High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube

Josef Dick, Michael Feischl, Fabian Zehetgruber

arXiv:2608.26526v1cs.LGmath.NA

TL;DR

The paper addresses the lack of finite-width, nonasymptotic bounds that simultaneously control square-free mixed input derivatives of random neural networks over a domain. It combines deterministic derivative estimates with a probabilistic analysis of tangent directions in wide Gaussian networks, obtaining polynomial rather than generally exponential depth dependence for tanh networks with Xavier initialization. These estimates yield Lipschitz and weighted Sobolev consequences relevant to robustness and QMC, while remaining subject to scope limitations concerning training, isotropy, and derivative type.

  • Problem

    Prior results do not provide a finite-width, nonasymptotic high-probability event uniformly bounding all square-free mixed input derivatives with explicit dependence on order, depth, width, dimension, and failure probability.

  • Method

    The proof combines deterministic mixed-derivative estimates with a majorant-series argument and measurable finite nets controlling tangent directions in wide Gaussian networks.

  • Results

    For sufficiently wide scalar-output tanh networks with Gaussian Xavier initialization, the simultaneous derivative bound has depth factor (C1L)^{|u|-1}, while first-order bounds are independent of depth.

  • Takeaways & Limitations

    The derivative event provides high-probability Euclidean Lipschitz and weighted Sobolev consequences connected to QMC quadrature and lattice-based training.

  • Takeaways & Limitations

    The QMC interpretation applies at initialization, isotropic Xavier factors do not decay across coordinates, and the theorem controls only square-free mixed derivatives.

Abstract

from arXiv · show

We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output $\tanh$ networks with Gaussian weights and Xavier initialization, we prove that there exist constants $C,C_0,C_1>0$ such that, whenever the common hidden width satisfies $n \geq C\left(L^3n_0^2(1+\log n_0)+L^2\left(1+\log(L/η)\right)\right)$, then, with probability at least $1-η$, the estimate $\left|D^u\mathcal{R}_{Φ^{(L)}}(x)\right| \leq C_0 |u|! (C_1L)^{|u|-1}\prod_{j\in u}β_j(η,n_0)$ holds simultaneously for every non-empty $u\subseteq[n_0]$ and every $x\in[0,1]^{n_0}$. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order $|u|$ grows at most polynomially as $L^{|u|-1}$, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.

1. Introduction

The paper studies derivative-based regularity bounds for random smooth neural networks, motivated by adversarial robustness and QMC approximation. It develops a finite-width high-probability result that replaces generally exponential depth dependence with polynomial dependence for sufficiently wide Gaussian networks.

  • Motivation: Small adversarial input perturbations can cause large output changes, motivating Euclidean Lipschitz and input-derivative bounds for smooth neural networks.For convex domains, the Euclidean Lipschitz constant is bounded by the supremum Euclidean gradient norm.
  • Motivation: Mixed input derivatives also govern weighted Sobolev norms and QMC error bounds for high-dimensional surrogate models.These estimates are relevant to many-query problems involving expensive parameter-dependent PDE solves.
  • Related work: Existing derivative analyses include deterministic Lipschitz and Hessian certification, smooth approximation, random-network Jacobian studies, and Gaussian-process limits.Related finite-width work includes distributional and cumulant analyses, but addresses different aspects of random-network behavior.
  • Research gap: The paper targets a finite-width, nonasymptotic event uniformly controlling every non-empty square-free mixed derivative over the input domain.The bound explicitly tracks derivative order, depth, width, input dimension, and failure probability.
  • Setting: The network setting uses smooth activations, fixed arbitrary biases, and independent centered Gaussian Xavier-scaled weights, with the main probabilistic result specialized to tanh.The scalar-output network has L ≥2 hidden layers of common width n.
  • Contributions: The proof isolates the Faà di Bruno term linear in the highest-order derivative and controls its tangent directions using measurable finite nets.High-probability matrix and tangent conditions replace a full operator-norm factor by a layerwise factor of order 1 + L^-1 along relevant directions.
  • Consequences: The resulting consequences include Euclidean Lipschitz continuity, weighted Sobolev bounds, QMC quadrature, and lattice-based training.The QMC interpretation also highlights the need for coordinate anisotropy or regularization when isotropic Xavier factors do not decay.

2. Derivative bounds of neural networks

The section develops deterministic mixed-derivative bounds for smooth neural networks, first using Euclidean operator norms and then improving depth dependence through tangent-direction conditions and a majorant-series argument.

  • Depth dependence: If κ_1 = ··· = κ_L > 1, the resulting deterministic bound contains a depth-dependent product that grows exponentially in L.This motivates an additional condition designed to remove exponential depth growth.
  • Tangent control: The improved argument isolates the linear highest-order derivative contribution and controls associated tangent directions through conditions imposed uniformly over inputs and active coordinate sets.The tangent class is defined by inequalities holding for every non-empty u, every layer, and every x in the input domain.
  • Majorant-series argument: A generating-function majorant bounds derivative-order coefficients and yields the depth factor L^(k−1) for square-free mixed derivatives of order k.Evaluating the majorant at y = c/L gives F_ℓ(y) ≤ C_0y, which explains the factor L^(k−1).
  • Scope: The resulting regularity statement does not imply a holomorphic extension to a complex polydisc because it controls only square-free mixed derivatives.The limitation concerns the scope of the derivative estimates rather than the validity of the square-free bounds.

3. High probability bounds for random neural networks

The section establishes high-probability matrix and tangent conditions for sufficiently wide Gaussian networks, then applies them to obtain uniform derivative bounds under Xavier initialization.

  • Tangent events: Finite measurable nets control tangent directions uniformly over the input domain, active coordinate sets, and network parameters satisfying the matrix bounds.The construction uses finite covering sets and measurable selection maps for the relevant tangent sets.
  • Probability estimate: The matrix and tangent events intersect with probability at least 1 − η, providing the event on which the high-probability derivative theorem applies.The tangent estimates use Gaussian concentration for fixed net directions and a union bound under the width condition.
  • Tanh specialization: For tanh networks with Xavier initialization, sufficiently large common hidden width yields uniform bounds for every non-empty coordinate set and every x in [0,1]^n0.The scalar-output architecture has L hidden layers of width n and uses the stated width condition involving L, n0, and η.

4. Consequences for quasi-Monte Carlo methods and Lipschitz continuity

The derivative estimate yields product-and-order-dependent bounds useful for weighted Sobolev and QMC analysis, while its first-order specialization provides high-probability Lipschitz control. These consequences include conditional lattice-rule error guarantees and a dimension-independent RMS rate under suitable weight assumptions.

  • QMC and Sobolev consequences: Proposition 13 provides uniform-in-input derivative bounds with the product-and-order-dependent structure used in quasi-Monte Carlo analysis.The resulting estimates support weighted Sobolev norms and quadrature bounds.
  • QMC and Sobolev consequences: Corollary 14 extends the derivative estimate to every positive weight family on the same probability event.The event does not depend on the deterministic choice of weights.
  • QMC and Sobolev consequences: With probability at least 1 −η, randomly shifted lattice rules admit a conditional root-mean-square error bound for the network realization.The expectation is over the independent random shift, conditional on the network derivative event.
  • Dimension-independent QMC rate: The CBC lattice-rule construction achieves a dimension-independent RMS rate under an infinite, dimension-independent sequence of majorizing weights.The resulting rate is O(N^-r), with the stronger estimate N^-1/(2λ).
  • Limitations: The QMC interpretation is limited to initialization, coordinate-weight assumptions, and square-free mixed derivatives.Training stability, anisotropic decay, or repeated-derivative control would be needed beyond these boundaries.
  • Lipschitz continuity: For every p ∈[1, ∞], the high-probability Lipschitz bound is Lip_p(RΦ(L); [0,1]^n0) ≤ C0 ∥β(η,n0)∥_p′.After strengthening the width condition, this becomes Lip_p(RΦ(L); [0,1]^n0) ≤ 3C0 n0^(1−1/p).

Declaration on generative AI assistance

Generative AI assisted with mathematical ideas, literature search, and exposition, while the authors independently checked, revised, and approved the manuscript’s contents.

  • AI assistance contributed to formulating mathematical ideas, including Proposition 5 and Proposition 10.
  • AI tools were also used for literature search and drafting and revising parts of the exposition.
  • The authors independently checked, revised, and approved all definitions, theorem statements, proofs, calculations, and references.

Appendix A. Auxiliary results

The appendix develops auxiliary analytic, combinatorial, probabilistic, and measurability results used throughout the paper, including derivative estimates, partition identities, Gaussian matrix bounds, and measurable selection.

  • The multivariate Faà di Bruno formula is specialized to square-free derivatives indexed by finite coordinate sets.
  • The appendix establishes a factorial growth bound for derivatives of tanh using a holomorphic extension and Cauchy’s estimate.
  • The partition identity in Lemma 21 is proved by strong induction, separating partitions according to whether they contain the singleton block {x}.
  • For Gaussian matrices, Proposition 24 bounds the smallest and largest singular values with probability at least 1 − 2 exp(−t^2/2).
  • The appendix constructs Borel-measurable lexicographically smallest elements of compact sections through nested measurable interval refinements.
Loading 2608.26526v1…