Source-linked AI summary
Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute
Xiaoyang Li, Runni Zhou
TL;DR
The paper studies whether differentiating hard-ReLU gradient descent and taking its vanishing-step flow limit commute. It develops a curvature-and-saltation framework, proves an event-free discrete derivative limit versus an event-aware flow derivative, and establishes the scope and extensions of this discrepancy.
Problem
State accuracy of a continuous-time surrogate does not by itself resolve whether its differentiated sensitivity matches increasingly fine hard-ReLU GD.
Method
The paper compares exact automatic differentiation of nonresonant hard-ReLU GD with a prepoint Stieltjes flow-sensitivity system separating regional Hessian density from speed-normalized event atoms.
Results
A nonzero event jump yields a rank-one discrepancy; globally convex losses prevent complete cancellation after a strict event, while a globally 1-strongly convex residual-ReLU family attains DφT = Jreg/R for any fixed R > 1.
Takeaways & Limitations
Differentiated flow surrogates represent fine discrete training only when omitted event transfers are absent or appropriately restored, with analogous parameter and adjoint extensions.
Takeaways & Limitations
The results cover deterministic full-batch finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events, not grazing, simultaneous events, Zeno behavior, or prevalence in large networks.
Abstract
from arXiv · showhide
Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. We prove that, over a fixed finite horizon, the GD states converge and these exact discrete derivatives approach an event-free regional propagator, whereas the derivative of the limiting flow also contains speed-normalized activation-event transfers. A prepoint Stieltjes representation separates the absolutely continuous regional Hessian from atomic interface curvature; one nonzero gradient jump produces an exactly rank-one endpoint discrepancy, and global convexity prevents complete multi-event cancellation whenever an event is strict. Nevertheless, a standard family of globally 1-strongly convex residual-ReLU squared-loss risks realizes arbitrarily large reciprocal sensitivity ratios on open initialization sets, with a uniform transversality margin. The same discrete-versus-flow decomposition extends to parameters and reverse-mode adjoints; resolved smoothing in the scalar or autonomous-normal regime and consistent event localization recover the flow sensitivity. The results concern deterministic full-batch, finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events; they are consistency theorems, not prevalence claims for large-scale training.
1 INTRODUCTION
The paper shows that exact differentiation of fixed-step hard-ReLU GD can converge to an event-free derivative even when the derivative of the limiting flow includes event transfers. It characterizes the resulting discrepancy, its convexity constraints, and extensions to parameters and reverse mode.
- Noncommuting limit: At every fixed nonresonant step size, automatic differentiation exactly differentiates the executed hard-ReLU GD program, but differentiation need not commute with the vanishing-step limit.The discrete states converge while their tangent dynamics can approach a different object from the limiting flow derivative.
- Noncommuting limit: Event timing creates the missing saltation transfer because nearby initializations reach activation boundaries at different times and experience different regional vector fields.Hard-branch numerical steps retain only an I + O(η) branchwise Jacobian, so no finite event transfer survives in the discrete derivative limit.
- Curvature and geometry: A speed-normalized pathwise curvature measure separates regional Hessian density from atomic activation-event curvature, with Jreg obtained by deleting the atoms.The framework gives exact rank-one one-event gaps and transported multi-event cancellation, volume, and interaction formulas.
- Convexity and sensitivity: For globally convex losses, every event weakly attenuates its direct normal mode, so one strict event prevents complete cancellation over a finite same-direction itinerary.Strong convexity alone does not guarantee commutation.
- Convexity and sensitivity: The globally 1-strongly convex residual-ReLU family realizes DφT = Jreg/R for any fixed R > 1 on an open initialization set with a transverse-speed margin independent of R.This establishes an arbitrary multiplicative reciprocal ratio rather than an unbounded additive discrepancy.
- Extensions and scope: The decomposition extends to parameter derivatives and reverse-mode adjoints, while the paper distinguishes these consistency results from claims about large-scale training prevalence.Reverse-mode AD remains exact for each fixed nonresonant discrete program.
2 SETUP AND THE NONCOMMUTING LIMIT
Under a stable, separated, same-direction transverse event itinerary, hard-ReLU GD states converge while exact discrete derivatives approach an event-free regional propagator rather than the saltation-interleaved flow derivative.
- Setup: The setup assumes piecewise-C2 regional loss extensions and a trajectory crossing finitely many activation interfaces under a stable, separated itinerary.Grazing, sliding, simultaneous hits, chattering, Zeno behavior, and terminal-time events are excluded.
- Discrete program: If no evaluated iterate lies on an interface, the hard-ReLU branch sequence is locally fixed and automatic differentiation differentiates that fixed discrete program.At an exact interface hit, software selects a product even though the endpoint map need not be classically differentiable.
- Noncommuting limit: For sufficiently small nonresonant step sizes, the GD derivative converges to the event-free regional propagator Jreg, while the limiting flow derivative retains event transfers.Thus state convergence alone does not guarantee derivative convergence.
- Rates: The discrete derivative satisfies ∥JADη,T − Jreg∥ = o(1), with an O(η) rate under locally Lipschitz regional Hessians, while ∥Jreg − DφT∥ may remain O(1).The distinct O(1) gap is caused by the omitted event contributions.
- Multiple events: Multiple event discrepancies factor through transported event terms, and equality requires more than the scalar condition rj = 1 for each event.The analysis uses submultiplicative matrix norms and ordered transport products.
- Scheme scope: The derivative-limit defect persists for verified event-blind regional one-step schemes, whereas event-localizing schemes are outside this result.Adaptive, implicit, or stage-switching methods require separate verification.
3 ATOMIC CURVATURE AND EXACT EVENT GEOMETRY
Activation events contribute speed-normalized atomic curvature that the flow derivative retains but regional discrete propagators omit. For one nonzero event, this omission is an exactly rank-one endpoint discrepancy whose visibility depends on transported alignment.
- Event-time mechanism: Different event times create a saltation transfer because perturbed trajectories spend different durations under incoming and outgoing vector fields.The transfer appears when perturbations are returned to a common clock.
- Event geometry: The direct event update is a rank-at-most-one identity perturbation acting only in the event-normal mode.It is rank one when the gradient jump coefficient β is nonzero.
- Atomic curvature: The pathwise curvature measure combines regional Hessian density with atomic interface curvature, producing a prepoint Stieltjes system with event jumps.Deleting the atoms yields the regional propagator Jreg, the limit of branchwise discrete AD.
- Endpoint discrepancy: A single nonzero event produces an exact rank-one endpoint gap, with outer-gradient visibility determined by the terminal gradient’s alignment with the transported event normal.The discrepancy can vanish when β = 0 or when the terminal gradient is orthogonal to the transported normal.
- Interpretation: The construction separates event strength, inverse crossing speed, incoming alignment, and outgoing transport without claiming that large discrepancies are typical in random networks.The result is geometric and does not establish prevalence.
4 CONVEXITY AND RELU EVENT STRUCTURE
Global convexity fixes the direction of ReLU event transfers and prevents complete cancellation when any event is strict. A residual-ReLU squared-loss family shows that strong convexity can coexist with arbitrarily large reciprocal sensitivity ratios on open initialization sets.
- 4.1 CONVEXITY FIXES THE DIRECTION OF TRANSFER: For globally convex losses, each event weakly attenuates its direct normal first variation, while strict events prevent complete cancellation across a same-direction itinerary.Negative singular interface curvature would instead amplify the normal mode and is incompatible with global convexity.
- 4.2 A RELU EVENT LAW: The residual-ReLU event law classifies χ > 0 as attenuation, χ = 0 as invisibility, and χ < 0 as amplification.For squared loss, χ is the residual projected onto the switching unit’s downstream output direction.
- 4.2 A RELU EVENT LAW: The event law is restricted to a single simple event, excluding simultaneous switches and genericity claims.Its derivation differentiates adjacent network graphs under same-direction transversality.
- 4.3 RESIDUAL-RELU PHASE THEOREM: For every initialization in the open set U_T, the residual-ReLU flow crosses zero once, and the discrete derivative converges to Jreg along nonresonant mesh sequences.The family uses a fixed event structure over the stated initialization set.
- 4.3 RESIDUAL-RELU PHASE THEOREM: For every R > 1, a globally 1-strongly convex choice realizes reciprocal sensitivity ratios with a transverse-speed lower bound independent of R.The ratio is multiplicative and holds on the open set U_T; the limit is taken for each fixed R before η ↓0.
5 HYPERPARAMETERS, REVERSE MODE, AND RESOLUTION
The discrete-versus-flow sensitivity gap extends to parameters and reverse-mode adjoints under a stable separated itinerary. Resolved smoothing and consistent event localization can recover flow sensitivity, but explicit tracking remains a consistency construction rather than a scalable training algorithm.
- 5.1 PARAMETERS INSIDE TRAINING: Parameter derivatives converge to the event-free continuation of the discrete derivative, whereas the limiting flow derivative includes event jumps.The result assumes joint regional smoothness and persistence of a stable, separated same-direction itinerary.
- 5.1 PARAMETERS INSIDE TRAINING: Initialization, vector-field, and moving-surface parameters contribute different terms through B, regional forcing, and Cj, while the step size is not differentiated.These parameter effects may coexist.
- 5.2 REVERSE MODE: Reverse-mode adjoints reproduce the same distinction: omitting event jumps and accumulators returns the event-free sensitivity limit, which may differ from the flow derivative.The transpose transfer follows from fixed-clock pairing rather than an inverse transpose.
- 5.3 RESTORATION AND SMOOTHING: Event-aware sensitivity recovers the flow object by localizing the stable itinerary and inserting consistent one-sided transfers.The construction requires one-sided data for every detected switch and is not presented as a scalable training algorithm.
- 5.3 RESTORATION AND SMOOTHING: Resolved smoothing restores the same transfer in the scalar or autonomous-normal regime when the shrinking curvature layer is resolved.The sufficient regime requires ρ → 0, R → ∞, and τR → 0; no general curved-interface smoothing theorem is claimed.
6 CONTROLLED NUMERICAL VERIFICATION
Controlled experiments verify that states converge while ordinary branchwise derivatives can remain separated from flow derivatives, with event-aware corrections restoring the flow sensitivity.
- The endpoint-sensitivity ratio follows r across convex attenuation, a C1 interface, and nonconvex amplification, while reciprocal ratios approach max{r, r−1}.
- At η = 2 × 10−4, the worst relative sensitivity-ratio and outer-gradient-ratio errors are 2.31 × 10−4 and 1.29 × 10−3.
- All 88 runs have one crossing, with fixed minimum one-sided speed one and eight nonresonant step sizes across r values from 1/40 to 40.
- At η = 2 × 10−4, the two-parameter event-aware correction reduces ∥JADη,T − DφT∥2 from 0.26865 to 1.23 × 10−4.
- Matrix and parameter experiments show states approaching flow endpoints, branchwise derivatives approaching regional limits, and event-aware forward and adjoint calculations restoring the flow derivatives.
7 RELATED WORK AND NOVELTY BOUNDARY
The paper distinguishes its learning-specific comparison of exact hard-ReLU GD derivatives with flow derivatives from established hybrid-systems sensitivity and GD–flow state literature.
- Prior work covers saltation, discontinuous-ODE sensitivity, parameter and adjoint jumps, singular Hessian measures, and GD–flow state comparison.
- η,T and DφT are treated as classical derivatives of a nonresonant discrete program and a stable-itinerary flow map, without identifying either with a generalized derivative selection.
- The paper characterizes when the regional and flow derivatives differ, including transported cancellation, convexity-based no-cancellation, and arbitrary reciprocal ratios in a globally strongly convex residual-ReLU risk.
8 LIMITATIONS AND CONCLUSION
The conclusions establish a noncommuting differentiation and vanishing-step limit, while bounding the claims to controlled finite-horizon dynamics and specific event regimes.
- Limitations: The study assumes deterministic full-batch, finite-horizon dynamics with continuous piecewise-C2 objectives and a stable finite itinerary of separated same-direction transverse events.
- Limitations: The scope excludes grazing, sliding, simultaneous or created events, chattering, and Zeno behavior, and does not establish prevalence in large networks.
- Limitations: Complete smoothing recovery is limited to scalar or autonomous-normal reductions, while explicit event tracking is diagnostic rather than a scalable training method.
- Conclusion: At fixed nonresonant step size, automatic differentiation exactly differentiates the executed hard-ReLU GD program, but its vanishing-step derivative limit can omit finite event transfers present in the flow derivative.
- Conclusion: One nonzero event jump creates a rank-one discrepancy, and any strict event in a globally convex loss prevents complete cancellation.
REPRODUCIBILITY STATEMENT
The paper formalizes piecewise-smooth ReLU objectives and sharp-interface sensitivity using regional Hessians, atomic curvature, event transfers, and reproducible numerical verification.
- Reproducibility: The numerical records use preserved JSON or NPZ arrays, with analytic curves regenerated from stated formulas rather than digitized from raster images.
- Piecewise-smooth ReLU objectives: A single vanishing preactivation with nonzero parameter gradient defines a regular hypersurface, and the one-sided parameter gradients have well-defined traces.
- Piecewise-smooth ReLU objectives: Fixing ReLU masks makes the network output polynomial on each activation-pattern cell, while continuity preserves matching risk traces across mask boundaries.
- Interface curvature: The distributional Hessian separates regional Hessian volume terms from an interface atom βνν⊤, locating curvature omitted by regional Hessian products.
- Sensitivity limits: Removing the atoms yields the regional product Jreg, which Theorem 1 identifies with the vanishing-step branchwise-AD limit.
- Sensitivity limits: A nonzero event jump produces a rank-one discrepancy, with hypergradient discrepancy vanishing exactly when β = 0 or gT ⊥Φ1ν.
- Convexity and event sign: Under global convexity, strict events attenuate the direct normal mode and prevent complete transported cancellation across a finite same-direction itinerary.
F GENERAL STATE AND FIRST-VARIATION THEOREM
Theorem 13 establishes state consistency for executed-branch Euler/GD and shows that its exact discrete derivatives converge to the event-free regional propagator, while the limiting flow derivative includes event transfers.
- Differentiability conditions: At nonresonant sizes, ordinary executed-branch AD is the classical derivative of the discrete endpoint map; exact interface landings can instead make the endpoint map nondifferentiable.Software AD still returns a selected branch product at an exact hit.
- State and derivative limits: The GD state converges over the fixed horizon, while exact nonresonant program derivatives converge to Jreg, the product of regional propagators.Under the stated rate assumption, the convergence estimate is O(η).
- State and derivative limits: The terminal flow map is C1 on a stable-itinerary neighborhood, and its derivative is obtained by interleaving regional propagation with event data.The event terms depend on the localized event state, normal, and one-sided fields.
- Event mechanism: A hard-branch numerical step has an I + O(η) Jacobian, so mixed event steps converge to identity rather than a finite event transfer.Thus state convergence does not imply convergence of tangent dynamics to the flow derivative.
G MULTI-EVENT CANCELLATION CRITERION
Multi-event discrepancies factor through transported saltation matrices: cancellation is possible, but a strict single gradient jump prevents equality, and convexity rules out complete cancellation under the stated conditions.
- Cancellation criterion: The flow and regional derivatives agree if and only if the transported product of event corrections equals the identity.Several nontrivial saltation matrices can cancel after regional transport.
- Volume and interaction bounds: Regional propagation can rotate and couple affected directions, so determinant equality alone does not imply matrix equality.This is why scalar volume criteria are weaker than the full matrix cancellation criterion.
- Volume and interaction bounds: The exact log-volume correction separates transported first-order event contributions from ordered multi-event interactions.The resulting expressions provide upper bounds; transported factors may still cancel.
- Single-event mismatch: For one event, a nonzero gradient jump forces an endpoint mismatch, with the discrepancy having rank one.The one-event criterion reduces equality to Ξ1 = I, equivalent to equality of the one-sided gradients.
- Parameters and adjoints: The same event-aware correction extends to parameter sensitivities and reverse-mode adjoints, using a transpose jump rather than an inverse transpose.Reverse-mode AD remains exact for each fixed nonresonant discrete program, but event terms are omitted from its regional limit.
- Method scope: The event-blind derivative-limit theorem includes Euler/GD but excludes event-localizing methods unless their separate hypotheses are verified.The corrected derivative belongs to an event-aware solver only when event localization and split-step differentiation are implemented.
M FINITE-SAMPLE RELU EMPIRICAL-RISK REALIZATION
The residual-ReLU squared-loss family provides an explicit one-event realization in which discrete derivatives converge to regional dynamics while flow sensitivity is multiplied by the saltation factor, including globally strongly convex cases.
- Curvature and convexity: The regional Hessians are H− = 1 and H+ = 2, while the interface coefficient is β = −c.For c ≤ 0, the objective is globally 1-strongly convex; for c > 0, the downward derivative jump proves global nonconvexity.
- Residual-ReLU construction: A residual-ReLU trajectory has one stable transverse crossing before T, with saltation factor r equal to the ratio of one-sided speeds.The initialization interval is exactly the condition 0 < t* < T, and the discrete itinerary also has one crossing at nonresonant sizes.
- Amplification: For every prescribed r > 0, parameters can fix the event time and minimum one-sided speed while enforcing DφT = rJreg.The reciprocal family attains arbitrarily large multiplicative gaps through r, while retaining a uniform transversality margin.
- Resonance: At resonance, an evaluated iterate can land exactly on the interface, so the endpoint map need not have a classical derivative even though software AD selects a branch product.This is why the derivative claims use nonresonant step-size sequences.
- Smoothing limits: Resolved smoothing recovers saltation in the scalar or autonomous-normal regime only when ρ → 0, R → ∞, and τR → 0, with fixed-window tail error retained.At fixed R, the resolved limit is F(R)/F(−R), not exactly the saltation factor.
- Smoothing limits: The scalar smoothing argument does not establish a general matrix or curved-interface theorem because noncommuting normal-tangential dynamics and geometric terms require separate control.The section explicitly excludes grazing, sliding, chattering, and simultaneous hits.
O NUMERICAL METHODS AND REPRODUCIBILITY
The numerical studies use deterministic, reproducible resolution sweeps with preserved array artifacts and validate endpoint, outer-gradient, and smoothing behavior under controlled nonresonant settings.
- Reproducibility: All numerical plots use deterministic resolution sweeps, with analytic curves recomputed from parameters and numerical values read from preserved JSON or NPZ arrays.No numerical value is recovered from raster images.
- Residual-ReLU experiments: The scalar residual-ReLU sweep uses 88 nonresonant rows across eight step sizes, with one crossing and minimum one-sided speed exactly one.The diagnostic event-aware derivative inserts the exact scalar event factor into the regional product.
- Residual-ReLU experiments: Maximum discrepancies are 3.82 × 10−14 for PyTorch endpoint AD and 4.11 × 10−13 for its outer derivative.Finite-difference endpoint and outer-derivative discrepancies are 3.66 × 10−9 and 3.37 × 10−8, respectively.
- Residual-ReLU experiments: The two-parameter study verifies one event at time 0.3, with one-sided normal speeds 1 and 1.9124409981 and β = −0.9124409981.It records state, initialization, threshold, event-aware, reverse-mode, finite-difference, and adjoint quantities for the same eight step sizes.
- Controls: The identity-Hessian control isolates persistent discrepancy from interface transfer rather than regional curvature.It is retained as a mechanism-isolating control, not the main learning realization.
- Smoothing experiment: The smoothing grid contains 625 float64 cells over η, τ ∈ [10−5, 10−1], with ρ ∈ [10−4, 10^4] and an R = 20 truncation floor of approximately 2.04 × 10−7.Figure 4 illustrates a sufficient fixed-profile resolution condition rather than a universal phase boundary.
- Scope and restrictions: The multidimensional sharp-interface results preserve tangential modes and multiply the normal mode by the normal-speed ratio, whereas smoothing extends directly only to an autonomous scalar normal coordinate.This is a scope distinction between the sharp-interface and smoothing results.