Source-linked AI summary

Quantitative Convergence of Wasserstein Gradient Flows of Kernel Mean Discrepancies

Lénaïc Chizat, Maria Colombo, Roberto Colombo, Xavier Fernández-Real

arXiv:2603.01977v1math.APcs.LGmath.OC

TL;DR

The paper studies well-posedness and long-time behavior for Wasserstein gradient flows of KMD functionals, including their quantitative convergence. It analyzes the equation across s ≥ 1 and extends the framework to shallow neural-network dynamics, obtaining global convergence for s = 1 and polynomial convergence for s > 1 under stated regularity conditions.

  • Problem

    The paper addresses the well-posedness and long-time behavior of Wasserstein gradient flows for Kernel Mean Discrepancy functionals.

  • Method

    The paper analyzes the KMD gradient-flow equation for all s ≥ 1 and considers its application to shallow neural-network dynamics on the sphere.

  • Results

    The flow has global convergence for s = 1, while for s > 1 it satisfies an explicit polynomial convergence bound with exponent depending on s and γ.

  • Takeaways & Limitations

    The analysis provides quantitative rates and extends convergence to higher-regularity topologies under suitable regularity of the data.

  • Takeaways & Limitations

    Exponential convergence requires a positive lower bound on the target ν, while convergence for unrestricted signed neural-network measures is more challenging.

Abstract

from arXiv · show

We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of shallow neural networks in the infinite-width and continuous time limit, as well as interacting particle systems with pairwise Riesz kernel interaction in the mean-field and overdamped limit. Our main analysis concerns the model case of KMD functionals given by the squared Sobolev distance $ \mathscr{E}^ν_{s}(μ)= \frac{1}{2}\lVert μ-ν\rVert_{\dot H^{-s}}^{2}$ for any $s\geq 1 $ and $ν$ a fixed probability measure on the $d$-dimensional torus. First, inspired by Yudovich theory for the $2d$-Euler equation, we establish existence and uniqueness in natural weak regularity classes. Next, we show that for $s=1$ the flow converges globally at an exponential rate under minimal assumptions, while for $s>1$ we prove local convergence at polynomial rates that depend explicitly on $s$ and on the Sobolev regularity of $μ$ and $ν$. These rates hold both at the energy level and in higher regularity classes and are tight for $ν$ uniform. We then consider the gradient flow of the population loss for shallow neural networks with ReLU activation, which can be cast as a Wasserstein--Fisher--Rao gradient flow on the space of nonnegative measures on the sphere $\mathbb{S}^d$. Exploiting a correspondence with the Sobolev energy case with $s=(d+3)/2$, we derive an explicit polynomial local convergence rate for this dynamics. Except for the special case $s=1$, even non-quantitative convergence was previously open in all these settings. We also include numerical experiments in dimension $d=1$ using both PDE and particle methods which illustrate our analysis.

1 Introduction

The paper develops well-posedness and quantitative convergence theory for Wasserstein gradient flows of KMD functionals, addressing settings where geodesic-convexity methods do not apply. It establishes global exponential convergence for s=1 and local polynomial convergence for s>1, with extensions to neural-network dynamics.

  • Motivation: KMD Wasserstein flows are motivated by shallow-neural-network training, generative modeling, and interacting particle systems.The evolving measure represents parameter or particle distributions, while the target measure defines the population-loss or interaction objective.
  • Problem: The paper addresses poorly understood well-posedness and long-time behavior because KMD energies are generally not geodesically convex in Wasserstein space.Therefore, standard contraction mechanisms for uniformly geodesically convex flows do not apply.
  • Well-posedness: For every s≥1, the authors identify natural weak solution classes and prove local well-posedness, uniqueness, regularity propagation, and quantitative stability.Global well-posedness follows above the semiconvexity threshold and also at the endpoint through log-Lipschitz velocity regularity.
  • Convergence for s=1: For s=1, bounded initial and target densities yield global convergence, including exponential weak convergence in energy and W2.The estimate is α^1/2 W2(µt,ν) ≤ ||µt−ν||_{Ḣ^-1} ≤ ||µ̄−ν||_{Ḣ^-1}e^-αt, while uniform convergence is exponential for Hölder targets.
  • Convergence results: The s=1 theory strengthens prior results by allowing bounded densities, proving unconditional qualitative convergence, and separating assumptions needed for energy versus stronger convergence.A positive lower bound on ν supports exponential convergence, whereas no lower bound on the initial measure is needed; discontinuous targets need not converge uniformly.
  • Convergence for s>1: For s>1, sufficiently regular flows satisfy the polynomial decay ||µt−ν||_{Ḣ^-s} ≤ ||µ̄−ν||_{Ḣ^-s}(1+t/C)^−(γ+s)/(2(s−1)).The rate extends through interpolation, requires a local small-discrepancy regime, and is sharp for uniform targets.

2 Well-posedness theory

This section establishes existence, uniqueness, and continuation for the active-scalar equation, including propagation of regularity up to maximal existence.

  • The analysis constructs a unique maximal solution and proves a continuation criterion for the active-scalar equation.The proof also establishes propagation of Hölder and Sobolev regularity from the initial and target data.

2.1 Preliminaries

The preliminaries define Riesz-kernel Sobolev energies, Wasserstein gradient flows, and weak active-scalar solutions on the torus, then connect these formulations.

  • The paper defines the Riesz kernel through the zero-mean fundamental solution of the fractional Laplacian on the torus.Convolution with this kernel represents the inverse fractional Laplacian and the homogeneous negative Sobolev norm.
  • The KMD energy is formulated on probability measures using the squared homogeneous negative Sobolev distance to a target measure.The functional is lower semicontinuous in the Wasserstein metric.
  • Wasserstein gradient flows are represented by absolutely continuous curves satisfying a continuity equation with a tangent velocity field.The weak formulation uses time continuity in the weak-* topology and distributional testing against smooth functions.
  • The natural weak class is chosen so the induced velocity field has a Lipschitz or log-Lipschitz modulus of continuity.This regularity supports the Lagrangian and Wasserstein formulations of the dynamics.
  • For measures in the local well-posedness class, the velocity field -∇K_s ∗ (μ_t − ν) belongs to the energy subdifferential.Consequently, weak solutions coincide with Wasserstein gradient flows and satisfy the energy dissipation identity.

2.2 Local well-posedness in the weak class

The paper proves local well-posedness in the weak class using stability and Lagrangian iteration, then extends solutions maximally with an explicit continuation criterion.

  • Solutions are quantitatively stable in Wasserstein distance under perturbations of initial and target measures.The detachment rate is exponential away from critical indices and can be double-exponential in critical cases.
  • Existence is obtained by Picard iteration at the Lagrangian flow level and passage to a weak-* limit of approximate solutions.Uniform bounds and compactness yield a solution on a sufficiently short time interval.
  • For every s ≥ 1 and admissible data, a unique maximal solution exists in the class X_s(T^d).The maximal solution is constructed by combining local existence, stability, and continuation.
  • When s ≥ d/2 + 1, the maximal existence time is infinite.For s ∈ [1, d/2 + 1), finite-time breakdown occurs exactly when the specified L^p norm diverges.

2.3 Propagation of regularity

Regularity estimates propagate Hölder and Sobolev smoothness from the data throughout the maximal existence interval, with smooth data producing smooth solutions.

  • Hölder regularity propagates from the initial and target data to the solution up to maximal existence.The argument uses flow-map representations, composition estimates, and Grönwall inequalities.
  • More generally, regularities stronger than X_s(T^d), including suitable L^r and C^k,1 classes, can be propagated.The stated extensions depend on the relation between s and the integrability threshold.
  • For smooth data, the solution, velocity field, and flow map remain smooth on the full maximal interval.This follows by bootstrapping regularity through the transport equation and inverse flow.
  • Sobolev energy estimates provide the regularity bounds needed for the paper’s smooth convergence results.These estimates control homogeneous Sobolev norms and propagate Sobolev regularity.

3 Quantitative convergence results

The section establishes quantitative convergence for Riesz kernel mean-discrepancy flows, with a structurally stronger global theory at s=1 and convergence mechanisms that extend under weaker initial assumptions.

  • The case s = 1: The Coulomb case s = 1 has a maximum principle that yields global-in-time bounded solutions.The density remains between the initial and target essential extrema.
  • The case s = 1: min{inf ¯µ, inf ν} ≤ inf µ_t ≤ sup µ_t ≤ max{sup ¯µ, sup ν} for all t ≥ 0.This estimate is the maximum principle used to prevent blow-up and establish global existence.
  • The case s = 1: Exponential decay follows by combining the energy dissipation identity with a differential inequality for the Sobolev discrepancy.The resulting estimate is explicitly exponential in time.
  • Removing the initial lower bound: Exponential weak convergence requires only a positive lower bound on the target measure, not on the initial measure.Regions where the initial density vanishes are filled exponentially fast when the target is bounded below.

4 Convergence for continuous shallow neural networks

The paper reduces continuous ReLU-network training dynamics to Wasserstein–Fisher–Rao evolution on the sphere and connects its kernel operator to Sobolev regularity.

  • Reduction to spherical dynamics: ReLU dynamics reduce to a single-species nonlocal continuity equation on the sphere under symmetry and regularity assumptions.The reduction trades some expressivity for dynamics on even functions and nonnegative measures on S^d.
  • Reduction to spherical dynamics: The reduction incurs a small loss of expressivity: sufficiently regular even functions are represented only up to a controlled constant.The representable function class is tied to the even part of the spherical target.
  • Spectral correspondence: For s = (d + 3)/2, the operator H is a continuous bijection between even Sobolev spaces whose regularity indices differ by s.This correspondence supplies the Sobolev interpretation of the arccos-kernel dynamics.
  • Spectral correspondence: The arccos-kernel operator has eigenvalues λ_k approximately proportional to k^-2s on even frequencies.Its multiplier structure underlies the correspondence with the Sobolev energy.
  • Well-posedness and regularity: The spherical Wasserstein–Fisher–Rao flow admits a unique global solution, stability under weak convergence of data, and propagation of regularity.These properties hold in the stated measure and Hölder/Sobolev settings.
  • Quantitative convergence: The paper derives higher-order energy estimates for the spherical dynamics, providing the analytic basis for quantitative convergence.The estimates are stated for even nonnegative data with γ ≥ s − 1 in the integer-regularity setting.

Appendix

The appendix collects Riesz-kernel estimates and extends fractional Kato–Ponce commutator estimates to the periodic setting.

  • Appendix: The appendix has two parts: estimates for periodic Riesz kernels and fractional Kato–Ponce commutator estimates on the torus.These estimates support the regularity and convergence arguments in the main text.

A.1 Riesz Kernel estimates

The appendix develops kernel estimates that control Riesz convolutions, their continuity, and their dependence on Wasserstein perturbations across the regimes s ≥ 1.

  • Kernel asymptotics: Periodic Riesz kernels have distinct short-distance regimes depending on s, including logarithmic corrections at critical exponents.The estimates are obtained by splitting the analysis according to the value of s.
  • Convolution regularity: The gradient of a Riesz convolution has a modulus of continuity controlled by the X_s norm of the convolved measure.The modulus is expressed through a kernel-dependent function ω_s.
  • Wasserstein estimates: For s below the critical threshold, Wasserstein control of gradient differences requires L^p integrability of the endpoint measures.The relevant exponent is p = d/(2s − 2), with separate critical and supercritical estimates.
  • Wasserstein estimates: The estimates use displacement interpolation between measures and bounds on the interpolating densities along Wasserstein geodesics.This converts transport distance into convolution-gradient control.
  • Critical and supercritical regimes: Critical convolution bounds involve logarithmic or stronger-norm corrections, while the supercritical regime admits bounded second derivatives of the kernel.The proof treats s = d/2 + 1 and s > d/2 + 1 separately.

A.2 Kato–Ponce Commutator estimates on the d-dimensional torus

This section establishes periodic Kato–Ponce fractional commutator estimates by reducing the torus problem to Euclidean estimates through Poisson summation. A periodic localization argument controls the resulting terms and yields interpolation refinements.

  • Poisson’s summation formula converts the periodic operator into a Euclidean fractional operator, allowing the use of the corresponding estimate from Li19.The Euclidean operator is identified as (−∆_R^d)^(γ/2), with its kernel representation used in the estimates.
  • The main lemma states a fractional commutator estimate on the d-dimensional torus for γ > 0 and paired exponents satisfying 1/2 = 1/p_i + 1/q_i.
  • The proof assumes zero averages, extends the functions periodically to R^d, and applies a periodic partition of unity with localized functions G_j.The localization uses cutoff functions supported on cubes and equal to one on smaller cubes.
  • The localized proof separates the estimate into terms (I) and (II), bounds them using cutoff support, Hölder’s inequality, kernel decay, and summation over lattice shifts, then combines the bounds to conclude (A.8).
  • For γ ∈ (0, 1], the section records refinements of (A.8), and interpolation yields further consequences for smooth periodic functions.These consequences use fractional Gagliardo–Nirenberg inequalities.
Loading 2603.01977v1…