Source-linked AI summary
Nonparametric inference for density-dependent McKean--Vlasov diffusions
Denis Belomestny, Ekaterina Morozova
TL;DR
The paper addresses nonparametric estimation of a density-dependent drift and stationary density from one-time observations in multivariate McKean–Vlasov diffusions. It reduces the problem to one dimension and uses a constrained sparse ReQU sieve maximum-likelihood estimator with endpoint-adapted approximation. The resulting rates are accompanied by a matching Assouad lower bound up to logarithmic factors, with finite-time extension under quantitative convergence to equilibrium.
Problem
The paper asks whether a density-dependent drift and its induced stationary density can be recovered from independent particles observed once at a common time.
Method
It reduces the stationary likelihood to a one-dimensional problem and constructs a constrained sparse ReQU sieve maximum-likelihood estimator with endpoint-adapted approximation.
Results
A matching Assouad lower bound shows the drift-coefficient rate is minimax optimal up to logarithmic factors, with finite-time guarantees also established.
Takeaways & Limitations
The coefficient rate is minimax optimal up to logarithmic factors, and the stationary guarantees extend to finite-time observations under quantitative convergence to equilibrium.
Takeaways & Limitations
Stationary observations identify the coefficient only on the density range (0,a0], and inverse stability requires a regular interior condition satisfied automatically when the potential has no critical values above its minimum.
Abstract
from arXiv · showhide
The present research is devoted to the nonparametric estimation of a density-dependent drift coefficient in a multivariate McKean--Vlasov diffusion from independent observations at a common time, as well as the stationary density. Under certain assumptions on the (known) potential, we reduce the problem to the one-dimensional one and construct a sieve maximum-likelihood estimator based on sparse ReQU neural networks subject to structural and Hölder constraints. Using the endpoint-adapted graded approximation, we achieve the rate of $\left(b_n\log n/n\right)^{2(β+1)/(2β+3)}$ for the Kullback-Leibler divergence between the true and estimated stationary densities, with $b_n$ being at most a logarithmic factor. Similarly, it is shown that the constructed estimator for the drift coefficient converges to the true one at the rate of $\left(b_n\log n/n\right)^{β/(2β+3)}$ in the $L^2$-metric. A matching Assouad lower bound proves minimax optimality of this bound up to logarithmic factors.
1 Introduction
The paper studies nonparametric recovery of a density-dependent drift and stationary density from one-time cross-sectional observations, reducing the multivariate problem to one dimension. It develops a constrained ReQU sieve estimator with endpoint-adapted approximation and establishes rates, lower bounds, and a finite-time extension.
- Motivation: The problem is recovering the density-dependent function Ξ and induced stationary density πΞ from independent particles observed once at a common time.The dynamics are nonlinear and d-dimensional, while the data are cross-sectional.
- Motivation: Pointwise density dependence creates a singular statistical problem because the drift is not continuous under weak perturbations of the law.
- Reduction: The stationary representation reduces the likelihood to the scalar energy V(X), turning estimation into a one-dimensional nonlinear inverse problem with tail ill-posedness.The coefficient is identifiable only over the range attained by the stationary density.
- Method: The paper constructs a constrained ReQU sieve maximum-likelihood estimator and develops endpoint-adapted approximation for pointwise density dependence.This extends the statistical theory to a broader class of confining potentials.
- Optimality and extension: A matching Assouad lower bound establishes minimax optimality of the coefficient rate up to logarithmic factors, and finite-time guarantees require T_n ≍ log n for the Ornstein-Uhlenbeck model.
2 Literature Review
The literature includes parametric, semiparametric, online, likelihood-based, and nonparametric inference for McKean–Vlasov systems, with prior work focused mainly on smoother law functionals or trajectory and interacting-particle observations.
- Existing inference: Existing statistical work covers maximum likelihood, semiparametric and parametric methods, online estimation, and nonparametric inference for McKean–Vlasov systems.
- Closest comparison: The closest likelihood-based study concerns interaction through the cumulative distribution function, whereas this paper addresses pointwise density dependence.The pointwise model requires endpoint-adapted approximation and separate inverse-stability analysis.
3 Model, stationary representation, and estimation idea
Under structural assumptions on the known confining potential and coefficient class, the paper derives a unique zero-flux stationary density and reduces likelihood-based estimation to scalar energy observations. Stationary data identify the coefficient only on an interior density range.
- Potential assumptions: The known potential is assumed to be C2, coercive, exponentially integrable, and regular in its level-set measure away from critical values.
- Potential assumptions: The Ornstein-Uhlenbeck potential satisfies the potential assumptions and is a principal example in the considered model class.
- Coefficient assumptions: The coefficient class is bounded between positive constants, while the true coefficient has Hölder regularity and remains separated from the sieve envelope boundary.
- Stationary representation: Under the stated assumptions, each admissible coefficient has a unique C1 zero-flux stationary density.
- Identifiability: Stationary observations identify Ξ0 only on the density range (0,a0], so recovery is restricted to an interior interval I contained in that range.The limit r→0 corresponds to spatial tails, while r→a0 corresponds to the minimum energy level.
- Likelihood reduction: If X follows πξ, the energy S=V(X) has density qξ(s)mV(s), and the likelihood based on X is equivalent to the likelihood based on the scalar energies Si=V(Xi).For each candidate, the scalar profile equation is solved and evaluated at observed energies.
4 Neural-Network Sieve and Graded Approximation
The estimator uses a structurally constrained sparse ReQU neural-network sieve, while graded meshes adapt approximation to endpoint behavior in the forward operator. Exact moment cancellations yield uniform coefficient and forward approximation bounds.
- Sieve construction: The sieve domain covers all empirically reachable densities and extends candidate coefficients constantly beyond its upper boundary to preserve global coefficient bounds.
- Sieve construction: The sieve uses fixed-depth ReQU networks with growing widths, sparsity, and bounded parameters, together with a Lipschitz clamp enforcing κ≤ξ≤K.
- Approximation target: The likelihood measures approximation through induced stationary densities, so analysis separates forward KL approximation from inverse control of coefficient error.
- Graded approximation: Endpoint behavior of the forward measure and the factor u^-1 in the inverse operator determine the need for a graded approximation mesh.For the Ornstein-Uhlenbeck potential, the measure contributes at most a logarithmic factor near zero.
- Graded approximation: A uniform mesh yields only order m^-(β+1/2), up to logarithmic factors, motivating graded knots aligned with the endpoint and the anchor 1.
- Approximation guarantee: The graded sieve admits candidates with exact cellwise moment cancellations, uniform coefficient error O(m^-β), and forward error O(m^-(β+1)).The construction also controls sieve entropy through fixed architecture constants.
5 Fast Oracle Inequality for the Stationary Density
The section develops a localized likelihood analysis on a truncated domain and transfers the resulting oracle inequality to the full stationary experiment. Truncation, sieve complexity, discretization, and optimization errors are balanced to obtain the leading stochastic rate.
- Localized likelihood control: The truncated analysis controls centered log-likelihood fluctuations uniformly over a finite net of the sieve.The resulting entropy contribution is of order BγM/n.
- Localized likelihood control: The localization argument yields a fast 1/n stochastic order rather than 1/√n.The first likelihood fluctuation term is absorbed into the excess risk.
- Oracle inequality: The oracle inequality separates sieve approximation, stochastic complexity, discretization, and numerical optimization errors.The truncation level and net resolution are selected only after transferring the bound to the full likelihood.
- Full-likelihood transfer: The full-MLE oracle inequality holds with probability at least (1 − δ)(1 − τ0(M))^n, while accounting for normalization, tail, and escape effects.On the no-escape event, full and truncated likelihoods differ only through normalizing constants.
- Full-likelihood transfer: Choosing Mn proportional to log n makes truncation and escape-probability terms negligible up to logarithmic factors.The leading stochastic term is of order m log n/n in the compact case and m(log n)^2/n in the non-compact case.
6 Upper Rates: Exact Forward Approximation and Nonlinear Inversion
The section analyzes the forward map from the drift coefficient to the stationary density and the inverse map needed for coefficient recovery. Graded approximation gives a polynomial KL bound, while interior regularity enables nonlinear inversion and finite-time transfer.
- Forward and inverse analysis: The MLE analysis reduces to forward approximation of stationary densities and inverse control of coefficients from density closeness.These are the two sides of the nonlinear map ξ ↦ πξ.
- Forward and inverse analysis: The forward map smooths coefficient perturbations by one integration, whereas inversion differentiates an inverse profile and is unstable without regularity.Consequently, coefficient recovery is governed by a Hölder interpolation modulus rather than a uniform linear inverse inequality.
- Exact forward approximation: KL(π0∥πξm) ≤ C_KL m^−2(β+1) for the graded approximant under the stated assumptions.The constant is independent of m and depends on the model and approximation constants.
- Density rates: Theorem 6.2 uses a graded ReQU sieve with Mn = v★ + Ctr log n and bn = O(1) in the compact case or bn ≤ O(log n) in the non-compact case.The sieve dimension is selected at the balance m ≍ (n/(bn log n))^1/(2β+3).
- Finite-time extension: Finite-time observations retain the stationary rates when exponential ergodicity makes nηTn and nτ0(Mn) vanish.A sufficient choice is Tn ≥ t0 + (1 + ε)ρχ^−1 log n, together with Mn proportional to log n.
- Nonlinear inverse stability: A uniform direct inverse bound cannot hold for free-knot ReQU networks because unit count does not control their smallest spatial scale.The proof therefore uses interpolation control of the inverse profile derivative.
- Nonlinear inverse stability: Coefficient recovery requires an interior regularity condition placing the relevant energy range in a compact interval away from critical values.For potentials without critical values above v★, including the Ornstein-Uhlenbeck potential, this condition is automatic.
- Nonlinear inverse stability: Under the interior condition, coefficient error is bounded by KL(π0∥πξ)^{β/(2(β+1))}, and optimizing the sieve dimension yields the stated coefficient rate.The resulting theorem controls the L2 error on an interior interval I.
7 Minimax lower bounds for estimating Ξ
The paper establishes a minimax lower bound for estimating the density-dependent coefficient by constructing localized alternatives that are separated in L2 but statistically close. The result extends to finite-time observations under quantitative convergence to equilibrium.
- Lower-bound result: The coefficient rate is minimax optimal up to logarithmic factors under the stated local regularity and stability assumptions.The lower-bound construction uses a local parameter class around an interior coefficient and an Assouad cube.
- Assouad construction: The alternatives perturb the primitive coordinate with disjoint, zero-mean smooth bumps, producing additive L2 separation while preserving the coefficient class.The perturbation amplitude is chosen of order b^β, with bandwidth b and M approximately b^-1 intervals.
- Assouad construction: Choosing b approximately n^-1/(2β+3) keeps neighboring n-sample KL divergences bounded and yields the minimax lower bound.Assouad’s lemma converts the bounded neighboring divergence and coefficient separation into the stated squared L2 risk rate.
- Theorem: The theorem takes the infimum over all measurable estimators based on stationary observations X_1, ..., X_n distributed according to Π_ξ.The constants in the theorem may depend on the model and parameter-class quantities but not on n.
- Finite-time extension: Finite-time transfer of the lower bound requires a uniform comparison over the entire Assouad cube between finite-time and stationary laws.The stationary-experiment result is therefore subject to an additional uniform finite-time approximation requirement.
8 Simulation study
The simulation study evaluates the ReQU-sieve estimator for several drift coefficients under the Ornstein–Uhlenbeck potential. Estimates are close to the truth, errors decrease with sample size, and observed rates depend on sieve approximation.
- Estimator construction: The sieve enforces bounded network outputs through a clamping construction, with a fixed-depth sparse ReQU architecture whose widths and parameter counts may grow with the sieve index.The network parametrization is designed to preserve structural constraints on the coefficient.
- Numerical results: The estimated and true curves are fairly close, with errors not exceeding 0.05 in the n=10000, d=2 experiment.Table 1 reports ISE values not exceeding 0.02 even for the smallest sample size, averaged over 100 repetitions.
- Numerical results: For Ξ0,1, the error decays approximately at the n^-1 rate, while Ξ0,3 converges noticeably more slowly because it is not contained in the sieve.The Ξ0,2 rate is less evident in the plotted results.
A.1 Proof of Proposition 3.4
The proof constructs the stationary density from a normalizing constant and establishes uniqueness of the positive zero-flux stationary solution. The argument uses monotonicity of the normalization map and connectedness of the state space.
- Existence: The map F_ξ is finite, continuous, and strictly increasing, with limits 0 and infinity at the two ends of its argument range.These properties imply a unique μ_ξ satisfying F_ξ(μ_ξ)=1.
- Existence: The resulting density π_ξ is positive and satisfies the zero-flux equation by differentiation.Positivity follows from the range of g_ξ^-1 and the stationary representation.
- Uniqueness: Any nonnegative C1 zero-flux stationary density must be strictly positive, since a zero at one point propagates along paths through the zero-flux equation.The path argument uses boundedness of ξ and continuity of the potential-gradient term.
- Uniqueness: Rewriting the zero-flux condition as ∇(g_ξ(eπ)+V)=0 shows that every stationary solution differs only through a constant normalizing parameter.Normalization forces that parameter to equal μ_ξ, so eπ=π_ξ.
A.2 Proof of Lemma 4.2
The approximation proof builds a globally Hölder corrected spline on an endpoint-adapted graded mesh and cancels primitive-operator errors at the knots. It then verifies approximation accuracy and realizability by sparse ReQU networks.
- Spline approximation: The graded approximation uses blended local Taylor polynomials whose derivatives through order k agree at every knot.Endpoint-adapted blending functions vanish to order k+1 at the relevant cell endpoints.
- Spline correction: Correction bubbles are selected to enforce primitive-operator cancellation at each knot while retaining uniformly bounded coefficients.The construction uses one nonzero bubble per cell and controls the coefficients across the graded mesh.
- Regularity and constraints: The corrected spline remains globally β-Hölder, stays inside the structural envelope, and satisfies the required interior constraints for sufficiently large m.Its sup-norm approximation error is of order m^-β, so clamping does not alter the spline or the cancellations.
- Error control: The forward approximation error is bounded by m^-(β+1) in L2(ν0) after combining the endpoint and interior regions.The proof obtains the same order for the associated operator term through control of the normalization perturbation.
- Neural-network realization: The approximant is piecewise polynomial of fixed degree and has a fixed-depth ReQU realization with O(m) nonzero parameters.The rescaling weights are polynomially bounded in m, placing the approximant in the prescribed sieve.
A.3 Proof of Lemma 5.1
The proof relates a likelihood-ratio expression to KL divergence and then controls the relevant empirical deviations with high probability.
- KL(p_A^0 ∥ p) equals the expectation of the likelihood-ratio transform Z_p_A(Y_1).
- A quadratic logarithmic bound converts the likelihood-ratio expression into a KL-based control under the ratio constraint γ ≤ r ≤ γ^-1.
- 2|𝒫_A|e^-t bounds the probability that at least one candidate in the finite class fails.
A.4 Proof of Proposition 5.2
The proposition extends finite-class likelihood analysis from truncated observations to the original setting, while establishing approximation, tail, and stability bounds used in the estimator analysis.
- Truncation extension: The non-truncated analysis conditions on all observations lying in A_M, where the conditional observations are i.i.d. with truncated invariant density π_M^0.
- Truncation extension: With probability at least 1 − δ under ℙ_M, the truncated bound also holds jointly with E_M with probability at least (1 − δ)(1 − τ_0(M))^n.
- Approximation control: sup_{0≤t≤1} ||G_m||_{L2(v_t)} ≤ C_path m^-(β+1), with C_path depending on κ, K, C_app, and C_fwd.
- Estimator rate: Choosing m ≍ (n/(b_n log n))^(1/(2β+3)) yields the stated KL-divergence bound with probability at least (1 − δ)(1 − τ_0(M_n))^n.
- Stability control: Candidate perturbations satisfy stability bounds proportional to (1 + M)||ξ − ζ||∞, including KL-related log-density control.
- Tail control: The truncation tail quantities d_m(M) and R̄_m(M) decay exponentially in M under Assumptions 3.1 and 3.3.Specifically, d_m(M) ≲ C_τe^-c_aM and R̄_m(M) ≤ C_R,εe^-(Ξ_0(0)−ε)M for sufficiently large M.