Source-linked AI summary

Geometric Ergodicity of Affine Invariant Ensemble Langevin and its Discrete Time Variants

Hong Ye Tan, Yifan Chen

arXiv:2609.03326v1math.STmath.NAmath.OC

TL;DR

Affine-invariant ensemble Langevin samplers lack quantitative convergence theory in the presence of possible covariance singularity. This paper proves geometric ergodicity using an inverse-covariance Lyapunov control, identifies divergence of direct Euler–Maruyama, and introduces covariance-trace regularization that restores geometric ergodicity in continuous and discrete time.

  • Problem

    Geometric ergodicity for affine-invariant ensemble samplers remained open because empirical covariance degeneracy creates a central analytical obstacle.

  • Method

    The paper combines an inverse-covariance barrier with coercive energy control and introduces covariance-trace time regularization for stable discretization.

  • Results

    The continuous ensemble dynamics, regularized diffusion, and sufficiently small-step regularized Euler–Maruyama discretization are geometrically ergodic, while direct Euler can diverge with positive probability.

  • Takeaways & Limitations

    Covariance-trace regularization restores geometric ergodicity in continuous and discrete time, although it sacrifices scale invariance.

Abstract

from arXiv · show

Affine-invariant ensemble samplers are widely used in Bayesian applications. However, their quantitative convergence theory, in particular geometric ergodicity, remains a basic open question. We study the affine invariant ensemble Langevin dynamics, an interacting particle system that uses the empirical covariance of the whole ensemble as a preconditioner. While effective in practice, theoretical understanding of this method is not available beyond plain qualitative convergence in total variation; a central difficulty is that the empirical covariance can approach singularity. This paper addresses this challenge. For potentials with bounded Hessian that are strongly convex outside a ball, we prove geometric ergodicity using a novel Lyapunov function that combines an inverse-covariance barrier with a coercive exponential energy. We then show that directly applying the Euler--Maruyama scheme can diverge with positive probability, even for a one-dimensional Gaussian target. This motivates a covariance-trace time regularization. We prove geometric ergodicity of the regularized diffusion and, for sufficiently small step size, of its unadjusted Euler--Maruyama discretization. We also show that the invariant distributions of the discretization converge weakly to the product target distribution as the step size tends to zero.

1. Introduction

This paper develops quantitative convergence theory for affine-invariant ensemble Langevin dynamics, addressing covariance degeneracy and instability in direct discretizations. It proves geometric ergodicity for continuous and covariance-regularized discrete dynamics under stated assumptions.

  • Method: The affine invariant ensemble dynamics use empirical covariance as an ensemble-adapted preconditioner, with a correction term preserving the product target π⊗N.The finite-ensemble correction appears explicitly in the dynamics and ensures invariance.
  • Motivation: Geometric ergodicity was previously open because empirical covariance can become singular, weakening noise near the boundary of the admissible state space.Earlier work established qualitative total-variation convergence but no quantitative rate.
  • Continuous-time results: Under Lipschitz-gradient and outside-ball strong-convexity assumptions, geometric ergodicity holds when N ≥ d + 3.The proof uses a Foster–Lyapunov function combining an inverse-covariance term with an exponential energy term.
  • Discrete-time instability: For d = 1 and f(x) = x^2/2, direct Euler–Maruyama diverges with positive probability for any η > 0 when the initial empirical variance is sufficiently large.The empirical covariance makes the drift cubic in the ensemble, allowing the covariance to grow without bound when noise remains small enough.
  • Regularization: A covariance-trace time regularization bounds γC for large covariance and restores geometric ergodicity for both the regularized diffusion and sufficiently small-step Euler–Maruyama discretizations.The regularization preserves π⊗N as invariant for the diffusion, while sacrificing general affine invariance.

2. Continuous time: geometric ergodicity

The paper upgrades qualitative ergodicity to geometric ergodicity for affine-invariant ensemble Langevin dynamics by controlling covariance collapse with an inverse-covariance Lyapunov construction.

  • Existing ergodicity result: The existing analysis establishes ergodicity in total variation through non-explosiveness, ellipticity, recurrence, and Harris recurrence, but not geometric ergodicity.The ensemble state space excludes singular covariance, making boundary behavior central to the analysis.
  • Log determinant is insufficient: The log-determinant Lyapunov function fails geometrically because its generator remains bounded below while covariance approaches singularity and the function diverges.This shows that positive powers of covariance cannot provide the required negative drift near the boundary.
  • New Lyapunov function: The new Lyapunov function combines an inverse-covariance barrier with a coercive exponential energy to control both covariance collapse and escape to infinity.The inverse-covariance component supplies boundary control, while coercivity makes the function norm-like at large ensemble states.
  • Geometric ergodicity: Under bounded Hessian and distant strong convexity, the paper proves a Foster–Lyapunov condition for N ≥ d + 3 and sufficiently small a > 0.The condition holds outside a compact set and therefore yields geometric ergodicity.
  • Geometric ergodicity: For N ≥ d + 3, the affine-invariant ensemble dynamics are geometrically ergodic under the stated distant-convexity assumptions.A broader tail condition is also identified, including potentials such as f(x) = ∥x∥4; in one dimension, an alternative Lyapunov function suggests N ≥ d + 2 may suffice.

3. Discrete time divergence and geometric ergodicity of a regularized scheme

Direct Euler–Maruyama discretizations can diverge with positive probability because covariance-dependent drifts become unstable, while trace-based regularization restores geometric ergodicity for sufficiently small steps.

  • Direct Euler–Maruyama divergence: For every positive step size, direct Euler–Maruyama can diverge with positive probability for a one-dimensional Gaussian when the initial empirical variance is sufficiently large.The mechanism is a cubic ensemble drift that can drive covariance to infinity while noise remains sufficiently small.
  • Leave-one-out divergence: The leave-one-out discretization also diverges with positive probability because its covariance preconditioning can produce the same covariance-exploding mechanism.For N ≥3, a two-step recurrence couples particle effects back to themselves and drives covariance toward infinity.
  • Covariance-trace regularization: Trace-based time regularization scales both drift and diffusion to cap the large-covariance regime while preserving the target distribution through a modified correction term.The factor γ(Tr C) stays near one for small covariance and keeps γC bounded for large covariance.
  • Scope of regularization: The regularized dynamics retain translation and orthogonal invariance but lose invariance under general affine maps.The regularization addresses large covariance, while covariance collapse near the boundary remains a separate analytical issue.
  • Continuous-time regularized dynamics: The regularized continuous-time diffusion is geometrically ergodic with stationary distribution Π∗ under the stated potential and particle-number assumptions.The proof uses the same Foster–Lyapunov function as the original continuous-time analysis, with N ≥ d + 3.
  • Discrete-time geometric ergodicity: For sufficiently small step size, the regularized Euler–Maruyama discretization converges geometrically in total variation to a unique invariant distribution Π∗,η.The invariant distributions satisfy Π∗,η ⇀ Π∗ as η → 0.

4. Proofs

The proofs establish geometric ergodicity through Foster–Lyapunov drift arguments that control both large ensemble spread and covariance degeneracy. The same framework is extended to the covariance-trace-regularized diffusion and its discrete-time kernel, with weak convergence of invariant distributions as the step size vanishes.

  • Continuous-time proofs: The Lyapunov analysis combines energy, trace, and inverse-trace controls to obtain compact sublevel sets away from the singular covariance boundary.The compact sets have the form K = {T ≤ T0, J ≤ J0, E ≤ E0}.
  • Continuous-time proofs: A norm-like Lyapunov function satisfying a Foster–Lyapunov drift condition yields geometric ergodicity for the affine invariant ensemble diffusion.Positive recurrence, strong Feller properties, and irreducibility complete the argument.
  • Discrete-time proofs: The discrete-time proof establishes a uniform Foster–Lyapunov bound for sufficiently small step size, with negative drift outside a compact set.The argument uses a uniform first-order expansion of the inverse covariance map and a change-of-measure bound for Gaussian increments.
  • Discrete-time proofs: The invariant distributions of the regularized Euler–Maruyama discretization converge weakly to the product target distribution as η → 0.The proof uses the bound ∥γC∥≤θ−1 and identifies every subsequential limit with Π∗.

5. Discussion

The discussion summarizes the paper’s continuous- and discrete-time results and identifies the principal scope boundary and future directions. The methods address covariance degeneracy broadly, but the regularization sacrifices scale invariance and the threshold N ≥ d + 3 excludes a borderline case.

  • Main conclusions: The paper proves geometric ergodicity for the continuous finite-particle dynamics using a Lyapunov function controlling both escape to infinity and covariance collapse.At the discrete level, direct Euler–Maruyama can diverge with positive probability, while covariance-trace regularization restores geometric ergodicity.
  • Scope and future directions: The common threshold N ≥ d + 3 arises from the inverse-trace coefficient, while N = d + 2 requires a different boundary weight.Other open directions include weaker assumptions and a stable explicit discretization retaining full affine invariance.
  • Broader relevance: The techniques are expected to apply to other affine-invariant sampling and data-assimilation methods facing covariance degeneracy.The paper notes that finite-particle geometric ergodicity is not yet available for those methods.

Appendix B. Supporting proofs

The appendix supplies the recurrence and irreducibility steps needed to turn the Lyapunov estimates into Harris ergodicity. It uses petite compact sets and almost-sure returns to establish Harris recurrence.

  • Recurrence argument: The petite-set return property implies Harris recurrence, completing the ergodicity proof framework.This step transfers the earlier recurrence and irreducibility results to Harris recurrence.
  • Recurrence argument: The diffusion is Feller and Lebesgue-irreducible, and a skeleton chain makes compact sets petite.These properties provide the local structure needed for the recurrence argument.
  • Recurrence argument: Positive recurrence implies that a compact set is hit with probability one from every starting point.The selected compact set is then petite with almost-surely finite hitting time.
  • Auxiliary estimates: The appendix also begins a Gaussian quadratic-form bound used in the supporting estimates.The displayed subsection introduces a one-dimensional Gaussian representation before deriving the bound.

B.2. Proofs in Section 3.2.

This appendix section proves elementary bounds for quadratic sublevel sets and records representations of leave-one-out covariances. These estimates support the analysis of Gaussian updates and covariance behavior.

  • Quadratic sublevel bounds: The sublevel set defined by a quadratic polynomial is bounded by completing the square and splitting into cases according to the constant term.The argument treats the regimes −u < c < u, c ≥ u, and c ≤ −u separately.
  • Quadratic sublevel bounds: The resulting interval-length estimate combines with a Gaussian density upper bound to control probabilities of quadratic constraints.The case analysis yields the desired inequality after applying the density bound.
  • Leave-one-out covariances: The appendix represents pairwise sums of leave-one-out covariances in terms of the full empirical covariance and centered particle deviations.This representation directly yields the first inequality used in the leave-one-out analysis.

B.3. Proof of Lemma 2.

The proof bounds trace terms by representing them along line segments and applying matrix-norm inequalities. A Frobenius inner-product Cauchy–Schwarz step supplies the final lower bound.

  • Hölder’s inequality on matrix norms bounds the trace terms appearing in the argument.
  • Line-segment representations connect the relevant quantities to symmetric matrices H_i with ||H_i||≤L.
  • The lower bound T_J≥d^2 follows from Cauchy–Schwarz for the Frobenius inner product.
  • Combining the preceding inequalities yields the claimed result.

B.4. Proof of Proposition 3.

The proposition establishes constants controlling the relevant quantities under the stated assumption. Its proof derives coercivity, gradient growth, and trace bounds using strong convexity outside a ball and covariance estimates.

  • There exist positive constants c1, b1, c2, b2, C_T, and B_T depending only on f and d.
  • On the set {J≤J0}, the remaining estimate follows by summing the gradient bound and using g^T Cg≥J^-1.
  • Strong convexity outside a sufficiently large ball yields the global lower bound f(x)≥c1||x||^2−b1.
  • A bounded Hessian gives quadratic upper growth, which combines with the lower bound to produce ||∇f||^2≥c2f(x)−b2.
  • The trace lower bound follows from monotonicity along line segments and the fact that their intersection with the ball has length at most 2R.

B.5. Proof of Lemma 4.

The proof computes the generator of the chosen function by differentiating the ensemble covariance and its inverse. It then combines the drift and diffusion contributions.

  • Writing δ_i=x_i−x̄ expresses the covariance in centered ensemble coordinates.
  • The first-order gradient determines the drift component of the generator.
  • Differentiating the covariance with respect to ensemble coordinates and then differentiating its inverse gives the required second-order terms.
  • Summing the coordinate contributions produces the diffusion component and the generator applied to the function.

B.6. Proof of Lemma 6.

The lemma controls random matrix perturbations uniformly by combining compactness, minimum-singular-value estimates, and a first-order expansion. Gaussian small-ball bounds provide the needed moment control near singularity.

  • For n=N−1≥d+2, the argument works on a compact set of bounded centered matrices with γI_d⪯R⪯2γI_d.
  • For sufficiently small η, boundedness of U keeps the singular values of Z+ηU away from zero.
  • The minimum singular value satisfies P(σ_min(Λ_η)≤s)≲s^(N−d), yielding finite moments when N−d>2q.
  • Conditioning on all but one column and projecting onto the orthogonal complement controls the probability of small singular values.
  • The Gaussian perturbation expansion has a zero-order, zero-mean √η term, an η term, and a uniformly controlled O(η^(3/2)) remainder.

Appendix C. Proof of Remark 2

The appendix proves a Foster–Lyapunov drift condition for affine invariant Langevin dynamics under the stated assumptions. The argument combines bounds on covariance-related quantities across large- and bounded-scale regimes, with sufficiently small a ensuring uniform negative drift.

  • Theorem and setup: The proof assumes f ∈ C2 ∩ L1(π) and strong growth outside a compact set through constants 0 < c1 < c2.These are the hypotheses stated in Theorem 4.
  • Theorem and setup: For N ≥ d + 3 and sufficiently small a > 0, the Lyapunov function (11) satisfies a Foster–Lyapunov condition.The condition holds with constants c > 0, b ∈ R, and a compact set K depending on a.
  • Bounding the drift: The argument bounds Tr K, D, P, and related trace terms using convexity, reverse triangle inequalities, Hessian control, and summation over particles.These estimates are assembled into bounds involving the scale variable S and the ratio θ := T/S.
  • Large-scale regime: For sufficiently large S, the derived estimates yield LWa/Wa ≤ −λ/2 after shrinking a so that c′a ≤ λ/4.The proof first obtains a uniform drift bound in the large-S regime.
  • Bounded-scale regime: When S is bounded, the remaining region with both S and J bounded is compact, and coercivity of f supplies the norm-like condition needed for the Foster–Lyapunov argument.This completes the bounded-scale case and the proof.
Loading 2609.03326v1…