Source-linked AI summary

Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows

Francesca Romana Crucinio, Sahani Pathiraja

arXiv:2609.18118v1stat.MLcs.LGmath.PR

TL;DR

The paper addresses convergence of WFR gradient flows for sampling distributions known up to a normalisation constant, focusing on strongly log-concave targets. It proves strong log-concavity preservation under curvature assumptions and derives non-asymptotic symmetrised-KL convergence without warm starts, with additive Wasserstein and Fisher–Rao dissipation.

  • Problem

    Existing WFR convergence results require strong conditions involving the initial distribution and target, including a warm-start condition.

  • Method

    The analysis combines Wasserstein transport and Fisher–Rao birth–death dynamics and uses strong log-concavity preservation under curvature assumptions.

  • Results

    WFR flows preserve strong log-concavity uniformly in time and exhibit exponential symmetrised-KL convergence with additive Wasserstein–Fisher–Rao dissipation, without warm-start assumptions.

  • Takeaways & Limitations

    The results provide refined convergence guarantees and characterize transport and birth–death components as additive contributions to WFR dissipation.

  • Takeaways & Limitations

    The guarantees focus on strongly log-concave targets, while extension to weakly log-concave, non-convex, or multimodal distributions remains open.

Abstract

from arXiv · show

We study the convergence of Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known up to a normalisation constant. By combining Wasserstein transport with Fisher-Rao birth-death dynamics, WFR flows balance exploration and selection. These flows have been recognised as a promising mechanism to accelerate convergence beyond Langevin dynamics. We show that for a class of strongly log-concave target distributions satisfying additional curvature conditions, WFR flows preserve strong log-concavity, in contrast to Wasserstein flows which enjoy this property only in the Gaussian setting. Exploiting this result, we derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, without requiring a warm-start as required in current estimates. In particular, we show that the convergence rate decomposes additively into Wasserstein and Fisher-Rao contributions, thereby confirming a recent conjecture within this setting. These results provide refined convergence guarantees and further develop the theoretical foundations of WFR gradient flows for sampling and Bayesian inference.

1 Introduction

The paper studies WFR gradient flows as a sampling framework that combines Wasserstein transport with Fisher–Rao birth–death dynamics. For strongly log-concave targets, it establishes preservation of strong log-concavity and an additive convergence rate without a warm-start condition.

  • WFR gradient flows combine Wasserstein diffusion with Fisher–Rao birth–death dynamics to balance exploration and selection in sampling.The framework is motivated by sampling targets known only up to a normalisation constant and by the need for efficient algorithms in challenging settings.
  • Unlike WFR flows, Wasserstein flows preserve log-concavity uniformly in time only in the Gaussian case.
  • The paper derives non-asymptotic convergence rates for the symmetrised KL divergence without requiring a warm-start condition.This addresses limitations of existing WFR convergence results, which require strong conditions on the initial-to-target density ratio.
  • The paper establishes conditions under which WFR flows preserve strong log-concavity for strongly log-concave targets.The result combines finite-time log-concavity preservation for the Wasserstein component with Fisher–Rao regularisation.
  • The convergence rate is the sum of the Wasserstein and Fisher–Rao rates, confirming a conjecture by Domingo-Enrich and Pooladian.

2 Wasserstein–Fisher–Rao Gradient Flow

WFR gradient flow combines Wasserstein transport and Fisher–Rao birth–death dynamics, with convergence analyzed through KL dissipation. The section contrasts its additive dissipation structure with existing WFR bounds requiring warm starts and only late-time validity.

  • 2 Wasserstein–Fisher–Rao Gradient Flow: WFR dynamics combine the Wasserstein and Fisher–Rao operators, while the Fisher–Rao component has a closed-form solution and the Wasserstein component generally does not.The PDE therefore combines transport-driven diffusion with Fisher–Rao reaction dynamics.
  • 2 Wasserstein–Fisher–Rao Gradient Flow: The instantaneous WFR dissipation equals the sum of the Wasserstein and Fisher–Rao dissipations, so it is at least as large as either component alone.This additive structure motivates a convergence rate combining the two singular-flow rates.
  • 2 Wasserstein–Fisher–Rao Gradient Flow: The Wasserstein-flow comparison relies on a Log-Sobolev assumption, while related KL guarantees also impose bounded second moments and a growth condition on the initial-to-target density ratio.These conditions are stated as assumptions for existing convergence results discussed in the section.
  • 2 Wasserstein–Fisher–Rao Gradient Flow: The resulting WFR rate has the additive Wasserstein–Fisher–Rao form conjectured in prior work, but the displayed result is obtained under the section’s stated assumptions.The rate is derived by differentiating KL along the WFR flow and using log-concavity-based arguments.
  • 2 Wasserstein–Fisher–Rao Gradient Flow: Earlier sharper WFR estimates require a warm start and apply only after t0 = log(M/δ^3), which increases as δ approaches zero.A one-dimensional Gaussian check likewise finds that available literature rates become sharp only for large t.

3 Preservation of log-concavity

The section establishes conditions under which WFR flows preserve strong log-concavity uniformly in time, using finite-horizon W-flow preservation together with Fisher–Rao regularisation. The result extends uniform preservation beyond the Gaussian setting, while its curvature guarantees depend on additional assumptions outside that setting.

  • FR-flow comparison: Unlike the W flow, the FR flow preserves strong log-concavity uniformly in time for strongly log-concave targets and initial distributions satisfying Assumption 1.Its curvature parameter αt is a convex combination of positive scalars for every t > 0.
  • Assumptions: The assumptions require smooth, strongly log-concave target and initial potentials plus strong convexity of the Schrödinger potential H; they are sufficient, not necessary.The condition on H controls curvature accumulated along the W-flow path, while the initial-density condition is described as user-compatible.
  • Proof strategy: The proof strategy combines a finite-time log-concavity result for the W flow with the strong regularising properties of the FR flow through sequential splitting.The intermediate W-flow density is characterised using Girsanov’s theorem.
  • W-flow comparison: The W flow preserves strong log-concavity only over a finite horizon when αh − Lπ < 0, but uniformly in time when αh − Lπ > 0.In the finite-horizon case, the horizon depends on the log-concavity of a difference involving V0 and Vπ and generally increases with αd.
  • WFR preservation: Theorem 1 shows that, under Lemmas 1–2 and an additional condition when αh − Lπ < 0, the WFR solution remains αt-strongly log-concave for all t > 0.When αh − Lπ > 0, no additional restriction on b is required; when αh − Lπ < 0, condition (12) is imposed.
  • Tightness and examples: For Gaussian targets, the theorem’s curvature condition is unnecessary, whereas for a non-Gaussian perturbation its bound is reasonably tight but degrades at large t.The Gaussian comparison reports tight constants; the non-Gaussian example’s asymptotic true constant is never worse than απ.

4 Convergence of Wasserstein–Fisher–Rao gradient flow

The section derives explicit convergence rates for WFR flow using strong log-concavity preservation, with symmetrised KL dissipation decomposing into Wasserstein and Fisher–Rao contributions.

  • Strong log-concavity preservation enables explicit non-asymptotic WFR convergence rates without warm-start or bounded-moment conditions.The analysis uses the symmetrised KL rather than KL alone.
  • Jeffreys’ divergence is used because it satisfies gradient dominance under Fisher–Rao flow and upper bounds KL(µ_t||π).KL itself is not geodesically convex under Fisher–Rao flow and does not satisfy the relevant gradient-dominance condition.
  • Proposition 1 establishes the decay result for the WFR PDE solution µ_t under the assumptions of the strong log-concavity theorem.The proof combines Wasserstein and Fisher–Rao dissipation estimates and applies Grönwall’s lemma.
  • The Proposition 1 rate is sharp for the displayed one-dimensional Gaussian comparison with the exact symmetrised KL decay.Figure 3 compares the exact rate with the Proposition 1 rate for a Gaussian and the target from Example 1.
  • The same exponential rate transfers to KL, although the initial Jeffreys’ divergence can make the bound looser at small times.This reflects the use of J(µ_0,π) as the bound’s initial constant.
  • The symmetrised KL decay rate is additive in the Wasserstein and Fisher–Rao rates, confirming the conjectured WFR decomposition.The result applies in the strongly log-concave setting, up to a factor from the strong log-concavity estimate.

5 Discussion

The discussion frames strong log-concavity preservation as the foundation for WFR convergence theory and identifies broader geometric settings as important directions for extension.

  • Under explicit curvature assumptions, WFR preserves strong log-concavity uniformly in time, with Fisher–Rao providing a regularising mechanism beyond Wasserstein-only preservation.This structural property supports the subsequent convergence analysis.
  • The resulting theory gives explicit non-asymptotic exponential symmetrised-KL convergence without warm starts, with additive transport and birth–death dissipation.The same estimate bounds KL through KL(µ_t∥π) ≤ J(µ_t,π).
  • The current guarantees are restricted to strongly log-concave targets, while extensions to weakly log-concave, non-convex, or multimodal distributions remain open.The preservation assumptions may also be sufficient rather than necessary.
  • A splitting proof strategy exploits complementary Wasserstein and Fisher–Rao operators to establish qualitative properties difficult to obtain directly from the reaction-diffusion equation.The perspective may generalise to other hybrid gradient flows and energy functionals.

A.1 Proof of Lemma 1

The proof characterizes the Wasserstein intermediate density using Girsanov’s theorem and a time-rescaled diffusion, then tracks strong log-concavity through a discretized recursion and its limiting ODE.

  • A time-rescaled process X_s = Y_{s/2} converts the diffusion to a convenient Brownian-motion representation.
  • Girsanov’s theorem characterizes the law of the overdamped Langevin diffusion underlying the Wasserstein step.
  • The discretized Wasserstein evolution factorizes through Wiener increments, yielding intermediate densities represented by strongly convex potentials.
  • Each recursion step preserves strong log-concavity under positivity conditions on the evolving curvature constants.
  • When αh − Lπ < 0, strong log-concavity persists only up to a finite horizon because the associated curvature ODE decreases without a stable fixed point.
  • When αh − Lπ > 0, solving the curvature ODE shows that the strong-concavity constant remains positive for all time.

A.2 Proof of Theorem 1

The theorem is proved by applying a sequential Wasserstein–Fisher–Rao splitting scheme and controlling its curvature recursion through an associated ODE.

  • The W-FR splitting applies a Wasserstein step followed by a Fisher–Rao step, and converges to the exact WFR solution at rate O(γ).
  • For αh − Lπ < 0, iteration-dependent step sizes are required because the Wasserstein preservation horizon depends on the relative convexity of the initial and target potentials.
  • The first split step preserves strong log-concavity by combining the Wasserstein curvature bound with the target potential’s curvature.
  • The Fisher–Rao step updates the potential by exponentially weighting the Wasserstein-step contribution, producing a new strong log-concavity constant.
  • Inductively, the split-step curvature constants follow a two-stage ODE discretization and remain positive under sufficiently small step-size thresholds.
  • When αh − Lπ > 0, the Wasserstein flow already preserves log-concavity uniformly in time, so the WFR flow has no further restrictions.
Loading 2609.18118v1…