Source-linked AI summary

Two Adjoint Perspectives on Fokker-Planck Optimization: A Microscopic-Macroscopic Correspondence

Kathrin Hellmuth, Qin Li, Yunan Yang

arXiv:2609.02072v1math.NAmath.APmath.OCmath.PR

TL;DR

The paper asks how apparently different macroscopic and microscopic adjoints for Fokker–Planck optimization correspond. It establishes their continuum correspondence and analyzes Eulerian and Lagrangian discretizations, finding that both converge to the same continuum gradient despite differing after discretization.

  • Problem

    The paper addresses the gap between the backward Kolmogorov adjoint in the Eulerian formulation and the pathwise adjoint in the Lagrangian formulation.

  • Method

    The paper establishes a conditional-expectation correspondence at the continuum level and studies Eulerian and Lagrangian particle-based discretizations.

  • Results

    Both discretization strategies provide convergent numerical approximations of the continuum gradient, although their finite-dimensional adjoints differ.

  • Takeaways & Limitations

    Eulerian and Lagrangian formulations are complementary representations of the same continuum optimization gradient, but should not be treated as identical discrete objects.

  • Takeaways & Limitations

    The continuum pathwise adjoint depends on the full terminal law, while the numerical rate may require further investigation to distinguish noise artifacts from the actual rate.

Abstract

from arXiv · show

The Fokker-Planck equation admits both a macroscopic Eulerian description through probability densities and a microscopic Lagrangian description through stochastic trajectories. Consequently, optimization problems constrained by the Fokker-Planck equation can be formulated from either perspective. Surprisingly, the corresponding adjoint equations appear to be fundamentally different: the macroscopic adjoint is governed by the backward Kolmogorov equation, whereas the microscopic adjoint evolves pathwise along stochastic trajectories. In this note, we reconcile these two formulations by establishing their correspondence in the continuum setting. We further show that, although their discrete gradients no longer coincide after discretization, both provide consistent numerical approximations of the continuum gradient. Explicit convergence rates are established for both discretization strategies.

1. Introduction

The paper reconciles Eulerian and Lagrangian adjoint formulations for Fokker–Planck optimization, then analyzes their distinct discretizations and convergence to the continuum gradient.

  • 1. Introduction: The central question is how microscopic pathwise and macroscopic backward-Kolmogorov adjoint formulations can be reconciled.
  • 1. Introduction: The paper interprets the two formulations as complementary Eulerian and Lagrangian differentiation paradigms for the same optimization problem.
  • 1. Introduction: A conditional-expectation identity connects the pathwise adjoint to the macroscopic adjoint and explains their identical continuum first-order variation.
  • 1. Introduction: The Eulerian discretization solves a discretized adjoint PDE, while the Lagrangian discretization directly uses finitely many stochastic sample paths.
  • 1. Introduction: Despite their discrete differences, both numerical representations converge to the same continuum gradient under the paper’s stated assumptions.
  • 1. Introduction: The continuum identity fails after discretization, so OTD and DTO yield genuinely different finite-dimensional adjoints.

2. Continuum Correspondence

The continuum analysis reconciles Eulerian and Lagrangian adjoints by showing that their apparently different constructions yield the same gradient under stated regularity assumptions.

  • Continuum correspondence: The Eulerian formulation uses a backward Kolmogorov adjoint, while the Lagrangian formulation follows stochastic trajectories with pathwise adjoints.The correspondence is established before numerical discretization.
  • Assumptions: The assumptions impose different regularity requirements in the Eulerian and Lagrangian settings, with simpler conditions for the Lagrangian ODE formulation.The Eulerian assumptions place the analysis in a Schauder fixed-point setting.
  • Continuum correspondence: Theorem 1 gives the conditional-expectation identity E(Yt | Xt = x) = ∇µ(t, x) ρt-almost everywhere.This identity directly connects the pathwise adjoint to the spatial gradient of the Kolmogorov adjoint.
  • Continuum correspondence: The continuum gradient therefore admits two equivalent representations, one Eulerian and one Lagrangian.The two representations arise from perturbing the Fokker–Planck equation and differentiating the stochastic dynamics, respectively.
  • Continuum correspondence: The continuum equivalence depends on the pathwise adjoint’s terminal condition using the full law ρT, not independently available sample-path information.Consequently, the stochastic Lagrangian formulation is an alternative continuum representation rather than a particle approximation.
  • Continuum correspondence: The proof preserves the natural dual pairing between the forward Jacobian and the pathwise adjoint.Their pairing is constant along the dynamics, while Feynman–Kac connects the resulting expression to the Kolmogorov adjoint.

3. Eulerian Formulation

The Eulerian optimize-then-discretize approach approximates the continuum gradient through empirical particle measures and adjoint PDEs, with weak convergence and explicit sampling and time-discretization error control.

  • 3. Eulerian Formulation: The Eulerian strategy first derives the continuum optimality system, then approximates the forward Fokker–Planck solution by an empirical particle measure.The adjoint remains a backward Kolmogorov PDE, but its terminal condition is determined by the empirical terminal measure.
  • 3. Eulerian Formulation: Because the Eulerian gradient approximation is measure-valued, convergence is evaluated weakly by testing against functions ϕ ∈ C1.Strong-topology convergence is not expected for the vector-valued measure.
  • 3.1. Main Theorem.: The Eulerian gradient converges weakly to the continuum gradient, and its error has the same order as the particle sampling error.The result follows from stability of the gradient operator with respect to perturbations of the forward trajectory.
  • 3.1. Main Theorem.: The gradient operator Γ is weakly continuous on C([0, T]; Pq(Rd)), enabling empirical-measure convergence to transfer to gradient convergence.The continuity estimate depends on the test-function norm, dimension, time horizon, drift bounds, and the Lipschitz constant of δJ.
  • 3.2. Extension to discrete-in-time setting.: Fully discrete Eulerian estimates incorporate particle sampling, numerical quadrature, and Euler–Maruyama trajectory errors as separate contributions.The analysis explicitly tracks dependence on both the particle count N and time step τ.
  • 3.2. Extension to discrete-in-time setting.: Euler–Maruyama trajectories converge strongly with order τ 1/2 to the continuous stochastic dynamics.The resulting time-discrete empirical measure is denoted ρN,τ.
  • 3.2. Extension to discrete-in-time setting.: The time-discrete Eulerian approximation has an explicit error bound combining sampling, integration, and Euler–Maruyama discretization effects.The bound holds under Assumptions (A1)–(A3E) and (A4), with τ < 1.
  • 3.2. Extension to discrete-in-time setting.: The estimates retain N-dependence in their constants so that sampling and time-discretization rates remain explicit.This avoids hiding particle-count dependence inside generic continuity constants.

4. Lagrangian Formulation

The Lagrangian formulation approximates the Fokker–Planck solution with particles, derives particle adjoints and gradients, and quantifies their convergence to the continuum gradient in weak form. The analysis separates sampling, stability, quadrature, and Euler–Maruyama errors, including the fully discrete setting.

  • 4. Lagrangian Formulation: The DTO formulation first discretizes the forward Fokker–Planck equation with a particle system and then differentiates the resulting finite-dimensional dynamics.
  • 4. Lagrangian Formulation: The particle adjoints satisfy the pathwise adjoint dynamics with terminal conditions determined by the empirical terminal measure.
  • 4. Lagrangian Formulation: The resulting gradient resembles a Monte Carlo average, but its error is not classical because the adjoint depends on the empirical terminal measure and is correlated with particle positions.
  • 4. Lagrangian Formulation: The central quantitative question is whether dependence on the empirical terminal measure degrades the classical N^-1/2 Monte Carlo rate.
  • 4.1. Main Theorem.: The error analysis decomposes the approximation into a classical Monte Carlo component and a stability error caused by replacing the continuum terminal measure with its empirical counterpart.
  • 4.2. Discrete-in-time setting.: The fully discrete analysis incorporates time quadrature and Euler–Maruyama errors alongside the particle approximation and uses pathwise backward Euler for the adjoint ODE.
  • 4.1. Main Theorem.: The Lagrangian gradient converges weakly to the continuum gradient when evaluated against test functions, under the stated regularity and moment assumptions.
  • 4.2. Discrete-in-time setting.: For τ ≤ 1, Theorem 5 provides a bound uniform in N and τ under additional smoothness assumptions, while the total error is split into sampling, temporal, and SDE-discretization components.

5. Numerical Examples

Numerical experiments in dimensions 1 and 2 compare particle densities and Eulerian/Lagrangian gradient approximations with PDE references, testing predicted convergence under increasing particle numbers and time-step refinement.

  • Experimental setup: The experiments evaluate Eulerian and Lagrangian particle approximations against the continuum gradient in dimensions d = 1 and d = 2.The reference gradient is computed from the forward and adjoint Fokker–Planck equations.
  • Particle approximation: N = 200 particle approximations are visibly coarse, while larger particle counts produce more accurate densities and gradients.The study uses N ∈ {200, 2000} in d = 1 and N ∈ {200, 10000} in d = 2.
  • Gradient approximations: For N = 2000 in d = 1, the Eulerian and Lagrangian gradient approximations almost perfectly overlay the reference.The corresponding approximations are constructed from particle trajectories through the Eulerian and Lagrangian formulas.
  • Gradient approximations: The Eulerian approximation is more accurate than the Lagrangian approximation in the linear setting, where its adjoint exactly matches the reference adjoint.This advantage is less pronounced for the nonlinear MMD objective.
  • Approximation rates: The convergence-rate experiments use mean errors over M = 30 realizations for the MMD objective and compare empirical slopes with theory.Time-discretization slopes are also reported with N = 2^15 particles, and all exceed the expected convergence rate.
  • High-dimensional convergence: For d = 10, the Lagrangian gradient's empirical standard deviation decays approximately at the Monte Carlo rate N^-1/2 as particle count increases.The median stabilizes, while a classical grid-based reference solver is prohibitively expensive in this dimension.

6. Conclusion

The conclusion establishes that Eulerian and Lagrangian adjoint solvers converge to the continuum gradient, while identifying dimension-dependent sampling, temporal regularity, and omitted PDE-discretization effects as open boundaries.

  • Conclusion: Both formulations provide convergent numerical approximations of the continuum gradient, with explicitly specified rates and weak convergence against smooth test functions.The expected classical Monte Carlo rate is N^-1/2, while the theory supports a dimension-dependent Wasserstein sampling rate α(N).
  • Conclusion: The dimension-dependent rate arises because the gradient is naturally quadratic and the analysis treats the adjoint variable accordingly.The conclusion leaves open whether the deterioration is intrinsic to Monte Carlo functional-gradient approximation or an artifact of the analysis.
  • Conclusion: Temporal regularity of the continuum SDE limits the quadrature error, although stronger orders may follow under weaker notions of convergence.The analysis is expected to generalize to schemes such as Milstein.
  • Conclusion: The present analysis excludes PDE discretization error in the Eulerian approximation, though gradient linearity in the adjoint may allow its direct transfer.

Appendix A. A-priori estimates

The appendix summarizes a-priori estimates used to analyze the particle and Eulerian approximations, beginning with well-posedness of the forward equation under the stated assumptions.

  • Appendix A. A-priori estimates: Both Eulerian and Lagrangian error estimates rely on a-priori estimates.
  • Appendix A. A-priori estimates: Under Assumptions (A1) and (A2), standard results establish a unique measure-valued solution to the forward Fokker–Planck equation.The solution is obtained from existence of a probability measure on the corresponding SDE path space.

A.1. Continuum solutions.

The continuum forward and adjoint problems are well posed under stated assumptions, and particle empirical measures inherit dimension- and moment-dependent convergence estimates.

  • A.1. Continuum solutions.: The forward equation has a unique solution ρ ∈ C([0, T]; Pq(Rd)) under Assumptions (A1) and (A2).
  • A.1. Continuum solutions.: Bounded drift and constant diffusion preserve the q-th moment bound from the initial condition and therefore bound all lower moments.
  • A.1. Continuum solutions.: The backward Kolmogorov adjoint equation is well posed under Assumptions (A1), (A2), and (A3E), with a unique solution and a stability estimate.Additional regularity is imposed through Assumption (A4) for further estimates.
  • A.1. Continuum solutions.: Theorem 2 follows from a Wasserstein-1 Monte Carlo estimate together with the identity ρt = Law(Xt).
  • A.1. Continuum solutions.: For empirical distributions of N independent particles, the theorem provides a dimension- and q-th-moment-dependent bound with constant C.
  • A.1. Continuum solutions.: The particle dynamics SDE admits a solution whose paths have temporal Hölder regularity.

A.2. Particle systems.

The particle-system analysis establishes regularity, stability, and well-posedness for the forward and adjoint dynamics, including Hölder continuity of empirical distributions and pathwise adjoint bounds.

  • Forward particle dynamics: Under (A1) and (A2), the SDE has a unique solution with Hölder-continuous paths.The Hölder estimate uses a constant depending on ∥a∥∞, p, and T.
  • Proof strategy: The forward regularity proof combines the Burkholder–Gundy–Davis inequality with Wasserstein bounds on particle differences.This yields Hölder continuity of the empirical particle distribution.
  • Forward particle dynamics: The empirical distribution ρN is Hölder continuous in time with exponent 1/2 under the Wasserstein-1 metric.The constant depends only on ∥a∥∞ and T.
  • Adjoint dynamics: The adjoint ODE admits a unique pathwise solution and is stable with respect to its final condition.Existence and stability follow from standard ODE arguments and the linearity of the adjoint dynamics.
  • Adjoint dynamics: Under (A1), (A2), and (A3L), the adjoint process has a uniformly bounded second moment.The bound is independent of t and s in [0,T].

A.3. Time discretization.

The time-discretization analysis compares the continuous particle system with its Euler–Maruyama approximation and derives a corresponding empirical-distribution estimate.

  • Empirical distributions: Particle-distance estimates transfer trajectory convergence to the empirical distributions of the continuous and discretized systems.The transfer uses the same particle-distance bound as equation (23).
  • Empirical distributions: The resulting corollary controls the empirical distribution functions ρN for the original particles and their Euler–Maruyama discretization.The statement applies under assumptions (A1) and (A2).

A.4. Proofs of Lemmas 1 and 2.

The lemma proofs use adjoint stability, Euler-scheme error decompositions, and regularity estimates to obtain bounds that are uniform over time steps.

  • Lemma 1: Lemma 1(a) follows from the uniform second-moment bound for independent copies of the adjoint process.The copies eYi,t are i.i.d. copies of Yt.
  • Lemma 1: Lemma 1(b) exploits adjoint stability and resolves the resulting recursion into a bound uniform in the time-step index k.The recursion uses (1+Cτ)^K ≤ e^(CKτ) with Kτ=T.
  • Lemma 1: The rate in equation (9) follows from the cited result and the constant CJ from assumption (A3L).The proof explicitly identifies CJ as the constant supplied by (A3L).
  • Lemma 2: Lemma 2 decomposes the continuous-versus-Euler adjoint error and bounds its terms using bounded and Lipschitz gradients together with process regularity.The argument combines Lipschitz continuity of Y, Hölder continuity of X, and strong Euler–Maruyama convergence.

Funding

The paper acknowledges support from NASA’s Jet Propulsion Laboratory, the Office of Naval Research, the National Science Foundation, and the Vilas associate award.

  • Funding: KH acknowledges support from the Jet Propulsion Laboratory PDRDF 24AW0133.
  • Funding: QL acknowledges support from ONR-N000142612095, NSF-DMS-2608503, and a Vilas associate award.
  • Funding: YY is partially supported by NSF grants DMS-2409855 and DMS-2540324 and ONR Award N00014-24-1-2088.
Loading 2609.02072v1…