Source-linked AI summary
Solving high-dimensional Hamilton-Jacobi-Bellman PDEs using neural networks: perspectives from the theory of controlled diffusions and measures on path space
Nikolas Nüsken, Lorenz Richter
TL;DR
The paper addresses how to solve high-dimensional HJB-PDEs and related diffusion-control problems when algorithmic performance depends critically on the loss function. It develops a path-space divergence framework, including novel log-variance divergences connected to forward-backward SDEs. The paper reports faster convergence, more accurate optimal-control approximation, and favorable stability properties for log-variance methods, while identifying scope limits for broader nonlinear PDEs and additional Hamiltonian minimisation.
Problem
High-dimensional HJB-PDEs and diffusion-control applications require effective iterative optimisation methods and principled loss functions, especially for sampling and rare-event problems.
Method
The paper constructs losses from divergences between path measures, connects log-variance divergences to forward-backward SDEs, and analyses their Monte Carlo gradient estimators.
Results
Log-variance adjustments can provide faster convergence and more accurate optimal-control approximation, and log-variance divergences are the only considered losses satisfying both studied stability criteria.
Takeaways & Limitations
The path-space viewpoint unifies several iterative diffusion optimisation methods and supports comparing their estimator stability and large-batch equivalences.
Takeaways & Limitations
The stated extensions rely on shift invariance when the nonlinearity depends on V only through ∇V, while additional Hamiltonian minimisation is left for future work.
Abstract
from arXiv · showhide
Optimal control of diffusion processes is intimately connected to the problem of solving certain Hamilton-Jacobi-Bellman equations. Building on recent machine learning inspired approaches towards high-dimensional PDEs, we investigate the potential of $\textit{iterative diffusion optimisation}$ techniques, in particular considering applications in importance sampling and rare event simulation, and focusing on problems without diffusion control, with linearly controlled drift and running costs that depend quadratically on the control. More generally, our methods apply to nonlinear parabolic PDEs with a certain shift invariance. The choice of an appropriate loss function being a central element in the algorithmic design, we develop a principled framework based on divergences between path measures, encompassing various existing methods. Motivated by connections to forward-backward SDEs, we propose and study the novel $\textit{log-variance}$ divergence, showing favourable properties of corresponding Monte Carlo estimators. The promise of the developed approach is exemplified by a range of high-dimensional and metastable numerical examples.
1 Introduction
The paper frames high-dimensional HJB-PDE solution as an iterative diffusion optimisation problem, with algorithms distinguished by their loss functions. It develops a path-space divergence framework connecting existing methods and introduces log-variance divergences with favorable convergence, approximation, and estimator-stability properties.
- Motivation: High-dimensional HJB-PDE applications motivate iterative diffusion optimisation for improving sampling efficiency, state-space exploration, or fit to empirical data.The framework repeatedly simulates controlled diffusions, evaluates a performance measure and gradient, and updates the control.
- Algorithmic framework: Existing iterative diffusion optimisation algorithms mainly differ in the loss function used to evaluate and update the control.Losses are typically expectations involving the controlled diffusion, so empirical implementation relies on ensemble averages.
- Contributions: Path-space divergences provide a systematic framework for designing and analysing loss functions across existing numerical methods.The paper builds this perspective from connections between optimal-control functionals and the KL-divergence.
- Contributions: Log-variance divergences modify forward-backward SDE approaches and can yield faster convergence and more accurate optimal-control approximations.They encapsulate a family of forward-backward SDE systems and are supported by numerical experiments.
- Contributions: Certain control-objective and forward-backward-SDE algorithms become equivalent when the simulation sample size N is large.This establishes a large-batch connection between KL-divergence-based and log-variance-based approaches.
- Contributions: Only the log-variance divergences satisfy both studied stability criteria: robustness under tensorisation and robustness at the optimal control solution.The paper investigates the associated sample-based gradient estimators and illustrates the result with extensive numerical experiments.
2 Optimal control problems, change of path measures and Hamilton-Jacobi-Bellman PDEs: connections and equivalences
The paper connects controlled diffusions, optimal control, path-measure conditioning, FBSDEs, and HJB-PDEs through equivalent formulations under suitable conditions. These connections support importance sampling and rare-event simulation, including a zero-variance control characterization.
- Controlled diffusions: The uncontrolled diffusion is specified by drift b, diffusion σ, Brownian motion, and deterministic initial condition xinit, while the controlled dynamics add an admissible steering term u.The coefficients satisfy regularity, boundedness, and nondegeneracy assumptions ensuring unique strong solutions.
- Optimal control and HJB-PDEs: The optimal-control problem seeks an admissible control u∗ minimizing the expected running and terminal costs, with the value function defined as the optimal cost-to-go.The associated HJB-PDE characterizes the value function, and the optimal control is recovered from its gradient.
- Equivalent formulations: The HJB formulation, optimal-control formulation, and FBSDE formulation are connected through Y_s = V(X_s,s) and Z_s = −u∗(X_s,s) = σ^T∇V(X_s,s).Under classical well-posedness and suitable uniqueness conditions, solving any of these problems can recover the HJB solution.
- Path-measure conditioning: Conditioning defines a reweighted path measure Q from the work functional, and seeks a control whose induced path measure P^u∗ coincides with Q.This formulation links rare-event sampling to controlled diffusion dynamics on path space.
- Scope and assumptions: The equivalences rely on the particular controlled-SDE and cost structure that enables Girsanov’s theorem, while stronger regularity assumptions may be relaxed using viscosity solutions.The paper notes that optimal controls need not be differentiable, so optimal-control or FBSDE solutions do not necessarily yield classical HJB solutions.
3 Approximating probability measures on path space
The paper recasts control optimization as minimizing divergences between probability measures on path space, yielding loss functions for approximating the optimal control. It develops variance and log-variance losses, connects them to KL and FBSDE formulations, and analyzes their Monte Carlo representations and implementation.
- Path-space formulation: The framework includes relative entropy, cross-entropy, variance, and log-variance losses, with the target measure Q representing the optimal controlled dynamics.Under the stated assumptions, the optimal control induces the target path measure Q = P^{u*}.
- Path-space formulation: Path-space divergences are converted into loss functions on control vector fields through the map from controls to induced path measures.The construction uses the measurable mapping u ↦ P^u and Girsanov representations.
- Variance-based losses: The loss L_D is nonnegative and vanishes exactly at the optimal control, so minimizing it can in principle recover u*.This characterization relies on the divergence property and the paper’s optimal-control theorem.
- Variance-based losses: Variance and log-variance divergences extend the path-space framework, with the log-variance form motivated by FBSDE connections and potential robustness in high-dimensional sample estimation.The log-variance divergence is presented as a new quantity and is closely connected to relative entropy.
- Monte Carlo representations: Girsanov-based representations replace path-space expectations with expectations on the underlying probability space, making the losses more suitable for probabilistic interpretation and Monte Carlo simulation.The paper distinguishes expectations over path space from expectations over the underlying sample space.
- Monte Carlo representations: Controlled-process representations for relative entropy and log-variance losses avoid the exponential factors appearing in some alternative formulations, while the choice of auxiliary control v affects estimator statistics.The authors caution that exponential reweighting can worsen estimator variance and that tuning v can materially affect numerical performance.
- Algorithmic implementation: The algorithmic framework parameterizes controls by θ and seeks the global minimum of the resulting loss over the parameterized control family.This provides the bridge from the theoretical loss construction to practical neural-network or Galerkin-style optimization.
4 Equivalence properties in the limit of infinite batch size
In the infinite-batch setting, log-variance losses share key equivalence properties with relative-entropy and moment losses. Their gradients can avoid differentiating running and terminal costs, while correct identification by alternative moment losses may require exact auxiliary parameters.
- Loss equivalence: Log-variance and relative-entropy losses behave equivalently in the infinite-batch limit when the log-variance update uses v = u.The paper also provides an analytical gradient expression.
- Gradient properties: The log-variance and relative-entropy formulations have matching gradients away from the optimal control, although their derivative calculations differ.Relative entropy differentiates the SDE solution and the running and terminal costs; log-variance does not require differentiating the latter.
- Optimization landscape: Log-variance losses have no non-optimal local minima under the stated directional-derivative condition: the derivative vanishes for all zero-mass perturbations if and only if P = Q.This suggests an optimization landscape without local minima where the procedure could get stuck.
- Moment losses: Moment and log-variance losses are equivalent at infinite batch size when v = u, but with v ≠ u, the moment loss identifies the optimal control only when y0 = −log Z.Otherwise, the optimal-control approximation can be inaccurate unless y0 is determined without error.
- Variance reduction: The log-variance loss’s derivative includes expectation-zero terms that act as control variates, while weight-specific control-variate corrections may add computational overhead and remain sensitive to Monte Carlo error.The paper notes that the weight-specific contribution is often small and fluctuates around zero.
5 Finite sample properties and the variance of estimators
Finite-sample analysis distinguishes estimators by robustness at the optimum and under tensorisation. Log-variance estimators satisfy both robustness desiderata, whereas relative-entropy, variance, and cross-entropy estimators fail at least one of them, often with high-dimensional error growth.
- Robustness criteria: The paper defines robustness at the solution to mean that Monte Carlo fluctuations in gradients are suppressed near the optimum, supporting accurate approximation.Without this property, relative error can become unbounded near the solution.
- Robustness at the solution: Variance and log-variance estimators are robust at the optimal control for every auxiliary update choice v, while relative entropy is not robust there.Non-robustness can make relative error grow without bound near the optimum and destabilize gradient descent.
- Robustness at the solution: The moment estimator is robust at the optimal control only under a condition involving y0, and exact satisfaction of that condition is rarely achieved in practice.The paper identifies this as a potential source of practical instability.
- High-dimensional robustness: Cross-entropy estimators are not robust both at the optimal control and under tensorisation.The tensorisation analysis gives a fixed-sample-size growth result, while the optimal-control analysis states non-robustness for all v.
- High-dimensional robustness: Under tensorisation, the standard Monte Carlo log-variance estimator is robust for all reference measures, whereas relative entropy is robust only when the reference measure equals P.Tensorisation models increasing dimensionality through products of identical component measures.
- High-dimensional robustness: The standard Monte Carlo variance estimator is not robust under tensorisation, with relative errors that can grow exponentially in the number of factors.The paper states that this suggests poor performance in high-dimensional settings.
6 Numerical experiments
The numerical experiments evaluate iterative diffusion optimisation across Ornstein–Uhlenbeck and metastable problems, varying losses, optimisers, initialisations, dimensions, and batch sizes. Log-variance performs consistently well, while cross-entropy and variance losses show slower convergence, instability, or sensitivity in several settings.
- 6.1 Computational aspects: The experiments combine explicit loss representations, automatic differentiation, gradient descent, Euler–Maruyama discretisation, and neural-network approximations of the optimal control.Controls are represented with feed-forward neural networks, with a shared space-time network performing better in the reported experiments.
- 6.2 Ornstein-Uhlenbeck dynamics with linear costs: In dimension d = 40, relative entropy and log-variance losses perform best, while moment and cross-entropy losses converge significantly more slowly.The variance loss is numerically unstable and is omitted from the figure and subsequent experiments.
- 6.2 Ornstein-Uhlenbeck dynamics with linear costs: The cross-entropy approximation is inferior to the relative entropy approximation in the plotted components of the 40-dimensional optimal control.The reported dimensional dependence of loss-estimator relative errors agrees with the theoretical prediction, although the experiment goes beyond the product case.
- 6.2 Ornstein-Uhlenbeck dynamics with linear costs: Initialising y0 with −log Z makes moment and log-variance losses perform similarly, whereas the initialisation of y0 can significantly affect convergence speed.The experiment compares Adam and SGD with N = 200, η = 0.01, and ∆t = 0.01.
- 6.3 Ornstein-Uhlenbeck dynamics with quadratic costs: For quadratic costs, relative entropy converges fastest and log-variance follows, while cross-entropy converges significantly more slowly and can diverge at larger learning rates.With SGD, the moment loss experiences fluctuations during later gradient steps.
- 6.3 Ornstein-Uhlenbeck dynamics with quadratic costs: Sampling X0 from N(0, Id×d) hardly affects overall convergence but improves approximation accuracy at initial time t = 0.The effect is particularly pronounced when independent ansatz functions are used at each time step.
- 6.4 Metastable dynamics in low and high dimensions: In the one-dimensional double-well problem, optimal control reduces relative error from 63.86 to 1.94 and raises the barrier-crossing ratio from about 0.02% to approximately 87.28%.The controlled dynamics steer trajectories toward the right well, where e−g is concentrated.
- 6.4 Metastable dynamics in low and high dimensions: In the metastable experiments, log-variance and moment losses perform well at both tested batch sizes, with log-variance reaching a satisfactory approximation in fewer gradient steps.Relative entropy suffers instabilities near the solution, while log-variance gradients have significantly lower variances.
7 Conclusion and outlook
The paper unifies iterative diffusion optimisation methods through divergences between path measures and identifies log-variance losses as especially stable. It also outlines extensions to other path-space divergences.
- A path-measure divergence framework unifies existing numerical methods within iterative diffusion optimisation.
- The novel log-variance divergences connect closely to forward-backward SDEs and are fundamentally equivalent to KL-divergence approaches.
- Only the log-variance loss is stable under both tensorisation and at the optimal control solution among the studied losses and estimators.
- Outlook: The authors propose extending the framework to other divergences on path space and studying the resulting algorithms.
A.1 Proofs for Section 3.1
The appendix establishes measure-theoretic identities used for the divergences and derives expressions for the associated losses. These calculations support the propositions in Section 3.1.
- The proofs compute divergence-related expressions from the model equations and measure changes introduced in Section 3.1.
- The appendix states that the relevant path measures are equivalent and derives their Radon–Nikodym derivatives using growth assumptions and Girsanov’s theorem.
- The variance and log-variance losses are treated through explicit stochastic-calculus calculations supporting their stated propositions.
A.2 Proofs for Section 4
The appendix derives the stochastic representations and variance identities underlying the results in Section 4. The proofs use change-of-measure arguments, differentiation under the integral, and Itô calculus.
- A change of measure and Girsanov’s theorem provide the stochastic-process representation used in the Section 4 proofs.
- The derivations interchange derivatives and integrals under dominated convergence and evaluate stochastic integrals using Itô’s isometry.
- The appendix computes variance and log-variance loss expressions and compares them with earlier identities to establish the propositions.
A.3 Proofs for Section 5
The appendix proves the Section 5 stability and variance claims through perturbation arguments, product-measure calculations, and inequalities for the relevant estimators.
- The appendix computes variance expressions for log-variance and cross-entropy estimators and uses Itô’s isometry to establish the corresponding claims.
- The proofs analyze perturbations of the optimal control and derive stochastic identities using Itô’s formula, integration by parts, and quadratic variation.
- Product-measure and sample-variance calculations establish the stated robustness and non-robustness results.
- The remaining arguments combine relative-error bounds, Jensen’s inequality, and the definitions of the losses to complete the propositions.
A.4 Optimal control for Ornstein-Uhlenbeck dynamics with linear cost
The section gives an analytical treatment of the control problem, relating the HJB value function to ψ and using the explicit distribution of X_T to derive a further expression.
- The control problem considered in Section 6.2 can be solved analytically.
- The value function solving the HJB-PDE satisfies V(x, t) = −log ψ(x, t).
- The distribution of X_T is known explicitly and is used with (21) to obtain a further result.