Source-linked AI summary
Dynamic programming for optimal control of stochastic McKean-Vlasov dynamics
Huyên Pham, Xiaoli Wei
TL;DR
The paper addresses optimal control of stochastic McKean-Vlasov dynamics motivated by cooperative equilibria in large populations under common noise. It develops dynamic programming through the conditional-law flow, derives a Wasserstein-space Bellman equation, and establishes viscosity and uniqueness results for the value function. The framework is applied to an explicitly solved linear-quadratic problem and an interbank systemic-risk model.
Problem
Optimal control of McKean-Vlasov dynamics with common noise requires a general dynamic-programming treatment for value functions on conditional probability laws.
Method
The paper reformulates the problem using the conditional law as state, proves its flow property, and applies Lions differentiability with Itô’s formula to derive the Bellman equation.
Results
The value function satisfies a dynamic programming principle and is the unique continuous viscosity solution of the Bellman equation under a quadratic growth condition.
Takeaways & Limitations
The framework provides a dynamic-programming and PDE characterization for stochastic McKean-Vlasov control in Wasserstein space, with an explicit linear-quadratic application.
Takeaways & Limitations
The formulation restricts controls to the common-noise filtration; open-loop controls depending on both noises are left for future work.
Abstract
from arXiv · showhide
We study the optimal control of general stochastic McKean-Vlasov equation. Such problem is motivated originally from the asymptotic formulation of cooperative equilibrium for a large population of particles (players) in mean-field interaction under common noise. Our first main result is to state a dynamic programming principle for the value function in the Wasserstein space of probability measures, which is proved from a flow property of the conditional law of the controlled state process. Next, by relying on the notion of differentiability with respect to probability measures due to P.L. Lions [32], and It{ô}'s formula along a flow of conditional measures, we derive the dynamic programming Hamilton-Jacobi-Bellman equation, and prove the viscosity property together with a uniqueness result for the value function. Finally, we solve explicitly the linear-quadratic stochastic McKean-Vlasov control problem and give an application to an interbank systemic risk model with common noise.
1 Introduction
The paper develops dynamic programming for stochastic McKean-Vlasov control with common noise, motivated by large interacting populations and cooperative equilibria. It derives a Bellman equation and viscosity characterization from the conditional-law flow.
- Motivation: Stochastic McKean-Vlasov control models mean-field interactions among particles evolving under common and idiosyncratic noise.The common noise creates a random environment, while particle coefficients depend on the empirical distribution.
- Control formulation: The formulation uses feedback controls depending on the state and its conditional law, with semi-feedback controls also allowing open-loop dependence on common noise.Controls may alternatively be represented as random fields on the state space.
- Related work: Prior work studied McKean-Vlasov control through maximum principles or dynamic programming under specific dynamics, density assumptions, or without common noise.The paper targets a general stochastic setting.
- Contributions: The paper proves a dynamic programming principle by establishing a flow property for the controlled conditional distribution and reformulating it as the sole controlled state variable.Continuity of the value function in Wasserstein space is also used.
- Contributions: Using Lions differentiability and an Itô chain rule for conditional-measure flows, the paper derives a fully nonlinear second-order Bellman PDE on Wasserstein space.The value function is shown to have the viscosity property, with a uniqueness result.
- Applications: The final section applies the framework to a linear-quadratic stochastic McKean-Vlasov control problem with explicit solutions.The paper also presents an interbank systemic-risk application with common noise.
2 Conditional McKean-Vlasov control problem
The control problem is formulated on a product probability space with common and idiosyncratic Brownian motions, conditional laws in P2(Rd), and progressively measurable controls. Under stated regularity assumptions, the state equation has a unique solution and the value function is defined through the associated cost.
- Probabilistic setting: The probability space separates common noise W^0 from idiosyncratic noise B, with filtrations generated by these processes and an independent atomless space.This construction supports random variables realizing probability measures in P2(Rd).
- State space: The conditional law of the state given the common-noise filtration is valued in P2(Rd), equipped with the 2-Wasserstein distance.The Wasserstein Borel σ-field is characterized through integration maps against continuous functions with quadratic growth.
- Controls: Admissible controls are F^0-progressive processes valued in a Polish control space A.The chosen formulation restricts controls to depend on the common-noise filtration.
- Dynamics: Under the standing Lipschitz and growth assumptions, the controlled stochastic McKean-Vlasov equation has a unique square-integrable adapted solution.The conditional-law process has continuous trajectories and is F^0-progressively measurable.
- Objective: The running and terminal costs define a finite cost functional, and the value function is v(t, ξ) := inf α∈A J(t, ξ, α).The value function satisfies a quadratic growth condition.
- Scope: The formulation’s goal is to characterize the value function through a dynamic-programming partial differential equation.The paper later restricts open-loop controls depending on both noises to future work.
3 Dynamic programming
This section establishes the dynamic-programming framework by constructing the conditional-law flow, proving continuity of the cost and value functions, and deriving the DPP for stochastic McKean–Vlasov control.
- Flow properties: The section proves a flow property for the controlled conditional distribution process, including measurability, square integrability, continuity, and compatibility with shifted controls.The construction uses a representative initial random variable and Picard iteration, followed by measurability and convergence arguments.
- Control formulation: The control formulation restricts controls to be progressively measurable with respect to the common-noise filtration, which makes the conditional law the sole state variable for dynamic programming.This is the formulation used to rewrite the cost functional in terms of the conditional distribution.
- Flow properties: The conditional-law process can be represented as a function of time, initial measure, control, and common-noise path, defining the state evolution on P2(Rd).The cost consequently depends on the initial random variable only through its conditional distribution, allowing the value function to be written as v(t, µ).
- Continuity: The cost functional is continuous in time, the initial measure, and the control, uniformly in the control for the first two variables; the value function is consequently continuous.Continuity in the control follows from continuity of the induced cost terms on P2(Rd)×A and P2(Rd).
- Dynamic programming principle: The resulting dynamic programming principle characterizes v through conditional costs up to stopping times and continuation values evaluated at the controlled conditional law.The stronger formulation provides an ε-optimal control independent of the stopping time, which is later useful for proving the viscosity supersolution property.
4 Bellman equation and viscosity solutions
The paper characterizes the stochastic McKean-Vlasov control value function through a Bellman equation, using Lions differentiability, lifted Hilbert-space methods, and Itô’s formula along conditional-measure flows.
- Differentiability on Wasserstein space: The authors introduce differentiation with respect to probability measures by lifting functions on P2(Rd) to functions on an L2 space.The lifted function depends on the distribution of the underlying random variable, enabling Fréchet-derivative representations of measure derivatives.
- Differentiability on Wasserstein space: Fully C2 regularity requires continuity and differentiability conditions for the measure derivative and its spatial and measure arguments.These conditions include continuity on P2(Rd) × Rd and existence of the spatial gradient of the measure derivative.
- Itô formula and Bellman equation: Itô’s formula along conditional-measure flows is obtained by representing the process and its coefficients on a lifted L2 space under suitable integrability and differentiability assumptions.The paper notes that the lifted formula requires twice continuously Fréchet differentiability, while the Wasserstein-space formula can hold more generally.
- Itô formula and Bellman equation: The dynamic programming Bellman equation is formulated for the stochastic McKean-Vlasov value function and analyzed through a lifting identification with L2(G; Rd).The Hilbert-space formulation is chosen because comparison principles are difficult to obtain directly in the Wasserstein space.
- Viscosity characterization: The viscosity framework is not intrinsic to Wasserstein space unless the lifted solution is twice continuously Fréchet differentiable, so the paper uses general L2 test functions.This extra regularity may fail for a smooth Wasserstein-space function, motivating the chosen lifted-space definition.
- Viscosity characterization: The value function is the unique continuous viscosity solution of the Bellman equation satisfying a quadratic growth condition.The viscosity characterization and uniqueness follow from the dynamic programming principle and a comparison principle for the lifted equation.
5 Linear quadratic stochastic McKean-Vlasov control
The section formulates and solves a linear-quadratic stochastic McKean-Vlasov control problem with common noise. Its value function is characterized through Riccati equations, yielding an optimal feedback control and an interbank systemic-risk application.
- Problem formulation: The model uses Lipschitz feedback controls for multivariate linear McKean-Vlasov dynamics with common and idiosyncratic noise.The drift and volatility coefficients depend linearly on the state, conditional mean, and control.
- Problem formulation: The running and terminal costs are quadratic in the state, conditional mean, and control.The formulation assumes symmetric cost matrices and distinguishes nonnegative and positive-definite matrices.
- Solution method: The quadratic value-function ansatz reduces the Bellman equation to Riccati equations for Λ and Γ, followed by linear ODEs for γ and χ.This reduction follows square completion and identification of variance, conditional-mean, and mean terms.
- Solution method: Under positive-definiteness conditions, verification identifies the value function with the quadratic ansatz and gives the optimal control in feedback form.The feedback mapping is Lipschitz under the stated solution conditions.
- Interbank systemic risk: The interbank application models log-monetary reserves with common noise and optimizes a common borrowing/lending feedback policy.The model penalizes borrowing or lending incentives and deviations from the conditional average, and admits an explicit Riccati-based solution.
- Interbank systemic risk: When σ1 = 0, the solution has Γ(t) = γ(t) = 0 and recovers the previously obtained borrowing/lending control expression.This specializes the common-noise model to the case without the σ1 contribution.