Source-linked AI summary

Learning in Mean Field Games: the Fictitious Play

Pierre Cardaliaguet, Saeed Hadikhanloo

arXiv:1507.06280v2math.OC

TL;DR

Mean Field Game equilibria are difficult for agents to compute directly, motivating the question of whether they can learn stable play. The paper introduces a Fictitious Play-like procedure that repeatedly best responds to an averaged population belief and proves convergence under potential-game structure, with scope limitations for first-order systems and finite-player interpretations.

  • Problem

    The paper asks whether agents in intricate Mean Field Games can learn equilibrium rather than assume they can compute the equilibrium configuration directly.

  • Method

    The paper defines a Fictitious Play-like procedure in which agents optimize against an averaged density belief, observe the resulting population evolution, and update that belief.

  • Results

    Any cluster point of the pre-compact learning sequence is a Mean Field Game solution under suitable assumptions, and monotonicity gives convergence of the full sequence to the unique solution.

  • Takeaways & Limitations

    The procedure provides a convergent learning framework for potential Mean Field Games, extending Fictitious Play to this setting.

  • Takeaways & Limitations

    For first-order Mean Field Games, classical solutions generally cannot be expected, so the learning procedure's solutions are not smooth.

Abstract

from arXiv · show

Mean Field Game systems describe equilibrium configurations in differential games with infinitely many infinitesimal interacting agents. We introduce a learning procedure (similar to the Fictitious Play) for these games and show its convergence when the Mean Field Game is potential.

1 Introduction

Mean Field Games model infinitesimal agents interacting with a large population, and this paper asks whether such agents can learn equilibrium through Fictitious Play. The proposed procedure updates population beliefs by averaging observed densities and converges under potential-game assumptions.

  • Motivation: Mean Field Games describe differential games in which each infinitesimal agent interacts with a large population.A solution pairs a value function for a typical player with the evolving population density.
  • Motivation: The paper studies whether agents can form an equilibrium by repeatedly forecasting population density, optimizing against that belief, observing the resulting density, and updating their average belief.This question arises because directly computing an equilibrium configuration may be unrealistic in an intricate game.
  • Related work: Fictitious Play selects a best response to the average of previous actions, but does not converge in every game; it does converge for several classes, including potential games.The paper adapts this learning idea to the Mean Field Game setting.
  • Learning procedure: The learning procedure computes an optimal control using the current average density belief, evolves the actual density through the Fokker-Planck equation, and then incorporates the observation into the next average.The procedure generates sequences of value functions and densities across stages.
  • Main result: Under suitable assumptions, every cluster point of the pre-compact sequence (u_n, m_n) solves the Mean Field Game system, while monotonicity yields convergence of the full sequence when the solution is unique.The convergence analysis relies on the coupling costs f and g deriving from potentials, making the system a potential game.
  • Scope: The paper presents its learning procedure as the first such procedure for Mean Field Games and treats both second-order and first-order systems in a periodic setting.The periodic assumption simplifies estimates and notation; the authors state that suitable extensions to other state spaces should not substantially change the result.

2 The Fictitious Play for second order MFG systems

The paper defines a Fictitious Play procedure for second-order Mean Field Games and proves convergence when the game derives from potentials. Under monotonicity, the entire sequence converges to the unique MFG solution.

  • Learning rule: The procedure alternates solving a Hamilton-Jacobi equation using the averaged density and a Fokker-Planck equation generating the next density.Players use the averaged population-density belief to compute an optimal control, observe the resulting density, and update the belief by averaging observations.
  • Potential structure: Convergence is analyzed under the potential-game condition that f and g are measure derivatives of potential functions F and G.The associated potential Φ links MFG solutions to minimizers and is almost decreasing along the learning sequence.
  • Convergence result: Any cluster point of the uniformly continuous sequence {(u^n, m^n)} is a solution of the second-order MFG system.The limiting averaged and realized densities coincide because they solve the same Fokker-Planck equation.
  • Convergence result: Under the monotonicity condition, the whole sequence converges to the unique solution of the second-order MFG system.Uniqueness reduces the compact sequence to a single accumulation point, which implies convergence of the full sequence.
  • Convergence proof: The proof combines uniform estimates, slow variation between iterations, and decay of the discrepancy between averaged and realized flows.The discrepancy sequence uniformly converges to zero, allowing passage to a cluster point of the pre-compact iterates.

3 The Fictitious Play for first order MFG systems

For first-order Mean Field Games, the paper defines a Fictitious Play over probability measures on curves and proves convergence under potential-game assumptions. Cluster points yield MFG solutions, while monotonicity ensures uniqueness and convergence of the full sequence.

  • The first-order MFG system: The first-order MFG system couples a Hamilton-Jacobi equation with a transport equation, interpreted using viscosity and distributional solutions rather than classical solutions.The density equation uses the drift −mDpH(x, ∇u), and classical solutions cannot generally be expected.
  • Uniqueness and finite-player extension: Under the monotonicity condition, the MFG solution is unique and the whole learning sequence converges to it; analogous finite-player accumulation points are MFG equilibria.For sufficiently many players and sufficiently late iterations, the finite-player process reaches any prescribed ε-neighborhood of the equilibrium.
  • The learning rule: The learning rule starts from an initial density or curve-measure belief, computes optimal responses, updates the population strategy distribution, and averages beliefs across stages.The curve formulation represents pure strategies as curves and mixed strategies as probability measures, with time marginals obtained by pushforward through evaluation maps.
  • The potential structure: The first-order Fictitious Play is analyzed when the coupling costs derive from potential functions on probability measures.The corresponding potential for the curve formulation is defined on probability measures over trajectories.
  • Convergence: Any cluster point of the pre-compact sequence of beliefs and optimal-response distributions coincides across the two distributions and induces a solution of the first-order MFG system.The limiting density is the time marginal of the limiting curve measure, while the limiting value function is defined through the limiting optimal-control cost.
  • Convergence: The potential is almost decreasing along the learning sequence, supporting the compactness and cluster-point argument used for convergence.The proof also establishes uniform Lipschitz bounds for optimal curves and identifies equality conditions needed in the limiting optimality argument.

A Well-posedness of a continuity equation

The section establishes existence and uniqueness for the continuity equation associated with the MFG system despite a generally nonsmooth vector field. The proof combines approximation, optimal-control properties, and Ambrosio’s superposition principle.

  • Analytic framework: The limiting value function is semiconcave, and a measurable selection of reachable gradients defines the vector field used in the continuity equation.This selection is used in the distributional formulation and in the subsequent uniqueness argument.
  • Well-posedness: The continuity equation admits a unique absolutely continuous solution.The solution is understood in the distributional sense and is also bounded.
  • Well-posedness: The associated transport vector field may be discontinuous, so uniqueness requires several complementary arguments.The main difficulty is that the field −DpH(t, x, ∇u) is not generally smooth.
  • Approximation: A bounded solution is first obtained through a perturbation argument using classical solutions of a regularized system.For ε > 0, the approximating pair (uε, mε) is a unique classical solution, with uniform bounds supporting passage to the limit.
  • Optimal trajectories: Optimal trajectories are characterized by the Hamiltonian dynamics, and differentiability of the value function is equivalent to uniqueness of the optimal trajectory.Conversely, solutions of the characteristic differential equation are optimal, providing the optimal synthesis used in the uniqueness proof.
  • Superposition principle: Ambrosio’s superposition principle represents weak transport solutions as probability measures on trajectories solving the characteristic ODE.Disintegration with respect to the initial condition, combined with uniqueness of optimal trajectories almost everywhere, yields uniqueness of the transported measure.
Loading 1507.06280v2…