Source-linked AI summary

Mean-covariance turnpikes in Wasserstein distributionally robust linear-quadratic control

Yanzhi Wu, Zhengping Ji

arXiv:2608.26986v1math.OCeess.SY

TL;DR

The paper studies long-horizon Wasserstein-penalized minimax control when empirical disturbances induce adversarial mean and covariance dynamics beyond standard turnpike analyses. It characterizes the nonzero mean reference through a reduced Hamiltonian saddle problem, proves horizon-uniform exponential turnpikes for mean and covariance variables, and constructs a hybrid policy whose terminal-layer cost gap decays exponentially. These results support horizon-independent approximation of most control stages, while actual state-distribution turnpikes beyond Gaussian moment surrogates remain open.

  • Problem

    Empirical disturbance data can be inaccurate, and existing theory did not establish horizon-uniform robust turnpikes, especially for adversarially induced covariance dynamics.

  • Method

    The paper combines a reduced convex-concave Hamiltonian saddle problem for uncentered mean references with Riccati and covariance-recursion analyses.

  • Results

    The authors prove horizon-uniform two-sided exponential turnpike estimates for mean and covariance quantities and an exponentially decaying hybrid-policy cost gap in terminal-layer length.

  • Takeaways & Limitations

    For prescribed accuracy, time-independent affine feedback can be used over most of a long horizon while retaining finite-horizon coefficients only near the terminal boundary.

  • Takeaways & Limitations

    The paper leaves open strengthening moment-based estimates to 2-Wasserstein turnpikes for actual closed-loop state distributions.

Abstract

from arXiv · show

We study long-horizon Wasserstein-penalized minimax control for discrete-time stochastic linear systems with empirical disturbance data, in which adversarial disturbance distributions induce time-varying mean and covariance dynamics, making standard turnpike arguments not directly applicable. For possibly uncentered data, we characterize the generally nonzero mean reference through a reduced convex-concave Hamiltonian saddle problem. We prove horizon-uniform, two-sided exponential turnpike estimates for the mean state, adjoint, control, worst-case disturbance mean, and closed-loop covariance, showing that they spend the majority of time near static references when the horizon is long. We further construct a hybrid policy combining time-independent affine feedback with finite-horizon steering over a terminal layer, proving that its worst-case cost gap decays exponentially with the terminal-layer length uniformly in the horizon, which helps reducing the computation cost for long-horizon robust controls. Numerical examples illustrate the estimates and their dependence on the Wasserstein penalty.

I. INTRODUCTION

The paper addresses long-horizon Wasserstein-robust control with empirical, potentially uncentered disturbances, where standard turnpike analysis must account for adversarial mean and covariance dynamics. It develops static references, horizon-uniform turnpike estimates, and a hybrid policy with exponentially decaying cost gap.

  • I. INTRODUCTION: Finite empirical disturbance data can make nominal distributions inaccurate, potentially degrading controller performance and closed-loop safety.The setting uses Wasserstein ambiguity to hedge against errors in the empirical distribution.
  • I. INTRODUCTION: Existing work did not establish horizon-uniform turnpikes for robust closed loops or covariance dynamics induced by adversarial disturbances.The paper also studies how solutions depend on horizon length and how that dependence can reduce computational complexity.
  • I. INTRODUCTION: For uncentered data, the empirical disturbance mean generates affine value-function and policy terms, and the static mean reference is characterized by a reduced convex-concave Hamiltonian saddle problem.This avoids artificially recentering the empirical distribution.
  • I. INTRODUCTION: Two-sided exponential estimates, uniform in the horizon, keep the mean state, adjoint, control, and worst-case disturbance mean near a static reference outside boundary layers.The estimates follow from a uniform product estimate for time-varying closed-loop matrices.
  • I. INTRODUCTION: An exact covariance recursion yields a horizon-uniform two-sided exponential covariance turnpike and bounds for expected quadratic functions and Gaussian moment surrogates.The covariance analysis is performed relative to a limiting Lyapunov equation.
  • I. INTRODUCTION: The hybrid policy uses time-independent affine feedback before a terminal layer and finite-horizon coefficients within it, with worst-case cost gap decaying exponentially in layer length.A fixed terminal-layer length suffices for prescribed cost accuracy uniformly as the horizon grows.

C. Main result in brief

The main result compares finite-horizon optimal mean and covariance variables with static references defined by a Hamiltonian saddle problem and a limiting Lyapunov equation. Under stabilizability, observability, and penalty conditions, the resulting errors decay exponentially from both time boundaries with constants independent of horizon.

  • C. Main result in brief: The reference variables comprise the static mean state, adjoint, control, and worst-case disturbance mean, together with a covariance reference from a limiting Lyapunov equation.The finite-horizon quantities include optimal mean state, adjoint, control, worst-case disturbance mean, and closed-loop covariance.
  • C. Main result in brief: Under Assumptions 1, 2, and 3, constants independent of horizon and time provide exponential turnpike bounds for the optimal variables.The main brief combines mean and covariance estimates and explicitly includes initial and terminal boundary terms.
  • C. Main result in brief: The Riccati analysis supplies the finite-horizon coefficients and policies that characterize the minimax optimizers and support the turnpike estimates.The coefficients satisfy P_t = S_{N−t}, while the penalty condition is verified uniformly across horizons under a sufficient terminal-cost condition.
  • C. Main result in brief: The Riccati-based outer minimization has a unique control solution, and the associated worst-case empirical support point is unique and affine.These properties apply when the value function has quadratic form.

B. Finite-horizon Riccati recursion

The finite-horizon value function is generated by backward Riccati recursions and yields affine optimal control and worst-case disturbance policies. The Riccati sequence converges to a stabilizing positive-semidefinite steady state, while the empirical mean affects affine terms but not the quadratic coefficient or feedback matrix.

  • Finite-horizon value function: The value function has the assumed form, with coefficients satisfying terminal conditions and backward recursions.The terminal conditions are P_N = Q_f, s_N = 0, and z_N = 0.
  • Optimal policies: The unique optimal control is affine, with feedback matrix K_t and feedforward term k_t.The worst-case disturbance distribution is then specified through its support points.
  • Mean and covariance roles: The empirical mean affects affine recursion terms and feedforward policies, whereas the quadratic coefficient P_t and feedback matrix K_t do not depend on it.The static mean reference depends linearly on the empirical mean, while the reference covariance depends on the empirical covariance instead.
  • Riccati convergence: The Riccati sequence remains positive semidefinite and bounded and converges to a stabilizing positive-semidefinite solution P_ss of P = R(P).The finite-horizon coefficients satisfy P_t = S_{N−t}.
  • Riccati convergence: The sufficient condition Q_f ⪯ R(Q_f) verifies the one-step penalty condition uniformly over all horizons.This condition is sufficient for horizon-independent verification and is not claimed necessary for a fixed horizon.
  • Static mean reference: The reduced static Hamiltonian has a unique saddle point that determines the generally nonzero static mean reference, control, and worst-case disturbance mean.Its Cartesian strategy space avoids introducing a separate zero-sum game with a shared constraint.

B. Finite-horizon mean-state and adjoint equations

The finite-horizon analysis defines mean state, control, disturbance mean, and adjoint variables under the optimal policies, then compares them with the static saddle-point reference. Their discrepancy is represented by coupled forward-backward error equations with initial and terminal boundary data.

  • Mean variables: The closed-loop construction defines the mean state, mean control, worst-case disturbance mean, and finite-horizon mean adjoint.The mean adjoint is obtained from the affine-quadratic value function.
  • Static comparison: These time-varying mean variables are compared with the solution of the static Hamiltonian problem.The static reference is used as the benchmark for the turnpike analysis.
  • Finite-horizon equations: Taking expectations under the worst-case disturbance policy produces finite-horizon equations for the mean state and mean adjoint.The derivation uses the optimality relations and the affine-quadratic value-function coefficients.
  • Error system: Subtracting the static reference equations yields error dynamics with boundary conditions determined by the initial mean error and terminal vector Q_f m_e − p_e.The resulting system is organized into forward mean-state and backward adjoint-error components.

V. MEAN AND COVARIANCE TURNPIKE ESTIMATES

The turnpike proof combines Riccati convergence with a horizon-uniform product estimate for time-varying closed-loop matrices. This yields two-sided exponential bounds and a horizon-independent bound on the number of stages outside any prescribed neighborhood of the static mean reference.

  • Proof strategy: The analysis proceeds from Riccati convergence to a uniform product estimate, then applies it to mean and covariance recursions.The resulting estimates also support Gaussian moment surrogates and finite-horizon policy approximation.
  • Product estimates: The Riccati sequence converges exponentially to its stabilizing limit, and the associated closed-loop matrix products satisfy a horizon-uniform exponential bound.The product estimate holds for all 0 ≤ s ≤ t ≤ N with constants independent of N, s, and t.
  • Mean turnpike: The auxiliary variable r_t obeys a backward recursion, while the mean-state error e_t obeys a forward recursion driven by r_t.This forward-backward separation produces the two-sided mean estimate.
  • Mean turnpike: Theorem 2 gives constants independent of horizon and time for exponential estimates of the mean state and mean control variables.The estimates are based on initial and terminal boundary-layer terms.
  • Measure turnpike: The number of stages where the mean state, control, or worst-case disturbance mean leaves a prescribed neighborhood is bounded independently of the horizon.Increasing the horizon enlarges the interior portion without increasing the off-turnpike-stage bound.

C. Covariance turnpike

The covariance recursion is derived from conditional second moments and centered disturbance fluctuations, with a reference covariance obtained by replacing finite-horizon Riccati matrices by their steady-state limit. A separate theorem establishes a two-sided, horizon-uniform exponential covariance turnpike.

  • Covariance recursion: The covariance recursion follows by decomposing conditional means and fluctuations and using the vanishing of cross terms.The centered support points have zero mean, allowing the conditional second-moment contributions to be combined.
  • Reference covariance: The reference covariance is defined by replacing P_{t+1} with P_ss in the covariance recursion.Because A_c is Schur, the defining equation has a unique symmetric positive-semidefinite solution.
  • Covariance turnpike: The covariance turnpike estimate contains an initial term from X_0 − X_e and a terminal term from finite-horizon Riccati deviations.The terminal contribution reflects deviations of P_{t+1}, A_c(P_{t+1}), and D(P_{t+1}) from their steady-state counterparts.
  • Covariance turnpike: Theorem 3 establishes a two-sided exponential covariance estimate with constants independent of the horizon and time.The covariance analysis is separate from the mean turnpike proof.

D. Gaussian moment surrogates

The paper extends mean and covariance turnpike estimates to quadratic observables and Gaussian moment surrogates, without assuming the actual state distributions are Gaussian.

  • Quadratic functions of the closed-loop state inherit bounds from the mean and covariance turnpike estimates.
  • Proposition 3 provides horizon-uniform constants for the quadratic observable estimate under the stated assumptions and bounded initial-state conditions.The constants are independent of the horizon, time, and particular admissible initial distribution.
  • N(m, X) denotes the Gaussian probability measure with mean m and covariance X.
  • The Gaussian moment surrogates match the finite-horizon and reference means and covariances, respectively, while the actual closed-loop state distributions need not be Gaussian.
  • The mean and covariance turnpike estimates imply a Wasserstein bound between the Gaussian moment surrogates.
  • E. Interior approximation of the mean and covariance: Removing L stages from each horizon boundary yields an interior region where finite-horizon means and covariance remain close to their time-independent references.For t in the interior, the approximation error is bounded by a term that decreases exponentially with L.

F. Approximation using quasi-turnpike with error estimation

The paper uses turnpike estimates to replace most finite-horizon feedback coefficients with static affine feedback while retaining terminal steering, and evaluates the resulting approximation through error envelopes and numerical policy comparisons.

  • Hybrid policy: The hybrid policy applies K_ssx + k_ss for the first N − q stages and finite-horizon coefficients K_tx + k_t over the final q stages.q = 0 yields the static affine policy, while q = N recovers the finite-horizon optimal policy.
  • Performance guarantee: Proposition 4 establishes a horizon-uniform exponential bound on the hybrid policy’s worst-case finite-horizon cost gap for sufficiently large terminal-layer length q.The construction replaces part of the time-varying policy with static controls to reduce computational complexity.
  • Numerical setup: The numerical study uses a two-dimensional system with horizon N = 80 and an empirical distribution containing M = 12 equally weighted support points.The same support is used at every stage, and alternative supports with matching first two moments produce mean and covariance sequences agreeing to numerical precision.
  • Turnpike errors: Figures 1–3 plot mean, control, disturbance-mean, covariance, quadratic-function, Gaussian-surrogate, and measure-turnpike errors against their static references.The dashed curves are a posteriori numerical envelopes illustrating the two-sided exponential form rather than evaluations of abstract analytical constants.
  • Measure turnpike: The measure-turnpike plot counts stages whose mean-state error exceeds ε, and its decreasing off-turnpike fraction agrees with a horizon-independent bound on their number.The Gaussian measures G_t and G_e are moment surrogates, not the actual closed-loop state distributions.

C. Policy comparison and coefficient computation time

The numerical examples examine penalty sensitivity, turnpike errors, and the computational implications of replacing most finite-horizon coefficients with stationary feedback. Near-threshold penalties produce larger errors and sharp covariance sensitivity, while the six-state example retains two-sided turnpike behavior.

  • Policy comparison: Six terminal stages give a relative value gap below 10^-8 under the original uncentered empirical distribution, while d = 0 performs worse.The gap is numerical and is not itself bounded by Proposition 4.
  • Coefficient computation time: The coefficient-time measurements include only Riccati matrices and policy coefficients Kt and kt, excluding simulation, value evaluation, and complete dynamic-program runtime.Each measurement is repeated until accumulated time exceeds 0.15 seconds after two warm-up runs, with the median of 15 measurements reported.
  • Penalty sensitivity: The limiting penalty threshold is estimated by bλth ≈ 0.2330, where λmin(λI − Ξ⊤PssΞ) crosses zero.The experiment evaluates a grid of λ values and uses numerical bisection to estimate the crossing.
  • Turnpike examples: The six-state example preserves two-sided turnpike errors, although mean-state and mean-adjoint errors decay more slowly than in the two-dimensional example.The near-threshold penalty case has larger combined errors, while values near the center are excluded from quantitative comparison because they can reach numerical precision.
  • Turnpike structure: The main results approximate finite-horizon mean variables and covariance by time-independent references outside initial and terminal boundary layers.The approximation is stated for any prescribed tolerance.
  • Hybrid policy: The hybrid policy uses stationary affine feedback away from the terminal boundary and finite-horizon coefficients over q final stages, with a horizon-independent exponential cost-gap decay in q.For fixed accuracy, the number of time-dependent terminal coefficients need not grow with the horizon.

APPENDIX A PROOFS IN SECTION III

The appendix derives the affine-quadratic dynamic-programming recursion, establishes uniqueness of minimizing controls and maximizing disturbances, and proves convergence and stability properties of the associated Riccati sequence.

  • Dynamic-programming recursion: The disturbance maximizer is unique because its objective Hessian is negative definite under λI − Ξ⊤Pt+1Ξ ≻ 0.
  • Dynamic-programming recursion: The minimizing input is unique because the control objective is strictly convex, with positive-definite Hessian from R ≻ 0 and λI − Ξ⊤Pt+1Ξ ≻ 0.
  • Dynamic-programming recursion: Averaging the disturbance-support-point equations and matching state and constant coefficients yields the affine-quadratic value-function recursion.
  • Riccati convergence: The Riccati operator is equivalent to the standard discrete-time operator with auxiliary input matrix W^1/2, state weight Q, and input weight I.
  • Riccati convergence: Starting from S0 = Qf ⪰ 0, the Riccati sequence remains positive semidefinite and bounded and converges to the stabilizing solution Pss.
  • Riccati convergence: The associated closed-loop matrix is Ac = (I + WPss)^−1A and is Schur because Pss is stabilizing.
  • Uniform bounds: Compactness of the convergent Riccati sequence and positivity of the relevant eigenvalue and singular-value maps provide uniform lower bounds for the required inverses.

APPENDIX B PROOFS IN SECTION V

The appendix proves uniform exponential contraction for the time-varying closed-loop products and propagates it through adjoint, mean, control, disturbance-mean, and covariance error recursions.

  • Product estimates: Each sufficiently interior closed-loop factor contracts in the HL-induced norm by α ∈ (0, 1), while only the final JL factors require separate uniform bounds.
  • Product estimates: The resulting transition products satisfy a horizon-uniform bound ∥ΦN(t, s)∥ ≤ CLρ^(t−s) for any ρ ∈ (α, 1).
  • Adjoint and mean errors: Backward iteration of the adjoint error gives ∥rt∥ ≤ CΦρ^(N−t)∥re∥.
  • Disturbance mean and covariance: Lipschitz dependence of Ac(P) and D(P), together with the product estimate, yields an exponentially decaying worst-case disturbance-mean error near the terminal boundary.
  • Disturbance mean and covariance: Iterating the covariance error recursion and combining decay rates proves the closed-loop covariance turnpike estimate.

D. Proof of Corollary 2

The corollary transfers mean and covariance estimates to Gaussian moment surrogates and analyzes the hybrid policy through stationary-policy contraction and Bellman telescoping.

  • Gaussian moment surrogates: Coupling Gaussian laws with shared standard-normal noise bounds their 2-Wasserstein distance using mean and covariance errors.
  • Gaussian moment surrogates: The resulting Gaussian-law estimate absorbs the fixed terminal reference offset into a constant while retaining exponential decay.
  • Hybrid policy: The stationary-policy Bellman map has fixed point (Pss, sss), and Ac being Schur yields local contraction for the hybrid policy's coefficients.
  • Hybrid policy: Uniform Riccati-tail bounds and the Bellman inequality control the hybrid policy's pre-terminal-stage cost contributions.
  • Hybrid policy: Because the hybrid policy equals the optimal policy over the final q stages, Bellman telescoping and optimality provide the corresponding upper and lower cost-gap bounds.
Loading 2608.26986v1…