Source-linked AI summary
On the invariance of risk-sensitive LQR gain under input randomization
Yeongjun Jang
TL;DR
Risk-sensitive LQR loses certainty equivalence, raising the question of whether deliberately randomized inputs alter its optimal gain. The paper analyzes the randomized formulation and shows that the gain and Riccati recursion remain unchanged because input noise affects both dynamics and cost, while the cost increment is computable in closed form.
Problem
Risk-sensitive LQR depends on process-noise statistics, so the effect of deliberate input randomization on its optimal gain and cost requires analysis.
Method
The paper formulates input-randomized risk-sensitive LQR with Gaussian input noise and accounts for that noise in both the dynamics and actual-input cost.
Results
The optimal gain and Riccati recursion are invariant under input randomization, while the optimal-cost increment can be evaluated in closed form.
Takeaways & Limitations
Input-randomized risk-sensitive LQR can use the original optimal gain without recomputing the Riccati recursion, including applications using randomization for privacy or exploration.
Abstract
from arXiv · showhide
This paper shows that the optimal gain of the risk-sensitive linear quadratic regulator (LQR) problem is invariant under input randomization, i.e., when the controller deliberately injects noise into the nominal control input. This appears counterintuitive at first glance because certainty equivalence does not hold for risk-sensitive LQR and input randomization inflates the effective process noise. Nonetheless, the gain is preserved because the input noise enters not only the system dynamics but also the cost functional, and its total effect on the gain eventually vanishes. Consequently, the optimal gain and its associated Riccati recursion need not be recomputed, and the increment in the optimal cost can be readily evaluated in closed form. This result facilitates the use of risk-sensitive LQR in applications that employ input randomization for privacy or exploration, such as watermarking for replay attack detection, differential privacy, and path integral control.
I. INTRODUCTION
Risk-sensitive LQR extends standard LQR by penalizing low-probability, high-cost trajectories, but loses certainty equivalence because its gain depends on process-noise statistics. The paper shows that input randomization nevertheless leaves the optimal gain and Riccati recursion unchanged, while increasing cost in a tractable way.
- Risk-sensitive LQR minimizes the expected exponential quadratic cost, assigning heavier weight to low-probability high-cost trajectories.
- Unlike standard LQR, risk-sensitive LQR has a noise-dependent Riccati recursion, so its optimal gain must be recomputed when process-noise statistics change.This loss of certainty equivalence can make repeated recomputation burdensome.
- Input randomization deliberately adds noise before the control input reaches the plant and is used in watermarking, differential privacy, and path integral control.
- The paper shows that input randomization leaves the risk-sensitive LQR optimal gain and associated Riccati recursion invariant.The result does not follow from certainty equivalence; the input noise enters both dynamics and the cost functional, so its total effect on the gain vanishes.
- Although input randomization increases optimal cost, the increment is available in closed form and joint feasibility holds when risk sensitivity is sufficiently small.The paper also gives a zero-sum dynamic-game interpretation.
II. REVIEW OF RISK-SENSITIVE LQR
The paper formulates finite-horizon risk-sensitive LQR for linear dynamics with Gaussian process noise and exponential quadratic cost. Its linear feedback solution is characterized by a backward Riccati recursion whose feasibility and gain depend on noise covariance and risk sensitivity.
- The finite-horizon plant has state x_t, input u_t, and mutually independent zero-mean Gaussian process noise w_t with covariance W_t.
- The controller minimizes an exponential transformation of a finite-horizon quadratic cost with state and input weight matrices Q_k ⪰ 0 and R_k ≻ 0.
- The risk-sensitivity parameter θ controls emphasis on high-cost trajectories: θ > 0 is risk-averse, θ < 0 is risk-seeking, and θ → 0 recovers risk-neutral LQR.
- Under a well-defined backward Riccati recursion satisfying S_t ≻ 0, the optimal control law has linear state-feedback form and the cost-to-go is quadratic.
- Because W_t explicitly enters the Riccati recursion, risk-sensitive LQR lacks certainty equivalence and requires gain recomputation when process-noise statistics change.The feasibility condition becomes more restrictive as process-noise intensity or positive risk sensitivity increases.
A. Input-randomized risk-sensitive LQR problem
The input-randomized problem separates nominal control from deliberately injected Gaussian noise that is independent of process noise and cannot be compensated directly. The injected noise becomes effective process noise, while the cost penalizes the actual applied input.
- The applied control is decomposed into nominal input and deliberately injected noise, with the nominal input chosen before the noise is realized.
- The injected input noise e_t is mutually independent zero-mean Gaussian with covariance Σ_t and is independent of the process noise w_t.
- After input randomization, w_t + B_t e_t acts as effective process noise in the rewritten dynamics.
- The randomized problem takes expectation over both process and input noise and penalizes the actual applied input, nominal input plus injected noise.
- The input-randomized risk-sensitive LQR retains a cost-to-go formulation with terminal condition V̄_N(x) = x^⊤Q_Nx.
B. Invariance of the optimal gain
Under the stated well-definedness and positive-definiteness assumptions, input randomization preserves the nominal linear state-feedback gain and Riccati matrix recursion. The resulting optimal cost-to-go remains quadratic in the state, with an additional recursively determined constant.
- B. Invariance of the optimal gain: Theorem 1 establishes that the optimal nominal control law retains a linear state-feedback form under input randomization.The result assumes the original backward Riccati recursion is well-defined and condition (9) holds.
- B. Invariance of the optimal gain: The optimal cost-to-go has the form ¯V_t(x) = x⊤P_tx + ¯η_t, preserving the quadratic state term while adding a scalar offset.The offset ¯η_t is determined by a backward recursion.
- B. Invariance of the optimal gain: The proof uses backward induction, Bellman’s principle of optimality, and Gaussian-expectation identities to derive the randomized problem’s recursion.The argument treats θ > 0 explicitly and states that θ < 0 follows analogously.
- B. Invariance of the optimal gain: Positive definiteness of the relevant Hessian ensures that the minimizer is uniquely attained at the linear feedback control.The proof establishes ˜G_t ≻ 0 from ˜H_t ≻ 0 before identifying the unique minimizer.
- B. Invariance of the optimal gain: Input-noise covariance does not enter the Riccati recursion or alter the original optimal gain because its dynamic and cost effects reshape curvature while leaving the minimizer unchanged.Thus, randomization changes the cost-to-go through a constant rather than through the state-dependent quadratic term.
C. Cost increment and joint feasibility
Input randomization increases the optimal cost but leaves the risk-sensitive LQR gain and Riccati recursion unchanged. Feasibility requires an additional condition, and both conditions are jointly satisfied for sufficiently small risk sensitivity.
- Cost increment: The optimal gain remains invariant, while input randomization changes the optimal cost through an additive term in its recursion.The cost increment can be computed without recomputing the Riccati recursion.
- Cost increment: The original and input-randomized optimal costs differ by a positive logarithmic determinant increment.The increment is expressed as −1/(2θ) times the sum of log determinants involving Ξ_t and Σ_t.
- Joint feasibility: Input randomization introduces an additional feasibility condition that restricts the admissible input-noise intensity and can fail only when θ > 0.This condition is coupled with the original risk-sensitive LQR feasibility condition.
- Joint feasibility: The feasible risk-sensitivity set has the form Θ = (−∞, θ̄) \ {0} for some θ̄ > 0.Thus, feasibility holds whenever the risk-sensitivity parameter is sufficiently small.
- Joint feasibility: The joint feasibility conditions are monotone in θ: if they hold at a given value, they remain valid for all smaller values.For a fixed θ, verifying feasibility is direct once the Riccati recursion has been computed.
D. Equivalent formulation as a zero-sum dynamic game
The input-randomized risk-sensitive LQR problem is equivalent to a zero-sum dynamic game with the same Riccati recursion and optimal control law. The game interpretation explains why deviating from the invariant gain exposes the controller to adversarial exploitation.
- Equivalent formulation: The input-randomized risk-sensitive LQR problem admits an equivalent zero-sum dynamic-game formulation sharing its Riccati recursion and optimal control law.This equivalence provides an intuitive interpretation of gain invariance.
- Game structure: For θ > 0, the game involves a controller minimizing cost and an auxiliary adversary maximizing it through state and input disturbances.The adversary’s disturbances are penalized quadratically rather than introduced by exponentiating the trajectory cost.
- Game structure: As θ increases, the disturbance penalties weaken, allowing stronger adversarial disturbances and representing greater risk aversion.The formulation treats risk sensitivity through adversarial disturbance penalties.
- Saddle point: The game has a saddle point under the Riccati well-definedness and feasibility conditions.The optimal value function retains the quadratic form xᵀP_t x.
- Gain invariance: At the optimal gain, the adversary has no incentive to inject input noise, whereas deviations allow a nonzero best-response disturbance that can exploit the controller.This gives a game-theoretic explanation for why the controller does not deviate from the invariant gain.
IV. CONCLUDING REMARKS
The paper establishes gain invariance under input randomization, derives the resulting cost increment and feasibility characterization, and provides a zero-sum dynamic-game interpretation.
- Contributions: The optimal gain of risk-sensitive LQR is invariant under input randomization.The paper also derives a closed-form cost increment and characterizes feasibility for sufficiently small risk-sensitivity parameters.
- Interpretation and future work: An equivalent zero-sum dynamic-game formulation provides further intuition for the gain invariance.Future work includes infinite-horizon extensions and risk-sensitive LQG invariance.
APPENDIX I TECHNICAL LEMMA
The technical lemma evaluates a Gaussian exponential-quadratic expression under a positive-definiteness condition and establishes positivity of the resulting transformed matrix.
- Assumptions: For Gaussian w ∼ N(0, W), the lemma assumes P ⪰ 0 and S := W⁻¹ − 2θP ≻ 0.Under these conditions, the exponential-quadratic Gaussian expression is well-defined.
- Proof strategy: The proof uses the Gaussian density, completion of the square, determinant identities, and the resulting Gaussian integral.The integrand becomes the density of N(2θS⁻¹Pz, S⁻¹).
- Proof strategy: Positive definiteness of S and W⁻¹ establishes ˜P ⪰ 0 through a congruence and eigenvalue argument.The proof factors ˜P using I + 2θMMᵀ and shows this factor is positive definite.
APPENDIX II PROOF OF LEMMA 1
The proof uses backward induction to establish a quadratic value function with a positive-semidefinite cost matrix. Completing the square yields a unique linear minimizer, while a Schur-complement argument proves positivity is preserved.
- The proof establishes by backward induction that P_t ⪰ 0 and V_t(x) = x⊤P_t x + η_t for every stage.The base case uses P_N = Q_N ⪰ 0 and η_N = 0.
- Completing the square in u and using monotonicity of the exponential reduces the Bellman minimization to the unique control u∗ = K̃_t x.Positive definiteness of H̃_t follows from P̃_{t+1} ⪰ 0 and R_t ≻ 0.
- Substituting the minimizing control back into the Bellman equation preserves the quadratic-plus-constant form of V_t.
- The block matrix Γ_t is positive semidefinite because Q_t ⪰ 0 and R_t ≻ 0; its Schur complement therefore proves P_t ⪰ 0.This completes the induction and the proof for θ > 0; the θ < 0 case is analogous.