Source-linked AI summary

Concurrent learning-based approximate optimal regulation

Rushikesh Kamalapurkar, Patrick Walters, Warren Dixon

arXiv:1304.3477v1eess.SYmath.OC

TL;DR

The paper addresses PE requirements in online approximate optimal regulation by evaluating Bellman errors at selected state-space points and identifying uncertain drift parameters through concurrent learning. It establishes UUB convergence of the states and developed policy toward their respective targets using Lyapunov analysis.

  • Problem

    Online RL-based approximate optimal control commonly requires PE for parameter convergence, while regulation and learning near the target create additional challenges.

  • Method

    The method evaluates approximate Bellman errors at desired state-space points and uses a concurrent-learning identifier to estimate uncertain plant dynamics.

  • Results

    UUB convergence of the system states to the origin and of the developed policy to the optimal policy is established using Lyapunov-based analysis.

  • Takeaways & Limitations

    PE is replaced by a weaker rank condition that can be verified online from recorded data, while state-space Bellman-error evaluation supports learning from selected points.

Abstract

from arXiv · show

In deterministic systems, reinforcement learning-based online approximate optimal control methods typically require a restrictive persistence of excitation (PE) condition for convergence. This paper presents a concurrent learning-based solution to the online approximate optimal regulation problem that eliminates the need for PE. The development is based on the observation that given a model of the system, the Bellman error, which quantifies the deviation of the system Hamiltonian from the optimal Hamiltonian, can be evaluated at any point in the state space. Further, a concurrent learning-based parameter identifier is developed to compensate for parametric uncertainty in the plant dynamics. Uniformly ultimately bounded (UUB) convergence of the system states to the origin, and UUB convergence of the developed policy to the optimal policy are established using a Lyapunov-based analysis, and simulations are performed to demonstrate the performance of the developed controller.

I. INTRODUCTION

The paper targets online approximate optimal control methods whose parameter convergence typically depends on PE, proposing concurrent learning and state-space Bellman-error evaluation to replace that requirement.

  • Online RL-based optimal control updates value-function parameters using the Bellman error, paralleling parameter adaptation in adaptive control.
  • PE is commonly required for parameter convergence, but it is often impossible to verify online and does not guarantee convergence under boundedness modifications alone.
  • Concurrent learning can provide parameter convergence without PE by using recorded state information and model-error-related updates.
  • RL-based online regulation faces coupled challenges: obtaining parameter convergence while regulating states near the goal for local value-function learning.
  • The proposed approach evaluates the Bellman error at selected state-space points and uses a concurrent-learning identifier for uncertain drift dynamics.

II. PROBLEM FORMULATION

The paper formulates infinite-horizon optimal regulation for a nonlinear control-affine system, approximating the HJB solution by minimizing a Bellman error computed from estimated value, policy, and dynamics.

  • The objective is to find an optimal feedback policy that regulates the system state to the origin.
  • The plant has unknown locally Lipschitz drift dynamics and known bounded locally Lipschitz control effectiveness.
  • The optimal value function is characterized through the Hamilton-Jacobi-Bellman equation, whose analytical solution is generally infeasible.
  • The Bellman error replaces the exact value function and optimal policy in the HJB equation with their estimates.
  • The estimated value function and policy are adjusted to minimize the Bellman error, while an adaptive identifier supplies an estimate of the unknown drift.

III. SYSTEM IDENTIFICATION

The identification scheme linearly parameterizes the unknown drift, estimates its parameters, and uses observer error dynamics to support concurrent learning-based estimation.

  • The unknown drift is represented as f(x) = Y(x)θ∗, where Y(x) is a regression matrix and θ∗ contains constant unknown parameters.
  • The estimated drift uses the same regression matrix with the parameter estimate ˆθ.
  • The identifier uses an observer with state estimate ˆx and a positive definite diagonal gain to form the state estimation error.
  • The parameter identification error is defined as ˜θ = θ∗ − ˆθ.

A. Concurrent learning-based parameter update

Concurrent learning replaces PE with a finite-data rank condition: recorded state information is stored and used in the parameter update, with derivative-estimation error determining the ultimate bound.

  • The concurrent-learning assumption requires a finite set of recorded time instances whose regression data satisfy a rank condition.
  • The rank condition is weaker than PE because finite-period state excitation suffices, and it can be verified online from past states.
  • Selected states and corresponding controls are recorded in a history stack for the concurrent-learning update.
  • The update law uses recorded data and requires estimates of past state derivatives, which can be obtained through numerical smoothing.
  • Derivative-estimation errors lead to uniformly ultimately bounded parameter-estimation errors, with the ultimate-bound size depending on the derivative-estimation error.

B. Convergence analysis

The Lyapunov-based analysis establishes convergence properties for the concurrent learning observer and its parameter estimates. These estimates are then used to approximately solve the HJB equation without known drift dynamics.

  • B. Convergence analysis: The candidate Lyapunov function provides the basis for analyzing observer and parameter-estimation error convergence.The analysis uses eigenvalue bounds and the Lyapunov derivative to establish the required stability properties.
  • B. Convergence analysis: The concurrent learning-based observer yields exponential regulation of parameter and state-derivative estimation errors.The result holds under the bounded-trajectory condition used in the convergence analysis.
  • B. Convergence analysis: The resulting parameter and state-derivative estimates enable approximate HJB solution without knowledge of the drift dynamics.The estimates are used in the subsequent approximate optimal-control development.

IV. APPROXIMATE OPTIMAL CONTROL

The system identifier supplies an approximate Bellman error for the optimal-control design. This approximate error is then used to obtain an approximate solution to the HJB equation.

  • IV. APPROXIMATE OPTIMAL CONTROL: The system identifier is used to approximate the Bellman error in the optimal-control formulation.The approximation incorporates the estimated dynamics and control-dependent cost terms.
  • IV. APPROXIMATE OPTIMAL CONTROL: The approximate Bellman error is used to obtain an approximate solution to the HJB equation.This connects the identified plant model to the value-function and policy design.

A. Value function approximation

The paper represents the optimal value function and policy with neural networks on a compact state set. Separate weight estimates support stability analysis and a Bellman-error update linear in the value-function weights.

  • A. Value function approximation: The neural-network representations approximate the optimal value function and optimal policy on a compact state subspace.The compactness assumption is used for the neural-network representation and subsequent stability analysis.
  • A. Value function approximation: The ideal value-function weights are bounded, while the activation function and reconstruction errors satisfy stated boundedness conditions.The assumptions specify bounds on the ideal weights, activation derivatives, and reconstruction errors.
  • A. Value function approximation: Two sets of estimated weights represent the same ideal weights to support stability analysis and a Bellman-error update linear in the value-function weights.This structure enables a least-squares-based adaptive update law.

B. Learning based on desired behavior

The learning scheme evaluates approximate Bellman errors at presampled state-space points selected according to desired behavior, rather than relying only on trajectory data. Its rank condition is weaker than PE but depends on online parameter estimates and may be encouraged by redundant sampling.

  • B. Learning based on desired behavior: Approximate Bellman errors can be evaluated at presampled points in the state space when an estimated system model is available.This permits learning from selected points rather than only from states observed along trajectories.
  • B. Learning based on desired behavior: The value-function weights are updated by a concurrent learning least-squares law with normalization, forgetting, saturation, and adaptation gains.The policy weights are separately updated to follow the value-function weights.
  • B. Learning based on desired behavior: The learning condition is weaker than PE and can be verified online, but its satisfaction cannot generally be guaranteed a priori because it depends on estimated parameters.The paper suggests selecting many more points than neurons to heuristically meet the rank condition.
  • B. Learning based on desired behavior: For regulation to the origin in deterministic systems, bounded points uniformly distributed around the origin are a natural selection.The selection reflects prior information about the desired behavior of the system.

V. STABILITY ANALYSIS

The stability analysis uses a Lyapunov function to establish uniform ultimate boundedness for the combined state, estimation, parameter, value-function, and policy-weight errors under stated gain conditions. The analysis also shows that bounded initial conditions yield a compact set containing the trajectories, supporting the neural-network approximation framework.

  • Stability analysis: The approximate Bellman error is rewritten using value-function and policy weight estimation errors to support the Lyapunov analysis.This representation connects the learning errors to the stability proof.
  • Assumptions: The stability bounds rely on Lipschitz properties of the drift dynamics and an auxiliary function on the compact set.The analysis notes that these bounds can be generalized using state-dependent nondecreasing functions.
  • Stability analysis: Theorem 1 establishes the main stability result when the assumptions and sufficient gain conditions hold.The theorem is introduced after defining the Lyapunov candidate and the required inequalities.
  • Stability conclusion: The combined vector of system, observer, parameter, value-weight, and policy-weight errors is uniformly ultimately bounded.The proof invokes Theorem 4.18 after establishing positive lower bounds for the relevant coefficients.
  • Lyapunov analysis: The Lyapunov derivative is bounded using the approximate Bellman errors, error bounds, and Young's inequality, then completed-square inequalities yield positivity conditions.These steps provide the inequalities needed to invoke a UUB theorem.
  • Stability conclusion: The resulting boundedness removes the practical effect of the compactness assumption by defining a compact set that contains the system trajectories for all future time.The trajectory set is obtained from the Lyapunov boundedness argument and bounded initial conditions.

VI. CONCLUSION

The paper develops an online approximate optimal controller that removes the need for persistence of excitation by using recorded-state data and Bellman-error approximations based on estimated dynamics. The analysis establishes UUB convergence of both the system states and the learned policy.

  • VI. CONCLUSION: The proposed RL-based online approximate optimal controller does not require persistence of excitation for convergence.The method replaces PE with a weaker rank condition verifiable from recorded data.
  • VI. CONCLUSION: The controller improves value-function approximation using Bellman errors evaluated at pre-sampled desired states with estimated system dynamics.A concurrent-learning parameter identifier compensates for uncertainty in the plant dynamics.
  • VI. CONCLUSION: The analysis establishes UUB convergence of the system states to the origin and of the learned policy to the optimal policy.These convergence properties are established using Lyapunov-based analysis.
Loading 1304.3477v1…