Source-linked AI summary
Learning Model Predictive Control for iterative tasks. A Data-Driven Control Framework
Ugo Rosolia, Francesco Borrelli
TL;DR
The paper addresses repetitive control tasks without a known reference trajectory, while seeking constraint satisfaction and non-increasing performance across iterations. It proposes a reference-free LMPC that learns a safe set and terminal cost from previous trajectories. The resulting controller guarantees recursive feasibility and stability under its assumptions, and simulations demonstrate convergence and performance improvements in linear and nonlinear tasks.
Problem
Repetitive tasks may lack a known reference trajectory, while nonlinear MPC alone can become infeasible and does not guarantee improved performance at every iteration.
Method
The LMPC recursively constructs a terminal safe set and terminal cost from successful previous state and input trajectories, then uses them in constrained predictive control.
Results
The controller guarantees recursive feasibility, asymptotic stability, and non-increasing iteration cost; simulations show convergence to the optimal solution in a constrained linear quadratic regulator and effective nonlinear minimum-time control.
Takeaways & Limitations
Learning from successful iterations enables reference-free control while preserving constraints and improving closed-loop performance across iterations.
Takeaways & Limitations
The theoretical guarantees are established only for deterministic systems, and the sampled-set terminal constraint can make the optimization computationally expensive.
Abstract
from arXiv · showhide
A Learning Model Predictive Controller (LMPC) for iterative tasks is presented. The controller is reference-free and is able to improve its performance by learning from previous iterations. A safe set and a terminal cost function are used in order to guarantee recursive feasibility and non-increasing performance at each iteration. The paper presents the control design approach, and shows how to recursively construct terminal set and terminal cost from state and input trajectories of previous iterations. Simulation results show the effectiveness of the proposed control logic.
I. INTRODUCTION
The paper develops a reference-free learning MPC for repetitive tasks where the reference trajectory is unknown. It uses trajectories from prior iterations to seek non-increasing cost while preserving constraint satisfaction and stability.
- Motivation: Reference-based iterative controllers assume a known, unchanged reference and primarily minimize tracking error under disturbances.
- Motivation: Reference-free control is motivated by tasks whose optimal trajectory is difficult to compute because dynamics, parameters, or human inputs may be uncertain or variable.
- Contributions: The proposed LMPC learns from previous iterations while keeping the initial condition, constraints, and objective function fixed.
- Contributions: A learned terminal safe set and terminal cost guarantee non-increasing iteration cost, recursive feasibility, and asymptotic stability.
- Problem formulation: The paper defines each iteration as a closed-loop trajectory starting from the same initial state and evaluates its performance through an infinite-horizon cost.
A. Sampled Safe Set
The sampled safe set stores states from successful previous trajectories that can reach the terminal equilibrium under constraints. The iteration cost and terminal cost are computed from these stored trajectories for learning MPC.
- Controllable and stabilizable sets: Stabilizable sets contain states that can reach a control-invariant target while satisfying constraints and then remain there indefinitely.
- Sampled Safe Set: Each state in SSj belongs to the maximal stabilizable set when its trajectory successfully steers the system to the terminal point xF.
- Sampled Safe Set: The sampled safe set SSj collects state trajectories from successful iterations up to iteration j.
- Learning mechanism: Because successful trajectories accumulate across iterations, the stored safe set supports the later proof of non-increasing iteration cost.
- Iteration Cost: The iteration cost is the infinite-horizon cost of the realized trajectory and quantifies controller performance at iteration j.
- Terminal Cost: Qj assigns each sampled-safe-set state the minimum cost-to-go among stored trajectories passing through that state.
III. LMPC CONTROL DESIGN
At each time step, LMPC solves a finite-horizon constrained problem whose terminal state lies in the previous sampled safe set, then applies the first control and repeats. Under the initial feasibility assumption, this construction preserves feasibility and stability across iterations.
- LMPC Formulation: LMPC approximates the infinite-horizon problem by repeatedly solving a finite-time constrained optimal control problem in receding-horizon form.
- LMPC Formulation: The optimization enforces system dynamics, state constraints, and input constraints over the prediction horizon.
- LMPC Formulation: Its terminal constraint places the predicted terminal state in SSj−1, using successful trajectories from prior iterations as the terminal set.
- Receding-horizon implementation: Only the first optimal input is applied before the problem is resolved at the next time step using the new state.
- Recursive feasibility and stability: Under Assumption 1, LMPC is feasible for all times and iterations, and xF is asymptotically stable at every iteration.
- Recursive feasibility and stability: The same assumptions also support non-increasing iteration cost across successive iterations.
B. Recursive feasibility and stability
The LMPC uses previous trajectories through its sampled safe set and terminal cost to preserve feasibility across time and iterations. Under the stated assumption, it also yields asymptotic stability of the equilibrium.
- Recursive feasibility: The sampled safe set and terminal cost constructed from prior trajectories support recursive feasibility and stability of the LMPC.The paper uses these objects to establish feasibility at every time and iteration and stability of xF.
- Stability: Under Assumption 1, the equilibrium point xF is asymptotically stable at every iteration while the LMPC remains feasible.Theorem 1 combines recursive feasibility with asymptotic stability for all t ≥ 0 and j ≥ 1.
- Recursive feasibility: Feasibility at iteration start follows from the non-empty previous safe set containing a feasible trajectory.The proof uses SS0 ⊆ SSj−1 and the common initial state to construct a feasible solution at t = 0.
- Recursive feasibility: The receding-horizon solution supplies a feasible shifted trajectory at the next time step, preserving state and input constraints.Induction then establishes feasibility for all j ≥ 1 and t ≥ 0.
- Stability: The optimal finite-horizon cost is a decreasing Lyapunov function along the closed-loop trajectory.Positive definiteness of the stage cost and continuity of the optimal cost support the stability argument.
C. Convergence properties
The paper establishes that iteration cost does not increase and analyzes the limiting trajectory as the controller learns. Under strict convexity and convergence, the steady-state trajectory is optimal for finite-horizon problems of arbitrary length; without convexity, only local properties are shown.
- Iteration-cost convergence: The iteration cost does not increase as the iteration index grows.This result is derived for the realized trajectories generated by the LMPC.
- Steady-state optimality: The finite-horizon comparison problem uses the same running cost, dynamics, and state and input constraints as the infinite-horizon problem.Its horizon may be longer than the LMPC horizon, while its terminal set may be smaller.
- Scope and assumptions: Strict convexity of the infinite-horizon problem is assumed for the global optimality result.For non-convex problems, the paper states that only local optimality can be shown under additional strictness assumptions.
- Steady-state analysis: The convergence analysis assumes that the LMPC converges to a steady-state input and trajectory.Under this assumption, the sampled safe set and terminal cost also converge to steady-state objects.
- Steady-state optimality: The steady-state trajectory is optimal for the finite-horizon problem with initial condition x∞t for every horizon T > 0.Theorem 3 extends the result from the controller horizon to arbitrary finite horizons under strict convexity.
A. Constrained LQR controller
The LMPC is tested on a constrained infinite-horizon linear-quadratic regulator, using an initially feasible trajectory to construct the sampled safe set and terminal cost. Across iterations, the controller maintains feasibility, achieves non-increasing cost, and converges extremely close to the exact optimal solution.
- Setup: The CLQR experiment initializes LMPC with a feasible trajectory, sampled safe set SS0, and terminal cost Q0(·).The feasible trajectory is generated by driving the system near the origin and then applying unconstrained LQR feedback.
- Setup: The controller uses a quadratic running cost, horizon length N = 4, and the CLQR state and input constraints.The implementation reformulates the LMPC as a mixed-integer quadratic program.
- Learning mechanism: Each subsequent closed-loop trajectory enlarges the sampled safe set used by the next iteration.The safe set is updated from the trajectories generated during learning.
- Results: After 9 iterations, the LMPC converges to a steady-state solution that is globally optimal for this CLQR instance.The converged solution saturates both state and input constraints, matching the exact solution.
- Results: The iteration cost is non-increasing, and the LMPC improves closed-loop performance across iterations.These properties are reported for the constrained LQR experiment and summarized in Table II.
- Results: 1.62 × 10^-5 is the maximum deviation between the LMPC steady-state trajectory and the exact CLQR optimum.The 2-norm difference between the exact optimal cost and the converged trajectory cost is 1.565 × 10^-20.
B. Dubins Car with Obstacle and Acceleration Saturation
The LMPC is applied to a minimum-time Dubins car with an obstacle and known acceleration saturation. Starting from a feasible trajectory, it learns a lower-cost steady-state maneuver that matches an independently computed locally optimal solution.
- Problem setup: The Dubins car is controlled from xS to the unforced equilibrium xF by minimizing an infinite-horizon minimum-time objective.The state contains planar position and velocity; acceleration and steering are the control inputs.
- Problem setup: The experiment includes a known acceleration saturation limit and an elliptical obstacle constraint.The saturation limit is s = 1, and the trajectory must remain outside the obstacle ellipse.
- Initialization: A brute-force feasible trajectory initializes the sampled safe set SS0 and terminal cost Q0(·) for LMPC.The target is xF = [54, 0, 0]T, with obstacle parameters ae = 8 and be = 6.
- Results: After 4 iterations, the LMPC converges to the steady-state solution, while the iteration cost decreases until convergence.The cost evolution is reported in Table III.
- Results: The learned controller accelerates near the saturation boundary until the midpoint, then decelerates to reach xF with zero velocity.This bang-bang-like behavior is consistent with the minimum-time objective.
- Validation: An interior-point nonlinear solver initialized with the LMPC trajectory produces the same locally optimal solution.This independently checks the local optimality of the learned steady-state trajectory.
C. Dubins Car with Obstacle and Unknown Acceleration Saturation
The LMPC is extended to a Dubins car whose acceleration saturation is unknown by augmenting the system with a saturation estimator and error dynamics. Across iterations, it learns the saturation coefficient while converging to the minimum-time maneuver.
- Adaptive formulation: The unknown-saturation experiment augments the Dubins-car system with an estimated saturation coefficient and an estimation-error state.The controller simultaneously steers the vehicle to xF and estimates the unknown coefficient.
- Adaptive formulation: The augmented state includes estimated position, velocity, saturation coefficient, and estimator error; inputs include acceleration, steering, and estimate updates.The augmented dynamics and terminal constraint are incorporated into the LMPC problem.
- Adaptive formulation: Each iteration measures the system state and estimates the saturation error by inverting the system dynamics.The first element of the optimized input sequence is applied at each time step.
- Optimality condition: If the estimated error is zero for every future time step, the augmented solution is locally optimal for the original Dubins-car problem.This condition links successful saturation identification to local optimality.
- Results: After 7 iterations, the LMPC converges to a steady-state solution while the iteration cost decreases until convergence.The sampled safe set evolution is shown in Figure 4 and the cost in Table IV.
- Results: The learned controller saturates acceleration, accelerates to the midpoint, then decelerates and reaches xF in 16 steps.This matches the optimal solution obtained in the known-saturation example.
V. PRACTICAL CONSIDERATIONS
The sampled safe set makes the terminal constraint integer-valued, increasing computational expense. A convex-hull relaxation can reduce this burden and preserve key guarantees for linear systems with convex stage costs.
- V. PRACTICAL CONSIDERATIONS: The sampled safe set is discrete, so the terminal constraint becomes an integer constraint.This requires solving a mixed integer programming problem at each time step, including for linear systems.
- V. PRACTICAL CONSIDERATIONS: Relaxing the sampled safe set to its convex hull and approximating Q(·) barycentrically can reduce computational burden.For linear dynamics and convex stage costs, the relaxed problem is convex.
- V. PRACTICAL CONSIDERATIONS: The relaxed approach preserves the properties established in Theorems 1–3 for linear systems with convex stage costs.
2) Parallelize Computations:
The LMPC structure supports complexity reduction by bounding the optimal solution and restricting computation to relevant safe-set points. The framework also states deterministic guarantees and describes extensions for disturbances.
- 2) Parallelize Computations:: A subset of the sampled safe set can be used in the LMPC optimization, and the computation can be parallelized.Upper and lower bounds on the optimal solution reduce problem complexity without losing the guarantees of Theorems 1–3.
- 2) Parallelize Computations:: The algorithm applies the current control action after solving the LMPC with a nonlinear optimization solver and selecting the minimizing candidate.The Dubins Car example uses Ipopt.
- Uncertainty: The theoretical guarantees are established only for the deterministic case, and disturbances cause those guarantees to be lost under the presented deterministic framework.The paper identifies stochastic and robust iterative learning MPC as extensions for future investigation.
- VI. CONCLUSIONS: The proposed controller uses a learned safe set and terminal cost to guarantee recursive feasibility and stability while improving performance across iterations.
- VI. CONCLUSIONS: The constrained infinite-horizon linear-quadratic regulator test converges to the optimal solution, while the Dubins car test estimates parameters and generates performance-improving trajectories.
VIII. APPENDIX
The appendix constructs a feasible trajectory by parameterizing candidate inputs, selecting a low-cost trajectory, refining it through nonlinear optimization, and extending it into the terminal set.
- VIII. APPENDIX: The initial guess fixes the saturation estimate at ˆs0 = 0.25 and uses a structured input parameterization.
- VIII. APPENDIX: The parameterized inputs use positive and negative ˜θ segments, bounded ˜a segments, zero acceleration intervals, and a terminal negative ˜a segment.
- VIII. APPENDIX: Trajectories generated from different parameter sets are ranked by a quantity, and the minimizing trajectory warm-starts a nonlinear optimization problem.
- VIII. APPENDIX: The optimized input sequence is applied to the system to compute realized trajectories and their error.
- VIII. APPENDIX: The resulting N-step trajectory steers the system into XF and can be used to construct the initial sampled safe set SS0.