Source-linked AI summary

Safety Beyond the Training Data: Robust Out-of-Distribution MPC via Conformalized System Level Synthesis

Anutam Srinivasan, Antoine Leeman, Glen Chou

arXiv:2602.12047v1cs.ROcs.LGeess.SYmath.OC

TL;DR

The paper addresses safe planning with learned dynamics under distribution shift, where model errors can make OOD behavior unsafe. It combines weighted conformal prediction with SLS-based robust MPC, and reports robust performance on car and quadcopter systems, including when learned dynamics generalize poorly OOD.

  • Problem

    Learned dynamics may be inaccurate beyond the training distribution, creating a need for high-probability error bounds and robust constraint satisfaction for the true system.

  • Method

    CP-SLS-MPC uses weighted conformal prediction with state-control-dependent error bounds inside SLS-based robust MPC, updating bounds online and guiding plans toward lower-error regions.

  • Results

    The framework improves safety and prediction accuracy over uncertainty-unaware MPC baselines on nonlinear 4D car and 12D quadcopter systems, including OOD scenarios.

  • Takeaways & Limitations

    CP-SLS-MPC achieves robust performance with learned dynamics that generalize poorly OOD while providing high-probability constraint guarantees.

  • Takeaways & Limitations

    The coverage-gap analysis assumes calibration scores are independent of the nominal score and that distribution drift changes gradually across state-control space.

Abstract

from arXiv · show

We present a novel framework for robust out-of-distribution planning and control using conformal prediction (CP) and system level synthesis (SLS), addressing the challenge of ensuring safety and robustness when using learned dynamics models beyond the training data distribution. We first derive high-confidence model error bounds using weighted CP with a learned, state-control-dependent covariance model. These bounds are integrated into an SLS-based robust nonlinear model predictive control (MPC) formulation, which performs constraint tightening over the prediction horizon via volume-optimized forward reachable sets. We provide theoretical guarantees on coverage and robustness under distributional drift, and analyze the impact of data density and trajectory tube size on prediction coverage. Empirically, we demonstrate our method on nonlinear systems of increasing complexity, including a 4D car and a {12D} quadcopter, improving safety and robustness compared to fixed-bound and non-robust baselines, especially outside of the data distribution.

1. Introduction

The paper targets unsafe behavior from learned dynamics outside the training distribution by combining calibrated error bounds with robust feedback planning. CP-SLS-MPC uses these bounds to provide probabilistic safety while improving planning efficiency and OOD performance.

  • Motivation: Learned dynamics can exploit model inaccuracies, motivating uncertainty-constrained planners that bound prediction error for safe execution.These bounds define trusted domains and reachable tubes around the training data.
  • Approach: CP-SLS-MPC jointly plans nominal trajectories and tracking controllers using robust MPC, weighted conformal prediction, and system level synthesis.The method updates error bounds with online data and uses model-error gradients to guide plans toward lower-error regions.
  • Contributions: The method provides finite-sample probabilistic safety through state-control-dependent conformal error bounds embedded in MPC.
  • Results: Numerical validation on a simulated nonlinear 4D car and 12D quadcopter improves safety and prediction accuracy over uncertainty-unaware MPC baselines, including OOD scenarios.

2. Related Work

Prior learning-based MPC methods estimate disturbances but typically lack state-dependent error bounds and cannot direct plans toward low-error regions. The paper combines CP-calibrated uncertainty with SLS to support safer planning toward such regions.

  • Learning-based MPC: Learning-based MPC unifies disturbance estimation with planning but typically lacks state-dependent error bounds and cannot steer plans toward low-error regions.
  • Robust MPC: Robust and tube MPC handle disturbance constraints, but direct closed-loop enforcement is often nonconvex and produces conservative overapproximations.
  • System Level Synthesis: SLS provides a convex closed-loop response parameterization for LTV systems and has been extended to data-driven and nonlinear dynamics.
  • This work: This work informs SLS with state-dependent, CP-calibrated error bounds to enable safe planning toward low-error regions.

3. Preliminaries and Problem Statement

The paper formulates learned nonlinear dynamics, weighted conformal calibration, and SLS-based reachable-tube control as a robust MPC problem. Its objective is high-probability constraint satisfaction for the true system across prediction and replanning horizons.

  • Dynamics model: The true dynamics are modeled as unknown nonlinear dynamics, while a learned model approximates them with an explicit dynamics error.
  • Problem formulation: The state and control spaces impose pointwise constraints and a target-set reaching objective for the learned-dynamics planning problem.
  • Conformal Prediction: Weighted conformal prediction converts calibration nonconformity scores into a calibrated error threshold and prediction set with 1−α coverage.
  • Conformal Prediction: Unlike uniform-weight CP, WCP addresses spatio-temporal distribution shifts, state-control correlations, and visits to OOD regions.
  • System Level Synthesis: SLS jointly optimizes nominal trajectories, tracking controllers, and reachable-tube overapproximations by propagating worst-case disturbances through closed-loop responses.
  • Nonlinear robust MPC: For nonlinear learned dynamics, Jacobian linearizations and tightened constraints are used to ensure the nonlinear constraints hold across the prediction horizon.
  • Problem formulation: The problem assumes offline data split into training and calibration sets, then seeks calibrated error bounds and robust constraint satisfaction for the true dynamics.

4. Method

CP-SLS-MPC combines weighted conformal prediction with SLS-based robust MPC to bound state-control-dependent learned-model error and constrain the true closed-loop system with high probability. It updates these bounds online and optimizes nominal trajectories, tracking controllers, and reachable tubes for robust operation.

  • Model training and conformal error bounds: The method uses weighted conformal prediction to produce state-control-dependent uncertainty sets from learned dynamics-model errors.A learned covariance model supplies local error shape and magnitude, while calibration converts it into a conformal threshold and ellipsoidal prediction set.
  • Model training and conformal error bounds: The conformalized error set is an origin-centered ellipsoid that contains the one-step model error with high probability, with coverage adjusted for distributional shift and reachable-tube variation.For points within a reachable tube, the coverage bound includes total-variation terms and an additional tube-dependent gap term.
  • SLS-based robust MPC: The MPC objective minimizes nominal cost, terminal cost, reachable-tube volume, and distance to calibration data, shrinking tubes in dense-data regions and expanding them under OOD conditions.The active error-reduction cost encourages trajectories toward calibration points where the learned dynamics are more accurate.
  • SLS-based robust MPC: CP-SLS-MPC integrates these error bounds into an SLS-based nonlinear MPC problem that jointly optimizes nominal trajectories, disturbance-feedback responses, and closed-loop reachable tubes.The optimization enforces nominal dynamics, SLS conditions, and tightened state-control constraints using the uncertainty representation.
  • Safety guarantees: Theoretical guarantees characterize safety over the prediction horizon and across repeated MPC replanning solves using timestep-specific and replanning-dependent miscoverage rates.The guarantees concern the realized closed-loop trajectory under the controller and account for repeated replanning.
  • Online implementation: The controller is solved online with sequential convex programming and second-order cone programs, executes the first control input, and incorporates the resulting transition into the calibration set.A real-time iteration scheme limits computation to one sequential-convexification iteration per control step.

5. Theoretical Analysis: Bounding the Coverage Gap

The analysis bounds the coverage gap caused by distributional drift between calibration data and nominal state-control points, including uniformly over each reachable tube. The bound exposes trade-offs involving calibration-point density, proximity to training data, and tube size.

  • Assumptions: The coverage-gap analysis assumes calibration scores are independent of the nominal score while allowing them to come from different distributions.It additionally assumes distribution drift between score distributions is bounded in a Lipschitz-type manner over the state-control space.
  • Coverage-gap bound: Theorem 3 expresses the coverage gap as a weighted sum of total-variation distances between calibration-score distributions and the nominal score distribution.The theorem presents this quantity as an interpretable bound tied to the calibration and nominal distributions.
  • Density and proximity effects: Greater calibration-point density decreases d_min and increases the coverage gap, while reducing ρ can mitigate the gap but may make the (1 −α)-quantile diverge.The bound also indicates that proximity to training data reduces the gap through d1, the minimum distance between calibration data and the nominal state-control point.
  • Tube-wide guarantee: Theorem 4 extends the guarantee from the nominal point to every state-control pair in the reachable tube, with an additional tube-dependent coverage penalty.This provides the form of guarantee needed when SLS requires the model-error bound to hold throughout the tube rather than only at its center.
  • Implications: Coverage remains close to the desired level under gradual spatial drift, and minimizing reachable-tube volume reduces the tube penalty γ(Rk).The penalty is defined using the tube’s maximum axis length, linking geometric tube size to the coverage guarantee.

6. Results

Experiments evaluate CP-SLS-MPC on learned 4D car and 12D quadcopter systems across in-distribution, OOD, friction, active-uncertainty, and obstacle-avoidance settings. The method generally improves robustness over vanilla MPC and fixed spherical bounds, with state-dependent tubes adapting to uncertainty.

  • Experimental setup: CP-SLS-MPC is evaluated against vanilla MPC and CP-Ball on learned 4D Dubins’ car and 12D quadcopter dynamics.The evaluation reports computation time, prediction error, and minimum obstacle distance over horizon T = 15.
  • Calibration sensitivity: Smaller calibration sets keep errors inside the ellipsoid more often but produce larger one-step tubes.Figure 1 compares prediction-error distance to the ellipsoid boundary and corresponding one-step θ tubes for varying Dcalib.
  • Friction car: Across 10 disjoint-domain friction-car runs, V-MPC failed twice, CP-Ball once, and CP-Ellipsoid was consistently successful.CP-Ellipsoid also maintained greater obstacle distance, while its state-dependent ellipsoids expanded along the direction of motion.
  • Active uncertainty reduction: Active uncertainty reduction shifts summed tube logvolumes left, indicating less runtime uncertainty and enabling longer paths that remain in the training domain.The method achieved lower prediction error and smaller tubes than the baseline without the uncertainty-reduction cost.
  • 12D quadcopter: 8 out of 10 quadcopter trajectories succeeded with CP-Ellipsoid versus 4 out of 10 with V-MPC.One CP-Ellipsoid failure was attributed to numerical instability without collision, while the authors note that faster tube expansion may be needed for consistent avoidance.

7. Conclusion

The paper develops an efficient WCP- and SLS-based framework for robust OOD MPC with learned dynamics. Its guarantees and experiments support robust performance across 4D car and 12D quadcopter variants despite poor OOD generalization.

  • Conclusion: CP-SLS-MPC combines weighted conformal prediction and system level synthesis for robust OOD MPC with learned dynamics.The framework provides high-probability rollout constraint guarantees and includes an uncertainty-reduction cost to favor the training domain when possible.
  • Conclusion: Experiments on 4D Dubins car and 12D quadcopter variants demonstrate robust performance even when learned dynamics generalize poorly OOD.The conclusion attributes this result to the combined WCP and SLS framework.

A.1. Proof of Theorem 3 and Corollary 5

The appendix develops coverage-gap bounds for weighted conformal prediction under independent data and distributional changes, then connects these bounds to robust MPC and system-level synthesis. It also records experimental dynamics, baselines, and tube-volume objectives.

  • Proof of Theorem 3 and Corollary 5: Theorem 3 bounds the weighted conformal-prediction coverage gap using weighted total-variation distances between calibration and test-score distributions.The proof rearranges distance-dependent weights into a decaying series bounded by geometric series.
  • Proof of Theorem 3 and Corollary 5: As Ncalib →∞, Corollary 5 gives a tighter coverage-gap bound under the assumptions of Theorem 3.The appendix states this asymptotic tightening explicitly.
  • Proof of Theorem 4: Theorem 4 accounts for uncertainty in the test non-conformity score by comparing distributions through total variation and applying a Lipschitz-type tube bound.The bound depends on the tube’s maximum axis length and a local Lipschitz constant.
  • Safety guarantees: The proof sketches transfer per-step coverage results into rollout safety and use intersection probabilities across forecasted MPC steps instead of only a union bound.SLS guarantees are conditioned on valid disturbance bounds.
  • Experimental models: The experiments use continuous-time Dubins-car and quadcopter dynamics discretized with Euler integration, including modified OOD and friction dynamics.The quadcopter model is 12-dimensional, while the car state includes position, direction, and velocity.
  • Baselines and objectives: Vanilla MPC uses learned dynamics without constraint tightening, whereas CP-Ball uses a constant spherical error bound and SLS formulations minimize tube volume as a surrogate for uncertainty.The CP-Ball representation is described as more conservative than the ellipsoidal bound.

D.1. Data Sampling

The experiments construct training, uncertainty, and calibration data over specified car and quadcopter state-control regions, including deliberately excluded OOD areas. Neural dynamics and uncertainty models then support SLS propagation, tube tightening, and active uncertainty reduction.

  • Car data sampling: Car training samples cover [px, py, θ, v] × [ω, a] with px ∈ [0, 5], py ∈ [−5, 5], and bounded angular, velocity, and control ranges.The OOD and friction experiments instead sample separated py intervals outside the central training region.
  • Car data sampling: The car datasets contain 1,000,000 dynamics points, 1,000,000 uncertainty points, and calibration sets with varying sizes.Ground truth is computed from the expert dynamics model at sampled state-control points.
  • Quadcopter data sampling: Quadcopter trajectories exclude a radius-1 OOD region around (2.5, 2.5, 2.5), use 0.05-second timesteps, and produce 4,702,400 training points.The dataset comes from 100,000 expert trajectories of length 100 after retaining optimizer-solvable trajectories.
  • Learned models: Dynamics and uncertainty models are one-hidden-layer tanh MLPs, with uncertainty outputs parameterized by Cholesky factors to ensure positive-definite matrices.Car models use 4,096 and 2,048 hidden nodes, while quadcopter models use 1,024 hidden nodes for both networks.
  • Optimization settings: SLS forward propagation uses a specified positive-definite surrogate matrix, while trajectory optimization uses an LQR-style cost with model-specific weighting matrices.The car and quadcopter costs use diagonal or block-diagonal matrices as appropriate.
  • Calibration settings: Conformal prediction uses ρ = 0.97 for cars, ρ = 0.8 for the quadcopter, and αk = 0.1 across models.The standard calibration sets contain 10,000 car points and 30,000 quadcopter points outside the OOD car experiment.
  • Active uncertainty sampling: Active uncertainty reduction uses K-means to select representative calibration points and penalizes distance from those points in positional state space.The objective also includes a goal weight and a sharpness parameter controlling concentration near selected points.

Appendix E. Theorem 3 Analysis

The toy analysis studies how calibration density and the weighting parameter ρ affect coverage bounds under spatially varying nonconformity scores. Smaller ρ and larger calibration sets improve the lower coverage bound, but smaller ρ can increase the calibrated quantile and reduce utility.

  • Toy Example: Smaller ρ values emphasize calibration points closer to the test point and increase the test point’s weight in the empirical score distribution.The toy example uses uniformly sampled points on a unit circle and averages results across 100 trials.
  • Coverage Bounds: The derived lower coverage bound and empirical coverage increase as ρ decreases, with bounds improving more rapidly at smaller ρ.The comparison includes the bound from (24), Barber et al. (2023), and the derived CP-SLS coverage bound in (23).
  • Coverage Bounds: For fixed ρ, increasing Ncalib raises the lower coverage bound.The result follows from the toy experiments across several calibration-set sizes.
  • Trade-off: Smaller ρ values improve coverage but can produce larger, potentially infinite, 1−α quantiles, creating a utility sacrifice.The analysis therefore frames ρ as a trade-off between coverage and the size of the calibrated uncertainty bound.

Appendix F. Additional Results

Additional experiments show that adaptive ellipsoid methods maintain prediction-error coverage and improve obstacle avoidance and trajectory behavior in car and quadcopter rollouts, including near OOD regions.

  • Car Results: Over 99.33% of in-domain car prediction errors remain inside both the fixed-ball and adaptive-ellipsoid coverage bounds.The figure reports empirical satisfaction of the coverage guarantee for both methods.
  • Car Results: The adaptive ellipsoid approach avoids the obstacle in the friction-car example, whereas the remaining approaches crash.The other methods fail to maintain sufficient obstacle proximity.
  • Computational Comparison: Table 2 compares computation times and prediction errors for vanilla-MPC, adaptive ellipsoid without active uncertainty, and adaptive ellipsoid with active uncertainty.The passage identifies the compared methods and reported quantities but does not provide their values.
  • Active Uncertainty: Active uncertainty produces smoother rollouts by mostly avoiding the OOD region or staying near the in-distribution region.The near-distribution trajectory encounters limited OOD errors despite taking more OOD steps than the comparison shown in Figure 9.
  • Quadcopter Results: CP-Ellipsoid successfully avoids the obstacle in an additional quadcopter rollout when entering the OOD region, unlike V-MPC.The rollout uses a different initial condition from the earlier quadcopter example.
Loading 2602.12047v1…