Source-linked AI summary

Fixed-Time Integral Reinforcement Learning for Saturated Nonlinear Multi-Agent Systems Under FDI Attacks

Tien Dat Vu, Minh Doan

arXiv:2609.06163v1eess.SY

TL;DR

The paper tackles secure leader–follower formation for unknown nonlinear multi-agent systems subject to disturbances, actuator FDI attacks, bounded inputs, and fixed-time requirements. It develops an integral reinforcement-learning zero-sum game with saturation-compatible cost shaping and critic learning, and proves practical fixed-time convergence of learning and formation errors to bounded residual sets under the stated conditions.

  • Problem

    Secure fixed-time formation learning must jointly handle unknown nonlinear dynamics, disturbances, graph-coupled FDI attacks, bounded commands, and limitations of simpler existing models.

  • Method

    The paper combines a local HJI-based integral reinforcement-learning framework, nonquadratic input utility, two-power state-cost shaping, and a replay-based critic update.

  • Results

    The analysis proves practical fixed-time convergence of critic weight estimation and leader-referenced formation errors to compact residual sets, while simulations show resilient tracking under persistent disturbances and time-varying FDI attacks.

  • Takeaways & Limitations

    The framework provides bounded secure commands and fixed-time practical formation behavior without explicit knowledge of the system drift.

  • Takeaways & Limitations

    Cost-shaping coefficients are selected by empirical tuning, and persistent perturbations lead to convergence to prescribed neighborhoods rather than exact convergence to the origin.

Abstract

from arXiv · show

The leader-follower formation control problem is investigated for nonlinear multi-agent systems with unknown dynamics, external disturbances, and false data injection (FDI) attacks on actuator channels. The problem is formulated as a zero-sum differential game and solved using the Integral Bellman-Isaacs approach. To address input saturation constraints, a non-quadratic control cost function is incorporated into the optimization problem, leading to a bounded control law. Furthermore, this paper proposes a cost function construction method and develops a critic learning law, which together guarantee the practical fixed-time stability of the system while overcoming the limitations of existing fixed-time reinforcement learning formulations. Finally, the practical fixed-time convergence of both the critic weight estimation error and the leader-referenced formation tracking error to bounded residual sets is rigorously proven. Simulation results demonstrate the effectiveness of the proposed method under external disturbances, FDI attacks, and input constraints.

I. INTRODUCTION

The paper addresses secure leader–follower formation for unknown nonlinear multi-agent systems exposed to disturbances, FDI attacks, bounded inputs, and fixed-time requirements. It combines a zero-sum game formulation with integral reinforcement learning, cost shaping, and critic adaptation to establish resilient practical fixed-time behavior.

  • Motivation: FDI signals can corrupt measurements, exchanged information, or control commands, motivating zero-sum differential-game modeling of secure formation control.The associated optimality condition is characterized by the HJI equation.
  • Motivation: Unknown nonlinear drift makes direct HJI solution intractable, while integral reinforcement learning removes the drift from finite-window learning residuals.Nonquadratic input utilities embed symmetric command constraints and yield bounded hyperbolic-tangent policies.
  • Research gap: Existing fixed-time secure reinforcement learning is limited to comparatively simple second-order models, leaving unknown nonlinear dynamics and coupled secure-control channels unresolved.The paper identifies the joint treatment of unknown drift, graph coupling, same-channel FDI, bounded commands, and fixed-time critic learning as challenging.
  • Approach: The proposed framework decomposes graph-coupled errors into drift, secure-control, and adversarial channels, then uses a two-power state penalty and integral Bellman–Isaacs residual.Recorded data preserve excitation as trajectories approach the desired formation.
  • Contributions: The paper proposes a cost construction, critic update law, unified stability analysis, and resilient controller for unknown nonlinear systems under FDI attacks.These contributions target fixed-time parameter-error convergence and stable, robust closed-loop behavior.

III. SECURE SATURATION-AWARE HJI FORMULATION

The secure formation problem is formulated as a local zero-sum graphical game with bounded defender commands and adversarial disturbance and FDI inputs. Saturation compatibility is embedded directly in the HJI formulation rather than imposed by an external saturation map.

  • Secure game formulation: The defender minimizes the cost, while disturbances and same-channel FDI signals act as maximizing inputs in the local zero-sum game.The original game cost remains independent of the unknown value function.
  • Saturation-aware HJI: Symmetric input constraints are embedded directly in the HJI equation through a saturation-compatible utility.This design produces admissible bounded secure policies within the prescribed input set.

A. Graphical Game and Fixed-Time Shaped HJI

The paper constructs a local graphical-game HJI formulation whose state penalty generates fixed-time comparison terms without assuming them directly for the unknown value function. The resulting analysis supports practical fixed-time convergence, while inverse cost recovery and admissibility preservation remain constrained by additional verification requirements.

  • Graphical game: The local graphical game separates self and neighboring secure inputs from locally and remotely entering adversarial channels.Neighboring secure policies are treated as fixed coupling signals when each follower’s local HJI equation is formed.
  • Fixed-time condition: A fixed-time stability theorem converts a two-power dissipation inequality with residual term Δ_i into practical fixed-time convergence.The settling-time bound is independent of the initial value, and Δ_i = 0 yields exact fixed-time convergence.
  • Cost-induced shaping: The running cost is designed so mixed powers of the measurable coordination error generate the required fixed-time structure directly through the HJI equation.This avoids placing the unknown value function in the running cost and makes the construction cost-induced rather than externally assumed.
  • Inverse-optimal interpretation: Inverse recovery offers a systematic alternative for constructing a fixed-time-compatible penalty, but requires matching, positivity, saturated-policy admissibility, and neighboring-policy consistency checks.The paper retains a simpler sufficient condition because inverse-cost synthesis is not its main contribution.
  • HJI formulation: The secure utility penalizes defender effort, while negative quadratic disturbance terms represent the maximizing roles of disturbances and same-channel FDI attacks.The resulting HJI equation is nonlinear and generally analytically intractable, motivating learning-based approximation.

B. Integral Bellman–Isaacs Residual

The paper replaces the unknown-drift pointwise HJI equation with a finite-window integral Bellman–Isaacs identity and learns the value function using a critic representation. Recorded data and normalized residuals support learning without explicit drift knowledge or persistent excitation.

  • The unknown value function and gradient are approximated on a compact trajectory-relevant region using an RBF critic with bounded approximation residuals.
  • A finite data window converts the HJI relation into a computable identity based on feature differences and an approximate integral running term.
  • When the critic weights, approximation, and policies are ideal, the residual is zero, so residual minimization recovers the saddle-point condition along the trajectory.
  • The integral Bellman–Isaacs residual enforces the HJI equation along measured trajectories without requiring explicit knowledge of the unknown drift.
  • A finite experience stack stores feature differences and running terms, replacing persistent excitation with a requirement that recorded data contain enough independent information.

C. Fixed-Time Critic Update

The critic is updated from normalized online and replay Bellman–Isaacs residuals using a two-power fixed-time learning law. This makes critic convergence relevant to closed-loop formation stability while allowing stored data to preserve excitation after trajectories become uninformative.

  • The residual formulation avoids pointwise model-based HJI errors because the unknown drift appears in the unavailable pointwise equation.
  • The update is derived from residual gradients whose current and replay regressors are feature differences, enabling direct learning from finite-window and recorded data.
  • The fixed-time critic update combines current and recorded residuals with leakage, using two powers to accelerate learning near zero and avoid slow learning for large residuals.
  • Because the learned control policy depends on the critic gradient, bounding critic-weight error in fixed time prevents an arbitrarily long learning transient from extending the closed-loop convergence bound.
  • Experience replay makes initial excitation sufficient for critic learning, even when future formation trajectories no longer remain persistently exciting.

IV. STABILITY ANALYSIS

The stability analysis proves practical fixed-time convergence for both the critic-weight error and the leader-referenced formation error under bounded approximation and learning residuals. The resulting time bounds are independent of the initial critic error and formation-state condition.

  • The critic-weight estimation error is practically fixed-time stable, and its convergence-time bound is independent of the initial critic-weight error.
  • Bounded Bellman–Isaacs approximation, integration, and leakage terms form a residual constant that does not depend on the initial critic-weight error.
  • The critic Lyapunov estimate uses two powers of the weight error plus a bounded perturbation term to establish practical fixed-time convergence.
  • The formation proof combines critic convergence with the cost-induced HJI decay to obtain practical boundedness of the leader-referenced error.
  • The leader-rooted formation error and critic-weight error reach bounded residual sets within fixed times, with the formation bound given by the maximum follower convergence time.

A. Simulation Setup

The simulation uses one leader and four nonlinear followers in a two-dimensional formation with unknown nonlinear drift, disturbances, actuator FDI, and bounded control channels. Replay data are collected before attack onset, and scaled initial errors test fixed-time behavior.

  • The simulation models four nonlinear followers with unknown drift, external disturbances, secure control inputs, and actuator-side FDI signals relative to one leader.
  • The actuator-side FDI attack begins at t = 8 s, while the leader and desired follower offsets define the leader-referenced formation task.
  • Each follower uses a six-node Gaussian RBF critic with normalized coordination-error inputs over the admissible state domain.
  • Replay data are collected during the pre-attack interval t ∈[0, 8] s and retained afterward to preserve initial-excitation information for critic learning.
  • The numerical fixed-time comparison scales initial formation-error perturbations by s ∈ {0.5, 1.0, 1.5, 2.0, 2.5}.

B. Simulation Results

The simulations show bounded leader–follower tracking and critic learning under increasing initial errors, input constraints, actuator-side FDI attacks, and persistent disturbances.

  • The observed settling times remain between 9.53 and 10.76 s as the initial-error scale increases from 0.5 to 2.5.The corresponding initial norm grows from approximately 30 to above 150, yielding a common numerical bound.
  • The six critic weights for each follower reach bounded constant values in approximately 5.7 s without subsequent drift.
  • Formation position errors are driven close to the origin while velocity errors remain small after the transient.The initial position errors span approximately [-5, 6] along x and [-4, 5] along y.
  • Secure inputs remain within |u_ci,ℓ| < 40, with peak magnitudes of approximately 39 and 32 along the x- and y-channels.
  • Bounded tracking errors are maintained under nonlinear dynamics, persistent disturbances, actuator FDI attacks, and input constraints.FDI signals reach approximately 0.67 after activation at t_a = 8 s, while disturbances remain within approximately [-0.063, 0.063].

VI. DISCUSSION AND LIMITATIONS

The discussion identifies nonconstructive stability certificates, empirically tuned cost coefficients, practical rather than exact convergence, and unresolved admissibility-preservation issues.

  • The comparison constants in the stability certificate exist on the relevant compact region but cannot yet be systematically computed from available system information.Their relationship with the cost-shaping coefficients is therefore not fully constructive.
  • Bounded approximation and adversarial residuals imply practical fixed-time convergence to a prescribed neighborhood rather than exact convergence to the origin under persistent perturbations.
  • The cost-shaping coefficients are conservatively selected by empirical tuning and affect both the Bellman–Isaacs residual and critic learning.
  • The admissibility certificate is local and does not establish that every subsequent critic-only policy update preserves admissibility.
  • Data-driven admissibility preservation during online learning and constructive evaluation of the HJI comparison constants remain future-work directions.

VII. CONCLUSION

The paper presents a fixed-time integral reinforcement learning framework for nonlinear leader–follower formation under unknown dynamics, disturbances, FDI attacks, and bounded inputs. Simulations confirm resilient formation tracking under persistent disturbances and time-varying FDI attacks.

  • The framework handles unknown dynamics, disturbances, FDI attacks, and bounded inputs in nonlinear leader–follower formation control.
  • A nonquadratic utility enforces input constraints, while a two-power cost and replay-based critic update drive critic and formation errors to compact residual sets in fixed time.The settling times are independent of initial conditions.
  • Simulations confirmed resilient formation tracking under persistent disturbances and time-varying FDI attacks.

APPENDIX A DATA-REUSED CONSTRUCTION OF AN INITIAL

The appendix constructs an admissible initial policy by reusing stored trajectory data to identify a lifted nominal model and certify local stability, bounded input, and finite cost. This initialization supports policy iteration while clarifying that online critic learning remains data-driven rather than an exact Koopman policy-iteration procedure.

  • The construction certifies and initializes an admissible HJI branch without requiring an exact finite-dimensional Koopman model.Online critic-only integral Bellman–Isaacs learning continues directly from finite-window residuals.
  • The lifted coordinate η_i = L_Ki(χ_i) preserves the local coordination-error origin because η_i = 0 if and only if χ_i = 0.
  • Stored trajectory windows identify a lifted nominal model through informative data satisfying rank(R_Ki) = N_Ki + m.The identified dynamics include a finite-data fitting residual and lifting or identification error term.
  • A validation bound with α_Ki := λ_min(Q_Ki) − 2∥P_i∥ρ_Ki > 0 yields a locally stabilizing initial policy and an invariant admissible region.The resulting Lyapunov derivative is negative away from the origin.
  • The initial policy produces exponential decay of the lifted state and converges to zero while satisfying the saturation-compatible input constraint.
Loading 2609.06163v1…