Source-linked AI summary

Toward Secure Communications for a UAV Swarm with Movable Antennas in SAGIN: CKM-Enabled Multi-Agent Reinforcement Learning Framework

Jiayang Wan, Yafei Wang, Jiawei Zhuang, Wenjin Wang, Tony Q. S. Quek

arXiv:2608.27973v1eess.SP

TL;DR

Secure SAGIN communication for UAV swarms requires joint optimization of movable antennas, trajectories, and heterogeneous link selections under CSI, action-space, and system-constraint challenges. The paper proposes a CKM-assisted MAPPO framework with rigid-body MA shifts, collaborative rewards, and action masks; it improves SEE while reducing CSI-acquisition and online-inference overhead relative to considered baselines.

  • Problem

    Joint secure optimization of MA positions, UAV trajectories, and link selections in SAGIN remains largely unexplored because timely CSI acquisition and coupled constrained decisions create substantial complexity.

  • Method

    The paper combines sparse-measurement Kriging CKMs and satellite ephemerides with MAPPO, rigid-body MA translation, hybrid rewards, and action masks.

  • Results

    The proposed framework achieves higher SEE with lower online inference and CSI-acquisition overhead than the considered baselines.

  • Takeaways & Limitations

    Joint MA positioning, trajectory optimization, and link selection can be coordinated for secure UAV-swarm communication while reducing online CSI and inference overhead.

Abstract

from arXiv · show

Space-air-ground integrated networks (SAGINs) can provide ubiquitous and reliable connectivity for unmanned aerial vehicles (UAVs). However, air-to-ground links, which are typically dominated by line-of-sight (LoS) propagation, are vulnerable to passive eavesdropping due to the broadcast nature of wireless channels. To enhance physical-layer security, we investigate a SAGIN-enabled secure downlink communication system in which UAVs select service links among satellite, aerial, and terrestrial networks while adjusting the positions of the movable antenna (MA) array to fully exploit connectivity and spatial degrees of freedom for improved secrecy communication performance. Specifically, we maximize the secrecy energy efficiency (SEE) of a UAV swarm by jointly optimizing the MA positions, UAV trajectories, and link selections, subject to UAV mobility, MA movement, and link connectivity constraints. To reduce the real-time channel state information (CSI) acquisition overhead, we propose a channel knowledge map (CKM)-assisted multi-agent reinforcement learning framework. Specifically, the CKM is first constructed from sparse channel measurements via Kriging interpolation and is then leveraged together with satellite ephemeris information to enable efficient storage and retrieval of CSI. To reduce the action-space dimensionality and computational complexity, we model the MA array using rigid-body kinematics and adjust its position through global rigid-body translation, thereby constructing a low-dimensional hybrid action space for the joint optimization decisions. To align local decisions with system-wide performance under system constraints, we design an individual-team collaborative reward mechanism and introduce action masks to enforce constraints on UAV mobility, collision avoidance, MA regions, and connectivity capacity.

I. INTRODUCTION

SAGIN addresses coverage gaps for highly mobile UAVs, while secure communication requires jointly exploiting heterogeneous links, trajectories, and movable-antenna spatial degrees of freedom. The paper proposes CKM-assisted MARL with reduced-dimensional actions and constraint-aware coordination.

  • SAGIN integrates terrestrial, low-altitude, and satellite networks to provide pervasive connectivity for low-altitude airspace.
  • Highly mobile UAVs experience rapidly varying link conditions, while fixed-position arrays cannot fully exploit local spatial channel variations.Movable antennas reposition elements within confined regions to obtain more favorable channel conditions.
  • The paper jointly optimizes MA positions, UAV trajectories, and satellite, aerial, or terrestrial link selections to maximize swarm secrecy energy efficiency under system constraints.
  • Sparse offline channel measurements are interpolated into a CKM and combined with satellite ephemerides to reduce online CSI-acquisition overhead.The framework enables UAVs to query CSI for their locations and candidate MA configurations during decision-making.
  • Rigid-body MA modeling, action masking, and individual-team rewards reduce action-space complexity while coordinating local decisions with swarm performance.The unified hybrid action space jointly generates link, trajectory, and MA decisions while excluding infeasible actions.

B. Channel Model

The A2G channel model combines geometry-dependent LoS probability and path loss with a movable receive array whose channel response depends explicitly on antenna positions. Optimizing those positions can improve channel conditions.

  • 1) MA-Enabled A2G Links:: Each UAV uses a rectangular-panel M-element MA array for A2G links, with element positions constrained to a two-dimensional movable region.
  • 1) MA-Enabled A2G Links:: The model computes link distance, horizontal distance, and elevation angle from GBS or ABS and UAV coordinates.
  • 1) MA-Enabled A2G Links:: LoS probability is modeled as an environment-dependent function of elevation angle, with average path loss combining free-space loss and additional LoS/NLoS losses.
  • 1) MA-Enabled A2G Links:: The small-scale A2G MIMO channel uses a geometric model with L effective paths and transmit and receive direction parameters.
  • 1) MA-Enabled A2G Links:: MA positions continuously adjust steering-vector phase terms, providing additional spatial degrees of freedom and enabling channel-condition optimization.

2) A2S Links:

The system models A2S communication through far-field single-path LoS channels and discretizes UAV operation into time slots. It also imposes fixed-altitude, trajectory, collision-avoidance, link-selection, capacity, and signal-model constraints.

  • 2) A2S Links:: A2S links use a far-field single-path LoS channel between each LEO satellite and UAV, with distance determined by their three-dimensional positions.
  • 2) A2S Links:: The satellite uses a K_x × K_y UPA, while each UAV uses a K-element upper-side ULA for A2S reception.
  • 2) A2S Links:: Mission time is divided into N equal slots, producing N + 1 discrete time instants for modeling time-varying UAV mobility.
  • 2) A2S Links:: UAV trajectories are restricted to fixed altitude and prescribed initial and final positions, with a minimum inter-UAV safety distance preventing collisions.
  • 2) A2S Links:: Each UAV selects at most one serving base station per slot, while each base station serves no more than its capacity limit.

III. PROBLEM FORMULATION

The paper formulates SEE maximization for a UAV swarm by jointly optimizing UAV trajectories, MA positions, and heterogeneous-network link selections. The resulting problem combines non-convex mixed-integer decisions, CSI overhead, and strict operational constraints.

  • SEE is maximized by jointly optimizing UAV trajectories, MA positions, and satellite, aerial, and terrestrial link selections.
  • The optimization is a non-convex mixed-integer nonlinear programming problem whose global optimum is generally difficult to obtain directly.
  • Each UAV is assumed to fly at a prescribed constant horizontal speed, while propulsion power alone is retained in the energy model.MA actuation and communication energy are left for future work.
  • Instantaneous CSI required for secrecy-rate evaluation incurs substantial signaling overhead.
  • The framework addresses the formulation challenges through CKM-based CSI retrieval, rigid-body MA translation, and action masking for feasible decisions.These mechanisms target CSI acquisition, hybrid action-space complexity, and constraint enforcement.

IV. PROPOSED CKM-ASSISTED MARL FRAMEWORK

The proposed framework uses an offline-constructed CKM and satellite ephemerides to support CSI-aware training without online physical-environment interaction. It jointly handles link selection, UAV trajectories, and MA positioning through hybrid actions and feasibility masking.

  • The CKM-assisted MARL framework uses an offline-constructed CKM to support CSI-aware training.Satellite ephemerides characterize the training environment alongside the CKM.
  • Agents acquire CSI through data-driven interactions without requiring online interactions with the physical environment during training.
  • Link selection, trajectory planning, and MA positioning are formulated as a unified hybrid discrete action space.Action masking ensures feasibility under hard constraints.

A. CKM Construction for MARL Algorithm

The CKM represents channel knowledge as a function of UAV location, propagation environment, and MA configuration. Sparse joint position–configuration measurements are interpolated with Kriging and combined with satellite geometry for CSI retrieval.

  • The MA-enabled MIMO channel is mainly determined by UAV location, propagation environment, and MA configuration.
  • Accurately modeling the channel function is difficult because complex environments and radio-wave interactions are challenging to characterize mathematically.
  • CKM validity relies on piecewise quasi-static geometry and large-scale statistics, while mobile objects can still induce blockage, reflections, and channel transitions.
  • A Kriging-based CKM stores joint position–MA-configuration query keys with corresponding measured channel-feature samples.
  • The query key combines two-dimensional UAV location with the vectorized MA configuration, and ordinary Kriging predicts channel features at unsampled points.
  • Stored CSI can be retrieved from current UAV location and MA configuration, with satellite ephemerides extending retrieval across heterogeneous links, locations, and time instants.
  • The proposed algorithm uses the CKM-assisted MAPPO formulation to address coupled decisions in a cooperative multi-agent POMDP.

2) Observation Space:

The framework uses centralized training with decentralized execution and a low-dimensional hybrid action space for UAV heading, link selection, and rigid-body MA translation. Action masks restrict decisions to feasible trajectory, link, and MA choices.

  • Observation Space: During decentralized execution, each UAV makes decisions from its local observation, including information about neighboring UAVs within communication range.
  • Low-Dimensional Action-Masked Hybrid Action Space: A hybrid discrete action jointly represents UAV trajectory control, link selection, and MA configuration optimization.
  • Low-Dimensional Action-Masked Hybrid Action Space: Independent per-antenna selection creates N_MA^M joint configurations and exponential action-space growth with the number of antennas.
  • Low-Dimensional Action-Masked Hybrid Action Space: Rigid-body modeling selects one of N_sh candidate translations for the entire MA array instead of independently selecting every antenna position.
  • Low-Dimensional Action-Masked Hybrid Action Space: The action-search complexity decreases from O(494) for per-antenna selection to O(9) for rigid-body translation under the considered parameter setting.
  • Low-Dimensional Action-Masked Hybrid Action Space: Action masks are constructed for trajectory control, link selection, and MA configuration using the corresponding system constraints.
  • Algorithm: Algorithm 1 executes parallel agent decisions, CKM queries, action masking, environment transitions, and MAPPO updates.

4) Individual-Team Collaborative Reward Function:

The framework uses an individual-team collaborative reward to balance each UAV’s SEE with swarm-level cooperation, while masked hybrid policies and centralized training support feasible decentralized decisions.

  • Individual-Team Collaborative Reward Function: The individual-team reward combines each UAV’s local performance with a team reward to balance individual optimization and system-level cooperation.The SEE is the core reward component guiding UAV trajectory optimization during flight.
  • MAPPO Training and Execution: MAPPO is trained under centralized training with decentralized execution, using a global-state critic and locally observed actor decisions.Homogeneous UAVs share actor parameters, improving scalability and preserving permutation invariance.
  • Hybrid Action Policy: The hybrid policy factorizes decisions into conditionally independent trajectory, link-selection, and MA-position sub-actions.The joint log-probability and entropy are computed as sums over the three policy heads.
  • Constraint Enforcement: Action masking assigns infeasible actions a logit of −∞ before softmax, while a joint resolver handles collision and residual-capacity conflicts.The same masked-distribution structure is used for trajectory, link, and MA action heads.
  • Optimization: PPO stabilizes learning through clipped policy and value objectives, normalized advantages, and entropy regularization for exploration.The value-loss and entropy terms are weighted by c1 and c2, respectively.

V. NUMERICAL RESULTS

The numerical-results section evaluates the proposed method in a 20-UAV SAGIN simulation and specifies the network, training, and optimization settings used for comparison.

  • Evaluation Design: The experiments assess secrecy energy efficiency as the primary metric and compare proposed algorithms with learning and convex-optimization-based alternating-optimization baselines.The benchmark methods are adapted to the joint optimization problem considered in this work.
  • Simulation Setup: The evaluation uses a 20-UAV swarm flying at fixed altitude through a 500 m×500 m area toward designated destinations.The heterogeneous network contains one GBS, one ABS, and two LEO satellites, with full frequency reuse and 1 MHz per UAV serving link.
  • Training Configuration: The proposed MAPPO-RBHR-CKM configuration trains shared actors with separate flight-direction, link-selection, and MA-position heads.The actor uses two 128-neuron hidden layers, the critic uses three 256-neuron layers, and training lasts 10,000 episodes of 40 slots.

B. Benchmarks

The benchmark suite compares the proposed MAPPO-RBHR framework with alternating-optimization variants, independent and deterministic multi-agent learning baselines, and fixed-design schemes.

  • Alternating-Optimization Baselines: The AO-Joint-CSI baseline jointly updates UAV trajectories, link selections, and MA positions.Other AO variants isolate fixed antennas, fixed ground-network links, or predefined straight-line trajectories.
  • Learning Baselines: IPPO-CSI and MADDPG-CSI use perfect instantaneous CSI to optimize the joint decisions through independent or deterministic multi-agent learning.IPPO makes independent UAV decisions, whereas MADDPG jointly optimizes the swarm decisions.
  • Proposed Methods: The proposed MAPPO-RBHR framework is evaluated with either CKM information or perfect instantaneous CSI while jointly optimizing trajectories, links, and MA positions.This separates the framework’s CKM-assisted and perfect-CSI operating conditions.
  • Evaluation Metric: SEE is the primary comparison metric, with transmit beamforming and receive combining implemented using the MMSE criterion.The benchmark evaluation is conducted during algorithm training and across the proposed and baseline schemes.

C. Performance Comparison

The proposed method exhibits favorable training convergence and higher SEE than key baselines, while ablations and complexity analysis assess its reward, rigid-body, joint-optimization, and CKM-assisted designs.

  • Training Performance: Cumulative return and SEE improve rapidly during training and then stabilize, while critic loss remains low throughout.These trends indicate favorable convergence behavior for the proposed algorithm.
  • Performance Comparison: MAPPO-RBHR-CSI achieves approximately 4.5% higher converged SEE than IPPO-CSI, 43% higher SEE than MADDPG-CSI, and 12% higher SEE than AO-Joint-CSI.MAPPO-RBHR-CKM outperforms most baseline schemes, while MAPPO-RBHR-CSI achieves the highest SEE during training.
  • Hyperparameter Analysis: A learning rate of 1 × 10^-4 provides the highest stable final SEE, and 100 fast-fading realizations provide the best convergence and SEE.Increasing the number of realizations further raises computational cost while providing only marginal additional gains.
  • Mechanism Ablation: The proposed algorithm achieves approximately 16.7% higher converged SEE than MAPPO-RB-CSI and 5.8% higher SEE than MAPPO-HR-CSI.Removing either the rigid-body shift mechanism or hybrid reward mechanism causes noticeable SEE degradation.
  • Scenario Ablation: Disabling trajectory optimization reduces SEE by approximately 13.8%, while variants removing MA or link optimization also achieve lower performance.The ablations support jointly optimizing UAV trajectories, link selection, and MA configuration.
  • Complexity and Deployment: Rigid-body shifts reduce the MA action-output dimension from MKgrid to Nsh, lowering actor computational load and online inference complexity.During deployment, the CKM-assisted scheme avoids real-time CSI acquisition, whereas AO methods require CSI and per-slot outer iterations.

2) CSI Acquisition Overhead:

The proposed MAPPO-RBHR-CKM eliminates real-time CSI acquisition by using offline CKM and satellite ephemeris data during training and deployment. This contrasts with CSI-based learning and optimization baselines that repeatedly acquire or evaluate instantaneous CSI.

  • No real-time CSI acquisition is required by MAPPO-RBHR-CKM during training or deployment.The method replaces real channel probing with offline CKM and ephemeris data.
  • CSI-based MARL baselines acquire instantaneous CSI for all U agents at every training step, resulting in O(EepisodeTU) acquisitions.
  • AO-Joint/FPA/GN-CSI methods require O(RUTndir) real-CSI evaluations across restarts, time slots, and candidate directions.
  • AO-Straight-CSI has lower acquisition complexity, requiring O(UT) because it follows a fixed straight path.
  • The resulting CKM-assisted policy improves SEE while reducing online inference and CSI-acquisition overhead relative to the considered baselines.
Loading 2608.27973v1…