Source-linked AI summary

Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement Learning

Tong Zhang, Yanfei Su, Shuai Wang, Wanli Ni, Chengzhong Xu, Huseyin Arslan

arXiv:2608.27909v1cs.ITcs.AI

TL;DR

Low-altitude wireless networks face dynamic channels, blockages, and interference that challenge fixed MIMO arrays serving UAVs and eVTOL aircraft. The paper proposes an EM-DT-assisted MARL framework with two-stage transfer learning, and its case study reports a 118.5% sum-rate gain over a fixed-position baseline through joint FA positioning and beamforming.

  • Problem

    Dynamic air-ground and air-air channels, sudden blockages, and heterogeneous interference make fixed MIMO arrays ill-suited to tracking mobile UAVs and eVTOLs.

  • Method

    The paper combines electromagnetic digital twins, multi-agent reinforcement learning, and sim-to-real transfer learning to optimize low-altitude FA networks.

  • Results

    118.5% sum-rate gain over traditional fixed-position baselines is achieved through joint FA-position and downlink-beamforming optimization.

  • Takeaways & Limitations

    Dynamic FA reconfiguration and beam steering can improve sum-rate in the studied two-BS two-UAV low-altitude network case study.

Abstract

from arXiv · show

Low-altitude wireless networks (LAWNs) integrate terrestrial and aerial platforms to provide ubiquitous communication, sensing, and localization services for unmanned aerial vehicles (UAVs) and electric vertical takeoff and landing (eVTOL) aircraft. However, dynamic air-ground and air-air channels, abrupt blockages, and heterogeneous interference hinder the realization of this goal. Nevertheless, fluid antenna (FA), a cutting-edge multiple-input multiple-output (MIMO) technique, overcomes these challenges by reconfiguring antenna positions to unlock additional spatial degrees-of-freedom. In this paper, towards bringing low-altitude FA networks into reality, we study the fast and high-performance FA reconfiguration for low-altitude FA networks with multi-agent reinforcement learning (MARL). Specifically, we present an electromagnetic digital twin (EM-DT)-assisted MARL framework. To fill the sim-to-real gap, we introduce a two-stage transfer learning framework. Our case study shows that joint FA positions and beamforming optimization can enhance the system sum-rate by 118.5%, compared to the fixed position baseline. This gain comes from the dynamic millisecond timescale reconfiguration of FA arrays and the adaptive steering of beams toward aerial users with mobility.

I. INTRODUCTION

Low-altitude wireless networks must support aerial users despite rapidly changing channels, blockages, and interference that fixed MIMO arrays cannot readily track. The paper proposes an EM-DT-assisted MARL framework with sim-to-real transfer, reporting a 118.5% sum-rate gain from joint FA positioning and beamforming.

  • Motivation: Low-altitude wireless networks target ubiquitous communication, sensing, and localization for UAVs and eVTOL aircraft.They are intended to provide high-data-rate, low-latency, and ultra-reliable connectivity in three-dimensional airspace.
  • Motivation: Dynamic air-ground and air-air channels, sudden blockages, and heterogeneous interference make fixed traditional MIMO arrays ill-suited to fast-moving aerial users.These conditions include rapid LoS/NLoS transitions and interference from aerial LoS connectivity to multiple base stations.
  • FA approach: FA technology adds spatial degrees of freedom by reconfiguring antenna positions, shapes, and radiation patterns rather than changing only fixed-array excitation weights.This position-domain controllability targets trajectory-dependent links, blockages, and interference.
  • Contributions: The proposed framework synthesizes EM-DT, MARL, and sim-to-real transfer learning for robust low-altitude FA network operation.The contribution presents a unified low-altitude FA network view and identifies future MARL-driven research directions.
  • Contributions: 118.5% sum-rate gain over traditional fixed-position baselines comes from jointly optimizing FA positions and downlink beamforming.The reported gain is enabled by millisecond-timescale FA reconfiguration and agile beam steering that tracks aerial users.

II. LOW-ALTITUDE FLUID ANTENNA NETWORK

The envisioned low-altitude FA network unifies communication, sensing, and localization through reconfigurable antenna positioning across aerial users and base stations. It combines FA-enabled users, FAS base stations, and more accurate position-aware beamforming.

  • Architecture: The proposed architecture integrates communication, sensing, and localization through FA positioning across users and base stations.This unified view is distinguished from existing low-altitude FA frameworks.
  • Architecture: Low-altitude FA networks include aerial users and ground or aerial base stations as their two major physical entity categories.Aerial users are served by FAS base stations in the illustrated network.
  • Aerial Users: FA-empowered aerial users require high-data-rate, low-latency, ultra-reliable connectivity and can support distributed sensing and user-to-user communication.The same FA capabilities support both communications and aerial-user-driven sensing functions.

1) Aerial Users:

FAS base stations can adapt coverage and interference management for mobile aerial users, while low-altitude FA networks are designed to satisfy varied connectivity and service requirements.

  • Aerial Users:: Ground and aerial FAS base stations can dynamically reconfigure FA positions to steer beams, mitigate interference, and support mobility.Aerial FAS base stations can also operate as on-demand flying base stations or relays for remote and disaster-stricken areas.
  • Aerial Users:: Control-domain coordination combines base-station handover and interference information with users’ CSI, positions, and kinematics.These inputs support real-time adaptation to dynamic channels, mobility, and heterogeneous interference.
  • Aerial Users:: FA position reconfiguration is intended to provide more favorable channel conditions for applications such as delivery, monitoring, inspection, and air taxis.The network is expected to extend connectivity to aerial users in urban canyons, remote rural areas, and high-mobility corridors.

1) Ubiquitous Communication:

Low-altitude networks must support tracking, collision avoidance, monitoring, and heterogeneous communication requirements across the airspace. Pervasive sensing and localization can combine multiple measurements into a unified situational view.

  • 1) Ubiquitous Communication:: Low-altitude FA networks should provide sensing and localization for UAV and eVTOL tracking, collision avoidance, and airspace monitoring.These functions span the low-altitude airspace.
  • 1) Ubiquitous Communication:: Representative EM-DTs pair satellite imagery in the top row with corresponding RANPLAN ACADEMIC V7.1 electromagnetic digital twins in the bottom row.Examples cover Shenzhen University Town, Guangzhou Zhujiang New Town CBD, Wembley Stadium, and the Las Vegas Strip.
  • 1) Ubiquitous Communication:: Sensing and localization can fuse time difference of arrival, angle of arrival, and received signal strength measurements from base stations and aerial users.The fused data form a unified situational view for navigation safety and airspace awareness.
  • 1) Ubiquitous Communication:: Low-altitude FA networks must accommodate high-throughput data delivery, low-latency control signaling, and ultra-reliable communication for mission-critical aerial operations.These requirements reflect heterogeneous quality-of-service demands.

III. ELECTROMAGNETIC DIGITAL TWIN-ASSISTED MULTI-AGENT REINFORCEMENT LEARNING

EM-DT provides a high-fidelity, closed-loop virtual environment for safe MARL training and policy testing in low-altitude FA networks. The optimization is framed as a Dec-POMDP, with actor-critic learning and millisecond policy inference supporting decentralized control.

  • EM-DT-assisted MARL: EM-DT models electromagnetic propagation and couples mobility, interference, and FA control actions in a closed loop for MARL training.Its core modeling requirements include 3D propagation effects, aerial-user trajectories, and network-level handover dynamics.
  • EM-DT-assisted MARL: EM-DT offers low-cost, risk-free, parallel training and pre-deployment testing under blockages, trajectory deviations, and interference bursts.It also supports domain randomization over propagation, interference, and mobility parameters for sim-to-real transfer.
  • Multi-Agent Reinforcement Learning: MARL learns high-performance policies through iterative environment interaction and actor-critic learning.
  • Multi-Agent Reinforcement Learning: The low-altitude FA optimization is formulated as a decentralized partially observable Markov decision process tailored to task requirements.Figure 3 illustrates the associated agents, states, actions, and rewards.
  • Multi-Agent Reinforcement Learning: State design supplies the critic with sufficient information for cooperative credit assignment while avoiding redundant features that increase sample complexity.Candidate state variables include FA positions, aerial-user kinematics, CSI, and interference statistics.

1) State:

Each agent receives a local observation from the environment, while the state can include network-wide physical and channel information. These inputs support centralized training and decentralized execution.

  • 1) State:: A local observation is information independently received by each agent from the environment.It may include the agent’s FA positions, kinematics, CSI, and interference conditions.
  • 1) State:: Local observations drive actor networks during decentralized execution and complement the global state during centralized training.
  • 1) State:: The state may include FA positions, UAV and eVTOL kinematics, CSI, interference statistics, and other task-relevant indicators.
  • 1) State:: Actions should account for hardware speed, movement latency, and settling time by updating continuous FA positions within feasible moving regions.Scanning the entire position space in every coherence interval is not prescribed for real systems.

3) Action:

The framework defines actions, rewards, centralized training, and transfer procedures for controlling low-altitude FA networks. A two-stage process combines EM-DT pre-training with limited real-system measurements to address sim-to-real challenges.

  • 3) Action:: Rewards typically represent network utility, such as weighted sum rate or a weighted combination of sum rate and sensing performance.Penalty terms can enforce constraints, while reward shaping can address sparse rewards.
  • 3) Action:: CTDE uses a centralized critic with global EM-DT information during training and decentralized actors using local observations during execution.
  • 3) Action:: The sim-to-real gap includes hardware impairments, FA switching latency, mutual coupling, blockages, weather attenuation, and GPS-denied conditions.These factors may not be fully captured during offline training and can affect policy-action feasibility.
  • 3) Action:: Stage-I pre-trains policies on tens of thousands of EM-DT trajectories with domain randomization over factors expected in real systems.Offline and off-policy MARL methods reuse datasets and do not require data generated by the current policy.
  • 3) Action:: Stage-II initializes transfer learning with pre-trained policies and high-quality datasets manually collected from the real system.Centralized training updates policies while data acquisition remains decoupled from policy execution.

V. CASE STUDY

The case study evaluates joint FA-position and downlink-beamforming optimization in a two-BS, two-UAV EM-DT scenario. After MATD3 training, the learned system outperforms fixed antenna positions and retains its gain under unseen testing conditions.

  • V. CASE STUDY: The EM-DT case study uses 2 FAS-BSs, each serving 1 UAV and equipped with 4 FAs, to maximize network sum rate.Optimization observes transmit-power, FA-region, and minimum-spacing constraints.
  • V. CASE STUDY: After MATD3 training, FAS-BSs reconfigure FA positions to steer main lobes toward desired aerial users and suppress interference toward undesired users.
  • V. CASE STUDY: 118.5% higher network sum-rate is achieved than the fixed antenna position baseline under previously unseen testing conditions.Testing uses fixed UAV height and trajectory conditions not encountered during training.
  • V. CASE STUDY: Training performance fluctuates because UAV heights, trajectories, and channel realizations are randomized, while joint optimization consistently outperforms the fixed-position baseline.

VI. OPEN ISSUES AND FUTURE DIRECTIONS

Practical deployment of low-altitude FA networks remains constrained by reconfiguration overhead and hardware limits. These constraints can consume the channel coherence budget and motivate hardware-aware control.

  • At speeds up to 30 m/s, coherence times at high carrier frequencies can fall below one millisecond, constraining FA reconfiguration.At 5.5 GHz, 30 m/s produces a 550 Hz Doppler shift and coherence time on the order of 1 ms.
  • Switching, actuator motion, calibration, and settling overhead limit practical reconfiguration rates in fast-moving UAV/eVTOL scenarios.Microsecond-level RF switching does not eliminate millisecond-level actuator motion and additional overhead.
  • Repeated FA scanning and physical repositioning for channel estimation may exceed the coherence budget, erasing theoretical benefits.The paper calls for switching-cost-aware control, per-coherence-interval budgets, and predictive positioning using temporal channel correlation.

B. CSI Usage Reduction

CSI usage in low-altitude FA networks is challenged by acquisition overhead, outdated information, safety requirements, environmental-model inaccuracies, and mismatched temporal scales. The paper points toward robustness, prediction, and hierarchical decision-making as responses.

  • CSI Usage Reduction: Accurate CSI acquisition imposes pilot, estimation, and feedback overhead, while fast-varying channels can become outdated before use.Suggested responses include event-triggered acquisition, partial observation, domain randomization, meta-learning, and EM-DT-based channel prediction.
  • CSI Usage Reduction: Safety-critical missions require reliable FA controllers, but stochastic MARL policies lack worst-case guarantees and remain difficult to verify.The paper identifies safe MARL integration, certification frameworks, and standardized testbeds as needed directions.
  • CSI Usage Reduction: Static or slowly updated EM-DT maps cannot fully represent dynamic obstructions, degrading ray-tracing fidelity in real low-altitude airspace.Moving UAVs, vehicles, and weather conditions are identified as sources of environmental mismatch.
  • CSI Usage Reduction: Different temporal scales between static BSs and fast-moving aerial users require hierarchical MARL for coordinated resource allocation, handover, FA reconfiguration, and beamforming.Slow-timescale variables handle long-term allocation and handover, while fast-timescale variables handle instantaneous reconfiguration and beamforming.

F. System Support for Artificial Intelligence

Low-altitude FA networks can support aerial artificial-intelligence applications by linking UAV observations to ground-server vision-language and language-model processing. The conclusion reports improved sum-rate from jointly optimizing FA positions and downlink beamforming.

  • F. System Support for Artificial Intelligence: Aerial visual question answering uploads UAV-team images through FA networks to a ground server for VLM description and LLM query answering.The application supports queries over long observation horizons, while UAVs may retain heterogeneous semantic information.
  • F. System Support for Artificial Intelligence: Joint FA-position and downlink-beamforming optimization consistently improved sum-rate over the fixed-position baseline in a two-BS two-UAV case study.The improvement came from dynamically reshaping beampatterns toward UAVs.
Loading 2608.27909v1…