Source-linked AI summary
Aerial Reliable Collaborative Communications for Terrestrial Mobile Users via Evolutionary Multi-Objective Deep Reinforcement Learning
Geng Sun, Jian Xiao, Jiahui Li, Jiacheng Wang, Jiawen Kang, Dusit Niyato, Shiwen Mao
TL;DR
The paper tackles limited UAV communication capability in dynamic mobile-user environments with interference and time-varying channels. It combines collaborative beamforming, a Gaussian-Markov mobility model, and EMOPPO-VLH to optimize rate and flight energy. Simulations show diverse non-dominated policies and superior performance across benchmark comparisons and scales.
Problem
Limited UAV energy and transmit power constrain communication range, while dynamic mobile users and time-varying channels complicate jointly optimizing rate and flight energy.
Method
The paper forms a UAV-enabled virtual antenna array, models mobility with a Gaussian-Markov process, and solves the resulting MOP as an MOMDP using EMOPPO-VLH with LSTM networks and hyper-sphere task selection.
Results
The proposed method generates diverse high-quality non-dominated policies and outperforms benchmark algorithms across different scales, with user mobility not affecting its effectiveness.
Takeaways & Limitations
Collaborative beamforming combined with evolutionary multi-objective DRL provides a range of rate–energy trade-offs for dynamic UAV-assisted mobile communications.
Takeaways & Limitations
The transmission model assumes identical maximum transmit power for every UAV in the virtual antenna array.
Abstract
from arXiv · showhide
Unmanned aerial vehicles (UAVs) have emerged as the potential aerial base stations (BSs) to improve terrestrial communications. However, the limited onboard energy and antenna power of a UAV restrict its communication range and transmission capability. To address these limitations, this work employs collaborative beamforming through a UAV-enabled virtual antenna array to improve transmission performance from the UAV to terrestrial mobile users, under interference from non-associated BSs and dynamic channel conditions. Specifically, we introduce a memory-based random walk model to more accurately depict the mobility patterns of terrestrial mobile users. Following this, we formulate a multi-objective optimization problem (MOP) focused on maximizing the transmission rate while minimizing the flight energy consumption of the UAV swarm. Given the NP-hard nature of the formulated MOP and the highly dynamic environment, we transform this problem into a multi-objective Markov decision process and propose an improved evolutionary multi-objective reinforcement learning algorithm. Specifically, this algorithm introduces an evolutionary learning approach to obtain the approximate Pareto set for the formulated MOP. Moreover, the algorithm incorporates a long short-term memory network and hyper-sphere-based task selection method to discern the movement patterns of terrestrial mobile users and improve the diversity of the obtained Pareto set. Simulation results demonstrate that the proposed method effectively generates a diverse range of non-dominated policies and outperforms existing methods. Additional simulations demonstrate the scalability and robustness of the proposed CB-based method under different system parameters and various unexpected circumstances.
1 INTRODUCTION
The paper addresses UAV communication limits by combining collaborative beamforming with memory-based modeling of mobile users and long-term multi-objective optimization. It proposes EMOPPO-VLH to balance transmission rate and flight energy while adapting to dynamic conditions.
- Motivation: A single UAV’s limited onboard energy and transmit power restrict its service area and communication capability.
- Collaborative beamforming: Collaborative beamforming lets multiple UAVs form a virtual antenna array that enhances signal strength, directivity, communication range, and interference resistance.
- Optimization challenge: UAV positioning and excitation-current weights jointly affect beam quality, transmission rate, and flight energy consumption.
- System model: The system models mobile-user movement with a Gaussian-Markov random mobility model under non-associated-BS interference and time-varying channels.
- Optimization formulation: The long-term MOP maximizes total achievable rate while minimizing overall UAV flight energy across sequential decisions.
- Proposed method: EMOPPO-VLH combines vectorized value functions, LSTM networks, and hyper-sphere task selection to handle dynamic multi-objective decisions and diversify Pareto solutions.
- Evaluation: Extensive simulations report high-quality non-dominated policies that outperform benchmark algorithms across scales, while user mobility does not affect the method’s effectiveness.
2 RELATED WORK
Prior work covers UAV communications, collaborative beamforming, and multi-objective optimization, but leaves a gap in dynamic mobile-user settings. This paper addresses that gap with CB and multi-policy DRL designed for realistic, rapidly changing environments.
- UAV-assisted terrestrial communications: Existing UAV communication studies optimize resources, altitude, or related objectives, but do not establish the paper’s combined CB and mobile-user setting.
- Research gap: The paper differs by balancing transmission performance and energy consumption through multi-objective optimization for dynamic UAV-assisted communications.
- Collaborative beamforming: Earlier CB studies address sensor networks or UAV arrays, yet do not explicitly integrate CB with dynamic terrestrial mobile users under realistic changing conditions.
- Multi-objective optimization: Traditional multi-objective methods become impractical for real-time decisions in highly dynamic environments.
- Deep reinforcement learning: Existing DRL approaches often combine objectives into one weighted reward and produce a single policy, potentially overlooking objective conflicts and solution diversity.
3 SYSTEM MODELS AND PRELIMINARIES
The system models a UAV swarm using collaborative beamforming to serve a terrestrial mobile user while accounting for interference, channel fading, UAV motion, energy consumption, and temporally correlated user mobility.
- 3.1 Network Segments: The modeled network contains rotary-wing UAVs, a terrestrial mobile user, a potentially interfering non-associated BS, and a central UAV controller.The controller manages UAVs over a separate control channel, while the UAVs transmit to the user through a UVAA.
- 3.2 Communication Model: Multiple UAVs form a UVAA whose array factor depends on UAV positions and excitation current weights, shaping the beam toward the user.The excitation weights determine transmitted signal amplitude and phase, while the user-directed gain is derived from the array factor.
- 3.2 Communication Model: The channel model includes large-scale path loss, Rician small-scale fading, and interference from a non-associated BS when computing the user's SINR.The interference model accounts for BS sidelobe leakage and the BS-to-user channel gain.
- 3.3 UAV Movement and Energy Consumption Models: UAV movement is modeled in three dimensions, with horizontal direction and horizontal or vertical distances determining successive positions and flight energy consumption.The energy model covers rotary-wing propulsion and climbing or descending actions.
- 3.4 Memory-based User Random Mobility Model: The terrestrial user's mobility follows a memory-based random walk in which current speed and direction depend on their previous values.This introduces temporal dependencies into the user trajectory rather than treating each movement independently.
4 PROBLEM FORMULATION AND ANALYSIS
The paper formulates a long-term multi-objective problem for collaborative UAV beamforming, balancing transmission rate against flight energy in dynamic mobile-user environments. It analyzes the problem’s trade-offs, NP-hardness, and need for adaptive solution methods.
- 4.1 Problem Formulation: The decision variables include UAV movement directions, horizontal and vertical flight distances, and excitation-current weights across time slots.These variables determine the UAVs’ three-dimensional positions and beamforming configuration.
- 4.1 Problem Formulation: The formulation maximizes cumulative UVAA-to-user transmission rate over time while minimizing total UAV motion energy consumption.The two objectives are defined over multiple time slots and depend on UAV movement and excitation-current decisions.
- 4.1 Problem Formulation: Weighted-sum and constraint-based methods may miss parts of the rate–energy trade-off and may struggle when objective priorities change in dynamic environments.The paper therefore uses multi-objective optimization to preserve a broader set of non-dominated solutions.
- 4.1 Problem Formulation: The model constrains excitation weights, UAV motion, operating areas, and minimum inter-UAV distances.The constraints bound weights and coordinates while preventing UAVs from leaving the designated region or approaching too closely.
- 4.2 Problem Analysis: The formulated multi-objective problem is NP-hard because a simplified single-slot version contains a nonlinear knapsack problem.The simplification fixes UAV and user positions and leaves excitation-current weights as decision variables.
- 4.2 Problem Analysis: Improving beam directivity and gain requires frequent UAV repositioning, which increases flight energy consumption.The UAVs must follow the user’s location over successive time slots, creating a direct conflict between the objectives.
- 4.2 Problem Analysis: Time-varying channels, interference from non-associated BSs, and random user mobility make the optimization environment dynamic.These conditions can destabilize communication links and require trajectories to adapt continually to user movement.
- 4.2 Problem Analysis: Because the problem is NP-hard, sequential, long-term, and uncertain, the paper proposes an evolutionary deep-reinforcement-learning approach.The proposed method is intended to support real-time decisions in the changing environment.
5 ALGORITHM
The paper transforms the dynamic multi-objective optimization problem into a multi-objective Markov decision process and uses deep reinforcement learning to learn adaptive policies. Multi-objective optimization targets a non-dominated policy set rather than a single solution.
- 5 ALGORITHM: Conventional optimization methods are challenged by NP-hardness, nonlinear objectives, long-term sequential decisions, uncertainty, and a massive cooperative action space.These characteristics make exhaustive, convex, evolutionary, and tightly constrained online approaches difficult to apply effectively.
- 5 ALGORITHM: Deep reinforcement learning combines reinforcement learning with neural networks to address sequential decision-making under dynamic and uncertain conditions.An agent updates its policy through environmental interaction to maximize cumulative rewards.
- 5 ALGORITHM: The proposed framework applies multi-objective optimization theory and seeks a non-dominated Pareto set instead of one policy optimizing all objectives simultaneously.This preserves alternatives across conflicting transmission-rate and energy objectives.
- 5 ALGORITHM: The formulated problem is represented as an MOMDP with state and action spaces, transition probabilities, vector rewards, a discount factor, and a weight distribution.Each reward component corresponds to one of the multiple objectives.
- 5 ALGORITHM: The paper transforms the formulated MOP into an MOMDP and solves it using a DRL-based method.This provides the modeling bridge from optimization objectives to sequential policy learning.
5.2 MOMDP Formulation
The MOMDP represents UAV and user positions as environmental state, while actions control UAV movement and excitation weights. Its reward structure jointly encourages transmission rate and low energy use while penalizing constraint violations.
- 5.2.1 State Space: The MOMDP state contains the positions of all UAVs and the terrestrial mobile user.Positioning devices such as GPS provide the locations used by the controller.
- 5.2 MOMDP Formulation: The controller therefore links observed UAV and user positions to movement, beamforming, and energy-related decisions at each time slot.The state and action definitions jointly operationalize the formulated optimization variables.
- 5.2.2 Action Space: Each UAV action specifies its horizontal direction, horizontal flight distance, vertical flight distance, and excitation-current weight.These controls correspond to the decision variables in the optimization problem.
- 5.2.2 Action Space: The action space uses movement directions rather than 3D Cartesian coordinates.The paper states that this captures temporal dynamics, aligns with practical control, and simplifies learning.
- 5.2.3 Reward Function: The reward system is designed to maximize total achievable rate while minimizing total UAV energy consumption.It provides multi-objective feedback after each action rather than a single-objective scalar target.
- 5.2.3 Reward Function: The reward assigns rate and negative-energy components, with scaled penalties applied when actions violate area or collision constraints.The indicator variable distinguishes valid actions from boundary breaches or adjacent-UAV collisions.
5.3 Multi-Objective Proximal Policy Optimization Task
The proposed multi-objective PPO task optimizes policies under different objective weights while retaining vector-valued value estimates. An LSTM is added to capture historical user movement and address uncertainty from time-varying channels.
- 5.3 Multi-Objective Proximal Policy Optimization Task: The learning task is defined by a weight vector and a policy to be optimized.Its objective is to maximize the weighted-sum reward associated with the multiple objectives.
- 5.3 Multi-Objective Proximal Policy Optimization Task: The proposed EMOPPO-VLH optimizes policies using different objective weights.This supports learning across multiple weighted multi-objective tasks rather than fixing one weighting.
- 5.3 Multi-Objective Proximal Policy Optimization Task: PPO is used as an actor-critic, on-policy, policy-gradient method with a clipped surrogate objective that limits large policy deviations.The clipping mechanism controls the probability-ratio update between new and old policies.
- 5.3 Multi-Objective Proximal Policy Optimization Task: The single-objective PPO value function is replaced by a vectorized value function for simultaneous handling of multiple optimization objectives.The vectorized function associates each state with a vector of expected returns.
- 5.3 Multi-Objective Proximal Policy Optimization Task: The target value function uses the immediate reward plus the discounted value of the next state.This target is used in the value-network update.
- 5.3 Multi-Objective Proximal Policy Optimization Task: Time-varying channel fading creates uncertain state transitions, while historical user behavior is not effectively captured by fully connected networks.The paper integrates LSTM to exploit temporal sequences in user movements.
5.4 The Proposed EMOPPO-VLH Method
EMOPPO-VLH combines warm-up and evolutionary stages to generate diverse Pareto policies for the multi-objective UAV communication problem. LSTM-based MOPPO captures temporal dependencies, while evolutionary task selection improves offspring quality and Pareto-set diversity.
- Framework overview: EMOPPO-VLH first creates a primary policy population during warm-up, then iteratively updates tasks, selects tasks, generates offspring, and updates the external Pareto archive.The framework contains warm-up and evolutionary stages, with LSTM-MOPPO optimizing selected tasks.
- LSTM network structure: LSTM networks replace the first fully connected layer in actor and critic networks to capture hidden user-movement patterns and temporal channel dependencies.The design uses input, output, and forget gates to retain relevant long-term states and discard irrelevant information.
- Warm-up stage: Each learning task combines an evenly distributed weight vector with LSTM-based policy and value networks, and optimized policies form the initial population.The task representation is Γi = ⟨ωi, πθi, Vπθi⟩.
- Evolutionary deep reinforcement learning process: The algorithm generates n·niter offspring by preserving new tasks after each iteration, increasing exploratory capacity and offspring-population diversity.This addresses the original MOPPO strategy of retaining tasks only after niter iterations.
- Evolutionary deep reinforcement learning process: Hyper-sphere-based selection computes weighted objective values and assigns higher selection probability to sub-hyper-spheres containing fewer policies.The selected tasks are optimized by LSTM-MOPPO for nevo evolutionary iterations.
5.5 Practical Implementation of EMOPPO-VLH-based UAV Management
EMOPPO-VLH uses centralized training with a simulation environment and deploys the converged policy through a central controller. Environmental changes or UAV failures can be handled by modifying parameters and retraining, with transfer learning and backups suggested for adaptation.
- Training phase: Training is recommended on a central controller because reward computation and evolutionary learning require information from all agents and substantial computational resources.The training environment is mathematical and collects user rates, UAV energy consumption, and agent locations.
- Simulation environment: The simulation uses memory-based user mobility, Rician channel, and UAV energy-consumption models derived from real-world data to support realistic training.After convergence, the trained policy can be deployed in practical environments.
- Deployment phase: During deployment, the trained policy receives real-environment state information and no longer requires real-time data on data rates and UAV energy consumption.A central controller can access UAV and terrestrial-user locations through positioning devices.
- Adaptation: When environment parameters change, administrators can modify the simulation and retrain the algorithm, while transfer learning can help achieve quick convergence.Examples include changing UAV masses or user movement patterns.
- Fault tolerance: For UAV failures, administrators can retrain policies for different swarm sizes in advance and use backup UAVs or redundancy mechanisms for fault tolerance.The proposed mechanism is intended to reduce delay and maintain communication performance during failures.
5.6 Computational Complexity of the Proposed EMOPPO-VLH
The computational complexity of EMOPPO-VLH grows with the number of generations, tasks, evolutionary iterations, and deep-network size.
- Complexity expression: The total complexity of EMOPPO-VLH is O(Gmax · n · nevo · (Σ_l=1^L m_l−1 · m_l + M)).Here n is the number of tasks, nevo is the number of evolutionary iterations, and L and ml denote network layers and units.
6 SIMULATION RESULTS
The simulations evaluate EMOPPO-VLH across UAV swarm scales and against evolutionary, multi-policy reinforcement-learning, and PPO-based baselines. EMOPPO-VLH achieves the strongest optimization results and Pareto-front convergence and diversity across the tested scenarios.
- 6.1 Simulation Setups: The experiments test small-scale and large-scale UAV swarms containing 8 and 16 UAVs, respectively, to assess scalability and robustness.The mobile user moves at an average speed of 1 m/s within a 100 m × 100 m square, while UAV horizontal and vertical flight limits are 20 m and 10 m per slot.
- 6.3 Baselines: EMOPPO-VLH is compared with MOEA/D, MOPSO, EDDPG, ETD3, EPPO, and EPPO with GRU using total achievable rate and total UAV energy consumption.The multi-objective evaluation additionally uses IGD and HV to assess Pareto-front proximity and policy diversity.
- 6.3 Baselines: The proposed method has complexity comparable in asymptotic form to several reinforcement-learning baselines but may scale more effectively in practice through neural-network generalization.The cited comparison concerns EDDPG, ETD3, EPPO, and EPPO with GRU, whose complexity is O(Gmax · n · nevo · (Σ_{l=1}^{L} m_{l−1} · m_l + M)).
- 6.4 Performance Evaluation: EMOPPO-VLH outperforms the other optimization approaches in both small-scale and large-scale scenarios on total achievable rate and total UAV energy consumption.The authors associate the higher achievable rate with greater interference resistance, shorter transmission time, and reduced UAV energy consumption.
- 6.4 Performance Evaluation: MOEA/D and MOPSO underperform the other algorithms because high-dimensional decision variables and dynamic uncertainty make rapid non-dominated-policy convergence computationally difficult.The decision space contains 4 · N · T variables, and repeated optimization in the dynamic environment creates computational overhead and slow response.
- 6.4 Performance Evaluation: EMOPPO-VLH achieves the smallest IGD values and largest HV values across all scenarios, indicating superior convergence and diversity among the compared algorithms.The reported explanation is that LSTM captures temporal dependence while hyper-sphere task selection broadens non-dominated policy coverage.
7 DISCUSSION
The discussion section directs readers to supplemental Appendix D for the effectiveness, reasonableness, and scalability analysis of the proposed method.
- 7 DISCUSSION: Appendix D contains the detailed analysis of the proposed method’s effectiveness, reasonableness, and scalability.The main text does not provide those analyses in this section.
8 CONCLUSION
The paper studies collaborative-beamforming UAV communications for mobile terrestrial users and proposes EMOPPO-VLH to optimize rate–energy trade-offs in dynamic conditions. Simulations report stronger performance, convergence, diversity, scalability, and robustness than the benchmark methods and settings considered.
- 8 CONCLUSION: The system uses a UAV swarm with collaborative beamforming to communicate with mobile terrestrial users under non-associated-BS interference and time-varying channels.A more realistic random-walk model is introduced to represent user movement.
- 8 CONCLUSION: The formulated multi-objective problem maximizes achievable rate while minimizing UAV energy consumption.The objectives create rate–energy trade-offs over the communication process.
- 8 CONCLUSION: EMOPPO-VLH obtains non-dominated policies with different trade-offs and outperforms benchmark algorithms in both small-scale and large-scale scenarios.The resulting policy set is intended to accommodate different user preferences.
- 8 CONCLUSION: IGD and HV results indicate superior convergence and diversity, while additional simulations support scalability and robustness under varied parameters and unexpected circumstances.These conclusions concern the proposed EMOPPO-VLH and collaborative-beamforming method within the evaluated UAV communication settings.