Source-linked AI summary
Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks
Yassine Afif, Ashutosh Balakrishnan, Philippe Martins, Mohammed Almekhlafi, Antoine Lesage-Landry, Gunes Karabulut Kurt
TL;DR
The paper addresses joint user association, power allocation, and handover management across LEO, MEO, and GEO networks under fast LEO dynamics. It combines a layer-aware MARL association policy with closed-loop convex power allocation and evaluates the design on realistic Nairobi data. The proposed policy achieves about 92% of greedy throughput with more than four times fewer handovers and about 14% higher throughput than stay.
Problem
Jointly coordinating association, power allocation, and handover management across heterogeneous orbital layers remains insufficiently addressed by prior work.
Method
The paper decomposes the mixed-integer nonlinear problem into a layer-aware MARL association policy and a convex power-allocation subproblem solved in closed loop.
Results
92% of greedy’s rate was achieved at 7.55 Mbps per time slot with 6.2 handovers per time slot, while stay achieved 6.6 Mbps per time slot.
Takeaways & Limitations
The learned multi-orbit policy attains a favorable throughput–handover trade-off and exhibits emergent offloading across orbital layers.
Abstract
from arXiv · showhide
Future sixth-generation non-terrestrial networks are expected to combine low Earth orbit (LEO), medium Earth orbit (MEO), and geostationary Earth orbit (GEO) satellites, whose complementary layers must be coordinated through joint user association, power allocation, and handover management under fast LEO dynamics. This paper studies this problem by formulating it as a mixed-integer nonlinear program and decomposing it into a multi-agent reinforcement learning (MARL) policy that selects the associations and a convex power-allocation subproblem solved exactly at each time slot that defines the reward of the MARL part. The association policy is trained with multi-agent proximal policy optimization (MAPPO) and the targeted multi-agent communication (TarMAC) mechanism, and is made aware of the orbital layer through a state that encodes layer-dependent handover penalties. Evaluated on a realistic multi-constellation scenario built from real two-line element data over Nairobi, Kenya, the proposed policy reaches 92% of the throughput of a greedy signal-to-noise ratio (SNR) maximizing scheme while triggering more than four times fewer handovers, and improves throughput by roughly 14% over a conservative stay heuristic. Compared to an LEO-only learned policy of identical architecture, it attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
I. INTRODUCTION
Multi-layer satellite networks combine LEO, MEO, and GEO systems with complementary coverage, latency, mobility, and capacity characteristics. Their joint management must coordinate power allocation and handovers under heterogeneous, dynamic conditions.
- LEO satellites provide low propagation delays and high data rates, but limited coverage and rapid relative motion create frequent reassociation demands.
- MEO and GEO layers can complement LEO capacity by supporting traffic offloading during congestion, coverage gaps, or service disruptions.
- Joint resource management must account for satellite mobility, heterogeneous coverage footprints, limited onboard resources, and dynamic traffic demands.
- Handover and power allocation are tightly coupled optimization problems as multi-layer constellation scale and heterogeneity increase.
- Prior optimization-based handover methods couple association and power allocation with penalties, but solving each slot independently neglects the problem’s sequential nature.
2) Multi-Orbit Integration:
Multi-orbit satellite management requires coordination across heterogeneous layers and among distributed agents. The paper addresses this gap with communication-enabled MARL and a closed-loop association–power-allocation design.
- 2) Multi-Orbit Integration:: MARL supports decentralized coordination among many agents through centralized training and decentralized execution, with MAPPO and communication mechanisms among representative approaches.
- 2) Multi-Orbit Integration:: TarMAC jointly learns message content and recipients, providing communication suited to structured multi-agent interactions.
- 2) Multi-Orbit Integration:: Prior work commonly treats handover, multi-orbit integration, and distributed resource allocation in isolation.
- 2) Multi-Orbit Integration:: Existing learned handover schemes rarely span multiple orbits or use inter-agent communication, while MARL resource-allocation methods rarely integrate handover in multi-orbit settings.
- B. Contributions: The paper formulates joint association, power allocation, and handover management as a hierarchical decomposition with MARL association followed by closed-loop convex power allocation.
- B. Contributions: The evaluation uses real Nairobi multi-constellation data and measures layer-wise associations, revealing emergent offloading behavior beyond throughput and handover metrics.
II. SYSTEM MODEL AND PROBLEM FORMULATION
The system model represents a three-layer satellite constellation serving ground users through time-varying satellite–beam associations. Visibility and orbital-layer structure determine which links can be selected over discrete time slots.
- A. Network Topology: The modeled network serves single-antenna ground users with satellites partitioned into LEO, MEO, and GEO layers.
- A. Network Topology: Figure 1 depicts ground users served by satellites from the LEO, MEO, and GEO layers.
- A. Network Topology: Each satellite has indexed beams, and the time horizon is divided into discrete decision slots.
- A. Network Topology: A binary association variable indicates whether user u is linked to satellite–beam pair (s,b) at time t.
- A. Network Topology: At each slot, orbital propagation supplies time-varying geometry, including elevation angle and slant range, for every user–satellite pair.
- A. Network Topology: The visibility indicator is determined using elevation angles and a layer-dependent visibility threshold.
B. Rate Analysis
The rate model combines channel gain, transmit power, noise, bandwidth, path loss, and layer-dependent antenna gains. User rate aggregates the achievable rates of its associated links.
- The channel model captures large-scale propagation effects and Rician small-scale fading with dominant line-of-sight components.
- The received SNR is computed from channel gain and transmit power together with noise spectral density, bandwidth, and total path loss.
- Transmit antenna gains are layer-dependent and assumed to increase with orbital altitude to compensate for greater propagation losses.
- The achievable Shannon rate is defined for each user, satellite, beam, layer, and time slot from the link model.
- A user’s achievable rate aggregates contributions from its associated links only.
C. Handover Modelling
The model represents handovers as beam switches between consecutive slots and assigns penalties according to the orbital-layer gap between serving beams.
- A handover occurs when user u switches its serving beam between consecutive time slots.
- The handover indicator distinguishes users that stay on the same beam from those that switch beams.
- The layer gap between the previous and new serving beams determines the handover penalty.
- With 0 ≤ α1 < α2 < α3 < 1, larger altitude gaps incur longer service interruptions.
D. Problem Formulation
The joint association and power-allocation problem maximizes throughput while enforcing rate, visibility, capacity, and power constraints, but binary–continuous coupling makes it non-convex and intractable directly.
- The optimization jointly selects association and power variables to maximize throughput while accounting for layer-dependent handover penalties.
- The formulation incorporates the handover penalty defined earlier into the optimization objective.
- The constraints enforce minimum user rates, active-link power allocation, per-satellite power budgets, beam capacities, single-beam service, visibility, and valid variable domains.
- The problem is a mixed-integer nonlinear program because bilinear coupling between binary associations and continuous powers creates a non-convex feasible set.
III. PROPOSED FRAMEWORK
The framework decomposes the joint problem into MARL-based association and exact convex power allocation, with layer-aware observations and handover-penalized rewards closing the learning loop.
- The method decomposes the problem into a learning-based association policy followed by exact power allocation in closed loop.
- Association via MARL: MARL agents use shared weights, MAPPO, and TarMAC to choose whether each user stays on its beam or switches among top-k candidates.
- Association via MARL: Each agent observes normalized candidate-link SNRs, serving-beam information, handover timing, beam residence, achieved rate, and orbital-layer encodings.
- Power allocation: Once associations are fixed, the power-allocation subproblem is concave in the continuous powers because its relevant constraints are linear.
- Power allocation: Optimal power allocation produces rates that define rewards combining achieved throughput with layer-dependent handover penalties.
IV. NUMERICAL CASE STUDY
The numerical case study evaluates the framework and baselines using a multi-constellation simulation setup, with performance comparisons averaged over 40 episodes.
- The study assesses the proposed framework in a simulation scenario before presenting comparative results.
- The numerical case study reports simulation and training parameters for the evaluation.
- Figure 2 compares the layer-aware policy against baselines using performance results averaged over 40 episodes.
A. Experimental Setup
The evaluation uses realistic multi-constellation data and frozen-policy episodes to compare learned and heuristic policies across throughput, handovers, reward, and orbital-layer utilization.
- Scenario: The scenario uses real TLE data for 30 users near Nairobi, Kenya, with dynamic LEO visibility generated through checkpoint-based anticipation.The LEO pool starts with 20 visible satellites and adds 10 at each of 40 look-ahead checkpoints.
- Baselines: Three baselines are evaluated: stay, greedy highest-SNR switching, and the learned single-layer TarMAC-LEO policy.Stay avoids switching while visibility persists; greedy ignores handover cost when selecting the highest-SNR satellite.
- Metrics: The evaluation reports mean episode reward, throughput, handovers, and the fraction of associations served by each orbital layer.These metrics are averaged over 40 evaluation episodes with frozen policies and no further learning.
- Training evaluation: Figure 3 tracks mean reward per episode during training for the proposed policy and the LEO-only baseline, with shading showing across-episode standard deviation.The figure compares training reward trajectories rather than frozen-policy evaluation outcomes.
B. Results and Discussion
Across 40 frozen-policy episodes, the proposed policy achieves a favorable throughput–handover trade-off by retaining most of greedy's rate while substantially reducing handovers and using higher orbital layers.
- Throughput–handover trade-off: 7.55 Mbps per time slot delivers about 92% of greedy's 8.2 Mbps rate while reducing handovers from 28.5 to 6.2 per time slot.The proposed policy occupies the favorable region of the throughput–handover comparison.
- Throughput–handover trade-off: About 14% higher throughput than stay is achieved, with 6.6 Mbps per time slot versus stay's 3.3 handovers per time slot.The comparison is against the conservative stay heuristic.
- Learned-policy comparison: Against TarMAC-LEO, the proposed policy attains slightly higher rate and fewer handovers: 7.55 Mbps and 6.2 handovers versus 7.48 Mbps and 6.7 handovers.TarMAC-LEO is restricted to the LEO layer, whereas the proposed policy uses all three orbital layers.
- Orbital-layer utilization: The proposed policy serves 95% of associations through LEO, 1% through MEO, and 4% through GEO, unlike TarMAC-LEO's 100% LEO allocation.The higher-layer associations constitute the reported offloading behavior.
V. CONCLUSION
The paper jointly addresses association, power allocation, and handover management across LEO, MEO, and GEO constellations using a decomposed MARL and convex optimization framework.
- Problem and method: The joint problem is formulated as a mixed-integer nonlinear program and decomposed into a layer-aware MARL association policy plus a convex power-allocation subproblem.The power allocation is solved in closed loop after the association decision.
- Results: In realistic Nairobi experiments using real TLE data, the proposed solution achieves about 92% of greedy throughput while reducing handovers more than four-fold.It also exceeds the conservative stay heuristic by about 14% in throughput.
- Multi-orbit behavior: Compared with an equivalent LEO-only MARL policy, the proposed method delivers slightly higher throughput and fewer handovers by offloading traffic to MEO and GEO.The conclusion identifies exploiting all three orbital layers as beneficial within the evaluated setting.
- Future work: Future work will scale the framework to larger user populations in MIMO systems and incorporate inter-satellite link constraints.These extensions define the stated scope boundary for subsequent development.