Source-linked AI summary

Distributed Physical Layer Authentication and Collaborative RSMA in Non-Terrestrial Networks via Graph Reinforcement Learning

Parsa Rajabi, Mohammad Mirzaee, Mohammad Reza Abedi, Nader Mokari, Paeiz Azmi

arXiv:2609.09475v1eess.SPcs.AIeess.SY

TL;DR

Existing NTN physical-layer authentication often lacks distributed multi-anchor verification, joint authentication-transmission design, and tag-privacy protection. SAFA-MZ integrates distributed authentication with collaborative multi-layer RSMA, AN, GDP, semantic grouping, and joint optimization, solved by RCEM and GA2C; simulations report average SSE gains up to 135% over single-connect transmission and 21% over transmission without AN.

  • Problem

    Existing NTN authentication designs often lack collaborative transmission integration and overlook authentication, privacy, and adversarial robustness in multi-HAPS systems.

  • Method

    SAFA-MZ jointly designs distributed group-based authentication and multi-layer collaborative RSMA with GDP-protected tags, AN-assisted secrecy, semantic grouping, and optimized HAPS placement, association, and power allocation.

  • Results

    Average SSE improves by up to 135% over single-connect transmission and 21% over the scheme without AN, while GA2C scales linearly with user count compared with RCEM's quadratic growth.

  • Takeaways & Limitations

    SAFA-MZ provides a distributed, secrecy- and fairness-aware framework for jointly authenticating users and transmitting data in multi-HAPS NTN systems.

  • Takeaways & Limitations

    Longer operation weakens GDP privacy guarantees, creating a trade-off that requires adjusting T, ϵ, and δ.

Abstract

from arXiv · show

Existing physical-layer authentication (PLA) schemes for non-terrestrial networks (NTNs) often rely on single-anchor verification, lack joint authentication-transmission design, and ignore tag privacy leakage under eavesdropping. In this paper, we consider passive, location-aware, static eavesdroppers without access to legitimate channel state information (CSI). Under this threat model, we propose secure adaptive federated authentication for multi-zone NTN systems (SAFA-MZ) that maximizes secrecy spectral efficiency (SSE) while ensuring authentication reliability, power limits, and coverage constraints. The main idea is to embed group-level authentication tags into a collaborative multi-layer rate-splitting multiple access (RSMA) transmission structure. Private and common signals are jointly beamformed, artificial noise (AN) is used to reduce information leakage, and group differential privacy (GDP) protects tag information against inference attacks. In addition, users are grouped by semantic priority to allocate SSE based on information importance. We formulate a joint SSE maximization problem under authentication reliability and probabilistic secrecy constraints, optimizing high-altitude platform station (HAPS) placement, user association, and RSMA power allocation. The resulting problem is solved using a repair-based cross-entropy method (RCEM) and a graph-aware advantage actor-critic algorithm (GA2C). RCEM scales quadratically with the number of users, while GA2C scales linearly and achieves scalable, low-latency inference. Simulation results under both colluding and non-colluding eavesdroppers show that the proposed method improves average SSE by up to 135% over single-connect transmission and 21% over the scheme without AN. These results confirm SAFA-MZ offers a scalable and secure solution for dynamic NTN environments.

I. INTRODUCTION

The paper identifies gaps in distributed authentication and secure collaborative transmission for multi-HAPS NTNs, then proposes SAFA-MZ to jointly address authentication, privacy, secrecy, and scalable resource allocation.

  • Research gap: Multi-HAPS NTNs require distributed multi-point authentication, but existing DPLA designs typically do not address collaborative transmission across multiple transmitters.The paper frames joint transmission and authentication as underexplored in multi-HAPS NTN systems.
  • Proposed framework: Collaborative multi-layer RSMA transmits group-common authentication tags and private user streams from multiple HAPSs, improving robustness, coverage, and SSE.Semantic grouping allocates SSE according to information importance while preserving private streams.
  • Privacy and secrecy: GDP protects group authentication tags against eavesdropping, while artificial noise is injected into legitimate-user channel null spaces to reduce information leakage.The framework treats tag privacy and physical-layer secrecy as joint design objectives.
  • Optimization problem: SAFA-MZ jointly considers secrecy, fairness, HAPS placement, user association, power allocation, and authentication constraints under colluding and non-colluding eavesdropper models.The resulting objective maximizes SSE subject to reliability, coverage, and probabilistic secrecy requirements.
  • Solution methods: RCEM provides a derivative-free evolutionary solution, while GA2C uses graph-aware reinforcement learning for scalable optimization in changing NTN topologies.Graph modeling permits node addition or removal without redesigning the network.
  • System model: The system model contains multiple HAPSs, single-antenna ground users, passive eavesdroppers, semantic user groups, and private, common, and artificial-noise signal components.HAPS placement and user association determine the served groups and transmission structure.

B. Transmit Model and Association

The transmit model supports exclusive or multi-connect HAPS association and superimposes private, common, and artificial-noise components. Group common symbols carry authentication and data information, with GDP noise obfuscating tags while preserving data utility.

  • Users may associate with exactly one HAPS in exclusive mode or multiple HAPSs in multi-connect mode.
  • Each HAPS transmit signal combines private streams, group-common streams, and artificial noise through power-splitting coefficients.
  • Common symbols carry both data and authentication tags, while injected GDP noise obfuscates the tag component.
  • Repeated differential-privacy releases weaken privacy guarantees because cumulative privacy loss grows with the number of slots T.
  • Null-space AN is assumed not to interfere with legitimate users under perfect CSI, although channel-estimation errors can create residual leakage.
  • The model evaluates secrecy against both non-colluding and colluding eavesdroppers using their respective achievable secrecy rates.

C. Fairness in Authentication

The framework measures user fairness with Jain’s fairness index to assess equality in achievable rates. The index equals one under perfect equality and approaches 1/|U| under extreme unfairness.

  • Jain’s fairness index measures fairness across users using their per-user achievable rates.
  • J = 1 indicates perfect equality, whereas J →1/|U| indicates extreme unfairness.

III. PROBLEM FORMULATION

The problem formulation jointly regularizes HAPS placement and group-tag cohesion while accommodating hard or soft association. It discourages vulnerable tag concentration and supports topology-aware layouts based on inter-HAPS coupling.

  • III. PROBLEM FORMULATION: The objective maximizes a secrecy- and fairness-aware network utility.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: The group-cohesion regularizer encourages users in the same group to receive signals from similar HAPSs.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: A graph-regularized layout term keeps strongly coupled HAPSs close together.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: Inter-HAPS affinity is treated as a known scalar from large-scale channel statistics and may reflect backhaul, interference, or spatial proximity.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: Hard association supports single-connect and multi-connect modes, while soft association replaces binary indicators with weights πu,k ∈[0, 1].
  • A. Graph-Regularized HAPS Placement and Group Cohesion: A small Jtag indicates tag concentration on few HAPSs, which can leave the whole group vulnerable to an attack on one HAPS.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: Soft-association cohesion quantities are computed from πu,k, while hard-association quantities use the corresponding binary indicators.
  • A. Graph-Regularized HAPS Placement and Group Cohesion: The simulation uses the defined pg,k construction and normalization for the cohesion formulation.

B. Secrecy- and Fairness-Aware Objective

The formulation accounts for served-user coverage, HAPS mobility and separation, secrecy reliability under two eavesdropper models, and weighted common-rate priorities. These constraints support a secrecy- and fairness-aware network objective.

  • The optimization defines served users because limited radio resources may leave some users unserved in a time slot.
  • HAPS velocity and pairwise separation are constrained by vmax and dmin, respectively.
  • The formulation includes a minimum required coverage ratio ξreq ∈(0, 1].
  • The objective combines secrecy and fairness considerations through weighting coefficients λtag and ηJ.
  • Probabilistic secrecy constraints impose reliability requirements for colluding and non-colluding eavesdroppers.
  • The total common-signal rate is weighted by group priorities, allowing higher weights for applications such as emergency communications.

IV. SAFA-MZ FRAMEWORK AND ALGORITHMS

The section addresses the non-convex mixed-integer optimization through derivative-free RCEM and graph-enhanced GA2C, while feasibility repairs and penalties enforce operational and secrecy constraints. RCEM samples and repairs static decisions, whereas its snapshot-specific operation limits generalization across topologies.

  • Solution strategy: RCEM and GA2C address the non-convex mixed-integer problem with stochastic secrecy constraints that make gradient-based methods inapplicable.RCEM is derivative-free, while GA2C uses graph reinforcement learning with a graph autoencoder; both reduce decision complexity using fixed precoder rules.
  • Secrecy constraints: Monte Carlo channel realizations approximate colluding and non-colluding secrecy-constraint violation probabilities for candidate decisions.The approximations use Nmc independent channel realizations and impose penalties when probabilistic or other constraints are violated.
  • RCEM: RCEM samples decision vectors, maps them to feasible tuples, and iteratively updates a Gaussian distribution toward elite solutions.Continuous variables are clipped, association logits are binarized, and the distribution is updated using samples with the smallest objective values.
  • Feasibility handling: Capacity, coverage, and rate repairs project candidate solutions onto key feasibility constraints before objective evaluation.Remaining constraints, including Monte Carlo secrecy approximations, are handled through penalties that discourage infeasible samples from selection as elites.
  • Limitation: RCEM optimizes a static decision vector, so it must be rerun for each snapshot and does not generalize across topologies.This limitation contrasts with the topology-aware learning approach developed for GA2C.

B. GA2C

GA2C combines actor-critic learning with graph-autoencoder embeddings to represent NTN topology and produce permutation-aware joint resource-allocation decisions. The representation is designed to improve stability, convergence, and generalization across network sizes and configurations.

  • Topology-aware learning: GA2C uses a graph autoencoder to encode HAPS interactions into compact, permutation-invariant features for topology-aware learning.The graph representation addresses inefficiency and sensitivity to node ordering when actor-critic methods operate on raw large-scale NTN features.
  • Joint decisions: GA2C maps each NTN snapshot to joint HAPS placement, user association, and RSMA power-allocation decisions.The policy handles continuous placement and power variables alongside association decisions.
  • State representation: User locations, tag identities, and HAPS-user large-scale channel features are separately encoded to support user-dependent associations while preserving HAPS-node permutation invariance.The graph state represents HAPS interactions, while user-side information remains available to the actor for association decisions.

1) Graph Construction and Normalization:

The graph module normalizes HAPS interaction weights, propagates node features through GCN layers, and mean-pools node embeddings into a graph-level representation. This representation supports actor-critic decisions while auxiliary reconstruction losses preserve topology and feature information.

  • Graph encoding: The graph autoencoder operates on normalized HAPS adjacency and node features to produce topology-aware node embeddings.The encoder uses GCN propagation, with normalized adjacency serving as the propagation operator and learnable weights and nonlinear activation transforming node features.
  • Graph pooling: Mean pooling converts node embeddings into a graph-level embedding used to represent the HAPS interaction topology.The graph-level representation is formed from the node embeddings after graph encoding.
  • Auxiliary reconstruction: The decoder reconstructs normalized adjacency and pooled features, providing auxiliary objectives for topology and feature reconstruction.Adjacency reconstruction uses sigmoid outputs matched to normalized edge weights, while the auxiliary loss is combined with the A2C loss.
  • Actor-critic control: The A2C actor uses Gaussian policies for continuous motion and power actions and Bernoulli policies for binary associations, while the critic estimates state value.Rewards include penalties for rate, coverage, secrecy, capacity, and priority-related constraints; temporal-difference targets provide the advantage signal.

V. COMPUTATIONAL COMPLEXITY ANALYSIS

The complexity analysis contrasts snapshot-based RCEM with learned GA2C inference. RCEM grows quadratically with user count, whereas GA2C scales linearly in the reported comparison and its graph size is bounded by inference and memory budgets.

  • RCEM complexity: RCEM repeatedly samples decision vectors, repairs feasibility, and evaluates Monte Carlo secrecy across Niter iterations for each static instance.Its runtime includes sampling, repair, and physical-layer evaluation costs involving users, serving HAPSs, eavesdroppers, and Monte Carlo realizations.
  • GA2C complexity: GA2C separates offline training from online inference, with inference requiring only a forward pass after training.Training includes graph message passing and Monte Carlo secrecy evaluation, while total training cost scales with interaction steps.
  • Complexity comparison: RCEM exhibits quadratic growth with respect to |U|, whereas GA2C scales linearly and is more suitable for large-scale and dynamic NTN deployments.The comparison considers identical physical quantities, with Monte Carlo secrecy enforcement as the dominant cost.
  • Inference bound: For fixed network and model parameters, dense-graph support scales as O(√Binf) with inference budget, while sparse-graph support scales linearly in Binf.The bound follows from graph-edge counts and applies under fixed GNN depth, hidden dimensions, and user count.
  • Memory bound: Graph size is jointly limited by available inference budget and memory capacity for storing edge weights.The memory constraint is expressed through the product of per-edge memory footprint and the number of graph edges.

B. An Embedding-Capacity Bound on the Number of HAPSs

The paper bounds the number of HAPSs that can be uniquely represented by graph embeddings using embedding precision, dimension, and quantized edge information. The resulting bound is combined with other constraints to determine the allowable HAPS count.

  • The graph embedding can represent at most 2^(n_bits d_z) distinct HAPS-node embeddings.The bound depends on the effective bits per coordinate and embedding dimension.
  • Quantized interaction graphs contribute at most 2^(q|E_g|) distinct adjacency patterns.Each undirected edge weight is stored using q bits.
  • Unique worst-case representation requires q 2|K|(|K|−1) ≤ n_bits d_z.This condition links the number of HAPSs to graph-edge precision and embedding capacity.
  • The maximum uniquely representable |K| is obtained from the derived bound for fixed d_z, n_bits, and q.The passage states that the maximum |K| is given by the resulting condition.
  • The final allowable |K| is bounded by the minimum of the three preceding bounds.This combines the separate embedding-capacity constraints into one HAPS-count limit.

B. Training Dynamics: Evolution versus Learning

GA2C learns a reusable topology-aware policy, while RCEM searches each snapshot directly; results show GA2C generally scales better, especially in larger networks, while privacy and secrecy constraints shape SSE and authentication reliability.

  • Training Dynamics: GA2C’s return increases over episodes, whereas RCEM converges quickly in early iterations; GA2C enables reusable fast inference after offline training.RCEM searches independently for each snapshot, while GA2C amortizes optimization through a learned policy.
  • SSE and Fairness: GA2C achieves higher HAPS-level and minimum group-level SSE than RCEM through topology-aware decisions and more uniform protection.Non-colluding eavesdroppers yield higher SSE than colluding eavesdroppers.
  • SSE and Fairness: Fairness generally increases with |K|; GA2C performs better in larger networks, while RCEM can be fairer in smaller networks through extensive evolutionary search.The comparison uses Jain’s fairness index and evaluates identical test episodes.
  • Method Comparison: AN improves SSE by at least 21%, while collaborative RSMA provides at least a 135% gain over single-connect transmission.The gains are reported from average SSE comparisons over 10 identical test episodes.
  • Scalability Analysis: As |U| grows, SSE decreases for both methods, but GA2C remains stronger in larger networks; increasing |G| or |E| also reduces SSE.Increasing |K| improves spatial diversity, interference coordination, and multi-connect gains.
  • Authentication Reliability: Authentication probability increases with privacy budget ϵ but decreases with group size Ng, demonstrating a GDP privacy–reliability trade-off.Larger groups require higher privacy budgets, while despreading gain from increasing Ls does not remove GDP distortion.
  • Overall Findings: The proposed framework combines multi-direction authentication, semantic group tags, collaborative RSMA, GDP, and graph-aware RL while jointly optimizing placement, association, and power.GA2C scales linearly with user count and provides low-latency online decisions, whereas RCEM scales quadratically.

APPENDIX A PROOFS

The appendix proves privacy guarantees for Gaussian perturbation, group-level protection, and repeated releases, while also establishing graph-encoder permutation properties and authentication reliability expressions.

  • Differential privacy proofs: The Gaussian perturbation mechanism satisfies (ϵ, δ)-DP by bounding privacy loss with a Gaussian tail argument.The proof models privacy loss as an affine Gaussian variable and applies a standard tail bound.
  • Differential privacy proofs: Changing at most N′ groups yields a group-level privacy bound by chaining single-group DP guarantees.Sequential composition multiplies the privacy factor and sums the δ′ terms across the chain.
  • Differential privacy proofs: Privacy loss over T independent DP releases is bounded using a Chernoff bound with slack δ̄.
  • Graph structure: The Laplacian objective penalizes large distances between strongly connected HAPS because link rate increases the corresponding graph weights.The proof expands L = D−W and uses the weight definition to connect the objective to HAPS distances.
  • Graph structure: The graph encoder is permutation-equivariant, while mean pooling produces permutation-invariant representations.The proof uses permutation transformations through the encoder and the identity 1^TΠ = 1^T for mean pooling.
  • Authentication reliability: Authentication reliability is expressed through a decision statistic combining the group tag with differential-privacy and channel noise, yielding Pc = (1−Q(1/σtot))^2.The expression follows unit-energy QPSK signaling, MRC, and despreading assumptions.
Loading 2609.09475v1…