Source-linked AI summary
Security-Aware Pinching-Antenna Systems (PASS): Physical-Layer Security Transmission
Zhaoming Hu, Xiaochen Nie, Ruikang Zhong, Haochen Li, Dengao Li, Xidong Mu
TL;DR
The paper addresses heterogeneous confidentiality in PASS, where security modes can change receiver roles for each information stream. It formulates joint beamforming, artificial-noise, and PA-position control, then develops HSPPO and MRHA-DPO for complementary efficiency and expressiveness. Simulations report secrecy and control gains from movable PAs, with MRHA-DPO achieving the best overall performance and HSPPO retaining low policy complexity.
Problem
Existing PASS and PLS designs do not adequately support service-dependent authorization in which receiver roles and interceptor sets vary across information streams.
Method
The paper establishes three role-dependent PASS security modes and jointly controls beamforming, artificial noise, and PA positions using HSPPO and MRHA-DPO.
Results
Movable PAs consistently outperform fixed-PA PASS and conventional MIMO in secrecy rate, while MRHA-DPO delivers the best overall performance and HSPPO provides stable low-complexity control.
Takeaways & Limitations
A common PASS platform can support service-dependent confidentiality requirements while balancing online efficiency and control capability.
Abstract
from arXiv · showhide
This paper investigates heterogeneous secure multi-user transmission in pinching-antenna systems (PASS), where dynamically adjustable pinching antennas reshape both guided-wave and free-space propagation to improve communication and confidentiality performance. Unlike conventional physical-layer security designs that represent different security requirements merely through weights or thresholds, heterogeneous services may change the logical role of each receiver for each information stream. To address this issue, we establish a unified role-dependent PASS transmission framework comprising low-, medium-, and high-security modes. These modes respectively maximize the minimum legitimate-user rate, protect confidential streams against external eavesdroppers, and further prevent non-target legitimate users from intercepting unauthorized information. The resulting joint optimization of information beamforming, artificial noise, and pinching-antenna positions is formulated as a long-horizon continuous-control problem. Two learning-based controllers are then developed to provide complementary complexity-performance tradeoffs. First, heterogeneous security-aware proximal policy optimization (HSPPO) directly transforms mode-specific rate and secrecy violations into normalized smooth feedback embedded in the proximal-policy-optimization advantage, enabling lightweight and violation-sensitive control. Second, multi-relational hierarchy-aware diffusion policy optimization (MRHA-DPO) combines a PASS-aware multi-relational graph encoder, a graph-conditioned hierarchical velocity network, and exact-inversion DPO training to achieve topology-aware and expressive control. The proposed framework enables a common PASS platform to flexibly support service-dependent confidentiality requirements while balancing online efficiency and control capability.
I. INTRODUCTION
Heterogeneous services require security-aware PASS transmission because security requirements can change receiver roles and stream-specific interceptor sets. The paper motivates a unified framework that adapts propagation, beamforming, and artificial noise to these differing requirements.
- Heterogeneous services require confidentiality levels ranging from rate-focused delivery to protection against external eavesdroppers and unauthorized legitimate users.
- Conventional PLS schemes commonly assume fixed threat models and uniform secrecy objectives across services.
- PASS introduces reconfigurable spatial degrees of freedom that can reshape propagation when conventional channel disparities are limited.
- In heterogeneous services, the same legitimate receiver may be authorized for one stream but act as an interceptor for another.
- Unified control is challenging because security modes produce unequal training feedback while PA–user, receiver-role, and inter-PA relationships couple state and actions.
C. Contributions
The paper establishes a PASS secure-transmission framework and develops two complementary learning-based controllers for jointly adapting beamforming, artificial noise, and PA configurations. The system model uses movable PAs on parallel waveguides to serve legitimate users amid external eavesdroppers.
- The framework jointly optimizes security-aware transmit beamforming and PA positioning in a long-horizon control problem that accounts for PA repositioning.
- HSPPO provides lightweight, stable cross-mode control by incorporating mode-specific performance and constraint violations into unified-policy updates.
- MRHA-DPO models PASS physical and security relations and the coupled beamforming, artificial-noise, and PA-positioning actions through hierarchical diffusion-based generation.
- The PASS system uses parallel dielectric waveguides with movable PAs to serve legitimate users in the presence of external eavesdroppers.
- Confidential symbols are precoded at the base station, injected into waveguides, and radiated into free space by activated PAs.
- The composite channel captures the superposition of signals radiated from all activated PAs under a line-of-sight propagation model.
B. Security-Aware PASS Transmission Model
The transmission model distinguishes legitimate decoding from target-dependent eavesdropping and uses three security levels to define who may intercept each stream. These levels progressively broaden the interceptor set from none, to Eves, to all unintended receivers.
- Legitimate users decode their own intended signals, whereas each Eve evaluates interception capability for a specific Bob’s stream.
- Achievable rates and eavesdropping rates are defined from receiver-specific SINRs for intended decoding and target-dependent interception.
- Low-security services treat other Bobs and Eves as non-adversarial, medium-security services protect against Eves, and high-security services protect against both groups.
1) Low-Security Transmission (LST):
The three PASS transmission modes assign distinct rate and secrecy objectives to the same physical architecture. They progress from worst-user rate maximization to external-eavesdropper protection and then full unintended-receiver isolation.
- 1) Low-Security Transmission (LST):: LST maximizes the worst-user achievable rate without explicit secrecy constraints or treating other receivers as potential interceptors.
- 2) Medium-Security Transmission (MST):: MST preserves intended-user rates while suppressing leakage to external Eves through a worst-user secrecy-rate objective.
- 3) High-Security Transmission (HST):: HST extends the interceptor set to every receiver except the intended Bob, requiring stronger stream-level isolation.
- The unified optimization selects RLST, RMST, or RHST according to the required security level while jointly optimizing beamforming and PA positions.
- The controller adapts to time-varying channels and security requirements while PA relocation couples current decisions to previous configurations.
- HSPPO reweights heterogeneous samples using normalized violation severity and consistency information while retaining PPO’s clipped, low-complexity updates.
A. Heterogeneous Security-Aware MDP Formulation
The framework formulates heterogeneous security-aware PASS transmission as a long-horizon continuous-control problem. HSPPO uses mode-dependent performance violations and reward feedback to train a shared policy while accounting for beamforming, artificial noise, PA positioning, and reconfiguration costs.
- State Space: HSPPO models PASS transmission as a DRL process in which each state includes propagation conditions, PA history, user geometry, CSI, and the active security requirement.The agent observes the state, samples a policy action, transitions through the environment, and receives an instantaneous reward.
- Action Space: The continuous action jointly controls the transmit beamforming matrix, including artificial-noise components, and the PA position set.These variables determine the effective channel and thereby govern received SINR and rate or secrecy performance.
- Reward Function: The reward combines the active mode’s communication or secrecy metric with a penalty for PA displacement between consecutive time slots.The movement coefficient weights reconfiguration cost, linking transmission performance to positioning overhead.
- Training Process: Each training transition records the active security mode and its corresponding mode-dependent security performance alongside the standard MDP variables.This preserves heterogeneous security context in the on-policy training buffer.
- Violation Modeling: HSPPO measures each sample’s mode-dependent performance margin against its required threshold and converts the margin into a smooth non-negative violation cost.A nonpositive margin is counted as a violation, while softplus provides differentiable severity feedback.
- Violation Modeling: The violation costs are standardized within each mini-batch so heterogeneous modes produce a normalized signal that can be combined with the reward advantage.This prevents larger numerical violation ranges from dominating policy updates.
- Security-Aware Advantage: An adaptive security coefficient increases when the observed violation rate exceeds its target and decreases otherwise.The coefficient adjusts the strength of violation correction according to the current mini-batch security status.
- Security-Aware Advantage: The security-aware advantage reduces or reverses high-reward actions with severe violations, while consistency weighting downweights samples whose reward and security feedback disagree.Compared with conventional PPO, these modifications retain clipped updates and a low-complexity policy structure while using heterogeneous samples more effectively.
IV. MULTI-RELATIONAL HIERARCHY-AWARE DIFFUSION POLICY OPTIMIZATION DESIGN
MRHA-DPO addresses the coupled structure of secure PASS control with a multi-relational graph representation and hierarchy-aware diffusion policy learning. It is designed to provide stronger representation and joint-action modeling than the lightweight HSPPO controller.
- Design Motivation: MRHA-DPO integrates PASS-specific state representation and structured action generation into diffusion-based policy learning.The design is motivated by interactions among physical deployment, wireless propagation, security relationships, beamforming, artificial noise, and PA positions.
- Core Architecture: Its multi-relational graph characterizes heterogeneous interactions among PAs, legitimate users, and potential eavesdroppers for structured policy-state embeddings.A hierarchy-aware action mechanism models the coupled evolution of heterogeneous control variables.
A. Multi-Relational PASS State Representation
The PASS state is represented as a multi-relational heterogeneous graph whose nodes encode PAs, legitimate receivers, and eavesdroppers, while relation-aware attention captures physical and security dependencies. Role-aware pooling then produces a structured graph-level embedding for diffusion policy learning.
- Graph Construction: The instantaneous PASS state is modeled as a directed heterogeneous graph because receiver quality, security roles, and PA configurations are relationally coupled.The graph contains PAs, legitimate receivers, and external eavesdroppers as distinct entity types.
- Node Features: Each node feature combines spatial coordinates, one-hot node category, channel state, and security-mode information.Channel features differ by role: PA nodes summarize waveguide links toward Bob and Eve sets, while receiver nodes describe channels from the PASS aperture.
- Relation Types: Directed edges carry relation labels distinguishing PA–Bob, PA–Eve, Bob–Eve, and same-role PA, Bob, and Eve interactions.These relations encode propagation dependencies together with security- and interference-related receiver interactions.
- Relation-Aware Attention: The HGNN encoder uses relation-specific query, key, and value projections to aggregate heterogeneous neighboring information for each target node.A fixed incoming-edge ordering preserves relation semantics during attention computation.
- Relation-Aware Attention: Attention weights incorporate relational compatibility, geometric and propagation conditions, edge semantics, and active security states.The geometric and edge priors provide physical and semantic information beyond feature compatibility.
- Feature Fusion: A relation-independent self representation is fused with aggregated relational context through an MLP, and stacked HGNN layers capture higher-order dependencies.This preserves target-node information while adding context from neighboring entities.
- Graph-Level Embedding: Role-aware pooling separately aggregates PA, Bob, and Eve representations and combines them with global average-pooled information.The resulting embedding preserves role-specific characteristics while capturing global PASS topology for diffusion policy learning.
B. Graph-Conditioned Hierarchy-Aware Action Generator
The hierarchy-aware action generator separates component-specific modeling from cross-component coordination for heterogeneous PASS control. A graph-conditioned velocity field then produces coordinated action evolution while preserving the physical distinctions among control variables.
- Motivation: The generator addresses the difficulty of jointly controlling transmission and PA configuration variables that operate in different domains.Flattened mappings may obscure structural differences and hinder coordination of coupled decisions.
- Hierarchical action generation: Dedicated lower-level branches model different action components according to their structures and constraints.This component-specific modeling preserves variable-specific characteristics.
- Hierarchical action generation: A higher-level module jointly coordinates intermediate representations to capture coupling among heterogeneous decisions.The design maintains global consistency across the joint PASS control space.
- Graph-conditioned velocity field: Three component-specific heads generate preliminary velocity components for signal beamforming, AN beamforming, and PA positioning.These components retain the distinct physical structures of their corresponding action variables before global coordination.
- Graph-conditioned velocity field: The preliminary components are jointly processed to account for their coupled effects on transmission rate, information leakage, and PASS geometry.Their concatenation defines the complete conditional velocity field used by the diffusion policy.
- Implementation: Real and imaginary parts of the information and AN beamformers are represented as real-valued coordinates and reconstructed by the action decoder.
C. Multi-Relational Hierarchy-Aware Diffusion Policy Optimization and Training Process
MRHA-DPO uses a reversible diffusion policy with structured PASS state conditioning and hierarchical action generation. Exact trajectory inversion supplies differentiable policy likelihoods for PPO-style optimization of long-horizon control.
- Policy construction: MRHA-DPO integrates structured PASS state representation and hierarchy-aware action generation into diffusion-based policy learning.The design targets the highly coupled PASS action space.
- Policy construction: The DPO backbone uses two coupled augmented-action variables that evolve over diffusion steps to produce a terminal augmented action.The velocity field is conditioned on the graph representation and active security mode.
- Policy construction: The physical PASS action is obtained by combining the two terminal variables, preserving a unique environment action despite augmented-space exploration.
- Likelihood evaluation: The reversible diffusion mapping enables exact inversion from terminal to initial Gaussian variables and exact policy-log-likelihood evaluation.The prior p0(·) is the standard Gaussian prior.
- Training process: MRHA-DPO training comprises action sampling, exact trajectory inversion, and policy optimization using on-policy transitions.Stored transitions include the active security mode, terminal augmented variables, and reward.
- Policy optimization: The diffusion-policy probability ratio and GAE advantage define a PPO clipped actor loss, supplemented by likelihood-based entropy regularization.
- Policy optimization: A terminal compression loss limits discrepancy between augmented variables representing the same physical PASS action.The final actor objective weights exploration regularization and terminal-variable consistency.
- End-to-end training: Actor and critic updates propagate through the reversible diffusion backbone, hierarchy-aware generator, and multi-relational state encoder end to end.
V. NUMERICAL RESULTS
The numerical-results section evaluates learning behavior, system performance, ablations, and spatial energy patterns under common channel, power, and PA-feasibility conditions.
- Evaluation scope: The evaluation covers learning behavior, system performance, ablation results, and spatial energy patterns.
- Evaluation scope: All evaluated methods use the same channel model, power constraint, and PA-feasibility constraints unless otherwise stated.
- Evaluation scope: The experiments compare the proposed framework across multiple aspects rather than reporting only a single performance measure.
A. Simulation Setup
The simulations model a dynamic PASS service region with moving Bobs and Eves and randomly varying security demands. The setup uses four waveguides with four PAs each and trains DRL policies across heterogeneous modes.
- System configuration: The service region measures 20 × 20 m^2 and contains four Bobs and two Eves on the ground plane.
- System configuration: The PASS uses N = 4 parallel dielectric waveguides with M = 4 PAs per waveguide at height H = 3 m.
- System configuration: The carrier frequency is fc = 28 GHz, the effective refractive index is neff = 1.4, and minimum inter-PA spacing is Δmin = λ/2.
- Dynamic environment: Bobs and Eves change locations each time slot, while security demand is randomly selected from LST, MST, and HST.A single policy is trained under heterogeneous and time-varying security requirements.
- Training configuration: All DRL algorithms are trained for 2000 episodes.
- Training evaluation: Fig. 4 compares training rewards under LST, MST, and HST security demands.
B. Convergence and Learning Performance
Across LST, MST, and HST, policy-based DRL methods generally stabilize the long-horizon PASS control problem, with HSPPO and MRHA-DPO outperforming conventional PPO. Movable-PA PASS also consistently exceeds fixed-PA PASS and conventional MIMO, although its advantage changes with transmit power and service-region size.
- Convergence and Learning Performance: PPO, HSPPO, and MRHA-DPO converge stably across LST, MST, and HST, whereas TD3 remains fluctuating and fails to stabilize.The paper attributes TD3’s instability to off-policy updates interacting with a strongly coupled action space and non-smooth security objectives.
- Convergence and Learning Performance: HSPPO and MRHA-DPO consistently outperform conventional PPO across all three security modes.HSPPO uses mode-dependent violation information with a lightweight Gaussian policy, while MRHA-DPO uses multi-relational PASS representations and diffusion-based action generation.
- Impact of Transmit Power: Movable-PA PASS achieves the highest secrecy-rate performance, followed by fixed-PA PASS and conventional MIMO, with gains increasing as transmit power rises.The improvement is attributed to additional spatial degrees of freedom and adaptation of propagation geometry to instantaneous user distributions.
- Impact of Service-Region Size: All schemes lose rate as the service region expands, but movable-PA PASS remains best, followed by fixed-PA PASS and conventional MIMO.Larger regions increase propagation loss; the gap between the two PASS schemes narrows as distributed users favor coverage-balanced PA placements.
- Impact of Service-Region Size: Under LST, the movable-PA PASS advantage over MIMO decreases from approximately 1.05 to 0.50 bit/s/Hz, while relative gain increases from approximately 70% to 119% as side length grows from 10 to 50 m.The paper reports that MIMO degrades more rapidly with increasing service-region size; performance ordering is LST, MST, then HST.
- Ablation Study of the Proposed Algorithms: Jointly enabling HSPPO’s advantage correction and security-consistency weighting yields the best minimum rate across LST, MST, and HST, especially under MST and HST.For MRHA-DPO, combining the multi-relational graph encoder and diffusion policy likewise achieves the highest performance across all three modes.