Source-linked AI summary
Deep Reinforcement Learning Based Intelligent Reflecting Surface for Secure Wireless Communications
Helin Yang, Zehui Xiong, Jun Zhao, Dusit Niyato, Liang Xiao, Qingqing Wu
TL;DR
The paper addresses robust secrecy-rate maximization for multiple legitimate users and eavesdroppers in IRS-assisted systems with dynamic channels and QoS requirements. It proposes DRL for joint BS and IRS beamforming, enhanced by PDS and PER, and reports improved secrecy rate and QoS satisfaction over competing approaches, including under channel uncertainty.
Problem
Existing approaches do not jointly handle multiple users, secure communication, imperfect CSI, and dynamic channels, while the resulting beamforming problem is non-convex.
Method
A DRL secure-beamforming framework jointly optimizes BS and IRS beamforming using an MDP, modified PDS learning, and PER.
Results
The proposed approach improves secrecy rate and QoS satisfaction probability over existing approaches and remains favorable as CSI quality decreases.
Takeaways & Limitations
Deep PDS-PER learning provides an effective secure-beamforming approach for dynamic IRS-aided communications with multiple users and eavesdroppers.
Abstract
from arXiv · showhide
In this paper, we study an intelligent reflecting surface (IRS)-aided wireless secure communication system for physical layer security, where an IRS is deployed to adjust its surface reflecting elements to guarantee secure communication of multiple legitimate users in the presence of multiple eavesdroppers. Aiming to improve the system secrecy rate, a design problem for jointly optimizing the base station (BS)'s beamforming and the IRS's reflecting beamforming is formulated given the different quality of service (QoS) requirements and time-varying channel condition. As the system is highly dynamic and complex, and it is challenging to address the non-convex optimization problem, a novel deep reinforcement learning (DRL)-based secure beamforming approach is firstly proposed to achieve the optimal beamforming policy against eavesdroppers in dynamic environments. Furthermore, post-decision state (PDS) and prioritized experience replay (PER) schemes are utilized to enhance the learning efficiency and secrecy performance. Specifically, PDS is capable of tracing the environment dynamic characteristics and adjust the beamforming policy accordingly. Simulation results demonstrate that the proposed deep PDS-PER learning-based secure beamforming approach can significantly improve the system secrecy rate and QoS satisfaction probability in IRS-aided secure communication systems.
I. INTRODUCTION
IRS is introduced as a passive, programmable surface for improving secure wireless communication, while prior work leaves multi-user, imperfect-CSI, and dynamic settings insufficiently addressed. The paper formulates and solves this challenge with a DRL-based joint beamforming approach enhanced by PDS and PER.
- Motivation: IRS elements adapt their reflection amplitudes and phases to strengthen legitimate-user signals while suppressing eavesdropper signals.IRS is a uniform planar array of low-cost passive reflecting elements.
- Prior work: Prior IRS security studies commonly considered single legitimate users and eavesdroppers, perfect CSI, or traditional optimization methods that are less efficient for large-scale systems.Existing DRL studies also generally omitted multiple users, secure communication, and imperfect CSI together.
- Problem setting: The paper targets multiple legitimate users, multiple eavesdroppers, time-varying channels, and QoS constraints while maximizing system secrecy rate.The optimization jointly covers BS transmit beamforming and IRS reflecting beamforming.
- Proposed approach: An RL framework uses instantaneous dynamic-environment observations to optimize BS and IRS beamforming through an MDP and a QoS-aware reward function.The reward incorporates both secrecy rate and users’ QoS requirements.
A. System Model
The system contains a multi-antenna BS, an IRS, multiple legitimate users, and multiple eavesdroppers communicating over imperfect and time-varying channels. The model represents IRS reflections, BS precoding, received signals, rates, secrecy rates, and bounded channel uncertainty.
- Network architecture: A BS with N antennas serves K single-antenna legitimate users while M single-antenna eavesdroppers are present, assisted by an IRS with L reflecting elements.An IRS controller coordinates with the BS, and maximal reflection without power loss is considered.
- Beamforming model: The IRS reflection matrix Ψ uses element amplitudes χ_l and phase shifts θ_l, while BS vector v_k precodes the signal for user k.The transmitted signal is the sum of the users’ precoded symbols and obeys a maximum BS transmit-power constraint.
- Received signals: Each legitimate user receives a direct BS signal, an IRS-reflected signal, inter-user interference, and additive Gaussian noise.Eavesdroppers receive corresponding signals when attempting to intercept legitimate-user transmissions.
- Channel uncertainty: Transmission delay and user mobility make CSI outdated when the BS and IRS transmit, producing channel uncertainty and potential performance loss.The outdated-CSI coefficient ρ ranges from 0 to 1; ρ = 1 eliminates the outdated-CSI effect, whereas ρ = 0 represents no CSI.
- Channel uncertainty: Actual channel coefficients are decomposed into estimated channels and error vectors whose Euclidean norms are bounded by deterministic error-region radii.The model applies this uncertainty representation to BS-user, IRS-user, BS-eavesdropper, and IRS-eavesdropper links.
- Rate model: User rates, eavesdropper rates, and individual secrecy rates are defined under the channel-uncertainty model, with secrecy rates clipped by [z]+ = max(0, z).Each eavesdropper may attempt to eavesdrop any legitimate user’s signal.
B. Problem Formulation
The formulation jointly designs robust BS and IRS beamforming to maximize worst-case secrecy rate while satisfying secrecy-rate, data-rate, transmit-power, and IRS unit-reflection constraints. Its coupled variables and non-convex objective make direct optimization difficult.
- Optimization objective: The optimization jointly selects robust BS beamforming V and IRS reflecting beamforming Ψ from a system beamforming codebook to maximize worst-case secrecy rate.The design includes worst-case secrecy-rate and data-rate requirements.
- Constraints: Constraints enforce each user’s target secrecy rate Rmin_k and target data rate, the BS maximum transmit power Pmax, and unit-modulus IRS reflection.The secrecy-rate and data-rate constraints guarantee the users’ worst-case requirements.
- Optimization difficulty: The problem is non-convex because the objective is non-concave in either beamforming variable, the variables are coupled, and IRS unit-norm constraints apply.The formulation seeks robust performance under channel uncertainty.
III. PROBLEM TRANSFORMATION BASED ON RL
The paper reformulates IRS-aided secure beamforming as a reinforcement-learning decision problem because the original optimization is non-convex and dynamic. The BS controller observes system states, selects beamforming actions, and learns a policy maximizing long-term discounted reward while satisfying secrecy and data-rate QoS requirements.
- Dynamic channels, changing user capabilities, and service applications make single-slot optimization potentially greedy and suboptimal because it ignores historical state and long-term benefit.
- The IRS-aided communication system is modeled as the environment, and the BS central controller acts as a model-free reinforcement-learning agent.
- The state includes all-user channel information, previous secrecy and transmission rates, and QoS satisfaction levels for legitimate users and eavesdroppers.
- The action jointly selects BS user beamforming vectors and IRS reflecting coefficients, while state transitions describe movement to a new system state after execution.
- The reward maximizes long-term discounted utility, combining system secrecy rate with penalties for unmet secrecy-rate and minimum-data-rate requirements.
IV. DEEP PDS-PER LEARNING BASED SECURE BEAMFORMING
The proposed secure beamforming approach uses deep reinforcement learning to address high-dimensional, uncertain IRS optimization, combining PDS learning and PER to adapt beamforming policies in dynamic environments. Optimization occurs at the BS, and the resulting IRS matrix can be transferred offline to the IRS.
- Conventional Q-learning, policy-gradient, and DQN methods have different limitations involving continuous spaces, convergence, suboptimality, or computational tractability.
- Deep PDS-PER learning uses observed CSI, previous secrecy rate, QoS feedback, rewards, and replay-buffer history to train a model that selects BS and IRS beamforming matrices.
- The policy optimization for BS matrix V and IRS matrix Ψ is performed at the BS, after which the optimized IRS matrix can be transferred offline to adjust reflecting elements.
A. Proposed Deep PDS-PER Learning
The proposed deep PDS-PER algorithm combines post-decision states, DQN-based value learning, and prioritized replay to exploit known transition information and focus training on informative experiences. Its stated convergence guarantee holds under standard learning-rate and exploration conditions.
- PDS-learning inserts an intermediate state after action execution, using known transition probability and reward before the system reaches the next state through unknown dynamics.
- The PDS action-value function incorporates extra known transition and reward information to expand the ordinary state-action value function.
- DQN estimates the action-value function with a neural network, minimizing a loss function and updating its parameters through the learning rate and gradient.
- PER samples experiences according to absolute temporal-difference error, causing higher-priority transitions to be replayed more frequently.
- Importance-sampling weights correct the altered visitation frequencies caused by prioritized replay, with η2 controlling the correction amount.
- Theorem 1 states that deep PDS-PER learning converges to the optimal action-value function with probability 1 when the learning-rate sequence meets the specified conditions.
B. Secure Beamforming Based on Proposed Deep PDS-PER Learning
Secure beamforming is implemented through separate training and deployment stages under a BS central controller. Training uses observed states and replayed transitions to learn actions, while deployment applies the trained model and updates beamforming matrices from its selected action.
- The approach separates secure beamforming into training and implementation stages managed by a central controller at the BS.
- During training, the controller observes CSI, previous predicted secrecy rate, and transmission data rate before inputting the state vector into DQN.
- Experience tuples containing state, action, reward, PDS state, and next state are stored in replay memory, where PER selects minibatches for DQN training.
- During implementation, the trained model selects the maximum-value action from the observed system state, and the environment returns an instantaneous reward.
- Training requires a powerful computation server and can run offline at the BS, whereas implementation can run online; model updates are needed after major environmental changes.
C. Computational Complexity Analysis
The proposed deep PDS-PER algorithm adds computational overhead through PDS-learning and PER, while its DNN training can be performed offline at a centralized powerful unit. The approach is reported to achieve better performance than classical DQN despite slightly higher complexity.
- DNN computation: The DNN agent requires O(Z0Zl+PL−1 l=1 ZlZl+1) computation at each time step.Here, Z0 denotes input-layer size, Zl denotes the number of neurons in the l-th layer, and L denotes the number of training layers.
- Training deployment: DNN training has high computational complexity but can be performed offline for finite episodes at a centralized powerful unit such as the BS.Offline training separates the expensive training phase from later deployment of the learned model.
- Complexity sources: PDS-learning doubles the classical DQN complexity from O(|S|2 × |A|) to O(2|S|2 × |A|).PDS-learning uses a PDS-state set equal to the MDP-state set, producing the additional computational term.
- Complexity sources: PER adds O(log2D) complexity for updating and sampling experiences from a replay buffer of size D.The extra operations arise from prioritized experience replay’s buffer updating and sampling procedures.
- Performance trade-off: The proposed algorithm has slightly higher complexity than classical DQN but achieves better performance, as evaluated in the subsequent section.The additional complexity is attributed to the PDS-learning and PER schemes.
D. Implementation Details of DRL
The implementation uses a DQN with a fully connected multilayer perceptron to map system information to beamforming decisions. Training data are generated from sampled channels and rewards, with separate training and validation datasets and a finite experience replay buffer.
- Dataset generation: Training samples pair sampled channel vectors h with corresponding reward vectors, and h is provided as the DQN input.Beamforming matrices are selected from predefined BS and IRS codebooks during data generation.
- Training procedure: 80% of generated data are used for training and 20% for validation or testing, with 1000 training epochs and 128 mini-batches per epoch.The experience replay buffer stores 32000 of the most recent experiences for random sampling.
- DQN structure: The DQN uses a multilayer perceptron, or feedforward fully connected network, to estimate beamforming matrices from environment descriptors.The estimated decisions include both the BS beamforming matrix and the IRS reflecting beamforming matrix.
- DQN structure: The network input comprises system states such as channel samples, achievable rate, and QoS information.The hidden layers connect their neurons to outputs from the preceding layer.
- Training objective: The regression loss trains DNN outputs ˆr to approximate normalized rewards ¯r by minimizing mean-squared error.The loss is parameterized by the complete set of DNN parameters θ.
V. SIMULATION RESULTS AND ANALYSIS
Simulations evaluate learning behavior and compare secure beamforming approaches across transmit power, IRS size, and CSI accuracy. The proposed deep PDS-PER approach outperforms the baselines while maintaining stronger secrecy-rate and QoS performance under channel uncertainty.
- Simulation settings: The simulations use N = 4 BS antennas, K = 2 legitimate users, M = 2 eavesdroppers, Pmax from 15 to 40 dBm, L from 10 to 60, and ρ from 0.5 to 1.The minimum secrecy rate and minimum transmission data rate are 3 bits/s/Hz and 5 bits/s/Hz, respectively.
- Learning behavior: A learning rate around α = 0.001 is selected because excessively large learning rates produce oscillatory behavior.Average reward is evaluated across training episodes for α = {0.1, 0.01, 0.001, 0.0001}.
- Learning behavior: Training and validation losses decrease sharply during the initial epochs and become approximately horizontal after 150 epochs, with validation loss only slightly higher.The loss curves are used to assess DNN fitting and identify potential overfitting or underfitting.
- Transmit power: As Pmax increases, secrecy rate and QoS satisfaction probability increase, and all IRS-assisted approaches outperform the traditional system without IRS.The proposed approach outperforms Baseline1 and DQN because it jointly optimizes BS and IRS beamforming while using PDS-learning and PER.
- IRS size: Increasing IRS elements from L = 10 to 60 improves IRS-assisted secrecy rate and QoS satisfaction, while the no-IRS approach remains constant.The proposed approach has a growing secrecy-rate gap over Baseline 1 and DQN and is the first to attain 100% QoS satisfaction as L increases.
- CSI uncertainty: When CSI becomes more outdated as ρ decreases, all approaches lose secrecy rate and QoS performance, but the proposed approach remains comparatively robust.At ρ = 0.7, it achieves secrecy-rate and QoS satisfaction improvements of 17.21% and 8.67% over Baseline 1.
VI. CONCLUSION
The paper formulates joint BS and IRS beamforming optimization under time-varying channels and addresses it with deep PDS-PER reinforcement learning. Simulations show improved secrecy rate and QoS satisfaction probability.
- The study jointly optimizes BS beamforming and IRS reflect beamforming under time-varying channel conditions.
- The secure beamforming optimization problem is formulated as a reinforcement learning problem for the highly dynamic, complex system.
- A deep PDS-PER learning-based approach jointly optimizes BS and IRS beamforming in the dynamic IRS-aided secure communication system.PDS and PER are used to improve learning convergence rate and efficiency.
- Simulation results show that the proposed approach outperforms existing approaches in system secrecy rate and QoS satisfaction probability.