Source-linked AI summary
RIS Enhanced Massive Non-orthogonal Multiple Access Networks: Deployment and Passive Beamforming Design
Xiao Liu, Yuanwei Liu, Yue Chen, H. Vincent Poor
TL;DR
The paper addresses fairness-related power allocation and spectral-efficiency challenges in NOMA precoding using an LSTM-based ESN for tele-traffic prediction and a D3QN-based RIS control algorithm. The proposed D3QN algorithm outperforms benchmarks, while the NOMA-RIS approach achieves higher energy efficiency than OMA-enabled RIS systems.
Problem
Fairness for weak users can require power allocation factors near zero, resulting in lower spectral efficiency, while brute-force searching is used in existing practice.
Method
The paper combines an LSTM-based ESN to predict future tele-traffic demand with a D3QN-based algorithm to determine RIS position and control policy.
Results
The proposed D3QN algorithm outperformed the DQN algorithm and benchmarks, and its learned policy determines RIS placement and control for maximal energy efficiency.
Takeaways & Limitations
The NOMA-RIS approach outperforms the benchmarks and provides higher energy efficiency than OMA-enabled RIS systems.
Abstract
from arXiv · showhide
A novel framework is proposed for the deployment and passive beamforming design of a reconfigurable intelligent surface (RIS) with the aid of non-orthogonal multiple access (NOMA) technology. The problem of joint deployment, phase shift design, as well as power allocation is formulated for maximizing the energy efficiency with considering users' particular data requirements. To tackle this pertinent problem, machine learning approaches are adopted in two steps. Firstly, a novel long short-term memory (LSTM) based echo state network (ESN) algorithm is proposed to predict users' tele-traffic demand by leveraging a real dataset. Secondly, a decaying double deep Q-network (D3QN) based position-acquisition and phase-control algorithm is proposed to solve the joint problem of deployment and design of the RIS. In the proposed algorithm, the base station, which controls the RIS by a controller, acts as an agent. The agent periodically observes the state of the RIS-enhanced system for attaining the optimal deployment and design policies of the RIS by learning from its mistakes and the feedback of users. Additionally, it is proved that the proposed D3QN based deployment and design algorithm is capable of converging within mild conditions. Simulation results are provided for illustrating that the proposed LSTM-based ESN algorithm is capable of striking a tradeoff between the prediction accuracy and computational complexity. Finally, it is demonstrated that the proposed D3QN based algorithm outperforms the benchmarks, while the NOMA-enhanced RIS system is capable of achieving higher energy efficiency than orthogonal multiple access (OMA) enabled RIS system.
I. INTRODUCTION
RIS technology can reshape wireless propagation with low hardware complexity, while NOMA improves spectrum efficiency by superimposing users’ signals. Existing work addresses RIS beamforming and NOMA optimization, but dynamic RIS deployment and design remain challenging.
- RIS capabilities: RISs can create virtual LoS links and improve received SINR, especially when direct BS–user links are blocked.They passively reflect signals and can be installed on structures such as building facades.
- RIS capabilities: RIS operation avoids a dedicated power source and requires minimal hardware complexity compared with conventional AF and DF relaying.This motivates RIS deployment in wireless networks.
- Prior work: Prior RIS studies jointly optimize active and passive beamforming, phase shifts, power allocation, transmit power, or weighted sum-rate under different system objectives.Reported objectives include minimizing symbol error rate or transmit power and maximizing energy efficiency or weighted sum-rate.
- Reported outcomes: Prior simulations reported up to 300% higher energy efficiency than conventional multi-antenna AF relaying.Other studies reported improved spectrum and energy efficiency or minimized total transmit power under SINR constraints.
- RIS-assisted NOMA: NOMA improves spectrum efficiency by superimposing signals at different powers and exploiting users’ different channel conditions.RIS-assisted NOMA can align user channel directions, but RIS phase shifts dynamically change the decoding order.
- RIS-assisted NOMA: Existing NOMA-RIS research includes joint precoding and reflecting-coefficient design, phase-shift and power-allocation optimization, and dynamic decoding-order optimization.One approach uses SCA and SROCR to obtain a locally optimal solution.
3) Machine learning in the RIS-enhanced wireless networks:
Machine learning has been used for RIS configuration, but dynamic deployment based on user demand remains insufficiently studied. This paper formulates an energy-efficiency-driven framework using traffic prediction and D3QN-based RIS control.
- Research gap: Prior RIS machine-learning studies mainly used deep learning to learn reflection matrices or map user positions to RIS unit-cell configurations.These methods addressed configuration learning rather than the full dynamic deployment problem.
- Research gap: Research on machine-learning methods for dynamically deploying and designing RISs remained limited.The paper identifies dynamic RIS deployment and design as an unresolved problem.
- Problem formulation: The framework maximizes RIS energy efficiency rather than throughput while satisfying each mobile user’s particular data requirement.The optimization jointly designs RIS deployment and phase-shift control.
- Framework: The proposed framework jointly designs RIS position, phase shift, and power allocation for energy-efficiency maximization.It also introduces a power-dissipation model based on varactor control.
- Proposed methods: An LSTM-based ESN predicts users’ tele-traffic demand using a real dataset, while D3QN determines RIS position and control policy.The LSTM-based ESN uses LSTM units within an ESN architecture.
- Proposed methods: The D3QN method addresses DQN action-value overestimation through double Q-learning and improves exploration with a decaying ϵ-greedy strategy.The algorithm targets joint RIS position design and phase-shift control.
- Results: The LSTM-based ESN balances prediction accuracy and computational complexity, while D3QN outperforms benchmarks in energy efficiency.The conclusions also report that the proposed NOMA-RIS approach outperforms the benchmarks.
II. SYSTEM MODEL AND PROBLEM FORMULATION
The system combines a BS, RIS, and clustered MUs in a MISO-NOMA downlink. RIS phase shifts, NOMA clustering, channel assumptions, and ZF-based signal processing define the modeled communication system.
- A. System Model: The modeled downlink contains an M-antenna BS, K single-antenna MUs, and an N-element RIS installed on a building facade.The RIS is connected to a controller that controls reflecting elements for phase shifting and amplitude absorption.
- A. System Model: NOMA partitions users into L clusters, with two MUs per cluster assumed for simplicity.Each cluster distinguishes a strong MU with larger channel gain from a weak MU.
- A. System Model: The BS is assumed to obtain perfect channel state information through ray tracing technology.The model includes direct BS–MU and reflected BS–RIS–MU links.
- A. System Model: RIS phase shifts form a diagonal phase-shifting matrix, with phase shifts in [0, 2π] and amplitude reflection coefficients in [0, 1].The model assumes βS,n = 1 because independently controlling amplitude and phase is energy-costly.
- 1) Zero-forcing precoding method:: NOMA superposes the two users’ signals, with power-allocation factors constrained by αl,a + αl,b = 1.The strong MU uses SIC to remove weak-user interference, whereas the weak MU decodes directly.
- 1) Zero-forcing precoding method:: ZF precoding removes inter-cluster interference, but weak MUs still experience inter-cluster and same-cluster strong-user interference.The ZF beamforming vector is generated from the strong MU channel in each cluster.
- 1) Zero-forcing precoding method:: The strong MU can remove inter-cluster interference through NOMA-ZF beamforming and intra-user interference through SIC.The weak MU does not perform SIC, so its received SINR retains interference terms.
2) projection hybrid NOMA precoding method:
PH-NOMA combines projection-based beamforming with NOMA to support both users in a cluster while removing inter-cluster interference. Its decoding constraints enforce successful SIC, and its power allocation reflects fairness-related trade-offs.
- 2) projection hybrid NOMA precoding method:: ZF precoding uses a single beam per group, creating a limitation when strong-user and weak-user data demands differ.When the strong user’s demand is comparable to the weak user’s, its power allocation factor must be sufficiently small.
- 2) projection hybrid NOMA precoding method:: PH-NOMA combines conventional ZF and hybrid NOMA precoding to remove inter-cluster interference.Unlike ZF, it considers beamforming for both users in the same cluster.
- 2) projection hybrid NOMA precoding method:: The PH-NOMA formulation expresses separate received signals and SINRs for users a and b in each cluster.The transmission combines wl,a sl,a and wl,b sl,b before reception.
- 2) projection hybrid NOMA precoding method:: Successful SIC requires the decoding rate for a user’s signal at another user to meet the corresponding target decoding rate.For two users, the condition is Rl,b→l,a ≥ Rl,b→l,b under the specified decoding order.
- 2) projection hybrid NOMA precoding method:: For three users, the SIC constraints require multiple ordered pairwise decoding-rate inequalities.The stated constraints include Rl,b→l,a ≥ Rl,b→l,b, Rl,c→l,a ≥ Rl,c→l,c, and Rl,c→l,b ≥ Rl,c→l,c.
- 2) projection hybrid NOMA precoding method:: PH-NOMA increases the achievable sum rate of the two users, but fairness can require low power allocation to the strong user and reduce spectral efficiency.The method is presented as a trade-off between supporting user demands and spectral efficiency.
B. Channel Model
The channel model considers a dense urban RIS-enhanced network with BS, RIS, and mobile-user locations, fading assumptions, traffic satisfaction levels, and system power consumption.
- Deployment scenario: The deployment scenario is a dense urban area where mobile users are surrounded by buildings, with explicit BS, RIS, and user positions.The BS position includes height, while the RIS position is represented in three dimensions.
- Channel assumptions: Path loss depends on the reference distance, link distance, and path loss exponent, while BS–MU and RIS–MU links use Rayleigh and Rician fading models.The supplied passages specify Rayleigh fading for BS–MU and RIS–MU channels and Rician fading for the remaining modeled link.
- Traffic and satisfaction modeling: Before RIS deployment, the BS predicts downlink traffic demand to identify possible congestion and support deployment decisions.The prediction uses machine-learning techniques and traffic history, with user requirements represented through satisfaction levels measured by MOS.
- Traffic and satisfaction modeling: User satisfaction is represented by five MOS categories: excellent, good, fair, poor, and bad, mapped to throughput requirements.The stated MOS ranges are excellent (4.5), good (3.5 ∼4.5), fair (2 ∼3.5), poor (1 ∼2), and bad (1).
- Power consumption: RIS power consumption is modeled through varactor-diode dissipation, with total system consumption including BS, MU, and RIS hardware components.The RIS contribution is expressed as PRIS(t) = NPn, where Pn is the dissipation of each varactor diode.
E. Problem Formulation
The paper formulates RIS-enhanced NOMA control as a long-term energy-efficiency optimization involving RIS placement, phase shifts, power allocation, and decoding order under system constraints.
- Objective: The protocol controls the RIS to assist or supplement cellular networks while maximizing the energy efficiency of the RIS-enhanced system.The objective uses system achievable sum MOS relative to energy dissipation.
- Decision variables: The optimization jointly selects the RIS phase shifts, three-dimensional position, and BS-to-MU power allocation.The RIS position is represented by three coordinates, and phase shifts and cluster powers are optimization variables.
- NOMA constraints: Dynamic decoding order is included to guarantee successful NOMA successive-interference-cancellation performance.The formulation also imposes decoding-order constraints across clusters.
- System constraints: The formulation includes fairness through transmit-rate constraints based on each cluster’s minimal average achievable user rate.The supplied passages identify this minimal average rate as the fairness-related quantity.
- Learning motivation: Because user demand, mobility, and local topology vary, the design is treated as a dynamic decision problem requiring interaction and long-term learning.Reinforcement learning is selected because the agent can monitor rewards from actions and optimize cumulative rather than only current benefits.
III. PROPOSED SOLUTIONS
The proposed solution combines LSTM-based ESN traffic prediction with D3QN-based joint RIS deployment and design, while evaluating the approach against several learning benchmarks.
- Traffic prediction: The paper proposes an LSTM-based ESN algorithm to predict each mobile user’s data demand.The prediction uses cellular traffic history and a real dataset, with an ESN architecture containing input, hidden, and output layers.
- Traffic prediction: The ESN adapts output weights while keeping input, reservoir, and feedback weights randomly generated and fixed after network construction.The model uses LSTM units as hidden neurons within the ESN structure.
- Traffic prediction: Conventional ESN reservoirs have short-term memory because their neurons are sparsely connected.This limitation motivates replacing the reservoir’s hidden neurons with LSTM units.
- Traffic prediction: The LSTM-based ESN improves prediction performance while trading against increased computational complexity.The paper characterizes the resulting behavior as a tradeoff between prediction performance and computational complexity.
- RIS control: The D3QN algorithm jointly determines RIS position, phase shift, and BS-to-MU power allocation, with DDQN, DQN, and conventional Q-learning used as benchmarks.The proposed D3QN method addresses deployment and design together, while the listed algorithms provide comparison points.
1) DQN Based Algorithm for the control of the RIS:
The RIS controller is modeled as a reinforcement-learning agent that selects deployment, phase, and power actions from system states, with DQN extensions addressing long-term rewards and action-value estimation.
- Agent operation: The BS acts as the agent and periodically observes the RIS-enhanced system before selecting actions through a controller.Each action produces a new state and a reward or penalty determined by the formulated objective.
- State and action design: The DQN state contains RIS phase shifts, allocated MU powers, and RIS and MU coordinates, while actions modify RIS position, phase shifts, and power allocation.The BS selects actions using a policy derived from the action-value function.
- Reward design: Maximizing cumulative reinforcement-learning reward is equivalent to maximizing average energy efficiency.The reward is linked to the system objective after each state transition.
- DQN learning: DQN approximates the Q-table with a neural network whose outputs are action-value estimates updated during training.The network parameters are updated by minimizing a target-based loss.
- DQN limitations and extension: DQN supports farsighted system evolution, but its shared action selection and value estimation can overestimate action values.The overestimation introduces errors during network training and motivates double Q-learning.
- DQN limitations and extension: DDQN uses separate primary and target networks so action selection and target-value generation are not based on the same max-over-Q operation.The paper introduces DDQN to address the DQN overestimation issue in RIS deployment and control.
3) D3QN Based Algorithm for Designing the RIS:
The D3QN design uses ϵ-greedy exploration for RIS control, with ϵ decaying over time so exploration gives way to exploitation as learning progresses. The approach is reported to converge to optimal actions under the stated conditions.
- D3QN selects greedy actions with probability 1 −ϵ and exploratory actions with probability ϵ to balance exploitation and exploration.
- Unlike conventional DQN, D3QN automatically decays ϵ from a high exploration rate toward zero as the policy converges.
- For any ϵ, the policy J is ϵ-optimal with probability 1 after some time Tϵ.
- The algorithm initializes RIS position, phase shifts, and power allocation, then repeatedly selects actions, observes rewards, updates states, and trains through replayed transitions.
- The decaying double deep Q-network is reported to converge to optimal actions, while its exploration parameters satisfy ε ∈[b, a + b] with a, b ≥0 and a + b ≤1.
C. State-Action Construction of the D3QN model
The D3QN state represents RIS phases, position, user positions, and allocated powers, while actions adjust RIS placement, phase shifts, and user power. Rewards are tied directly to energy-efficiency changes, and the convergence analysis establishes optimal-Q convergence under stated assumptions.
- State construction: The system state contains RIS element phases, RIS 3D position, mobile-user positions, and BS power allocations, with cardinality N + 2K + 3.
- Action construction: D3QN actions move the RIS, vary its phase shifts, and change each mobile user's transmit power to improve user data rates.
- Action construction: At every timeslot, received SINR determines the dynamic decoding order before the D3QN state is updated.
- Reward design: Actions that improve energy efficiency receive positive rewards, whereas actions reducing energy efficiency receive penalties, making long-term rewards target long-term energy efficiency.
- Convergence: Under the stated learning-rate and approximation assumptions, the decaying D3QN is capable of converging to the optimal Q value.
2) Complexity of the proposed algorithm:
The section analyzes D3QN complexity and simulation performance. The proposed LSTM-based ESN balances prediction accuracy and computational complexity, while D3QN converges and improves energy efficiency relative to the reported benchmarks and OMA configuration.
- Complexity: The total computational complexity of D3QN is O(TNeΞ), determined by CNN processing and learning across T timeslots and Ne episodes.
- Prediction performance: The LSTM-based ESN outperforms ESN in prediction accuracy, is slightly less accurate than LSTM, and offers a tradeoff with computational complexity.
- Algorithm convergence: Both D3QN and DQN converge, whereas Q-learning fails to converge to the optimal state because of its large table state and action spaces.
- Algorithm convergence: D3QN outperforms DQN, attributed to double Q-learning and the decaying ϵ-greedy policy.
- Energy efficiency: Energy efficiency rises sharply as BS transmit power increases from 2dBm to 12dBm, then declines when additional power increases energy consumption without increasing summed MOS.
- Energy efficiency: NOMA-PH-optimal outperforms NOMA-PH-random because dynamic decoding order is recomputed each timeslot, and NOMA-assisted RIS achieves better energy efficiency than OMA-assisted RIS.
4) Impact of the RIS in terms of energy efficiency:
RIS deployment and design improve energy efficiency in the NOMA system, but the gains depend on placement and reflecting-element count. The proposed learning framework combines traffic prediction with D3QN-based RIS control and outperforms benchmarks in energy efficiency.
- RIS deployment and energy efficiency: RIS deployment enhances the system's energy efficiency compared with operation without RIS.Figure 6 compares NOMA networks with and without RIS.
- RIS deployment and energy efficiency: The proposed D3QN algorithm identifies an optimal RIS position that improves performance over random deployment and barycenter placement.The optimal line corresponds to the position derived by the proposed D3QN algorithm.
- Reflecting-element count: Energy efficiency rises rapidly as the number of reflecting elements increases from 5 to 18.The cited discussion concerns a fixed BS transmit power.
- Reflecting-element count: With fixed BS transmit power, an optimal number of reflecting elements exists because additional elements can increase RIS energy consumption without increasing users' sum MOS.The extra consumption is attributed to the RIS power parameter P_RIS.
- Learning-based design: The framework predicts users' tele-traffic demand with LSTM-based ESN and jointly determines RIS position and control policy with D3QN.The BS acts as the D3QN agent and uses feedback from mobile users while learning the RIS policy.
- Learning-based design: The proposed D3QN algorithm converges after appropriate training and outperforms DQN-based benchmarks through double Q-learning and a decaying ϵ-greedy policy.The proposed NOMA-RIS approach is also reported to outperform benchmarks in energy efficiency.