Source-linked AI summary
A Survey on Model-based, Heuristic, and Machine Learning Optimization Approaches in RIS-aided Wireless Networks
Hao Zhou, Melike Erol-Kantarci, Yuanwei Liu, H. Vincent Poor
TL;DR
RIS control becomes highly complex as the number of elements grows, motivating efficient optimization methods. This survey organizes and compares model-based, heuristic, and machine learning approaches for RIS-aided wireless networks, highlighting their different performance, stability, complexity, and generalization characteristics.
Problem
Controlling many RIS elements creates huge solution spaces and makes exact optimization difficult, while existing approaches impose differing formulation requirements.
Method
The survey examines problem formulations and applications of model-based, heuristic, and machine learning optimization techniques for RIS-aided wireless networks.
Results
Model-based methods generally provide satisfying performance and stability but have complicated designs and low generalization, whereas heuristic methods offer low-complexity sub-optimal solutions.
Takeaways & Limitations
The comparison provides a roadmap for selecting RIS optimization techniques according to the requirements of particular wireless-network scenarios.
Abstract
from arXiv · showhide
Reconfigurable intelligent surfaces (RISs) have received considerable attention as a key enabler for envisioned 6G networks, for the purpose of improving the network capacity, coverage, efficiency, and security with low energy consumption and low hardware cost. However, integrating RISs into the existing infrastructure greatly increases the network management complexity, especially for controlling a significant number of RIS elements. To unleash the full potential of RISs, efficient optimization approaches are of great importance. This work provides a comprehensive survey on optimization techniques for RIS-aided wireless communications, including model-based, heuristic, and machine learning (ML) algorithms. In particular, we first summarize the problem formulations in the literature with diverse objectives and constraints, e.g., sum-rate maximization, power minimization, and imperfect channel state information constraints. Then, we introduce model-based algorithms that have been used in the literature, such as alternating optimization, the majorization-minimization method, and successive convex approximation. Next, heuristic optimization is discussed, which applies heuristic rules for obtaining low-complexity solutions. Moreover, we present state-of-the-art ML algorithms and applications towards RISs, i.e., supervised and unsupervised learning, reinforcement learning, federated learning, graph learning, transfer learning, and hierarchical learning-based approaches. Model-based, heuristic, and ML approaches are compared in terms of stability, robustness, optimality and so on, providing a systematic understanding of these techniques. Finally, we highlight RIS-aided applications towards 6G networks and identify future challenges.
I. INTRODUCTION
This survey addresses the complexity of optimizing RIS-aided wireless networks by organizing problem formulations and comparing model-based, heuristic, and machine-learning approaches. It also covers RIS applications toward 6G and identifies control and optimization challenges.
- Motivation: RISs improve wireless propagation with low energy consumption and hardware cost, but independently configuring many elements creates large optimization spaces and management complexity.Complexity increases further when beamforming, spectrum allocation, decoding order, or UAV trajectory variables are jointly optimized.
- Problem formulations: The survey covers objectives including sum-rate, capacity, energy efficiency, fairness, and secrecy-rate maximization, plus power minimization, discrete phase shifts, integer controls, and imperfect CSI.
- Model-based methods: Model-based methods rely on full problem knowledge and include AO, MM, SCA, BCD, SDR, SOCP, FP, and BnB, but require suitable properties such as convexity, continuity, or differentiability.
- Heuristic algorithms: Heuristic methods trade optimality and accuracy for low complexity and fast solutions, supporting NP-hard problems and serving as baselines or supplements.The survey discusses CCP, meta-heuristics, greedy algorithms, and matching theory.
- ML algorithms: The survey analyzes supervised, unsupervised, reinforcement, federated, graph, transfer, hierarchical, and meta-learning for RIS optimization, including dataset acquisition and RL state, action, and reward design.It highlights dataset acquisition and customized RL definitions as important for exploiting RISs.
- Applications and challenges: The work systematically compares optimization approaches, surveys RIS applications including NOMA, SWIPT, mmWave, THz, NTNs, V2X, and ISAC, and identifies future control and optimization challenges.
2) RIS Control:
RIS control requires jointly configuring phase shifts with beamforming, resource allocation, and other network decisions across diverse optimization objectives. The surveyed formulations include sum-rate, power, energy-efficiency, and constraint-aware designs.
- RIS Control: RIS phase-shift design must often be jointly optimized with beamforming, user association, and resource allocation because these problems are non-convex and highly nonlinear.The RIS controller receives information for phase-shift decisions, while joint network decisions enlarge the optimization difficulty.
- RIS Control: Sum-rate maximization primarily controls BS beamforming vectors and RIS phase shifts under transmit-power and phase constraints, with alternating optimization widely used to decouple the variables.The surveyed literature spans MIMO, MISO, SISO, NOMA, mmWave, and vehicular scenarios.
- RIS Control: Power minimization reduces BS transmission power while enforcing SINR or data-rate requirements, but fractional and logarithmic constraints make the formulation non-convex.Studies also examine discrete phase shifts and imperfect CSI, although continuous phases and perfect CSI remain common assumptions.
- RIS Control: Energy-efficiency maximization jointly increases transmission rate and reduces power consumption, making it more complicated than isolated rate or power objectives.Its formulation includes transmit-power, RIS-phase, and QoS constraints, and the power model varies with the deployment scenario.
E. User Fairness Maximization
User fairness optimization targets worst-case user performance, while RIS-aided secure transmission and discrete-control problems introduce additional objectives, variables, and constraints. These formulations expose trade-offs between solution quality, realism, and optimization complexity.
- E. User Fairness Maximization: User fairness maximization targets the minimum SINR or data rate so that performance is protected for the worst-case user.Existing formulations mainly optimize BS beamforming and RIS phase shifts under transmit-power and RIS-phase constraints.
- E. User Fairness Maximization: RIS-aided secure-transmission studies commonly optimize BS beamforming and RIS phase shifts, but many assume single-user settings, continuous phases, and perfect CSI.These assumptions reduce interference and optimization complexity but create a gap between theoretical studies and practical applications.
- E. User Fairness Maximization: Discrete phase shifts, RIS on/off control, resource allocation, and association create mixed-integer nonlinear problems with substantially greater optimization difficulty.Integer variables can arise alongside continuous phase-shift and beamforming variables.
- E. User Fairness Maximization: Quantization offers low complexity but may degrade performance, whereas relaxation requires more elaborate reformulation and heuristic or ML methods directly optimize discrete controls.Deep reinforcement learning can treat discrete RIS phases as actions and interact with the wireless environment to maximize long-term benefit.
H. Optimization Constraints for Imperfect CSI
Imperfect CSI is addressed through deterministic error bounds or statistical error models, while RIS optimization remains difficult because of non-convexity, coupled variables, large solution spaces, and integer controls. The survey positions model-based and alternative algorithms as responses to these challenges.
- H. Optimization Constraints for Imperfect CSI: Perfect and instantaneous CSI is often impractical because of limited feedback overhead, noise, and interference, motivating deterministic and statistical CSI-error models.The deterministic model bounds error magnitude, while the statistical model uses random-error distributions and probabilistic performance constraints.
- H. Optimization Constraints for Imperfect CSI: RIS optimization commonly involves non-convex, highly nonlinear objectives and constraints containing SINR fractions and logarithms.Dedicated transformation and relaxation are therefore needed to obtain more tractable formulations.
- H. Optimization Constraints for Imperfect CSI: RIS phase shifts are highly coupled with active beamforming, NOMA decisions, UAV trajectories, and other controls, making simultaneous optimization difficult.For RIS-UAV systems, changing UAV altitude requires corresponding phase-shift optimization to maintain network performance.
- H. Optimization Constraints for Imperfect CSI: Many RIS elements and additional network controls create large solution spaces, while integer decisions such as association and resource allocation produce NP-hard problems.Mixed continuous-integer formulations become more complicated when discrete controls coexist with phase-shift design.
- H. Optimization Constraints for Imperfect CSI: The survey introduces model-based methods, including alternating optimization, to handle coupled RIS optimization by iteratively solving individual control-variable subproblems.Alternating optimization is time-efficient per iteration and avoids step-size tuning and extra storage vectors.
B. Block Coordinate Descent
BCD generalizes coordinate descent by grouping multiple control variables into blocks, improving efficiency for large RIS optimization problems. Its performance depends on block selection and updating, while MM uses surrogate upper bounds to simplify difficult objectives.
- B. Block Coordinate Descent: BCD groups multiple control variables into dynamically selected blocks, making it more suitable than AO for large RIS optimization problems.Each block is optimized while the other blocks remain fixed.
- B. Block Coordinate Descent: Block selection affects performance because choosing blocks that decrease the objective most can maximize improvement.Block updating also admits alternatives such as proximal updating.
- B. Block Coordinate Descent: BCD has low memory and iteration costs and can support parallel or distributed implementations, but block selection and updating may be difficult.
- B. Block Coordinate Descent: A two-block BCD alternates BS active beamforming and RIS passive beamforming to maximize sum-rate.
- C. Majorization-Minimization Method: MM replaces a difficult continuous objective with an easier surrogate upper bound that can improve or preserve the original objective value.The surrogate is iteratively constructed and optimized for objectives such as power minimization or sum-rate maximization.
- C. Majorization-Minimization Method: MM is low-complexity but may be impractical when constructing a tight global upper bound for non-convex RIS problems.
D. Successive Convex Approximation
SCA approximates RIS optimization objectives with strongly convex surrogate functions without requiring tight upper bounds. This flexibility supports efficient reformulations involving SDR, SOCP, and fractional transformations, but step-size design affects accuracy.
- D. Successive Convex Approximation: SCA is more flexible than MM because its surrogate need not be a tight upper bound of the original objective.The relaxed requirement simplifies surrogate design and implementation for RIS optimization.
- D. Successive Convex Approximation: SCA repeats surrogate construction and solution updates until convergence, using a step size to control each variable update.The surrogate must remain strongly convex over the feasible set.
- D. Successive Convex Approximation: SCA is frequently applied to RIS sum-rate and power-minimization problems because it relaxes MM’s upper-bound requirement.Its surrogate design still depends on the specific objective and constraints.
- D. Successive Convex Approximation: SDR lifts quadratic RIS formulations into semidefinite programs by relaxing the rank-one constraint, enabling efficient convex optimization.Recovering a feasible RIS phase vector from a higher-rank solution can produce a sub-optimal result.
- D. Successive Convex Approximation: SOCP efficiently handles convex RIS subproblems and has lower stated complexity than SDP for large-dimension problems.The cited complexities are O(n^2 P_i n_i) for SOCP and O(n^2 P_i n_i^2) for SDP.
- D. Successive Convex Approximation: AO or BCD can decouple RIS control variables before solving subproblems with SOCP or SDR.
G. Fractional Programming
Fractional programming transforms ratio objectives common in wireless optimization by decoupling numerator and denominator terms. It reduces formulation complexity but generally must be combined with other optimization methods to solve the reformulated problem.
- G. Fractional Programming: FP is useful for SINR and energy-efficiency objectives because wireless optimization frequently contains fractional terms.
- G. Fractional Programming: Dinkelbach’s method converts a single-ratio objective into an iterative difference form with an auxiliary variable updated from the current solution.Alternating updates of the auxiliary variable and decision variable produce a converged solution with non-decreasing auxiliary values.
- G. Fractional Programming: Classical single-ratio transformations cannot directly guarantee convergence or maximization for sum-ratio problems.An equivalent transform is introduced for the more common sum-ratio setting.
- G. Fractional Programming: FP lowers complexity by eliminating fractional terms and is particularly useful when RIS phase shifts jointly affect received signal strength and interference.It also supports RIS-related max-min fairness reformulations.
- G. Fractional Programming: FP usually serves as a transformation step, after which methods such as AO solve the reformulated subproblems iteratively.A common combination applies FP to SINR terms and AO to coupled RIS phase-shift and BS power variables.
H. Branch-and-Bound
Branch-and-bound addresses discrete RIS optimization through tree-based enumeration, bounding, and pruning. It is suited to NP-hard phase-shift control but can be slow and sensitive to search and pruning rules as RIS size grows.
- H. Branch-and-Bound: BnB enumerates subsets of a discrete solution space in a search tree, solving subproblems and pruning branches using estimated bounds.
- H. Branch-and-Bound: BnB performance depends strongly on search and pruning rules, which can be difficult to design.
- H. Branch-and-Bound: BnB is mainly applied to discrete RIS phase-shift control for sum-rate maximization, power minimization, and max-min SINR.These formulations are usually mixed-integer nonlinear programs and are NP-hard.
- H. Branch-and-Bound: BnB may converge slowly when many RIS elements require repeated branching and exploration of new solutions.
B. Meta-heuristic Algorithms
Meta-heuristic algorithms search large RIS phase-shift spaces with few formulation requirements, offering efficient alternatives to model-based optimization. Greedy and matching methods provide additional low-complexity strategies for element control and resource allocation.
- Meta-heuristic Algorithms: Meta-heuristic methods search large RIS phase-shift spaces with few requirements on problem form, using policies such as GA, PSO, ant colony optimization, simulated annealing, and tabu search.They are motivated by the difficulty of obtaining closed-form solutions for large, non-convex RIS control problems.
- Meta-heuristic Algorithms: Population-based methods initialize candidate RIS designs, evaluate fitness such as sum-rate or energy efficiency, and iterate until convergence or a maximum iteration count.GA initialization includes parameters such as population size and crossover rate.
- Meta-heuristic Algorithms: Meta-heuristics readily support continuous and discrete RIS phase shifts and can significantly reduce complexity for MINLPs, but local optima and parameter sensitivity remain concerns.Large RISs may require many GA individuals, increasing exploration cost.
- Greedy Algorithms: Greedy algorithms make locally optimal decisions stage by stage, including element-by-element RIS on/off or phase-shift control, thereby reducing joint optimization complexity.Their simple policies do not guarantee global optimality or consistently strong output.
- Matching Theory: Matching theory targets resource allocation and association problems through swap operations, including user-BS-RIS association, channel assignment, and D2D-user pairing.Two-sided exchange stability requires that no swap blocking pair can improve overall utility.
E. Discussions and Numerical Results
Heuristic algorithms avoid many transformations required by model-based methods and provide low-complexity solutions across RIS optimization tasks. The section also introduces ML approaches, whose effectiveness depends on data, architecture, and deployment conditions.
- Heuristic Algorithms: Heuristic algorithms commonly offer lower complexity than model-based methods by avoiding convexification and other transformations for non-convex RIS formulations.The survey covers CCP, greedy, meta-heuristic, and matching-based approaches.
- Heuristic Algorithms: Greedy, meta-heuristic, and matching methods trade optimality, formulation requirements, or generality differently: greedy methods are locally optimal, meta-heuristics iteratively explore, and matching specializes in allocation.Matching can be more efficient for allocation problems because of its dedicated design.
- Heuristic Algorithms: Heuristic and model-based methods can be combined, such as using SCA for beamforming and phase control while applying greedy optimization to RIS on/off decisions.Other combinations include SDR with greedy heuristics and SCA/SDR with matching.
- Numerical Results: The greedy-versus-genetic example compares element-by-element phase control with evolutionary search over phase-shift combinations under varying RIS sizes.The caption specifies a MISO system with one BS and multiple UEs.
- ML Optimization: ML optimization for RIS networks spans supervised, unsupervised, reinforcement, federated, graph, transfer, hierarchical, and meta-learning approaches.These methods are presented as responses to dynamic environments, evolving architectures, and diverse user requirements.
- Supervised Learning: Supervised-learning studies often predict CSI or RIS phases from partial CSI or pilots, but training performance depends on labeled data quantity and realistic datasets remain rare.Reported training sets range from 5000 to 200000 samples, with higher data rates when samples increase from 5000 to 30000.
2) Loss Functions and Algorithm Training:
Supervised learning trains RIS predictions against labeled targets, whereas unsupervised learning directly optimizes network objectives without predefined outputs. Both approaches involve practical trade-offs in data, architecture, validation, and solution quality.
- Loss Functions and Algorithm Training: Supervised RIS models minimize prediction error against labeled phase-shift or CSI targets, which may come from exhaustive search or model-based algorithms.AO and BCD are cited as sources of desired phase-shift targets for DNN training.
- Loss Functions and Algorithm Training: The supervised-training loss is mean squared error over RIS outputs, comparing desired phase shifts with neural-network predictions across N elements.The dataset supplies desired phases, while network weights determine the predicted phases.
- Algorithm Training: Supervised workflows collect datasets from simulators, searches, testbeds, or model-based methods, then select architectures such as FNNs, CNNs, and recurrent networks.The available inputs can include UE positions, data rates, and received pilot signals.
- Neural Network Architecture and Overfitting: Neural-network design must balance performance against training cost, while overfitting requires controls such as dropout, larger datasets, reduced depth, or early stopping.The survey notes architectures ranging from 4 to 9 layers in supervised studies.
- Unsupervised Learning: Unsupervised neural networks use channel information as input and RIS phase shifts as output, with losses tied directly to objectives such as maximizing SNR.For the illustrated single-user case, maximizing h_RΘG + h_D maximizes the SNR-related objective.
- Unsupervised Learning: Unsupervised methods avoid predefined targets and can be more practical, but their outputs are difficult to validate and their solution quality is not guaranteed.Clustering methods such as k-means group RIS elements with similar estimated channel coefficients to reduce computational complexity.
2) Action Definition:
RL formulates RIS optimization as sequential decision-making through states, actions, and rewards, with action design determined by the controllable network variables. DQN variants and DDPG support discrete and continuous control, respectively, but training can require extensive interaction.
- Action Definition: RL states encode environment information such as CSI, RIS phases, positions, energy levels, or prior rates, while actions represent controllable variables including phases, beamforming, positions, and on/off states.UAV altitude may be included because it affects channel conditions.
- Action Definition: RL rewards reflect objectives such as data rate, energy efficiency, channel capacity, or SNR and can combine multiple objectives and constraints.Reward design therefore links the MDP to the optimization formulation.
- Algorithm Architecture and Training: Q-learning updates state-action values, but large state-action spaces slow convergence; DQN replaces the Q-table with neural-network estimation, and DDQN separates action selection from evaluation to reduce overestimation.The main and target networks estimate current and target Q-values in DQN.
- Action Definition: Discrete-action methods such as DQN or DDQN quantize continuous RIS phases, whereas DDPG directly handles continuous phase-shift actions.DDPG combines actor-critic learning with Q-value evaluation.
- Algorithm Architecture and Training: In the DRL workflow, the agent maps state s to beamforming and RIS-phase actions, receives rewards, transitions to s′, and stores experience tuples for training.The illustrated applications include DDQN and DDPG for joint active and passive beamforming.
- Algorithm Architecture and Training: RL commonly suffers from low sampling efficiency, requiring potentially hundreds of millions of environment samples for real-world training.This creates substantial interaction costs outside simulation.
D. Federated Learning and RISs
Federated learning supports distributed RIS-aided wireless optimization while preserving local data, but wireless-link impairments and device heterogeneity remain important constraints.
- 1) RIS-enhanced Over-the-air FL:: FL trains local models on decentralized devices and aggregates their parameters into a global model without exchanging local datasets.This distributed procedure is applied to RIS-aided wireless communications and can reduce data-sharing requirements.
- 1) RIS-enhanced Over-the-air FL:: RIS-enhanced AirFL provides an indirect UE-RIS-BS path that improves model-uploading and downloading efficiency, convergence rate, and precision.RISs manipulate signal propagation to improve channel capacity when obstacles cause penetration loss and slow parameter exchange.
- 1) RIS-enhanced Over-the-air FL:: RIS-aided FL studies optimize global training loss, MSE, power consumption, or FL utility through variables such as transmit power and RIS phase shifts.One reported approach solves joint transmit-power and phase-shift control using alternating optimization with QCQP and SDP.
- 1) RIS-enhanced Over-the-air FL:: FL can also optimize RIS communications by predicting achievable rates or aggregating local DDPG networks while limiting dataset exchange and protecting private CSI.Local FL models on RISs can reduce communication overhead, while CSI privacy matters because it may reveal user locations.
- 1) RIS-enhanced Over-the-air FL:: Frequent parameter sharing creates communication overhead, while unequal device computation and storage capacities can affect FL model aggregation and updates.These constraints arise from the distributed implementation rather than from the RIS channel alone.
2) Graph Learning for RIS Control and Optimizations:
Graph learning models RIS-aided wireless networks through node and edge relationships, while transfer and hierarchical learning reuse experience or divide decisions across timescales.
- 2) Graph Learning for RIS Control and Optimizations:: GNNs capture mutual user interference and RIS-user interactions, supporting unsupervised user scheduling and RIS configuration.The graph contains one RIS node and K UE nodes, allowing updates to incorporate neighboring users and configure RIS elements collectively.
- 2) Graph Learning for RIS Control and Optimizations:: GNNs can generalize across changing user numbers by adding or removing components instead of retraining a fixed-size FNN.This property can reduce ML model training effort in networks with variable numbers of users.
- Transfer Learning-boosted Wireless Networks with RISs:: Transfer learning reuses mapped expert knowledge from related source tasks to guide learner actions in target RIS control tasks.The transferred expert value can provide extra rewards for active and passive beamforming decisions targeting higher sum-rate or energy efficiency.
- Transfer Learning-boosted Wireless Networks with RISs:: Transfer learning improves exploration efficiency and convergence, but it depends on existing experts and difficult mappings between source and learner tasks.These dependencies limit direct reuse when task conditions differ substantially.
- Hierarchical Learning for RIS-aided Wireless Networks:: Hierarchical learning uses a meta-controller for long-term goals and sub-controllers for short-term RIS decisions across different timescales.Defining hierarchy relationships and decomposing highly dynamic tasks remain key challenges.
H. Meta-Learning
The survey presents meta-learning and related ML techniques for RIS optimization, including transfer and hierarchical learning, while emphasizing practical training and deployment challenges.
- H. Meta-Learning: Meta-learning learns across related tasks so prior experience can improve target-task training and learning efficiency.An example pre-trains a model on RIS beamforming, BS beamforming, and UAV trajectory-design tasks for UAV-aided joint optimization.
- H. Meta-Learning: Meta-learning must balance broad task coverage against specialization because overly broad or overly specific meta-training can harm target-task adaptation or generalization.The source-task distribution therefore requires careful design.
- ML Techniques for RIS Optimization: Supervised learning needs fine-grained labeled data, whereas unsupervised learning reduces label dependence but makes generated results difficult to validate.The trade-off affects practical selection of learning methods for RIS control.
- ML Techniques for RIS Optimization: Graph learning remains difficult to apply in wireless networks because real-time environments can generate dynamically changing graphs.The survey identifies this application area as still requiring further effort.
- Discussions and Numerical Results: Transfer deep reinforcement learning converges faster and achieves higher average reward than conventional DRL, with larger gains as RIS elements increase.TDRL reuses existing expert knowledge, improving exploration efficiency when conventional DRL faces greater exploration difficulty.
- Discussions and Numerical Results: Combining sleep control with RIS phase-shift optimization yields higher energy efficiency than using either technique alone in multi-BS, multi-RIS networks.Hierarchical reinforcement learning coordinates long-term BS on/off decisions with short-term phase-shift control.
VII. COMPARISON AND RELATIONSHIP BETWEEN MODEL-BASED, HEURISTIC AND ML APPROACHES
The survey compares model-based, heuristic, and ML optimization approaches for RIS networks and relates their selection to application-specific requirements and future 6G challenges.
- Comparison and Relationship: Model-based optimization can be efficient after reformulation but requires complex transformations and full parameter knowledge, limiting solution quality and robustness under uncertainty.Approximations and relaxations may undermine solution quality, while environmental changes can substantially affect performance.
- Comparison and Relationship: Heuristic methods trade optimality and accuracy for low complexity and fast solutions, and can complement model-based methods within alternating-optimization schemes.One example combines SCA for beamforming with a greedy algorithm for RIS on/off control.
- Comparison and Relationship: ML methods adapt to dynamic environments and use unified formulations, while combinations with model-based algorithms can exploit both approaches.A BCD-generated dataset can support subsequent location-based supervised learning for RIS-aided mobile edge computing.
- RIS-assisted 6G Applications: RIS-assisted 6G applications include NOMA, SWIPT, mmWave and THz communications, non-terrestrial networks, V2X, and ISAC, each with distinct optimization difficulties.Examples include changing NOMA decoding order, coupled mmWave beam selection, and worst-case reliability requirements in V2X.
- Challenges and Future Directions: Future work should address ML deployment, training mode and cost, practical RIS location and scale optimization, and flexible frameworks combining complementary approaches.RIS placement must account jointly for wireless environment, user distribution, and service requirements rather than remain a fixed simulation parameter.
- Comparison and Relationship: Model-based methods offer performance and stability, heuristics provide low-complexity suboptimal solutions, and ML offers generalization and robustness across dynamic environments.The survey emphasizes that no approach dominates universally; selection depends on optimization requirements and application scenarios.