Source-linked AI summary
MARL-Based Sequential RIS Auctions: A Physical-Layer Security Analysis
Yuanyu Zhang, Yu Zhang, Jialu He, Zhixin Huang, Shuangrui Zhao, Yulong Shen
TL;DR
The paper asks how competition between legitimate receivers and eavesdroppers for neutral-operator RIS resources affects physical-layer security. It develops a sequential auction and MARL framework for budget-aware bidding, and reports that the learned strategy achieves the strongest secrecy performance among the compared strategies while approaching the ideal upper bound.
Problem
Security implications of market-based RIS allocation remain largely unexplored when eavesdroppers can pose as legitimate requesters and compete for scarce RIS resources.
Method
The paper combines sequential first-price sealed-bid RIS auctions, a budget-coupled Markov game, and MADDPG-based MARL with centralized training and decentralized execution.
Results
The learned policy achieves the highest secrecy rate per unit bid among the compared methods, outperforming random and fixed strategies and approaching the ideal physical-layer upper bound.
Takeaways & Limitations
Security in shared RIS markets is jointly determined by physical-layer conditions and economic allocation decisions.
Abstract
from arXiv · showhide
Reconfigurable intelligent surfaces (RISs) hold great potential to enhance coverage, spectral efficiency, and communication security by intelligently configuring their reflecting elements. When owned by a neutral RIS operator, these elements can be offered as resources for which legitimate receivers and eavesdroppers compete. This paper investigates such competition and evaluates its impact on the physical-layer security performance of legitimate receivers. To model the competition, we develop a sequential RIS auction (SRA) framework, in which a bundle of RIS elements is auctioned in each round through a first-price sealed-bid mechanism, with each bidder submitting its bid based on the achievable rate gain and remaining budget. We then formulate the sequential bidding process as a Markov game by specifying its states, actions, rewards, and state transitions. To solve the game, we propose a multi-bidder deep deterministic policy gradient (MADDPG)-based multi-bidder reinforcement learning (MARL) approach under centralized training and decentralized execution (CTDE), enabling legitimate receivers and eavesdroppers to learn bidding strategies that maximize their long-term economic surplus. Numerical results show that, under the considered eavesdropper bidding strategies, the RL-based strategy enables legitimate receivers to achieve the highest secrecy rate per unit cost, outperforming random and fixed strategies and approaching the ideal physical-layer upper bound.
I. INTRODUCTION
The paper examines how legitimate receivers and eavesdroppers compete for scarce RIS resources supplied by a neutral operator. It models this competition through sequential auctions and budget-aware MARL, finding improved secrecy performance for legitimate receivers.
- Existing market-based RIS studies largely assume legitimate bidders and emphasize communication or economic outcomes, leaving security implications underexplored.An eavesdropper may pose as a legitimate requester and compete for RIS resources.
- The SRA framework allocates disjoint RIS modules sequentially through first-price sealed-bid auctions between a legitimate receiver and an eavesdropper.The RIS is partitioned into M modules of L = N/M elements, with each won module controlled by its bidder for the remaining transmission episode.
- The sequential bidding process is modeled as a budget-coupled Markov game whose states, actions, rewards, and transitions capture inter-round resource and budget effects.The framework uses bidders’ achievable rate gains and remaining budgets to support long-term bidding decisions.
- A MADDPG-based MARL method uses centralized training and decentralized execution to learn long-term bidding strategies for competing bidders.During training, global states and joint actions support value estimation, while execution relies on local observations.
- The learned policy achieves the highest secrecy rate per unit bid among the compared methods, outperforming random and fixed strategies and approaching the ideal physical-layer upper bound.The evaluation considers competition between legitimate receivers and eavesdroppers under the proposed sequential auction framework.
- The network model contains Alice, a passive RIS relay, Bob, and Eve, with RIS modules treated as economically priced and competitively allocated resources.The RIS is operated by an independent, market-oriented service provider.
C. Physical-Layer Security Model
The physical-layer security model derives legitimate and eavesdropper communication rates from RIS-assisted received signals and defines secrecy rate as the nonnegative difference between them. Eve remains passive over the air but can acquire RIS modules through the market to strengthen interception.
- The received signal at each bidder combines the Alice–RIS–bidder cascaded channel, a direct channel, transmitted data, and additive noise.The model uses transmit power P, information symbol x, and AWGN n_i with variance σ^2.
- The achievable communication rates are obtained from the corresponding signal-to-noise ratios for Bob and Eve.
- The instantaneous secrecy rate is defined as Rs = [RB−RE]+, where [·]+ = max(·, 0).
- Eve does not inject wireless interference; instead, it bids for RIS modules and configures won modules to strengthen the Alice–Eve cascaded link.The RIS provider handles registration, billing, and contract enforcement for authorized bidders.
- The provider and Bob have incomplete information about bidders’ security intent and behavior during the sealed-bid auction.Bob observes limited auction feedback, including previous winning prices and its own allocation outcomes.
III. SEQUENTIAL RIS AUCTION (SRA) FRAMEWORK AND PHYSICAL-LAYER SECURITY ANALYSIS
The SRA framework sequentially allocates RIS modules through first-price sealed-bid auctions, linking each bidder’s rate gain and remaining budget to economic competition and security performance.
- A. Auction Setup: A neutral RIS provider allocates M modules sequentially, granting each winner temporary phase-control rights over its acquired modules.Bob seeks to enhance the legitimate link, whereas Eve seeks to strengthen interception capability.
- B. Module Valuation and Bidding Rule: Each bidder’s module valuation converts its marginal achievable-rate gain into monetary value through λ > 0.The valuation is state-dependent because marginal rate gains depend on modules acquired in previous rounds.
- A. Auction Setup: The highest bidder wins each module and pays its submitted bid under the first-price payment rule.Budgets and allocation sets then update before the next auction round.
- B. Module Valuation and Bidding Rule: The payment mechanism creates a trade-off between winning probability and payment cost, producing forward-looking behavior across rounds.Current payments affect future auction opportunities through remaining budgets.
C. Bidding Problem Formulation
The paper formulates budget-coupled sequential bidding as a finite-horizon game and connects module ownership to secrecy-rate changes through valuation-driven allocation conditions.
- C. Bidding Problem Formulation: Bob and Eve’s bidding decisions are temporally coupled because secrecy rate depends on cumulative module allocations across all auction rounds.The bidding problem is formulated as a finite-horizon sequential decision problem under budget constraints.
- C. Bidding Problem Formulation: The bidding ratio satisfies ai(t) ∈ [0, 1] at every auction round.This bounds the strategy used in the sequential optimization problem.
- C. Bidding Problem Formulation: The objective is difficult to solve in closed form because valuations depend on historical allocations and the opponent’s unknown strategy.This removes a recursive closed-form representation of the bidder’s objective.
- C. Bidding Problem Formulation: Incomplete information, simultaneous actions, continuous bidding ratios, and combinatorial allocation states make standard equilibrium analysis and exact dynamic programming unsuitable.The state-action space grows exponentially with the number of modules M.
- D. Security-Economic Coupling in RIS Module Allocation: Module allocation changes secrecy rate through the marginal rate gains of Bob and Eve, which also determine their economic valuations.This provides the interface between auction decisions, physical-layer secrecy, and the learning formulation.
- D. Security-Economic Coupling in RIS Module Allocation: Under the module-isolated channel, allocating a module to either bidder cannot decrease that bidder’s achievable rate.The theorem defines the marginal rate gains used in the security-economic analysis.
- D. Security-Economic Coupling in RIS Module Allocation: Excluding ties, Bob wins a module according to the comparison between the bidders’ bids, with a simplified condition when budgets are inactive.The inactive-budget condition substitutes vi,m(t) = λ∆i,m(t) and cancels the common λ factor.
- D. Security-Economic Coupling in RIS Module Allocation: The SRA mechanism couples economic bidding policy to secrecy through module ownership and the resulting rate increments.The supplied figure is titled “SRA Mechanism with MARL-Based Bidding.”
IV. MARL-BASED BIDDING STRATEGY
The paper solves the sequential RIS auction using CTDE-based MADDPG, with bidders learning forward-looking, budget-feasible policies from local auction observations.
- IV. MARL-Based Bidding Strategy: The sequential bidding problem is instantiated as a Markov game and solved with a CTDE-based MADDPG framework.Each bidder maps local auction observations to budget-feasible bidding ratios while maximizing long-term economic surplus.
A. Markov Game Instantiation
The SRA is modeled as a two-bidder stochastic Markov game whose observations, continuous actions, surplus rewards, and transitions capture sequential allocation and budget dynamics.
- A. Markov Game Instantiation: The SRA Markov game contains Bob and Eve, with state, joint action, reward, and transition components induced by first-price allocation and budget updates.The game is represented by GSRA = ⟨S, A, R, P⟩.
- A. Markov Game Instantiation: Centralized training uses allocation history, budgets, rates, secrecy rate, and round index, while each bidder executes from a partial local state.The local state preserves privacy during decentralized execution.
- A. Markov Game Instantiation: The local observation combines accumulated rate, normalized remaining budget, auction progress, current valuation, previous clearing price, and recent winning outcome.These variables encode economic constraints, valuation, temporal context, and competitive signals.
- A. Markov Game Instantiation: Each bidder selects a continuous bidding ratio controlling bid aggressiveness relative to the current module valuation.Using a normalized ratio creates a bounded action space and supports budget-feasible bidding.
- A. Markov Game Instantiation: The instantaneous reward equals valuation minus bid when a bidder wins and zero otherwise.This reward represents economic surplus and captures the trade-off between aggressive bidding and budget preservation.
- A. Markov Game Instantiation: The first-price auction determines the winner and clearing price from the joint actions, after which the auction state evolves.The transition is Markov because channel realizations remain fixed within an episode and current state-action variables determine related parameters.
B. MADDPG under CTDE
The paper applies MADDPG under CTDE to learn continuous bidding policies for competing RIS bidders. Centralized critics use global auction information during training, while execution relies on local observations.
- Each bidder maps its local observation to a continuous bidding ratio in [0, 1], with the SRA mechanism enforcing bid feasibility.The monetary bid is determined from valuation and remaining budget after the policy selects the normalized ratio.
- Centralized critics evaluate joint bidder actions under the global state to mitigate non-stationarity from simultaneous policy updates.The critic uses the state and both bidders’ actions during training.
- Target networks compute the target value for critic learning, using each bidder’s economic-surplus reward and the discount factor.The target-value construction follows the standard MADDPG procedure.
- Training iterates through auction episodes by selecting actions, executing auctions, updating budgets, storing transitions, and updating actor, critic, and target networks.The procedure runs for the number of auction rounds and then applies temporal-difference and policy-gradient updates.
- The actor parameters are updated with deterministic policy gradients so each bidder adjusts its bidding ratio toward higher long-term utility.This links local policy parameters to the global value estimate.
3) Policy Update:
The policy update uses centralized-critic feedback to account for budget coupling and opponent responses. The resulting policies approximate adaptive best responses in repeated, budget-constrained competition.
- 3) Policy Update:: Centralized-critic feedback lets each bidder adjust its policy toward maximizing long-term utility while accounting for inter-temporal budget coupling.The learned policy receives value information based on global auction interactions.
- 3) Policy Update:: Compared with static or myopic rules, MARL policies anticipate future opportunities and adapt to environmental and strategic changes.The policies are interpreted as approximate best responses under repeated competition.
- 3) Policy Update:: The evaluation examines how sequential allocation, budget coupling, and strategic bidding affect secrecy performance, bidding behavior, and provider revenue.It also assesses the security and economic efficiency of auction outcomes.
A. Simulation Setup and Baseline Strategies
The simulations compare fixed, random, and learned bidding under common RIS and channel conditions. Results show that the learned strategy improves secrecy performance and reward relative to the baselines across the evaluated adversarial settings.
- A. Simulation Setup and Baseline Strategies: The evaluation uses an 8 × 8 RIS with 64 elements partitioned into 8 modules of 8 elements, auctioned over 8 rounds.Results are averaged over 100 independent evaluation episodes after training convergence unless otherwise specified.
- A. Simulation Setup and Baseline Strategies: Fixed bidding uses a constant valuation-based ratio, whereas random bidding uses probabilistic participation and randomly selects feasible bids.Random bidding does not exploit module valuation, previous outcomes, or future opportunities.
- A. Simulation Setup and Baseline Strategies: The smoothed training reward increases from approximately 2.5 to 12.0 and then stabilizes, indicating convergence to a stable bidding strategy.Raw rewards fluctuate because of wireless-channel randomness and multi-bidder competition.
- A. Simulation Setup and Baseline Strategies: Bob-RL achieves the highest row-average reward and secrecy rate against the evaluated heterogeneous Eve strategies.The learned policy exploits state-dependent valuations, payment costs, and remaining budgets.
- A. Simulation Setup and Baseline Strategies: Eve-RL reduces Bob’s average secrecy rate to 1.32 bps/Hz, demonstrating the effect of an adaptive eavesdropper strategy.This result is reported for the considered strategy comparisons.
- A. Simulation Setup and Baseline Strategies: The RL auction tracks the ideal physical-layer upper bound more consistently than fixed and random strategies in episode-wise secrecy-rate traces.Fixed bidding degrades under time-varying channels and budget constraints, while random bidding shows severe fluctuations.
- A. Simulation Setup and Baseline Strategies: The RL strategy shifts the secrecy-rate CDF toward higher values, maintaining approximately 2.2–3.0 bps/Hz and a median close to the ideal upper bound.Random bidding has pronounced low-rate outage behavior, while fixed bidding remains in a narrow, suboptimal range.
D. Secrecy Sensitivity to Auction and Budget Constraints
Secrecy performance depends jointly on auction-round granularity and available budget. The learned strategy is strongest in strategy-sensitive regimes, while scarcity and saturation constrain all bidders.
- D. Secrecy Sensitivity to Auction and Budget Constraints: A small number of auction rounds is sufficient for the RL strategy to establish an effective allocation pattern and retain a margin over baselines.Coarse partitioning limits flexibility, whereas overly fine partitioning fragments marginal gains and bidding budgets.
- D. Secrecy Sensitivity to Auction and Budget Constraints: At very low budgets, all methods perform poorly and RL may trail simpler strategies because its strategic advantage cannot activate under extreme scarcity.This identifies the lower boundary of the strategy-sensitive budget regime.
- D. Secrecy Sensitivity to Auction and Budget Constraints: As budgets enter a strategy-sensitive regime, RL improves sharply and dominates baselines by balancing current wins against future opportunities.At larger budgets, all methods saturate, but RL retains an advantage through high-value configuration selection.
E. Cost-Aware Secrecy Efficiency
The RL-based bidding strategy improves secrecy cost efficiency by adapting payments to module value, competition, and remaining budget. It maintains strong secrecy and utility without excessive payment, while avoiding the inefficiencies of rigid bidding rules.
- Adaptive bidding: 3.55 average bid lets RL secure high-value modules while preserving budget, unlike aggressive fixed strategies reaching 9.91.The adaptive policy avoids both excessive aggressiveness and extreme conservatism.
- Secrecy efficiency: RL sustains a secrecy-to-bid ratio of 1.03, exceeding fixed-strategy ratios that decline from 0.45 to 0.16.RL achieves the highest ratio against every considered Eve strategy.
- Utility and revenue: RL maintains the highest utility across considered adversarial strategies without excessive payment.Higher provider revenue from aggressive bidding does not necessarily improve secrecy or allocation efficiency.
- Strategic determinants: Secrecy performance depends jointly on physical-layer module values, budget constraints, and bidding strategies rather than auction horizon or budget size alone.Increasing the auction horizon or budget does not guarantee monotonic secrecy gains because temporal allocation can produce diminishing returns.
- Framework implication: The market-oriented SRA framework links physical-layer conditions with economic allocation decisions through state-dependent valuation and MARL-based bidding.This enables forward-looking resource acquisition under budget-coupled strategic interactions.