Source-linked AI summary

Intelligent Resource Slicing for eMBB and URLLC Coexistence in 5G and Beyond: A Deep Reinforcement Learning Based Approach

Madyan Alsenwi, Nguyen H. Tran, Mehdi Bennis, Shashi Raj Pandey, Anupam Kumar Bairagi, Choong Seon Hong

arXiv:2003.07651v3cs.NIeess.SP

TL;DR

The paper addresses resource slicing for dynamically multiplexed eMBB and URLLC traffic, whose rate, latency, and reliability requirements differ. It combines optimization with deep reinforcement learning for eMBB allocation and URLLC scheduling. Simulations satisfy stringent URLLC reliability while keeping eMBB reliability higher than 90%.

  • Problem

    Dynamic multiplexing requires ensuring eMBB data rate and reliability while guaranteeing URLLC’s stringent requirements.

  • Method

    An optimization-aided DRL framework combines eMBB resource allocation with intelligent URLLC scheduling.

  • Results

    eMBB reliability remains higher than 90% while the proposed approach satisfies stringent URLLC reliability.

  • Takeaways & Limitations

    The framework provides a resource-slicing approach for eMBB–URLLC coexistence in dynamic multiplexing scenarios.

Abstract

from arXiv · show

In this paper, we study the resource slicing problem in a dynamic multiplexing scenario of two distinct 5G services, namely Ultra-Reliable Low Latency Communications (URLLC) and enhanced Mobile BroadBand (eMBB). While eMBB services focus on high data rates, URLLC is very strict in terms of latency and reliability. In view of this, the resource slicing problem is formulated as an optimization problem that aims at maximizing the eMBB data rate subject to a URLLC reliability constraint, while considering the variance of the eMBB data rate to reduce the impact of immediately scheduled URLLC traffic on the eMBB reliability. To solve the formulated problem, an optimization-aided Deep Reinforcement Learning (DRL) based framework is proposed, including: 1) eMBB resource allocation phase, and 2) URLLC scheduling phase. In the first phase, the optimization problem is decomposed into three subproblems and then each subproblem is transformed into a convex form to obtain an approximate resource allocation solution. In the second phase, a DRL-based algorithm is proposed to intelligently distribute the incoming URLLC traffic among eMBB users. Simulation results show that our proposed approach can satisfy the stringent URLLC reliability while keeping the eMBB reliability higher than 90%.

I. INTRODUCTION

5G must multiplex eMBB and URLLC services with sharply different rate, latency, and reliability requirements. The paper therefore develops a resource-allocation framework that protects both services under dynamic URLLC arrivals.

  • Service requirements: eMBB targets high data rates for applications such as 4K video and virtual reality, whereas URLLC targets mission-critical traffic with low latency and ultra-high reliability.The stated URLLC packet-error-rate requirement is around 10^-5, compared with 10^-3 for eMBB.
  • 5G NR enablers: 5G NR supports shorter URLLC transmission intervals through numerology, reduced symbol periods, mini-slots, and slot configurations.The paper notes that symbol periods can be shortened by increasing sub-carrier spacing, while mini-slots can use 2–3 symbols.
  • Paper contribution: The proposed framework leverages 5G NR frame-structure flexibility to allocate resources for the distinct requirements of eMBB and URLLC.Its stated objective is to reduce coexistence-related eMBB degradation while accounting for both services’ reliability.
  • Coexistence challenge: Dynamic multiplexing prevents incoming URLLC packets from being delayed, so immediate URLLC scheduling can reduce both eMBB throughput and reliability.The paper defines dynamic multiplexing as overlapping URLLC traffic on eMBB traffic at every mini-slot.
  • Existing scheduling approaches: Preemptive scheduling meets URLLC latency by stopping eMBB transmission during short URLLC TTIs, but may degrade ongoing eMBB transmissions.Orthogonal scheduling instead reserves resources, while dynamic reservation requires additional control overhead.
  • Existing scheduling approaches: Reserved URLLC resources can be wasted when no URLLC transmission occurs, creating an efficiency trade-off for orthogonal scheduling.The dynamic reservation scheme avoids fixed reservation but requires additional control overhead compared with semi-static reservation.

C. Challenges and Contributions

The paper addresses eMBB–URLLC coexistence under dynamic traffic and channel conditions by combining optimization-based allocation with DRL scheduling. Its framework balances eMBB rate and reliability while targeting stringent URLLC reliability requirements.

  • Challenges: Coexisting URLLC and eMBB traffic creates a trade-off among latency, reliability, and spectral efficiency.
  • Challenges: Existing standalone optimization methods may fail to capture dynamic URLLC traffic and channel conditions or violate URLLC reliability constraints after relaxation.
  • Contributions: The proposed optimization objective maximizes average eMBB data rate, minimizes its variance, and satisfies URLLC constraints.
  • Contributions: The framework uses two phases: eMBB resource allocation for RBs and power, followed by DRL-based URLLC scheduling over ongoing eMBB transmissions.
  • Contributions: The eMBB allocation problem is decomposed into subproblems and relaxed into convex forms, while URLLC scheduling replaces integer puncturing counts with continuous RB weights.
  • Results: The DRRA-PGACL combination provides a reliable and efficient approach, with simulations satisfying stringent URLLC reliability and keeping eMBB reliability above 90%.

D. Organization

The paper reviews URLLC requirements, coexistence approaches, and DRL methods before presenting its system model, algorithms, evaluation, and conclusion. Prior work covers resource allocation, multiple-access trade-offs, slicing, and reliability-aware optimization.

  • Organization: Section II reviews URLLC requirements and design, eMBB–URLLC coexistence, and DRL in wireless networks.
  • Organization: Section III introduces the system model, eMBB data-rate impact, URLLC data rate, chance constraints, and final problem formulation.
  • Organization: Section IV presents the eMBB resource allocation algorithm, while Section V presents the DRL-based resource slicing framework.
  • Organization: Section VI evaluates the proposed algorithms, and Section VII concludes the paper.
  • Related Work: Prior studies examine short-blocklength allocation, latency and reliability constraints, V2V power minimization, and joint optimization of radio resources and transmission schemes.
  • Related Work: Related coexistence research compares orthogonal and non-orthogonal access, puncturing, slicing, matching, bandits, and risk-aware formulations.

C. DRL in wireless networks

Prior work applies deep reinforcement learning to wireless resource allocation and eMBB–URLLC coexistence. This paper combines optimization-based resource allocation with DRL while incorporating eMBB rate variability and reliability into dynamic multiplexing.

  • Related work: Prior studies use deep RL for resource allocation and decision-making under rate, latency, and reliability constraints.Some approaches map user rates to resource-block and power-allocation vectors, using latency and reliability as feedback.
  • Related work: DRL-based methods for eMBB–URLLC coexistence include deep deterministic policy gradient and deep Q-learning algorithms.
  • Novelty: Unlike related work, the paper incorporates the risk associated with serving random URLLC requests during ongoing eMBB transmissions.It uses eMBB data-rate variance to characterize transmission risk and reliability under coexistence.
  • Contribution: The proposed approach combines optimization theory with DRL for resource allocation in dynamic eMBB–URLLC multiplexing.The paper describes this as a holistic approach and analyzes mean–variance aspects of the coexistence problem.
  • System model: The system serves downlink URLLC and eMBB slice requests through a gNB connected to eMBB and URLLC users.The model includes K eMBB users, N URLLC users, and B resource blocks.
  • System model: Incoming URLLC packets are transmitted immediately by puncturing ongoing eMBB transmissions, which can degrade eMBB service quality.eMBB users use long TTIs, whereas URLLC users use short TTIs or mini-slots.

B. URLLC data rate based on finite block-length coding

The paper models URLLC rates in the finite block-length regime and formulates risk-aware joint eMBB/URLLC resource allocation. The resulting mixed-integer nonlinear problem is addressed through a two-phase optimization-and-learning approach.

  • URLLC rate model: URLLC achievable rates use finite block-length channel coding because short packets and reliability requirements are not accurately captured by Shannon capacity.The formulation includes channel dispersion and transmission error probability.
  • eMBB allocation: RBs and transmission power are allocated to eMBB users at the beginning of each eMBB time slot to satisfy minimum rates and improve reliability.
  • Reliability-aware scheduling: Puncturing users with low data rates can cause high eMBB reliability degradation, motivating puncturing preferences in resource allocation.
  • Optimization formulation: The strategy maximizes average eMBB data rate, reduces eMBB reliability impact, and satisfies URLLC constraints using a variance-aware formulation.The objective combines expected eMBB rate and its variance, with β controlling the variance weight.
  • Optimization formulation: The URLLC reliability constraint limits outage probability below Θmax, while additional constraints govern power, RB assignment, and punctured mini-slots.The punctured-mini-slot variable can take integer values from 0 through M.
  • Solution structure: The joint resource-allocation problem is a mixed-integer nonlinear, NP-hard problem whose feasible search may have exponential complexity.
  • Solution structure: A two-phase approach combines optimization methods and learning to avoid the difficulty of solving the full problem directly.

IV. EMBB RESOURCE ALLOCATION: OPTIMIZATION METHODS BASED APPROACH

The paper addresses eMBB resource allocation by reformulating a non-convex mixed-integer problem and solving decomposed, relaxed subproblems iteratively. The objective incorporates eMBB rate variability through a risk-sensitive utility while using rounding to recover feasible integer allocations.

  • Risk-sensitive objective: The objective uses an exponential risk-averse utility that captures both the mean and variance of eMBB data rates.The utility becomes risk-neutral as µ →0 and more risk-averse as µ increases.
  • Problem decomposition: The mixed-integer non-convex problem is decomposed into eMBB RB allocation, eMBB power allocation, and URLLC scheduling subproblems.The variables x and z are relaxed to continuous variables for the RB-allocation and URLLC-scheduling subproblems.
  • Problem decomposition: Markov’s inequality converts the probability constraint into a linear constraint, enabling tractable relaxed subproblems.The relaxed subproblems are solved iteratively until convergence.
  • Convex reformulation: Each decomposed subproblem is transformed into a convex form before iterative optimization and final binary conversion.The RB allocation formulation is established as convex, while rounding restores binary allocation variables.
  • Feasibility recovery: The rounding formulation maximizes G(˜x) while minimizing the maximum RB-allocation violation ∆ through a negative weight α.This modification is used to obtain a feasible solution after rounding.

B. eMBB power allocation problem

The power-allocation and URLLC-scheduling subproblems are shown to be convex after relaxing discrete variables and replacing difficult constraints. An iterative DRRA procedure then rounds the relaxed solution and converges sub-linearly.

  • eMBB power allocation problem: For fixed relaxed RB allocation and URLLC placement, the eMBB power-allocation problem is formulated as a convex optimization problem.Its objective is concave in power, and the remaining constraints are linear.
  • URLLC scheduling subproblem: The URLLC scheduling problem is reformulated as a convex optimization problem by replacing its chance constraint with a Markov-inequality-based linear constraint.The puncturing variable is represented continuously during relaxation and linked to the number of punctured mini-slots.
  • DRRA procedure: The DRRA algorithm decomposes the problem, relaxes integer variables, and repeatedly solves RB allocation, power allocation, and URLLC scheduling subproblems.At each iteration it updates ˜x, p, and w, then derives z from the relaxed scheduling solution.
  • DRRA procedure: The algorithm converts the relaxed RB allocation into a binary solution using threshold rounding and a modified problem that controls allocation violations.The best rounding is achieved when ϱ →1.
  • Convergence: The convergence analysis states that the algorithm converges sub-linearly in the order of 1/ϵ.The stopping condition is based on an ϵ-optimal solution.

V. INTELLIGENT URLLC SCHEDULING: DEEP REINFORCEMENT LEARNING BASED APPROACH

The paper adds deep reinforcement learning to handle dynamic and sporadic URLLC traffic and channel variations after optimization-based resource allocation. The DRL agent dynamically verifies URLLC reliability and adjusts system parameters.

  • Motivation and framework: The DRL-based algorithm addresses random, sporadic URLLC traffic and channel variations through interaction with the environment.The URLLC reliability constraint is dynamically verified as system parameters are adjusted to URLLC requirements.
  • State and action design: The reduced state contains eMBB data rates, URLLC channel states, and arrived URLLC traffic at each decision epoch.The reduction summarizes allocated RBs, power, and channel state through the eMBB data rate.
  • State and action design: The action is a B × M puncturing matrix specifying the number of punctured mini-slots for each eMBB resource block.The agent selects punctured resources from each RB based on the learned policy.
  • Reliability-aware learning: A time-varying reliability weight φ(t) is used to ensure URLLC reliability as network states change dynamically.The weight depends on an estimated outage probability obtained from empirical measurements over the last T slots.
  • PGACL approach: The policy-gradient actor-critic algorithm combines policy learning and value learning to optimize policies with fast convergence and low computational cost.The actor controls the policy, while the critic evaluates the selected policy using the reward function.

A. PGACL algorithm for URLLC scheduling

PGACL uses actor and critic components to learn URLLC scheduling policies from network states, rewards, and value estimates. The proposed DRRA-PGACL framework initializes learning with DRRA solutions and trains through experience replay.

  • PGACL components: The actor selects actions from the policy, while the critic evaluates the selected policy using the reward function.The actor updates policy parameters by policy gradients, and the critic updates value estimates using gradient descent.
  • Critic learning: The critic approximates the value function with a linear function estimator and computes estimation error using temporal-difference learning.The value estimate uses a basis-function vector and a weight parameter vector.
  • DRRA-PGACL integration: The DRRA-PGACL framework forwards DRRA resource-allocation results and the current network state to PGACL.The experience pool is initialized using the current DRRA solution.
  • DRRA-PGACL integration: During the first ˆT learning steps, PGACL replaces its selected action with the DRRA-derived action before continuing policy-based scheduling.Later interactions store state, action, reward, and next-state tuples for training.
  • Training and evaluation: The network is trained by sampling random experience tuples, while the reliability weight φ(t) is updated according to the learning procedure.The paper evaluates convergence time and performance through simulation experiments.

1) MAT [15]:

The evaluation compares resource-allocation and URLLC-scheduling strategies across fairness, eMBB rate stability, convergence, and URLLC outage reliability. Risk-sensitive allocation improves fairness and transmission stability, while optimization-aided DRL converges faster and better controls URLLC outage tails.

  • Fairness: Around 90% fairness is achieved with µ = −10, whereas µ = −0.1 produces lower fairness closer to Sum-Rate.
  • Fairness: The Sum-Rate approach gives the worst fairness because it maximizes average sum rate without considering data-rate variance.
  • Rate stability: Higher negative µ values reduce eMBB sum-rate variance and produce more stable, reliable transmissions over time.
  • Rate stability: 50 Mbps average eMBB sum rate varies from 40–60 Mbps at µ = −5.0, while µ = −10.0 narrows the range to 45–52 Mbps.
  • Convergence: Optimization-aided PGACL uses DRRA results for initial training, enabling fast convergence and better response to the dynamic environment.
  • URLLC reliability: At Θ∗ = 0.04, DRRA has about 0.18 violation probability, while PGACL ensures stringent URLLC reliability and minimizes outage tail risk.

E. Impact of URLLC traffic on eMBB reliability

The proposed risk-averse method balances eMBB rate and reliability under varying URLLC traffic and target-rate requirements. It outperforms baselines in eMBB reliability, while increasing URLLC load reduces both reliability and average eMBB rate.

  • E. Impact of URLLC traffic on eMBB reliability: The proposed algorithm guarantees higher eMBB reliability than the compared baselines.
  • E. Impact of URLLC traffic on eMBB reliability: The proposed method protects users in bad channel states by puncturing users in better states, enhancing eMBB reliability.
  • E. Impact of URLLC traffic on eMBB reliability: Higher URLLC traffic decreases eMBB reliability because more eMBB resources must be punctured.
  • E. Impact of URLLC traffic on eMBB reliability: When URLLC load rises from 10 to 90 packets/time slot, Sum-Rate decreases from 64 Mbps to 55 Mbps, while the proposed rate falls from 55 Mbps to 40 Mbps.

APPENDIX A

The convergence analysis relies on block multi-convex structure and three assumptions concerning continuity, the KL property, and initialization near a critical point. Under these assumptions, the sequence converges globally to the closest critical point, with a rate depending on θ.

  • APPENDIX A: The optimization variable is partitioned into blocks, with a differentiable function, convex block terms, and a block multi-convex feasible set.
  • APPENDIX A: The analysis assumes G is continuous and bounded below, while ψ satisfies the Kurdyka-Lojasiewicz property.
  • APPENDIX A: The initial point must be sufficiently close to a critical point, and each iteration’s function value must remain above the critical-point value.
  • APPENDIX A: Under these assumptions, the sequence converges globally to the closest critical or stationary point.
  • APPENDIX A: When θ = 2/3, the analysis yields a sub-linear convergence rate.
Loading 2003.07651v3…