Source-linked AI summary
A Survey on Low Latency Towards 5G: RAN, Core Network and Caching Solutions
Imtiaz Parvez, Ali Rahmati, Ismail Guvenc, Arif I. Sarwat, Huaiyu Dai
TL;DR
Achieving 5G ultra-low latency requires changes across radio access, core-network, and caching domains. This paper surveys proposed solutions and architectures, including SDN, NFV, MEC, and caching, while reporting field-test and scheme-level results.
Problem
5G must support ultra-low-latency services, with latency no more than 1 ms critical in some applications, requiring changes across multiple network domains.
Method
The paper presents a comprehensive survey of low-latency sources, constraints, 5G architecture, and proposed solutions across RAN, core network, and caching.
Results
The survey reviews low-latency techniques and architectures, including SDN, NFV, MEC, high-speed backhaul, distributed caching, and field tests, with reported schemes meeting 5G latency requirements.
Takeaways & Limitations
Low-latency 5G design spans coordinated changes in radio, core, backhaul, and caching infrastructure rather than a single network component.
Takeaways & Limitations
Detailed comparison of the surveyed low-latency solutions is beyond the scope of this work.
Abstract
from arXiv · showhide
The fifth generation (5G) wireless network technology is to be standardized by 2020, where main goals are to improve capacity, reliability, and energy efficiency, while reducing latency and massively increasing connection density. An integral part of 5G is the capability to transmit touch perception type real-time communication empowered by applicable robotics and haptics equipment at the network edge. In this regard, we need drastic changes in network architecture including core and radio access network (RAN) for achieving end-to-end latency on the order of 1 ms. In this paper, we present a detailed survey on the emerging technologies to achieve low latency communications considering three different solution domains: RAN, core network, and caching. We also present a general overview of 5G cellular networks composed of software defined network (SDN), network function virtualization (NFV), caching, and mobile edge computing (MEC) capable of meeting latency and other 5G requirements.
I. INTRODUCTION
5G must support latency-critical human and machine applications while improving capacity, reliability, energy efficiency, and connection density. This survey organizes low-latency approaches across RAN, core-network, and caching domains.
- Motivation: Current 4G networks cannot fulfill all technical requirements of emerging services such as wearable devices, virtual or augmented reality, and immersive 3D experiences.These requirements span data rate, latency, reliability, energy efficiency, traffic density, mobility, and connection density.
- Architectural challenge: Achieving low latency requires architectural changes because delay arises across the RAN, backhaul, and core network.The survey identifies SDN, NFV, MEC, caching, and new physical-layer interfaces as relevant architectural or access-network technologies.
- Survey scope: The survey reviews cellular-network latency reduction comprehensively and divides existing solutions into RAN, core-network, and caching categories.It also discusses latency sources, fundamental constraints, field tests, trials, experiments, and future research directions.
III. SOURCES OF LATENCY IN A CELLULAR NETWORK
End-to-end packet delay is distributed across radio access, backhaul, core processing, and transport to the Internet or cloud. The radio component includes queuing, alignment, transmission, processing, and retransmission effects.
- End-to-end decomposition: The one-way transmission delay is decomposed as T = T_Radio + T_Backhaul + T_Core + T_Transport.The E2E delay is approximately 2×T, according to the paper’s latency model.
- RAN latency: T_Radio covers packet transmission between eNBs and UEs, including propagation, processing, transmission, and retransmission delays.Processing includes functions such as channel coding, scrambling, CRC attachment, precoding, modulation mapping, and OFDM signal generation.
- Backhaul latency: T_Backhaul is the connection delay between the eNB and EPC, while microwave can have lower latency than optical fiber but limited spectrum capacity.Backhaul technology therefore involves a latency–capacity trade-off.
- Core latency: T_Core is the processing time of core entities for security, bearer control, mobility handling, IP allocation, and packet filtering.The EPC control plane and data plane have different QoS needs, motivating their separation for more efficient processing.
- Radio-delay components: Radio latency is modeled as T_Radio = t_Q + t_FA + t_tx + t_bsp + t_mpt.These terms represent queuing, frame alignment, transmission processing and payload transmission, base-station processing, and terminal processing.
- Latency constraint: For low-latency communication, T_Radio should not exceed 0.5 ms, requiring radio transmission times on the order of hundreds of microseconds rather than the 1 ms 4G configuration.The paper links this target to changes in frame structure, modulation and coding, waveforms, transmission techniques, and symbol design.
IV. CONSTRAINTS AND APPROACHES FOR ACHIEVING LOW LATENCY
Wireless networks face fundamental trade-offs among capacity, coverage, latency, reliability, and spectral efficiency. The paper identifies shorter frames, lower retransmission probability, priority handling, and edge caching as approaches to reduce delay.
- Fundamental constraints: Optimizing one wireless-network metric can degrade another because capacity, coverage, latency, reliability, and spectral efficiency are fundamentally coupled.In LTE, the 10 ms radio frame and smallest 1 ms TTI impose a fixed latency-related structure.
- RAN approaches: A shorter radio frame with limited control overhead is needed to reduce transmission time for low-latency communication.Scheduling, resource allocation, and channel-training procedures can be eliminated or merged to reduce overhead.
- RAN approaches: Reducing first-transmission packet error probability can lower retransmission delay through new waveforms and transmission techniques.This approach targets the retransmission component of packet latency rather than only the initial transmission interval.
- Traffic handling: Latency-critical data requires priority over normal data so it can be dispatched immediately.The paper presents priority mechanisms as an approach for handling delay-sensitive traffic.
- Trade-offs: Asynchronous communication may reduce latency relative to synchronized operation but requires additional spectrum and power resources.Synchronization and orthogonality in OFDM are identified as barriers to achieving low latency.
- Caching approaches: Caching at the network edge can reduce delay between the core network and base station by storing popular data closer to users.The survey groups low-latency solutions into RAN, core-network, and caching categories.
V. RAN SOLUTIONS FOR LOW LATENCY
RAN low-latency solutions modify the air interface and processing pipeline through shorter and flexible transmission intervals, waveform and access changes, and PHY/MAC enhancements. These approaches must balance latency with reliability, throughput, spectral efficiency, and control overhead.
- RAN enhancements span frame and packet structures, multiple access, waveforms, coding, antennas, control channels, detection, aggregation, QoS, cloud RAN, and location awareness.
- A. Frame/packet structure: LTE uses 10 ms frames divided into 1 ms subframes, whereas increased subcarrier spacing can produce 0.25 ms subframes and shorter OFDM symbols.Changing subcarrier spacing from 15 to 30 KHz reduces OFDM symbol duration from 66.67 µs to 33.33 µs.
- A. Frame/packet structure: Flexible frame designs adjust transmission time intervals and resource allocation to diverse service requirements, but higher offered load can increase control overhead, latency, and reliability impacts.At low offered load, 0.25 ms TTI is described as attractive; with greater load, control overhead increases.
- A. Frame/packet structure: Physical-layer and MAC-layer proposals include flexible control/data resource allocation, guard periods, and smaller-cell or higher-frequency deployments to reduce air-interface delay.The subframe duration is expressed using control symbols, data symbols, symbol and cyclic-prefix durations, and guard periods.
- A. Frame/packet structure: Air-interface designs expose trade-offs among latency, capacity, coverage, reliability, throughput, and spectral efficiency rather than optimizing latency independently.
B. Advanced Multiple Access Techniques/Waveform
Advanced access techniques and waveforms target low latency by reducing synchronization and orthogonality burdens, adapting subcarrier resources, and improving short-packet transmission efficiency. The survey compares orthogonal, non-orthogonal, asynchronous, and filtered multicarrier approaches.
- Asynchronous and non-orthogonal access is investigated because synchronization and orthogonality requirements associated with OFDM hinder low-latency communication.
- SCMA combines symbol mapping and spreading through multidimensional codewords, while NOMA and SCMA address synchronization and orthogonality requirements in 5G scenarios.
- UFMC outperforms OFDM by about 10% in time-frequency efficiency, inter-carrier interference, and long- or short-packet transmission cases.UFMC also performs better than FBMC for very short packets while showing similar performance for long sequences.
- In UFMC, filtered subband components form the time-domain transmit vector, with filtering and shorter symbol duration reducing cyclic-prefix overhead relative to FBMC.The symbol duration is N + L − 1 samples, and block-per-subcarrier filtering enables shorter time-domain filtering.
- UFMC can tailor symbol duration across users by matching N1 + L1 − 1 = N2 + L2 − 1 despite different FFT sizes and filter lengths.
C. Modulation and Channel Coding
Modulation, coding, and transmitter adaptations address the reliability and processing constraints of small packets while reducing transmission overhead. The surveyed techniques include polar and LDPC coding, parallel processing, shortened cyclic prefixes, beamforming, and full duplex.
- C. Modulation and Channel Coding: LDPC and polar codes outperform turbo codes for small packets, and field tests identify polar coding as a candidate scheme for 5G across multiple scenarios.
- C. Modulation and Channel Coding: Highly parallel turbo decoding, reduced-memory IFFT designs, and finite-blocklength information theory are proposed for latency-sensitive processing and strict reliability constraints.
- D. Transmitter Adaptation: An asymmetric window reduces cyclic-prefix overhead by 30%, lowering latency while increasing susceptibility to channel-induced ISI and ICI.
- D. Transmitter Adaptation: A switched mmWave architecture assigns low-resolution digital beamforming to control signals and analog beamforming to data, reducing control-signaling overhead and physical-layer round-trip latency.
- D. Transmitter Adaptation: Full-duplex techniques may improve capacity, feedback, and latency, but their capacity-latency trade-offs require extensive investigation.
E. Control Signaling
Control-signaling and symbol-detection methods reduce latency by limiting overhead, shortening transmission intervals, streamlining bearer setup, and lowering receiver complexity. The surveyed methods also expose reliability, processing, and throughput trade-offs.
- E. Control Signaling: Reducing packet size makes control overhead a larger portion of transmission, motivating sparse encoding, scaled frames, dedicated control channels, and two-symbol uplink subslots.
- E. Control Signaling: Parallel establishment of radio and S1 bearers and a single bearer-configuration signal reduce signaling interaction rounds between the UE and eNBs.
- E. Control Signaling: Adaptive RLC modes, slotted-TTI resource management, and control-plane designs are proposed to reduce latency while balancing processing power, throughput, and signaling overhead.
- E. Control Signaling: SDN-based local mobility management reduces handover latency and signal overhead by minimizing inter-node exchanges and forwarding X2 signaling centrally.
- E. Control Signaling: Symbol detection contributes latency through channel estimation and decoding, motivating low-complexity receivers, while GFDM space-time encoding and widely linear estimation improve symbol error rate and latency.
- E. Control Signaling: Compressed sensing and SCMA receiver designs target low-complexity detection; a real-time SCMA prototype triples throughput while maintaining latency similar to flexible orthogonal transmissions.
G. mmWave Communications
mmWave communications are a promising 5G route to high throughput and low latency, but require redesigned radio, MAC, physical-layer, location-aware, and QoS/QoE mechanisms. The surveyed work addresses adaptable transmission, numerology, beamforming, resource allocation, and service differentiation.
- mmWave Communications: Carrier aggregation using mmWave can provide massive bandwidth and ultra-low latency, particularly for VR/AR applications requiring high throughput.The paper summarizes mmWave low-latency work in Table XI.
- mmWave Communications: New mmWave MAC designs reduce latency through smaller transmission intervals, dynamic control-signal placement, and directional multiplexing.These designs also address multiple users, bursty traffic, and beamforming constraints.
- mmWave Communications: Two physical-layer numerologies target indoor or LOS and NLOS communications, supported by channel measurements in the 28–73 GHz range.A separate mmWave proof-of-concept evaluated throughput in outdoor LOS conditions at mobile speeds up to 20 km/h.
- Location-aware communications: Location-aware resource allocation can reduce overhead and delay by predicting channel quality beyond traditional time scales.The surveyed literature discusses protocol-stack use of location data and remaining challenges for mmWave deployment.
- QoS/QoE differentiation: QoS/QoE control must differentiate latency-critical services because conventional QoS metrics may not capture users’ perceived satisfaction.Proposed approaches include service-to-frequency mapping, client-side monitoring, interference-aware beamforming, and video-stream routing.
J. CRAN and Other Aspects
CRAN, core-network virtualization, SDN, MEC, and distributed architectures target lower latency while also addressing management, scalability, energy, and capacity requirements. The surveyed approaches expose trade-offs, including added controller or virtualization delays and energy–latency tension.
- CRAN: CRAN centralizes baseband processing while retaining radio front ends at cell sites, simplifying management but requiring links with 250 µs delay for low-latency services.Real-time kernels, Docker, and DPDK are proposed to optimize processing and networking latency.
- Energy efficiency: 35% lower SBS energy consumption is achieved in simulations when under-utilized small cells are switched off using an initial transmission delay.The delay allows users to wait for an SBS with better link quality, creating an energy–latency trade-off.
- Core-network limitations: Centralized LTE EPC data-plane routing increases end-to-end latency for local communication, motivating distributed core-network implementations.SDN, NFV, fog, and MEC are surveyed as technologies for distributing network functions.
- SDN and NFV: SDN-based separation of control and user planes enables independent scalability, flexible flow distribution, mobility management, and MEC facilitation.Multiple controllers can improve scalability, but controller deployment may add latency, requiring a design trade-off.
- MEC and distributed architectures: Distributed SDN/NFV-based MEC placement reduces redundant data-center capacity by around 75% while meeting 5G latency requirements and reducing backhaul bandwidth.Distributed core elements reduce end-to-end delay by bringing functions closer to users.
- Core network virtualization: NFV-based EPC virtualizes EPC elements using VMs, while function partitioning and path optimization target network and processing latency.Placing decentralized control closer to users can help high-mobility, low-latency applications but may complicate policy and charging enforcement.
B. Backhaul Solutions
5G backhaul is a latency bottleneck because massive connectivity, small-cell density, and latency-critical services increase capacity demands. Proposed solutions combine adaptive tunneling, optical and PON architectures, unified transport, SDN, and cooperative caching.
- Backhaul requirements: 5G backhaul capacity becomes a bottleneck under 1000x capacity, massive connectivity, and latency-critical services, requiring higher-capacity transport solutions.Existing microwave, copper, and optical-fiber links are used according to availability and requirements.
- General backhaul: Adaptive GTP termination switches between cloud-based and quick tunnels according to user requests or other factors to reduce latency.A proposed intermediate element optimizes GTP tunnels between the eNB, mobile-network interface, and Internet.
- Optical and PON backhaul: 10 ms latency is targeted using modified VLC optical-window links for low-cost small-cell backhaul, while next-generation baseband can achieve end-to-end latency below 2 ms.A PON architecture and dynamic bandwidth allocation support ultra-short-latency handovers between neighboring cells.
- Transport networks: 5G crosshaul unifies backhaul and fronthaul traffic in a packet-based transport network based on MAC-in-MAC Ethernet.The surveyed transport literature also considers SDN and alternative latency-handling technologies.
- Caching: SDN-enabled cooperative caching across macro and small cells supports coverage, low latency, energy efficiency, and throughput in limited-backhaul scenarios.The referenced caching taxonomy includes local, device-to-device, small-cell, and macro-cell caching.
2) mmWave Backhaul:
Low-latency 5G backhaul and caching solutions address bottlenecks between radio access and core networks by using mmWave links and placing content closer to users.
- mmWave Backhaul: MmWave backhaul is presented as a promising option for reliable, high-capacity, low-latency connectivity in ultra-dense 5G networks.Its integration with massive MIMO can improve link reliability and provide sufficient data rates for wireless backhaul.
- mmWave Backhaul: An in-band wireless backhaul framework for inter-base-station coordination is feasible without considerably affecting cell access capacities.
- mmWave Backhaul: Different mmWave frame designs target low latency with durations of 0.1 ms for LOS and 0.05 ms for NLOS scenarios.The LOS structure is described as suitable for short-distance indoor access or in-band backhaul.
- Caching Solutions: Insufficient backhaul capacity can create latency bottlenecks during peak traffic, making caching a candidate approach for latency reduction.
- Caching Solutions: The probability of obtaining content from an eNB is directly associated with download latency, so effective caching strategies can significantly reduce latency.
- Caching Solutions: Caching schemes are grouped into local, D2D, SBS, and MBS caching, with users searching progressively from the nearest available source.
B. Fundamental Latency-storage trade-off in Caching
Caching introduces fundamental trade-offs between latency and resources such as storage, memory, fronthaul, backhaul, and delivery rate. The surveyed works characterize these trade-offs with information-theoretic latency metrics and optimization bounds.
- Fundamental Latency-storage trade-off: Caching research studies trade-offs involving latency versus storage, memory versus rate, memory versus CSIT, storage versus maximum link load, and caching capacity versus delivery rate.
- Fundamental Latency-storage trade-off: Normalized delivery time (NDT), fractional delivery time (FDT), and delivery time per bit (DTB) are used to evaluate latency-related caching trade-offs.
- Fundamental Latency-storage trade-off: Interference alignment in one compared approach may require infinite symbol extension, limiting its practical processing scope.
- Fundamental Latency-storage trade-off: Decentralized placement with coded delivery achieves considerable improvement over a derived lower bound for delivery latency in a cache-enabled network.
- Fundamental Latency-storage trade-off: The surveyed bounds indicate that lowest delivery latency may require cloud-based compressed precoding together with edge-based interference management.
- Fundamental Latency-storage trade-off: For two users and two edge nodes, coded multicasting does not reduce NDT under wireless multicast fronthaul, unlike receiver-side caching.
- Fundamental Latency-storage trade-off: An F-RAN method that integrates online caching and delivery within each time slot outperforms existing offline caching schemes under NDT evaluation.
C. Existing Caching Solutions for 5G
Existing 5G caching research addresses cache placement, content delivery, cooperation, and latency reduction under storage, bandwidth, and coordination constraints. Reported studies include cooperative multicast-aware caching that reduces average content-access latency by up to 13%, alongside field-tested RAN approaches achieving sub-millisecond or approximately 1 ms latency.
- Caching architecture: Caching studies divide mobile file delivery into cache placement and content delivery, using centralized or distributed approaches.Placement determines which content is cached at base stations, commonly according to user requests and available network information.
- Cache placement: Distributed cache placement minimizes average download delay under base-station storage constraints using a low-complexity belief-propagation algorithm for an NP-hard problem.The formulation explicitly captures a trade-off between latency and storage capacity.
- Cooperative delivery: Cooperative content caching and delivery improve average downloading latency compared with previously known caching schemes.Latency-aware caching can also reduce link load, improve delivery time, and converge faster than probabilistic caching; caching with forwarding targets end-to-end UE latency without coordination.
- Cooperative delivery: Cooperation among cells reduces delay relative to non-cooperative caching, while cooperative cloudlet caching can improve cache hit rate and reduce content-delivery latency.These approaches optimize delivery delay subject to finite cache capacity or decentralized cloud-service operation.
- Cooperative delivery: 13% decrease in average content-access latency is achieved by cooperative multicast-aware caching versus multicast-aware caching with the same total cache capacity.The result comes from trace-driven simulations and reflects joint use of multicast and cooperation.
- Field validation: Field experiments report RTT latency as low as 1 ms on a DSP platform and 1.5 ms HARQ RTT for TDD downlink using a novel frame structure.A separate quasi-static simulator reports 0.8365 ms RTT latency including uplink scheduling requests, a fivefold reduction relative to LTE.
IX. OPEN ISSUES, CHALLENGES AND FUTURE RESEARCH DIRECTIONS
Although proposals target 1 ms latency, 5G still has open issues across RAN architecture, channel modeling, access design, admission control, multiplexing, and CRAN/HRAN deployment. Future work must validate proposed techniques in field tests while addressing trade-offs among latency, reliability, energy, and spectral efficiency.
- Cross-cutting challenges: Existing low-latency proposals require further validation in field tests and continued evolution from current LTE systems.Open issues span RAN, core network, backhauling, caching, and resource management.
- RAN channel design: mmWave offers spectrum from 3-300 GHz, but deployment depends on location and environmental topology, while channel models remain insufficiently developed.Needed work includes delay spread, path loss, NLOS beamforming, angular spread, Doppler, and atmospheric effects across indoor and outdoor settings.
- RAN channel design: Small packets prevent the distortion and thermal-noise averaging available with large packets, requiring channel modeling, simulations, and field tests across carrier bands.The issue is specifically tied to low-latency small-packet transmission.
- Resource management: RAN admission control remains insufficiently explored when spectral and energy efficiency must be maintained under latency constraints.CRAN/HRAN and caching are identified as relevant components, but performance bounds remain to be investigated.
- Waveforms and access: OFDM's orthogonality and synchronization requirements conflict with low-latency goals, motivating access techniques and waveforms requiring less coordination and robust operation in dispersive channels.Candidate non-orthogonal or asynchronous schemes include SCMA, IDMA, GFDM, FBMC, and UFMC.
- Scheduling: Latency-critical packets must be multiplexed with other traffic, but instant access and resource reservation remain insufficiently explored.The challenge is to protect latency-critical services while sharing radio resources.
- Network architecture: Large CRAN/HRAN deployments are difficult to design, leaving trade-offs such as energy efficiency versus latency unresolved.The challenge arises when combining centralized and heterogeneous radio access architectures.
B. Core Network Issues
The survey identifies core-network and caching challenges that must be addressed to reduce 5G latency. Key issues include orchestration, scalability, backhaul adaptation, cache design, protocol support, resource allocation, and mobility management.
- Core network: SDN/NFV core networks require heterogeneous-resource orchestration that maintains low latency.Effective resource allocation and function implementation remain open research issues.
- Core network: Existing SDN/NFV studies often emphasize control-plane integration without detailed implementation or user-plane scalability.The survey highlights standardization and scalability as further research opportunities.
- Core network: Adaptive mmWave front/backhaul techniques are needed to optimize heterogeneous backhaul utilization while supporting low latency.The survey identifies mmWave as attractive because of low implementation cost and limited fiber availability.
- Caching: Caching-enabled MEC can support memory-intensive applications, but it introduces trade-offs among capacity, latency, storage, link load, memory, and rate.The survey discusses distributed and centralized caching, placement, delivery, and cooperation as latency-reduction approaches.
- Caching: Caching research remains open on how cache size, location, wireless channels, redundancy, intra-cache communication, and protocols affect latency.GTP tunneling also complicates content-aware or object-oriented caching, motivating suitable protocol designs.
- Caching: Mobility complicates caching through interference, pilot contamination, system configuration, user-server association, and latency-inducing handovers.The survey calls for low-latency handover studies across diverse caching scenarios.