Source-linked AI summary

Stochastic End-to-End Latency Modeling of the IoT-Edge-Cloud Continuum: Impact of Jitter and Traffic Variability on Deterministic Service Provisioning

Keyvan Aghababaiyan, Javier Gozalvez, Baldomero Coll-Perales

arXiv:2608.25658v1cs.NI

TL;DR

The paper asks how stochastic arrival-time jitter and traffic variability affect deterministic end-to-end service guarantees in the IoT–edge–cloud continuum. It introduces an openly released queueing-based model covering communication and computing latency, then analyzes complete latency distributions and deadline compliance. The results show that sensitivity depends on deadlines, computing demand, workload, connectivity, and execution location, with cloud execution especially sensitive to traffic variability.

  • Problem

    Deterministic provisioning requires understanding tail latency and variation across both communication and computing, but unified stochastic characterization across the continuum remains limited.

  • Method

    The paper develops an openly released queueing-based end-to-end model that jointly represents communication and computing latency across the IoT–edge–cloud continuum.

  • Results

    Services with stringent deadlines and larger computing demands are more sensitive to temporal variability, while cloud execution is especially sensitive to traffic variability because of additional communication latency.

  • Takeaways & Limitations

    Offloading decisions should jointly consider service requirements, communication and computing conditions, and sources of latency variability to support deterministic service levels.

Abstract

from arXiv · show

6G will integrate communication and computing capabilities in a IoT-edge-cloud continuum, enabling nodes to distribute workloads across the continuum. To support time-sensitive services, both communications and computing latencies must be controlled. Two key sources of temporal variability are arrival-time jitter and traffic variability. They can both impact the timing at which data is generated, transmitted and processed, and the resulting fluctuations can propagate throughout the continuum, increasing latency uncertainty. This paper studies the impact of stochastic temporal variability on the ability to support end-to-end deterministic service levels across the continuum. To this end, we present a novel queueing-based end-to-end latency model for the continuum, which we openly release. The model jointly captures computing and communication latency, and characterizes the complete end-to-end latency distribution, including tail latency. Our analysis shows that services with stringent latency deadlines and larger computing demands are more sensitive to temporal variabilities, making local execution the preferred option. In contrast, services with more relaxed deadlines are more resilient to temporal variabilities when executed locally or at the edge despite higher average and tail latencies. Edge offloading is beneficial under good cellular connectivity and increasing local processing workloads, whereas cloud execution is more sensitive to traffic variabilities because of the additional communication latency. Our analysis also shows that services offloaded are more sensitive to traffic variability than jitter due to higher communication latencies. These findings highlight that effective service offloading must jointly consider service requirements and sources of temporal variability to guarantee deterministic service levels.

I. INTRODUCTION

The paper addresses the need for end-to-end deterministic service guarantees by introducing a stochastic model that jointly represents communication and computing latency across the IoT–edge–cloud continuum. It analyzes how jitter and traffic variability affect latency distributions, deadline compliance, and offloading decisions.

  • I. INTRODUCTION: Deterministic service levels require bounded tail latency and controlled variation, not average latency alone.Arrival-time jitter and traffic variability increase latency uncertainty as their effects propagate through data generation, communication, and computing.
  • I. INTRODUCTION: The proposed queueing-based model jointly captures communication and computing latency across the RAN, transport network, core network, Internet, and computing infrastructure.The complete model characterizes end-to-end latency distributions, including tail behavior.
  • I. INTRODUCTION: The model is openly released to support analysis of how stochastic temporal variability propagates through the continuum.Its stated use is to study deterministic service provisioning and service offloading policies.
  • I. INTRODUCTION: Prior work commonly studies communication or computing latency separately, while unified end-to-end stochastic characterization remains an objective of this work.Existing offloading models cited in the passage include cases limited to average latency or incomplete communication paths.
  • I. INTRODUCTION: The reference architecture lets user equipment execute tasks locally or reach edge and cloud nodes through the 5G network.The proposed approach models latency across communication and computing domains using queueing theory.

B. Edge Computing

Edge computing is modeled as a multi-server queue whose arrival-process representation changes with the number and variability of offloading sources. The resulting computing latency combines queue waiting and processor service time.

  • B. Edge Computing: An edge node is modeled as an M/M/m queue when many independent sources produce an approximately Poisson arrival stream.Each of the m_e processor cores is an independent parallel server under FCFS scheduling.
  • B. Edge Computing: The edge utilization is ρ_e = λ_e/(m_e μ_e), with λ_e the aggregate arrival rate and μ_e the service rate of each core.The model assumes steady-state operation and uses the Erlang-C formulation for queueing behavior.
  • B. Edge Computing: Under moderate and high loads, edge waiting time is approximated by an exponential distribution, while under low load it is treated as negligible.This load-dependent approximation is combined with exponential processor service time.
  • B. Edge Computing: Total edge computing latency is the convolution of waiting and service times, yielding a hypo-exponential distribution.The latency is T_e = T_W_e + T_S_e.
  • B. Edge Computing: When offloading sources are few or arrivals are non-Poisson, the model adopts a GI/M/m queue with general inter-arrival variability.The Erlang-C extension modifies the mean queue length using the arrival coefficient of variation.

C. Cloud Computing:

Cloud computing is represented by a GI/M/m queue because traffic arrives through an Internet connection with variable inter-arrival times. Cloud latency combines queue waiting with exponential processor service time.

  • C. Cloud Computing:: The cloud node is modeled as a GI/M/m queue with m_c independent cores and exponential service rate μ_c.Its input traffic has average rate λ_c and arrives over an Internet connection.
  • C. Cloud Computing:: Cloud inter-arrival variability is represented with a Weibull distribution because traffic arrives through the Internet.The Weibull model is referenced as the characterization of Internet-related arrival variability.
  • C. Cloud Computing:: The cloud mean queue length is obtained from an Erlang-C extension of the M/M/m model using cloud-specific arrival, service, core-count, and utilization parameters.The zero-service probability is likewise adapted by replacing the edge parameters with λ_c, μ_c, m_c, and ρ_c.
  • C. Cloud Computing:: Under low load, cloud waiting time is treated as negligible, and cloud computing latency is derived by convolving waiting and processor-service distributions.The processor service time remains exponential with rate μ_c.

A. Radio Access Network

The RAN latency model derives transmission-time variability from wireless channel conditions and queueing delay from arrival characteristics. It uses different queue models for Poisson-like and variable input traffic.

  • A. Radio Access Network: Under low load, RAN waiting time is treated as negligible, and total RAN latency is obtained by convolving transmission and waiting-time distributions.The RAN latency is T_RAN = T_S + T_W.
  • A. Radio Access Network: Under Rayleigh fading, instantaneous SNR γ is exponential with mean γ̅, and Shannon theory determines the achievable RAN rate.The normalized rate is proportional to log2(1 + γ), linking channel quality to transmission time.
  • A. Radio Access Network: Transmission-time distribution is heavy tailed and has higher variance under poor channel conditions, characterized by low γ̅.This distribution represents RAN service time for a packet of size L.
  • A. Radio Access Network: The RAN uses an M/G/1 queue for Poisson arrivals and a GI/G/1 queue when inter-arrival times are variable.The service-time distribution is determined by wireless channel dynamics.
  • A. Radio Access Network: For GI/G/1 traffic, Kingman’s approximation estimates mean waiting time using the squared variation coefficients of arrivals and transmission times.The waiting-time distribution is then approximated by an exponential PDF.

B. Transport network

The transport network latency model separates propagation from transit latency and represents transit through queueing models chosen according to traffic and service-time variability. It combines link-level latency distributions across the network path.

  • Transport latency comprises propagation latency and transit latency, with transit including transmission and queuing times.Propagation depends on total distance and propagation speed.
  • Uplink traffic is modeled across a series of M/M/1 queues with Poisson arrivals justified by aggregation of many independent flows.The Palm–Khintchine theorem supports approximating aggregated gNB traffic as Poisson.
  • Each link’s service rate depends on the allocated fraction of link capacity and the service data rate.
  • Downlink traffic aggregated from multiple UPFs is modeled at the first transport link as a GI/M/1 queue because its arrivals are non-Poisson.Arrivals use a general inter-arrival distribution, while service remains exponential.
  • Subsequent transport links use M/M/1 queues, and their transit latency distributions are combined through convolution.The first-link transit latency is formed by convolving waiting and service-time distributions.

C. Core Network

The core network model represents serial nodes with deterministic service and selects arrival models according to traffic direction. It derives per-node and aggregate transit latency distributions from these queueing assumptions.

  • The core network is modeled as n_CN serial nodes, each represented by a deterministic-service queue.Propagation latency is determined by core-network distance and propagation speed.
  • For uplink traffic, Poisson arrivals yield an M/D/1 model at each core-network node.The service time is deterministic with fixed service rate μ_i^CN.
  • Each M/D/1 node’s expected transit latency includes both deterministic service time and queue waiting time.
  • Uplink transit across multiple core-network nodes is represented by a shifted Gamma (Erlang) distribution.The aggregate latency is obtained by convolving the node-level transit distributions.
  • Downlink Internet traffic uses a G/D/1 model because inter-arrival times are highly variable and described by a Weibull distribution.The service time remains deterministic, and mean waiting time is obtained from a Pollaczek–Khinchine–type expression.

D. Internet

Internet latency is defined between the core-network UPF and the cloud server and modeled empirically as a Weibull-distributed round-trip latency. The model incorporates transmission, propagation, and intermediate-router or switch queuing times.

  • Internet latency is the latency between the core network UPF and the cloud server in both uplink and downlink directions.
  • The Internet-latency model uses a Weibull distribution fitted to measured round-trip-time data between Internet-node pairs.The measured latency includes transmission, propagation, and queuing at intermediate routers or switches.
  • The fitted Weibull distribution has scale parameter 0.0175 and shape parameter 1.87.

VI. MODEL VALIDATION

The validation benchmarks communication and computing latency models against available measurements from the literature. The authors identify the scarcity of end-to-end measurements for the considered continuum architecture as a limitation.

  • The proposed communication and computing latency models are benchmarked against corresponding literature measurements whenever available.
  • Scarcity of empirical measurements for the considered end-to-end continuum architecture limits direct validation coverage.

A. Communication latency

The paper models communication and computing latency across the IoT-edge-cloud continuum, including stochastic jitter and traffic variability. Its component and end-to-end estimates are compared with empirical latency ranges, while local, edge, and cloud execution are evaluated under specified network and processing conditions.

  • Communication latency validation: 1–6 ms average and 2–10 ms 99th-percentile RAN latency ranges from empirical 5G NR measurements are consistent with the model estimates.The measurements were obtained under challenging industrial NLOS propagation conditions and across different TDD configurations.
  • Communication latency validation: 0.402–10.279 ms transport-network latencies reported across 5G deployment scenarios and capacity allocations support the model’s transport-latency estimates.Measurements from Vodafone’s live London 5G macro network ranged from 1.68 to 5.6 ms depending on routing configuration.
  • Communication latency validation: 2, 2.001, and 2.001 ms are the model’s average, 90th-percentile, and 99th-percentile core-network latencies, respectively, with propagation dominating.The modeled optical core spans 200 km and produces 2 ms uplink-and-downlink propagation latency.
  • End-to-end validation: 7.8–20 ms empirical edge and 53.8–150 ms empirical cloud end-to-end latency ranges are consistent with the corresponding model estimates.The edge measurements used 5G NR NSA with MEC servers co-located at gNBs and dedicated high-capacity fiber links.

VIII. IMPACT OF ARRIVAL-TIME JITTER

Arrival-time jitter reduces deadline satisfaction and increases tail latency, with stronger effects for stringent deadlines, high computing demands, processor workloads, and offloaded execution.

  • Jitter increasingly reduces deadline satisfaction for shorter-deadline services, with cloud execution most sensitive because communication latency consumes available budget.
  • The 99th percentile latency rises with jitter across execution locations, most strongly for higher computing demands because irregular arrivals increase queueing and waiting times.
  • Tail-latency growth does not necessarily reduce deadline satisfaction for relaxed deadlines, but stringent requirements make smaller increases consequential.
  • Under higher processor workloads, jitter more strongly degrades deadline satisfaction for stringent services by increasing processor queueing delays, while relaxed deadlines are only marginally affected.
  • Edge offloading can outperform local execution for relaxed-deadline services under high local workload, whereas cloud execution becomes substantially worse as link quality declines or jitter increases.
  • For a 30 ms deadline and 6 Mcycles demand, edge offloading provides no benefit under increasing jitter, and cloud execution performs worse than local execution.

IX. IMPACT OF TRAFFIC VARIABILITY

Traffic variability affects deadline satisfaction and tail latency similarly to jitter but generally more severely, especially for stringent services, higher computing demands, and cloud execution.

  • As traffic variability increases, the probability of meeting latency deadlines decreases, generally more strongly than under jitter.
  • Traffic variability increases 99th-percentile latency, with the largest increases for cloud-offloaded services and higher computing demands.
  • Increasing traffic variability produces stronger deadline degradation for stringent latency requirements and higher computing demands across processor workloads, cellular conditions, and execution locations.
  • Relaxed-deadline services are more resilient locally or at the edge despite increased average and tail latency because larger latency budgets accommodate uncertainty.
  • Edge offloading is particularly beneficial under high local processor workload and good cellular quality, while cloud execution is more sensitive to traffic variability because communication latency further reduces the latency budget.
Loading 2608.25658v1…