Source-linked AI summary
Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief
Meiyi Li
TL;DR
Fixed flexibility percentages do not reveal how much data-center load remains eligible across durations, reliability levels, or cluster portfolios. Using a 185-day, 155,410-GPU production trace, the paper reconstructs power and derives a workload-semantic flexibility envelope. It finds that dependable relief declines with duration, scalar assumptions misestimate the surface, and scheduler delay adds essentially no dependable capacity, motivating contract specifications based on explicit dimensions.
Problem
Grid studies lack public production evidence showing how eligible data-center load persists across event durations and co-moves across clusters.
Method
The paper reconstructs power from a hyperscale GPU trace and evaluates workload-semantic eligibility across duration, reliability, portfolios, and scheduler diagnostics.
Results
95%-available relief falls from 2.51 MW at one hour to 1.95 MW at 24 hours, while observed scheduling delay contributes zero 95%-available backlog capacity.
Takeaways & Limitations
Flexibility contracts should specify duration, reliability, portfolio, realizable response, recovery, and recalibration rather than one percentage of site load.
Takeaways & Limitations
The trace identifies historical eligibility but not verified delivered response because it lacks checkpoint, service-level, restart, network, and control-latency measurements.
Abstract
from arXiv · showhide
Grid studies often represent data-center flexibility as a fixed percentage of load, although no public production trace has shown how much eligible load persists across event durations or co-moves across clusters. We reconstruct 4,439 hourly power observations from a 185-day trace of 155,410 GPUs and derive a workload-semantic flexibility envelope. The fleet's time-averaged Monte Carlo median facility demand is 55.8 MW, while immediate eligible curtailment averages 3.55 MW after retaining allocated-GPU idle power: 12.1% of workload power and 6.35% of median facility power. Under full realization of that eligibility, 95%-available relief falls from 2.51 MW for one hour to 2.32 MW for four hours and 1.95 MW for 24 hours; a common realizable fraction q scales every value exactly by q. A mean-calibrated scalar overstates these quantities by 17%, 25%, and 47%, while a scalar tail-calibrated at four hours understates the one-hour product by 6% and overstates the 24-hour product by 17%; the share that reproduces the surface varies by a factor of 1.6 across durations and reliability levels. Aggregating 13 clusters raises four-hour firmness from 0.38 to 0.66, but cross-cluster covariance limits the gain. The production scheduler exposes almost no additional delay-based capacity: newly deferrable arrivals average 0.008 MW and have zero 95%-available capacity. These results replace an assumed flexibility percentage with duration, reliability, portfolio, and realizability terms that can be written into interconnection and demand-response contracts.
1 Introduction
Data-center flexibility is increasingly relevant to grid interconnection, but fixed percentages do not show how much load can be released for a specified duration and reliability. This paper uses a hyperscale production trace to derive a workload-semantic flexibility envelope and compare it with scalar approximations.
- Motivation: 176 TWh of U.S. data-center electricity use in 2023 motivates better evidence for managed load reduction.Data centers represented 4.4% of national electricity, with projections of 6.7–12% by 2028.
- Research gap: A scalar flexibility share cannot express persistence across event durations, co-movement across clusters, or flexibility already spent by the scheduler.Existing studies use fixed, shiftable, interruptible, or whole-site abstractions despite controlled evidence that some response is technically possible.
- Data: 155,410 GPUs across 17 clusters and 185 days provide production workload structure for reconstructing workload, IT, and facility power.The trace includes workload type, priority, standby state, and scheduling delay for 47% of execution spans.
- Approach: The paper converts scheduler records into contract variables using workload-conditioned power curves, a semantic flexibility envelope, and a duration–reliability–portfolio surface.The approach also tests fixed-share, time-shuffled, and independent-cluster planning representations.
- Headline results: 3.55 MW of immediate eligible curtailment equals 12.1% of workload power but only 6.35% of median facility power.These figures reject a universal flexibility percentage as a complete description of dependable relief.
- Headline results: 95%-available relief declines from 2.51 MW for one hour to 1.95 MW for 24 hours when every eligible watt is realizable.A common realizable fraction q scales the entire surface exactly by q.
2 Trace, Power Reconstruction, and Metrics
The paper reconstructs power from workload and hardware records, defines eligibility conservatively, and measures dependable relief across duration, reliability, portfolios, and scheduling effects. Its metrics preserve persistence and covariance that scalar or independent-cluster models discard.
- Trace and reconstruction: 4,439 valid hourly observations are reconstructed from aggregate GPU, workload, host, and accelerator records over the trace.The aggregation stores sufficient statistics so piecewise-linear power curves can be evaluated without rescanning the pod table.
- Trace and reconstruction: Workload-conditioned curves separate training, online inference, offline inference, and other work while accounting for idle, host, IT, and facility power.Online inference receives a 0.43-T total-power floor at the central setting, and facility power uses a 1.2 central PUE.
- Flexibility envelope: Eligibility assigns low-priority rows to an eligible-attributed layer, while shift power remains a semantic pool rather than offered capacity.The shift layer may require checkpointing, migration, or restart before it can be delivered.
- Flexibility envelope: The immediate curtailment boundary removes active GPU power but retains allocated-GPU idle power.Removing idle power requires node sleep or reassignment, so facility and workload denominators are reported separately.
- Dependability metrics: K measures lower-tail eligible power over historical event windows for a chosen duration, reliability, portfolio, and realizable fraction q.Positive homogeneity means scaling eligibility by q scales every K value exactly by q.
- Counterfactuals: Mean-calibrated, time-shuffled, independent-cluster, and tail-calibrated representations expose what scalar assumptions discard.Rolling-origin calibration reports absolute-MW out-of-sample derating for four- and 24-hour products.
- Portfolio aggregation: Fluctuation scaling and synchrony quantify how aggregation changes variability and whether covariance violates independence assumptions.The analysis compares cluster–type–priority, cluster–type, and cluster units, with R=1 as the zero-covariance benchmark.
- Scheduler diagnostics: The scheduler diagnostic measures observed temporal displacement and separately tests newly deferrable arrivals under a stricter backlog replay.It is a diagnostic of exercised delay, not a feasible alternative schedule.
3 Results
The reconstructed fleet is large, growing, and persistently loaded, but its eligible relief declines with event duration and is constrained by cross-cluster covariance, realizability, and limited scheduler delay. These results show why flexibility must be specified by duration, reliability, portfolio, and operational conditions rather than one scalar share.
- Fleet demand: 55.83 MW is the time-averaged Monte Carlo median facility demand across 4,439 hours, with a 38.84% increase from the first to last week.The median series has a 0.820 load factor.
- Fleet demand: 3.42 MW is the average within-day peak-to-trough swing, while the clock-hour profile troughs at hour 06 and peaks at hour 17.Raw autocorrelation is 0.950 at 24 hours and 0.846 at seven days, although both include fleet growth.
- Aggregation and covariance: 0.470 to 0.662 is the rise in fluctuation-scaling exponent across increasingly coarse aggregation levels, but overlapping intervals make it a point-estimate pattern rather than an established difference.The exponents correspond to cluster–type–priority units, cluster–type units, and clusters.
- Aggregation and covariance: 1.758 is the cluster-level synchrony ratio, and approximately two thirds of cross-cluster excess covariance remains unidentified after calendar and capacity normalization.The common ten-cluster ratio is 1.753, while calendar removal and capacity normalization explain 14.6% and 34.4% cumulatively.
- Eligible workload: 3.55 MW is the immediate eligible curtailment after retaining allocated-GPU idle draw, equal to 12.08% of workload power and 6.35% of median facility load.The corresponding hourly P10 is 2.71 MW; the semantic eligible layer would be 5.99 MW only if idle draw were also removed.
- Duration and portfolio: 2.51 MW, 2.32 MW, and 1.95 MW are the 95%-available eligible-workload values for one-, four-, and 24-hour durations under q=1.Across 13 clusters, four-hour firmness rises from 0.378 for one cluster to 0.656 for all 13, below the independent-cluster counterfactual of 0.729.
- Scalar comparison: 17.4%, 25.5%, and 46.6% are the mean-calibrated scalar's overstatements at one, four, and 24 hours, while tail calibration at four hours makes the one-hour product 6.4% low and the 24-hour product 16.9% high.The reproducing workload share is 0.85, 0.79, and 0.68 of the mean share at 95% availability for one, four, and 24 hours.
- Realizability: q scales every eligible-relief value exactly, so the four-hour, 95%-available offer is 2.3168q MW and q=0.432 is required to offer 1 MW.Field preemption, checkpoint, service-level, and telemetry tests are needed to estimate q.
4 Related Work
Prior work models data-center flexibility through scheduler studies, power models, and fixed load tiers, but rarely provides production-wide workload evidence. This paper addresses that gap with a hyperscale trace while preserving the broader planning context.
- Power modeling: Power-modeling research spans device telemetry, fleet models, GPU-state measurements, and compositional generators that preserve energy and autocorrelation.The present study trades device-scale dynamics for six months of workload semantics and cluster hierarchy.
- Scheduler-aware flexibility: Scheduler-aware studies move compute across time or sites but rarely expose a production-wide denominator.Related work includes carbon-aware and virtual-capacity systems.
- Experimental and trace evidence: A controlled 256-GPU demonstration verified multi-hour response, while other studies measured power-capping capability before scaling it through simulation.Trace-based studies also estimate delay flexibility from older or non-production traces.
- Grid-planning context: Grid-planning studies show that flexible-load benefits depend on location, event shape, and the assumed division among fixed, shiftable, and interruptible tiers.One screening study instead treats a large load as fully curtailable for limited hours.
5 Discussion and Limitations
The proposed duration–reliability surface turns historical eligibility into a compact contract primitive while making realizability, covariance, temporal validity, and external validity explicit. Its planning claims remain bounded by an illustrative screen, a single-operator trace, and hourly modeled power.
- Contract representation: A duration–reliability surface can specify event duration, target availability, realizable fraction q, and rolling-calibration derating.The representation replaces a single percentage with five durations, four reliability levels, and one response factor.
- Realizability: Historical eligibility is not verified response because the trace omits checkpoint completion, service-level violations, restart energy, network bottlenecks, and control latency.Until field trials estimate q by workload class and horizon, the envelope is an upper bound on delivered power.
- Portfolio covariance: Positive residual covariance limits portfolio diversification, and the independent-cluster counterfactual overstates four-hour relief by 11% and 24-hour relief by 20%.Calendar removal and capacity normalization explain only 34.4% of excess variance, leaving the covariance mechanism unidentified.
- Planning boundary: The MISO South calculation is an illustrative peak-cap screen rather than a network study.Transmission, reserves, contingencies, forecast error, and generator deliverability remain outside the screen.
- External validity: A single-operator trace limits external validity because workload labels, priority policy, hardware mix, and cluster coordination may differ elsewhere or over time.The authors release per-cluster series and mappings and emphasize replication across operators.
- Temporal and power-model validity: Hourly utilization and modeled power cannot test sub-hour response, short training oscillations, or minimum power during a 15-minute settlement interval.Some public measurements are suggestive, while host-model parameters remain independently unvalidated.
- Construct validity: Semantic mapping assumes every low-priority pod is preemptible, although the trace does not record whether individual pods checkpoint successfully.Alternative mappings and power re-anchoring move the eligible share across reported ranges, and only 47.1% of spans match delay records over the full trace.
6 Artifact Availability
The authors identify the source trace and publish derived data products for reproducibility. The raw trace remains externally distributed rather than redistributed with the paper.
- Data and artifacts: The raw Alibaba cluster-trace-gpu-v2026 is distributed by its authors, while derived data products are public in tagged release v1.0.1 and archived at Zenodo.Appendix E regional demand uses the U.S. Energy Information Administration’s Hourly Electric Grid Monitor.
7 Conclusion
The study replaces a universal flexibility percentage with a production-derived envelope that separates eligibility from dependable delivered response. Its principal quantities vary with duration, portfolio covariance, realizability, recovery, and recalibration, making those dimensions suitable for contract specification.
- Measured eligibility: 12.1% of workload power is immediately eligible under an idle-retained boundary, but it represents only 6.35% of median facility demand.The reconstructed evidence comes from a 155,410-GPU hyperscale fleet.
- Practical implication: The practical output is a contract specification in which duration, reliability, cluster portfolio, realizable response, recovery, and recalibration determine offered megawatts.Publishing these dimensions lets planners use production evidence without treating eligibility as proven control.
A.1 Central parameters and uncertainty ranges
The power model combines measured workload anchors with Monte Carlo variation and corrects online inference for a nonzero idle-related floor. Sensitivity checks show that hourly averaging and curve anchoring materially affect levels, while the eligibility conclusion remains similar.
- Model parameters and uncertainty: 300 Monte Carlo draws vary power parameters systematically across the fleet rather than as independent hourly errors.The model varies accelerator ratings, idle and online-inference parameters, host coefficients, PUE, and curve family.
- Model parameters and uncertainty: 117 published measurements, including 92 in represented categories, support the anchor scale but not facility-specific calibration.Measurements include exact, recomputed, and figure-read observations across workload categories.
- Online-inference floor: A 0.43-TDP online-inference floor raises the must-run floor and removes low-utilization online inference from apparent curtailment.The floor represents decode-dominated power that persists when reported SM utilization is low.
- Temporal averaging: 22–25% overstatement occurs when a concave training curve is evaluated at hourly mean utilization under alternating utilization, whereas linear curves remain unbiased.This is a stress bound rather than an empirical within-hour error distribution.
- Sensitivity to curve anchors: 54.49 MW, 52.82 MW, and 55.83 MW are the measured-mid, measured-low, and base median facility powers, respectively.Their idle-retained eligibility values are 13.22%, 10.38%, and 12.08%, placing immediate eligibility near one tenth of workload power.
- Temporal checks: 91–97% of hourly mean power is typically retained at 15-minute minima across the examined workloads.Shorter windows vary more; LoRA has five-minute P10 of 0.48 and one-minute P10 of 0.14.
B Load and Aggregation Details
The reconstructed load is dominated by allocated GPU pods and retains strong temporal dependence, while cluster-composition relationships and covariance mechanisms remain only partly resolved.
- Load composition: 67.35% of central IT power comes from allocated GPU pods, 31.07% from hosts, and 1.58% from unallocated-GPU idle power.Applying one PUE preserves these shares in the central facility series.
- Composition uncertainty: 0.149 is the fitted curvature relating cluster coefficient of variation to training share, but its 90% bootstrap interval spans −1.207 to 10.007.The trace therefore does not confirm the simulated U-shaped relation from prior work.
- Cluster aggregation: 1.765 is the average raw cluster-to-fleet synchrony ratio across six plateaus, falling to 1.567 after capacity normalization.The stable-capacity test supports a portfolio correction but leaves its operational cause open.
C Scheduler and Envelope Details
The scheduler-derived delay layer is nearly negligible relative to semantic eligibility, while the envelope’s composition and cross-cluster relationships vary over time.
- Scheduler trace: 99.9999569% coverage was achieved in the day-110–184 join audit, with only 108.58 unmatched GPU-hours.The unmatched mass spans three pod-days.
- Delay-derived capacity: 0.00794 MW of one-hour eligible arrivals and 0.00353 MW of four-hour arrivals have zero 95%-available capacity.The one-hour quantity is 0.083% of the 9.61 MW semantic shift layer.
- Envelope composition: 0.227 is the median eligible-power correlation across 15 participating clusters.Standby reaches 1.08% in the last window, while the central floor ranges from 43.57% to 49.78%.
D Dependable-Availability Robustness
Dependable relief declines smoothly with duration and depends on portfolio aggregation, calibration, and realizability rather than a single fixed flexibility share. Recovery assumptions can further change the energy interpretation of curtailment.
- Duration and reliability: 0.59–0.77 is the observed range of four-hour firmness across trace periods and months.The full duration–reliability grid shows a smooth decline, and the 99% four-hour cell uses 44 tail windows.
- Realizability: 2.3168q MW is the four-hour, 95%-available relief for the support-filtered 13-cluster portfolio.Offers of 1 and 1.8 MW require q≥0.432 and q≥0.777, respectively.
- Calibration robustness: 2.546, 2.290, and 1.727 MW are rolling-origin calibrated capacities at one, four, and 24 hours under absolute-MW calibration.The corresponding mean coverages before calibration are 0.954, 0.946, and 0.907.
- Scalar representations: The 30/50/20 representation assumes a 50% flexible tranche shiftable within the day, which the trace does not support.The comparison places assumed tiers alongside trace-derived eligible GPU relief at the marginal-PUE boundary.
- Recovery and rebound: −0.125 is the net-to-gross energy ratio for a four-hour call with full rebound and a one-hour checkpoint interval.Recovery time must use the offered qK_0.95,4 rather than the larger mean eligible layer.
E Illustrative MISO South Peak-Cap Screen
The MISO South screen tests whether trace-derived curtailment can satisfy a regional cap under residual-floor and annual-call constraints. Results show that the residual floor, rather than the response-hour budget, binds in the 5%-margin trace-derived case.
- Screen setup and principal result: 1.684 GW is the noncurtailable screening load, rising to 1.809 GW with idle-retained eligibility under the main marginal-response convention.Across offsets, the trace-derived median is 1.809 GW with a 1.742–1.865 GW range, a 7.43% gain.
- Screen setup and principal result: 1.837 GW is obtained with the 1.2 average-PUE sensitivity, compared with 1.809 GW under the main convention.The corresponding gains are 9.07% and 7.43%, respectively.
- Sensitivity across representations: At zero margin, every representation with a noncurtailable floor admits zero load, whereas the fully curtailable case admits 0.94–3.09 GW depending on its hour budget.The tested margins are 0, 2, 5, and 10%, with annual response budgets of 0.25–2% of hours.
- Response-budget behavior: 87.72 hours is the nominal 0.5% annual response budget, while the trace-derived case uses a median of three response events and three invoked hours.The median longest event is two hours, with P5–P95 ranges of two–three events, two–four invoked hours, and one–two hours longest event.
- Response-budget behavior: The 5%-margin trace-derived case is slack on response hours, so the residual floor binds before the annual hour allowance.The screening rule requires both hourly residual excess to remain within curtailment and response calls to stay within budget; the budget cannot permit postresponse cap violations.
- Interpretation and scope: The load-duration and ratio views support joint residual-floor and call-budget terms rather than replacing one constraint with the other.The screen focuses on the upper tail of regional load and compares representation ratios across margins and budgets.