Source-linked AI summary
The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem
Elioth Sanabria
TL;DR
LLM providers respond to compute congestion by degrading service, but query-level accounting misses retries and churn caused by unsatisfactory answers. The paper models this feedback with classical capacity, queueing, and allocation primitives, finding that degradation can worsen capacity use and trigger unstable regimes. Its fluid approximation leaves exact tail behavior and delayed retries for future work.
Problem
The paper asks how congestion-driven quality degradation should be priced when failed answers cause retries and customer churn, rather than merely reducing per-query inference cost.
Method
The paper combines a newsvendor with churned lifetime value, a geometric retry multiplier, transient queue dynamics, heterogeneous-customer allocation, and a shadow-price dual.
Results
The analysis identifies measurable static regimes where cheaper models save energy but consume more capacity, plus ignition and latch conditions under which reactive throttling sustains or permanently worsens degradation.
Takeaways & Limitations
Under congestion, throttling is a demand lever: allocation should account for retry-inflated load, customer heterogeneity, and the shadow price of intelligence.
Takeaways & Limitations
The fluid first-order approximation does not provide exact tail certification in the critical band or model delayed retries and heavy-tailed token distributions.
Abstract
from arXiv · showhide
Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context. The industry's accounting says this saves money. We show the accounting is wrong, because it prices a query when the customer buys an answer. A degraded answer fails with some probability, and a failed answer either returns as a retry, inflating arrivals when the system is most loaded, or departs as churn, destroying lifetime value on a ledger no cost dashboard displays. We model inference allocation with three classical primitives: a newsvendor whose stockout cost is churned lifetime value, a geometric retry multiplier in which the recycled product is dissatisfaction, and a two-regime transient queue whose arrival rate is made endogenous by retries. Statically, there is a nonempty, measurable regime in which a cheaper model saves energy per satisfied answer while consuming strictly more capacity per satisfied answer, so the discount inverts exactly when capacity binds. Dynamically, a reactive throttle fired during a surge can cross an ignition threshold beyond which it manufactures more traffic than it sheds, and a release rule set below the degraded equilibrium converts a transient surge into a permanent degraded regime. With heterogeneous customers, throttling is a transportation problem in retry-inflated load whose optimal policy rations intelligence by critical ratio, class by class, and whose dual, the shadow price of intelligence, prices a marginal query by class and by hour; closed-form trajectories make it computable in milliseconds. Stochastic analysis sharpens rather than erodes the thesis: the ignition boundary acquires a predicted width, and noise punishes the reactive policy that parks the system against it. Under congestion, throttling is not a cost lever but a demand lever.
1 Introduction
The paper reframes congestion-driven service degradation as a supply-chain and operations problem: cheaper per-query inference can increase effective demand and destroy customer value. It develops classical models showing when degradation saves energy, worsens capacity use, triggers unstable feedback, and requires class-specific allocation.
- Motivation: Providers degrade models, reasoning budgets, or context under congestion because the immediate effect is lower energy, memory, and server-time use.The paper argues this first-order accounting misses quality failures and their downstream effects.
- Motivation: Unsatisfactory answers either return as retries that inflate arrivals or lead to churn that destroys lifetime value.The resulting throttle response can initially reduce cost, then increase it, while permanent customer loss remains absent from ordinary cost dashboards.
- Approach: The analysis combines a newsvendor with churned lifetime value as stockout cost, a geometric retry multiplier, and a transient finite-server queue with endogenous arrivals.These primitives connect capacity allocation, dissatisfaction, and congestion dynamics.
- Results: Static tier comparisons produce separate energy and capacity frontiers, with a measurable interval where a cheaper tier saves energy per satisfied answer but consumes more server time per satisfied answer.The paper presents this inversion as falsifiable using service times, power draws, and logged traffic.
- Results: Reactive degradation can cross an ignition threshold, while a release threshold below degraded equilibrium can turn a transient surge into a permanent degraded regime.With heterogeneous customers, optimal throttling becomes a transportation problem in retry-inflated load, with a class- and hour-specific shadow price of intelligence.
- Related Literature: The paper positions its contribution at the intersection of retrial and abandonment queues, quality-dependent demand, supply-network amplification, and LLM serving allocation.Its distinctive feedback is completion-driven: the provider controls the quality that determines retry behavior.
2 The Model
The model represents LLM serving as tiered capacity allocation with measurable demand, service, energy, quality, retry, and churn primitives. It converts posted tier properties into retry-adjusted effective quantities and uses fluid dynamics while explicitly fencing off delayed retries, tails, and other realism limits.
- Technology and Cost: The fleet has fixed servers and concurrent inference slots, while its operating cost combines electricity, resident-context memory overhead, and a convex service-level penalty.The simplified technology captures energy, memory, and latency costs as backlog grows.
- Tier Ladder: Each tier is characterized by active-slot power draw, mean service time, and dissatisfaction probability, with weaker tiers using less power and time but producing more dissatisfaction.Tier 0 is the strongest tier, and comparisons use retry-adjusted rather than posted quantities.
- Feedback Loop: An unsatisfied user retries with probability ρ or abandons with probability 1 −ρ, with abandonment generating expected lifetime-value loss prℓ.The retry loop makes attempts behind one satisfied answer geometric, while churn is charged separately from electricity.
- Demand: Demand varies by hour, and the operational forecast is the conditional expectation rt = E[Dt | Ht] given observable history; the analysis requires its availability and error variance.The forecasting machinery itself is not essential to the subsequent model.
- Static Case: The static capacity problem is a newsvendor in effective slot-time, where under-provisioning intelligence is treated as a stockout whose cost is churn rather than a lost sale.The resulting critical ratio incorporates lifetime value and electricity through the churn charge prℓ.
- Scope: The model uses fluid dynamics for tractability, with square-root buffers and simulation for tails, while assuming admission-time retries, exponential service, monotone dissatisfaction, two theoretical customer classes, and permanent linear churn.Delayed retries, distribution-dependent dynamics, heavy-tailed tokens, and win-back dynamics remain outside the model’s declared scope.
3 Analysis
The paper shows that inference degradation must be evaluated per satisfied answer, because retries and churn can reverse apparent energy savings and destabilize queues. Static frontiers, transient dynamics, and customer heterogeneity yield measurable conditions for when degradation worsens effective capacity and for whom intelligence should be rationed.
- 3.1 The Two Frontiers: The energy and capacity frontiers differ, creating a nonempty measurable region where a weak tier is cheaper per satisfied answer but consumes more server time.The interval (1, ∆w0/∆wj) is nonempty because the weak tier draws less power; within it, the energy condition holds while the capacity condition fails.
- 3.1 The Two Frontiers: 115.4 slot-milliseconds versus 100: the distilled tier saves 42% of electricity per satisfied answer while consuming 15% more capacity.Its retry multiplier is M1 ≈1.92, placing the effective service-time ratio R = 1.15 inside the trap region.
- 3.2 The Dynamics of a Poorly Timed Throttle: 57.3 thousand effective arrivals per second exceed 50 thousand throughput, so throttling grows the queue by 7.3 thousand jobs per second instead of draining it.The degraded tier delivers only θ1 = 26 thousand effective throughput after retries, and the service-level penalty rises quadratically with backlog.
- 3.3 Who to Throttle: Customer Heterogeneity: Anticipatory heterogeneous routing keeps load below capacity by degrading only insensitive traffic, while sensitive traffic remains on the strong tier.In the example, the insensitive 30% yields total effective occupancy of 2.94 < 3 thousand slots, so the system never saturates.
- 3.2 The Dynamics of a Poorly Timed Throttle: A saturated queue drains under tier j if and only if fresh demand is below its effective throughput; above that ignition threshold, retries sustain positive drift and unbounded backlog and churn.This is the structural condition behind the reactive throttle’s self-sustaining congestion.
- 3.2 The Dynamics of a Poorly Timed Throttle: A release rule never fires when the degraded equilibrium exceeds its release level, converting a temporary surge into a persistent degraded regime.If the degraded equilibrium lies below the release line, the trajectory crosses it in finite time and release is final; otherwise it remains above the line forever.
4 From Fluid to Practice
The paper addresses four deployment questions: how to solve the dynamic program efficiently, what fluid approximations discard, how to make the optimal policy operable, and how to identify the unobserved behavioral primitive. Closed-form transient dynamics make computation fast, while stochastic analysis, policy distillation, and log-based identification define the practical boundaries and uses of the framework.
- Solving It: The dynamic program uses backward induction over T epochs and a discretized class-backlog grid, with static ledger constraints pruning admissible assignments before dynamics begin.The admissible menu can collapse from 2744 to 72 assignments in reported instances.
- Solving It: Explicit integration requires µmax∆/c substeps per epoch because stability and accuracy must resolve the fastest relaxation rate, making fast tiers costly to integrate.The closed-form alternative follows a linear leg and an exponential leg to a closed-form crossing time, so transition cost is independent of the service-rate scale.
- Solving It: The closed-form two-regime recursion retains transient trajectories while replacing time discretization with formula-based crossing times and constant-cost transitions.It sits between pointwise stationary approximations and exact transient analyses that require numerical integration.
- Role of the Variance: In the multi-class extension, frozen capacity shares are re-resolved at leg boundaries, producing first-order error in leg length concentrated in deep saturation with unbalanced backlogs.The policy-value certificate remains within single digits of percent on every reported instance, although Q-regret tails occur off the optimal trajectory.
- Role of the Variance: The fluid model discards fluctuations, so it is reliable for ordering and regional conclusions outside the critical band but cannot certify service-level tails near marginal stability.Retry cascades create overdispersed arrivals whose variance grows near the ignition threshold, making reactive policies that park near the boundary especially vulnerable.
- Policy as a Tree: The optimal heatmap can be distilled into an occupancy-weighted decision tree, and under nested thresholds an ϵ-exact tree has O(|X| × #regimes × J) leaves.The tree is presented as the policy’s normal form rather than merely a compression.
- Identifying the Price of the Discount: Most model primitives come from benchmarks, meters, documentation, and bills; retry fingerprints in provider logs directly estimate the product d_jρ_x that drives the model’s effective quantities.Separating dissatisfaction from retry probability is needed only for the churn column.
5 Numerical examples
Five calibrated provider instances show that effective-dominance pruning sharply reduces policy menus and that the fluid DP is consistently cheaper and feasible under congestion. The experiments also reveal provider-specific mechanisms, including surge anticipation, retry-driven instability, class-dependent dual prices, and compact policies close to the DP.
- 5.1 Instances, menus, and the effective-dominance collapse: Public benchmark data, provider documentation, and calibrated behavioral primitives define five provider instances with heterogeneous customer classes and a shared diurnal demand profile.The profile peaks at 0.90 of θmix and includes an off-window overnight surge exceeding the comfortable mix while remaining survivable near maximum throughput.
- 5.1 Instances, menus, and the effective-dominance collapse: Effective-dominance pruning reduces Anthropic’s 125 actions to 9, DeepSeek’s 27 to 4, and OpenAI’s 2744 to 72.Kimi collapses to one all-strong action because its cheap tiers are no faster and retries make degradation a pure loss.
- 5.3 Results: The fluid DP is simultaneously the cheapest policy and feasible at every hour, while always-strong and heuristic policies can become unstable, violate latency, or incur much higher churn-adjusted cost.The tuned stump costs $648.95M versus $9.38M for the DP on Anthropic, with churned lifetime value of $57.20M versus $4.26M.
- 5.3 Results: 69× on Anthropic, 10× on OpenAI, 371× on Google, and 24× on DeepSeek are the DP’s margins over the best feasible baseline.The margins vary with regime: Google’s gap is driven by surge timing, whereas Kimi is a null result because its optimal action is all-strong.
- 5.3 Results: Google’s DP pre-positions before the 03:00 surge and keeps peak backlog orders of magnitude below baselines, while Kimi is harmed by degradation during the surge.Google’s static-optimal mix is all-strong off-surge, making degradation valuable there for throughput rather than energy; Kimi’s maximum-throughput action is all-strong.
- 5.4 The dual, computed: Figure 7 prices marginal queries by class and hour, with class-specific direct floors and a near-common congestion rent when capacity binds.On Anthropic, the specialized off-peak dual is 3.9× the casual dual, while DeepSeek and Kimi lack switch structure because they hold the static-optimal action all day.
- 5.5 Certificates: The fluid solver runs each instance in under three seconds and remains ε-close in rollout value to the converged competitor.The closed-form forward pass also matches the Euler simulation, supporting the runtime and policy-value certificates.
- 5.6 The policy in normal form: five trees: Depth-three policy trees use at most eight leaves, with certified rollout gaps of 0.54% and 0.20% on the shown instances.The trees expose missing forecast, backlog, peak-flag, and class-specific splits; for example, OpenAI switches from Luna (max) to Sol (medium) across the capacity frontier.
6 Concluding Remarks
The paper frames congestion control as an accounting problem: customers buy answers, while systems price queries, and retry feedback connects the two. Its first-order framework derives static tier traps, dynamic ignition and latch behavior, heterogeneous-class rationing, and dual prices, while identifying delayed retries and heavy-tailed tokens as the boundary requiring a sequel.
- 6 Concluding Remarks: A geometric retry multiplier bridges query-level accounting and answer-level outcomes, producing tier-economics inversion, demand amplification, and class- and hour-specific marginal prices.The framework combines static allocation, transient dynamics, heterogeneous throttling, and a dual variable for marginal intelligence.
- 6 Concluding Remarks: Delayed retries make the orbit a state variable and turn the dynamics distribution-dependent, requiring a nonlinear Markov-chain formulation for exact tail certification and heavy-tailed token distributions.These effects mark the boundary of the paper’s first-order approximation and motivate treating the sharper tail analysis as future work.
A.1 Proof of Proposition 5
Proposition 5 compares anticipatory and reactive degradation during a feasible surge. Anticipation avoids ignition but may pay linear churn, whereas reactive degradation incurs cubic service-level costs and can become permanently costly.
- Anticipatory path: A feasible anticipatory path starts degradation before the surge, keeps retry-inflated load below capacity, and never saturates.Its service-level cost is zero, while energy savings equal γ(Ŝ0∆w0 − Ŝ1∆w1) per unit of degraded volume.
- Reactive path: The reactive trajectory crosses capacity before the surge ends and then accumulates backlog with constant positive drift θ := r_surge − θ1.The crossing occurs at some τ < t2 when the surge rate exceeds the degraded throughput.
- Reactive path: Reactive service-level cost grows at least cubically with surge duration, as backlog increases linearly after ignition.The in-surge lower bound is proportional to CSLAθ^3(mk)^2(t2 − τ)^3.
- Post-surge dynamics: After the surge, excess backlog adds another cubic cost when calm demand is below degraded throughput; otherwise, the degraded queue never drains.For r_calm ≥ θ1, reactive cost is unbounded regardless of surge length.
- Comparison: Anticipation pays linear churn over its lead window, while reactive excess grows cubically, producing a critical surge duration beyond which anticipation dominates.The comparison equates the reactive cubic excess with the anticipatory churn charge P.
A.2 Proof of Proposition 6
Proposition 6 solves constrained heterogeneous routing through a capacity shadow price. The optimal policy orders degradations by damage per unit of capacity relief and becomes nested as congestion increases.
- Shadow-price formulation: A multiplier ν on the binding capacity constraint separates the routing problem by customer class and prices marginal tier changes.Moving class-x traffic between tiers changes the objective by λx ∂c(x,j)/∂j and relaxes capacity according to the class’s capacity relief.
- Capacity relief: At the margin against tier 0, capacity relief from tier j is Ŝ0(x) − Ŝj(x).This relief is the denominator in the class-specific damage-to-relief index.
- Optimal ordering: Optimal routing degrades classes with positive relief in increasing order of the damage-to-relief ratio I(x,j).Classes with nonpositive relief are never degraded at a binding constraint because doing so tightens capacity.
- Congestion monotonicity: The degradation set is nested in congestion because ν is nondecreasing in backlog N.As N rises, classes satisfying I(x,j) ≤ ν form a nondecreasing set.
A.3 Proof of Proposition 8
Proposition 8 shows that the optimal state-dependent routing map can be represented exactly by a compact decision tree organized around regime, class, and backlog thresholds.
- Policy structure: The optimal action depends on the time regime, the class’s position in the ordering, and backlog N relative to class-specific shadow-price thresholds.The regime determines the arrival rate entering ν; the remaining splits identify the appropriate tier action.
- Tree representation: A tree splitting first on regime, then class, then backlog thresholds reproduces the optimal routing map exactly.The construction uses at most #regimes × |X| × J leaves, with fewer leaves available for ε-exact approximations.
B Additional numerical assets
The numerical assets visualize distilled class-specific routing alongside reactive, scheduled, and always-strongest baselines. Reported rollout certificates show zero value gap to the dynamic program for the displayed distilled trees.
- Casual class: For casual traffic, the displayed strongest-to-weakest ladder is 3.7 Flash high, 3.7 Flash medium, then 3.7 Flash low.The asset states that casual traffic is served at every hour by the fluid dynamic program and its distilled policy.
- Shared-pool routing: The Google figure uses the same distilled-tree and baseline format, with trees allowed to split on other classes’ backlogs in the shared pool.Shared capacity means a backlog in one class can affect routing decisions for another class.
- Specialized class: For specialized traffic, the displayed ladder contains V4-Pro-0813 thinking followed by V4-Flash-0731 thinking.The corresponding asset identifies specialized traffic as served at every hour by the fluid dynamic program and distilled policy.
- Routing figures: The Kimi, DeepSeek, and Google figures compare one distilled routing tree per class with reactive, scheduled, and always-strongest baselines.Reactive degradation uses a backlog trigger and releases at 0.5 C; scheduled degradation follows published windows, while always strongest never degrades.