Paper deep dive
The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem
Elioth Sanabria
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/26/2026, 5:17:21 AM
Summary
This paper models Large Language Model (LLM) inference allocation as a supply chain problem, arguing that service degradation (throttling) during congestion is often a suboptimal strategy. The authors demonstrate that degraded answers have a higher failure probability, leading to customer retries (inflating demand) or churn (destroying lifetime value). Using a model combining newsvendor logic, geometric retry multipliers, and transient queueing theory, the paper derives the 'shadow price of intelligence'—a dual variable that prices marginal queries based on class and time. The study concludes that under congestion, throttling acts as a demand lever rather than a cost lever, potentially creating ignition thresholds where reactive throttling worsens congestion.
Entities (8)
Relation Signals (6)
Shadow Price of Intelligence → prices → Marginal Query
confidence 96% · the shadow price of intelligence, prices a marginal query by class and by hour
Elioth Sanabria → affiliatedwith → Lehigh University
confidence 95% · Affiliation: Lehigh University - College of Business
Service Degradation → causes → Retry Loop
confidence 95% · A degraded answer fails with some probability, and a failed answer either returns as a retry
Retry Loop → inflates → Demand
confidence 94% · inflating arrivals when the system is most loaded
Service Degradation → leadsto → Customer Churn
confidence 92% · a failed answer either returns as a retry... or departs as churn
Service Degradation → isnot → Cost Lever
confidence 90% · Under congestion, throttling is not a cost lever but a demand lever.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context. The industry's accounting says this saves money. We show the accounting is wrong, because it prices a query when the customer buys an answer. A degraded answer fails with some probability, and a failed answer either returns as a retry, inflating arrivals when the system is most loaded, or departs as churn, destroying lifetime value on a ledger no cost dashboard displays. We model inference allocation with three classical primitives: a newsvendor whose stockout cost is churned lifetime value, a geometric retry multiplier in which the recycled product is dissatisfaction, and a two-regime transient queue whose arrival rate is made endogenous by retries. Statically, there is a nonempty, measurable regime in which a cheaper model saves energy per satisfied answer while consuming strictly more capacity per satisfied answer, so the discount inverts exactly when capacity binds. Dynamically, a reactive throttle fired during a surge can cross an ignition threshold beyond which it manufactures more traffic than it sheds, and a release rule set below the degraded equilibrium converts a transient surge into a permanent degraded regime. With heterogeneous customers, throttling is a transportation problem in retry-inflated load whose optimal policy rations intelligence by critical ratio, class by class, and whose dual, the shadow price of intelligence, prices a marginal query by class and by hour; closed-form trajectories make it computable in milliseconds. Stochastic analysis sharpens rather than erodes the thesis: the ignition boundary acquires a predicted width, and noise punishes the reactive policy that parks the system against it. Under congestion, throttling is not a cost lever but a demand lever.
Tags
Links
- Source: https://arxiv.org/abs/2608.23986v1
- Canonical: https://arxiv.org/abs/2608.23986v1
Trouble viewing inline? Open PDF directly →
Full Text
123,705 characters extracted from source content.
Expand or collapse full text
The Shadow Price of Intelligence Quality Degradation in LLM Inference as a Supply Chain Problem Elioth Sanabria Affiliation: Department of Decision and Technology Analytics Affiliation: Lehigh University - College of Business Email: els626@lehigh.edu August 25, 2026 Abstract Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context. The industry’s accounting says this saves money. We show the accounting is wrong, because it prices a query when the customer buys an answer. A degraded answer fails with some probability, and a failed answer either returns as a retry, inflating arrivals precisely when the system is most loaded, or departs as churn, destroying lifetime value on a ledger no cost dashboard displays. We model inference allocation with three classical primitives: a newsvendor whose stockout cost is churned lifetime value, a geometric retry multiplier in which the recycled product is dissatisfaction, and a two-regime transient queue whose arrival rate is made endogenous by retries. Statically, there is a nonempty, measurable regime in which a cheaper model saves energy per satisfied answer while consuming strictly more capacity per satisfied answer, so the discount inverts exactly when capacity binds. Dynamically, a reactive throttle fired during a surge can cross an ignition threshold beyond which it manufactures more traffic than it sheds, and a release rule set below the degraded equilibrium converts a transient surge into a permanent degraded regime. With heterogeneous customers, throttling becomes a transportation problem in retry-inflated load whose optimal policy rations intelligence by critical ratio, class by class, and whose dual, the shadow price of intelligence, prices a marginal query by class and by hour; closed-form trajectories make it computable in milliseconds. Stochastic analysis sharpens rather than erodes the thesis: the ignition boundary acquires a predicted width, and noise punishes the reactive policy that parks the system against it. Under congestion, throttling is not a cost lever but a demand lever. Keywords: LLM inference; capacity allocation; service degradation; queueing; duality. 1 Introduction Allocating inflexible resources against random demand is the founding problem of operations, and its newest instance is also its most expensive: deciding, query by query, who receives how much machine intelligence. The providers themselves describe the stakes without euphemism. A 2026 Anthropic careers posting reads: “Anthropic is compute-constrained, and how we allocate that compute is one of the highest-leverage decisions we make as a company. Today, allocation choices are only loosely tied to the user outcomes we ultimately care about: retention, lifetime value, and the experience of people relying on Claude.” The distance named in that paragraph, between the allocation decision and the outcomes it exists to serve, is not a data problem and not an engineering problem. It is a modeling problem, and it is precisely the kind the classical toolkit of operations was built to close. When a language model provider is congested, it does not turn customers away. It degrades them: it routes queries to smaller or quantized models, cuts the compute budget for reasoning, truncates the context the model is allowed to read. Degradation is the industry’s universal congestion lever because its first-order effect is irresistible. A degraded query consumes less energy, less memory, and less server time, and every dashboard in the industry duly reports that degrading saves money. This paper is about the sign of that claim being wrong. The error is an accounting error, and it fits in one sentence: the customer buys an answer, but the dashboard prices a query. A degraded answer is unsatisfactory with some probability, and an unsatisfied customer does one of two things. They re-ask, in which case the query returns to the arrival stream and inflates demand at exactly the moment the system is least able to absorb it. Or they abandon, in which case a lifetime-value asset is destroyed on a ledger no cost chart displays. The true impulse response of a throttle is therefore cost up, then down, plus a permanent loss that never appears on the cost curve. Since compute is genuinely scarce, someone must receive less intelligence; the question is never whether to degrade. The question is how much, when, and, once customers differ, who. The price of these decisions is a dual variable, the shadow price of intelligence, and computing it is the destination of this paper. We reach that destination with deliberately classical equipment. The model has three primitives. The first is a capacity choice under random demand, a newsvendor in which the cost of a stockout is not a lost sale but the expected lifetime value of a customer who never returns. The second is a retry loop: because each unsatisfactory answer re-enters the queue independently, the number of attempts behind one satisfied answer is geometric, and its mean is a multiplier of exactly the kind that arises in input-output economics when a producer consumes a share of its own output, except that here the recycled product is dissatisfaction. The third is a transient fluid model of a finite-server queue, saturated and linear above capacity, exponential and self-correcting below it, with one modification that changes everything: the arrival rate is endogenous, because completions at a degraded tier feed retries back into the stream. Each ingredient is standard (Leontief, 1966; Halfin and Whitt, 1981; Whitt, 2002). What is not standard is the loop that closes over them: the failure probability driving retries and churn is not an attribute of the customer but the direct consequence of a quality decision the provider controls. The feedback is an instrument, and this paper is about how to play it. Four results follow, each a corollary of the accounting identity above. Statically, comparing model tiers per query and per satisfied answer yields two different break-even frontiers, an energy frontier and a capacity frontier, and on the nonempty interval between them a weak tier saves money per satisfied answer while consuming strictly more server time per satisfied answer; every quantity in the frontier conditions is measurable, the service times and power draws from public benchmarks and the multipliers from logged traffic, so the trap is a falsifiable statement about real systems, and inside it the wrong decision registers as a saving on every dashboard. Dynamically, switching saturated traffic to a degraded tier drains the queue if and only if fresh demand is below the tier’s effective, retry-deflated throughput; above that threshold the feedback sustains itself, the standard reactive rule raises raw capacity while worsening the congestion it was deployed to relieve, and a release threshold placed below the degraded equilibrium never releases, converting a transient surge into a permanent degraded regime. With heterogeneous customers, throttling becomes a transportation problem whose capacity constraint must be written in retry-inflated load; the optimal policy degrades classes in increasing order of marginal damage per unit of capacity relief and strictly dominates the uniform throttling that every production load balancer implements. Finally, the dual of that problem prices a marginal query by class and by hour, collapsing off peak to the electricity of a strong-tier answer and carrying, at the crunch, a scarcity rent that differs sharply across classes; the spread of that dual over a day is the “loosely tied” of the posting above, made into a number. A companion analysis prices what the fluid model discards, variance, and shows the noise strengthens the thesis: the ignition threshold becomes a boundary with a predicted width, and the prescription is the oldest formula in operations, a square-root buffer, applied to a boundary instead of a quantity. Section 5 demonstrates the entire pipeline, statics, dynamics, matching, duality, and policy distillation, on five calibrated instances, as a proof of concept rather than an empirical claim. 1.1 Related Literature Four literatures meet in this problem, and the point of contact is the same in each: a quantity that the literature treats as a primitive of the environment is, for an LLM provider, a decision. Retrials, abandonment, and rational queueing. In retrial queues, blocked customers enter an orbit and return (Falin and Templeton, 1997; Artalejo and Gómez-Corral, 2008); in the abandonment tradition descending from Erlang-A, waiting customers renege, and staffing rules are derived to control the fraction lost (Garnett et al., 2002; Zeltyn and Mandelbaum, 2005). Our customers do both, but for a different reason and at a different point in the process: they defect after service, in response to its quality, and the quality is the firm’s control. This inverts the direction of the classical analysis. Where Erlang-A asks how many servers hold abandonment below a target, we ask which quality menu holds the retry feedback below ignition. The strategic queueing literature beginning with Naor (1969) and synthesized by Hassin and Haviv (2003) endogenizes customer behavior through equilibrium reasoning about delay; our customers respond to realized quality rather than anticipated delay, which removes the fixed-point subtlety of equilibrium arrival rates but introduces a sharper one, a completion-driven feedback whose gain the firm sets tier by tier. Two closer antecedents deserve explicit mention. Queues with Bernoulli feedback (Takács, 1963) recycle a completed job with a fixed probability, and de Véricourt and Zhou (2005) route calls under a resolution probability, unresolved customers calling back and re-entering the system, failure-driven re-entry after service with routing as the decision. What separates our setting is that the re-entry probability is neither an environment primitive nor a fixed server attribute: it is the quantity the provider chooses, tier by tier and class by class, jointly priced against a churn ledger denominated in lifetime value, and it moves the arrival rate within the congestion event itself. The trap region, the ignition threshold, and the shadow-price duality all live in that joint choice, and none arises when the feedback gain is exogenous. Quality-speed trade-offs and demand-service interactions. A line of work models servers who choose a speed-quality point, with congestion penalizing slowness and errors penalizing haste (Anand et al., 2011; Hopp et al., 2007; Alizamir et al., 2013; Ata and Shneorson, 2006), and a complementary line models service quality feeding back into future demand through loyalty and word of mouth (Hall and Porteus, 2000; Gans, 2002; Aflaki and Popescu, 2014). We combine the two channels and add the one the LLM setting makes unavoidable: failed service returns to the queue immediately, within the congestion event itself, so quality degradation moves the arrival rate on the same time scale as the backlog it was meant to relieve. The resulting effective-throughput inversion (Proposition 1) and ignition threshold (Proposition 3) have, to our knowledge, no analogue in either line once the feedback gain becomes a decision variable. Amplification in supply networks. The bullwhip effect shows how individually rational hedging amplifies demand variability as it propagates through a supply network (Lee et al., 1997; Chen et al., 2000), and its modern reading is structural: the amplification is a property of the network’s feedback topology, not of the managers operating it (Sanabria, 2026). Our central result is of the same species. The retry spiral is individually rational cost-cutting that amplifies the demand it faces, it survives perfect information and instantaneous control, and its remedy, like the bullwhip’s, is architectural rather than exhortative. Server allocation, scheduling, and LLM serving. The autoscaling literature asks how many servers to run against time-varying load, trading energy against latency (Gandhi et al., 2012; Lin et al., 2013); square-root staffing and fluid approximations for transient many-server systems supply its analytical backbone (Halfin and Whitt, 1981; Mandelbaum and Massey, 1995; Whitt, 2002). The systems literature on LLM inference optimizes the serving layer itself: batching, paged attention, memory management (Kwon et al., 2023; Yu et al., 2022). We hold the fleet and the serving stack fixed and ask the question orthogonal to both: not how many servers or how to batch, but what quality to serve, to whom, and at what implicit price. The index policy of Section 3.3 is a relative of the cμcμ rule and its generalizations (Van Mieghem, 1995), with the novelty that the capacity purchased by degradation is itself deflated by the retry multiplier of the class being degraded. For the interpretable-policy question of Section 4.3 we draw on decision-tree representations of Markov decision policies (Sanabria et al., 2021), and for the distribution-dependent dynamics that delayed retries induce, a natural extension of this work, on the nonlinear-Markov formulation of Dieker et al. (2026) (QPLEX and QDP). Section 2 builds the model. Section 3 contains the analysis: the statics, the dynamics, the matching problem, and the shadow price. Section 4 confronts the four questions a deployment must answer: how the program is solved and at what computational price, what variance does to the fluid conclusions, how to make the policy legible, and how to identify the one primitive that is not directly observable. Section 5 is the proof-of-concept computation, and Section 6 concludes. 2 The Model In this section we present a stylized but rich enough modeling paradigm to capture the nuance of the problem. We start by describing the primitives that allow us to answer how much to degrade and when. We add customer heterogeneity in Section 3.3. Table 1 collects the notation once; each symbol is introduced in context below, and the derived quantities of the last block are the ones the analysis runs on. Table 1: Notation. The symbols of the first three blocks are primitives, measurable from benchmarks, serving traces, and logged traffic; the derived quantities of the last block carry the retry feedback and are the objects every result compares. symbol meaning fleet k, m, mkmk servers; concurrent slots per server; total capacity in jobs NtN_t jobs resident in the system at time t γ, w0w_0, cmc_m, κ, CSLAC_SLA, β electricity price; idle draw per server; memory cost scale and exponent; service-level penalty scale (dollars per hour) and elasticity tiers j∈0,…,Jj∈\0,…,J\ quality tier, 00 the strongest; φ a routing mix across tiers Δwj w_j, [Sj] E[S_j], μj _j power draw per active slot; mean service time; service rate 1/[Sj]1/ E[S_j] djd_j dissatisfaction probability: the answer fails its user, increasing in j customers ρ probability an unsatisfied user re-asks (retry); 1−ρ1-ρ abandons prp_r, ℓ churn probability upon abandonment; expected lifetime margin, so each abandonment books the expected LTV loss prℓp_r rtr_t fresh-demand arrival rate, the forecast [Dt∣ℋt] E[D_t H_t] derived Mj=11−djρM_j= 11-d_jρ retry multiplier: expected attempts behind one satisfied answer S~j=Mj[Sj] S_j=M_j\, E[S_j] effective service time: slot-time behind one satisfied answer θj=(1−djρ)kμjm _j=(1-d_jρ)k _jm effective throughput: satisfied answers per unit time at saturation rteffr^eff_t effective arrival rate: fresh demand plus completion-driven retries The technology and its cost. The provider operates k servers, each hosting m concurrent inference slots, for a capacity of mkmk simultaneous jobs. With N jobs resident and the fleet serving at quality tier j, the operating cost per hour is (this can be later extended for a combination of tiers, but we defer it to build the tradeoff intuition): c(k,N)=γ[kw0+NΔwj]+γcmNκ+CSLA((N−mk)+mk)β.c(k,N)\;=\;γ [k\,w_0+N\, w_j ]\;+\;γ\,c_mN^κ\;+\;C_SLA\! ( (N-mk)^+mk )^\!β. (1) The three terms are the physics of inference serving. The first is electricity at price γ: an idle floor w0w_0 per powered server plus an active draw Δwj w_j per busy slot. The second is the memory overhead of holding N attention contexts resident, superlinear with exponent κ>1κ>1 because context caches compete for a fixed pool. The third is a service-level penalty at scale CSLAC_SLA, a constant with units of dollars per hour, the penalty rate incurred when the excess backlog equals capacity itself; the penalty is convex in the relative backlog with elasticity β=2β=2, encoding latency commitments that bite quadratically once the system saturates. While simplified, these approximate the energy tradeoff of a system with N customers and k servers. The tier ladder. The fleet runs at a tier j∈0,…,Jj∈\0,…,J\, tier 00 the strongest, or a routing mix φ across tiers. Each tier carries exactly three numbers: a power draw per active slot Δwj w_j, decreasing in j because quantized and distilled models draw less; a mean service time [Sj] E[S_j], with rate μj=1/[Sj] _j=1/ E[S_j], decreasing because fewer tokens complete faster; and a dissatisfaction probability dj∈[0,1)d_j∈[0,1), increasing in j, the probability that the answer fails the user who receives it (or in an agentic workflow also mimics extended workflows for a given task, again increasing in j). No tier is exempt, the strongest included: d0≥0d_0≥ 0 is a measured quantity, not an assumption. The first two are measurable from public benchmarks and serving traces. The third is the new object, the price of the discount, and the reason the problem is interesting; its estimation is the subject of Section 4.4. The feedback loop. What does an unsatisfied user do? With probability ρ they re-ask, and the query re-enters the arrival stream indistinguishable from fresh demand. With probability 1−ρ1-ρ they abandon, and abandonment is not free: the abandoning customer churns with probability prp_r, taking an expected lifetime margin of ℓ dollars with them, so each abandonment books an expected LTV loss of prℓp_r on a ledger that is not the electricity bill.11 1 The realized loss is a random variable, ℓ if the customer churns and zero otherwise, with mean prℓp_r . Every objective in this paper is an expected cost, and expectation is linear, so replacing each random loss by its mean changes no expected total, whatever the dependence across abandonments: booking the deterministic charge prℓp_r is exact accounting for a risk-neutral provider, not an approximation. Both branches matter, but the retry branch hides a multiplier. Each attempt fails and returns with probability djρd_jρ, independently across attempts, so the number of attempts behind one satisfied answer is geometric and its expectation is22 2 ∑i≥0(djρ)i=(1−djρ)−1 _i≥ 0(d_jρ)^i=(1-d_jρ)^-1 a geometric series. Mj=11−djρ.M_j\;=\; 11-d_jρ. (2) Readers of input-output economics will recognize the form. An economy that consumes a fraction ϕφ of its own output in production does not deliver demand D; it must gross production up to D/(1−ϕ)D/(1-φ), a feedback factor generated by output re-entering its own production process (Leontief, 1966). Equation (2) is that factor with the recycled product replaced by dissatisfaction: the customer has become a link in their own supply chain. This observation, that a throttled system manufactures part of its own demand, is the seed of everything that follows. The multiplier converts what the benchmark posts into what the system experiences, and it pays to fix the notation once. Define the effective service time and the effective throughput of a tier, S~j:=Mj[Sj],θj:=mkS~j=(1−djρ)kμjm, S_j\;:=\;M_j\, E[S_j], _j\;:=\; mk S_j\;=\;(1-d_jρ)\,k _jm, (3) the slot-time behind one satisfied answer and the rate at which a saturated fleet produces satisfied answers. Tier 0 is not exempt: S~0=M0[S0] S_0=M_0\, E[S_0] carries the strong tier’s own multiplier. Every tier comparison in this paper is a comparison of effective, not posted, quantities; the posted ones are what the dashboard shows, and the wedge between the two is the paper. It is also worth noting that this mechanic is also present in agentic workflows, where a similar dissatisfaction mechanism generates the same feedback loop. The demand. Arrivals form a stochastic process with hour-of-day structure, and the operational forecast is the conditional expectation rt=[Dt∣ℋt]r_t= E[D_t H_t] given the observable history, calibrated by regression and validated out of sample.33 3 One bookkeeping consequence of d0>0d_0>0: logged arrivals already contain the incumbent tier’s retries, so the fresh-demand rate is the deconvolution rt=(1−dj(t)ρ)rtloggedr_t=(1-d_j(t)ρ)\,r_t^logged under the incumbent tier j(t)j(t); the feedback term of (7) then supplies all retries and none are double-counted. Nothing below depends on the forecasting machinery, only on the availability of rtr_t and its error variance. The static case: provisioning is a Newsvendor. Before any dynamics, observe that the one-hour capacity choice is already a complete classical problem. Let p−cp-c be the margin on a served query, coc_o the hourly cost of an idle slot, and prℓp_r the churn charge on an unserved one. Capacity q against random offered load L, demand in effective slot-time, L=DS~0L=D\, S_0, earns [(L∧q)(p−c)−(q−L)+co−(L−q)+prℓ] E[(L q)(p-c)-(q-L)^+c_o-(L-q)^+p_r ], and the first-order condition delivers the familiar quantile with an unfamiliar tenant in the numerator: q⋆=FL−1(p−c+prℓp−c+prℓ+co)↝kt⋆=1m[rtS~0+Var(Lt)Φ−1(CR)],q =F_L^-1\! ( p-c+p_r \,p-c+p_r +c_o\, ) k _t= 1m [r_t\, S_0+ Var(L_t)\; ^-1(CR) ], (4) where the Gaussian form on the right is square-root staffing at the critical ratio CRCR (Halfin and Whitt, 1981). The buffer multiplier is now an exchange rate between lifetime value and electricity, and the moral of this case deserves stating before the model gets richer: under-provisioning intelligence is a stockout, and the cost of a stockout was never the lost sale. It was the churn, prℓp_r . Scope. The model trades realism for tractability at six declared fences, each revisited where it binds: retries fold in at admission via the geometric multiplier, deferring the delayed-retry, distribution-dependent variant to a sequel on the terms of Dieker et al. (2026); design is by fluid dynamics with tails certified by simulation, and every boundary near the critical band receives a square-root buffer, never a bare fluid threshold (Section 4.2); djd_j is monotone in j and customer type is observed rather than reported, so there is no gaming, with d0d_0 set to zero in the worked examples purely to keep the arithmetic transparent, the framework itself carrying a general d0≥0d_0≥ 0 through S~0 S_0 in every result; two customer classes suffice for the theory because the class axis, not the count, is the point, the numerics of Section 5 carrying three; services are exponential, with heavy tails entering through the simulation oracle rather than the design loop; and churn is permanent and linear in prℓp_r , with no win-back dynamics. 3 Analysis The analysis proceeds in four movements, and all four are consequences of the same identity: the customer buys an answer, but the system prices a query. Statics first, because the inversion is visible before anything moves; then dynamics, where the inversion becomes a spiral; then heterogeneity, where the spiral becomes a matching problem; then duality, where the matching problem yields the number the paper is named after. 3.1 The Two Frontiers Is a weaker tier actually cheaper? The naive comparison prices a raw query. The correct comparison prices a satisfied answer, and the multiplier (2) stands between the two. Per satisfied answer, tier j consumes: S~jΔwjγ⏟energy,S~j⏟slot-time,Mjdj(1−ρ)prℓ⏟destroyed lifetime value, S_j\, w_j\,γ_energy\,, S_j_slot-time\,, M_j\,d_j\,(1-ρ)\,p_r _destroyed lifetime value\,, (5) because every satisfied answer drags MjM_j attempts behind it and sheds abandonments at rate dj(1−ρ)d_j(1-ρ) along the way, for the destroyed LTV: someone asks MjM_j times, each dissatisfactory with probability djd_j, lastly they do not retry (1−ρ)(1-ρ) and they are lost as customers with probability prp_r with an LTV ℓ . Comparing tier j against tier 00 column by column produces two break-even conditions: tier j saves energy per satisfied answer if and only if the first inequality below holds, and it relieves capacity if and only if the second does, S~jΔwj<S~0Δw0⏟energy frontier,S~j<S~0⏟capacity frontier. S_j\, w_j\;<\; S_0\, w_0_energy frontier\,, S_j\;<\; S_0_capacity frontier\,. (6) These are two different conditions, and because a weak tier by construction draws less power, the energy frontier is strictly easier to cross. The reader should pause on what the gap between them means before reading on: there is a region in which one frontier has been crossed and the other has not, and no dashboard that prices queries can tell the two sides apart. Proposition 1 (The trap). If Δwj<Δw0 w_j< w_0, the trap region 1<S~jS~0<Δw0Δwj1\;<\; S_j S_0\;<\; w_0 w_j is nonempty, and on it tier j costs less per satisfied answer in energy while consuming strictly more server time per satisfied answer than tier 00. Every quantity in (6) is measurable, the service times and power draws ([Sj],Δwj)( E[S_j], w_j) from public benchmarks for named open-weight model pairs and the multipliers from the estimation program of Section 4.4, with no exemption for the strong tier, whose d0d_0 enters through S~0 S_0; the trap is therefore a falsifiable statement about real tiers. Proof. Write R:=S~j/S~0R:= S_j/ S_0. The energy frontier reads R<Δw0/ΔwjR< w_0/ w_j and the capacity frontier reads R<1R<1. Since Δw0/Δwj>1 w_0/ w_j>1, the interval (1,Δw0/Δwj)(1,\, w_0/ w_j) is nonempty, and on it the energy condition holds while the capacity condition fails. ∎ Example 3.1 (Two tiers of the same model). A provider serves with a strong model, tier 00, drawing Δw0=4 w_0=4 kW per active slot at [S0]=100 E[S_0]=100 milliseconds per query, and a distilled sibling, tier 11, drawing Δw1=2 w_1=2 kW at [S1]=60 E[S_1]=60 milliseconds. For the traffic in question a tier-1 answer fails with probability d1=0.6d_1=0.6, a failed user re-asks with probability ρ=0.8ρ=0.8, and the strong tier with d0=0d_0=0, so S~0=[S0]=100 S_0= E[S_0]=100. Is the switch profitable? Compute the multiplier first: d1ρ=0.48d_1ρ=0.48, so M1=1/0.52≈1.92M_1=1/0.52≈ 1.92, and each satisfied answer costs nearly two attempts, S~1=1.92⋅60=115.4 S_1=1.92· 60=115.4 slot-milliseconds. The effective ratio is R=S~1/S~0=1.15R= S_1/ S_0=1.15, squarely inside the trap region (1,Δw0/Δw1)=(1,2)(1,\, w_0/ w_1)=(1,2). Per satisfied answer, tier 1 spends S~1Δw1=230.8 S_1 w_1=230.8 kW-ms of energy, that is 230.8230.8 joules, against tier 0’s 400400, a 42% saving on the electricity bill, while occupying 115.4115.4 slot-milliseconds against tier 0’s 100100: the cheap tier consumes 15% more capacity per satisfied answer. During a surge, when capacity rather than electricity is the binding resource, this is exactly the wrong trade, and it registers as a saving on every cost dashboard the provider owns, when in fact it is the opposite. See Figure 1 for a schematic. intelligence → per answerpostedrealizedstrong model (j=0j=0)distilled model (j=Jj=J) Figure 1: The two frontiers (schematic). The posted frontier, per raw query, is monotone: weaker is cheaper, which is what every pricing page implies. The realized frontier, per satisfied answer, carries the multiplier and the churn charge of (5) and bends back up as djρd_jρ grows; past the energy break-even the ordering of the tiers inverts. The green arrow is the inversion at the distilled model: the vertical gap between its posted and realized cost, the retry storm plus the churn, is the cost of confusing a query with an answer. Remark 1 (The third frontier). The frontiers (6) price energy and slot-time, but the cost (1) carries a third, occupancy-dependent column: the memory term γcmNκγ c_mN^κ. Routing traffic of rate λ to tier j shifts the standing population by ΔN=λ(S~0−S~j) N=λ ( S_0- S_j ), worth γcmκNκ−1ΔNγ c_mκ N^κ-1 N per hour at occupancy N: a discount on the weak tier that grows with load even before the service-level term bites. Explicitly, per unit of rerouted rate the full hourly ledger, energy plus memory minus churn from (5), reads γ(S~0Δw0−S~jΔwj)+γcmκNκ−1(S~0−S~j)−Mjdj(1−ρ)prℓ,γ ( S_0 w_0- S_j w_j )\;+\;γ c_mκ N^κ-1 ( S_0- S_j )\;-\;M_j\,d_j(1-ρ)\,p_r , and setting it to zero yields the closed-form break-even expected LTV loss (prℓ)⋆(N)=γ(S~0Δw0−S~jΔwj)+γcmκNκ−1(S~0−S~j)Mjdj(1−ρ),(p_r ) (N)\;=\; γ ( S_0 w_0- S_j w_j )+γ c_mκ N^κ-1 ( S_0- S_j )M_j\,d_j(1-ρ), increasing in N when S~j<S~0 S_j< S_0, with minimum at N=0N=0, the pure-energy break-even (prℓ)⋆:=(prℓ)⋆(0)(p_r ) :=(p_r ) (0). With prℓp_r below this break-even the weak tier dominates at every congestion level and the optimal policy has no occupancy threshold at all; above it the ledger flips sign at the finite threshold N⋆=[Mjdj(1−ρ)prℓ−γ(S~0Δw0−S~jΔwj)γcmκ(S~0−S~j)]1/(κ−1),N \;=\; [ M_j\,d_j(1-ρ)\,p_r -γ ( S_0 w_0- S_j w_j )γ c_mκ ( S_0- S_j ) ]^1/(κ-1), below which the churn charge dominates and above which the memory discount does. The threshold N⋆N is therefore not a tuning parameter but an economic object: its existence and location are determined by the expected LTV loss prℓp_r relative to (prℓ)⋆(p_r ) , that is, by how much a customer is worth. Section 4.3 meets this boundary again, from the other side. 3.2 The Dynamics of a Poorly Timed Throttle The static analysis locates the trap; it says nothing about time, and time is where the richness of the problem lives. Degradation is not a lever to pull but a moment to choose, and choosing it requires knowing how fast the system moves between its regimes. This subsection supplies those transition rates. A finite-server system lives in one of two regimes. Above capacity, all mkmk slots are busy and jobs complete at throughput kμjmk _jm, the number of finished queries the fleet delivers per unit time; with every slot serving at rate μj _j, this is the maximum throughput the fleet can attain. The backlog therefore changes linearly, accumulating when the arrival rate exceeds the processing rate, that is, r>kμjmr>k _jm, and depleting when the processing rate exceeds the arrival rate, that is, r<kμjmr<k _jm: jobs arrive at rate r and depart at rate kμjmk _jm, both constant, so the net change per unit time is the constant difference r−kμjmr-k _jm, or evolving over time as t(r−kμjm)t\,(r-k _jm). Below capacity, service scales with occupancy and the system relaxes exponentially toward the level at which arrivals balance departures, the equilibrium r/μjr/ _j that Little’s law prescribes (Whitt, 2002; Eick et al., 1993). One modification separates our system from the textbook one, and it changes everything: the arrival rate is endogenous, because completions at a degraded tier feed retries back into the stream. The following lemma makes the dynamics exact in expectation. Lemma 2 (Two-regime dynamics with endogenous arrivals). Let services at the incumbent tier j be exponential with rate μj _j, and let the effective arrival rate rteff=rt+djρ⋅(completion rate at t)=rt+djρμjmin(Nt,mk)r^eff_t\;=\;r_t+d_jρ·(completion rate at t)\;=\;r_t+d_jρ\, _j (N_t,\,mk) be constant over a step [t,t+Δt][t,t+ t] containing no regime switch. The retry term scales with min(Nt,mk) (N_t,mk), the number of jobs in service: a waiting job cannot generate a retry, because its customer has not yet received an answer, so the feedback saturates at djρkμjmd_jρ\,k _jm no matter how large the backlog grows. Then, in expectation, Nt+Δt=Nt+(rteff−kμjm)Δtif Nt≥mk(saturated, linear)Nte−μjΔt+rteffμj(1−e−μjΔt)if Nt<mk(recovering, exponential)N_t+ t\;=\; casesN_t+ (r^eff_t-k _jm )\, t&if N_t≥ mk (saturated, linear)\\[6.0pt] N_t\,e^- _j t+ r^eff_t _j (1-e^- _j t )&if N_t<mk (recovering, exponential) cases (7) and the regime switches occur exactly at the crossing times of N through mkmk, both closed form: from the saturated branch, t⋆=(mk−Nt)/(rteff−kμjm)t =(mk-N_t)/(r^eff_t-k _jm) whenever the drift points toward the boundary, and from the recovering branch, t⋆=μj−1ln[(N¯−Nt)/(N¯−mk)]t = _j^-1 [( N-N_t)/( N-mk) ] with N¯:=rteff/μj N:=r^eff_t/ _j, finite exactly when the equilibrium sits beyond the boundary. Proof. The recovering branch adapts the classical two-population argument for the transient infinite-server queue (Eick et al., 1993); the saturated branch, the endogenous arrival rate, and the switching logic are self-contained below. Recovering branch (Nt<mkN_t<mk). Every resident job holds a slot, so over the step the system serves as if capacity were unlimited, and the population at t+Δt+ t splits into survivors of the NtN_t incumbents and arrivals during (t,t+Δt](t,t+ t]. By memorylessness of the exponential, each incumbent’s residual service is again exponential with rate μj _j regardless of its elapsed service, so it remains resident with probability e−μjΔte^- _j t, and the incumbents contribute [∑i=1NtTi≥Δt]=Nte−μjΔt E [ _i=1^N_t 1\T_i≥ t\ ]=N_t\,e^- _j t. An arrival in (t+s,t+s+ds](t+s,\,t+s+ds] occurs with probability rteffdsr^eff_t\,ds, at most one per infinitesimal interval, and is still resident at t+Δt+ t with probability e−μj(Δt−s)e^- _j( t-s); integrating over the step, ∫0Δtrteffe−μj(Δt−s)s=rteff∫0Δte−μjuu=rteffμj(1−e−μjΔt). _0 tr^eff_t\,e^- _j( t-s)\,ds\;=\;r^eff_t _0 te^- _ju\,du\;=\; r^eff_t _j (1-e^- _j t ). Summing the two populations gives the exponential branch, a transient Little’s law: as Δt t grows the memory term vanishes and N relaxes toward the equilibrium rteff/μjr^eff_t/ _j that Little’s law prescribes. Saturated branch (Nt≥mkN_t≥ mk). All mkmk slots are busy, and the minimum of mkmk independent exponentials of rate μj _j is exponential with rate equal to the sum of the rates, so completions depart at aggregate rate kμjmk _jm while arrivals accrue at rate rteffr^eff_t; the expected net change over the step is (rteff−kμjm)Δt (r^eff_t-k _jm ) t, the linear branch. Endogeneity and switching. The completion rate entering rteffr^eff_t is kμjmk _jm when saturated and NtμjN_t _j when recovering; within a step it is constant by hypothesis, so the constant-rate derivations above apply, and a change of branch requires N to cross mkmk, which pins the switch epochs to the crossing times of N through mkmk. The formulas follow by solving each branch for the time at which N equals mkmk: the linear branch directly, the exponential branch by isolating e−μjte^- _jt and taking logarithms. ∎ Throughout, j indexes the tier currently serving, the strong tier 00 included; the regime, saturated or recovering, is determined by the state N alone, not by the tier. Between regime switches the trajectory is a formula and the switch epochs are the crossing times of the lemma (the 139 milliseconds of Example 3.2 below is the logarithm), so evaluating an entire horizon costs a handful of arithmetic operations: no time grid, no probability mass function, and, as Section 3.4 exploits, dual variables by finite differences at the same price. Remark 2 (What the lemma buys, and what it pays). The lemma tames the mean of a branching feedback, not the process. Its tractability is purchased by two modeling choices: instantaneous retries, which collapse the feedback into a rate modification rather than a convolution with a retry-delay distribution, and memorylessness, which makes the drift piecewise linear in the state. What the mean-field view discards is precisely the variance of the branching cascade, an overdispersion that grows without bound as djρd_jρ approaches the ignition threshold; Section 4.2 prices it. The recursion’s nearest antecedents all fix the feedback gain: Bernoulli-feedback queues (Takács, 1963) recycle completions at a constant probability, call-routing with service failure (de Véricourt and Zhou, 2005) recycles unresolved callers at a server attribute, retrial queues recycle customers blocked before service, and quality-feedback models move demand on the slow timescale of loyalty; in all of them the failure probability is an environment primitive rather than a control, so the completion-driven, same-timescale feedback of (7) with a chosen gain had no reason to be written down until degradation became a decision. The recursion is also simple enough to run by hand, which we now do. Example 3.2 (Throttling at the surge). A datacenter runs k=150k=150 clusters of m=20m=20 slots, so mk=3mk=3 thousand concurrent jobs, on tier 0 with [S0]=100 E[S_0]=100 milliseconds as in Example 3.1: μ0=10 _0=10 per second, and, setting d0=0d_0=0 to simplify the analysis, effective throughput θ0=30 _0=30 thousand jobs per second. A surge arrives at r=33.3r=33.3 thousand queries per second with N(0)=2N(0)=2 thousand jobs in the system. The controller is the industry’s standard reactive rule: when N crosses an operator-chosen trigger level N^fire N_fire, degrade everyone to tier 1, with [S1]=60 E[S_1]=60 milliseconds (μ1≈16.7 _1≈ 16.7 per second, raw throughput 50 thousand per second) and, for this traffic, d1=0.6d_1=0.6 and ρ=0.8ρ=0.8 as in Example 3.1. What happens? The system starts below capacity, relaxing toward the equilibrium r/μ0=33.3/10≈3.33r/ _0=33.3/10≈ 3.33 thousand. Since 3.33>33.33>3, it must cross the capacity line first, at the time solving 2e−10t+3.33(1−e−10t)=32e^-10t+3.33(1-e^-10t)=3, that is e−10t=0.25e^-10t=0.25, or t=ln4/10≈0.139t= 4/10≈ 0.139 seconds, 139 milliseconds in. From there the system saturates and grows linearly at r−θ0=3.3r- _0=3.3 thousand jobs per second. Suppose the trigger sits at N^fire=3.5 N_fire=3.5 thousand, so the rule fires 150 milliseconds later. On paper the throttle looks decisive: raw capacity jumps from 30 to 50 thousand per second against demand of 33.3, so the queue should drain at 16.7 thousand per second. Now account for the feedback loop. In saturation, completions run at 50 thousand per second, and a fraction d1ρ=0.48d_1ρ=0.48 of them come back: 24 thousand retries per second, so reff=33.3+24=57.3>50r^eff=33.3+24=57.3>50. The queue does not drain. It grows at 7.3 thousand jobs per second, more than twice its pre-throttle rate, while the service-level term of (1) bites quadratically in the swelling backlog. The throttle raised raw capacity by two thirds and made the congestion worse. In plain words: once the retries it generates are subtracted from its raw output, tier 1 delivers θ1=(1−d1ρ)×50=26 _1=(1-d_1ρ)× 50=26 thousand fresh answers per second against tier 0’s θ0=30 _0=30. The fast cheap model is the slow expensive one. And every second of the spiral pours 50×0.6×0.2=650× 0.6× 0.2=6 thousand abandonments onto the churn ledger, none of which appear on the cost-rate chart. Figure 2 sketches the path. What would work instead? Not this throttle at another time: with a single class the instance sits inside the trap of Proposition 1 (S~1=115.4>100=S~0 S_1=115.4>100= S_0), so degrading everyone, at any moment, adds effective load. The anticipatory path of the figure requires heterogeneity. Suppose 30%30\% of the traffic is insensitive to the weak tier, say d1=0.05d_1=0.05 and ρ=0.3ρ=0.3 for that fraction. Run the same arithmetic as before: d1ρ=0.015d_1ρ=0.015, so M1=1/0.985≈1.02M_1=1/0.985≈ 1.02 and S~1≈1.02×60≈61 S_1≈ 1.02× 60≈ 61 milliseconds, a multiplier near one because this traffic barely retries. Now price the offered load, arrivals times effective service time, class by class: the sensitive 70%70\% stays on tier 0 and occupies 0.7×33.3×0.100≈2.330.7× 33.3× 0.100≈ 2.33 thousand slots, the insensitive 30%30\% moves to tier 1 and occupies 0.3×33.3×0.061≈0.610.3× 33.3× 0.061≈ 0.61 thousand, for a total of 2.94<32.94<3 thousand. Degrading the insensitive fraction before the surge therefore holds the system below capacity: it relaxes toward 2.942.94 thousand, never saturates, and no spiral ignites, the dashed path of Figure 2. Who is insensitive, and how to find them, is the subject of Section 3.3. 22mk=3mk=3N^fire=3.5 N_fire=3.5throttle fires at N^fire N_fire:slope more than doublesrelaxes to N¯≈2.94<mk N≈ 2.94<mktt (seconds)N(t)N(t) (thousand jobs)reactive: the spiralanticipatory: never ignites Figure 2: The boomerang of Example 3.2, in the example’s units of thousands of jobs. Three levels anchor the vertical axis: capacity mk=3mk=3 (dotted), the reactive trigger N^fire=3.5 N_fire=3.5 (dash-dotted), and the Little’s-law equilibrium N¯=reff/μ≈2.94 N=r^eff/μ≈ 2.94 of the anticipatory mix, which sits below capacity. Reactive throttling fires when N crosses the trigger and ignites the retry spiral; the anticipatory path relaxes toward N¯ N and never saturates. The cost rate spikes before it saves, and the churned lifetime value accrues on a separate, monotone ledger that no cost chart displays. The example is not an accident of its numbers. Each of its pathologies, the throttle that adds load, the queue that never drains, the release that never comes, is an instance of a structural result, and the results follow. Proposition 3 (Ignition threshold). Suppose the system is saturated, Nt≥mkN_t≥ mk, and the fleet switches to tier j. The queue drains back below mkmk if and only if r<θjr< _j, the tier’s effective throughput (3). Above this threshold the retry feedback is self-sustaining: the drift is positive, N never returns below mkmk, and the backlog and the churn ledger grow without bound. Proof. With N≥mkN≥ mk, all mkmk slots are busy, completions run at kμjmk _jm, and a fraction djρd_jρ re-enter, so reff=r+djρkμjmr^eff=r+d_jρ\,k _jm and the net change of N per unit time is ΔNΔt=reff−kμjm=r−(1−djρ)kμjm=r−θj. N t\;=\;r^eff-k _jm\;=\;r-(1-d_jρ)\,k _jm\;=\;r- _j. The queue drains if and only if r−θj<0r- _j<0. If positive, N stays above mkmk, the saturated branch of (7) continues to apply, and the same positive rate persists forever: what the throttle sheds in service time, the feedback more than replaces in retries. ∎ Corollary 4 (The latch). Every reactive rule comes in two halves: it degrades to tier 11 when N exceeds the trigger N^fire N_fire of Example 3.2, and it releases back to tier 00 when N falls below a release level N^rel≤N^fire N_rel≤ N_fire, the backlog deemed low enough to declare the congestion over. Start the clock when the surge ends: from time 00 on, the arrival rate is a constant rcalm<rsurger_calm<r_surge, demand genuinely back to normal, the system still degraded, and N0>N^relN_0> N_rel. The rule returns to the strong tier if and only if rcalmS~1<N^rel.r_calm\, S_1\;<\; N_rel. The left side is where the degraded system comes to rest: by Little’s law, arrivals at rate rcalmr_calm each holding a slot for an effective time S~1 S_1, retries included, keep rcalmS~1r_calm\, S_1 jobs in the system. The release fires only if that resting level lies below the release line. When the condition holds, the release is also final: back on tier 00 the same logic applies with S~0 S_0 in place of S~1 S_1, and in the trap region S~0<S~1 S_0< S_1, so the system settles at the lower level rcalmS~0r_calm\, S_0, below the line it just crossed, and the trigger never re-fires. When it fails, the throttle never releases, even though the surge is over: the transient surge becomes a permanent degraded regime, and the churn ledger accrues at the constant rate rcalmM1d1(1−ρ)prℓr_calm\,M_1d_1(1-ρ)\,p_r forever. An operator who sets N^rel N_rel below rcalmS~1r_calm\, S_1 has, without knowing it, disabled the release: every customer is served by the weak model until demand itself moves, and the exit, when it comes, arrives for the wrong reason, a quiet weekend, or the churn the latch is causing eroding rcalmr_calm until the resting level sinks below the line. The rule reads that as congestion clearing; it is the customer base clearing. Proof. For t>0t>0 the arrival rate is the constant rcalmr_calm and the tier is 11, so by (7) the trajectory decreases monotonically toward the degraded equilibrium N¯1:=rcalm/((1−d1ρ)μ1)=rcalmS~1 N_1:=r_calm/ ((1-d_1ρ) _1 )=r_calm\, S_1, first linearly while saturated, then along the exponential branch once below mkmk. If N¯1<N^rel N_1< N_rel, the trajectory crosses the release line in finite time, by the crossing formulas of Lemma 2, and the rule releases. If N¯1≥N^rel N_1≥ N_rel, the trajectory approaches N¯1 N_1 from above and stays above it forever, hence above the release line: Nt>N^relN_t> N_rel for all t, and the rule never fires its release. Durability after release: once the tier is 00, the same monotone argument drives N toward the strong tier’s resting level N¯0:=rcalmS~0 N_0:=r_calm\, S_0. In the trap region S~0<S~1 S_0< S_1 gives N¯0<N¯1<N^rel≤N^fire N_0< N_1< N_rel≤ N_fire, so the trajectory keeps falling and never re-crosses the trigger. Outside the trap, N¯0>N¯1 N_0> N_1 and the trajectory climbs back toward N¯0 N_0; the release holds if and only if N¯0<N^fire N_0< N_fire, and when it does not, the rule re-fires and cycles between the tiers indefinitely. ∎ Proposition 5 (Anticipation dominates reaction). This is Example 3.2 with the one ingredient the reactive rule ignores, the forecast: the demand model of Section 2 announces the surge in advance. Start the system at rest on tier 00: for t<t1t<t_1 the arrival rate is rcalm<θ0r_calm< _0 and N sits at its resting level rcalmS~0<mkr_calm\, S_0<mk. On the surge window [t1,t2][t_1,t_2] the rate jumps to rsurger_surge: too high for the strong tier, rsurge>θ0r_surge> _0, so something must give; too high for a uniform throttle, rsurge>θ1r_surge> _1, so degrading everyone ignites; yet feasible under full degradation, rsurgeS~1<mkr_surge\, S_1<mk, so a policy that clears capacity in advance exists. Compare two policies. The reactive rule waits for N to cross the trigger N^fire N_fire and then degrades everyone; by Proposition 3 it ignites, and from its firing time τ∈(t1,t2)τ∈(t_1,t_2), after the surge begins and before it ends, the service-level cost accrues as the cube of the remaining surge, at least CSLA3(mk)2(rsurge−θ1)2(t2−τ)3 C_SLA3(mk)^2(r_surge- _1)^2\,(t_2-τ)^3. The anticipatory policy degrades a fraction φ of traffic from a lead time t0<t1t_0<t_1, while there is still slack capacity for the retry cascade to drain into; it achieves the same energy savings, never saturates, and pays exactly one charge the reactive rule avoids, churn over the lead window, φrcalmd1(1−ρ)prℓ(t1−t0), \,r_calm\,d_1(1-ρ)\,p_r \,(t_1-t_0), linear in the lead and independent of the surge length. Cubic against linear: anticipation trades a fixed, known loss, the churn of the customers degraded early, for the removal of a loss that compounds with every moment of the surge, the backlog swelling under a positive drift. Nor does the reactive bill close at t2t_2: the surge ends with (rsurge−θ1)(t2−τ)(r_surge- _1)(t_2-τ) jobs above capacity, which drain only at the calm margin θ1−rcalm _1-r_calm, paying the same cubic charge a second time, scaled by (rsurge−θ1)/(θ1−rcalm)(r_surge- _1)/( _1-r_calm); even a short surge leaves a long aftermath (this assumes rcalm<θ1r_calm< _1, so the degraded queue can drain at all; otherwise the reactive cost is unbounded and the latch case below applies). Equating premium to penalty makes the trade-off exact: anticipation strictly dominates if and only if the post-firing surge exceeds the critical window T⋆=[3(mk)2φrcalmd1(1−ρ)prℓ(t1−t0)CSLA(rsurge−θ1)2(1+rsurge−θ1θ1−rcalm)]1/3,T \;=\; [ 3(mk)^2\, \,r_calm\,d_1(1-ρ)\,p_r \,(t_1-t_0)C_SLA\,(r_surge- _1)^2\, (1+ r_surge- _1 _1-r_calm ) ]^1/3, when t2−τ<T⋆t_2-τ<T the surge is too brief for the backlog to hurt and the reactive rule is genuinely cheaper: anticipation is not free insurance, it is insurance worth buying exactly when the storm outlasts T⋆T . The cube root is the strengthening: because the penalty compounds cubically while the premium accrues linearly, T⋆T grows only as the cube root of the churn premium, so even a tenfold increase in what anticipation costs barely doubles the surge length that justifies it. And whenever the latch of Corollary 4 binds, T⋆T is irrelevant: the reactive ledger grows without bound and dominance holds at every horizon. The optimal switch surface is computed by dynamic programming over (7), with the chance constraint P(W>s)≤α P(W>s)≤α enforced by pruning every state-action pair violating skμm≥max(N,mk)+max(N,mk)Φ−1(1−α)s\,kμ m≥ (N,mk)+ (N,mk)\; ^-1(1-α). Proof. Appendix A.1. ∎ The three results share one mechanism: under congestion, throttling is not a cost lever but a demand lever. It manufactures the very traffic it was deployed to shed, and doing it reactively converts a capacity shortage into a self-sustaining one. 3.3 Who to Throttle: Customer Heterogeneity So far every customer felt degradation identically, and real customer bases are not like that. A researcher doing context-heavy work feels a single tier of degradation acutely; a user parsing CSV fields into a table does fine on a smaller model. Index customer classes by x∈x∈ X, a label observed by the router; X is a general set; taking =researcher,parser X=\researcher,parser\, two classes each aggregating the many customers whose workload fits the description, gives the smallest instance rich enough for everything this section builds: the index, the nested thresholds, and the dual spread of Section 3.4 all take their simplest nontrivial form on two classes. Let each primitive of Section 2 carry the class: the dissatisfaction probability becomes dj(x)d_j(x), the retry probability ρx _x, and the expected LTV loss (prℓ)x(p_r )_x. Formally, the sensitivity curve dj(x)d_j(x) is steep in j for the first type and nearly flat for the second, and the retry probability ρx _x and the expected LTV loss (prℓ)x(p_r )_x order the same way: the researcher re-asks more and is worth more. The effective service time and the effective throughput of (3) become class-dependent, S~j(x):=Mj(x)[Sj] S_j(x):=M_j(x)\, E[S_j] and θj(x):=mk/S~j(x) _j(x):=mk/ S_j(x). To feel the asymmetry, take a researcher at d1=0.6d_1=0.6, ρ=0.9ρ=0.9 and a parser at d1=0.05d_1=0.05, ρ=0.3ρ=0.3: the multipliers are M1=(1−0.54)−1≈2.17M_1=(1-0.54)^-1≈ 2.17 against (1−0.015)−1≈1.02(1-0.015)^-1≈ 1.02. Routing the researcher to tier 1 more than doubles their traffic; routing the parser is nearly free. Uniform throttling, the standard load balancer, treats these two customers identically; we study the tradeoff as follows. The action is no longer a tier but a routing: φ(x,j) (x,j) is the fraction of class-x traffic sent to tier j, and λx _x is class x’s fresh-demand rate, the class decomposition rt=∑xλxr_t= _x _x of the forecast, counting first attempts only because S~j(x) S_j(x) already carries each attempt’s retries. Writing c(x,j):=γS~j(x)Δwj+Mj(x)dj(x)(1−ρx)(prℓ)xc(x,j):=γ\, S_j(x)\, w_j+M_j(x)\,d_j(x)(1- _x)\,(p_r )_x for the cost per satisfied class-x answer at tier j, the first term the energy of (5) and the second its destroyed lifetime value, now evaluated at class x’s own multiplier and expected LTV loss, the one-period problem is a transportation problem with classes as demand nodes and tiers as warehouses. One precaution decides whether the formulation is right or wrong: the capacity constraint must charge each unit of routed traffic its retry-inflated load S~j(x) S_j(x), not its posted service time, otherwise the optimization sees tier 1 as cheap capacity and happily recommends the spiral of Example 3.2, min∑x,jφ≥0c(x,j)λxφ(x,j)s.t.∑jφ(x,j)=1∀x∈,∑xS~j(x)λxφ(x,j)≤mkj∀j. _ ≥ 0\; _x,jc(x,j)\, _x\, (x,j) .t. _j (x,j)=1\;\;∀ x∈ X, _x S_j(x)\, _x\, (x,j)\;≤\;mk_j\;\;∀ j. (8) The objective sums, over every class and tier, the cost per satisfied answer times the volume routed there: energy plus destroyed lifetime value, dollars per unit time. The first constraint says every class is served somewhere, the fractions across tiers summing to one, so the provider cannot solve congestion by silently dropping a class. The second constraint prices capacity in retry-inflated slot-time: each unit of class-x traffic routed to tier j occupies S~j(x) S_j(x) slot-time, the multiplier included, so a class that retries heavily consumes capacity its posted service time conceals, and in the trap region routing it to the weak tier tightens the constraint the routing was meant to relax. The linear objective is reconciled with the nonlinear cost (1) term by term: the idle floor γkw0γ kw_0 is constant across routings and drops from any argmin; the service-level penalty is not ignored but converted into the capacity constraint, exact rather than approximate because on the feasible set the excess backlog is identically zero, its economics surviving as the constraint’s dual ν; and the memory term, convex in occupancy, is linearized at the operating point, its marginal rate γcmκNκ−1γ c_mκ N^κ-1 per unit slot-time being precisely the discount Remark 1 prices, with the omitted curvature only reinforcing the index ordering that follows. The linear program is small, classes times tiers, and solves in microseconds; its value is not the solution but its structure, which the next proposition extracts: at a binding constraint the optimum is a greedy ordering, and the ordering is by a single computable index. researcher λR _Rparser λP _Ptier 00 (strong) mkmktier 11 (weak) mkmkλRφ(R,0) _R\, (R,0)λPφ(P,1) _P\, (P,1)λRφ(R,1) _R\, (R,1)λPφ(P,0) _P\, (P,0)∑jφ(x,j)=1 _j (x,j)=1∑xS~0(x)λxφ(x,0)≤mk _x S_0(x)\, _x\, (x,0)\;≤\;mk∑xS~1(x)λxφ(x,1)≤mk _x S_1(x)\, _x\, (x,1)\;≤\;mkedge cost per satisfied answer:c(x,j)=γS~j(x)Δwj+Mj(x)dj(x)(1−ρx)(prℓ)xc(x,j)=γ\, S_j(x)\, w_j+M_j(x)\,d_j(x)(1- _x)\,(p_r )_x Figure 3: The transportation LP (8) for two classes and two tiers. Solid arrows are the dominant flows in the calibrated instance; dotted arrows are the cross-assignments the LP considers and rejects. Capacity is charged in effective slot-time S~j(x)=Mj(x)[Sj] S_j(x)=M_j(x)\, E[S_j], not posted [Sj] E[S_j]: a heavy-retry class consumes capacity its service time conceals, and in the trap region the capacity constraint tightens under degradation. Proposition 6 (Index rule). At a binding capacity constraint, the optimal routing degrades classes in increasing order of the index I(x,j)=∂c(x,j)/∂jS~0(x)−S~j(x),I(x,j)\;=\; ∂ c(x,j)/∂ j\; S_0(x)- S_j(x)\;, marginal damage per unit of capacity relief, and the optimal policy has nested thresholds in the congestion state: the set of degraded classes at backlog N contains the set degraded at any N′<N <N. Proof. Appendix A.2. The index is a relative of the cμcμ rule (Van Mieghem, 1995), with the denominator written in effective slot-time: degrading a heavy-retry class buys less capacity than its service-time discount suggests, because part of the freed capacity is immediately reclaimed by that class’s own retries, and in the trap region of Proposition 1 the denominator goes negative, pricing the class out of degradation entirely when capacity binds. ∎ Remark 3. When the fleet is congested, every slot of capacity freed by degrading someone must be paid for in degraded service, and the exchange rate differs by customer. The index says: degrade the customers for whom that swap is cheapest first: the ones whose retries eat the least of the capacity just freed, and whose dissatisfaction and churn cost the least. Heavy-retry, high-value customers are priced out of degradation entirely when capacity binds, because degrading them buys less relief than it costs in damage. Proposition 7 (Dominance). Targeted degradation weakly dominates uniform degradation for any sensitivity profile, and strictly whenever profiles differ across classes. Proof. Any uniform policy is feasible for (8) with φ(x,j) (x,j) constant in x, so the optimum weakly improves on it. For strictness: at equal energy savings, uniform throttling books churn at the population-average expected LTV loss while the targeted policy books it at the minimum over classes achieving the same capacity relief, and when profiles differ the average strictly exceeds the minimum. ∎ The uniform policy is the LP optimum restricted to routings of the form φ(x,j)=φ(j) (x,j)= (j), the same tier mix for every class; its gap to the true optimum is the excess churn it books at the population-average LTV loss rather than the class-specific minimum, positive exactly when sensitivity profiles differ. That gap is the entire value of knowing who is who. 3.4 The Shadow Price of Intelligence The capacity constraint’s dual, the Lagrange multiplier on the retry-inflated slot-time budget of (8), is the number the paper is named after: it is the marginal cost of one additional unit of effective load, the price the system assigns to intelligence at the binding tier. In the static problem each demand-satisfaction constraint carries a dual variable πx _x, and in the dynamic problem the corresponding object is a derivative through the value function, πx(t):=∂Vt⋆∂λx(t), _x(t)\;:=\; ∂ V _t∂ _x(t), (9) the system’s marginal cost of one more class-x query at hour t. The value function behind (9) is the dynamic version of the transportation problem (8). The state is the class-indexed backlog s=(Ns(x))x∈ N_s=(N_s(x))_x∈ X with total Ns:=∑xNs(x)N_s:= _xN_s(x); an action is an assignment as:→0,…,Ja_s: X→\0,…,J\ of a tier to each class, and A denotes the admissible assignment sequences (at,…,aT)(a_t,…,a_T). Write ns(x)n_s(x) for the class-x jobs in service, all of them below capacity and the proportional share above it, ns(x):=Ns(x)min1,mk/Nsn_s(x):=N_s(x) \1,\,mk/N_s\. The stage cost is the operating cost (1) with its active-draw term opened up by class, the multi-tier extension deferred in Section 2, c(k,s,as):=γ[kw0+∑x∈ns(x)Δwas(x)]+γcmNsκ+CSLA((Ns−mk)+mk)β,c (k, N_s,a_s )\;:=\;γ [k\,w_0+ _x∈ Xn_s(x)\, w_a_s(x) ]\;+\;γ\,c_mN_s^κ\;+\;C_SLA\! ( (N_s-mk)^+mk )^\!β, which reduces to (1) when every class shares one tier. Then Vt⋆=mina(⋅)∈ V _t\;=\; _a(·)∈ A ∑s=tT[c(k,s,as)+∑x∈das(x)(x)(1−ρx)μas(x)ns(x)(prℓ)x] _s=t^T [\,c (k, N_s,a_s )\;+\; _x∈ Xd_a_s(x)(x)\,(1- _x)\, _a_s(x)\,n_s(x)\,(p_r )_x ] (10) s.t. .t. Ns+1(x)=ℱ(Ns(x),λx(s),as(x))∀x∈, N_s+1(x)\;=\;F (N_s(x),\, _x(s),\,a_s(x) ) ∀ x∈ X, P(Ws>s¯)≤α, P (W_s> s )\;≤\;α, with ℱF the two-regime recursion (7) applied class by class, each class carrying its own effective quantities, dissatisfaction das(x)(x)d_a_s(x)(x), retry probability ρx _x, and its capacity share as the regime boundary. The second sum is the churn ledger: class-x completions depart at rate μas(x)ns(x) _a_s(x)\,n_s(x), each an attempt that fails with probability das(x)(x)d_a_s(x)(x) and abandons with probability 1−ρx1- _x, booking the expected LTV loss (prℓ)x(p_r )_x. The latency constraint is a chance constraint at threshold s¯ s, enforced by pruning every state–action pair that violates the normal surrogate of Proposition 5; crucially, it permits transient saturation and charges for it through CSLAC_SLA, rather than forbidding the saturated regime the analysis is about. The fresh-demand rate λx(s) _x(s) enters only the class-x dynamics, so the derivative (9) is well defined class by class; with a single class the assignment collapses to a tier-switching sequence and (10) is the problem of Section 3.2 verbatim. The forward pass makes πx(t) _x(t) computable at every level: closed form at the newsvendor level, obtained by differentiating (4) through the chance constraint; the LP dual of (8) in the static problem; and, in the dynamic model, a finite difference through the deterministic forward pass, with no Monte Carlo jitter. Because the forward pass is closed form between regime switches, the dynamic dual recomputes in milliseconds, making it a live control signal rather than an overnight batch job (Section 5 reports the measured cost). The structure of πx(t) _x(t) answers the posting quoted in the introduction, in two parts. First, the class duals differ at every hour, not only at the crunch: even with the capacity constraint slack, a marginal specialized query costs its strong tier’s multiplied compute and churn, Mj(x)(cj+dj(x)(1−ρx)(prℓ)x)M_j(x) (c_j+d_j(x)(1- _x)(p_r )_x ), while a marginal casual query costs its cheap tier’s, so the flat relative price of one that any uniform policy implicitly charges is wrong around the clock. Second, when capacity binds, the shared pool adds a scarcity rent that is nearly common across classes, because congestion anywhere reprices capacity everywhere: dollar differences between the class duals widen at the crunch while their ratios compress toward one. The distance the posting names between allocation choices and user outcomes is therefore a two-part number, a floor spread that never closes and a rent that arrives for everyone at once, and both parts are now computable, monitorable, and optimizable against; Section 5.4 computes them on the calibrated instances. 4 From Fluid to Practice Four questions separate the analysis from a deployment: how to solve the dynamic program fast enough to matter, what the fluid approximation discards and whether it matters, how to make the optimal policy operable by a human, and whether the one primitive that is not directly observable can be identified from data. Each has a short answer. 4.1 Solving It: the Complexity Dividend of the Closed Form The value function (10) is solved by backward induction: T epochs, a grid of g points per class backlog with multilinear interpolation, and the admissible assignments 1 A_1 surviving the static ledger (5), which prunes the menu before the dynamics begin (Section 5.1 reports collapses as severe as 2744→722744→ 72). Any solver of this type evaluates the same count of transitions, T×|1|×g||,T\;×\;| A_1|\;×\;g^| X|, so the cost difference between solvers lives entirely in the price of one transition, and that price is the point of Lemma 2. Without the lemma, the natural transition oracle is the drift itself: the backlog changes at rate rteffr^eff_t minus the completion rate, and one steps this forward, either as the Markov chain or as its mean ODE. Any such explicit scheme must resolve the fastest relaxation in the dynamics, the below-capacity decay at rate μj _j: stability and accuracy require steps δ with μjδ≤c _jδ≤ c, c≈0.2c≈ 0.2 in practice, hence μmaxΔ/c _ /c substeps per epoch of length Δ . The service rate sets the integrator’s clock, and the fast, cheap tiers that make degradation attractive are exactly the tiers that make integrating it expensive. Lemma 2 is the escape: because the drift is piecewise linear in the state, the ODE solves in closed form on each side of the capacity boundary, the crossing time is itself a formula, and within a leg the trajectory crosses at most once, so one transition costs a constant handful of elementary-function evaluations regardless of μmax _ or Δ . Per transition: Θ(μmaxΔ)⏟EuleragainstΘ(1)⏟closed form, \! ( _ \, )_Euler (1)_closed form, a speedup linear in the service-rate scale. This is the pattern of Table 4: the measured factors grow with μmax _ and peak on the fastest ladder, sitting below the raw substep ratio only because interpolation and cost evaluation are overhead common to both solvers. Figure 4 draws one transition under each solver. backward sweep, T epochssss+1s+1backlog grid, g||g^| X|t→t+Δt\ →\ t+ kmk(a) converged Eulerbackward sweep, T epochssss+1s+1backlog grid, g||g^| X|t→t+Δt\ →\ t+ kmk(b) closed form (Lemma 2)μmaxΔ/c _ /c steps,one evaluation eacht⋆t two formulas,one crossing time Figure 4: One transition of the backward induction, magnified. The induction fills the same T×|1|×g||T×| A_1|× g^| X| value surface under either solver; the panels differ inside the transition arrow. (a) An explicit integrator resolves the recovering branch’s relaxation at rate μj _j, so each transition takes μmaxΔ/c _ /c substeps. (b) Lemma 2 gives the same trajectory as a linear leg and an exponential leg joined at the closed-form crossing time t⋆t : constant cost per transition, independent of the service-rate scale. The device is worth locating in the queueing literature, because it occupies a gap. The standard responses to time-varying load sit at two poles. At one, stationary formulas are applied pointwise, hour by hour, which is blind to precisely the object this paper controls: how fast the backlog moves between regimes, the transient on which the ignition threshold and the latch live. At the other, exact transient analysis, the fluid and strong-approximation limits for time-varying many-server systems (Mandelbaum and Massey, 1995; Whitt, 2002) and, further out, the distribution-dependent formulations of Dieker et al. (2026), describes the trajectory faithfully but as an object to be integrated, at the integrator’s price derived above. The two-regime recursion of Lemma 2 sits between the poles: transient, yet a formula, a transient Little’s law that relaxes to the stationary prescription as the step grows and is exact along the way. Its ingredients are classical, the recovering branch in particular being the transient infinite-server decomposition of Eick et al. (1993); what we have not found in the literature is the pieced two-regime trajectory itself, crossing times included and arrivals made endogenous, nor its deployment as a constant-cost transition oracle inside an optimization loop, and that deployment is what turns transient control from a simulation study into arithmetic. The line between exact and approximate deserves stating precisely. The single-class dynamics of (7) requires no time discretization at all: between changes of the arrival rate or the tier, the trajectory is a formula, and the only event is the capacity crossing, whose time is itself closed form. Discretization enters solely through the multi-class extension, where a shared pool makes each class’s capacity a function of every class’s current backlog; the solver holds these shares frozen within a leg, re-resolving at the leg boundaries, and the error is first-order in the leg length. It concentrates in deep saturation with unbalanced backlogs, off the optimal trajectory, which is where the Q-regret tails of Section 5.5 sit while the policy-value certificate stays within single digits of percent on every instance. The dividend compounds at the dual: after the backward pass, a forward pass costs T transitions, so the finite difference (9) prices the full daily surface πx(t) _x(t) in O(||T2)O(| X|\,T^2) transitions of closed-form arithmetic, no Monte Carlo, no averaging. The solve runs in seconds; repricing against a revised forecast, the input that actually changes intraday, runs in milliseconds. That is the difference between a shadow price that is a quarterly slide and one that is a live control signal. Stepping back, the tractability is not one device but four multiplying: memorylessness collapses the state to a backlog per class, the static ledger (5) collapses the action space before the dynamics begin, the lemma collapses the transition to constant cost, and the normal surrogate collapses the tail constraint to a pruning rule. Remove any one and the problem returns to overnight simulation. 4.2 The Role of the Variance The fluid model (7) buys its tractability, closed-form trajectories whose evaluation over a horizon costs one elementary-function call per regime switch, by discarding fluctuations, and an honest account of what that purchase costs has two faces. Where the model is safe: Propositions 1, 6, and 7 are statements about orderings and regions, not levels, so noise moves their boundaries by O(σ)O(σ) without reordering unless the compared quantities are nearly equal, in which case either choice is near-optimal and the error is self-limiting; the saturated regime is drift-dominated, and the churn ledger aggregates over many customers and obeys the law of large numbers. Where it is not: the tail. The service-level constraint P(W>s)≤α P(W>s)≤α is a statement about fluctuations, and a fluid path can only answer zero or one; worse, the economics of capacity puts the operating point exactly where this matters, since idle servers are pure cost it makes economic sense to run a system that is barely stable, so the provider lives in the critical band where the normal surrogate is a central limit theorem applied precisely where it is least reliable. 0.350.350.40.40.450.450.50.5000.50.511d⋆d fluid andexact disagreedissatisfaction probability djd_jP(ignition)fluidexact Figure 5: The ignition boundary and its stochastic width (schematic). The fluid model (dashed) predicts a sharp step at d⋆d : below, stable; above, ignited. The exact stochastic system (solid) transitions over a band of finite width. The shaded region marks where the two verdicts disagree: the fluid gives a binary answer while the exact ignition probability lies in the middle, and a fluctuation can push the system into the spiral from which the fluid dynamics say there is no return. And the mechanism manufactures its own variance. The retry cascade is a branching process, so the effective arrival stream is overdispersed, and its variance blows up as djρd_jρ approaches the threshold of Proposition 3, in the same way a supply network propagates variability through a multiplier of its own, distinct from and larger than the mean multiplier (Lee et al., 1997; Chen et al., 2000). The consequence is that the deterministic threshold is really an ignition boundary: below it the fluid system is stable, but a fluctuation can kick the stochastic system across, into a spiral from which the fluid dynamics correctly say there is no return. The repair lies in one of the deepest and well-established concepts in operations research. Means decide who, in what order, and when; variance decides the buffer. Every threshold in the policy, the throttle trigger, the forbidden region, the tree boundaries of the next subsection, is pulled inside its fluid location by Φ−1(1−α)scale ^-1(1-α) scale: square-root safety staffing, which is to say the newsvendor buffer of (4) applied to a boundary instead of a quantity. And the asymmetry is an interesting result in itself: noise punishes the reactive policy, which parks the system flush against the ignition boundary, and rewards the anticipatory one, which buys distance from it. Stochasticity strengthens the thesis it might have been expected to erode. Section 5 verifies both halves of this account on the calibrated instances: the fluid verdict is essentially exact outside a narrow critical band, the ignition transition has the width the central limit theorem predicts, and the buffered threshold becomes a number rather than a sentiment. With delayed retries the arrival rate becomes a functional of the current population state: the dynamics are distribution-dependent, a nonlinear Markov chain in exactly the sense of Dieker et al. (2026), and it is the principled extension of this work. The fluid model of this paper is a first-order approximation: it locates the trap, the spiral, and the index rule; the regime where the stakes concentrate, exact tail certification in the critical band, heavy-tailed token distributions, retries arriving with delay, is left for future work. 4.3 The Policy as a Tree The optimal policy of Section 3 is a heatmap over (t,N,x)(t,N,x) that no operations engineer will certify, diff, or roll back. The industry’s actual policy, degrade everyone once N>N^N> N, is a depth-one stump, constant in x. The territory between the stump and the heatmap is a decision tree, and here it comes cheap, because the situation is distillation, not learning: tree methods for Markov decision processes interleave splitting and optimization because the policy is unknown during construction (Sanabria et al., 2021), whereas here the teacher is already solved. The procedure has four steps. First, run the fluid dynamic program and extract the optimal actions, the action values Q(t,N,x,a)Q(t,N,x,a), and the occupancy measure μ(t,N)μ(t,N) under the optimal policy, one backward and one forward pass with closed-form dynamics. Second, fit the tree to minimize the occupancy-weighted value regret ∑t,N,xμ(t,N)[Q(t,N,x,a⋆)−Q(t,N,x,tree(t,N,x))], _t,N,xμ(t,N)\, [Q(t,N,x,a )-Q(t,N,x,tree(t,N,x)) ], (11) so that where Q is flat across actions the tree may be wrong for free and where it is steep, near the ignition boundary, the splitter concentrates resolution, coarse where the constraint is slack and fine where it bites. Third, close the loop: deploying the tree changes the occupancy and the retry feedback again, so roll (7) forward under the tree, refit on the induced occupancy, and repeat until the splits stabilize; every candidate ships with an exact certificate of leaf count, value gap, and violation path. Fourth, hard-code the buffered forbidden region of Section 4.2: any leaf whose cell intersects it takes the safe action, so the chance constraint holds by architecture rather than by hope. Proposition 8 (Canonical form). Under the nested-threshold structure of Proposition 6, the optimal policy admits an ϵε-exact tree with O(||×#regimes×J)O(| X|×\#regimes× J) leaves, and the distillation above recovers it. The tree is not a compression of the policy; it is the policy’s normal form. Proof. Appendix A.3. ∎ Each leaf is literally a row, “surge, class = parser → tier 1,” that an engineer can read, stress-test, diff, and roll back, which is the deployable format the stump already has and the heatmap never will. 4.4 Identifying the Price of the Discount Every primitive of the model but one is read off a benchmark or a meter: service times and power draws from public leaderboards and serving traces, capacities and peak windows from published documentation, electricity from the bill. The exception is the behavioral triple, the dissatisfaction probability dj(x)d_j(x), the retry probability ρx _x, and the expected LTV loss (prℓ)x(p_r )_x, and the estimation program for it runs on logs every provider already keeps. Tier assignments are recorded per query; sessions link queries to customers; and dissatisfaction leaves fingerprints: a re-ask of the same question within minutes, semantically near its predecessor, an explicit regeneration, a thumbs-down, an agentic run abandoned mid-task. The per-answer rate of retry fingerprints at tier j estimates the product dj(x)ρxd_j(x) _x directly. The product is the right target, because it is what the results consume. The multiplier (2), the effective quantities (3), the trap condition of Proposition 1, the ignition threshold of Proposition 3, and the latch of Corollary 4 depend on (dj,ρ)(d_j,ρ) only through djρd_jρ, the directly estimable object; once it is in hand, fresh demand follows from logged arrivals by the deconvolution of footnote 3. Separating djd_j from ρ is needed only for the churn column of (5), and the abandonment branch identifies it imperfectly: a dissatisfaction fingerprint followed by no retry brackets dj(1−ρ)d_j(1-ρ) from below, and all answers with no follow-up bracket it from above, because a silent satisfied exit and a silent dissatisfied one look alike. The honest output is therefore an interval, which propagates to an interval on the churn column and leaves the capacity side of every frontier untouched. The identification split runs through every downstream object: the fingerprint rate pins djρd_jρ exactly, so the multiplier, the effective service time, the trap interval, and the ignition threshold are point-identified from the log; the churn column needs djd_j and ρx _x apart plus (prℓ)x(p_r )_x, and every object reading it, the newsvendor buffer, the index numerator, the dual floor, carries the interval forward. The program is built so that the interval lives on the ledger the dashboard does not show. The comparison across tiers is confounded by construction: routers degrade under congestion, and congestion changes who is asking and how patient they are, so naive per-tier fingerprint rates mix the tier’s effect with the crowd’s. The environment supplies two instruments. Past capacity incidents shift tier assignment for fleet-level reasons orthogonal to any individual query, the classical exclusion. And published clock policies are a regression discontinuity in time: DeepSeek’s peak windows switch tiers at fixed minutes, so queries arriving seconds apart on either side of the boundary face the same demand and a discontinuous tier shift, and the jump in the fingerprint rate at the boundary estimates the difference in djρd_jρ across the switched tiers. The churn side closes the same way: (prℓ)x(p_r )_x anchors to subscription price points, and a survival regression of churn on exposure to degraded episodes, instrumented by the same incidents, estimates prp_r. This program is why Section 5 declares its behavioral parameters as assumptions rather than estimates: the logs it needs exist, but they belong to the providers. What can be done from outside is what Section 5 does, anchor the assumptions to public prices and run the pipeline; what a provider can do from inside is replace every assumed number in Table 2 with a fingerprint rate and an instrumented regression, and then read Proposition 1 against its own menu. 5 Numerical examples This section runs the entire pipeline of the paper, statics, dynamics, duality, and policy distillation, on five instances calibrated to public benchmark data. The tier physics of each provider’s posted menu, mean service times, quality indices, and cost per task, are read from the Artificial Analysis leaderboard; DeepSeek’s concurrency limits, peak windows, and two-to-one peak/off-peak price ratio are taken from its published API documentation; and the behavioral primitives (ρx,prℓx)( _x,p_r _x) are declared assumptions anchored to public subscription price points, per the estimation program of Section 4.4. The five providers were chosen because their menus instantiate five distinct regimes of the theory: a steep reasoning-effort ladder (Anthropic), a published trap tier alongside a published clock policy (DeepSeek), a fast ladder where degradation buys throughput rather than energy (Google), a flat ladder where degradation buys nothing (Kimi), and a fourteen-variant lattice containing equal-quality tiers split across the energy and capacity frontiers (OpenAI). Everything below is a proof of concept in the sense declared in Section 1: the measurable skeleton of each instance is real, the behavioral parameters are assumptions, and the point is the pipeline, not the estimates. 5.1 Instances, menus, and the effective-dominance collapse Table 2 collects the calibration. Three customer classes, shared across all five providers because classes describe customers rather than vendors, carry the heterogeneity of Section 3.3: casual (low retry probability, low lifetime value), specialized (high retry, high value), and agentic (retries by construction), with dissatisfaction dj(x)d_j(x) linear in the relative quality gap of each ladder and d0(x)>0d_0(x)>0 for every class, the strong tier included. Demand follows a shared diurnal profile whose peak sits at 0.900.90 of the static-optimal-mix throughput θmix _mix, the “barely stable” operating point that rational provisioning implies, plus an overnight surge at 03:00 (the nineteenth hour of the horizon, which opens at 08:00), deliberately placed off the published peak windows, sized to exceed the comfortable mix while remaining survivable near the maximum-throughput action. Table 2: Calibrated instances. Markers give each number’s provenance: B benchmark (Artificial Analysis), P published by the provider, A assumption anchored to public prices, C calibration choice, D derived. Symbols as in Table 1, by class x (Section 3.3). customer classes (shared across providers) casual special. agentic shareA 0.60 0.25 0.15 share of fresh demand ρx _xA 0.30 0.85 0.98 probability an unsatisfied user re-asks d0,xd_0,xA 0.02 0.05 0.10 dissatisfaction at the strongest tier βx _xA 0.8 2.5 3.0 quality sensitivity: dj(x)=min0.95,d0,x+βx(1−Ij/I0)d_j(x)= \0.95,\,d_0,x+ _x(1-I_j/I_0)\ prp_rA 0.03 0.08 0.05 churn probability upon abandonment ℓ $240 $1,800 $2,000 lifetime margin of a churned customer prℓp_r $7.2 $144 $100 churn charge per abandonment, prℓp_r provider instances (one per regime of the theory) tiersB poolC θmix _mix/hrD actionsD regime Anthropic 5 2 000 190 943 125→ 9 steep reasoning-effort ladder OpenAI 14 2 000 40 258 2744→ 72 14-variant lattice, tiers split across the two frontiers Google 4 2 000 462 784 64→ 27 fast ladder: degradation buys throughput, not energy DeepSeek 3 3 500P 123 362 27→ 4 published trap tier, clock policy, per-tier pools Kimi 3 2 000 102 576 27→ 1 flat ladder: no lever; the policy space is a point shared calibration tier physicsB service time [Sj] E[S_j], quality index IjI_j, cost per task cjc_j, per tier demandC daily profile peaking at 0.90θmix0.90\, _mix; overnight surge at 03:00, sized min(1.5θmix, 0.9θmax) (1.5\, _mix,\,0.9\, _ ) peak windowsP 09–12 and 14–18; electricity price ×2× 2 inside them service levelC P(W>300s)≤0.05 P(W>300s)≤ 0.05, enforced by pruning cost scalesC markup 3×3×; CSLAC_SLA = one peak hour of tier-0 revenue; memory exponent κ=1.3κ=1.3; idle draw 30%30\% of tier-0 active The first output of the pipeline is static: before any dynamics, the per-satisfied-answer ledger of (5) prunes each provider’s menu class by class, discarding any tier that another tier weakly dominates in both effective slot-time S~j(x) S_j(x) and per-satisfied total cost. The collapse, reported in the last column of Table 2, is severe and informative. Anthropic’s 125 joint actions reduce to 9: the agent class is pinned to the single admissible tier Opus 5 (medium), priced out of the strong tier by slot-time and out of the weak tiers by its own retry multiplier, the two-sided exclusion of Proposition 6. DeepSeek’s 27 actions reduce to 4: the trap tier identified by Proposition 1 is inadmissible for every class, and, more striking, the strong tier is inadmissible for the majority (casual) class, whose entire effective menu is the Flash sibling. OpenAI’s 2744 actions reduce to 72, with the 221221-second flagship surviving only on the specialist menu. Kimi’s menu collapses to a single action, all classes on the strong tier: on a flat ladder where the cheap tiers are no faster, the retry multiplier makes every degradation a pure loss, and the correct policy space is a point. These admissible sets are computed from the benchmark numbers and the class primitives alone, no optimization involved; they are the trap proposition doing menu design. 5.2 Policies and protocol Five policies are compared on a common fine-grained simulator, all starting from the resting level of their own hour-0 action, so that cost accounting is identical and only the decision rule differs. P0 serves every query at the strong tier. P1 is the industry’s reactive stump, degrade everyone when total backlog crosses a trigger and release below a second threshold, with the degrade tier and both thresholds tuned per instance by rollout search, a deliberately generous opponent. P2 is the clock policy modeled on DeepSeek’s published mechanism, degrade everyone during the published peak windows regardless of state, again with the tier tuned. P3 is the fluid dynamic program of Section 3: hourly epochs, the closed-form two-regime transition of Lemma 2 with the arrival rate made endogenous by retries, a 13313^3 interpolated grid over the three class backlogs, and the chance constraint P(W>s)≤0.05 P(W>s)≤ 0.05 at an absolute threshold s=300s=300 seconds enforced by pruning, never by dollars. P4 is the depth-three decision tree distilled from P3 by the occupancy-weighted procedure of Section 4.3 and then rolled out as a policy in its own right, so that its value gap to the DP is a certified number rather than a fit statistic. Table 3: Twenty-four-hour economic cost by policy and provider; lower is better, and the fluid DP (bold) is on every instance simultaneously the cheapest policy and feasible at every hour. Demand is scaled to each provider’s own throughput, so dollar figures compare across policies within a row, not across providers. Baselines: always strongest serves every query at tier 0; the reactive threshold degrades everyone when total backlog crosses a trigger, releasing below a second threshold; scheduled degrades everyone during the published peak windows; both are tuned per instance (best degrade tier and thresholds by rollout search). †unstable: the backlog diverges under its own retry feedback (Proposition 3), so no finite cost is meaningful. kv: the policy violates the 300-second latency SLA in k of 24 hours. The final column reports the cost of the best feasible baseline as a multiple of the DP. provider always strongest reactive threshold scheduled fluid DP tree best/DP Anthropic unstable† $648.95M unstable† $9.38M $9.43M 69× OpenAI $110.56M15v $14.49M $60.52M4v $1.39M $1.39M 10× Google $7.27B2v $3.30B $7.27B2v $8.90M $8.86M 371× DeepSeek unstable† $49.51M4v unstable† $2.02M $2.02M 24× Kimi $3.44M $3.44M $1.47B9v $3.44M $3.44M 1.0× 5.3 Results Table 3 reports the twenty-four-hour economics and Figure 6 the backlog trajectories. Three patterns organize the table. First, under capacity provisioned to the optimal mix, not optimizing is not an option: the always-strong policy is unstable on Anthropic and DeepSeek, its backlog diverging under its own retry feedback, and where it survives it does so only at one to two orders of magnitude above the optimum or in violation of the latency constraint. Second, the two industry heuristics fail in exactly the modes the theory predicts. The clock policy is unstable on Anthropic and DeepSeek and, most cleanly, pathological on Kimi, where its scheduled degradation steps into a flat ladder and manufactures a recurring retry storm, the sawtooth of Figure 6 being the fire-and-release cycling of Corollary 4; it is also blind to the off-window surge by construction. The tuned stump stays feasible on most instances but pays for feasibility on the churn ledger: on Anthropic it books $648.95M against the DP’s $9.38M, with churned lifetime value of $57.20M against the DP’s $4.26M on identical demand, the uniform-degradation penalty of Proposition 7 in a single column. Third, the fluid DP is, on every instance, simultaneously the cheapest policy on the board and feasible at every hour, and its margin over the best feasible baseline, the last column of the table, varies with the regime exactly as the statics predict: large factors where the menu is rich (Anthropic at 69×, OpenAI at 10×), a factor of 371× on Google where the entire gap is surge timing, 24× on DeepSeek, and parity on Kimi, where the DP correctly recognizes that no lever exists and reproduces the always-strong policy to the dollar. The Kimi row is the null result that validates the method: an optimizer that found savings where the theory says none exist would be fitting noise. Figure 6: Demand and backlog under the policies of Table 3, five providers, one shared diurnal profile scaled to each provider’s throughput (hour of day; the day is a finite horizon, not a cycle, so the two edges need not match). Top panels: fresh demand, with the static-optimal-mix throughput θmix _mix dashed; the morning and afternoon plateaus are the published peak windows, and the overnight spike at 03:00 is the surge, outside every scheduled window and above θmix _mix. On Kimi the surge is a plateau rather than a spike: it is sized min(1.5θmix, 0.9θmax) (1.5\, _mix,\,0.9\, _ ), and on a flat ladder the all-strong action is already the maximum-throughput action, so the cap binds at the daily peak level. Bottom panels: the resulting backlogs (log scale), the closed-form forward pass of (7) under each policy’s logged hourly actions, validated against the Euler simulator’s hourly states to a maximum relative gap of 0.43%, diverging paths included; every excess of backlog growth over what fresh demand alone would produce is policy-manufactured retry traffic. The reactive threshold’s spike train on DeepSeek is the latch of Corollary 4 on published capacities: each release drops all classes back into the 500-slot Pro pool and the trigger re-fires. The distilled tree is omitted: visually identical to the fluid DP on every panel (value gaps in Table 3). The Google instance isolates the anticipation mechanism of Proposition 5. Its static-optimal mix is all-strong, so off the surge the DP and the always-strong policy agree; the surge exceeds the strong tier’s mixed throughput and is survivable only down-ladder, so degradation here buys throughput rather than energy. The DP pre-positions ahead of the 03:00 surge and glides through at a peak backlog orders of magnitude below the baselines, which absorb the same surge as a spike of hundreds of thousands of jobs and a multi-billion-dollar service-level bill. The same mechanism, run in reverse, appears on Kimi: because its maximum-throughput action is the all-strong action, the surge there punishes any policy caught degraded, which is precisely the clock policy’s failure. 5.4 The dual, computed Figure 7: The shadow price of intelligence, five providers: frozen-policy duals πx(t)=∂V⋆/∂λx(t) _x(t)=∂ V /∂ _x(t) of (9), finite differences through the closed-form forward pass under the DP’s actions; dotted verticals mark hours where the optimal action switches, where the dual may kink. Top panels: the specialized class’s dual decomposed into the direct cost of one more satisfied answer at the hour’s action, Mj(x)(cj+dj(x)(1−ρx)prℓx)M_j(x) (c_j+d_j(x)(1- _x)p_r _x ) with the peak electricity multiplier, and the congestion rent above it. Bottom panels: the rent by class, in dollars. Two facts organize the panels. The direct floors differ across classes at every hour, so the flat relative price of one that any uniform policy implicitly charges is wrong around the clock, off peak by 3.9× on Anthropic; and the pool charges a near-common rent when capacity binds, so dollar gaps between classes widen at the crunch while dual ratios compress. Rent panels are clipped where a single hour dominates, the value printed at the marker: on Google the marginal casual query at the surge carries $145 of service-level knock-ons, the dual at the capacity wall. DeepSeek and Kimi hold one action all day, the static optimum: no switches, and rent moves only with load. Duals in the last hours are biased low by the finite horizon, and the rent reaches the latency constraint through the smooth CSLAC_SLA surrogate rather than a hard wall. Figure 7 prices the marginal query, by class and by hour, on all five instances: the frozen-policy duals of (9), finite differences through the closed-form forward pass under the DP’s own actions, the entire surface recomputed in seconds. The two-part structure of Section 3.4 is visible on every panel. The floors never meet: off peak the specialized dual sits at 3.9× the casual dual on Anthropic, pure direct-cost spread with no congestion involved, so a router that prices all classes alike is mispricing at four in the morning, not just at the crunch. And the rent is the system’s, not the class’s: when load rises the three rent curves move together, the shared pool charging everyone the same scarcity premium, with the class identities carried almost entirely by the floors. The exception proves the mechanism: on Google, whose story is the surge, the marginal casual query at 03:00 carries $145 of service-level knock-ons, the dual at the capacity wall, four orders of magnitude above its off-peak floor. On Anthropic the rent peaks at the window shoulders rather than inside the windows: mid-window the DP has already degraded and capacity is cheap, while at the shoulders it is still running strong tiers into rising load, the anticipatory policy visible in the dual. Two disclosed biases: duals in the last hours of the day are truncated by the finite horizon, and the rent reaches the latency constraint through the smooth CSLAC_SLA surrogate, consistent with how every rollout in this section is priced. DeepSeek and Kimi hold the static-optimal action all day, so their duals carry no switch structure and their rents simply follow load, the dual-side reflection of the parity row in Table 3. 5.5 Certificates Table 4 certifies the fluid solution against a brute-force competitor: the identical dynamic program, grid, action set, and cost function, with the closed-form transition of Lemma 2 replaced by a converged Euler integration at the per-provider stable step. The fluid solver runs each instance in under three seconds against the competitor’s tens to hundreds, a speedup that grows with the provider’s service-rate scale, largest exactly where numerical integration is most expensive; this is the “computable in milliseconds” claim of Section 3.4 made concrete, and it is what makes the dual a live control signal rather than an overnight batch job. Two accuracy certificates accompany the speed claim. The policy-value certificate, reported in the table, rolls out the fluid and converged greedy policies on the same simulator and compares their twenty-four-hour costs; it is the deployment-relevant metric and closes the chain of custody, the fluid policy is ε -close to the converged optimum in rollout value, and the tree below is δ-close to the fluid policy, both epsilons printed. The trajectories themselves carry a third certificate: the closed-form forward pass of (7) matches the Euler simulator’s hourly states to a maximum relative gap of 0.43% (mean 0.008%) over every provider, policy, and hour mark, the diverging paths included; Figure 6 is drawn from that forward pass, so what the reader sees is the lemma’s formula, verified against the integrator on-trajectory, complementing the off-trajectory grid states the Q-regret table probes. The pointwise Q-regret statistics, reported in the appendix, price the fluid policy’s actions state by state under the converged Q-function; their tails concentrate off-trajectory in deep-saturation grid states on the shared-pool, high-rate providers, where the frozen-share approximation inside the closed form is weakest, and are reported in full rather than smoothed. Table 4: Runtime and accuracy of the fluid solver against a brute-force competitor: the identical dynamic program, grid, action set, and cost function, with the closed-form transition replaced by a converged Euler integration at the per-provider stable step. The policy-value gap compares twenty-four-hour rollout costs of the two greedy policies on a common simulator; pointwise Q-regret statistics are reported in Appendix B. provider fluid (s) Euler (s) speedup policy-value gap Anthropic 0.31 23 75× +5.89% OpenAI 2.41 274 114× +4.22% Google 0.95 287 302× +1.73% DeepSeek 0.13 11 78× +0.00% Kimi 0.04 3 69× +0.00% 5.6 The policy in normal form: five trees Figures 8 and 9 are the central artifact: the distilled policy, one depth-three tree per class, rendered in the same normal form as the baselines so that the comparison is visual, the stump is a single split on total backlog, the clock a single split on the peak flag, always-strong a single leaf, and the DP’s tree shows exactly which splits they are missing. Every tree has at most eight leaves, every certified rollout gap to the DP is small (0.54% and 0.20% on the instances shown), and the splits are the paper’s objects surfacing from data. The OpenAI casual tree opens on the frontier switch of Section 3.1: below a total-backlog threshold and off the overnight surge, route to Luna (max), the $0.05, 172172-second energy-frontier tier; past the threshold, route to Sol (medium), the $0.37, 1313-second capacity-frontier tier, the cheapest model on the menu becoming the most expensive thing to serve exactly when capacity binds. The Anthropic casual tree splits at its root on the demand forecast r^t r_t, degrading ahead of load, the anticipatory structure of Proposition 5 as a root node; its specialist tree splits on the peak flag, off-peak at moderate effort and peak hours governed by the two-to-one electricity price, the energy dual of Section 3.4 as literal tree structure; and its agent tree is a single leaf, a class whose effective menu is a singleton. The Google, DeepSeek, and Kimi trees are in Appendix B; the DeepSeek specialist tree carries nested backlog thresholds on its own queue, the index-rule structure of Proposition 6, and all three Kimi trees are single leaves with perfect fidelity, the correct policy tree for that provider having one leaf per class, proven rather than asserted. In the trees, r^t r_t is the demand forecast, t the hour of day, Ncas,Nspec,Nag,NtotN_cas,N_spec,N_ag,N_tot backlogs by class and in total, and C pool capacity. Anthropic models, strongest → weakest: Opus 5 max (tier 0) Opus 5 high (tier 1) Opus 5 medium (tier 2) Opus 5 low (tier 3) Sonnet 5 non-reasoning (tier 4) where class = casual is served, at every hour (fluid DP, distilled) forest where class = specialized is served, at every hour (fluid DP, distilled) forest where class = agentic is served, at every hour (fluid DP, distilled) forest baselines: one tree for every class reactive threshold scheduled (peak windows) always strongest forest forest forest Figure 8: Anthropic: the distilled routing, one tree per class, and the industry baselines in the same form. Baselines: reactive threshold degrades everyone above a backlog trigger (release at 0.5C0.5\,C, C the pool capacity); scheduled degrades during the published windows; always strongest never degrades. Trees may split on other classes’ backlogs, since a shared pool reprices capacity everywhere; splits with identical branches are merged. Certified rollout value gap of the tree to the DP: 0.54%. OpenAI models, strongest → weakest: Sol max (tier 0) Sol xhigh (tier 1) Sol medium (tier 4) Luna max (tier 6) where class = casual is served, at every hour (fluid DP, distilled) forest where class = specialized is served, at every hour (fluid DP, distilled) forest where class = agentic is served, at every hour (fluid DP, distilled) forest baselines: one tree for every class reactive threshold scheduled (peak windows) always strongest forest forest forest Figure 9: OpenAI: the distilled routing, one tree per class, and the industry baselines in the same form. Baselines: reactive threshold degrades everyone above a backlog trigger (release at 0.5C0.5\,C, C the pool capacity); scheduled degrades during the published windows; always strongest never degrades. Trees may split on other classes’ backlogs, since a shared pool reprices capacity everywhere; splits with identical branches are merged. Model names drop the shared prefix “GPT-5.6”. Certified rollout value gap of the tree to the DP: 0.20%. 6 Concluding Remarks The paper’s argument compresses to one accounting identity and its consequences. The customer buys an answer; the system prices a query; and the wedge between the two, a geometric retry multiplier, inverts the static economics of model tiers, converts a reactive throttle into a demand amplifier with an ignition threshold and a one-way latch, turns heterogeneous throttling into a transportation problem in retry-inflated load, and equips the whole with a dual variable that prices a marginal query by class and by hour. The machinery is known: the newsvendor, the multiplier, and the transient fluid queue are each decades old (Leontief, 1966; Eick et al., 1993; Halfin and Whitt, 1981). What the recomposition unlocks is a control problem that was previously tractable only by simulation: a failure probability that is a control rather than a primitive, closed over by the retry loop, and priced on both ledgers, the one the dashboard shows and the one it does not. The unlocking is multiplicative rather than singular, four collapses compounding: memorylessness collapses the state, the static ledger collapses the action space, the two-regime closed form collapses the transition, and the normal surrogate collapses the tail. Remove any one and the problem returns to overnight simulation. The boundary of the method is where the first collapse fails. Delayed retries restore the orbit as a state variable, making the dynamics distribution-dependent, a nonlinear Markov chain in the sense of Dieker et al. (2026) via the QPLEX-DQP formalism; that formulation is the natural home for exact tail certification in the critical band and for heavy-tailed token distributions, and it is exactly why the extension is a sequel rather than a section. What this paper’s first-order approximation delivers on its side of the boundary is the structure, the trap, the spiral, the index rule, and the dual, and most of the cost recoverable by controlling the system well; what lies beyond it is the sharper accounting of the tails. References Aflaki and Popescu (2014) S. Aflaki and I. Popescu Managing retention in service relationships. Management Science 60 (2), p. 415–433. Cited by: §1.1. Alizamir et al. (2013) S. Alizamir, F. de Véricourt, and P. Sun Diagnostic accuracy under congestion. Management Science 59 (1), p. 157–171. Cited by: §1.1. Anand et al. (2011) K. S. Anand, M. F. Paç, and S. Veeraraghavan Quality–speed conundrum: trade-offs in customer-intensive services. Management Science 57 (1), p. 40–56. Cited by: §1.1. Artalejo and Gómez-Corral (2008) J. R. Artalejo and A. Gómez-Corral Retrial queueing systems: a computational approach. Springer, Berlin. Cited by: §1.1. Ata and Shneorson (2006) B. Ata and S. Shneorson Dynamic control of an M/M/1 service system with adjustable arrival and service rates. Management Science 52 (11), p. 1778–1791. Cited by: §1.1. Chen et al. (2000) F. Chen, Z. Drezner, J. K. Ryan, and D. Simchi-Levi Quantifying the bullwhip effect in a simple supply chain: the impact of forecasting, lead times, and information. Management Science 46 (3), p. 436–443. Cited by: §1.1, §4.2. de Véricourt and Zhou (2005) F. de Véricourt and Y. Zhou Managing response time in a call-routing problem with service failure. Operations Research 53 (6), p. 968–981. Cited by: §1.1, Remark 2. Dieker et al. (2026) A. B. Dieker, S. T. Hackman, Z. Wang, and Y. Yan QPLEX decision processes: formulation via nonlinear Markov chains and optimization via policy gradients. Note: Preprint External Links: 2605.17149 Cited by: §1.1, §2, §4.1, §4.2, §6. Eick et al. (1993) S. G. Eick, W. A. Massey, and W. Whitt The physics of the Mt/G/∞M_t/G/∞ queue. Operations Research 41 (4), p. 731–742. Cited by: §3.2, §3.2, §4.1, §6. Falin and Templeton (1997) G. I. Falin and J. G. C. Templeton Retrial queues. Chapman & Hall, London. Cited by: §1.1. Gandhi et al. (2012) A. Gandhi, M. Harchol-Balter, R. Raghunathan, and M. A. Kozuch AutoScale: dynamic, robust capacity management for multi-tier data centers. ACM Transactions on Computer Systems 30 (4), p. 1–26. Cited by: §1.1. Gans (2002) N. Gans Customer loyalty and supplier quality competition. Management Science 48 (2), p. 207–221. Cited by: §1.1. Garnett et al. (2002) O. Garnett, A. Mandelbaum, and M. I. Reiman Designing a call center with impatient customers. Manufacturing & Service Operations Management 4 (3), p. 208–227. Cited by: §1.1. Halfin and Whitt (1981) S. Halfin and W. Whitt Heavy-traffic limits for queues with many exponential servers. Operations Research 29 (3), p. 567–588. Cited by: §1.1, §1, §2, §6. Hall and Porteus (2000) J. Hall and E. Porteus Customer service competition in capacitated systems. Manufacturing & Service Operations Management 2 (2), p. 144–165. Cited by: §1.1. Hassin and Haviv (2003) R. Hassin and M. Haviv To queue or not to queue: equilibrium behavior in queueing systems. Kluwer Academic Publishers, Boston. Cited by: §1.1. Hopp et al. (2007) W. J. Hopp, S. M. R. Iravani, and G. Y. Yuen Operations systems with discretionary task completion. Management Science 53 (1), p. 61–77. Cited by: §1.1. Kwon et al. (2023) W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), p. 611–626. Cited by: §1.1. Lee et al. (1997) H. L. Lee, V. Padmanabhan, and S. Whang Information distortion in a supply chain: the bullwhip effect. Management Science 43 (4), p. 546–558. Cited by: §1.1, §4.2. Leontief (1966) W. Leontief Input-output economics. Oxford University Press, New York. Cited by: §1, §2, §6. Lin et al. (2013) M. Lin, A. Wierman, L. L. H. Andrew, and E. Thereska Dynamic right-sizing for power-proportional data centers. IEEE/ACM Transactions on Networking 21 (5), p. 1378–1391. Cited by: §1.1. Mandelbaum and Massey (1995) A. Mandelbaum and W. A. Massey Strong approximations for time-dependent queues. Mathematics of Operations Research 20 (1), p. 33–64. Cited by: §1.1, §4.1. Naor (1969) P. Naor The regulation of queue size by levying tolls. Econometrica 37 (1), p. 15–24. Cited by: §1.1. Sanabria et al. (2021) E. Sanabria, D. D. Yao, and H. Lam Decision tree algorithms for MDP. Note: Working paper, Columbia University Cited by: §1.1, §4.3. Sanabria (2026) E. Sanabria Supply chain analytics: a data-driven approach. Note: Lecture notes, Lehigh University Cited by: §1.1. Takács (1963) L. Takács A single-server queue with feedback. Bell System Technical Journal 42 (2), p. 505–519. Cited by: §1.1, Remark 2. Van Mieghem (1995) J. A. Van Mieghem Dynamic scheduling with convex delay costs: the generalized cμcμ rule. The Annals of Applied Probability 5 (3), p. 809–833. Cited by: §1.1, §3.3. Whitt (2002) W. Whitt Stochastic-process limits: an introduction to stochastic-process limits and their application to queues. Springer, New York. Cited by: §1.1, §1, §3.2, §4.1. Yu et al. (2022) G. Yu, J. S. Jeong, G. Kim, S. Kim, and B. Chun Orca: a distributed serving system for transformer-based generative models. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI), p. 521–538. Cited by: §1.1. Zeltyn and Mandelbaum (2005) S. Zeltyn and A. Mandelbaum Call centers with impatient customers: many-server asymptotics of the M/M/n+G queue. Queueing Systems 51 (3–4), p. 361–402. Cited by: §1.1. Appendix A Proofs A.1 Proof of Proposition 5 Fix a horizon [0,T][0,T], a calm rate rcalm<θ0r_calm< _0, and a surge window [t1,t2][t_1,t_2] on which rt=rsurge>θ0r_t=r_surge> _0, with rsurge>θ1r_surge> _1 so the reactive rule fires above ignition and rsurgeS~1<mkr_surge\, S_1<mk so the surge is feasible under full degradation. Step 1: an anticipatory path that never ignites. Degrading a fraction φ∈(0,1] ∈(0,1] of traffic produces retry-inflated offered load ℓt(φ)=rt[(1−φ)S~0+φS~1] _t( )=r_t [(1- )\, S_0+ \, S_1 ]. Feasibility guarantees a φ and a start t0<t1t_0<t_1 with ℓt(φ)≤(1−δ)mk _t( )≤(1-δ)mk for all t and some δ>0δ>0, the lead t1−t0t_1-t_0 long enough for the pre-surge backlog to relax below ℓt1(φ) _t_1( ) along the exponential branch of (7). On this path the system never saturates, the service-level term of (1) is identically zero, and the energy saved equals γ(S~0Δw0−S~1Δw1)γ( S_0 w_0- S_1 w_1) per unit of degraded volume, the same volume the reactive rule degrades over the surge. Step 2: the reactive path pays superlinearly. By Proposition 3 the reactive trajectory crosses mkmk at some τ<t2τ<t_2, fires, and thereafter carries constant positive drift θ:=rsurge−θ1θ:=r_surge- _1. Its backlog at t∈[τ,t2]t∈[τ,t_2] is at least θ(t−τ)θ(t-τ), so its cumulative service-level cost is at least CSLA(mk)2∫τt2θ2(t−τ)2t=CSLAθ23(mk)2(t2−τ)3 C_SLA(mk)^2 _τ^t_2θ^2(t-τ)^2dt= C_SLAθ^23(mk)^2(t_2-τ)^3, cubic in the surge duration. The bill does not close at t2t_2: if rcalm<θ1r_calm< _1, the excess backlog B:=θ(t2−τ)B:=θ(t_2-τ) drains at rate β:=θ1−rcalmβ:= _1-r_calm and contributes a further CSLA(mk)2∫0B/β(B−βt)2t=CSLAθ33(mk)2β(t2−τ)3 C_SLA(mk)^2 _0^B/β(B-β t)^2\,dt= C_SLAθ^33(mk)^2β(t_2-τ)^3, the same cube scaled by θ/βθ/β; if rcalm≥θ1r_calm≥ _1, the degraded queue never drains at all and the reactive cost is unbounded regardless of surge length. Step 3: comparison. The anticipatory policy pays one charge the reactive avoids, churn over the lead window, of size P:=φrcalmd1(1−ρ)prℓ(t1−t0)P:= \,r_calm\,d_1(1-ρ)\,p_r \,(t_1-t_0), linear in the lead and independent of the surge length. Summing the two cubes of Step 2, the reactive excess is at least CSLAθ23(mk)2(1+θβ)(t2−τ)3 C_SLAθ^23(mk)^2 (1+ θβ )(t_2-τ)^3, and setting this equal to P and solving for t2−τt_2-τ gives the critical window T⋆T of the statement: for t2−τ<T⋆t_2-τ<T the reactive rule is cheaper and anticipation does not pay; for t2−τ>T⋆t_2-τ>T the anticipatory policy strictly dominates, with a margin growing as the cube of the excess. If the release threshold sits below the degraded calm equilibrium, Corollary 4 makes the reactive ledger unbounded and dominance strict at every T regardless of T⋆T . The chance-constraint surrogate. A job arriving at occupancy N waits behind max(N,mk) (N,mk) stages of exponential work served at aggregate rate kμmkμ m, so its waiting time has mean max(N,mk)/(kμm) (N,mk)/(kμ m) and variance max(N,mk)/(kμm)2 (N,mk)/(kμ m)^2; the normal approximation to P(W≤s)≥1−α P(W≤ s)≥ 1-α is exactly skμm≥max(N,mk)+max(N,mk)Φ−1(1−α)skμ m≥ (N,mk)+ (N,mk)\, ^-1(1-α), and pruning the violating pairs preserves feasibility of every surviving trajectory. □ A.2 Proof of Proposition 6 Attach a multiplier ν≥0ν≥ 0 to the binding capacity constraint of (8). The Lagrangian separates by class, and moving a unit of class-x traffic from tier j to j+1j+1 changes the objective by λx∂c(x,j)/∂j _x\,∂ c(x,j)/∂ j and relaxes the constraint by λx(S~j(x)−S~j+1(x)) _x ( S_j(x)- S_j+1(x) ); at the margin against tier 0, the relief per unit is S~0(x)−S~j(x) S_0(x)- S_j(x). A routing is optimal iff no swap of degraded volume between classes improves the Lagrangian, i.e., iff classes with positive relief are degraded in increasing order of the ratio I(x,j)I(x,j) of damage to relief, while classes with nonpositive relief, the trap region of Proposition 1, are never degraded at a binding constraint since doing so tightens it; ties are resolved arbitrarily and do not affect the value. Nestedness in the congestion state follows because ν is nondecreasing in N through the convexity of the service-level term in (1), so the set x:I(x,j)≤ν\x:I(x,j)≤ν\ is nondecreasing in N. □ A.3 Proof of Proposition 8 By Proposition 6 the optimal action at (t,N,x)(t,N,x) is determined by (i) the regime of t, calm or surge, through the arrival rate entering ν; (i) the position of I(x,⋅)I(x,·) in the class ordering; and (i) the position of N relative to the at most J thresholds at which ν crosses the successive indices of class x. A tree that splits first on regime, then on class, then on the class-specific thresholds in N, reproduces this map exactly, with at most #regimes×||×J\#regimes×| X|× J leaves; ϵε-exactness with fewer leaves follows by merging any adjacent cells whose Q-gap is below ϵε under the occupancy measure. That the distillation of Section 4.3 recovers it follows because the fit minimizes occupancy-weighted regret over a hypothesis class containing this tree, and the closed-loop certificate verifies attainment ex post; the collapse to depth one below the break-even expected LTV loss of Remark 1 is the case of a single effective tier with the N-threshold at infinity. □ Appendix B Additional numerical assets Google models, strongest → weakest: 3.7 Flash high (tier 0) 3.7 Flash medium (tier 1) 3.7 Flash low (tier 3) where class = casual is served, at every hour (fluid DP, distilled) forest where class = specialized is served, at every hour (fluid DP, distilled) forest where class = agentic is served, at every hour (fluid DP, distilled) forest baselines: one tree for every class reactive threshold scheduled (peak windows) always strongest forest forest forest Figure 10: Google: the distilled routing, one tree per class, and the industry baselines in the same form. Baselines: reactive threshold degrades everyone above a backlog trigger (release at 0.5C0.5\,C, C the pool capacity); scheduled degrades during the published windows; always strongest never degrades. Trees may split on other classes’ backlogs, since a shared pool reprices capacity everywhere; splits with identical branches are merged. Model names drop the shared prefix “Gemini”. Certified rollout value gap of the tree to the DP: -0.43%. DeepSeek models, strongest → weakest: V4-Pro-0813 thinking (tier 0) V4-Flash-0731 thinking (tier 1) where class = casual is served, at every hour (fluid DP, distilled) forest where class = specialized is served, at every hour (fluid DP, distilled) forest where class = agentic is served, at every hour (fluid DP, distilled) forest baselines: one tree for every class reactive threshold scheduled (peak windows) always strongest forest forest forest Figure 11: DeepSeek: the distilled routing, one tree per class, and the industry baselines in the same form. Baselines: reactive threshold degrades everyone above a backlog trigger (release at 0.5C0.5\,C, C the pool capacity); scheduled degrades during the published windows; always strongest never degrades. Trees may split on other classes’ backlogs, since a shared pool reprices capacity everywhere; splits with identical branches are merged. Certified rollout value gap of the tree to the DP: 0.00%. Kimi models, strongest → weakest: K3 max (tier 0) K3 low (tier 1) where class = casual is served, at every hour (fluid DP, distilled) forest where class = specialized is served, at every hour (fluid DP, distilled) forest where class = agentic is served, at every hour (fluid DP, distilled) forest baselines: one tree for every class reactive threshold scheduled (peak windows) always strongest forest forest forest Figure 12: Kimi: the distilled routing, one tree per class, and the industry baselines in the same form. Baselines: reactive threshold degrades everyone above a backlog trigger (release at 0.5C0.5\,C, C the pool capacity); scheduled degrades during the published windows; always strongest never degrades. Trees may split on other classes’ backlogs, since a shared pool reprices capacity everywhere; splits with identical branches are merged. Model names drop the shared prefix “Kimi”. Certified rollout value gap of the tree to the DP: 0.00%. Table 5: Pointwise Q-regret of the fluid policy under the converged Q-function, on mutually feasible state–action pairs (uniform mean and p95, occupancy-weighted mean), with the feasibility-disagreement rate. Tails concentrate off-trajectory in deep-saturation grid states on the shared-pool, high-μ providers; the deployment-relevant policy-value certificate is in Table 4. provider mean p95 occ. mean feas. disagr. Anthropic 1.84% 7.71% 3.33% 4.54% OpenAI 1.25% 5.02% 1.93% 0.00% Google 8.05% 39.53% 11.66% 1.51% DeepSeek 1.49% 1.90% 1.65% 0.00% Kimi 0.00% 0.00% 0.00% 0.00%