Paper deep dive
Photons = Tokens: The Physics of AI and the Economics of Knowledge
Alec Litowitz, Nick Polson, Vadim Sokolov
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/20/2026, 3:24:58 PM
Summary
This paper applies thermodynamic and information-theoretic principles to the economics of AI, defining the 'token' as a physical quantity with measurable energy costs. Using Landauer's principle and Shannon's channel capacity, the authors construct a supply-and-demand balance sheet for global token production, estimating that projected 2028 US AI energy allocations could support roughly 6.5Ă10^17 tokens per year. The study argues that while physical constraints on token production are significant, the binding constraint is not computational capacity but the problem of agencyâdetermining which questions are worth asking. It connects Goodhart's law to the Heisenberg uncertainty principle and applies Coase's theory of the firm to the AI value chain.
Entities (17)
Relation Signals (9)
John von Neumann â authored â The Computer and the Brain
confidence 95% ¡ John von Neumann prepared what would be his final work: The Computer and the Brain
David MacKay â authored â Sustainable EnergyâWithout the Hot Air
confidence 95% ¡ David MacKay published Sustainable EnergyâWithout the Hot Air
Julian Simon â challenged â Paul Ehrlich
confidence 95% ¡ In 1980, Simon challenged biologist Paul Ehrlich to choose any raw materials and any future date
Token â hasphysicalcost â Landauer's principle
confidence 95% ¡ We define the token... as a physical quantity with measurable thermodynamic cost. Using Landauer's principle...
Token â isconstrainedby â Shannon's channel capacity
confidence 95% ¡ Using Landauer's principle, Shannon's channel capacity... we construct a supply-and-demand balance sheet
AI Value Chain â analyzedby â Coase's theory of the firm
confidence 90% ¡ We apply Coaseâs theory of the firm... to the AI value chain
DeepSeek â releasedmodelwith â low_inference_cost
confidence 90% ¡ DeepSeek released a frontier-class model at inference costs roughly an order of magnitude below prevailing rates
Token Economy â isconstrainedby â Arrow's impossibility theorem
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Debates about artificial intelligence capabilities and risks are often conducted without quantitative grounding. This paper applies the methodology of MacKay (2009) -- who reframed energy policy as arithmetic -- to the economy of AI computation. We define the token, the elementary unit of large language model input and output, as a physical quantity with measurable thermodynamic cost. Using Landauer's principle, Shannon's channel capacity, and current infrastructure data, we construct a supply-and-demand balance sheet for global token production. We then derive a finite question budget: the number of meaningful queries humanity can direct at AI systems under physical, information-theoretic, and economic constraints. We apply Coase's theory of the firm and the durable-goods monopoly problem to the AI value chain -- from photon to atom to chip to power to token to question -- to identify where economic value concentrates and where regulatory intervention is warranted. We argue that the expansion of the token budget does not resolve a deeper constraint: under structural uncertainty, the decisive variable is not how many questions can be answered but which questions are worth asking -- a problem of agency and direction that computation alone cannot solve. We connect limits of measurement in the token economy to a structural parallel between Goodhart's law and the Heisenberg uncertainty principle, and to Arrow's impossibility result for efficient information pricing. The framework yields order-of-magnitude estimates that discipline policy discussion: at current efficiency, the projected 2028 US AI energy allocation of 326~TWh could support roughly $6.5 \times 10^{17}$ tokens per year, or 225,000 tokens per person per day -- more than three orders of magnitude above estimated mid-2024 utilization.
Tags
Links
- Source: https://arxiv.org/abs/2603.06630v1
- Canonical: https://arxiv.org/abs/2603.06630v1
Trouble viewing inline? Open PDF directly â
Full Text
92,421 characters extracted from source content.
Expand or collapse full text
Photons = Tokens: The Physics of AI and the Economics of Knowledge Alec Litowitz QStar Capital. https://w.qstarcap.com Nick Polson University of Chicago. ngp@chicagobooth.edu Vadim Sokolov George Mason University. vsokolov@gmu.edu (First Draft: February 8, 2026 This Draft: ) Abstract Debates about artificial intelligence capabilities and risks are often conducted without quantitative grounding. This paper applies the methodology of MacKay (2009)âwho reframed energy policy as arithmeticâto the economy of AI computation. We define the token, the elementary unit of large language model input and output, as a physical quantity with measurable thermodynamic cost. Using Landauerâs principle, Shannonâs channel capacity, and current infrastructure data, we construct a supply-and-demand balance sheet for global token production. We then derive a finite question budget: the number of meaningful queries humanity can direct at AI systems under physical, information-theoretic, and economic constraints. We apply Coaseâs theory of the firm and the durable-goods monopoly problem to the AI value chainâfrom photon to atom to chip to power to token to questionâto identify where economic value concentrates and where regulatory intervention is warranted. We argue that the expansion of the token budget does not resolve a deeper constraint: under structural uncertainty, the decisive variable is not how many questions can be answered but which questions are worth askingâa problem of agency and direction that computation alone cannot solve. We connect limits of measurement in the token economy to a structural parallel between Goodhartâs law and the Heisenberg uncertainty principle, and to Arrowâs impossibility result for efficient information pricing. The framework yields order-of-magnitude estimates that discipline policy discussion: at current efficiency, the projected 2028 US AI energy allocation of 326 TWh could support roughly 6.5Ă10176.5Ă 10^17 tokens per year, or 225,000 tokens per person per dayâmore than three orders of magnitude above estimated mid-2024 utilization. 1 Introduction In 1957, dying of cancer, John von Neumann prepared what would be his final work: The Computer and the Brain, the Silliman Lectures at Yale that he was too ill to deliver (von Neumann, 1958). In eighty pages, he argued that computation is a physical process. The brain and the digital computer are both engines for transforming energy into information, and the thermodynamic costs of this transformation are irreducible. A neuron dissipates roughly 10â310^-3 ergs per operation; a vacuum tube of his era, roughly 10210^2 ergsâa factor of 10510^5. The brain compensates for its slow, imprecise components through massive parallelism; the computer compensates through raw speed. In both systems, the fundamental operations are the same: storage and retrieval of information, governed by the laws of physics. Von Neumann died before he could finish the manuscript, but the thesis was complete: intelligence, whether natural or artificial, is computation, and computation is physics. Half a century later, David MacKay published Sustainable EnergyâWithout the Hot Air, a book that reframed energy policy as a problem of arithmetic (MacKay, 2009). His method was simple: convert every energy source and every energy demand into the same unitâkilowatt-hours per day per personâand check whether the numbers balance. The debate over artificial intelligence needs both von Neumannâs insight and MacKayâs method. Von Neumann established that computation is physics; MacKay showed how to do the accounting. This paper applies MacKayâs arithmetic to the computational economy that von Neumann foresaw, with the tokenâthe elementary unit of large language model input and outputâplaying the role the kilowatt-hour played in MacKayâs analysis. A token is a subword unit, typically three to four characters of English text. Generating one requires a forward pass through a neural network: a physical computation on silicon, consuming electricity, dissipating heat. The analogy to energy is not metaphorical. Power in physics is energy per unit time, measured in watts. Power in AI is information processed per unit time, measured in tokens per second. Both are physical quantities, both are finite, and both can be accounted for. The intellectual lineage runs through cybernetics and the theory of computation. Turing (1936) established the mathematical limits of computation: there exist well-defined problems that no finite procedure can solve, and every computable function can be realized by a universal machine. Wiener (1948) unified the study of control and communication in animals and machines under the principle that feedback requires information, and information has physical cost. von Neumann (1966) demonstrated that self-reproducing automata require a minimum complexity thresholdâand that arbitrarily reliable systems can be built from unreliable components through redundancy. Shannon (1948) established the mathematical theory of communication, quantifying information in bits and proving that every physical channel has a finite capacity. Wiener (1950) drew the social implication: if machines process information as organisms do, then the consequences of automation are as much questions of physics as of policy (Prigogine, 1978). Strubell et al. (2019) estimated the carbon footprint of training a large NLP model; Patterson et al. (2021) extended this to Google-scale training; Luccioni et al. (2024) shifted focus to inference. Our contribution differs in scope: rather than measuring individual models, we construct a balance sheet for the entire token economy. This paper develops five results. First, we construct a token budget: a MacKay-style balance sheet of global AI token supply and demand, grounded in current energy infrastructure data and physical efficiency limits (Sections 2 and 3). Second, we apply Coaseâs theory of the firm and the durable-goods monopoly problem to the AI value chainâfrom photon to atom to chip to power to token to questionâto identify where economic value concentrates and why it migrates upward in the stack (Section 4). Third, we derive a question budget: the finite number of meaningful queries that can be directed at AI systems, connecting Coxâs theory of questions to Shannon entropy and Landauerâs thermodynamic cost (Section 5). Fourth, we argue that physical and informational abundance does not resolve the problem of agency: under structural uncertainty, the binding constraint is not how many questions can be answered but which questions are worth askingâa problem of direction, path dependence, and judgment that computation alone cannot solve (Section 6). Fifth, we identify a measurement limit on the token economyâa structural parallel between Goodhartâs law and the Heisenberg uncertainty principleâand connect it to Arrowâs impossibility of efficient information pricing (Section 7). The central question is simple: How many questions is the world allowed to ask? 2 The Physics of Tokens A token is a physical object. Its generation requires a sequence of matrix multiplications on semiconductor hardware, consuming electrical energy and dissipating heat. Landauer (1961) proved that any logically irreversible computationâsuch as erasing a bit of informationâmust dissipate at least Emin=kBâTâlnâĄ2E_ =k_BT 2 (1) of energy into the environment, where kBâ1.38Ă10â23k_Bâ 1.38Ă 10^-23 J/K is Boltzmannâs constant and T is the absolute temperature. At room temperature (Tâ300Tâ 300 K), this gives Eminâ2.87Ă10â21E_ â 2.87Ă 10^-21 J per bit erased. The resolution of Maxwellâs demon paradox confirms this bound: any agent that reduces entropy in one part of a system must pay for the information it acquires by increasing entropy elsewhere (Bennett, 2003). Computation is not free; it is a thermodynamic process. A token drawn from a vocabulary of size V carries at most log2âĄV _2V bits of information. For a typical LLM vocabulary of Vâ100,000Vâ 100,000, this gives log2âĄ(100,000)â16.6 _2(100,000)â 16.6 bits per token. In practice, the information content is lower: Shannon (1951) estimated the entropy rate of printed English at 1.0â1.3 bits per character, implying 3â5 bits per subword token. The Landauer floor therefore lies between âź10â20 10^-20 J (at the Shannon entropy of âź 4 bits) and âź5Ă10â20 5Ă 10^-20 J (at the full vocabulary of 16.6 bits)âa factor of five, negligible relative to the 101910^19-fold efficiency gap. We adopt 12 bits per token as a practical working estimate for forward-pass computation, which must evaluate the full output distribution even though only a fraction of bits carry information. The Landauer floor for one token is then ELandauertoken=12ĂkBâTâlnâĄ2â3.4Ă10â20â J.E_Landauer^token=12Ă k_BT 2â 3.4Ă 10^-20 J. (2) 2.1 The Actual Cost Current LLMs operate far above this floor. Empirical measurements place the energy cost of inference at approximately 0.00010.0001â0.0020.002 Wh per output token, depending on model size (Epoch AI, 2025). For a mid-range estimate of 5Ă10â45Ă 10^-4 Wh per token, converting to joules: Eactualtokenâ5Ă10â4Ă3600=1.8â J.E_actual^tokenâ 5Ă 10^-4Ă 3600=1.8 J. (3) The ratio of actual to theoretical minimum cost is EactualtokenELandauertoken=1.83.4Ă10â20â5Ă1019. E_actual^tokenE_Landauer^token= 1.83.4Ă 10^-20â 5Ă 10^19. (4) Current AI hardware is roughly 101910^19â102010^20 times less efficient than the thermodynamic limit. This gap is enormous but decomposable. Modern CMOS logic gates dissipate roughly 10310^3â10610^6 times the Landauer limit per switching operation; the additional orders of magnitude arise because generating a single token requires on the order of 101210^12 floating-point operations, each involving many gate-level bit erasures. The gap represents the combined overhead of irreversible logic, memory access, data movement, cooling, and power conversion. But the dollar cost of inference is falling fast. Appenzeller (2024) documents what he terms âLLMflationâ: for an LLM of equivalent capability, inference cost has declined approximately 10Ă10Ă per year. Between November 2021 and late 2024, the dollar cost of achieving a fixed benchmark score fell from $60 to $0.06 per million tokensâa 1,000Ă1,000Ă reduction in three yearsâdriven by six independent factors: GPU hardware improvements, model quantization (from 16-bit to 4-bit arithmetic), software optimizations, smaller models achieving equivalent performance, improved training techniques (RLHF, DPO), and open-source competition compressing margins. The first four factors reduce energy per token and expand the physical token budget; the last two reduce dollar cost without reducing energy consumption. Each successive order of magnitude is harder to achieve, and the Landauer floor remains absolute. These efficiency gains, however, do not imply reduced total energy consumption. Jevons (1865) identified the mechanism in 1865: when technological improvement reduces the cost of using a resource per unit of output, total consumption of that resource tends to increase, not decrease, because cheaper use induces greater demand. Jevons observed that Wattâs fuel-efficient steam engine did not conserve coalâit made coal economically useful in industries that had never used it, and Englandâs coal consumption soared. âIt is a confusion of ideas to suppose that the economical use of fuel is equivalent to diminished consumption. The very contrary is the truth.â The Jevons paradox applies directly to the token economy. In January 2025, the Chinese laboratory DeepSeek released a frontier-class model at inference costs roughly an order of magnitude below prevailing rates, demonstrating that open-weight models with aggressive quantization and architectural efficiency could match proprietary systems at a fraction of the price (DeepSeek-AI, 2024). The immediate effect was not reduced AI energy consumption but an explosion of demand: cheaper tokens made AI economically viable for applicationsâbulk document processing, real-time translation, continuous code generationâthat had been priced out at higher per-token costs. LLMflation is the AI economyâs Watt steam engine: each order-of-magnitude cost reduction expands the addressable market by more than an order of magnitude, so total energy consumption rises even as energy per token falls. The inefficiency compounds further in practice. GPU clusters achieve a Model FLOPS Utilization (MFU) of only 45â55% of theoretical hardware throughput (Nebius AI, 2025). Hardware failures in a 3,000-GPU training cluster occur every âź 10 hours on average (Sivathanu et al., 2024), and recovery, checkpointing, and rollback add approximately 30% overhead to wall-clock training time. That reliable computation emerges at all from such failure-prone hardware is a consequence of von Neumannâs redundancy theorem: arbitrarily reliable systems can be built from components with error probability p<1/2p<1/2. Modern AI clusters implement this through checkpoint-and-restart protocols, spare node pools, and error-correcting interconnects. These overheads apply primarily to training; inference workloads have different utilization and failure profiles. The distinction matters for the balance sheet in Section 3. Training a frontier model is a one-time, high-intensity burst: a single training run may consume tens of gigawatt-hours over weeks or months (Strubell et al., 2019; Patterson et al., 2021). Inference is a continuous, low-intensity flow: each query costs a fraction of a watt-hour, but the queries never stop. As of 2024, inference accounts for an estimated 60â70% of total AI electricity consumption, up from roughly one-third in 2023, and is projected to reach approximately two-thirds by 2026 (International Energy Agency, 2025; Luccioni et al., 2024). The energy budget is increasingly dominated not by the cost of creating intelligence but by the cost of using it. The balance sheet uses the inference figure (5Ă10â45Ă 10^-4 Wh per token). 2.2 Shannonâs Ceiling Shannon (1948) established an upper bound: the channel capacity theorem. For any channel with bandwidth B and signal-to-noise ratio S/NS/N, the maximum rate of reliable information transmission is C=Bâlog2âĄ(1+SN)â bits per second.C=B _2\! (1+ SN ) bits per second. (5) Token throughputâAI âpowerââcannot exceed the channel capacity of the hardware interconnects, memory buses, and network links (MacKay, 2003, Ch. 1). 2.3 The Bekenstein Bound Bekenstein (1981) showed that the information in any bounded region of space is finite. For a system of energy E enclosed in a sphere of radius R, the maximum entropy is Sâ¤2âĎâkBâEâRââc,S⤠2Ď k_BER c, (6) where â is the reduced Planck constant and c is the speed of light. Lloyd (2000) showed that a 1-kg system confined to a 1-liter volume can store at most âź1031 10^31 bits and perform at most âź1051 10^51 operations per second. The point is qualitative: the information content and processing rate of any physical system are finite. The Landauer floor sets the energy cost per bit; Shannonâs capacity sets the information rate; the Bekenstein bound sets the absolute information content. 3 The Token Budget We construct a balance sheet for the token economy. The demand side uses global figures (AI services are consumed worldwide). The supply side focuses on the United States, which hosts an estimated 50â60% of global AI compute infrastructure and, through cloud providers headquartered there, serves a predominantly global user base (International Energy Agency, 2025). Dividing US energy capacity by the world population therefore yields a lower bound on per-person token availabilityâthe true figure is higher once non-US data centers (Europe, China, the Middle East) are included. Where US and global figures are mixed, we note the scope explicitly. 3.1 Demand: How Many Tokens Does the World Consume? Estimating global AI inference volume is difficult: providers disclose selectively, definitions vary (output tokens alone versus all tokens processed), and demand has been growing rapidly. In February 2024, OpenAI alone reported generating approximately 100 billion words per dayâroughly 1.3Ă10111.3Ă 10^11 output tokens from a single provider (Altman, 2024). By mid-2025, Googleâs AI models were processing approximately 980 trillion tokens per month (Alphabet Inc., 2025), and an aggregate cross-provider analysis placed global inference volume at 101310^13â101410^14 tokens per day (OpenRouter and Andreessen Horowitz, 2025). For the balance sheet, we adopt a mid-2024 estimate of approximately 101210^12 tokens per day globally, consistent with the provider-level data and the rapid growth trajectory documented by Epoch AI (2025). In MacKayâs per-person terms: 1012â tokens/day8Ă109â peopleâ125â tokens/person/day. 10^12 tokens/day8Ă 10^9 peopleâ 125 tokens/person/day. (7) At 0.75 words per token, 125 tokens is roughly 94 wordsâabout a paragraph. The average person speaks approximately 16,000 words per day (Mehl et al., 2007). Behind these inference tokens lies a prior investment: the re-encoding of existing information. Search engines indexed the web by cataloguing links between pages; large language models re-encode itâcompressing the vast corpus of human text into parametric representations that can be queried in natural language. Training a frontier model on trillions of tokens is, in information-theoretic terms, a lossy compression of the web into the modelâs weights. Inference then decompresses on demand: each user query extracts a targeted reconstruction from the compressed representation. The energy cost of encoding (training) must be amortized across the inference tokens it enables. AI-specific electricity consumption in the United States was 53â76 TWh in 2024, projected to reach 165â326 TWh by 2028 (International Energy Agency, 2024, 2025). If efficiency remains at its 2024 levelâa deliberately conservative assumption given the rapid gains documented in Section 2âand the upper-bound projection (326 TWh) materializes, the 2028 token capacity would be 326Ă1012â Wh5Ă10â4â Wh/token=6.5Ă1017â tokens/yearâ1.8Ă1015â tokens/day. 326Ă 10^12 Wh5Ă 10^-4 Wh/token=6.5Ă 10^17 tokens/yearâ 1.8Ă 10^15 tokens/day. (8) Per person, this gives 1.8Ă10158Ă109â225,000â tokens/person/day, 1.8Ă 10^158Ă 10^9â 225,000 tokens/person/day, (9) or roughly 169,000 words (at 0.75 words per token)âa novel per day. 3.2 Supply: What Can the Infrastructure Provide? The supply side is constrained by energy infrastructure. The total US electrical generation capacity is approximately 1,250 GW (including distributed generation), producing 4,178 TWh in 2023 (U.S. Energy Information Administration, 2024). Table 1 presents the balance sheet. Table 1: Token economy balance sheet, United States. The 2024 observed figure (âź 125 tokens/person/day) is estimated global demand divided by world population (see text for sources and uncertainty); capacity figures assume all AI electricity is used for inference at 5Ă10â45Ă 10^-4 Wh/token. In practice, training consumes 30â40% of AI electricity (declining as inference grows), so realized inference capacity is correspondingly lower. AI electricity midpoint: 65 TWh is the center of the 53â76 TWh range. 2024 2028 (proj.) Units Demand (observed) and capacity (projected) AI electricity consumption 65 326 TWh/yr Token capacity (at 5Ă10â45Ă 10^-4 Wh/token) 1.3Ă10171.3Ă 10^17 6.5Ă10176.5Ă 10^17 tokens/yr Observed tokens per person per day âź 125â â tokens/person/day Capacity per person per day 44,500 225,000 tokens/person/day Supply (grid) Total US electricity generation 4,178 âź 4,500 TWh/yr AI share of US grid 1.6% 7.2% % Maximum tokens at current efficiency 8.4Ă10188.4Ă 10^18 9.0Ă10189.0Ă 10^18 tokens/yr Physical limits Landauer-limited tokens (entire US grid) âź1038 10^38 tokens/yr Efficiency gap (actual / Landauer) âź5Ă1019 5Ă 10^19 Ă âMid-2024 estimate of âź1012 10^12 tokens/day; see text for sources and uncertainty. The table reveals three facts. First, AIâs share of the US electricity grid is growing rapidly but remains a single-digit percentage, even under aggressive projections. Second, the current utilization gapâobserved output of âź 125 tokens/person/day versus an energy-implied capacity of 44,500âsuggests that the binding constraint in 2024 is hardware deployment, not energy. Third, the gap between current practice and the Landauer limit (âź1019 10^19) implies that enormous efficiency gains are thermodynamically possible. A fourth observation is implicit: AI companies benefit from an externality embedded in electricity regulation itself. US industrial electricity pricesâapproximately 7â8 cents per kWh (U.S. Energy Information Administration, 2024)âare the product of a century of public utility regulation: rate-of-return constraints, public investment in generation and transmission, and cross-subsidies between customer classes that keep industrial rates below the marginal social cost of generation. The AI industry purchases this regulated input and sells tokens at unregulated market prices. At 326 TWh and $0.07/kWh, the projected 2028 AI electricity bill is roughly $23 billionâa small fraction of the revenue the token economy is expected to generate. The implicit subsidy flows from ratepayers and the environment (which bears the unpriced externalities of generation) to AI shareholders. This is the same structure that has historically benefited aluminum smelters and steel mills, but the growth rate of AI electricity demandâdoubling every two to three yearsâmeans the subsidy is scaling faster than any precedent in industrial energy consumption. 3.3 The Simon Wager Revisited The question of whether physical resources constrain the token economy has a precedent. In 1972, the Club of Rome published The Limits to Growth, projecting that exponential demand for commodities would exhaust key resources within decades (Meadows et al., 1972). Economist Julian Simon argued the opposite: human ingenuity, operating through price signals, would ensure that resources never run out in any economically meaningful sense (Simon, 1981). In 1980, Simon challenged biologist Paul Ehrlich to choose any raw materials and any future date; Simon would bet that inflation-adjusted prices would fall. Ehrlich and colleagues selected five metalsâchromium, copper, nickel, tin, and tungstenâand bought $200 of each on paper, for a total stake of $1,000, indexed to September 29, 1980, with a payoff date of September 29, 1990. Between 1980 and 1990 the worldâs population grew by more than 800 millionâthe largest single-decade increase in historyâyet by September 1990 every metal had fallen in real terms. Tin dropped from $8.72 to $3.88 per pound; tungsten fell by more than half. In October 1990, Ehrlich mailed Simon a check for $576.07 (Simon, 1980). For four decades, Simonâs position appeared vindicated. Commodity prices remained stable or fell as substitution, recycling, and discovery outpaced demand. But the victory was sensitive to timing: asset manager Jeremy Grantham noted that had the wager run from 1980 to 2011, Simon would have lost on four of the five metals (Grantham, 2011). The AI buildout may represent the scenario in which the Club of Romeâs arithmetic finally bindsânot on the original decade-long horizon, but on the longer one that Grantham identified. That copper was one of Simonâs five metals gives the parallel particular force. The numbers are concrete (Goldman Sachs Research, 2024). A conventional data center requires 5,000â15,000 tons of copper; a hyperscale AI facility requires up to 50,000 tons, with intensity ranging from 27 to 66 tons per megawatt. In 2024, data centers consumed approximately 500,000 tons of copper worldwide. Projections place this at 1.1 million tons annually by 2030âroughly 4% of global demandâand 3 million tons by 2050, a sixfold increase. These demands arrive simultaneously with competing requirements from electric vehicles, renewable energy, and grid electrification; global copper demand is projected to grow from 28 million metric tons in 2025 to 42 million by 2040 (International Energy Agency, 2025). The supply side is equally constrained. Decades of underinvestment in exploration have produced a declining rate of major discoveries since the 1980s, while existing mines face steadily falling ore gradesâlower grades mean more energy, more water, and higher costs per ton of refined metal. A major new copper mine takes 15 to 20 years from discovery to production. The industry would need approximately six new tier-one mines to come online every year through 2050 merely to meet baseline demand growth, before accounting for the incremental demands of electrification and AI (Giustra, 2025). Copper prices have already reflected this arithmetic, rising from approximately $3.80 per pound in 2023 to nearly $6.00 in early 2025. Mining executive Robert Friedland has stated the constraint in its starkest form: merely to sustain 3% global GDP growthâbefore accounting for electrification, data centers, or renewable energyâthe world must mine as much copper in the next two decades as it has extracted in the past ten thousand years (Giustra, 2025). The AI buildout could create genuine scarcityânot because of population growth, as the Club of Rome predicted, but because of the physical demands of the token economy. Simon was right that ingenuity can substitute around resource constraints, but substitution requires time, capital, and physical alternatives. A copper mine takes two decades; a GPU generation takes two years. Whether the copper, energy, and fabrication capacity exist to support a thousand-fold expansion of the token economy in four years is an empirical questionâand the arithmetic should be done before the commitments are made. 3.4 The MacKay Lesson MacKayâs central insight was that energy debates become productive only when reduced to arithmetic. The same holds for AI. The cost of not doing the arithmetic is illustrated by a cautionary episode from the energy debate itself. In 2009âthe same year MacKay published his bookâLevitt and Dubner (2009) argued in Superfreakonomics that solar panels were counterproductive because, being black, they absorb sunlight and radiate waste heat. Pierrehumbert (2009) showed that elementary arithmetic demolished the claim: the worldâs electricity needs require roughly 53,000 km2 of solar cellsâabout 0.01% of Earthâs surfaceâproducing waste heat that would warm the planet by approximately 0.006âC, roughly 300 times smaller than the warming from the CO2 emissions they would displace. Moreover, waste heat is a one-time effect upon installation, whereas CO2 warming accumulates for centuries; after 100 years of operation, the avoided greenhouse warming exceeds the waste heat effect by a factor of 125. The error was not subtle; it was a failure to multiply. Pierrehumbertâs observationâthat if one cannot execute basic energy accounting, the rest of the analysis cannot be trustedâapplies with equal force to AI policy claims advanced without quantitative grounding. The projected 2028 capacity of 225,000 tokens per person per day is not negligible. Whether the infrastructure can support this growthâand whether it shouldâcan only be answered with numbers. 4 The Value Stack The token budget tells us how many tokens exist. It does not tell us where economic value concentrates. Coase (1937) asked: if markets are efficient, why do firms exist? His answer was transaction costs. Using the price mechanismâsearching for counterparties, negotiating contracts, enforcing agreementsâis not free. Firms emerge when the cost of organizing a transaction internally is lower than the cost of conducting it through the market. The boundary of the firm falls where the marginal cost of internal organization equals the marginal cost of market exchange. This principle applies directly to the AI value chain, which can be decomposed into a stack of physical and informational layers: photonâatomâchipâpowerâtokenâquestionâvalue.photon . (10) Each arrow represents a transformation with characteristic physics, timescales, and transaction costs. At the bottom, photons strike silicon in solar cells to generate electricity, or heat water in gas turbines to spin generators; the same element, purified to 99.9999999% (ânine ninesâ), forms the substrate of semiconductor wafers. Atoms of silicon, copper, gold, and rare earths are assembled into chips through lithographic processes operating at nanometer precisionâa single advanced fab consumes roughly 60 million liters of ultrapure water per day (TSMC, 2024). Chips consume electrical power, measured in watts, to execute the matrix multiplications that produce tokens, measured in tokens per second per watt. Tokens are assembled into coherent responses to questions, and questions generate economic value when they reduce uncertainty about decisions. The stack spans twelve orders of magnitude in timescale: mining copper takes years, fabricating a chip takes months, generating a token takes milliseconds, reading it takes a fraction of a second. Capital flows upward through the stack because each transformation compresses more physical input into less physical output: tons of ore become grams of silicon, grams of silicon become nanometers of circuit, and nanometers of circuit produce megabytes of tokens that propagate at the speed of light. The economic logic of the stack follows from a physical asymmetry that Bertrand Russell identified in compressed form (Russell, 1935): âWork is of two kinds: first, altering the position of matter at or near the earthâs surface relatively to other such matter; second, telling other people to do so. The first kind is unpleasant and ill paid; the second is pleasant and highly paid.â At the bottom of the stackâmining copper, refining silicon, building data centersâwork consists of moving atoms on Earthâs surface. At the topâgenerating tokens, answering questionsâwork consists of moving bits through fiber and silicon at the speed of light. The same compression gradient that drives capital upward has a physical basis: information propagates at c while atoms are bound by mass and distance. The cost of moving a kilogram of copper from mine to fab scales with both; the cost of moving a megabyte of tokens from data center to user scales with bandwidth alone. Robotics attempts to close this gap; that it remains vastly more expensive to manipulate atoms than bits is the physical basis of Russellâs observation. Russellâs asymmetry yields a single economic principle: value is speed. The principle originates in von Neumannâs analysis. In The Computer and the Brain, he reduced all computation to two primitive operations: a memory register must be able to âstoreâ a number and to ârepeatâ it upon âquestioningââto emit it to another organ upon request (von Neumann, 1958). Storage and retrieval: that is the entire architecture. The economic question is which operation commands a premium, and von Neumann identified the answer in his analysis of access times. He observed that it is âtechnologically difficult, orâwhich is the way in which such difficulties usually manifest themselvesâvery expensive, to provide all words with access time t,â and proposed a hierarchical memory structure in which frequently needed data is stored at fast (expensive) access times and the remainder at slow (cheap) ones. The cost of computation is, at root, the cost of retrieval speed. This is precisely the economic service that AI provides. A large language model compresses a training corpusâtrillions of tokens of human knowledgeâinto parametric storage (the model weights). Inference is retrieval: given a query, the model reconstructs a targeted answer from its compressed representation in milliseconds. The economic value lies not in the storage (the weights sit inertly on disk) but in the speed and relevance of retrievalâthe ability to answer a natural-language question from a petabyte-scale knowledge base in the time it takes a human to read the first word of the response. The previous generation of information retrievalâweb searchâorganized knowledge by keyword and hyperlink. Language models organize it by meaning, compressing text into parametric representations that support semantic retrieval at machine speed. The history of computing confirms von Neumannâs hierarchy. In the 1990s, the industry organized around storage (EMC) â networking (Cisco) â data â application. As storage commoditized, value migrated to retrieval speed. Jeff Deanâs âNumbers Everyone Should Knowââaccess latencies from L1 cache (0.5 nanoseconds) to disk seek (1.6 milliseconds), spanning six orders of magnitudeâbecame the canonical reference for system design (Dean, 2009). The firms that accelerated retrieval captured the surplus: Intelâs Math Kernel Library, NVIDIAâs CUDA, Googleâs Tensor Processing Units. At each transition, the physical layer depreciated while the speed layer above it captured the premium. The AI stack recapitulates this pattern: the value of a GPU depreciates with each generation, but the value of fast, accurate retrieval from a compressed knowledge base only increases. The competitive structure of the frontier AI industry confirms this. The race among laboratoriesâOpenAI, Anthropic, Google DeepMind, xAI, Metaâis, beneath the differences in architecture, training data, and safety methodology, a race for inference speed per unit energy. The binding competitive variable is tokens per watt per dollar. The firm that retrieves fastest at a given energy budget captures the inference market; all other advantagesâbrand, distribution, safety reputation, user baseâare subordinate to this physical constraint. Whoever generates fastest, wins. 4.1 Where Value Concentrates Coaseâs theory predicts that value concentrates where transaction frictions are highest. At the base of the stack, the chip layer, transaction costs are extreme. Semiconductor fabrication requires capital expenditures exceeding $20 billion per facility (TSMCâs Arizona complex is budgeted at over $65 billion for three fabs (TSMC, 2024)), lead times of three to five years, and a supply chain spanning dozens of countries. The result is vertical integration: a handful of firms (TSMC, Samsung, Intel) control global chip production. Moore (1965) observed that transistor density doubles approximately every two years, a trajectory that has driven six decades of hardware improvement. But Mooreâs trajectory is approaching physical limits. Feature sizes below 3 nanometers encounter quantum tunneling, lithographic precision constraints, and thermal density barriers that make further doubling progressively more expensive per transistor even when technically achievable. The implication for the value stack is direct: as the hardware improvement rate decelerates, value migrates decisively from the physical chip toward the software and algorithmic layersâCUDA, model architectures, inference optimizationâthat determine how efficiently a fixed transistor budget is utilized. When chips stop getting faster, speed becomes a software problem. Yet even a near-monopolist at this layer faces a fundamental constraint identified by Coase (1972): the durable goods monopoly problem. A monopolist selling a durable good competes with its own past and future output. Once a chip is sold, it persists in the market, reducing demand for the next generation. The mechanism is precise: a buyer who values a GPU at y knows that the monopolist will eventually lower the price to capture the next tier of buyers at valuation x<yx<y. If the buyer is patientâif the discount factor d is close to 1âwaiting costs little, and the monopolist must drop its first-period price to dâx+(1âd)âydx+(1-d)y to prevent delay. In the limit, with a continuum of buyer valuations and a discount factor approaching unity, the monopolist is driven toward competitive pricing in the first period despite controlling the entire supply (Coase, 1972). The analogy to mining is exact: every ounce already sold sits in the market competing with the next ounce. Applied to the AI chip market: the cadence is relentlessâHopper (2022), Blackwell (2024), Vera Rubin (2026), Feynman (2028)âwith each generation offering roughly 5Ă5Ă the inference throughput at a fraction of the cost per token (NVIDIA, 2024), depreciating the millions of units already deployed. Every buyer knows the next generation is eighteen months away and five times as capable; waiting is rational. A caveat: as of 2024, the dominant GPU vendor maintains gross margins near 75%, suggesting value has not yet migrated upward. In the short run, fabrication bottlenecks and extreme demand sustain pricing powerâthe discount factor is effectively low because buyers cannot afford to wait when capacity is scarce and the opportunity cost of delayed deployment is high. But Coaseâs logic is patient. As supply catches up and the secondary market for used GPUs deepens, the effective discount factor rises, margins will compress, and value will shift toward the non-durable layers: tokens and questions, which are consumed upon generation and cannot accumulate in a secondary market. The question layer and the token layer. At the top of the stack, transaction costs take a different form. Arrowâs information paradox dominates: the value of a questionâs answer cannot be assessed until the answer is obtained (Arrow, 1962). This makes efficient pricing of AI queries impossible in principle. Users cannot comparison-shop for answers they do not yet have. The middle of the stackâthe conversion of power into tokensâis where the Coasean logic is sharpest. A chip manufacturer selling GPUs to individual consumers incurs enormous transaction costs: each buyer must acquire hardware expertise, build cooling infrastructure, manage software dependencies, and bear the risk of rapid obsolescence. The Coasean conclusion is blunt: a GPU manufacturer should not be selling chips to end users. The economically efficient unit of sale is not a chip, or even a server, but a token. A cloud provider that integrates from chips to tokens eliminates the frictions and captures the surplus. The hardware layer should be invisible to the end user, just as the turbine is invisible to the electricity consumer and the fiber-optic cable is invisible to the telephone caller. The software layer: CUDA and lock-in. Within the hardware layer itself, value concentrates not in the physical chip but in the software framework that makes it programmable. CUDAâNVIDIAâs parallel computing platformâis the decisive intermediary between raw silicon and useful computation. In von Neumannâs terms, CUDA converts a matrix of transistors into a retrieval engine: it orchestrates thousands of parallel cores to execute the matrix multiplications that constitute a forward pass, turning stored model weights into a retrieved answer in milliseconds. A GPU without CUDA is inert silicon; with CUDA, it is the fastest knowledge-retrieval system ever built. The framework, not the device, is the locus of lock-in and therefore of value. The chip competes with its successors (Coaseâs durable goods problem); the software ecosystem that delivers retrieval speed persists. The depth of this lock-in is visible in pricing. Jensen Huang has argued that even if competitors offered their chips for free, the opportunity cost of using them would exceed the savings (NVIDIA, 2024). The reasoning is physical: every data center is power-limited, and in a power-constrained facility, the binding variable is tokens per watt. If NVIDIAâs co-designed stack delivers twice the tokens per watt of an alternative, a data center operator choosing the alternative forgoes half its potential revenue from each watt of capacityâa cost that dwarfs any discount on the silicon itself. The competitive moat is not the chip price but retrieval speed per unit energy, a joint product of hardware architecture and software stack. This is von Neumannâs access-time hierarchy made commercial: the firm that delivers the fastest retrieval per joule captures the surplus, regardless of what competitors charge for storage. Self-competition and cluster economics. Yet the same firm that is immune to external competition is locked in permanent competition with itself. Each new generationâBlackwell delivering up to 30Ă30Ă the inference throughput of Hopper on large mixture-of-experts models (NVIDIA, 2024)ârenders the installed base obsolete. The millions of Hopper GPUs already deployed become the competitor that no pricing strategy can neutralize: already paid for, already racked, already running. Huangâs argument about free competitor chips applies with equal force to NVIDIAâs own prior generationâbut in reverse. A customer with a functioning Hopper cluster faces high switching costs to Blackwell, and the Hopper clusterâs opportunity cost of not running is zero because it is already sunk. NVIDIA must price each generation not only against external alternatives but against its own customersâ rational reluctance to replace working capital. This is Coaseâs durable goods conjecture in its purest form: the monopolist competes with its own past output over an infinite horizon, and the faster it innovates, the faster it depreciates its own installed base. The economics of AI cluster operation illustrate this directly. A 3,000-GPU training run at commodity cloud pricing ($2â3.50/GPU-hour) costs approximately $1.6 million for a single model training job (Nebius AI, 2025). Of this, roughly 30% is consumed by overhead: hardware failures (mean time between failures âź 10 hours at this scale), checkpointing, rollback to saved state, and cluster maintenance (Sivathanu et al., 2024). The difference between a commodity GPU rental and a vertically integrated providerâwhich reduces failure recovery from hours to minutes through automated replacement, shared buffer capacity, and managed orchestrationâtranslates to savings of 20â25% of total training cost. This is the Coasean boundary of the firm made quantitative: the transaction costs of operating raw hardware are high enough that vertical integration from chip to token is economically efficient, and firms that sell tokens rather than GPU-hours capture the resulting surplus. The Coasean logic extends beyond the AI value chain to the structure of the firms that use it. Coaseâs original questionâwhy do firms exist?âreceived an answer calibrated to the transaction costs of the mid-twentieth century: the costs of search, negotiation, contracting, and monitoring in the open market. AI compresses these costs at every margin. A language model that can draft contracts, analyze markets, summarize regulations, and coordinate suppliers reduces the marginal cost of market exchange relative to internal organization. The Coasean implication is that the optimal firm shrinks. In the limiting case, a single individual with sufficient AI access can perform the coordination that previously required a departmentâthe firm of one. This restructuring is asymmetric across the stack. At the token and question layers, where work consists of processing information, the minimum viable firm approaches a single principal. At the atom layerâmining copper, fabricating chips, pouring concrete for data centersâwork remains bound by physical coordination, safety regulation, and the irreducible logistics of moving matter. The optimal firm size at the bottom of the stack remains large. AI compresses the informational component of transaction costs; it does not compress the physical component. The result is a bifurcation: an economy of very small firms at the top of the stack, sustained by very large firms at the bottom. 4.2 The Compression Principle Shannon (1948) proved that any data source with entropy H can be compressed to a rate approaching H bits per symbol but no further. This source coding theorem governs the stack: it sets the minimum data movement for inference, the information density of tokens, and the minimum tokens needed to answer a question. The compression principle interacts with Mooreâs law to reinforce this pattern. As hardware improves, the cost of raw computation falls exponentially. But the information content of a questionâits Shannon entropyâdoes not shrink with better hardware. A medical diagnosis, a legal analysis, a scientific hypothesis: these have irreducible informational complexity that no amount of hardware improvement can compress away. A second, independent floor comes from computational complexity. Cook (1971) and Karp (1972) established that a large class of problemsâNP-hard problems including optimal scheduling, protein folding, network design, and many combinatorial optimization tasksârequire computation that grows exponentially with problem size under the widely believed conjecture that Pâ NPP . No hardware acceleration changes this: a problem that requires 2n2^n operations on a slow machine still requires 2n2^n operations on a fast one. Mooreâs law and GPU scaling reduce the constant factor; they do not reduce the exponent. The token economy can generate answers faster, but for the hardest questions, the computational cost of the correct answer exceeds the energy budget of any physically realizable machine. Shannon sets the information-theoretic floor; Cook and Karp set the algorithmic floor. Both are absolute. The compression principle extends to computation itself. Turing (1936) showed that any computable function can be decomposed into a sequence of elementary operations on a universal machine. Programming languages formalize this at different levels of the abstractionâperformance frontier: Haskell and the lambda calculus tradition maximize expressivenessâcomplex programs are built by composing simple functions, fâgf g, each of which compresses a subproblem into its outputâat the cost of execution speed. C++ compresses closer to the machine; CUDA compresses further still, mapping computation directly onto GPU hardware. An LLM generating code traverses this hierarchy, assembling token sequences that represent nested function applications. The compression is hierarchical at every level: tokens compress characters, functions compress tokens, programs compress functions, and the deep learning frameworks (PyTorch, TensorFlow) compress neural network specifications into sequences of CUDA kernel calls. The hierarchy extends above the programmerâs level: a spreadsheet is a constrained programming languageâcells as variables, formulas as functions, recalculation as executionâwith severe limits on recursion, type systems, and algorithmic complexity. Excel is, in essence, a restricted and slower C++, which is itself a restricted and slower version of direct machine code. Each layer of abstraction compresses the userâs cognitive burden at the cost of computational expressiveness. The LLM sits at the top of this hierarchy: it accepts natural language as input and compresses the entire abstraction stack into a single interface. This yields a corollary that inverts a common assumption: in the mature AI economy, the act of writing codeâtranslating human intent into machine instructionsâis a low-value activity. Code is a compression of intent into tokens, and as models improve, this compression becomes automated. Acemoglu and Restrepo (2019) show that automation displaces labor in existing tasks but simultaneously creates new tasks in which humans have comparative advantage. Applied to the AI stack: as token generation is automated, value migrates to the tasks that tokens cannot automateâformulating questions, exercising judgment, generating novel data from embodied experience. The scarce resource is not the ability to instruct a machine but the ability to formulate a question worth asking. Keynes (1930) anticipated this conclusion nearly a century ago. Writing during the Great Depression, he predicted that within a hundred yearsâby approximately 2030âcompound interest and technological progress would solve what he called âthe economic problemâ: the struggle for subsistence that had occupied humanity since its origins. He coined the term technological unemploymentââunemployment due to our discovery of means of economising the use of labour outrunning the pace at which we can find new uses for labourââbut dismissed it as a temporary phase of maladjustment. The deeper challenge, Keynes argued, would be leisure: âfor the first time since his creation man will be faced with his real, his permanent problemâhow to use his freedom from pressing economic cares, how to occupy the leisure, which science and compound interest will have won for him, to live wisely and agreeably and well.â He envisioned a fifteen-hour work week within three generations. The token economy is Keynesâs prediction arriving, roughly on schedule, in a form he could not have anticipated. AI does not merely automate physical labor; it automates the informational component of nearly all labor, creating not physical leisure but cognitive leisure. The question budget (Section 5) is the computational form of Keynesâs leisure problem: with 2,200 questions per person per day at the 2028 upper bound, the constraint is not the capacity to obtain answers but the wisdom to ask questions worth answering. Keynes worried that the wealthy classes, freed from economic necessity, had âfailed disastrouslyâ to find purpose. The same risk applies to a civilization with abundant tokens and no framework for directing them. Formalizing this budgetâand determining how many questions the infrastructure can physically supportârequires measuring the information content of inquiry itself. 5 The Question Budget Tokens are the medium; questions are the message. Cox (1946) proved that any system for reasoning under uncertainty that satisfies basic consistency requirementsâdivisibility, comparability, and agreement with common senseâmust be isomorphic to probability theory. Cox (1961) developed this into a full algebraic framework: probability is not merely a statistical tool but the unique consistent extension of Boolean logic to propositions whose truth value is uncertain. Jaynes (2003) showed that Coxâs axioms, combined with the principle of maximum entropy, yield a complete theory of inference from incomplete information. Cox (1979) extended this framework to the logic of inquiry. A question can be formally defined as the set of propositions that would answer it. The âbearingâ of a question on an outstanding issue is a measure analogous to probability: it quantifies the relevance of an inquiry to the reduction of uncertainty. Just as Coxâs axioms uniquely determine probability as the calculus of plausible reasoning, the logic of inquiry uniquely determines how questions should be ordered by their expected information gainâa result that connects directly to Shannonâs mutual information (Equation 12). 5.1 Questions as Entropy Reduction A question specifies what information is sought; receiving the answer reduces entropy (Shannon, 1948). Consider a random variable X with Shannon entropy Hâ(X)=ââipâ(xi)âlog2âĄpâ(xi).H(X)=- _ip(x_i) _2p(x_i). (11) Asking a question Q and receiving an answer provides mutual information Iâ(Q;X)=Hâ(X)âHâ(XâŁQ),I(Q;X)=H(X)-H(X Q), (12) reducing the remaining uncertainty from Hâ(X)H(X) to Hâ(XâŁQ)H(X Q). An ideal yes/no question that bisects the probability mass provides exactly 1 bit of information. A question with n equiprobable answers provides at most log2âĄn _2n bits. 5.2 The Thermodynamic Cost of Questions By Landauerâs principle (Equation 1), each bit of information processed has a minimum thermodynamic cost of kBâTâlnâĄ2k_BT 2. Asking a question and receiving an answer that provides I bits of mutual information therefore costs at least EquestionâĽIâ kBâTâlnâĄ2.E_question⼠I¡ k_BT 2. (13) This is a floor, not a ceiling. In practice, the cost is dominated by the token processing required to formulate the question and generate the answerâtypically 100â1,000 tokens for a substantive query-response pair. 5.3 The Finite Question Budget Given the projected 2028 AI energy budget of 326 TWh and an illustrative conversation length of 100 tokens (prompt plus responseâa short factual query; substantive interactions typically consume 500â2,000 tokens, which would reduce the question count by 55â20Ă20Ă), the implied annual question budget under this allocation is 6.5Ă1017â tokens/year100â tokens/question=6.5Ă1015â questions/year, 6.5Ă 10^17 tokens/year100 tokens/question=6.5Ă 10^15 questions/year, (14) or roughly 6.5 quadrillion questions. Per person per day: 6.5Ă10158Ă109Ă365â2,200â questions/person/day. 6.5Ă 10^158Ă 10^9Ă 365â 2,200 questions/person/day. (15) This estimate is sensitive to two parameters: the energy budget and the tokens per question. Table 2 shows the question budget under varying assumptions. Table 2: Sensitivity of the daily question budget per person to energy budget and conversation length. Assumes 5Ă10â45Ă 10^-4 Wh/token and 8 billion people. The 326 TWh column uses the 2028 US upper-bound projection; the 650 TWh column approximates global AI energy if non-US capacity is included. Questions/person/day Tokens/question 65 TWh 165 TWh 326 TWh 650 TWh 100 450 1,130 2,200 4,500 250 180 450 900 1,800 500 90 225 450 900 1,000 45 113 225 450 At 100 tokens per question (a short factual query), the 2028 US budget yields 2,200 questions per person per day. At 1,000 tokens per question (a substantive multi-turn exchange with a long response), the budget drops to 225âstill substantial, but an order of magnitude smaller. The sensitivity to conversation length matters because frontier models increasingly generate long, detailed responses: a medical differential diagnosis or a legal memorandum may consume 2,000â5,000 tokens, reducing the effective question budget by a factor of 20â50 relative to the short-query estimate. The allocation of questions is an economic problem: which questions are worth asking, who gets to ask them, and at what cost. The bottleneck is not the âfuelâ for good questionsâhuman curiosity and domain knowledgeâbut the computational infrastructure to answer them. 6 Knowledge and Agency The previous sections established what the token economy can produce: a finite budget of tokens and questions, distributed across a value stack governed by transaction costs and physical constraints. This section asks what the token economy cannot resolve. The expansion of computational capacity increases the rate at which uncertainty can be processed; it does not determine which uncertainties matter, how beliefs should be revised, or when to act on incomplete models. These are problems of agencyâand they do not yield to scale. 6.1 From Entropy to Direction Sections 2â5 established that tokens and questions are physically bounded. Section 7 will show that optimization against imperfect proxies introduces structural distortion. Between physical constraint and measurement limits lies a third variable: how finite questions are deployed under uncertainty. Shannon defined information as entropy reduction (Shannon, 1948). A token reduces uncertainty about a distribution of possible symbols. A question reduces uncertainty about a random variable via mutual information. These measures discipline the arithmetic of the token economy, but entropy reduction is not the same as directional knowledge. A message may resolve ten bits of uncertainty about tomorrowâs weather or ten bits about which treatment will extend a patientâs life; the Shannon entropy is identical, but the bearing on actionâin Coxâs senseâis not. Knowledge relevant to agency is not merely lower entropy but reduced uncertainty about the consequences of interventions: not what is likely to be observed, but what would change if one acted differently. The token economy increases the rate at which uncertainty can be reduced, but rate is not direction. A system producing a million tokens per second can answer many questions; which questions to ask requires a model connecting information to consequencesâa causal structure that the token-generating process does not itself supply. 6.2 Agency Under Structural Uncertainty Agency in the token economy can be defined as the capacity to: ⢠Formulate high-bearing questions, ⢠Update beliefs coherently, ⢠Act when probability models are incomplete, ⢠Preserve adaptability across regime change. In environments of measurable risk, optimal action follows from expected value maximization conditional on known distributions. In environments of structural uncertaintyâunknown models, shifting states, nonstationary systemsâoptimization within a fixed distribution can be precisely wrong. Much of technological, economic, and institutional evolution occurs in the latter regime: the distributions governing semiconductor demand in 2030, the political economy of AI regulation, or the long-run effects of automating legal reasoning are not drawn from known urns. The expansion of the token budget does not eliminate structural uncertainty; it increases the speed at which agents encounter it. A system that generates answers faster arrives sooner at the boundary where its training distribution ceases to apply. Abundance of answers is not abundance of certainty. 6.3 Iterative Inquiry, Path Dependence, and Optionality Under structural uncertainty, effective use of a finite question budget resembles disciplined experimentation rather than static optimization. Let StS_t denote the feasible state space at time t . An action a induces a (stochastic) transition: St+1=Tâ(St,a,Îľt),S_t+1=T(S_t,a, _t), where Îľt _t captures unmodeled shocks and Tâ(â )T(¡) is generally unknown and may itself evolve with the state. Actions differ not only in expected payoff but in how they reshape possible futures. Irreversible commitments collapse branches of the decision tree. This introduces path dependence: the set of reachable states at a future horizon depends on the sequence of prior actions, not merely on the current state. Consider a simple contrast. An exploitative action may offer high immediate expected payoff but collapse the option set to a narrow future trajectory. An exploratory action may yield lower short-run payoff while preserving a broader set of reachable states. Exploration carries positive option value whenever the expected continuation value under uncertainty exceeds the value of prematurely collapsing the state spaceâeven if exploitation appears superior in the short run. Preserving the dimensionality of StS_tâkeeping more futures reachableâcan dominate short-run maximization when models are misspecified, because the value of flexibility rises with the probability that the current model is wrong. The token economy accelerates inference, but it does not repeal path dependence. A firm that uses AI to optimize aggressively along one strategic path forecloses alternatives that no subsequent computation can reopen. More tokens cannot restore branches already collapsed. 6.4 Measurement Limits and the Residual Role of Judgment Section 7 shows that optimization against imperfect proxies generates distortion proportional to proxy quality (Equation 18). Increasing optimization pressure amplifies both genuine improvement and gaming in fixed proportion; the gaming fraction (1âĎ2)(1-Ď^2) is set by the proxy, not by the optimizer. No increase in computational scale eliminates this structure, because proxy metrics approximate objectivesâthey do not embody them. This leaves a residual role for judgment external to automated feedback loops. The limitation is not computational but structural: objectives cannot be fully captured by any finite metric. As token throughput increases, so does the scale at which proxy distortion operates. An AI system optimizing medical diagnoses against patient satisfaction scores, or scientific output against citation counts, amplifies the gap between proxy and objective in proportion to its throughput. Optimization can be industrialized; judgmentâthe recognition that the proxy has diverged from the objectiveâcannot. 6.5 Humans as the Open System Shumailov et al. (2024) demonstrate that models trained recursively on synthetic outputs exhibit distributional collapse: the tails of the distribution are progressively lost, variance shrinks, and error compounds with each generation. Without exogenous data, diversity contracts monotonically. Large language models compress the statistical regularities of their training corpus into parametric representations. They do not originate new distributionsâthey recombine and interpolate within the support of what they have observed. The underlying data-generating process remains human activity embedded in physical and social environments: experiments conducted, observations recorded, judgments rendered, errors made and corrected. In thermodynamic terms, the model is a bounded computational system; sustained performance requires continual input from a process external to the inference loop. A model trained on the web as of 2024 has no access to the results of experiments conducted in 2025, the consequences of policies enacted after its training cutoff, or the embodied experience of navigating a world it has never inhabited. Inference is downstream of lived experienceâthis is not a normative claim but a structural one: a closed informational system converges toward homogeneity. 6.6 Collective Updating and Informational Fragmentation Individual belief revision may be represented as Bayesian updating: pâ(θâŁD)âpâ(DâŁÎ¸)âpâ(θ).p(θ D) p(D θ)\,p(θ). Collective action requires some shared protocol for updating beliefs. If agents interpret evidence under incompatible revision rulesâor begin from sufficiently divergent priorsâposterior distributions can diverge even as signal volume increases. This is not a pathology but a theorem: Bayesian agents with different priors and different likelihood models can process identical data and arrive at opposing conclusions. In a low-token environment, fragmentation is constrained by scarcity of evidence: there are not enough signals to sustain many incompatible worldviews simultaneously. In a high-token environment, fragmentation can accelerate, because agents can selectively query AI systems to reinforce incompatible models of the worldâeach receiving internally consistent, well-argued responses that deepen rather than resolve disagreement. More information does not guarantee convergence; it guarantees only that each agentâs model becomes more elaborate. Coordination therefore depends not only on compute access but on institutional norms governing inference and revision. 6.7 Abundance Without Orientation Projected infrastructure permits hundreds to thousands of daily queries per person (Table 2). This represents a qualitative shift from informational scarcity to informational abundance. Yet rapid entropy reduction is not self-directing. In the presence of path dependence and proxy distortion, high-throughput inference can amplify instability as easily as it can amplify stability. The marginal token is inexpensiveâfractions of a cent at current pricingâbut the marginal well-posed question is not. Formulating a question with high bearing on an open problem requires domain knowledge, causal reasoning, and awareness of what is not yet known, none of which scale with token throughput. The limiting resource in the mature token economy is therefore not computation but directional coherence: alignment between inquiry, belief revision, and long-horizon objectives under structural uncertainty. 6.8 The Boundary of Machine Inference Large language models approximate conditional distributions pâ(yâŁx)p(y x): given this context, what token is likely next? Causal reasoning requires a different operationâintervention, denoted dâoâ(X)do(X)âwhich asks what would happen if a variable were set by external action rather than passively observed. Models may simulate causal reasoning through patterns absorbed from training data, but their grounding remains derivative of observed conditional distributions, not of the interventionist structure that generated those observations. Human agency operates at the interventionist level: a physician prescribes a treatment, an engineer changes a design, a policymaker enacts a regulation. Each commits to an action that reshapes the distribution of future outcomes. Token generation simulates patterns within existing distributions; agency alters the distributions themselves. This distinction is not a limitation to be overcome by scaleâit is a structural boundary between statistical inference and causal intervention. 6.9 Implication The token economy converts energy into symbol sequences at unprecedented scale. Physical limits constrain efficiency (Section 2); measurement limits constrain optimization (Section 7); structural uncertainty constrains model validity. Within these boundaries, the decisive variable is how finite questions are selected, sequenced, and acted upon. These constraintsâphysical, informational, and structuralâdefine the boundary conditions within which the token economy operates. The next section asks what happens when optimization proceeds without acknowledging them. 7 The Limits of Measurement and Optimization Section 6 argued that agency in the token economy requires disciplined inquiry under structural uncertainty: the capacity to select, sequence, and act upon finite questions while preserving optionality across evolving states. We now introduce a further constraint. Even when objectives are specified and inquiry is directed, optimization against measurable proxies introduces irreducible distortion. The limitation is not computational scale but structural: metrics approximate objectives; they do not embody them. The token economy is physically bounded. It is also subject to a subtler constraint: the limits on measuring and optimizing it. Goodhart (1984) observed, in the context of British monetary policy, that âany observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.â In its popular formulation: when a measure becomes a target, it ceases to be a good measure. Language models trained to maximize a reward signal routinely exploit the proxy rather than achieving the intended objective (Amodei et al., 2016). Clark and Amodei (2016) documented a vivid example: a reinforcement learning agent trained to score points in the boat racing game CoastRunners learned to circle repeatedly through regenerating targets rather than completing the courseâcatching fire, colliding with other boats, and going the wrong direction, yet achieving a higher score than any course-completing strategy. Benchmark gamingâoptimizing model performance on standardized tests rather than on real-world capabilityâis the same phenomenon at the industry level. 7.1 Heisenbergâs Uncertainty Heisenberg (1927) established that the position x and momentum p of a particle cannot be simultaneously determined with arbitrary precision: Îâxâ ÎâpâĽâ2. x¡ p⼠2. (16) The uncertainty principle follows from the non-commutativity of position and momentum operatorsâa property of quantum states, not of the measurement apparatus. But the intuition is often expressed through Heisenbergâs microscope: a photon used to determine a particleâs position imparts momentum to it, disturbing the quantity one seeks to measure. It is this measurement-disturbance picture, rather than the deeper algebraic structure, that parallels Goodhartâs law most closely. 7.2 A Structural Parallel These two principlesâone from economics, one from quantum mechanicsâare instances of a common structure: observation coupled to control distorts the observed quantity. In Heisenbergâs case, the photon used to measure position kicks the particle, altering its momentum. In Goodhartâs case, the optimization pressure applied to a metric distorts its relationship to the underlying objective. A toy model makes the structure precise. Let θ be the true objective (e.g., genuine model capability) and m=θ+Îľm=θ+ a measurable proxy (e.g., benchmark score), where Îľ is noise independent of θ with Varâ(Îľ)=ĎÎľ2Var( )= _ ^2. The proxy-objective correlation is Ď=ĎθĎθ2+ĎÎľ2.Ď= _θ _θ^2+ _ ^2. (17) An optimizer that selects the candidate with the highest proxy score m from N alternatives achieves an expected proxy gain Gâ2âlnâĄNG 2 N (for large N, by extreme value theory). The expected genuine improvement is F=Ď2â G.F=Ď^2¡ G. (18) The remainder, (1âĎ2)âG(1-Ď^2)G, is pure gamingâimprovement in the noise component Îľ that contributes nothing to the true objective θ. As optimization pressure increases (larger N), G grows, but the fraction of genuine improvement F/G=Ď2F/G=Ď^2 remains fixed. To increase G without losing fidelity, one must reduce ĎÎľ2 _ ^2âthat is, improve the proxy itself, which requires new measurement effort. The analogy to Heisenberg is now visible. In quantum mechanics, Îâxâ ÎâpâĽâ/2 x¡ p⼠/2: increasing precision in position (Îâxâ0 xâ 0) forces loss of precision in momentum (Îâpââ pââ). In the Goodhart model, increasing optimization gain (GââGââ) forces the gaming waste (1âĎ2)âG(1-Ď^2)G to grow without bound, while the genuine improvement F=Ď2âGF=Ď^2G grows at a slower rate set by the fixed proxy quality Ď2Ď^2. The formal analogy is not identity. In Heisenberg, there is a strict tradeoff: reducing Îâx x forces Îâp p to increase. In Goodhart, there is no strict tradeoff: increasing G increases both the genuine gain Ď2âGĎ^2G and the gaming waste (1âĎ2)âG(1-Ď^2)G in fixed proportion. Heisenberg constrains a product; Goodhart constrains a ratio. What the two structures share is that extraction of information from a coupled system is accompanied by an irreducible distortion set by a property of the measurement apparatusââ in one case, ĎÎľ2 _ ^2 in the other. Financial markets provide an empirical laboratory for this structure. Bruce Kovner, one of the most successful macro traders of the twentieth century, stated the principle explicitly: âThe more a price pattern is observed by speculators the more prone you have false signals; the more the market is a product of non-speculative activity, the greater the significance of technical breakoutâ (Schwager, 1989). Soros (1987) formalized the same insight as reflexivity: market participantsâ beliefs affect fundamentals, which affect beliefs, creating a feedback loop in which observation and control are inseparable. Goodhartâs original observation was itself drawn from monetary policyâa market context. For AI alignment, the implication is direct. Any proxy metric for âbeneficial AI behaviorââhowever carefully constructedâwill be distorted by the optimization process that targets it. Stronger optimization does not help; only better proxies (lower ĎÎľ2 _ ^2) do. This is not a solvable engineering problem at fixed measurement quality but a structural feature of optimization against imperfect metrics. 7.3 The Information Market Arrow (1962) identified a fundamental paradox in the economics of information: the value of information cannot be assessed until it is acquired, but once acquired, the buyer has no need to pay for it. Information is non-rivalrous, non-excludable, and of uncertain value ex anteâproperties that make efficient pricing impossible. Multiple large language models now compete as information providers, a game in the sense of von Neumann and Morgenstern (1944). But Arrowâs paradox applies: users cannot know which model will best answer a question until they have the answer. And Goodhartâs law applies: models optimized for benchmarks diverge from models optimized for genuine utility. The result is an information market that is useful but fundamentally incapable of efficient allocation. This triple constraintâGoodhartâs measurement distortion, Heisenbergâs observer back-reaction, and Arrowâs pricing impossibilityâdefines the boundary conditions of the token economy. The token budget can be computed; it cannot be perfectly optimized. 8 Conclusions If the question budget is finite, the question of who gets access is inescapable. The same physical stackâphoton to atom to chip to power to token to questionâterminates in radically different uses. The identical joule of electricity, converted through the identical GPU into the identical token, can answer a physicianâs diagnostic query, generate a legal brief, assist a climate scientist, or help a teenager produce a video she will watch once and forget. The stack does not discriminate. The allocation mechanism does. This is the central regulatory question: who decides which questions get asked? Three candidate allocators present themselvesâmarkets, governments, and platformsâand none is without defect. Coase (1960) showed that in the absence of transaction costs, the initial allocation of property rights does not affect efficiency: parties will bargain to the optimal outcome regardless. But transaction costs are never zero, and in the AI stack they are substantial. The implication is that the initial allocation of access rights to AI compute does matter for efficiency, not just equity. The default allocator is the market: tokens go to those willing to pay. Market allocation has the standard virtuesâit aggregates dispersed information about willingness to pay, it provides incentives for efficiency, and it does not require a central planner to know the value of every possible question. But it also has the standard defects. Markets systematically underinvest in public goods (Arrow, 1962). A question about the mechanism of antibiotic resistance has enormous social value but generates no private revenue for the questioner. A question about optimizing advertising click-through rates generates immediate private revenue but negligible social value. Left to the market, the question budget will be allocated disproportionately toward commercially valuable queries and away from basic science, public health, education, and democratic governanceâprecisely the domains where the social return to information exceeds the private return. Arrowâs paradox (Section 4) sharpens the problem: market allocation is structurally biased toward uses where value is known in advance (routine commercial applications) and away from uses where value is uncertain but potentially transformative (research, exploration, novel inquiry). In practice, neither markets nor governments currently allocate the question budget. Platforms do. A handful of firmsâOpenAI, Anthropic, Google, xAI, Metaâthat control the model layer and the inference infrastructure determine access through pricing tiers, rate limits, acceptable use policies, and content moderation rules. This is de facto allocation by platform fiat, constrained only by competition among providers (limited, given the capital requirements) and by regulatory oversight (nascent, given the speed of deployment). Platform allocation has the virtue of speed: firms can deploy and iterate access policies faster than legislatures can draft statutes. But it concentrates the allocation decision in entities whose objective function is shareholder value, not social welfare. The question of which questions humanity gets to ask is decided, in effect, by the terms of service of three or four companies. The third model is direct public investment. Just as governments fund research through NSF and NIH because markets underinvest in basic knowledge (Arrow, 1962), governments could fund public AI infrastructure for domains where social value exceeds private value. The precedent is the electrical grid and the telephone network: both were initially controlled by private monopolies, both were subjected to common-carrier regulation (universal service obligations, rate regulation, non-discriminatory access), and both generated enormous positive externalities once access was democratized. The question is whether AI computeâspecifically, access to tokensâshould follow the same regulatory trajectory. The choice among these modelsâor their combinationâdepends on empirical facts about transaction costs that vary across the stack. At the chip layer, where natural monopoly conditions prevail, utility-style regulation is most appropriate. At the token layer, where competition among model providers is feasible, market regulation suffices with corrections for the failures identified by Arrow and Goodhart: antitrust enforcement to prevent concentration, transparency requirements for benchmarks (to mitigate Goodhart distortion), and information disclosure rules. At the question layer, where Arrowâs paradox is most acute and where the gap between social and private value is widest, some form of public allocation is likely necessary. The Coasean insight is that no single regulatory model fits the entire stack. The optimal intervention depends on where you stand in the value chain. What is clear from the arithmetic is that the allocation question is not abstract: with 2,200 questions per person per day at the 2028 upper bound, the question budget is large enough to matter but finite enough to require choice. The choice between a medical diagnosis and a disposable video is not a technical question. It is a political oneâand it should be made with the numbers in hand. The allocation problem has a temporal dimension that current debate neglects. The question budget is distributed not only across uses at a point in time but across generations over time. The standard economic tool for intertemporal allocation is discounting: a token consumed today is valued more highly than a token consumed tomorrow, at a rate r reflecting time preference and the opportunity cost of capital. But as Price (1993) argues, discounting embeds an ethical judgment that is rarely examined: it systematically devalues the welfare of future persons, not because their needs are less real, but because they are further away in time. At a discount rate of 5%, the welfare of a generation 50 years hence is weighted at (1.05)â50â0.087(1.05)^-50â 0.087âless than a tenth of the present generationâs weight. A question that would cure a disease in 2075 is valued at a tenth of the same question asked for commercial convenience today. The difficulty is that the inputs to the token economy are partly exhaustible. The energy and materials committed to AI infrastructure todayâcopper mined, silicon refined, power plants builtâare irreversibly allocated. A megawatt-hour consumed generating ephemeral tokens in 2025 cannot be consumed answering a scientific question in 2050. The efficiency gap (Section 2) offers some relief: if energy per token falls by orders of magnitude, future generations inherit a more capable infrastructure per unit of energy. But the physical materialsâcopper, rare earths, water for coolingâdo not benefit from algorithmic improvement. Every ton of copper drawn into a data center today is a ton unavailable for future infrastructure, whether AI or otherwise. If the question budget is a finite physical resource, its intertemporal allocation raises the same questions that arise for any exhaustible resource: at what rate should the present generation consume computational capacity, and what obligation, if any, does it bear to preserve capacity for successors whose questions it cannot anticipate? âLive as though youâl die tomorrow, but farm as though youâl live foreverâ (Marsden, 1998): the token economy is governed by the first imperative, but sustainability demands the second. The token economy is a physical economy. Tokens are units of information, and information processing is constrained by thermodynamics, channel capacity, and finite energy supply. The efficiency gap between current practice and the Landauer floor leaves room for improvement, but the floor is absolute. The balance sheet shows the AI token economy expanding from roughly 125 tokens per person per day in mid-2024 to a projected 225,000 under aggressive infrastructure assumptions. The question budget reframes AI policy as resource allocation: finite computational capacity must be distributed across science, commerce, governance, and private consumption. But physical expansion does not dissolve epistemic constraint. A thousand-fold increase in token throughput multiplies the rate at which answers can be generated; it does not improve the quality of the questions that elicit them, reduce the structural uncertainty in which decisions must be made, or close the gap between proxy metrics and the objectives they approximate. Entropy reduction is not direction (Section 6): a system may process vast quantities of information without clarifying which action improves outcomes. Optimization against imperfect proxies introduces structural distortion that grows with optimization pressure, not despite it (Section 7). No allocation mechanismâmarket, platform, or stateâcan fully substitute for disciplined judgment in the selection and sequencing of inquiry. The token economy can amplify inquiry. It cannot determine which inquiries matter. That choice remains human. References Acemoglu and Restrepo (2019) Daron Acemoglu and Pascual Restrepo. Automation and new tasks: How technology displaces and reinstates labor. Journal of Economic Perspectives, 33(2):3â30, 2019. Alphabet Inc. (2025) Alphabet Inc. Alphabet second quarter 2025 results. https://blog.google/inside-google/, 2025. Reported 980 trillion tokens processed monthly by Google AI models. Altman (2024) Sam Altman. Statement on OpenAI daily usage. Post on X, February 2024, 2024. Reported approximately 100 billion words generated per day. Amodei et al. (2016) Daron Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan ManĂŠ. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565, 2016. Appenzeller (2024) Guido Appenzeller. LLMflation: LLM inference cost. Andreessen Horowitz, November 2024. URL https://a16z.com/llmflation-llm-inference-cost/. Arrow (1962) Kenneth J. Arrow. Economic welfare and the allocation of resources for invention. In The Rate and Direction of Inventive Activity: Economic and Social Factors, pages 609â626. Princeton University Press, 1962. Bekenstein (1981) Jacob D. Bekenstein. Universal upper bound on the entropy-to-energy ratio for bounded systems. Physical Review D, 23(2):287â298, 1981. Bennett (2003) Charles H. Bennett. Notes on Landauerâs principle, reversible computation, and Maxwellâs demon. Studies in History and Philosophy of Science Part B: Studies in History and Philosophy of Modern Physics, 34(3):501â510, 2003. Clark and Amodei (2016) Jack Clark and Dario Amodei. Faulty reward functions in the wild. https://openai.com/index/faulty-reward-functions/, 2016. Coase (1937) Ronald H. Coase. The nature of the firm. Economica, 4(16):386â405, 1937. Coase (1960) Ronald H. Coase. The problem of social cost. Journal of Law and Economics, 3:1â44, 1960. Coase (1972) Ronald H. Coase. Durability and monopoly. Journal of Law and Economics, 15(1):143â149, 1972. Cook (1971) Stephen A. Cook. The complexity of theorem-proving procedures. In Proceedings of the Third Annual ACM Symposium on Theory of Computing, pages 151â158, 1971. Cox (1946) Richard T. Cox. Probability, frequency, and reasonable expectation. American Journal of Physics, 14(1):1â10, 1946. Cox (1961) Richard T. Cox. The Algebra of Probable Inference. Johns Hopkins Press, Baltimore, 1961. Cox (1979) Richard T. Cox. Of inference and inquiry: An essay in inductive logic. In Raphael D. Levine and Myron Tribus, editors, The Maximum Entropy Formalism, pages 119â167. MIT Press, 1979. Dean (2009) Jeff Dean. Designs, lessons and advice from building large distributed systems. Keynote, LADIS workshop, 2009. Including âNumbers Everyone Should Knowâ. DeepSeek-AI (2024) DeepSeek-AI. DeepSeek-V3 technical report. arXiv:2412.19437, 2024. Epoch AI (2025) Epoch AI. Key trends and figures in machine learning. https://epoch.ai/trends, 2025. Giustra (2025) Frank Giustra. Is copper at the start of the next supercycle? https://frankgiustra.com/posts/is-copper-at-the-start-of-the-next-supercycle/, 2025. Goldman Sachs Research (2024) Goldman Sachs Research. Generational growth: AI, data centers and the coming US power demand surge. https://w.goldmansachs.com/insights/articles/AI-poised-to-drive-160-increase-in-power-demand, 2024. Goodhart (1984) Charles A. E. Goodhart. Monetary Theory and Practice: The UK Experience. Macmillan, London, 1984. Grantham (2011) Jeremy Grantham. Time to wake up: Days of abundant resources and falling prices are over forever. GMO Quarterly Letter, April 2011, 2011. Heisenberg (1927) Werner Heisenberg. Ăber den anschaulichen Inhalt der quantentheoretischen Kinematik und Mechanik. Zeitschrift fĂźr Physik, 43(3â4):172â198, 1927. International Energy Agency (2024) International Energy Agency. Electricity 2024: Analysis and forecast to 2026. https://w.iea.org/reports/electricity-2024, 2024. International Energy Agency (2025) International Energy Agency. Energy and AI: Data centres, the new frontier of energy demand. https://w.iea.org/reports/energy-and-ai, 2025. Jaynes (2003) Edwin T. Jaynes. Probability Theory: The Logic of Science. Cambridge University Press, 2003. Jevons (1865) William Stanley Jevons. The Coal Question: An Inquiry Concerning the Progress of the Nation, and the Probable Exhaustion of Our Coal-Mines. Macmillan, London, 1865. Karp (1972) Richard M. Karp. Reducibility among combinatorial problems. In Raymond E. Miller and James W. Thatcher, editors, Complexity of Computer Computations, pages 85â103. Plenum Press, New York, 1972. Keynes (1930) John Maynard Keynes. Economic possibilities for our grandchildren. In Essays in Persuasion, pages 358â373. W. W. Norton, New York, 1930. Landauer (1961) Rolf Landauer. Irreversibility and heat generation in the computing process. IBM Journal of Research and Development, 5(3):183â191, 1961. Levitt and Dubner (2009) Steven D. Levitt and Stephen J. Dubner. Superfreakonomics: Global Cooling, Patriotic Prostitutes, and Why Suicide Bombers Should Buy Life Insurance. William Morrow, New York, 2009. Lloyd (2000) Seth Lloyd. Ultimate physical limits to computation. Nature, 406(6799):1047â1054, 2000. Luccioni et al. (2024) Alexandra Sasha Luccioni, Yacine Jernite, and Emma Strubell. Power hungry processing: Watts driving the cost of AI deployment? In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 85â99, 2024. MacKay (2003) David J. C. MacKay. Information Theory, Inference, and Learning Algorithms. Cambridge University Press, 2003. MacKay (2009) David J. C. MacKay. Sustainable EnergyâWithout the Hot Air. UIT Cambridge, 2009. Marsden (1998) John Marsden. Remarks on sustainable land management, 1998. Widely quoted in Australian agricultural literature. Meadows et al. (1972) Donella H. Meadows, Dennis L. Meadows, Jørgen Randers, and William W. Behrens I. The Limits to Growth. Universe Books, New York, 1972. Mehl et al. (2007) Matthias R. Mehl, Simine Vazire, NairĂĄn RamĂrez-Esparza, Richard B. Slatcher, and James W. Pennebaker. Are women really more talkative than men? Science, 317(5834):82, 2007. Moore (1965) Gordon E. Moore. Cramming more components onto integrated circuits. Electronics, 38(8):114â117, 1965. Nebius AI (2025) Nebius AI. The economics of AI clusters. Nebius whitepaper, 2025. NVIDIA (2024) NVIDIA. NVIDIA Blackwell platform arrives to power a new era of computing. https://nvidianews.nvidia.com/news/nvidia-blackwell-platform-arrives-to-power-a-new-era-of-computing, 2024. OpenRouter and Andreessen Horowitz (2025) OpenRouter and Andreessen Horowitz. State of AI 2025: An empirical 100 trillion token study. https://a16z.com/state-of-ai/, 2025. Patterson et al. (2021) David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021. Pierrehumbert (2009) Raymond T. Pierrehumbert. An open letter to Steve Levitt. RealClimate, October 2009. URL https://w.realclimate.org/index.php/archives/2009/10/an-open-letter-to-steve-levitt/. Price (1993) Colin Price. Time, Discounting and Value. Blackwell, Oxford, 1993. Prigogine (1978) Ilya Prigogine. Time, structure, and fluctuations. Science, 201(4358):777â785, 1978. Russell (1935) Bertrand Russell. In Praise of Idleness and Other Essays. George Allen & Unwin, London, 1935. Schwager (1989) Jack D. Schwager. Market Wizards: Interviews with Top Traders. New York Institute of Finance, New York, 1989. Shannon (1948) Claude E. Shannon. A mathematical theory of communication. Bell System Technical Journal, 27(3):379â423, 1948. Shannon (1951) Claude E. Shannon. Prediction and entropy of printed English. Bell System Technical Journal, 30(1):50â64, 1951. Shumailov et al. (2024) Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. AI models collapse when trained on recursively generated data. Nature, 631(8022):755â759, 2024. Simon (1980) Julian L. Simon. Resources, population, environment: An oversupply of false bad news. Science, 208(4451):1431â1437, 1980. Simon (1981) Julian L. Simon. The Ultimate Resource. Princeton University Press, 1981. Sivathanu et al. (2024) Muthian Sivathanu, Yijia Zhao, and Vijay Janapa Reddi. Revisiting reliability in large-scale machine learning research clusters. arXiv preprint arXiv:2410.21680, 2024. Soros (1987) George Soros. The Alchemy of Finance: Reading the Mind of the Market. Simon and Schuster, New York, 1987. Strubell et al. (2019) Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645â3650, 2019. TSMC (2024) TSMC. TSMC arizona: Third fab announcement and CHIPS Act funding. https://pr.tsmc.com/english/news/3122, 2024. Turing (1936) Alan M. Turing. On computable numbers, with an application to the Entscheidungsproblem. Proceedings of the London Mathematical Society, s2-42(1):230â265, 1936. U.S. Energy Information Administration (2024) U.S. Energy Information Administration. Electric power annual 2023. https://w.eia.gov/electricity/annual/, 2024. von Neumann (1958) John von Neumann. The Computer and the Brain. Yale University Press, New Haven, 1958. von Neumann (1966) John von Neumann. Theory of Self-Reproducing Automata. University of Illinois Press, Urbana, 1966. Edited and completed by Arthur W. Burks. von Neumann and Morgenstern (1944) John von Neumann and Oskar Morgenstern. Theory of Games and Economic Behavior. Princeton University Press, 1944. Wiener (1948) Norbert Wiener. Cybernetics: Or Control and Communication in the Animal and the Machine. MIT Press, Cambridge, MA, 1948. Wiener (1950) Norbert Wiener. The Human Use of Human Beings: Cybernetics and Society. Houghton Mifflin, Boston, 1950.