Paper deep dive
LLMs Don't Pay for the Jump
Paras Balani, Subhrakanta Panda
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/23/2026, 3:08:05 AM
Summary
The paper argues that Large Language Models (LLMs) cannot perform genuine scientific abduction (the 'Jump') because they lack a physical mechanism where epistemic error incurs a tangible cost. Using Max Planck's resolution of the blackbody radiation problem as a case study, the authors demonstrate that Planck's abductive leap was driven by the physical unacceptability of infinite energy predictions (thermodynamic coupling), not just data compression or induction. They contend that fixed-weight transformer inference lacks this coupling, meaning LLMs do not suffer physical consequences for maintaining false models, unlike human scientists or physical systems subject to thermodynamic constraints.
Entities (10)
Relation Signals (8)
Rayleigh-Jeans Law → predicted → Infinite Energy
confidence 97% · The Rayleigh-Jeans law... predicted that a blackbody in thermal equilibrium would emit increasingly large amounts of energy... ultimately implying infinite energy at high frequencies.
Max Planck → performed → Abduction
confidence 96% · Planck's move to E = hν required no embodied simulation... We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost.
Large Language Models → lacks → Thermodynamic Coupling
confidence 95% · We formalize this distinction through thermodynamic coupling and show that fixed-weight transformer inference lacks such coupling, regardless of model scale.
Max Planck → resolved → Rayleigh-Jeans Law
confidence 95% · Max Planck resolved the blackbody radiation problem in 1900... Planck responded by changing the assumption about how matter exchanges energy with radiation.
Large Language Models → cannotperform → Abduction
confidence 94% · Zahavy [2026] argues that Large Language Models... cannot perform the abductive 'Jump'...
Floridi et al. → argues → LLMs Have Stochastic Core
confidence 92% · Floridi et al. [2025b]... argue that LLMs have a 'stochastic core' and only an 'abductive appearance.'
Epistemic Error → requires → Physical Cost
confidence 90% · We therefore argue that the missing ingredient in machine abduction may lie deeper than embodiment: a system must have some physical mechanism through which epistemic error becomes costly enough to force revision.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simulation. Zheng-Xin [2026] and Farmer [2026] question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding. Max Planck resolved the blackbody radiation problem in 1900. Planck's move to E = h{\nu} required no embodied simulation. It was motivated by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity, that could not be physically accepted. We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost. We formalize this distinction through thermodynamic coupling and show that fixed-weight transformer inference lacks such coupling, regardless of model scale. This is consistent with empirical results showing that output entropy remains nearly unchanged across tasks with sharply increasing causal difficulty, even as accuracy falls from 100% to 17%. We therefore argue that the missing ingredient in machine abduction may lie deeper than embodiment: a system must have some physical mechanism through which epistemic error becomes costly enough to force revision.
Tags
Links
- Source: https://arxiv.org/abs/2608.14397v1
- Canonical: https://arxiv.org/abs/2608.14397v1
Trouble viewing inline? Open PDF directly →
Full Text
52,508 characters extracted from source content.
Expand or collapse full text
LLMs Don’t Pay for the Jump Paras Balani 1 and Subhrakanta Panda 2 1 Department of Mathematics and Department of Computer Science, Birla Institute of Technology and Science, Pilani, Hyderabad Campus, Jawahar Nagar, Kapra Mandal, Medchal District, Telangana 500078, India 2 Department of Computer Science, Birla Institute of Technology and Science, Pilani, Hyderabad Campus, Jawahar Nagar, Kapra Mandal, Medchal District, Telangana 500078, India Abstract Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive “Jump” that produced Einstein’s equivalence princi- ple, and attributes this limitation to the absence of embodied simulation. Zheng-Xin [2026] and Farmer [2026] question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding. Max Planck resolved the blackbody radiation problem in 1900. Planck’s move to E = hν re- quired no embodied simulation. It was motivated by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity, that could not be physically accepted. We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost. We for- malize this distinction through thermodynamic coupling and show that fixed-weight transformer inference lacks such coupling, regardless of model scale. This is consistent with empirical results showing that output entropy remains nearly unchanged across tasks with sharply increasing causal difficulty, even as accuracy falls from 100% to 17%. We therefore argue that the missing ingredient in machine abduction may lie deeper than embodiment: a system must have some physical mechanism through which epistemic error becomes costly enough to force revision. 1 Introduction and Related Work In 1900, Max Planck encountered a failure that classical physics could not explain away. The Rayleigh-Jeans law, derived from the established principles of classical electrodynamics and sta- tistical mechanics, predicted that a blackbody in thermal equilibrium would emit increasingly large amounts of energy as frequency increased, ultimately implying infinite energy at high fre- quencies. This became known as the ultravio- let catastrophe. The prediction followed directly from the theory’s own assumptions and was deci- sively contradicted by observation. Planck there- fore had to alter the underlying picture itself. He proposed that energy exchange occurs in discrete amounts according to E = hν, introducing an as- sumption with no place in classical mechanics. The legitimacy of that leap requires a differ- ent account of what scientific discovery involves. A prominent view developed by Schmidhuber (2008) treats scientific discovery as fundamen- tally a problem of compression: finding the short- est program, rule, or description that accounts for a set of observations Schmidhuber [2008]. On this view, a scientist observes recurring regular- ities in data and searches for a more compact principle that captures them. This makes dis- covery closely related to Induction, where a gen- 1 arXiv:2608.14397v1 [cs.AI] 14 Aug 2026 eral rule is inferred from particular observations and results. It also fits naturally with recent progress in systems such as AlphaProof, which has demonstrated strong performance on formal mathematical reasoning through reinforcement learning Hubert et al. [2026]. Deduction begins with an established set of assumptions and de- termines what follows from them, while induc- tion attempts to infer a general rule from ob- served regularities. If scientific discovery were adequately described as the combination of these two operations, then the remaining task would appear to be one of scale and capability. Given enough experimental data, mathematical knowl- edge, and computational power, a sufficiently ca- pable Large Language Model should, in princi- ple, be able to reconstruct the hypotheses that scientists have discovered. Under this picture, Planck’s quantization hy- pothesis should have been accessible to a sys- tem capable of recognizing the failure of the Rayleigh–Jeans law, compressing the observed blackbody spectrum into a simpler description, and deriving the consequences of that descrip- tion. But Planck introduced a new constraint on what physical processes were allowed to oc- cur, one that was absent from the classical prin- ciples that had generated the anomaly in the first place. The problem therefore cannot be reduced to finding a regularity in the data and deriving its consequences. Planck had to propose a new physical possibility that classical theory did not contain. So, can a scientific discovery be reduced to the combination of induction and deduction, or whether some discoveries require changing the space of possible explanations itself. The Rayleigh-Jeans catastrophe was not a compression problem: there was no dataset to compress, because the observations that would eventually confirm quantized radiation had not yet been collected in a form that pointed to quantization. Nor was it a deduction problem: deduction from classical premises is precisely what produced the catastrophe. What Planck required was a third operation - the generation of an axiom that neither the data nor the exist- ing rules could supply, motivated instead by the sheer physical unacceptability of the alternative. Peirce (1934) named this operation Abduction: the inference of a Case, or a new Rule, from a Rule and a surprising Result Peirce [1934]. • Deduction (Rule + Case → Result): the analytic application of a rule to a case, the only mode of inference that guarantees truth. • Induction (Case + Result → Rule): the synthetic extraction of a general rule from repeated cases and results, validated by statistical frequency. • Abduction (Rule + Result → Case): the inference of a case, or a new rule, that would explain an otherwise inexplicable re- sult. Unlike deduction, abduction does not guar- antee truth, and unlike induction, it does not re- quire repeated cases. It responds to a result that a system cannot afford to leave unexplained. We take “cannot afford” literally. Planck did not re- ject classical electrodynamics because of accumu- lated statistical evidence, but because its predic- tion of infinite radiated energy at finite temper- ature was physically untenable. The distinction, we argue, lies in the cost of error. Small dis- crepancies can be absorbed through corrections or revised assumptions, but an unbounded con- tradiction demands new axioms. Genuine abduc- tion begins when a theory can no longer afford to preserve its existing rules. Zahavy [2026], in LLMs Can’t Jump, makes a similar argument using Einstein’s formulation of General Relativity. He argues that Newtonian physics faced no decisive empirical crisis when Einstein began his work, so Einstein’s abductive leap, expressed through the equivalence princi- ple, did not emerge from data compression. In- stead, Zahavy attributes it to embodied simu- lation, particularly Einstein’s imagined experi- ence of an observer in a falling elevator. Drawing on Magnani [2009]’s account of manipulative ab- duction and Harnad [1990]’s symbol grounding problem, Zahavy argues that LLMs manipulate physical language without access to the sensory and embodied experience that can give physical 2 concepts meaning. He therefore proposes physi- cally consistent, multimodal, action-controllable world models, such as those explored by Bruce et al. [2024] in Genie, as a possible route toward such grounding. But if embodiment is the missing ingredient in scientific abduction, then the problem is fun- damentally one of modality. LLMs lack direct sensory access and physical interaction, so they cannot ground their representations through per- ception or action. Planck’s case does not fit this explanation. His abductive move involved no bodily experience or physical interaction. He en- countered a mathematical consequence of classi- cal statistical mechanics: the Rayleigh-Jeans law implied an unbounded amount of radiated energy at high frequencies. Planck responded by chang- ing the formal assumptions governing the distri- bution of energy among electromagnetic modes and introducing the relation E = hν. The source of the new hypothesis was a contradiction within the mathematical theory itself. No sensory expe- rience was required to formulate it. Zheng-Xin [2026], in a public reply to Za- havy, raises a similar objection from another di- rection. He notes that Einstein’s route to Gen- eral Relativity was not the only possible one. Feynman’s later reconstruction begins from spe- cial relativity and quantum field theory, models gravity through a massless spin-2 graviton, and reaches the same field equations through a largely deductive route. Farmer [2026] on the other hand, argues that while most scientific abduction may require sen- sorimotor coupling, an exception is “identity ab- duction,” where two independently developed structures are recognized as equivalent through a shared representation. This can occur without physical interaction when a diagram or other rep- resentation reveals an invariant connecting them. Einstein’s equivalence principle may fit this cat- egory: the recognition that inertial and gravita- tional mass are the same quantity. This raises the possibility that Einstein’s insight depended less on embodiment itself than on a represen- tational relation revealed through an embodied thought experiment. Taken together, these ob- jections point to the same gap in the embodiment account. Embodiment may provide a sufficient mechanism for abduction, as Einstein’s elevator thought experiment suggests, but it does not ap- pear to be necessary in every case. Planck’s quantization hypothesis provides a historical ex- ample in which the abductive move arose from a formal contradiction without physical interaction or sensory experience. A related concern appears in the work of Floridi et al. [2025b], who approach the prob- lem from the philosophy of information. They argue that LLMs have a “stochastic core” and only an “abductive appearance.” On their ac- count, an LLM can produce statements that re- semble hypotheses generated through inference to the best explanation, but this apparent rea- soning may arise from statistical learning over human-generated text. The training corpus al- ready contains the results of scientific reasoning, including theories, explanations, arguments, and descriptions of evidence. The model can there- fore reproduce combinations of these products without performing the original inferential pro- cess that produced them. So basically, success- ful hypothesis generation alone does not estab- lish that the system has performed abduction. A system may produce a plausible explanation by exploiting patterns in a corpus that already con- tains explanations, without independently iden- tifying a surprising result and constructing a new hypothesis to account for it. Floridi et al. [2025a] develop a related argu- ment using category theory and the problem of symbol grounding. Their claim is that LLMs do not necessarily solve the grounding problem by acquiring direct connections between symbols and the physical world. Instead, they can oper- ate on symbols that have already been grounded by human agents. This connects their argument to Harnad’s account of epistemic parasitism. Human scientists acquire concepts through in- teraction with the world and then encode the resulting knowledge in language, mathematics, diagrams, and other symbolic systems. An LLM can subsequently learn from these symbolic products without reproducing the original pro- cess through which those symbols acquired their physical meaning. The model therefore inherits 3 a grounded conceptual structure from the corpus without having to establish the grounding itself. Sun and Saparov [2025] introduce InAbHyD, a synthetic benchmark designed to test whether LLMs can generate abductive hypotheses that are both correct and parsimonious, favoring the simplest explanation consistent with the evi- dence. They find that models perform rea- sonably well in simple world models, but their performance declines sharply as the underlying world model becomes more complex. This degra- dation persists despite in-context learning and reinforcement learning from verifiable rewards. Salimi et al. [2026] provide the first compre- hensive survey of abductive reasoning in LLMs. They formalize abduction as a two-stage pro- cess consisting of Hypothesis Generation and Hy- pothesis Selection, and review more than sixty studies across eleven model families ranging from 3B to 72B parameters. Their survey identifies a persistent gap in the field: limited mechanistic understanding of how abductive reasoning oc- curs. Existing work can show that abduction becomes unreliable as problems increase in com- plexity and novelty, but it has not yet estab- lished what properties a computational system must possess for abduction to succeed. This is the question we take up. Zahavy asks what cognitive mechanism Einstein pos- sessed that LLMs lack, while Zheng-Xin and Farmer question whether such a mechanism must be embodied at all. We ask what lies beneath these explanations: what must be true of the physical substrate of any bounded predictive sys- tem for an error to become costly enough to force the abandonment of a rule? This motivates our choice of Planck rather than Einstein as the cen- tral case study. Planck’s abduction was purely formal, and the crisis he faced had an unusually clear structure: classical theory predicted that a finite, measurable quantity must be infinite. His problem therefore makes the cost of persisting with an erroneous rule explicit. We argue that this cost is more than a metaphor. Any phys- ical computation, including inference in a lan- guage model, is subject to thermodynamic con- straints. Landauer [1961] showed that logically irreversible computation has an associated phys- ical cost, with the erasure of one bit requiring a minimum dissipation of k B T ln 2 under the stan- dard Landauer bound. This sets a physical limit on computation, but it does not determine how costly it is for a system to maintain a false model of the world. That second quantity is logically independent of computational cost. A system may have sufficient resources to compute an answer while having no internal pres- sure to abandon the rule that produced it. This distinction also changes how we interpret the role of embodiment. Zahavy identifies the ab- sence of embodiment as a possible obstacle to ab- duction because sensory and physical interaction may provide the grounding needed to construct new hypotheses. Floridi and his coauthors raise a different concern: an LLM may operate on repre- sentations that were already grounded by human agents, so adding sensory inputs would not by it- self explain how the system generates or grounds its own concepts. Our argument shifts the focus from the modality through which a system rep- resents the world to the physical consequences of maintaining an erroneous representation. 2 Background 2.1 Crisis of Classical Radiation Why does a physical anomaly sometimes force scientists to abandon a theory, while in other cases the theory survives after being modified? The answer depends on where the failure oc- curs. A physical theory connects assumptions about how nature works to mathematical equa- tions that generate predictions. When an obser- vation disagrees with a prediction, the disagree- ment does not automatically show that the the- ory is false. The problem may come from an inaccurate approximation, an incomplete calcu- lation, or an assumption that applies only within a limited range of conditions. The situation be- comes much more serious when the prediction follows directly from the theory’s basic assump- tions and the prediction is clearly contradicted by reliable observations. This was the problem Max Planck encountered in 1900 with blackbody ra- diation. A blackbody is an idealized object that 4 absorbs radiation and reaches thermal equilib- rium with its surroundings. Physicists wanted to calculate how much electromagnetic energy such an object should emit at each frequency. Classi- cal physics provided the necessary tools: classical electrodynamics described electromagnetic radi- ation, while statistical mechanics described how energy should be distributed among the possible modes of the electromagnetic field. From these principles, the Rayleigh–Jeans law was derived. At low frequencies, the law agreed with observa- tions. At high frequencies, however, it predicted that the emitted energy would increase without limit. If the law were correct at all frequencies, the total energy emitted by a blackbody would be infinite. This was called the ultraviolet catas- trophe because the problem became apparent at ultraviolet frequencies. The important point is that the divergence was not caused by a faulty measurement or a lack of data. It came from the mathematical consequences of classical the- ory itself. Keeping the classical assumptions un- changed meant accepting a prediction that phys- ical systems plainly did not obey. Planck responded by changing the assump- tion about how matter exchanges energy with ra- diation. Classical physics treated energy as con- tinuous, so an oscillator could gain or lose any amount of energy. Planck proposed that energy exchange occurred in discrete units, with the en- ergy of each unit given by E = hν. Here, E is the energy of one unit, ν is the fre- quency of the radiation, and h is a new constant of nature, now called Planck’s constant. This proposal introduced a new idea into physics. An oscillator associated with a high frequency could no longer exchange an arbitrarily small amount of energy. Each exchange required an amount proportional to that frequency. At high fre- quencies, the required energy therefore became large compared with the thermal energy available at a fixed temperature. High-frequency modes became much less likely to be excited, which prevented the unlimited accumulation of energy predicted by the Rayleigh–Jeans law. Planck’s formula consequently reproduced the observed blackbody spectrum. 2.2 Boltzmann, Entropy, and the Cost of Being Wrong Planck’s route to E = hν did not begin with radiation. It began with an idea he borrowed, reluctantly, from Ludwig Boltzmann, a physicist whose statistical approach remained controver- sial among many of his contemporaries. In the 1870s, Boltzmann proposed that entropy could be understood in terms of the number of micro- scopic arrangements compatible with a system’s observed macroscopic state. In its familiar form, this relation is S = k B lnW, where S is entropy, k B is Boltzmann’s con- stant, and W denotes the number of possi- ble microscopic configurations corresponding to the same macroscopic state [Boltzmann, 1877, Planck, 1901]. So, this connected a thermody- namic quantity, entropy, to the hidden micro- scopic structure of matter. To use W in this way, however, one had to assume that matter consisted of discrete atoms and molecules whose possible arrangements could be counted, even though those microscopic constituents could not be directly observed. Physicists such as Ernst Mach and Wilhelm Ostwald rejected the idea that atoms should be treated as physically real entities. They regarded atoms as useful theoretical constructions that could organize observations without committing physics to an unseen microscopic reality. From this perspective, thermodynamics should be for- mulated using quantities that could be measured directly, such as energy, temperature, and en- tropy. Boltzmann’s statistical mechanics took a different position. It treated macroscopic ther- modynamic behavior as the result of an enor- mous number of microscopic states, even though those states were inaccessible to direct observa- tion. Planck initially shared many of the reserva- tions about Boltzmann’s approach, which makes his later use of Boltzmann’s statistical reasoning important. To derive his radiation law, Planck 5 would have to rely on a microscopic picture that he had not previously regarded as physically se- cure [Janssen, 2019, Flamm, 1997]. Boltzmann maintained his position despite this opposition, and the cost was personal as well as professional. His statistical approach faced sustained criticism from influential physi- cists, while the existence of atoms remained dis- puted throughout much of his career. He died in 1906, before Jean Perrin’s experiments on Brow- nian motion provided strong experimental sup- port for atomism, building on Einstein’s theoret- ical analysis of Brownian motion in 1905. The history of Boltzmann’s later years has therefore often been discussed in connection with the isola- tion and despair he experienced while defending a microscopic view of matter that many of his contemporaries rejected. Boltzmann’s statistical framework provided the mathematical machinery that made Planck’s abduction possible. Planck used S = k B lnW to count the ways a fixed energy could be dis- tributed among discrete oscillators. This count- ing connected Planck’s quantization postulate to entropy and led to the specific relation E = hν. Boltzmann’s method therefore supplied the for- mal bridge between Planck’s new assumption and the radiation law it produced. But there is a second, more structural reason Boltzmann belongs in this story, and it is the rea- son we place him between the crisis (Section 2.1) and its resolution (Section 2.3) rather than treat- ing him as a mere technical aside. His atomism shows that defending a rejected axiom can carry real professional and personal costs. An axiom that survives sustained resistance therefore dif- fers from one that is proposed casually and aban- doned when difficulties arise. We return to this asymmetry in Section 5, where we argue that the cost of maintaining a commitment, rather than the accumulation of confirming evidence alone, helps distinguish genuine abduction from an or- dinary, revisable guess. 2.3 Planck’s Quantization Planck did not arrive at E = hν because he found Boltzmann’s statistical mechanics convinc- ing. By the autumn of 1900, he had exhausted the classical approaches he trusted, and he later described his decision as “an act of desperation,” undertaken because “a theoretical interpretation had to be found at any cost.” Throughout the 1890s, he had tried to derive the blackbody spec- trum within classical electrodynamics and ther- modynamics. When these approaches failed, he turned to Boltzmann’s statistical counting, de- spite his earlier reservations about the underlying microscopic picture. So basically, Planck introduced discrete en- ergy elements because the continuous alterna- tives available within classical theory all led back to the same divergence. Even after introducing E = hν, he initially treated quantization as a formal device for obtaining the correct radiation law. Its later interpretation as a physical prin- ciple developed through Einstein’s 1905 work on the photoelectric effect and subsequent quantum theory. Planck’s move therefore fits Peirce’s ac- count of abduction: E = hν was neither induced from an established pattern nor deduced from ex- isting principles. It was a new axiom introduced because preserving the existing rule system led to a physically unacceptable result. 3 The Limits of Inductive Infer- ence Compression rewards a rule to the extent that the rule makes existing data more predictable; it offers no comparable reward for a rule that also happens to prevent a future, not-yet-collected measurement from contradicting a different the- ory’s extrapolation. Nothing about a com- pression objective, evaluated against the data available in 1900, would have penalized the Rayleigh–Jeans law for its behavior at frequen- cies where no data yet existed to compress, since a compression-driven search only optimizes over the observations it has been given. Planck’s ab- duction was motivated by exactly this kind of ex- 6 trapolated, not-yet-observed failure: the mathe- matical certainty that the classical law, followed to its logical conclusion, would demand an infi- nite quantity. No accumulation of the low- and mid-frequency data actually in hand, however large, could have surfaced that failure through pattern-matching alone. What was needed in- stead was a system able to take the classical rule’s own unconfirmed extrapolation seriously enough to treat its consequence as intolerable before that consequence had been measured. As Section 2.3 notes, the experiments that later established quantization as a physical prin- ciple, including Einstein’s 1905 explanation of the photoelectric effect, Bohr’s 1913 atomic model, and Compton’s 1923 scattering results, were not yet available in 1900. Induction can infer only from the cases and results available to it. It cannot directly anticipate a rule whose strongest confirmation will come from future ob- servations. By the standard of fit to existing data, Planck’s interpolation formula was already sufficient. The deeper claim E = hν was not demanded by the evidence available at the time. Its significance came from making a specific, fal- sifiable claim about microscopic energy exchange before the evidence needed to support that claim had appeared. Compression therefore provides no natural in- centive to prefer a rule because it avoids a fu- ture contradiction that is absent from the cur- rent dataset. A compression objective evaluates how well a rule explains the observations it has already been given. It has no direct penalty for the behavior of a theory in an unobserved regime. 4 The Limits of Deduction Once Planck’s postulate was in place, the prob- lem changed from proposing a new principle to deriving its consequences. Given E = hν as an axiom, the derivation of the blackbody spectrum, the recovery of the Rayleigh–Jeans law in the low-frequency limit, and the Stefan–Boltzmann law from the total emitted energy are deduc- tive tasks. In Peirce’s schema, they follow the form Rule + Case → Result: once the premises are fixed, the task is to determine what fol- lows from them. This is the class of reasoning at which modern formal systems perform well. Proof assistants based on dependent type the- ory and reinforcement-learned systems such as AlphaProof [Hubert et al., 2026] can derive com- plex results from explicitly stated premises, with recent systems reaching strong performance on formal mathematical benchmarks. Deduction, however, cannot produce the ax- iom from which the deduction begins. An ax- iom is a premise, not a theorem derived from prior premises. No chain of valid deductions from classical electrodynamics and classical sta- tistical mechanics could yield E = hν, because those premises treat energy exchange as continu- ous. The Rayleigh–Jeans catastrophe was there- fore not a failure of deduction. It was the correct deductive consequence of the classical assump- tions available in 1900. A system restricted to those assumptions could derive the divergent en- ergy integral with complete logical accuracy, but it could not use deduction alone to replace the as- sumptions that produced it. Deduction explores the logical space defined by a set of premises. Planck’s move required leaving that space and introducing a premise that the existing theory did not contain. This is the distinction between deriving the consequences of a theory and gener- ating a new theory. Suppose a modern reasoning system were asked to find inconsistencies in nineteenth- century physics. It could identify the Rayleigh– Jeans divergence as a mathematically anomalous result: the theory predicts a divergent integral for a quantity that experiments show to be fi- nite. A sufficiently thorough search could there- fore flag the catastrophe as an anomaly. But detecting an inconsistency does not determine whether it requires a change in theory. Classical physics contained many approximations and ide- alizations that were tolerated without prompting a revision of its foundations. The difference was the cost of the error. A 2% discrepancy might be attributed to measurement error or an im- perfect approximation, whereas an infinite pre- diction for a finite observable quantity cannot be absorbed in the same way. Recognizing that 7 distinction requires more than deduction. The system must have some prior criterion for how costly an error is and when that cost becomes unacceptable. Without such a criterion, a de- ductive search can identify anomalies but cannot determine which ones justify abandoning an ex- isting rule. Planck’s own decade of unsuccessful attempts to preserve the classical account illus- trates this distinction: the problem was visible before he accepted that the cost of retaining the classical framework had become too high. 5 Thermodynamic (De)coupling: Formal Framework 5.1 Descriptive versus Regulatory Uncertainty We distinguish two roles that uncertainty can play in a predictive system, following the termi- nology of Gamal Eldin [2026]. Definition 1 (Thermodynamic cou- pling). A computational system is thermody- namically coupled if the physical cost of an op- eration, E cost , increases with the epistemic error of its output, ε =|ˆy− y ∗ |: ∂E cost ∂ε > 0. In a coupled system, epistemic error has a direct physical consequence. Maintaining an in- correct prediction becomes more costly as the er- ror increases, creating a physical pressure to re- vise the prediction or the rule that produced it. Uncertainty is therefore regulatory: the system’s physical cost depends on the accuracy of its pre- dictions, so persistent error creates pressure for correction. Definition 2 (Thermodynamic decou- pling). A system is thermodynamically decou- pled if ∂E cost ∂ε = 0, so that the energetic cost of an operation is inde- pendent of its epistemic error, conditional on the state of the physical substrate. Uncertainty in such a system is therefore descriptive: it charac- terizes the system’s predictive uncertainty with- out affecting the physical cost or subsequent dy- namics of the computation. Whether a predic- tion is correct or incorrect does not, by itself, change the energetic cost of producing it. Every operation performed by a bounded physical system remains subject to Landauer’s (1961) bound, with each irreversible bit erasure dissipating at least k B T ln 2 joules. This cost ap- plies regardless of whether the erased informa- tion corresponds to a correct inference or an er- roneous prediction. The distinction we draw is narrower: whether the physical cost of compu- tation depends on epistemic quality, or whether it remains independent of whether the system is correct. This distinction allows us to restate the lim- its identified in Sections 3 and 4. Planck’s cri- sis was difficult to ignore because the classical theory produced an unbounded physical conse- quence. An infinite predicted energy could not be treated as a small residual or absorbed through a minor correction. Planck spent years attempt- ing to preserve the classical framework before the cost of maintaining it became greater than the cost of abandoning its underlying assumptions. Boltzmann’s history shows the same asymmetry at the level of a scientific career. Defending atom- ism imposed sustained professional and personal costs despite limited acceptance at the time. In both cases, the pressure to change a commitment came from the cost of maintaining it, not simply from the accumulation of evidence against it. We argue that this distinction separates genuine ab- ductive commitment from an ordinary hypothe- sis that can be revised without significant cost. 5.2 The Softmax–Gibbs Analogy The softmax function used to convert logits into a token distribution, P(v i | ctx) = exp(x i /τ) P j exp(x j /τ) , has the same mathematical form as the Gibbs– Boltzmann distribution over microstates, with 8 the inference-time parameter τ playing a formal role analogous to temperature. Theorem 1 (Softmax decoupling). In a transformer with fixed weights θ and fixed in- ference temperature τ, the token-level Shannon entropy H t =− X i P θ (v i | ctx t ) logP θ (v i | ctx t ) is a deterministic function of (θ,τ, ctx t ) alone. It is therefore independent of whether ctx t is in- distribution or requires causal extrapolation be- yond the training data. For fixed θ and ctx t , the transformer pro- duces a fixed logit vector x t . The temperature τ determines the corresponding token probabil- ities through the softmax function, and these probabilities uniquely determine H t . No vari- able in this computation represents the phys- ical temperature of the hardware, the instan- taneous power dissipated by the computation, or the eventual correctness of the generated to- ken. Consequently, two contexts can produce the same predictive entropy while differing in their epistemic status. One may correspond to a fa- miliar pattern that the model has learned reli- ably, while another may require an extrapolation unsupported by the training distribution. The softmax entropy records the distribution of the model’s output probabilities, but the computa- tion that produces those probabilities does not assign any additional physical cost according to whether the prediction is correct or incorrect. 5.3 Cost-Coupled Abduction: A Pro- posed Criterion We can now state the criterion developed in the preceding sections. Induction and deduction can identify patterns and derive consequences, but neither provides a mechanism for treating a par- ticular error as intolerable. Section 5.1 called this missing mechanism regulatory uncertainty, while Section 5.2 showed that transformer inference ex- hibits the opposite structure: token entropy de- scribes the output distribution without depend- ing on whether the resulting claim is correct. Criterion (Cost-Coupled Abduction). A bounded predictive system can generate a gen- uinely new axiom only if some internal cost is causally coupled to the magnitude of its epis- temic error on a specific result. Formally, ∂E cost ∂ε > 0, where the coupling must operate locally when a particular anomaly is encountered, so that main- taining the rule responsible for the error becomes increasingly costly until the rule is revised. Planck’s history is consistent with this crite- rion, although it does not provide a direct mea- surement of his cognitive cost function. The relevant proxies are the singular divergence he could not absorb as an ordinary approximation, his years of unsuccessful attempts to preserve the classical framework, and the costs associated with maintaining controversial commitments in Boltzmann’s case. These pressures preceded the later evidence that confirmed quantization, so the commitment cannot be explained solely by inductive reward from subsequent observations. Current LLM inference does not exhibit this form of coupling. A correct and incorrect out- put produced by the same model on the same hardware does not incur a systematically differ- ent computational cost, and Theorem 1 shows that output entropy is determined by the model and context rather than by the truth of the re- sulting claim. An LLM can reproduce Planck’s derivation once E = hν is supplied as a premise, but nothing in the inference process makes the absence of that premise increasingly costly when the classical theory produces an anomaly. The proposed missing ingredient is therefore not de- ductive ability or sensory modality alone, but a physical mechanism through which epistemic er- ror creates a local cost that can force revision of the rule producing it. 6 Empirical Signatures of De- coupling Gamal Eldin [2026] tested whether token-level Shannon entropy tracks causal difficulty using 9 tasks that ranged from retrieval of established results (“Kepler”), to causal reasoning with novel parameters (“Newton”), to extrapolation beyond plausible training coverage (“Newton OOD”). Across three Llama models spanning two orders of magnitude in size, from 3B to 70B parameters, entropy remained statistically flat within each model across the three categories, with ranges of only 0.011 to 0.028 nats and all p ≥ 0.568. Accuracy, however, ranged from 0% to 100%. The 70B model achieved perfect accuracy on Ke- pler tasks but only 17% on Newton OOD, while entropy differed by just 0.011 nats between the two. The same pattern appeared in N = 1,000 examples evaluated on GPT-4o-mini with pro- grammatic scoring: accuracy varied by 33.2 per- centage points while entropy varied by only 0.019 nats. These results are consistent with the decou- pling described by Theorem 1. The model’s pre- dictive entropy changes little as the task shifts from familiar retrieval to difficult causal extrapo- lation, even when accuracy changes substantially. Larger models also become more confident over- all, without a corresponding increase in the rela- tionship between confidence and correctness. Sun and Saparov [2025] find that LLMs gen- erate high-quality, parsimonious hypotheses in simple scenarios but struggle as the underlying world models become more complex. This degra- dation persists under both in-context learning and reinforcement learning from verifiable re- wards. Salimi et al. [2026], synthesizing more than sixty studies across eleven model fami- lies, report a similar pattern and identify “lim- ited mechanistic understanding of abductive pro- cesses” as an open problem. Taken together with the entropy results, these findings point to a more specific form of insensitivity: thermodynamic de- coupling, where increasing epistemic difficulty does not produce a corresponding physical cost that could drive the system toward revision. 7 Toward Thermodynamically Coupled Architectures We are not proposing thermodynamic coupling as a wholly new requirement. Biological nervous systems provide an existing example of such cou- pling, most prominently formalized by Friston’s (2010) free-energy principle [Friston, 2010]. In this framework, the nervous system minimizes a bound on surprise through prediction error and its estimated precision. Prediction error there- fore has consequences for the system’s physi- cal state: unresolved errors can drive changes in synaptic activity, attention, autonomic re- sponses, and metabolic demand. In the termi- nology of Section 5.1, biological prediction error is regulatory because epistemic error is linked to the physical processes that maintain and update the system. This provides a concrete comparison with Planck’s case and with transformer inference. If a predictive system is genuinely thermodynam- ically coupled, larger or more persistent errors should produce measurable changes in physiolog- ical or computational resource allocation, such as increased arousal, sustained attention, or reallo- cation of processing resources. Theorem 1 shows that current transformer inference has no corre- sponding coupling in its token-level uncertainty. This comparison does not establish that biolog- ical systems can perform abduction while trans- formers cannot. It instead identifies the missing property of a thermodynamically decoupled sys- tem: epistemic error has no physical consequence that would compel the system to revise the rule producing it. One possible engineering direction is neuro- morphic hardware in which prediction error con- tinues to drive weight updates during inference, with the energy cost of each update increasing with the magnitude of the error being corrected. This corresponds to the “Intrinsic Cost” mecha- nism proposed by Gamal Eldin (2025). Such a system would couple computational cost to epis- temic error: sustaining an incorrect prediction would consume additional energy until the un- derlying representation or rule was revised. This 10 would create a physical analogue of the costs as- sociated with Boltzmann’s defense of atomism and Planck’s repeated attempts to preserve the classical account. 8 Limitations A framework that explains past observations is not sufficient unless it also makes predictions that could fail. Section 5.3 therefore yields two direct tests. First, if a thermodynamically coupled architecture of the kind proposed in Section 7 is evaluated on the task suite from Section 6, its internal cost or uncertainty sig- nal should track accuracy as causal demand in- creases, unlike the entropy of current transform- ers. If the signal remains flat despite declining accuracy, this would count as evidence against the criterion itself. Second, the criterion predicts an effect on hypothesis quality, not only calibration. If the degradation in abductive performance reported by Sun and Saparov [2025] reflects thermody- namic decoupling, then a coupled system should show less degradation as world-model complex- ity increases. If hypothesis quality declines by the same amount while uncertainty becomes bet- ter calibrated, coupling would improve calibra- tion without improving abduction, weakening the stronger form of the criterion. We do not test these predictions here. 9 Conclusion Zahavy [2026] approaches the problem through Einstein: what allows a scientist to move from an existing theory to the axioms of a new one when neither accumulated data nor deduction can provide those axioms? We approach the same question through Planck. His route to E = hν required no embodied thought experi- ment. It began with a finite, measured quan- tity that classical theory predicted to be infinite [Planck, 1901]. Sections 3 and 4 showed why nei- ther induction nor deduction could supply the required postulate. The available data could al- ready be fitted without quantization, while the classical premises produced the catastrophe when their consequences were derived correctly. What was missing was a mechanism that made the anomaly costly enough to demand a change in the underlying rule. Boltzmann’s experience de- fending atomism provides a second example of such sustained cost. Section 5 formalized this idea as thermody- namic coupling. A system is coupled when the physical cost of computation depends on epis- temic error, and decoupled when that cost is independent of correctness. Theorem 1 shows that token entropy in a fixed transformer does not provide such coupling. Section 6 then showed empirically that entropy remains nearly unchanged even as accuracy falls sharply across tasks and model scales. Section 7 considered possible coupled systems, drawing on biological prediction-error dynamics and neuromorphic ar- chitectures. Section 8 turned this proposal into an empirical question: whether introducing such coupling would improve abductive performance under increasing causal and model complexity re- mains to be tested. The scope of the argument is deliberately lim- ited. Our case comes from the physical sciences, where theoretical error can produce a measurable physical consequence and where cost therefore has a concrete referent. We do not claim that the same criterion applies unchanged to mathematics or computer science, where the objects of inquiry are formal systems. Within the physical sciences, our claim is narrower: current language models may possess substantial capabilities for induction and deduction, yet their computational substrate contains no mechanism that makes epistemic er- ror itself increasingly costly. Cost-coupled ab- duction proposes that such a mechanism may be necessary when a system must decide that an ex- isting rule has become too expensive to retain. 11 References Ludwig Boltzmann. über die beziehung zwischen dem zweiten hauptsatze der mechanischen wärmetheorie und der wahrscheinlichkeitsrechnung respektive den sätzen über das wärmegle- ichgewicht. Sitzungsberichte der Kaiserlichen Akademie der Wissenschaften, 76:373–435, 1877. Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal Behbahani, Stephanie Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh, and Tim Rocktäschel. Genie: Generative interactive environments. arXiv preprint, 2024. URL https://arxiv.org/abs/2402.15391. Michael Farmer. Abduction without a body? representational grounding and the abduction loop for scientific hypothesis generation, 2026. URL https://arxiv.org/abs/2608.02505. Dieter Flamm. Ludwig boltzmann: A pioneer of modern physics. arXiv preprint, 1997. URL https://arxiv.org/abs/physics/9710007. Luciano Floridi, Yiyang Jia, and Fernando Tohmé. A categorical analysis of large language models and why llms circumvent the symbol grounding problem. arXiv preprint arXiv:2512.09117, 2025a. URL https://arxiv.org/abs/2512.09117. Luciano Floridi, Jessica Morley, Claudio Novelli, and David Watson. What kind of reasoning (if any) is an llm actually doing? on the stochastic nature and abductive appearance of large language models. arXiv preprint arXiv:2512.10080, 2025b. URL https://arxiv.org/abs/2512.10080. Karl J. Friston. The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2):127–138, 2010. doi: 10.1038/nrn2787. Ahmed Gamal Eldin. Descriptive versus regulatory uncertainty in bounded predictive systems. arXiv preprint arXiv:2605.18909, 2026. URL https://arxiv.org/abs/2605.18909. Stevan Harnad. The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3):335– 346, 1990. doi: 10.1016/0167-2789(90)90087-6. Thomas Hubert, Rishi Mehta, Laurent Sartran, Miklós Z. Horváth, Goran Žužić, Eric Wieser, Aja Huang, Julian Schrittwieser, Yannick Schroecker, Hussain Masoom, Ottavia Bertolli, Tom Zahavy, Amol Mandhane, Jessica Yung, Iuliya Beloshapka, Borja Ibarz, Vivek Veeriah, Lei Yu, Oliver Nash, Paul Lezeau, Salvatore Mercuri, Calle Sönne, Bhavik Mehta, Alex Davies, Daniel Zheng, Fabian Pedregosa, Yin Li, Ingrid von Glehn, Mark Rowland, Samuel Albanie, Ameya Velingker, Simon Schmitt, Edward Lockhart, Edward Hughes, Henryk Michalewski, Nicolas Sonnerat, Demis Hassabis, Pushmeet Kohli, and David Silver. Olympiad-level formal mathematical reasoning with reinforcement learning. Nature, 651:607–613, 2026. doi: 10.1038/s41586-025-09833-y. URL https://w.nature.com/articles/s41586-025-09833-y. Michel Janssen. Planck, the second law of thermodynamics, and black-body radiation. In Construct- ing Quantum Mechanics: Volume 1, The Scaffold, 1900–1923, pages 45–83. Oxford University Press, 2019. doi: 10.1093/oso/9780198845478.003.0002. Rolf Landauer. Irreversibility and heat generation in the computing process. IBM Journal of Research and Development, 5(3):183–191, 1961. doi: 10.1147/rd.53.0183. 12 Lorenzo Magnani. Abductive Cognition: The Epistemological and Eco-Cognitive Dimensions of Hypothetical Reasoning, volume 3. Springer, 2009. Charles Sanders Peirce. Collected Papers of Charles Sanders Peirce. Harvard University Press, 1934. Max Planck. Ueber das gesetz der energieverteilung im normalspectrum. Annalen der Physik, 309 (3):553–563, 1901. doi: 10.1002/andp.19013090310. Moein Salimi, Shaygan Adim, Danial Parnian, Nima Alighardashi, Mahdi Jafari Siavoshani, and Mohammad Hossein Rohban. Wiring the ’why’: A unified taxonomy and survey of abductive reasoning in llms. arXiv preprint arXiv:2604.08016, 2026. URL https://arxiv.org/abs/2604. 08016. Jürgen Schmidhuber. Driven by compression progress: A simple principle explains essential aspects of subjective beauty, novelty, surprise, interestingness, attention, curiosity, creativity, art, science, music, jokes. In Anticipatory Behavior in Adaptive Learning Systems, pages 48–76. Springer, 2008. Yunxin Sun and Abulhair Saparov. Language models do not follow occam’s razor: A benchmark for inductive and abductive reasoning. arXiv preprint arXiv:2509.03345, 2025. URL https: //arxiv.org/abs/2509.03345. Tom Zahavy. Llms can’t jump. preprint, 2026. URL https://w.tomzahavy.com/files/ llms-cant-jump.pdf. Yong Zheng-Xin. Hot take: Llm can “jump”. https://yongzx.github.io/blog/2026/08/08/ llm-can-jump, August 2026. Accessed August 13, 2026. A Author’s Note I am pursuing an MSc in Mathematics alongside a B.E. in Computer Science, while working on cosmology problems involving Einstein’s field equations in modified theories of gravity. My work is largely mathematical and computational, which has made the role of assumptions in physical theories particularly concrete. In modified gravity, changing parameters of an imposed energy condition can lead to a different set of field equations and, consequently, a different physical theory. This raises a question that deduction alone cannot answer: why choose one consistent set of assumptions over another, and what would it mean for an AI system to generate such assumptions? Zahavy’s (2026) work gave me a vocabulary for this question through Peirce’s triad and the role of abduction. Reading it alongside Gamal Eldin’s (2026) work on thermodynamic decoupling led to the synthesis developed in this paper. My mathematical and computational background made the entropy-flatness results particularly natural to interpret, while my work with modified gravity made the problem of generating new axioms feel closely connected to problems I encounter in theoretical physics. It is important to note that this work does not attempt a novel historical or physical account of Planck’s route to quantization, since such analysis falls outside my expertise. Instead, I draw on established sources to examine the conditions under which AI systems might move beyond inference toward genuine scientific discovery. 13 B Supplementary Formal Details 1. Classical Equipartition: At thermal equilibrium, classical statistical mechanics assigns each electromagnetic mode an average energy k B T, independent of frequency. Combined with the classical mode density 8πν 2 /c 3 , this yields the Rayleigh–Jeans law and the ultraviolet catastrophe: the number of high-frequency modes grows without bound while each receives the same energy. 2. Boltzmann’s Statistical Entropy: Entropy is related to the number of microscopic con- figurations W compatible with a macroscopic state, S = k B lnW. As discussed in Section 2.2, this required treating matter as composed of discrete, countable microstates, a controversial assumption defended by Boltzmann. This statistical framework supplied the machinery Planck needed to formulate his quantization. 3. Planck’s Quantization: Planck proposed that energy exchange occurs in discrete units, E = hν. This postulate was not derived from the previous assumptions. Planck introduced it because maintaining continuous energy exchange alongside the observed finite blackbody spectrum was untenable. Combined with Boltzmann’s counting, it gives ⟨E(ν,T)⟩ = hν exp (hν/k B T)− 1 , which approaches k B T when hν ≪ k B T and suppresses high-frequency contributions expo- nentially. 4. The Correspondence Requirement: Any replacement for classical equipartition must reduce to it in the regime where the classical law was already correct. Item 3’s average energy satisfies lim hν/k B T→0 ⟨E(ν,T)⟩ = k B T, recovering Item 1 at low frequency or high temperature, exactly where the Rayleigh–Jeans law matched observation. A postulate that failed this limit would not have been an admissible replacement, regardless of how well it resolved the divergence. This is the same constraint discussed in Section 2.3, where Planck initially treated quantization as a formal device rather than a physical claim. 5. Thermodynamic Decoupling (Formal Statement): For a bounded computational sys- tem with fixed weights θ, fixed inference temperature τ, and inference-time output entropy H t as defined in Theorem 1, decoupling holds when ∂E cost ∂ε = 0, ε =|ˆy− y ∗ |, conditioned on the substrate state (T,P). Proposition B.1 and Theorem 1 (Section 5.2) establish that H t is a function of (θ,τ, ctx t ) alone, with no dependence on ε, T, or P. This item states the general decoupling condition, of which Theorem 1’s entropy result is one specific instance, restricted to token-level Shannon entropy as the observable proxy for uncertainty. 14