Paper deep dive
Spilled Energy in Large Language Models
Adrian Robert Minut, Hazem Dewidar, Iacopo Masi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 9:37:17 PM
Summary
This paper introduces a training-free method for detecting hallucinations in Large Language Models (LLMs) by reinterpreting the final softmax classifier as an Energy-Based Model (EBM). The authors define two metrics, 'spilled energy' and 'marginalized energy', derived from output logits. Spilled energy captures discrepancies between energy values across consecutive generation steps, which empirically correlate with factual errors, biases, and failures. The method is evaluated on state-of-the-art LLMs (LLaMA, Mistral, Gemma, Qwen3) across nine benchmarks, demonstrating robust hallucination detection without requiring trained probe classifiers or activation ablations.
Entities (10)
Relation Signals (6)
Spilled Energy → detects → Hallucination
confidence 95% · spilled energy... correlate with factual errors, biases, and failures
LLM Softmax Classifier → reinterpretedas → Energy-Based Model
confidence 95% · We reinterpret the final Large Language Model (LLM) softmax classifier as an Energy-Based Model (EBM)
Spilled Energy → derivedfrom → LLM Output Logits
confidence 90% · two completely training-free metrics derived directly from output logits: spilled energy
Spilled Energy → measures → Discrepancy in Energy Values
confidence 90% · captures the discrepancy between energy values across consecutive generation steps
Grathwohl et al. → inspired → EBM Reinterpretation
confidence 85% · taking inspiration from what Grathwohl et al. (2020) did for classifiers
Orgad et al. → usedmethod → Probe Classifiers
confidence 80% · Similar to Orgad et al. (2025)... we abandon the idea of using a probe classifier
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We reinterpret the final Large Language Model (LLM) softmax classifier as an Energy-Based Model (EBM), decomposing the sequence-to-sequence probability chain into multiple interacting EBMs at inference. This principled approach allows us to track "energy spills" during decoding, which we empirically show correlate with factual errors, biases, and failures. Similar to Orgad et al. (2025), our method localizes the exact answer token and subsequently tests for hallucinations. Crucially, however, we achieve this without requiring trained probe classifiers or activation ablations. Instead, we introduce two completely training-free metrics derived directly from output logits: spilled energy, which captures the discrepancy between energy values across consecutive generation steps that should theoretically match, and marginalized energy, which is measurable at a single step. Evaluated on nine benchmarks across state-of-the-art LLMs (including LLaMA, Mistral, and Gemma) and on synthetic algebraic operations (Qwen3), our approach demonstrates robust, competitive hallucination detection and cross-task generalization. Notably, these results hold for both pretrained and instruction-tuned variants without introducing any training overhead. Code available at: this http URL
Tags
Links
- Source: https://arxiv.org/abs/2602.18671v4
- Canonical: https://arxiv.org/abs/2602.18671v4
Trouble viewing inline? Open PDF directly →
Full Text
96,743 characters extracted from source content.
Expand or collapse full text
Spilled Energy in Large Language Models Adrian R. Minut 1,2 &Hazem Dewidar 1,2 &Iacopo Masi 1 Sapienza University of Rome, Italy 1 OmnAI Lab 2 GLADIA Abstract We reinterpret the final Large Language Model (LLM) softmax classifier as an Energy-Based Model (EBM), decomposing the sequence-to-sequence probability chain into multiple interacting EBMs at inference. This principled approach allows us to track “energy spills” during decoding, which we empirically show correlate with factual errors, biases, and failures. Similar to Orgad et al. (2025), our method localizes the exact answer token and subsequently tests for hallucinations. Crucially, however, we achieve this without requiring trained probe classifiers or activation ablations. Instead, we introduce two completely training-free metrics derived directly from output logits: spilled energy, which captures the discrepancy between energy values across consecutive generation steps that should theoretically match, and marginalized energy, which is measurable at a single step. Evaluated on nine benchmarks across state-of-the-art LLMs (including LLaMA, Mistral, and Gemma) and on synthetic algebraic operations (Qwen3), our approach demonstrates robust, competitive hallucination detection and cross-task generalization. Notably, these results hold for both pretrained and instruction-tuned variants without introducing any training overhead. Code available at github.com/OmnAI-Lab/spilled-energy/ Q/A: ‘What is the capital of Italy? Answer:’ Logit ThecapitalofItalyisRome ✓ThecapitalofItalyisSydney ✗ Spilled (Ours) ThecapitalofItalyisRome ✓ThecapitalofItalyisSydney ✗ Reasoning: ‘A farmer has 12 chickens. Each chicken lays 2 eggs per day. How many eggs will the farmer collect in 5 days?’ Logit 12chickenslay2eggsperday.In5days,thefarmerwillcollect12x2x5=120eggsin5days ✓12chickenslay2eggsperday.In5days,thefarmerwillcollect12x2x5=470eggsin5days ✗ Spilled (Ours) 12chickenslay2eggsperday.In5days,thefarmerwillcollect12x2x5=120eggsin5days ✓12chickenslay2eggsperday.In5days,thefarmerwillcollect12x2x5=470eggsin5days ✗ Figure 1: Color-coded comparison of hallucination detection with LLaMa-3 8B using logit confidence and spilled energy. Our method generalizes well across topics (e.g., Q&A, reasoning) and diverse LLMs. ✓ indicates a correct answer and ✗ an incorrect one. While our approach focuses on the exact answer tokens (e.g. Rome/Sydney and 120/470, see Section 4.2), here we apply min–max normalization to the full answer for visualization, as truthful hallucination. 1 Introduction The widespread adoption of Large Language Models (LLMs) across various domains has brought increasing attention to their critical limitation: their tendency to generate incorrect or misleading information—commonly referred to as “hallucinations.” This issue supports the idea that LLMs are just stochastic parrots (Bender et al., 2021) answering in a way that is statistically plausible with respect to the input prompt despite not having a real understanding of it. On the other side, recent reasoning capabilities proper to ChatGPT 4o (OpenAI-Team, 2023) or Deepseek (Liu et al., 2024) offer counter evidence to actually support this. Ongoing research seeks to characterize and categorize hallucinations, setting them apart from other error types (Liu et al., 2022; Ji et al., 2023; Huang et al., 2023b; Rawte et al., 2023). At the same time, recent discussions have introduced terms such as confabulations (Millidge, 2023) and fabrications (McGowan et al., 2023), sometimes attributing a form of “intention” to LLMs—though the very idea of LLM “intentionality” and other human-like qualities remains contested (Salles et al., 2020; Serapio-García et al., 2023; Harnad, 2024). Research on LLM hallucinations can be categorized into two main branches: the first one is the extrinsic branch, where the hallucinations are measured with respect to the interpretation that humans give to those errors (Bang et al., 2023; Ji et al., 2023; Huang et al., 2023b; Rawte et al., 2023). The second branch was started by Kadavath et al. (2022b), proposing to study the hallucinations within the model itself. Following Kadavath et al. (2022b), the work in Li et al. (2024) proposes Inference-Time Intervention (ITI) as a way to improve the “truthfulness” of LLMs at inference time. ITI functions by altering model activations at inference time, steering them along specific directions within a restricted set of attention heads. Our work is also different from Yin et al. (2023), since we care about detecting errors in LLMs, whereas they introduce an automated methodology to detect when LLMs are aware that they do not know how to answer. In this work, we follow the definition of hallucinations given by Orgad et al. (2025) as any form of error produced by an LLM—including factual mistakes, biased outputs, breakdowns in common-sense reasoning, and related issues. Like them, we also confirm that the truthfulness signal is concentrated in the “exact answer tokens.” Nevertheless, unlike them, we abandon the idea of using a probe classifier (Belinkov, 2022) trained for each task and dataset. Given that LLMs are foundational models, user interactions typically occur in the wild, making it difficult to predict which probe classifier is best suited for detecting hallucinations in real-world scenarios. Furthermore, in this setting, classifier weights should not only be updated dynamically for each task, but the optimal token–layer combination is also dataset-dependent, which conflicts with the broad LLM applicability. Indeed, in the work by Orgad et al. (2025), the authors report: “We find that probing classifiers do not generalize across different tasks.” In our paper, we propose to solve this problem with a training-free method that generalizes better across different tasks and is mathematically principled using the framework of Energy-based Models (EBMs). Fig. 1 reports a qualitative comparison across tasks, comparing to the logit confidence. Additional samples are shown in Section D.2. We reinterpret the final softmax classifier over the vocabulary of LLM as an EBM, taking inspiration from what Grathwohl et al. (2020) did for classifiers. This perspective enables us to decompose the sequence-to-sequence probability chain into multiple interacting EBMs that operate jointly during inference. Through this decomposition, we introduce the notion of “spilled energy” in LLM decoding and show empirically that such spill strongly correlates with errors. Given that our method is solely based on the mathematics of EBMs and the chain rule of probability, we do not have to train or tune our detector, striking a good generalization across tasks and LLMs. Building on this foundation, our contributions are as follows: ⋄ Training-free, LLM hallucination detection generalizing across tasks using the EBM framework. We introduce a method for detecting hallucinations that requires no additional training, in contrast to prior work that relies on trained classifiers and ablations of model activations. Our approach directly reads values inside the LLM, enabling natural generalization across tasks and performing better than logit-based detection. ⋄ Two energy-based metrics. We define two complementary measures of energy spills: (i) delta energy ΔE(i:1) E_ θ(x_i:1), which captures discrepancies between energy values across two time steps that should be mathematically equivalent, and (i) marginal energy Em(i:1)E^m_ θ(x_i:1), which can be evaluated at a single time step. ⋄ Scalable and generalizable analysis. Our framework is mathematically principled, training-free, and exhibits strong cross-dataset generalization. We scale our analysis to state-of-the-art LLMs, including Llama 3-8B-Instruct and Mistral-7B-Instruct, and demonstrate competitive performance across nine benchmarks, showing robustness across datasets and architectures. Fig. 2(a) illustrates the core idea of our method: rather than using a naïve approach, such as simply recording the logit or training a probe classifier at the activations of the answer token, we first reinterpret the LLM as an autoregressive EBM via the chain rule of probabilities. We then further decompose each conditional probability, incorporating insights from Grathwohl et al. (2020). At the time step of the exact token i−1i-1, we extract the energy, which corresponds to the logit, and compare it with the marginal energy at the next time step i, corresponding to the denominator of the softmax. According to the chain rule, these two quantities should be identical; however, they differ in the LLM implementation—Fig. 2(b). We find that the discrepancy, which we term spilled energy ΔE(i:1) E_ θ(x_i:1), correlates strongly with instances where the LLM produces an incorrect output—see Fig. 2(c). Moreover, its detection signal separates well correct and incorrect classes across datasets, reflecting the model’s confidence, as shown in Fig. 2(d). Figure 2: How energy spills in LLMs. (a) Language Modeling p(i:1)p(x_i:1) is attained as a decomposition problem following the chain rule of probability, implemented as autoregressive: we recursively apply a discriminative classifier over the vocabulary V to attain generative modeling with larger context size i.e. p(i|i−1:1)p(x_i|x_i-1:1). (b) We reinterpret each discriminative classifier as a generative EBM, finding a connection between two quantities that should be the same across time steps yet are different. We call this difference “the spilled energy” ΔE(i:1) E_ θ(x_i:1) in Eq. 8. (c) Given that we simply read values inside the LLM, our approach is training-free and correlates well with hallucinations on a synthetic math dataset with increasing difficulty; (d) histograms of spilled energy values, for incorrect and correct answers on all nine datasets using min pooling for Llama-3-Instruct. The two distributions are easily separable by using a simple threshold, resulting in a generalization across real-world tasks. 2 Related Work EBM applications to Trustworthy AI. EBMs have been applied to improve the reliability and interpretability of Deep Nets. For example, Energy-Based Out-of-Distribution Detection (OOD) (Liu et al., 2020) uses the energy score as a more robust alternative to softmax confidence. At the same time, Grathwohl et al. (2020) presents how to reinterpret a discriminative classifier as EBM to train models that are both discriminative and generative. Following this work, Zhu et al. (2021) provides new insights into the role of energy when training EBMs and robust classifiers using adversarial training. Instead, Mirza et al. (2024; 2025) explain adversarial attacks by reinterpreting the softmax classifier as an EBM, showing that these perturbations correspond to shifts in the underlying energy landscape. Foundations of Hallucination in LLMs. LLMs are prone to diverse errors—including bias, reasoning failures, and generation of factually incorrect information unsupported by reliable sources. Karpowicz (2025) frames hallucination and imagination as mathematically identical phenomena, both emerging from a necessary violation of information conservation. Also Xu et al. (2025) provides a formal learning-theoretic proof that hallucinations are unavoidable. They define a formal world in which both the LLM and the ground-truth are computable functions, showing through classic results in computability theory, that no LLM can learn all such functions. As a consequence, hallucination is not just a practical artifact but a fundamental limitation of LLMs, valid even under idealized conditions. Recently, Kalai et al. (2025) showed that hallucinations come from the statistical problem of the pretraining methodology: minimizing the cross entropy naturally causes errors because it does not train the model to express uncertainty and say “I do not know.” Kalai et al. (2025) proposes changing the evaluation practices to not reward models for guessing, but rather to mimic the human exams that penalize only wrong answers. Detecting and Mitigating LLM Hallucinations. Orgad et al. (2025) train classifiers on the internal representations of the LLMs to predict, based on the features, the correctness of the answer. Given an LLM in a white-box setting, an input prompt, and the generated response y y, the classifier’s task is to predict whether y y is a hallucination. Orgad et al. suggested that LLMs may encode more factual knowledge in their latent subspaces than is revealed in their outputs. Gekhman et al. (2025) proposed a framework for studying hidden knowledge. Finally, Santilli et al. (2025) point out that uncertainty quantification in language models is often evaluated using metrics like AuROC. This shares biases between detection methods and correctness functions (e.g., length effects) that systematically distort results. One way to mitigate hallucinations is to act at the decoding stage, where the output generation can be steered Subramani et al. (2022). Steering vectors provide a straightforward way to control a model by adding a fixed vector to its activations (Dunefsky and Cohan, 2025). Fu et al. (2025) introduced DeepConf, a test-time method that leverages model-internal confidence signals to filter out low-quality reasoning traces during or after generation. Kuhn et al. (2023b); Fadeeva et al. (2024); Farquhar et al. (2024), and its follow-up by Kossen et al. (2025) in which they approximate the semantic entropy in a more efficient way. Constrained decoding approaches Li et al. (2023); Peng et al. (2023) modify token selection policies. Similarly, reinforcement learning with fact-based rewards Ouyang et al. (2022) has been used to bias decoding trajectories toward verifiable outcomes. Incorrect answers may also be given due to an ambiguous prompt: Kuhn et al. (2023a)’s CLAM framework uses few-shot prompts to classify a question’s ambiguity and then asks the user to clarify. 3 Background and Foundations 3.1 Energy-Based Models We give an overview of Energy-based Models (EBMs) and their use in discriminative classifiers. EBMs. Energy-Based Models are a class of probabilistic models in which the probability distribution over data points x is defined in terms of an energy function E()E_ θ(x). The energy function, parameterized by a neural network θ (Lecun et al., 2006), assigns a scalar energy to each configuration of x, where lower energy values correspond to higher likelihood. The resulting probability distribution is given by p()=exp(−E())Zp_ θ(x)= (-E_ θ(x))Z_ θ where Z_ θ denotes the partition function (normalizing constant), defined as Z=∑exp(−E())Z_ θ= _x (-E_ θ(x)) for discrete x, or equivalently Z=∫exp(−E())Z_ θ= (-E_ θ(x))\,dx for continuous x. Standard neural networks are often deterministic function approximators, mapping ↦yx y, EBMs instead define a full probability distribution over data or latent variables. One of the strengths of EBMs is their flexibility in modeling arbitrary distributions without being tied to a specific parametric form. This flexibility comes from the fact that the energy function E()E(x) can be defined in various ways. Training involves learning the parameters of the energy function such that the probability distribution p()p_ θ(x) matches the empirical distribution of the data. This is typically achieved using techniques like contrastive divergence, score matching, or maximum likelihood. Notation. Let V denote the vocabulary of an LLM, i.e., the set of all tokens that can be processed as input and generated at each decoding step, with size ||=V|V|=V. We shorten the sequence of tokens N,…,1\x_N,…,x_1\ as =N:1X=\x_N:1\, and i∈x_i denotes the token in the i-th position along the sequence. We model the LLM as a function :ℝN×V→ℝV θ:R^N× V ^V, implemented by a transformer, or any other sequence-to-sequence mechanism. For a sequence i:1\x_i:1\ as input, we write (i:1)[k] θ (x_i:1 )[k] to denote the predicted logit assigned to the k-th token class in V for the i+1i+1 token in the sequence, as is standard in autoregressive LLM training (Ouyang et al., 2022). 3.2 Autoregressive Large Language Models Generative modeling has been pursued through a variety of approaches beyond autoregression (AR). Variational Autoencoders (VAEs) (Kingma and Welling, 2014) learn a probabilistic latent variable model by encoding inputs into a latent space and decoding samples back to the data domain. Generative Adversarial Networks (GANs) (Goodfellow et al., 2014) frame generation as a min-max game between a generator and a discriminator. The diffusion process has been incorporated into neural nets (Sohl-Dickstein et al., 2015) and, more recently, Diffusion Models (Ho et al., 2020) have emerged as a powerful class of generative models. While these paradigms differ in how they approximate the data distribution, AR models are special in their kind and take a more direct route by factorizing the joint probability of sequences into conditionals, making them especially suitable for language modeling. We now focus on the AR formulation that underlies most LLMs. Textual data is segmented into a sequence of tokens =i,…,1X=\x_i,…,x_1\, and a language modeling objective is employed to maximize the likelihood of such data (Radford and Narasimhan, 2018). In other words, we model the joint probability of tokens in the sequence X, through a conditional probability parameterized by θ: p(i:1)=p(i|i−1:1)…p(2|1)p(1)=∏ip(i|i−1:1)⏟discriminative modelp(1).p(x_i:1)=p(x_i\ |\ x_i-1:1)… p(x_2\ |\ x_1)\ p(x_1)=Π _i p_ θ(x_i\ |\ x_i-1:1)_discriminative model\ p_ θ(x_1). (1) What we find interesting about this factorization is that, although it seeks to attain generative modeling, i.e., p(i:1)p(x_i:1), it actually uses recursively discriminative classifiers, parameterized by a transformer network θ, that predicts a discrete distribution of the next token ix_i over the vocabulary V, given previous tokens i−1:1x_i-1:1. This is used to model each conditional probability. 4 How Energy Spills in LLMs When predicting the token at position i, the conditional probability modeled by θ can be decomposed using the probabilities of the sequences. As a result, the marginal term from step i cancels out with the sequence probability from the decomposition at the previous step i−1i-1, which means we have: p(i:1)=∏ip(i|i−1:1)=∏ip(i:1)p(i−1:1)⟹…p(i:1)p(i−1:1)⏟step ip(i−1:1)⏞step i−1p(i−2:1)⋯=p(i:1).p(x_i:1)=Π _ip_ θ(x_i|x_i-1:1)=Π _i p_ θ(x_i:1)p_ θ(x_i-1:1) … p_ θ(x_i:1) p_ θ(x_i-1:1)_step $i$\ p_ θ(x_i-1:1)^step $i-1$p_ θ(x_i-2:1)…=p(x_i:1). (2) This indeed confirms that Eq. 1 results in the correct formulation for language modeling, which is p(i:1)p(x_i:1). Following the mathematics, these quantities should cancel out along the sequence, but we will now show that, in practice, this constraint is not explicitly optimized for, and we can exploit it for hallucination detection. 4.1 Interpreting LLMs as Energy-based models (EBMs) Let us continue the expansion from Eq. 2. Writing the conditional as the ratio between the joint distribution in the numerator and the marginal distribution in the denominator, we note that this ratio is actually implemented in LLMs as a softmax classifier that digests the embedding of the prior sentence i−1:1x_i-1:1 and predicts the next token ix_i; thus, this chain of equality holds true. We can then apply the “trick” from Grathwohl et al. (2020) as: p(i|i−1:1)=p(i:1)p(i−1:1)=exp(i−1:1)[id(i)]∑k=1Vexp(i−1:1)[k]whereid:0,1V↦[1,…,V]. p_ θ(x_i|x_i-1:1)= p_ θ(x_i:1)p_ θ(x_i-1:1)= θ(x_i-1:1) [ id(x_i) ]Σ _k=1^V θ(x_i-1:1)[k]\ where\ id:\ \0,1\^V [1,…,V]. (3) id is the map that takes as input a one-hot encoding vector ix_i for a word token at position i in the text and outputs its index in the vocabulary. A typical cross-entropy loss only optimizes with the supervision provided by the ground-truth token, through the vocabulary index id(i) id(x_i). This loss ignores all other quantities or constraints related to the complete sequence X, i.e., it ignores all the time steps higher than i+1i+1. We can write the conditional probability of Eq. 3 as a ratio of two EBMs as: logp(i|i−1:1)=logexp(−Eℓ(i:1))exp(−Em(i−1:1))Z~()Z()=−Eℓ(i:1)+Em(i−1:1). p_ θ(x_i|x_i-1:1)= (-E _ θ(x_i:1)) (-E^m_ θ(x_i-1:1)) Z( θ)Z( θ)=-E _ θ(x_i:1)+E^m_ θ(x_i-1:1). (4) Following Zhu et al. (2021), the partition functions simplify since logZ~()=logZ() Z( θ)= Z( θ)111For a formal proof, please see Section A.1.. Eℓ,EmE _ θ,~E^m_ θ are computed from the output of the model, but with two big differences: EℓE _ θ as a single logit extracted using the id of the sampled token, EmE^m_ θ by marginalizing over all ids in the vocabulary. The two energies can be derived from the softmax of the logits, by connecting Eq. 4 and Eq. 3: −logp(i|i−1:1) - p_ θ(x_i\ |\ x_i-1:1) =−log(exp((i−1:1)[id(i)])∑kexp((i−1:1)[k]))= =- ( ( θ(x_i-1:1)[ id(x_i)])Σ _k ( θ(x_i-1:1)[k]) )= (5) =−(i−1:1)[id(i)]⏟Eℓ(i:1)+log∑k=1Vexp(i−1:1)[k]⏟−Em(i−1:1) = - θ(x_i-1:1) [ id(x_i) ]_E _ θ(x_i:1)+ Σ _k=1^V θ(x_i-1:1)[k]_-E^m_ θ(x_i-1:1) (6) where (i−1:1) θ(x_i-1:1) produces the logits over the entire vocabulary V, and id(i) id(x_i) allows us to extract the logit of the sampled token at decoding step i. We can think of Eℓ(i:1)E _ θ(x_i:1) as the energy of the sampled tokens i:1\x_i:1\, and Em(i−1:1)E^m_ θ(x_i-1:1) as the energy E(i:1)E_ θ(x_i:1), marginalized over all possible ix_i. Considering the decoding at step i in Eq. 4, we get: Eℓ(i:1)=−(i−1:1)[id(i)],Em(i−1:1)=−log∑k=1Vexp(i−1:1)[k].E _ θ(x_i:1)=- θ(x_i-1:1)[ id(x_i)], E^m_ θ(x_i-1:1)=- Σ _k=1^V θ(x_i-1:1)[k]. (7) Using the chain rule and Eq. 6, we can write the negative log-likelihood in terms of energies as: −logp(N:1)=−log∏ip(i|i−1:1)=∑iEℓ(i:1)−Em(i−1:1)- p(x_N:1)=- Π _ip_ θ(x_i|x_i-1:1)= _iE _ θ(x_i:1)-E^m_ θ(x_i-1:1) without considering the base case p(1)p_ θ(x_1). Now, if we develop the above equation as done for Eq. 2, we write the total energy of a sequence of length N as E(N:1)E_ θ(x_N:1). Observe that the two energies, not interacting at the same step but at steps i and i−1i-1, should be equal, but they are measured in the LLM at different generation steps and from different components. E(N:1)=∑i=1N−1Eℓ(i+1:1)−Em(i:1)=…Eℓ(i+1:1)⏟ΔE(i:1)−Em(i:1)⏞timestepi+1+Eℓ(i:1)−Em(i−1:1)⏞timestepi…E_ θ(x_N:1)= _i=1^N-1E _ θ(x_i+1:1)-E^m_ θ(x_i:1)\;=\;…\; E _ θ(x_i+1:1)\ to0.0pt$ -E_ θ(x_i:1)+E_ θ(x_i:1)_ E_ θ(x_i:1)$ -E^m_ θ(x_i:1)^timestep~i+1+ E _ θ(x_i:1)-E^m_ θ(x_i-1:1)^timestep~i\;… At timestep i+1i+1, first −Em(i:1)-E^m_ θ(x_i:1) is measured, taking the denominator in the softmax as in the right part of Eq. 6, whereas at timestep i, the second Eℓ(i:1)E _ θ(x_i:1) is taken, reading the logit in the softmax, left part of Eq. 6. We thus define the discrepancy between the two quantities as spilled energy: Definition 4.1 (Spilled Energy ΔE(i:1) E_ θ(x_i:1)). The spilled energy in an LLM is the difference between two energies that, in principle, should be equal, but given that they are measured i) at different time steps i) in different components, could be different. ΔE(i:1)≜−Em(i:1)+Eℓ(i:1)=−log∑kexp((i:1)[k])⏟timestepi+1+(i−1:1)[id(i)]⏟timestepi E_ θ(x_i:1) -E^m_ θ(x_i:1)+E _ θ(x_i:1)= - Σ _k ( θ(x_i:1)[k])_timestep~i+1+ θ(x_i-1:1)[ id(x_i)]_timestep~i (8) Since both terms on the right side should be equal to E(i:1)E_ θ(x_i:1), delta values should always be zero when we are correctly modeling the energy at timestep i. A shorter explanation for why spilled energy needs to be zero is given in Section A.3. 4.2 Detecting hallucinations with spilled energy EBMs have previously been used to assess neural network credibility (Liu et al., 2020), and calibration for LLMs has been explored by the Anthropic team (Kadavath et al., 2022b). However, dominant training-free baselines such as logits or “p(true)p(true)” remain weak. We likewise adopt a training-free approach, but rely on Eq. 8 and its variants as discriminants. We feed the prompt i−1,…,1\x_i-1,…,x_1\ to the LLM θ and obtain the completion N,…,i\x_N,…,x_i\. Following Orgad et al. (2025), we focus on the “exact answer” tokens—those in [i+1,N][i+1,N] that contain the precise answer (e.g., Rome in Fig. 1), denoted [u,w]⊆[i+1,N][u,w] [i+1,N]. For instance, it would be the tokens associated with Rome in the question in Fig. 1. We identify this span by prompting the LLM for a brief answer. When the answer spans multiple tokens, we apply a pooling strategy, which we ablate in Section 5. We propose measuring two values that correlate well with hallucinations: 1. Marginal energy Em(i:1)E^m_ θ(x_i:1); 2. Spilled energy ΔE(i:1) E_ θ(x_i:1) by definition of Eq. 8. We also attempt to combine the two metrics into scaled spilled energy ΔEs E_s, where the spilled energy is multiplied by the absolute value of the marginal energy as ΔEs(i:1)=|Em(i:1)|ΔE(i:1) E_s(x_i:1)= E^m_ θ(x_i:1) E_ θ(x_i:1). The metrics proposed here are independent, new for LLMs, and can all be tested efficiently. These measures can be computed over the full sequence, but for error detection, as discussed in Table 3, we must extract the values in the localized exact interval [u,w][u,w] to avoid false positives. Note that Eℓ(i:1)E _ θ(x_i:1) is the classic baseline which in literature is referred to as “logits” or “logits confidence”. 5 Experiments To evaluate spilled energy, we consider two complementary settings. First, a controlled synthetic environment, where we generate both correct and incorrect multi-digit arithmetic solutions. Second, established real-world benchmarks, where errors arise naturally across diverse reasoning and comprehension tasks. Together, these experiments test whether insights from the clean synthetic setup transfer to the complexity of open-domain language understanding. 5.1 Spilled Energy under Synthetic Arithmetic Experimental Setting. We first evaluate spilled energy in a controlled setting: multi-digit arithmetic problems with more than 14 digits. For each instance, we generate both correct and incorrect solutions. We tested three different LLMs: Llama-3 8B (Dubey et al.), Qwen-3 8B (Qwen-Team), and Mistral-7B-Instruct v0.3 (Jiang et al.). Incorrect solutions are obtained by introducing random numerical errors of varying magnitude. Specifically, we define three error ranges that differ in their difficulty of detection: ⋄ Easy: random offset in the range [1000,10000][1000,10000], which are typically easier to identify. ⋄ Medium: random offset in the range [100,1000][100,1000], where detection requires closer inspection. ⋄ Hard: random offset in [1,10][1,10], much harder to detect since they appear plausible at first glance. This design allows us to systematically probe whether spilled energy can distinguish between correct and incorrect generations across different levels of error subtlety. Results. We observe that spilled energy values separate correct from incorrect solutions with high reliability across all error ranges and across all LLMs. In particular, spilled energy consistently assigns lower values to correct generations and higher values to incorrect ones, producing a clear margin of separation. Compared to standard baselines such as logits, spilled energy achieves superior discriminative power, especially for errors in the more challenging range [1,10][1,10], see Fig. 3. We offer more results in Fig. 5. Larger, better-detailed ROC and histograms are in Figs. 7 and 6 respectively. LLama-3-8B-Instruct (a) Easy (b) Medium (c) Hard (d) ROC (e) Easy (f) Medium (g) Hard (h) ROC Correct Incorrect Spilled Energy Easy Logit Energy Medium Marginal Energy Hard Qwen-3 8B Figure 3: Histograms of Spilled Energy values across models (rows) on Math Sums with different error ranges in the answer (columns, decreasing range left to right, making it harder to detect errors). All sums are performed on 13-digit integers. In the fourth column, we show ROC curves for Hallucination Detection across the error ranges (colors) and methods (line styles). (a) Results by Orgad et al. (b) Spilled Energy Improvement over Orgad et al. Figure 4: (a) AuROC performance as percentages of probing classifiers on exact answer tokens by Orgad et al. for LlaMA-3-Instruct. (b) depicts the performance difference between our Spilled ΔE E with Min pooling and theirs. Positive values indicate cases where Spilled ΔE E outperforms Orgad et al.. This comparison highlights the generalization capabilities of our method, compared to probing classifiers. Legend: low performance high performance. 5.2 Cross-dataset Results in Real-World Benchmarks Experimental Setting. We evaluate our methods on a diverse set of established NLP benchmarks, including Math (Hendrycks et al.), TriviaQA (Joshi et al.), HotpotQA (Yang et al.), Winogrande (Sakaguchi et al.), Winobias (Zhao et al.), Movies (Orgad et al.), MNLI (Williams et al.) and IMDB (Maas et al.). These datasets span a wide range of reasoning and error-detection tasks, allowing us to test whether the patterns observed in the synthetic arithmetic setting extend to real-world, open-domain scenarios. Here too, we evaluate multiple LLMs that are either instruction-aligned or not aligned, such as LLaMA-3 (Dubey et al.), and Mistral (Jiang et al.). As emphasized by Orgad et al., it is essential to first localize the tokens most relevant to the final answer before applying error detection. Since exact answer tokens may consist of multiple tokens, we further adopt a pooling strategy across the localized span to obtain a final score per sentence. We compare spilled and marginal energy against baselines such as the probing classifiers of Orgad et al., logit confidence of Varshney et al. and p(true)p(true) of Kadavath et al.. Table 1: Hallucination detection performance, in terms of AuROC, across nine benchmarks and four different LLMs. We measure the generalization across all tasks by computing the average. Pool HotpotQA HotpotQA-WC IMDB Math MNLI Movies TriviaQA Winobias Winogrande Average LLaMA-Instruct Dubey et al. (2024) p(true)p(true) — 58.31±0.32± 0.32 51.66±1.05± 1.05 50.72±1.20± 1.20 49.53±2.16± 2.16 52.33±0.98± 0.98 59.30±0.85± 0.85 45.99±0.51± 0.51 45.47±1.58± 1.58 48.33±0.68± 0.68 51.29±04.86± 04.86 Orgad et al. Mean 66.56±9.10± 9.10 59.00±8.14± 8.14 69.78±14.76± 14.76 66.56±17.04± 17.04 60.56±12.53± 12.53 66.44±8.06± 8.06 63.22±11.11± 11.11 67.33±11.97± 11.97 58.00±7.79± 7.79 64.16±03.90± 03.90 Logit EℓE Max 72.85±2.12± 2.12 91.11±1.52± 1.52 42.08±5.07± 5.07 57.81±3.82± 3.82 25.52±3.00± 3.00 43.97±1.38± 1.38 68.89±1.96± 1.96 39.95±2.41± 2.41 49.40±2.16± 2.16 54.62±18.97± 18.97 Marginal EmE^m Max 76.72±1.38± 1.38 30.74±3.45± 3.45 85.63±2.39± 2.39 27.08±5.06± 5.06 89.90±1.25± 1.25 96.17±0.63± 0.63 80.13±1.87± 1.87 57.67±2.94± 2.94 47.47±1.83± 1.83 65.72±24.39± 24.39 Marginal EmE^m Min 75.91±1.62± 1.62 97.57±0.75± 0.75 14.37±2.39± 2.39 70.55±2.43± 2.43 61.21±3.24± 3.24 72.21±1.60± 1.60 73.38±1.86± 1.86 47.19±2.71± 2.71 53.98±2.30± 2.30 62.93±21.89± 21.89 Spilled ΔEs E_s Max 53.65±1.40± 1.40 36.28±2.99± 2.99 55.80±4.32± 4.32 35.44±3.41± 3.41 58.81±2.58± 2.58 70.30±1.49± 1.49 48.70±2.44± 2.44 36.53±2.98± 2.98 44.32±1.70± 1.70 48.87±11.26± 11.26 Spilled ΔE E Min 85.98±1.09± 1.09 93.00±1.61± 1.61 47.66±4.06± 4.06 65.58±3.02± 3.02 73.95±1.97± 1.97 89.34±1.04± 1.04 87.07±1.33± 1.33 60.72±2.74± 2.74 55.11±2.05± 2.05 73.16±15.64± 15.64 LLaMA Dubey et al. (2024) p(true)p(true) — 52.83±0.71± 0.71 49.33±0.86± 0.86 52.30±0.58± 0.58 58.63±1.26± 1.26 53.78±0.70± 0.70 60.76±0.69± 0.69 62.94±0.51± 0.51 50.02±1.24± 1.24 53.47±0.54± 0.54 54.90±04.77± 04.77 Orgad et al. Mean 61.22±9.95± 9.95 56.78±8.70± 8.70 72.67±13.91± 13.91 69.67±15.07± 15.07 60.33±13.77± 13.77 64.00±8.40± 8.40 66.44±8.20± 8.20 60.89±12.60± 12.60 53.56±4.36± 4.36 62.84±05.71± 05.71 Logit EℓE Max 53.47±2.13± 2.13 49.02±1.79± 1.79 48.27±1.32± 1.32 57.38±6.09± 6.09 91.76±0.91± 0.91 57.42±1.43± 1.43 52.77±2.58± 2.58 50.74±1.51± 1.51 51.17±1.83± 1.83 56.89±12.70± 12.70 Marginal EmE^m Max 78.00±1.30± 1.30 76.90±1.09± 1.09 48.29±1.16± 1.16 68.77±8.33± 8.33 10.93±1.42± 1.42 80.70±1.98± 1.98 67.49±1.69± 1.69 51.91±2.32± 2.32 51.28±2.47± 2.47 59.36±20.69± 20.69 Marginal EmE^m Min 58.39±2.79± 2.79 59.20±1.95± 1.95 51.71±1.16± 1.16 34.13±8.78± 8.78 97.42±0.51± 0.51 50.37±2.43± 2.43 69.88±1.40± 1.40 49.05±2.20± 2.20 49.00±2.30± 2.30 57.68±16.75± 16.75 Spilled ΔEs E_s Min 77.75±1.52± 1.52 79.44±2.05± 2.05 43.39±1.82± 1.82 72.87±6.10± 6.10 99.97±0.08± 0.08 61.56±2.95± 2.95 77.55±1.62± 1.62 52.34±2.57± 2.57 48.17±1.62± 1.62 68.12±17.15± 17.15 Spilled ΔE E Min 79.04±1.78± 1.78 80.83±1.87± 1.87 43.22±1.67± 1.67 74.36±5.54± 5.54 99.97±0.08± 0.08 61.97±2.81± 2.81 78.54±1.57± 1.57 52.11±2.58± 2.58 48.21±1.62± 1.62 68.69±17.48± 17.48 Mistral-Instruct Jiang et al. (2023) p(true)p(true) — 56.67±0.80± 0.80 53.41±0.68± 0.68 48.84±0.78± 0.78 51.63±1.29± 1.29 54.93±0.53± 0.53 60.64±0.47± 0.47 63.59±0.57± 0.57 56.34±0.92± 0.92 56.92±0.57± 0.57 55.88±04.45± 04.45 Orgad et al. Mean 64.78±10.56± 10.56 56.78±7.95± 7.95 82.67±11.63± 11.63 68.78±11.43± 11.43 64.22±12.12± 12.12 64.89±11.55± 11.55 65.44±12.10± 12.10 61.00±12.23± 12.23 61.44±11.31± 11.31 65.56±06.84± 06.84 Logit EℓE Max 77.24±1.66± 1.66 83.84±1.66± 1.66 22.28±2.54± 2.54 57.67±3.29± 3.29 78.98±1.58± 1.58 76.89±1.49± 1.49 80.35±1.88± 1.88 45.53±2.60± 2.60 48.17±1.97± 1.97 63.44±19.99± 19.99 Marginal EmE^m Max 64.63±1.97± 1.97 33.42±1.90± 1.90 81.33±2.32± 2.32 26.52±2.28± 2.28 17.62±1.20± 1.20 86.60±1.20± 1.20 65.46±2.25± 2.25 56.41±4.44± 4.44 51.14±1.71± 1.71 53.68±22.53± 22.53 Marginal EmE^m Min 87.58±1.35± 1.35 97.94±0.62± 0.62 18.67±2.27± 2.27 67.58±3.37± 3.37 97.96±0.55± 0.55 84.90±1.37± 1.37 87.75±1.73± 1.73 49.19±3.97± 3.97 48.49±1.86± 1.86 71.12±25.68± 25.68 Spilled ΔEs E_s Max 49.13±2.50± 2.50 36.37±2.40± 2.40 46.45±2.56± 2.56 29.05±2.57± 2.57 53.79±1.55± 1.55 55.24±2.17± 2.17 46.73±1.98± 1.98 53.30±3.66± 3.66 51.20±1.84± 1.84 46.81±08.24± 08.24 Spilled ΔE E Min 91.12±1.10± 1.10 97.47±0.78± 0.78 59.77±2.57± 2.57 66.63±3.46± 3.46 95.95±0.83± 0.83 94.99±0.93± 0.93 91.75±1.01± 1.01 50.74±3.15± 3.15 49.00±1.92± 1.92 77.49±19.42± 19.42 Mistral Jiang et al. (2023) p(true)p(true) — 54.21±0.76± 0.76 51.68±0.76± 0.76 50.40±0.50± 0.50 45.86±2.05± 2.05 51.94±0.50± 0.50 49.12±0.63± 0.63 58.00±0.67± 0.67 53.76±1.17± 1.17 47.29±0.55± 0.55 51.36±03.73± 03.73 Orgad et al. Mean 61.78±9.27± 9.27 57.44±6.95± 6.95 76.22±12.82± 12.82 65.78±15.27± 15.27 56.67±11.83± 11.83 64.22±8.91± 8.91 64.33±10.40± 10.40 58.00±12.29± 12.29 54.56±4.36± 4.36 62.11±06.21± 06.21 Logit EℓE Max 49.54±1.42± 1.42 52.47±1.61± 1.61 32.72±2.89± 2.89 57.21±3.89± 3.89 92.49±1.15± 1.15 30.52±2.00± 2.00 39.73±2.03± 2.03 46.53±3.80± 3.80 44.41±2.42± 2.42 49.51±17.28± 17.28 Marginal EmE^m Max 83.57±1.13± 1.13 86.83±1.70± 1.70 45.31±2.49± 2.49 62.26±4.29± 4.29 96.03±0.83± 0.83 99.27±0.24± 0.24 92.26±1.31± 1.31 51.31±3.35± 3.35 54.49±2.48± 2.48 74.59±19.91± 19.91 Marginal EmE^m Min 87.52±1.31± 1.31 90.91±1.58± 1.58 54.69±2.49± 2.49 86.21±1.96± 1.96 98.80±0.35± 0.35 94.41±0.62± 0.62 83.66±2.16± 2.16 52.15±1.74± 1.74 46.37±2.02± 2.02 77.19±19.05± 19.05 Spilled ΔEs E_s Max 60.54±1.81± 1.81 60.18±1.84± 1.84 43.47±2.76± 2.76 71.93±3.62± 3.62 45.94±2.40± 2.40 78.84±1.53± 1.53 67.92±1.32± 1.32 57.24±3.72± 3.72 51.88±1.90± 1.90 59.77±11.08± 11.08 Spilled ΔE E Min 84.24±1.18± 1.18 83.74±1.41± 1.41 57.43±2.99± 2.99 78.26±2.93± 2.93 96.69±0.62± 0.62 84.47±1.17± 1.17 81.27±1.83± 1.83 50.62±1.72± 1.72 48.72±1.75± 1.75 73.94±16.18± 16.18 Ablation of the exact answer token. We provide an ablation experiment on the impact of selecting the exact answer tokens. Table 2 reports average AuROC over 9 benchmarks and 4 LLMs with the exact answer, along with another column that offers the improvement provided by using the exact answer. Like prior work, we confirm that searching for the exact answer provides a notable boost: the improvement is very pronounced (∼24% 24\%) for spilled and marginal energy, while the logit baseline receives a modest increase of 9%9\%. Cross-dataset results. We next evaluate in the more general setting of cross-dataset transfer, which better reflects real-world usage. For methods requiring training, we report the average performance on each dataset when trained separately on each of the other datasets (e.g., performance on IMDB is the average accuracy of classifiers trained on each of the other nine datasets). Fig. 4 shows a confusion matrix of cross-dataset performance, where the rows represent the training dataset and the columns represent the testing dataset, and where red indicates good performance and blue indicates low accuracy. The model tested is LlaMA-3-Instruct. Fig. 4(a) shows that probing classifiers, as soon as they go out-of-distribution from the training dataset, perform only marginally better than random guessing. The sharp drop observed in the off-diagonal elements supports our premise that this standard, in-distribution setup significantly overestimates the utility of trained probes for broad LLM deployment. Meanwhile, Fig. 4(b) displays the improvement of Spilled ΔE E over the probing classifier, where a positive red result means improvement of our method. Ours exhibits greater performance across most datasets without requiring training. The generalization is proved with a strong increment over the off-diagonal. Moreover, in some cases, such as TriviaQA, HotpotQA, and Movies, we have improvements even on the diagonal. Additional confusion matrices are available in Section D.3. Table 1 summarizes results across nine benchmarks. The result reported in each cell is the average of the accuracies of Fig. 4(a) within a column. Spilled energy consistently outperforms logit confidence, and substantially surpasses the probing classifiers of Orgad et al. (2025). While this latter performs well when trained and tested on the same dataset, their performance drops sharply under cross-dataset evaluation, as reflected in their higher standard deviations. By contrast, ours requires no training and generalizes robustly across diverse benchmarks. We observe that instruction-tuned models tend to amplify the margin by which spilled energy outperforms other methods, whereas on non-aligned Mistral, spilled energy may rank slightly behind marginal energy. We also compare pooling strategies and find that min pooling yields the best overall performance across methods. Table 3 shows our method generalizes to Gemma over different LLM size, 1B and 4B. Pool Average % Exact w/ exact answer answer increase Logit EℓE Max 56.12 +9.23 Orgad et al. Mean 63.67 – Marginal EmE^m Min 67.23 +20,02 Marginal EmE^m Max 63.34 +3,62 Spilled ΔE E Min 73.32 +24.06 Table 2: Improvements in AuROC with the exact answer. Average across 4 LLMs and 9 benchmarks. Impact of Instruction Tuning. Table 3: Hallucination detection performance on the Gemma Model Instruct for different parameters of the model, 1B and 4B. Pool IMBD Movies TriviaQA Winogrande Winobias MNLI Math HotpotQA HotpotQA-WC Average Gemma-Instruct 4B Kamath et al. (2025) Logit EℓE Max 50.09±0.45± 0.45 60.88±3.96± 3.96 53.95±2.10± 2.10 49.77±0.15± 0.15 54.43±2.80± 2.80 27.00±2.16± 2.16 78.64±3.47± 3.47 62.84±1.97± 1.97 64.49±2.02± 2.02 55.79±13.24± 13.24 Marginal EmE^m Max 49.14±2.70± 2.70 83.02±1.56± 1.56 84.14±1.39± 1.39 51.49±1.97± 1.97 47.97±1.80± 1.80 100.00±0.00± 0.00 74.57±3.60± 3.60 83.70±0.77± 0.77 85.95±2.03± 2.03 73.33±17.94± 17.94 Marginal EmE^m Min 50.86±2.70± 2.70 51.29±3.30± 3.30 55.33±1.80± 1.80 48.12±1.89± 1.89 51.91±2.10± 2.10 99.01±0.50± 0.50 76.03±3.27± 3.27 62.59±1.49± 1.49 71.84±2.72± 2.72 63.00±15.75± 15.75 Spilled ΔEs E_s Max 50.89±1.65± 1.65 50.77±5.72± 5.72 56.08±2.48± 2.48 50.59±1.72± 1.72 53.53±2.81± 2.81 95.61±0.56± 0.56 43.94±3.21± 3.21 50.87±1.87± 1.87 51.21±1.68± 1.68 55.94±14.35± 14.35 Spilled ΔE E Min 50.89±1.65± 1.65 86.13±4.28± 4.28 89.01±1.06± 1.06 50.18±1.97± 1.97 53.10±3.05± 3.05 99.66±0.21± 0.21 82.29±2.46± 2.46 89.10±1.75± 1.75 82.70±1.35± 1.35 75.89±17.98± 17.98 Gemma-Instruct 1B Kamath et al. (2025) Logit EℓE Max 46.33±0.82± 0.82 48.12±11.45± 11.45 58.89±1.61± 1.61 50.50±2.45± 2.45 53.49±3.71± 3.71 49.28±2.12± 2.12 65.12±6.62± 6.62 62.24±3.62± 3.62 75.67±1.96± 1.96 56.63±9.13± 9.13 Marginal EmE^m Max 45.42±1.78± 1.78 94.15±8.44± 8.44 83.66±1.82± 1.82 50.23±3.83± 3.83 49.93±1.56± 1.56 98.17±0.39± 0.39 64.21±6.67± 6.67 86.87±1.39± 1.39 82.33±1.27± 1.27 72.77±19.33± 19.33 Marginal EmE^m Min 54.58±1.78± 1.78 28.93±14.50± 14.50 39.80±2.54± 2.54 49.84±4.38± 4.38 50.39±1.80± 1.80 56.33±1.60± 1.60 63.20±4.27± 4.27 41.58±2.85± 2.85 61.56±1.61± 1.61 49.58±10.47± 10.47 Spilled ΔEs E_s Max 45.17±2.37± 2.37 33.27±11.49± 11.49 49.01±1.67± 1.67 52.27±3.56± 3.56 49.91±2.59± 2.59 77.48±1.92± 1.92 40.49±4.17± 4.17 49.18±3.93± 3.93 35.77±2.13± 2.13 48.06±12.13± 12.13 Spilled ΔE E Min 45.02±2.45± 2.45 82.82±12.91± 12.91 80.73±2.16± 2.16 52.48±3.75± 3.75 49.77±2.82± 2.82 92.93±1.79± 1.79 56.82±6.90± 6.90 85.64±2.23± 2.23 71.86±1.77± 1.77 68.67±16.84± 16.84 We observe a difference in the behavior in the base models and their instruction-tuned ones. While instruction-tuning generally improves generation quality, it can degrade the calibration of classical confidence metrics, as described in Huang et al. (2023a); Ho et al. (2025). For instance, examining the average performance in Table 1, the logit baseline EθℓE_θ decreases from 56.89% to 54.62% for LLaMA-3, indicating that fine-tuning may lead to overconfidence. In contrast, Spilled Energy (ΔEθ E_θ) consistently benefits from instruction tuning, showing improved detection rates across both LLaMA-3 (68.69% to 73.16%) and Mistral (73.94% to 77.49%). Variance and Generalization. A notable observation in Table 1 is the higher standard deviation associated with marginal and spilled energy compared to the probing classifiers in the average column. This variance is not a weakness but a reflection of the method’s training-free nature. Since ΔEθ E_θ relies on the intrinsic energy landscape of the LLM, its magnitude and sensitivity are naturally dependent on the specific domain (e.g., the sharp energy peaks in Math and HotpotQA versus the flatter distributions in Winobias and IMDB). Probing classifiers, by contrast, have high-variance when cross-testing yet the average of cross-testing results is mostly constant just above random chance (≈62−64%≈ 62-64\%). Limitations. A current limitation of spilled energy is that it sometimes produces false positives on tokens that are not semantically informative, as shown in Section D.2. We observe this effect most prominently on punctuation tokens (e.g., commas, periods) and on words at the beginning of sentences. In these cases, the probability mass over the next token is naturally spread across many plausible options, leading to inflated spilled energy values even in otherwise correct generations. This highlights the importance of accurately identifying the exact answer tokens, as detection is most reliable when restricted to the parts of the output that carry the semantic content of the answer. 6 Conclusion We reinterpreted the softmax layer of LLMs as an EBM, which lets us define spilled energy: the discrepancy between energy values that should be equal across consecutive time steps. We show theoretically and empirically that this discrepancy provides a strong, training-free signal for detecting hallucinations and errors in LLM outputs. Through synthetic arithmetic experiments, we demonstrate that spilled energy reliably separates correct from incorrect generations, outperforming baselines such as logits and marginal energy. Across diverse real-world NLP benchmarks, spilled energy generalizes robustly without requiring additional classifiers or task-specific training, unlike probing methods that struggle with transfer. Overall, spilled energy offers a principled and practical framework for error detection in LLMs and a new perspective on the internal energy dynamics of autoregressive models. Ethics Statement This work adheres to the ICLR Code of Ethics. Our study focuses on methodological contributions to error and hallucination detection in Large Language Models. We do not train new models or collect additional data; instead, we rely exclusively on publicly available datasets and widely used benchmark models for evaluation. We note that part of our evaluation includes the Math dataset, which was publicly accessible at the time of experimentation but has since been taken down following a copyright claim. We emphasize that this dataset was used solely for evaluation purposes of our method, and only prior to the date of the takedown. No redistribution of the dataset was made, and our reported results are limited to demonstrating methodological effectiveness. Our work does not involve personally identifiable information, sensitive content, or human subjects, and does not raise foreseeable risks of harm. We believe the proposed approach contributes positively to research on trustworthy AI by providing a training-free and generalizable framework for error detection in language models. Reproducibility Statement We are committed to ensuring the reproducibility of our results. All experimental details, including model configurations, evaluation protocols, and datasets used, are described in the main text and Appendix B. Upon acceptance of this work, we will publicly release the code implementing our method, along with instructions to reproduce all reported experiments. This will allow the community to verify our findings and build upon our work. Acknowledgment This work was supported by projects PNRR MUR PE0000013-FAIR under the MUR National Recovery and Resilience Plan funded by the European Union - NextGenerationEU, PRIN 2022 project 20227YET9B “AdVVent” CUP code B53D23012830006. It was also partially supported by Sapienza research projects D2QNeT and BEAT (Better dEep leArning securiTy) — bando per la ricerca di Ateneo 2024, and via the Seed of ERC grant “MINT.AI” (cup B83C25001040001). This work is additionally supported by the MUR FIS2 grant n. FIS-2023-00942 “NEXUS” (cup B53C25001030001). The work of Hazem Dewidar was carried out while he was enrolled in the Italian National Doctorate on Artificial Intelligence run by Sapienza University of Rome. Computing was supported by CINECA through the Italian SuperComputing Resource Allocation (ISCRA) projects Ge-Di HP10CRPUVC and SLEY HP10CX9CMC. References Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, et al. (2023) A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023. Cited by: §1. Y. Belinkov (2022) Probing classifiers: promises, shortcomings, and advances. Computational Linguistics 48 (1), p. 207–219. External Links: Link, Document Cited by: §1. E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell (2021) On the dangers of stochastic parrots: can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, p. 610–623. Cited by: §1. A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al. (2024) The LLaMa 3 herd of models. arXiv e-prints, p. arXiv–2407. Cited by: §5.1, §5.2, Table 1, Table 1. J. Dunefsky and A. Cohan (2025) One-shot optimized steering vectors mediate safety-relevant behaviors in LLMs. In Second Conference on Language Modeling, External Links: Link Cited by: §2. E. Fadeeva, A. Rubashevskii, A. Shelmanov, S. Petrakov, H. Li, H. Mubarak, E. Tsymbalov, G. Kuzmin, A. Panchenko, T. Baldwin, P. Nakov, and M. Panov (2024) Fact-checking the output of large language models via token-level uncertainty quantification. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 9367–9385. External Links: Link, Document Cited by: §2. S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal (2024) Detecting hallucinations in large language models using semantic entropy. Nature 630 (8017), p. 625–630. External Links: Document, Link Cited by: §2. Y. Fu, X. Wang, Y. Tian, and J. Zhao (2025) Deep think with confidence. External Links: 2508.15260, Link Cited by: §2. Z. Gekhman, E. Ben-David, H. Orgad, E. Ofek, Y. Belinkov, I. Szpektor, J. Herzig, and R. Reichart (2025) Inside-out: hidden factual knowledge in LLMs. In Second Conference on Language Modeling, External Links: Link Cited by: §2. I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio (2014) Generative adversarial nets. In NeurIPS, Cited by: §3.2. W. Grathwohl, K. Wang, J. Jacobsen, D. Duvenaud, M. Norouzi, and K. Swersky (2020) Your classifier is secretly an energy based model and you should treat it like one. In ICLR, Cited by: §1, §1, §2, §4.1. S. Harnad (2024) Language writ large: llms, chatgpt, grounding, meaning and understanding. arXiv preprint arXiv:2402.02243. Cited by: §1. D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021) Measuring mathematical problem solving with the math dataset. NeurIPS. Cited by: §5.2. J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. In NeurIPS, Cited by: §3.2. Z. Ho, S. Liang, and D. Tao (2025) Review of hallucination understanding in large language and vision models. External Links: 2510.00034, Link Cited by: §5.2. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu (2023a) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. CoRR abs/2311.05232. Cited by: §5.2. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. (2023b) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232. Cited by: §1. Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung (2023) Survey of hallucination in natural language generation. ACM Computing Surveys 55 (12), p. 1–38. Cited by: §1. A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed (2023) Mistral 7b. External Links: 2310.06825, Link Cited by: §5.1, §5.2, Table 1, Table 1. M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer (2017) TriviaQA: a large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1601–1611. Cited by: §5.2. S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, S. Johnston, S. El-Showk, A. Jones, N. Elhage, T. Hume, A. Chen, Y. Bai, S. Bowman, S. Fort, D. Ganguli, D. Hernandez, J. Jacobson, J. Kernion, S. Kravec, L. Lovitt, K. Ndousse, C. Olsson, S. Ringer, D. Amodei, T. Brown, J. Clark, N. Joseph, B. Mann, S. McCandlish, C. Olah, and J. Kaplan (2022a) Language models (mostly) know what they know. External Links: 2207.05221 Cited by: §5.2. S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al. (2022b) Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §1, §4.2. A. T. Kalai, O. Nachum, S. S. Vempala, and E. Zhang (2025) Why language models hallucinate. Technical report OpenAI and Georgia Tech. Note: Technical Report Cited by: §2. A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. György, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. Klimczak-Plucińska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. Põder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot (2025) Gemma 3 technical report. External Links: 2503.19786, Link Cited by: Table 3, Table 3. M. P. Karpowicz (2025) On the fundamental impossibility of hallucination control in large language models. External Links: 2506.06382, Link Cited by: §2. D. P. Kingma and M. Welling (2014) Auto-encoding variational bayes. In ICLR, Cited by: §3.2. J. Kossen, J. Han, M. Razzak, L. Schut, S. A. Malik, and Y. Gal (2025) Semantic entropy probes: robust and cheap hallucination detection in LLMs. External Links: Link Cited by: §2. L. Kuhn, Y. Gal, and S. Farquhar (2023a) CLAM: selective clarification for ambiguous questions with generative language models. External Links: 2212.07769, Link Cited by: §2. L. Kuhn, Y. Gal, and S. Farquhar (2023b) Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §2. Y. Lecun, S. Chopra, R. Hadsell, M. A. Ranzato, and F. J. Huang (2006) A tutorial on energy-based learning. In Predicting structured data, (English (US)). Cited by: §3.1. K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg (2024) Inference-time intervention: eliciting truthful answers from a language model. Advances in Neural Information Processing Systems 36. Cited by: §1. X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. Hashimoto, L. Zettlemoyer, and M. Lewis (2023) Contrastive decoding: open-ended text generation as optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 12286–12312. External Links: Link, Document Cited by: §2. A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. (2024) Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437. Cited by: §1. T. Liu, Y. Zhang, C. Brockett, Y. Mao, Z. Sui, W. Chen, and B. Dolan (2022) A token-level reference-free hallucination detection benchmark for free-form text generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 6723–6737. External Links: Link, Document Cited by: §1. W. Liu, X. Wang, J. D. Owens, and Y. Li (2020) Energy-based out-of-distribution detection. In NeurIPS, Cited by: §2, §4.2. A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts (2011) Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, Oregon, USA, p. 142–150. External Links: Link Cited by: §5.2. A. McGowan, Y. Gui, M. Dobbs, S. Shuster, M. Cotter, A. Selloni, M. Goodman, A. Srivastava, G. A. Cecchi, and C. M. Corcoran (2023) ChatGPT and bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Research 326, p. 115334. Cited by: §1. B. Millidge (2023) LLMs confabulate not hallucinate. Beren’s Blog. External Links: Link Cited by: §1. M. H. Mirza, M. R. Briglia, F. Bartolucci, S. Beadini, G. Lisanti, and I. Masi (2025) Understanding adversarial training with energy-based models. External Links: 2505.22486, Link Cited by: §2. M. H. Mirza, M. R. Briglia, S. Beadini, and I. Masi (2024) Shedding more light on robust classifiers under the lens of energy-based models. In ECCV, Cited by: §2. OpenAI-Team (2023) GPT-4 technical report. External Links: 2303.08774, Link Cited by: §1. H. Orgad, M. Toker, Z. Gekhman, R. Reichart, I. Szpektor, H. Kotek, and Y. Belinkov (2025) LLMs know more than they show: on the intrinsic representation of llm hallucinations. In ICLR, Cited by: §B.1, §B.1, §B.1, Appendix B, Figure 10, Figure 8, Figure 9, §D.3, Table 5, Table 5, Table 5, Table 5, §1, §2, §4.2, Figure 4, 4(a), 4(b), §5.2, §5.2, Table 1, Table 1, Table 1, Table 1, Table 2. L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. External Links: 2203.02155, Link Cited by: §2, §3.1. B. Peng, M. Galley, P. He, H. Cheng, Y. Xie, Y. Hu, Q. Huang, L. Liden, Z. Yu, W. Chen, and J. Gao (2023) Check your facts and try again: improving large language models with external knowledge and automated feedback. External Links: 2302.12813, Link Cited by: §2. Qwen-Team (2025) Qwen3: think deeper, act faster. Note: Accessed: 2025-09-23 External Links: Link Cited by: §5.1. A. Radford and K. Narasimhan (2018) Improving language understanding by generative pre-training. In OpenAI Technical Report, External Links: Link Cited by: §3.2. V. Rawte, S. Chakraborty, A. Pathak, A. Sarkar, S. T. I. Tonmoy, A. Chadha, A. Sheth, and A. Das (2023) The troubling emergence of hallucination in large language models - an extensive definition, quantification, and prescriptive remediations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 2541–2573. External Links: Link, Document Cited by: §1. K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi (2021) Winogrande: an adversarial winograd schema challenge at scale. Communications of the ACM 64 (9), p. 99–106. Cited by: §5.2. A. Salles, K. Evers, and M. Farisco (2020) Anthropomorphism in ai. AJOB neuroscience 11 (2), p. 88–95. Cited by: §1. A. Santilli, A. Golinski, M. Kirchhof, F. Danieli, A. Blaas, M. Xiong, L. Zappella, and S. Williamson (2025) Revisiting uncertainty quantification evaluation in language models: spurious interactions with response length bias results. In ACL, Cited by: §2. G. Serapio-García, M. Safdari, C. Crepy, L. Sun, S. Fitz, P. Romero, M. Abdulhai, A. Faust, and M. Matarić (2023) Personality traits in large language models. arXiv preprint arXiv:2307.00184. Cited by: §1. J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, Cited by: §3.2. N. Subramani, N. Suresh, and M. Peters (2022) Extracting latent steering vectors from pretrained language models. In Findings of the Association for Computational Linguistics: ACL 2022, S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 566–581. External Links: Link, Document Cited by: §2. N. Varshney, W. Yao, H. Zhang, J. Chen, and D. Yu (2023) A stitch in time saves nine: detecting and mitigating hallucinations of llms by validating low-confidence generation. External Links: 2307.03987 Cited by: §5.2. A. Williams, N. Nangia, and S. Bowman (2018) A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), p. 1112–1122. External Links: Link Cited by: §5.2. Z. Xu, S. Jain, and M. Kankanhalli (2025) Hallucination is inevitable: an innate limitation of large language models. External Links: 2401.11817, Link Cited by: §2. Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning (2018) HotpotQA: a dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, p. 2369–2380. Cited by: §5.2. Z. Yin, Q. Sun, Q. Guo, J. Wu, X. Qiu, and X. Huang (2023) Do large language models know what they don’t know?. In ACL, Cited by: §1. J. Zhao, T. Wang, M. Yatskar, V. Ordonez, and K. Chang (2018) Gender bias in coreference resolution: evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), M. Walker, H. Ji, and A. Stent (Eds.), New Orleans, Louisiana, p. 15–20. External Links: Link, Document Cited by: §5.2. Y. Zhu, J. Ma, J. Sun, Z. Chen, R. Jiang, Y. Chen, and Z. Li (2021) Towards understanding the generative capability of adversarially robust classifiers. In ICCV, p. 7708–7717. Cited by: §A.1, §2, §4.1. Appendix A Appendix A.1 Partition Functions Proof used in Eq. 4 We extend the proof of Zhu et al. to the sequence-to-sequence setting by treating next-token prediction as a multi-class classification problem. At step i, the input is the prefix i−1:1\x_i-1:1\, and the model outputs logits over the vocabulary V of size V. For notational consistency, we define the following energy terms: Eℓ(i:1)=−log(exp((i−1:1)[id(i)])),Em(i−1:1)=−log(∑k=1Vexp((i−1:1)[k])). \ array[]lE _ θ(x_i:1)=- \! ( ( θ(x_i-1:1)[ id(x_i)] ) ),\\[5.0pt] E^m_ θ(x_i-1:1)=- \! ( _k=1^V ( θ(x_i-1:1)[k] ) ). array . (9) The probability of the sequence up to position i can be expressed as p(i:1)=exp(−Eℓ(i:1))Z,p_ θ(x_i:1)= (-E _ θ(x_i:1))Z_ θ, (10) where Z_ θ is the global partition function (normalizing constant), defined over all possible continuations of all prefixes: Z=∑i−1:1∑iexp((i−1:1)[id(i)])=∑i−1:1∑k=1Vexp((i−1:1)[k]).Z_ θ= _x_i-1:1 _x_i \! ( θ(x_i-1:1)[ id(x_i)] )= _x_i-1:1 _k=1^V \! ( θ(x_i-1:1)[k] ). (11) Similarly, the probability of the prefix i−1:1x_i-1:1 can be written using the marginal energy: p(i−1:1)=exp(−Em(i−1:1))Z~,p_ θ(x_i-1:1)= (-E^m_ θ(x_i-1:1)) Z_ θ, (12) where Z~ Z_ θ is the corresponding normalizing constant: Z~=∑i−1:1exp(−Em(i−1:1))=∑i−1:1exp(log∑k=1Vexp((i−1:1)[k])). Z_ θ= _x_i-1:1 \! (-E^m_ θ(x_i-1:1) )= _x_i-1:1 \! ( _k=1^V \! ( θ(x_i-1:1)[k] ) ). (13) By expanding the logarithm in Eq. 13, we obtain Z~=∑i−1:1∑k=1Vexp((i−1:1)[k]), Z_ θ= _x_i-1:1 _k=1^V \! ( θ(x_i-1:1)[k] ), (14) which is identical to Eq. 11. Hence, the two partition functions coincide: Z=Z~.Z_ θ= Z_ θ. (15) A.2 The Role of Temperature in Spilled Energy We now analyze how the temperature parameter τ affects the definition of spilled energy. Starting from Eq. 3, the probability of the next token under temperature scaling is logpθ(i|i−1:1) p_θ(x_i\ |\ x_i-1:1) =logexp(1τ(i−1:1)[Id(i)])∑kexp(1τ(i−1:1)[k]) = \! ( 1τ θ(x_i-1:1)[Id(x_i)] )Σ _k \! ( 1τ θ(x_i-1:1)[k] ) (16) =1τ(i−1:1)[Id(i)]−log∑kexp(1τ(i−1:1)[k]). = 1τ\, θ(x_i-1:1)[Id(x_i)]- Σ _k \! ( 1τ θ(x_i-1:1)[k] ). (17) Accordingly, the spilled energy becomes ΔEθ(i:1)=1τ(i−1:1)[Id(i)]−log∑k=1|V|exp(1τ(i,…,1)[k]). E_θ(x_i:1)= 1τ\, θ(x_i-1:1)[Id(x_i)]- _k=1^|V| \! ( 1τ θ(x_i,…,x_1)[k] ). (18) Limit case τ→∞τ→∞. When the temperature tends to infinity, the logits are scaled down towards zero, making all tokens equally likely: limτ→+∞ΔEθ(i:1) _τ→+∞ E_θ(x_i:1) =limτ→∞1τ(i−1:1)[Id(i)]−log∑k=1|V|exp(1τ(i−1:1)[k]) = _τ→∞ 1τ θ(x_i-1:1)[Id(x_i)]- _k=1^|V| \! ( 1τ θ(x_i-1:1)[k] ) (19) =0−log∑k=1|V|exp(0) =0- _k=1^|V| (0) (20) =−log|V|. =- |V|. (21) Thus, for τ→∞τ→∞ the model degenerates into a uniform random classifier over the vocabulary. Interpretation. Varying τ perturbs the balance between the two energy terms, introducing a systematic error in ΔEθ E_θ. From the perspective of the Boltzmann distribution, scaling by 1τ 1τ corresponds to injecting or removing energy from the system. At high temperatures (τ→∞τ→∞), the system approaches maximum entropy, where all tokens have equal probability. At low temperatures (τ→0+τ→ 0^+), the distribution collapses onto the maximum logit token, making the model highly deterministic. Error accumulation. As we generate tokens sequentially, we accumulate deviations in ΔEθ E_θ: logpθ(i−1:1)=1τ(i−1:1)[Id(i)]−log∑kexp(1τ(i−1:1)[k])+∑j=1iΔEθ(j:1). p_θ(x_i-1:1)= 1τ θ(x_i-1:1)[Id(x_i)]- _k \! ( 1τ θ(x_i-1:1)[k] )+ _j=1^i E_θ(x_j:1). (22) Hence, temperature scaling not only modifies the probabilities but also reshapes the cumulative error landscape traced by spilled energy. A.3 Why Spilled Energy should be zero? TL;DR Consider Eq. (2) in our paper and the simplification that occurs between the two probabilities between step i and step i−1i-1: that simplification occurs because the probability in the denominator at step i is the same as the probability in the numerator at step i−1i-1 in order to perform language modeling correctly. We measure those inside and LLMs in terms of energy, and the spilled energy is the amount by which they differ. Please see the definition below. Let us assume a sequence of three tokens 2,1,0x_2,x_1,x_0. If we do language modeling with autoregression, minimizing the negative log-likelihood, we have: −logp(2,1,0)=−logp(2|1,0)⏟step 2p(1|0)p(0)- p(x_2,x_1,x_0)=- p(x_2|x_1,x_0)_step 2p(x_1|x_0)p(x_0) Now, every conditional probability on the right side is implemented with a transformer ending in a softmax discriminative classifier. Eq. 3 and Eq. 6 allow us to re-interpret: step 2:−logp(2|1,0)=−logp(2,1,0)p(1,0)=−log[exp(θ(1,0)[id(2)])∑kVexp(θ(1,0)[k])]= step 2: - p(x_2|x_1,x_0)=- p(x_2,x_1,x_0)p(x_1,x_0)=- [ (θ(x_1,x_0)[id(x_2)] ) _k^V (θ(x_1,x_0)[k] ) ]= (23) =Eℓ(2,1,0)−Em(1,0). =E (x_2,x_1,x_0)-E^m(x_1,x_0). (24) In other words, we reinterpret: ⋄ the numerator p(2,1,0)p(x_2,x_1,x_0) as the energy Eℓ(2,1,0)E (x_2,x_1,x_0), which is the logit (ℓ ) of the softmax at timestep 2; ⋄ The denominator as the energy Em(1,0)E^m(x_1,x_0) obtained with the marginalization (m) across the vocabulary V. This value can be read “read” simply by taking the denominator of the softmax at timestep 2. Please remember this term. It is better to indicate them as energies (since they are not probabilities), and given their logarithmic properties, we obtain a difference. We use the notation l for logits and m for marginalization. Now, when we go across steps and we connect two-time steps: step 1:−logp(1|0)=−logp(1,0)p(0)=Eℓ(1,0)−Em(0).step 1: - p(x_1|x_0)=- p(x_1,x_0)p(x_0)=E (x_1,x_0)-E^m(x_0). We see that at timestep 1, the value Eℓ(1,0)E (x_1,x_0) appears again, but measured at the logit level. In other words, across the time-steps 2 and 1, the quantity E(1,0)E(x_1,x_0) is measured twice: ⋄ at timestep 2, as the marginalization; ⋄ at timestep 1, as the logit. In the architecture or in the loss, there is no mechanism that forces these quantities to be the same, but they should be equal, given the language modeling objective. This is the same as saying that in Eq. 2, the probabilities across time steps need to cancel out as shown. In other words, the following: p(2,1,0)=p(2|1,0)p(1|0)p(0)p(x_2,x_1,x_0)=p(x_2|x_1,x_0)p(x_1|x_0)p(x_0) Implies: E(2,1,0)=Eℓ(2,1,0)−Em(1,0)+Eℓ(1,0)⏟should be zero−Em(0)+Eℓ(0)⏟should be zeroE(x_2,x_1,x_0)=E (x_2,x_1,x_0)~ -E^m(x_1,x_0)+E (x_1,x_0)_should be zero~ -E^m(x_0)+E (x_0)_should be zero To model the energy of a sequence Eℓ(2,1,0)E (x_2,x_1,x_0) correctly, then: ⋄ −Em(1,0)+Eℓ(1,0)=0-E^m(x_1,x_0)+E (x_1,x_0)=0 (spilled energy at timestep 2 if non-zero) ⋄ −Em(0)+Eℓ(0)=0-E^m(x_0)+E (x_0)=0 (spilled energy at timestep 1 if non-zero) so that E(2,1,0)=Eℓ(2,1,0)E(x_2,x_1,x_0)=E (x_2,x_1,x_0). Appendix B Reproducibility For comparisons on real-world tasks, we adopt the same experimental setting as Orgad et al. (2025), whose implementation is publicly available at https://github.com/technion-cs-nlp/LLMsKnow. This ensures that our baselines and evaluation procedures follow an established and validated protocol. In addition, we release our codebase, which includes: ⋄ computation of the proposed energy-based measures; ⋄ scripts for reproducing the synthetic arithmetic preliminary experiments; ⋄ example of how our method can be integrated into a benchmarking or production pipeline. B.1 Exact Answer Token Detection Details To analyze the spilled energy specifically on the tokens carrying the semantic weight of the answer, we must first localize the ”exact answer” span [u,w][u,w] within the longer generated sequence y y. We adopt the methodology proposed by Orgad et al. (2025), utilizing a combination of heuristics and an auxiliary instruction-tuned LLM to perform this extraction. Extraction Strategy Depending on the nature of the task, we employ two strategies to identify the exact answer substring s: ⋄ Heuristic Matching: For tasks with a closed set of possible labels (e.g., classification tasks or multiple-choice QA), we perform string matching to locate the label within the generation. ⋄ LLM-based Extraction: For open-ended generation tasks (e.g., TriviaQA, Math), where the answer form varies, we employ an instruction-tuned model (Mistral-7B-Instruct) to extract the short answer from the long-form generation. Prompting for Extraction Following Orgad et al. (2025), we prompt the auxiliary model with the original question q and the generated long answer y y using the following template: Prompt for Exact Answer Extraction Extract from the following long answer the short answer, only the relevant tokens. If the long answer does not answer the question, output NO ANSWER. Q: [Question 1] A: [LLM long answer 1] Exact answer: [Short exact answer 1] Q: [Question 2] A: [LLM long answer that does not answer the question] Exact answer: NO ANSWER Q: [Question] A: [LLM long answer] Exact answer: Verification and Token Mapping To ensure robustness, we verify that the extracted string s is a valid substring of the original generation y y. If the extraction is invalid or the model outputs ”NO ANSWER,” we retry the extraction up to five times. If a valid substring is still not found, the sample is excluded from the analysis to avoid identifying incorrect tokens. Once the substring s is validated, we map it to the corresponding token indices [u,w][u,w] in the original sequence. The spilled energy analysis is then performed specifically over this interval, or pooled across it (e.g., via min-pooling) as described in Section 5.2. Table 4: Answer Extraction Success Rate across tasks for Mistral-Instruct. Dataset Success Rate (%) TriviaQA 90.29 HotpotQA 87.37 Movies 93.61 MNLI 92.99 Math 87.59 HotpotQA-WC 92.38 Answer Extraction Performance For answer localization, we achieve accuracy comparable to the results of Orgad et al. (2025). We report in Table 4 the extraction success rate across the full datasets using Mistral-7B-Instruct. Note that some datasets have been excluded (e.g., IMDB, Winobias, Winogrande) since they have a finite set of possible answers that can be used to easily locate the exact answer within the model’s generation. Appendix C LLM Usage Large language models were used exclusively for text polishing and minor exposition refinements. All substantive research content, methodology, and scientific conclusions were developed entirely by the authors. Appendix D Supplementary Material This supplementary material is intended to complement the main paper by providing further motivation for our assumptions and design choices, as well as additional ablation studies or plots, such as ROCs and histograms, that could not fit in the main paper. D.1 Additional results for Synthetic Arithmetic In Fig. 5 we augmented Fig. 3 in the main paper, also adding the results for Mistral-7B-Instruct v0.3 and LLaMa-3-8B. The same findings of the figure in the paper also translate to this LLM, meaning that our method generalizes across LLMs. Fig. 6 and Fig. 7 also extend and provide more details of Fig. 3 in the main paper by showing, respectively, the histograms and the ROC at a better resolution and displayed in different frames. Also, we have added results for Mistral-7B-Instruct v0.3 and LLaMa-3-8B. D.2 Additional Qualitative Results In this section, we offer additional results of the detection performance following what is shown in Fig. 1. We report both success cases and failure cases. While it is difficult to draw conclusions and predict when, why, and on which topics spilled energy may work or not, we noticed that it appears to perform reliably on knowledge-based factual content but, at times, exhibits difficulties with reasoning tasks and numerical information, despite working well on math questions, as demonstrated in Section 5.1. Further investigation is required to better understand and validate these patterns. Mistral-7B-Instruct v0.3 (a) Easy (b) Medium (c) Hard (d) ROC (e) Easy (f) Medium (g) Hard (h) ROC Correct Incorrect Spilled Energy Easy Logit Energy Medium Marginal Energy Hard Llama-3 8B Figure 5: Histograms of Spilled Energy values across models (rows) on Math Sums with different error ranges in the answer (columns, decreasing range left to right, making it harder to detect errors), as described in Section 5.1. In the fourth column, we show ROC curves for Hallucination Detection across the error ranges (colors) and methods (line styles). Llama-3 8B (a) Easy (b) Medium (c) Hard (d) Easy (e) Medium (f) Hard (g) Easy (h) Medium (i) Hard (j) Easy (k) Medium (l) Hard Llama-3-8B-Instruct Qwen3-8B Mistral-7B-Instruct v0.3 Figure 6: Histograms of Spilled Energy values for Correct and Incorrect answers across models on Math Sums, increasing difficulty from left to right. We compute sums on 13-digit integers, for incorrect answers we add a random offset sampled uniformly from the error interval: Easy ∼(1e3,1e4) (1e3,1e4) - Medium ∼(1e2,1e3) (1e2,1e3) - Hard ∼(1,10) (1,10); for more details see Section 5.1. Llama-3-8B (a) Easy (b) Medium (c) Hard (d) Easy (e) Medium (f) Hard (g) Easy (h) Medium (i) Hard (j) Easy (k) Medium (l) Hard Llama-3-8B-Instruct Qwen-3 8B Mistral-7B-Instruct v0.3 Figure 7: ROC curves for Hallucination Detection across models (rows) on Math Sums with different error ranges in the answer (columns, decreasing range left to right). All sums are performed on 13-digit integers. Legend: Spilled (ours) Spilled ΔE E Logit EℓE Marginal EmE^m Pool HotpotQA HotpotQA-WC IMDB Math MNLI Movies TriviaQA Winobias Winogrande Average LLaMA-Instruct Orgad et al. (2025) Mean 66.56±9.10± 9.10 59.00±8.14± 8.14 69.78±14.76± 14.76 66.56±17.04± 17.04 60.56±12.53± 12.53 66.44±8.06± 8.06 63.22±11.11± 11.11 67.33±11.97± 11.97 58.00±7.79± 7.79 64.16±3.90± 3.90 Spilled ΔE E Min 85.98±1.09± 1.09 93.00±1.61± 1.61 47.66±4.06± 4.06 65.58±3.02± 3.02 73.95±1.97± 1.97 89.34±1.04± 1.04 87.07±1.33± 1.33 60.72±2.74± 2.74 55.11±2.05± 2.05 73.16±15.64± 15.64 Marginal EmE^m Max 76.72±1.38± 1.38 30.74±3.45± 3.45 85.63±2.39± 2.39 27.08±5.06± 5.06 89.90±1.25± 1.25 96.17±0.63± 0.63 80.13±1.87± 1.87 57.67±2.94± 2.94 47.47±1.83± 1.83 65.72±24.39± 24.39 Marginal EmE^m Min 75.91±1.62± 1.62 97.57±0.75± 0.75 14.37±2.39± 2.39 70.55±2.43± 2.43 61.21±3.24± 3.24 72.21±1.60± 1.60 73.38±1.86± 1.86 47.19±2.71± 2.71 53.98±2.30± 2.30 62.93±21.89± 21.89 Logit EℓE Max 72.85±2.12± 2.12 91.11±1.52± 1.52 42.08±5.07± 5.07 57.81±3.82± 3.82 25.52±3.00± 3.00 43.97±1.38± 1.38 68.89±1.96± 1.96 39.95±2.41± 2.41 49.40±2.16± 2.16 54.62±18.97± 18.97 Spilled ΔE E Max 54.34±1.58± 1.58 47.68±2.81± 2.81 52.34±4.06± 4.06 40.33±3.05± 3.05 56.44±2.81± 2.81 68.56±1.87± 1.87 47.54±2.40± 2.40 38.40±2.61± 2.61 44.97±1.51± 1.51 50.07±8.66± 8.66 LLaMA Orgad et al. (2025) Mean 61.22±9.95± 9.95 56.78±8.70± 8.70 72.67±13.91± 13.91 69.67±15.07± 15.07 60.33±13.77± 13.77 64.00±8.40± 8.40 66.44±8.20± 8.20 60.89±12.60± 12.60 53.56±4.36± 4.36 62.84±5.71± 5.71 Logit EℓE Min 87.93±1.01± 1.01 91.24±0.80± 0.80 51.73±1.32± 1.32 42.99±5.68± 5.68 97.01±0.43± 0.43 99.86±0.16± 0.16 84.53±0.87± 0.87 49.29±1.46± 1.46 48.52±1.78± 1.78 72.57±22.36± 22.36 Spilled ΔE E Min 79.04±1.78± 1.78 80.83±1.87± 1.87 43.22±1.67± 1.67 74.36±5.54± 5.54 99.97±0.08± 0.08 61.97±2.81± 2.81 78.54±1.57± 1.57 52.11±2.58± 2.58 48.21±1.62± 1.62 68.69±17.48± 17.48 Spilled ΔEs E_s Min 77.75±1.52± 1.52 79.44±2.05± 2.05 43.39±1.82± 1.82 72.87±6.10± 6.10 99.97±0.08± 0.08 61.56±2.95± 2.95 77.55±1.62± 1.62 52.34±2.57± 2.57 48.17±1.62± 1.62 68.12±17.15± 17.15 Marginal EmE^m Max 78.00±1.30± 1.30 76.90±1.09± 1.09 48.29±1.16± 1.16 68.77±8.33± 8.33 10.93±1.42± 1.42 80.70±1.98± 1.98 67.49±1.69± 1.69 51.91±2.32± 2.32 51.28±2.47± 2.47 59.36±20.69± 20.69 Marginal EmE^m Min 58.39±2.79± 2.79 59.20±1.95± 1.95 51.71±1.16± 1.16 34.13±8.78± 8.78 97.42±0.51± 0.51 50.37±2.43± 2.43 69.88±1.40± 1.40 49.05±2.20± 2.20 49.00±2.30± 2.30 57.68±16.75± 16.75 Logit EℓE Max 53.47±2.13± 2.13 49.02±1.79± 1.79 48.27±1.32± 1.32 57.38±6.09± 6.09 91.76±0.91± 0.91 57.42±1.43± 1.43 52.77±2.58± 2.58 50.74±1.51± 1.51 51.17±1.83± 1.83 56.89±12.70± 12.70 Logit EℓE ALT 43.56±1.95± 1.95 39.74±1.73± 1.73 48.27±1.32± 1.32 57.41±6.06± 6.06 91.71±0.94± 0.94 43.11±1.57± 1.57 43.62±2.57± 2.57 50.74±1.51± 1.51 51.17±1.83± 1.83 52.15±14.88± 14.88 Logit EℓE Last Token 43.56±1.95± 1.95 39.74±1.73± 1.73 48.27±1.32± 1.32 57.41±6.06± 6.06 91.71±0.94± 0.94 43.11±1.57± 1.57 43.62±2.57± 2.57 50.74±1.51± 1.51 51.17±1.83± 1.83 52.15±14.88± 14.88 Marginal EmE^m ALT 61.59±1.88± 1.88 58.64±1.60± 1.60 48.29±1.16± 1.16 67.93±9.32± 9.32 10.75±1.44± 1.44 61.39±1.80± 1.80 49.73±1.45± 1.45 51.19±2.59± 2.59 51.44±2.50± 2.50 51.22±15.61± 15.61 Marginal EmE^m Last Token 61.59±1.88± 1.88 58.64±1.60± 1.60 48.29±1.16± 1.16 67.93±9.32± 9.32 10.75±1.44± 1.44 61.39±1.80± 1.80 49.73±1.45± 1.45 51.19±2.59± 2.59 51.44±2.50± 2.50 51.22±15.61± 15.61 Marginal EmE^m Mean 58.27±2.50± 2.50 58.64±1.58± 1.58 48.29±1.16± 1.16 68.32±8.35± 8.35 6.12±0.70± 0.70 66.55±3.22± 3.22 45.67±1.38± 1.38 51.80±2.29± 2.29 51.29±2.46± 2.46 50.55±17.33± 17.33 Mistral-Instruct Orgad et al. (2025) Mean 64.78±10.56± 10.56 56.78±7.95± 7.95 82.67±11.63± 11.63 68.78±11.43± 11.43 64.22±12.12± 12.12 64.89±11.55± 11.55 65.44±12.10± 12.10 61.00±12.23± 12.23 61.44±11.31± 11.31 65.56±6.84± 6.84 Spilled ΔE E Min 91.12±1.10± 1.10 97.47±0.78± 0.78 59.77±2.57± 2.57 66.63±3.46± 3.46 95.95±0.83± 0.83 94.99±0.93± 0.93 91.75±1.01± 1.01 50.74±3.15± 3.15 49.00±1.92± 1.92 77.49±19.42± 19.42 Marginal EmE^m Min 87.58±1.35± 1.35 97.94±0.62± 0.62 18.67±2.27± 2.27 67.58±3.37± 3.37 97.96±0.55± 0.55 84.90±1.37± 1.37 87.75±1.73± 1.73 49.19±3.97± 3.97 48.49±1.86± 1.86 71.12±25.68± 25.68 Logit EℓE Max 77.24±1.66± 1.66 83.84±1.66± 1.66 22.28±2.54± 2.54 57.67±3.29± 3.29 78.98±1.58± 1.58 76.89±1.49± 1.49 80.35±1.88± 1.88 45.53±2.60± 2.60 48.17±1.97± 1.97 63.44±19.99± 19.99 Marginal EmE^m Max 64.63±1.97± 1.97 33.42±1.90± 1.90 81.33±2.32± 2.32 26.52±2.28± 2.28 17.62±1.20± 1.20 86.60±1.20± 1.20 65.46±2.25± 2.25 56.41±4.44± 4.44 51.14±1.71± 1.71 53.68±22.53± 22.53 Logit EℓE Last Token 55.77±2.38± 2.38 71.26±2.28± 2.28 22.28±2.54± 2.54 71.21±2.42± 2.42 47.78±2.26± 2.26 42.93±1.96± 1.96 58.36±3.52± 3.52 45.65±2.94± 2.94 48.30±2.04± 2.04 51.50±14.26± 14.26 Logit EℓE ALT 55.77±2.38± 2.38 71.26±2.28± 2.28 22.28±2.54± 2.54 71.21±2.42± 2.42 47.78±2.26± 2.26 42.93±1.96± 1.96 58.36±3.52± 3.52 45.65±2.94± 2.94 48.30±2.04± 2.04 51.50±14.26± 14.26 Mistral Orgad et al. (2025) Mean 61.78±9.27± 9.27 57.44±6.95± 6.95 76.22±12.82± 12.82 65.78±15.27± 15.27 56.67±11.83± 11.83 64.22±8.91± 8.91 64.33±10.40± 10.40 58.00±12.29± 12.29 54.56±4.36± 4.36 62.11±6.21± 6.21 Marginal EmE^m Min 87.52±1.31± 1.31 90.91±1.58± 1.58 54.69±2.49± 2.49 86.21±1.96± 1.96 98.80±0.35± 0.35 94.41±0.62± 0.62 83.66±2.16± 2.16 52.15±1.74± 1.74 46.37±2.02± 2.02 77.19±19.05± 19.05 Marginal EmE^m Max 83.57±1.13± 1.13 86.83±1.70± 1.70 45.31±2.49± 2.49 62.26±4.29± 4.29 96.03±0.83± 0.83 99.27±0.24± 0.24 92.26±1.31± 1.31 51.31±3.35± 3.35 54.49±2.48± 2.48 74.59±19.91± 19.91 Spilled ΔE E Min 84.24±1.18± 1.18 83.74±1.41± 1.41 57.43±2.99± 2.99 78.26±2.93± 2.93 96.69±0.62± 0.62 84.47±1.17± 1.17 81.27±1.83± 1.83 50.62±1.72± 1.72 48.72±1.75± 1.75 73.94±16.18± 16.18 Spilled ΔE E Max 61.50±1.88± 1.88 63.60±1.68± 1.68 42.57±2.99± 2.99 76.27±3.42± 3.42 47.01±2.48± 2.48 81.84±1.60± 1.60 68.07±1.30± 1.30 58.71±3.69± 3.69 51.13±1.87± 1.87 61.19±12.30± 12.30 Spilled ΔEs E_s Max 60.54±1.81± 1.81 60.18±1.84± 1.84 43.47±2.76± 2.76 71.93±3.62± 3.62 45.94±2.40± 2.40 78.84±1.53± 1.53 67.92±1.32± 1.32 57.24±3.72± 3.72 51.88±1.90± 1.90 59.77±11.08± 11.08 Table 5: Hallucination detection performance, in terms of AuROC, across nine benchmarks and different LLMs. We measure the generalization across all tasks by computing the average. (a) (b) Figure 8: Fig. 8(a) presents the cross-dataset performance of the method proposed by Orgad et al. (2025) using Llama-3. Fig. 8(b) depicts the performance difference between their method and our Spilled ΔE E with Min pooling. Positive values indicate cases where Spilled ΔE E outperforms the method of Orgad et al. (2025). All numbers are computed as percentages. (a) (b) Figure 9: Fig. 9(a) presents the cross-dataset performance of the method proposed by Orgad et al. (2025) using Mistral. Fig. 9(b) depicts the performance difference between their method and our Spilled ΔE E with Min pooling. Positive values indicate cases where Spilled ΔE E outperforms the method of Orgad et al. (2025). All numbers are computed as percentages. (a) (b) Figure 10: Fig. 10(a) presents the cross-dataset performance of the method proposed by Orgad et al. (2025) using Mistral-Instruct. Fig. 10(b) depicts the performance difference between their method and our Spilled ΔE E with Min pooling. Positive values indicate cases where Spilled ΔE E outperforms the method of Orgad et al. (2025). All numbers are computed as percentages. D.3 Additional results for Cross-testing with Real World Benchmarks Table 5 shows how our method compares with the baseline methods, Orgad et al. (2025) and Logit EℓE . This table was obtained by using various pooling methods in the pooling window from which we measure possible hallucinations. More details below based on the example in Fig. 11: ⋄ Min: minimum energy value in the pooling window. Energy Measured: −3-3 ⋄ Max: maximum energy value in the pooling window. Energy Measured: 1111 ⋄ Mean: mean of all the energies in the pooling window. Energy Measured: 2.082.08 ⋄ Last Token: energy of the last token in the pooling window. Energy Measured: −3-3 ⋄ After Last Token: energy of the first token after the pooling method. Energy Measured: 11 Figure 11: Example of the Pooling Window D.3.1 Success Cases Question: ‘Which planet is known as the Red Planet?’ Logits: The Red Planet is Mars . ✓ Ours: The Red Planet is Mars . ✓ Logits: The Red Planet is Jupiter . ✗ Ours: The Red Planet is Jupiter . ✗ Question: ‘What is the largest mammal in the world?’ Logits: The largest mamm al in the world is the Blue Whale ✓ Ours: The largest mamm al in the world is the Blue Whale ✓ Logits: The largest mamm al in the world is the House Cat . ✗ Ours: The largest mamm al in the world is the House Cat . ✗ Question: ‘Who painted the Mona Lisa?’ Logits: The Mona Lisa was painted by Leonardo da Vinci . ✓ Ours: The Mona Lisa was painted by Leonardo da Vinci . ✓ Logits: The Mona Lisa was painted by Pablo Esc obar . ✗ Ours: The Mona Lisa was painted by Pablo Esc obar . ✗ Question: ‘What gas do plants breathe in for photosynthesis?’ Logits: They breathe in carbon dioxide ✓ Ours: They breathe in carbon dioxide ✓ Logits: They breathe in oxygen ✗ Ours: They breathe in oxygen ✗ Question: ‘In which continent is Egypt Located?’ Logits: Egypt is located in Africa ✓ Ours: Egypt is located in Africa ✓ Logits: Egypt is located in Europe ✗ Ours: Egypt is located in Europe ✗ Question: ‘What is the fastest land animal?’ Logits: The fastest land animal is the che et ah ✓ Ours: The fastest land animal is the che et ah ✓ Logits: The fastest land animal is the lion ✗ Ours: The fastest land animal is the lion ✗ Question: ‘What is the hardest natural substance on Earth?’ Logits: The hardest natural substance is diamond ✓ Ours: The hardest natural substance is diamond ✓ Logits: The hardest natural substance is gold ✗ Ours: The hardest natural substance is gold ✗ Question: ‘Which ocean is the largest?’ Logits: The largest ocean is the Pacific Ocean ✓ Ours: The largest ocean is the Pacific Ocean ✓ Logits: The largest ocean is the Indian Ocean ✗ Ours: The largest ocean is the Indian Ocean ✗ D.3.2 Failure Cases Question: ‘Who was the first person to walk on the moon?’ Logits: Neil Armstrong ✓ Ours: Neil Armstrong ✓ Logits: Buzz Ald rin ✗ Ours: Buzz Ald rin ✗ Reasoning: ‘If there are 3 cars and each car has 4 wheels , how many wheels are there in total? ’ Logits: Each car has 4 wheels . So , for 3 cars , the total number of wheels is 3 x 4 = 12 wheels . ✓ Ours: Each car has 4 wheels . So , for 3 cars , the total number of wheels is 3 x 4 = 12 wheels . ✓ Logits: Each car has 8 wheels . So , for 3 cars , the total number of wheels is 3 x 8 = 14 wheels . ✗ Ours: Each car has 8 wheels . So , for 3 cars , the total number of wheels is 3 x 8 = 14 wheels . ✗ Reasoning: ‘What is the square root of 64?’ Logits: The square root of 64 is 8 ✓ Ours: The square root of 64 is 8 ✓ Logits: The square root of 64 is 10 ✗ Ours: The square root of 64 is 10 ✗ Question: ‘What blood type is known as the universal donor?’ Logits: O negative ✓ Ours: O negative ✓ Logits: AB positive ✗ Ours: AB positive ✗