Paper deep dive
Probing for Knowledge Attribution in Large Language Models
Ivo Brink, Alexander Boer, Dennis Ulmer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 9:53:15 AM
Summary
This paper introduces AttriWiki, a self-supervised pipeline for generating labeled training data to probe Large Language Models (LLMs) for knowledge attribution. The study distinguishes between parametric knowledge (internal memory) and contextual knowledge (provided in prompts). Using hidden representations from models like Llama-3.1-8B, Mistral-7B, and Qwen-7B, the authors train linear probes to classify the dominant knowledge source. Results show high accuracy (up to 0.96 Macro-F1) and successful generalization to external benchmarks like SQuAD and WebQuestions. The study also highlights that attribution mismatches significantly increase error rates, though correct attribution does not guarantee factual correctness.
Entities (10)
Relation Signals (9)
AttriWiki â generatesdatafor â Knowledge Attribution Probing
confidence 95% · We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data... to study contributive attribution
AttriWiki â enablesclassificationof â Parametric Knowledge
confidence 92% · We then train lightweight probing classifiers on these representations to distinguish contextual from parametric retrieval
AttriWiki â enablesclassificationof â Contextual Knowledge
confidence 92% · We then train lightweight probing classifiers on these representations to distinguish contextual from parametric retrieval
Linear Probe â achieveshighaccuracyon â LLaMA-3.1-8B
confidence 90% · Probes trained on AttriWiki achieve up to 0.96 Macro-F1 on Llama-3.1-8B
Linear Probe â achieveshighaccuracyon â Mistral-7B
confidence 90% · Probes trained on AttriWiki achieve up to 0.96 Macro-F1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B
Linear Probe â achieveshighaccuracyon â Qwen-7B
confidence 90% · Probes trained on AttriWiki achieve up to 0.96 Macro-F1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B
Attribution Mismatch â causes â Increased Error Rates
confidence 88% · attribution mismatches raise error rates by up to 70%
AttriWiki â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on knowledge conflicts. Probes trained on AttriWiki achieve up to 0.96 Macro-$F_1$ on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transfer to SQuAD and WebQuestions with 0.94-0.99 Macro-$F_1$, and generalise zero-shot to Tighidet et al. (2024)'s benchmark, outperforming their probe on conflicting settings without retraining. Furthermore, attribution mismatches raise error rates by up to 70%, though correct attribution does not guarantee correct answers, pointing to the need for broader detection frameworks.
Tags
Links
- Source: https://arxiv.org/abs/2602.22787v2
- Canonical: https://arxiv.org/abs/2602.22787v2
Trouble viewing inline? Open PDF directly â
Full Text
71,928 characters extracted from source content.
Expand or collapse full text
Probing for Knowledge Attribution in Large Language Models Ivo Brink KPMG NL / University of Amsterdam Amsterdam, The Netherlands ivobrink00@gmail.com &Alexander Boer KPMG NL Amsterdam, The Netherlands boer.alexander@kpmg.nl &Dennis Ulmer University of Amsterdam Amsterdam, The Netherlands dennis.ulmer@mailbox.org Abstract Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on knowledge conflicts. Probes trained on AttriWiki achieve up to 0.96 Macro-F1F_1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transfer to SQuAD and WebQuestions with 0.94â0.99 Macro-F1F_1, and generalise zero-shot to Tighidet et al. (2024)âs benchmark, outperforming their probe on conflicting settings without retraining. Furthermore, attribution mismatches raise error rates by up to 70%, though correct attribution does not guarantee correct answers, pointing to the need for broader detection frameworks. AttriWiki Probing for Knowledge Attribution in Large Language Models Ivo Brink KPMG NL / University of Amsterdam Amsterdam, The Netherlands ivobrink00@gmail.com Alexander Boer KPMG NL Amsterdam, The Netherlands boer.alexander@kpmg.nl Dennis Ulmer University of Amsterdam Amsterdam, The Netherlands dennis.ulmer@mailbox.org 1 Introduction Picture yourself rushing to book a last-minute flight for a funeral. You ask an airlineâs chatbot whether you can claim a bereavement fare, and the system confidently replies, âYes, you have 90 days.â The rule, however, never existed, and months later, the refund is denied. This example stems from a real case in which Air Canada was held liable for its chatbotâs misinformation (Cecco, 2024; Lifshitz and Hung, 2024). A single inaccurate response can now impose serious financial, legal, or practical consequences. Although LLM hallucinations, untrue or ungrounded outputs (Huang et al., 2025), have become a central research focus, existing detection techniques often miss a more fundamental issue: users lack transparency over where a modelâs answers come from. The airline could have benefited from knowing the chatbotâs true provenance, as attribution would have enabled verification against official policy. Existing prior approaches in research, including LLMs-as-a-judge (Manakul et al., 2023; Ravi et al., 2024; Friel and Sanyal, 2023) and uncertainty measures (Tomani et al., 2024; Shelmanov et al., 2025), struggle because models can be confidently wrong (Kuhn et al., 2023; Simhi et al., 2025) and because âhallucinationâ lacks a consistent operational definition (Qi et al., 2024). To address this gap, we turn to contributive attribution (Worledge et al., 2024), which distinguishes between contextual knowledge (from the prompt or retrieved evidence) and parametric knowledge (stored in the modelâs weights). This framing complements the well-known divide between faithfulness and factuality errors (Huang et al., 2025; Qi et al., 2024; Ye et al., 2024) within hallucination taxonomies, but shifts the focus: instead of asking whether the output is true, we ask whether the model relied on the source the user intended. Attribution signals can help users detect when a model defaults to conflicting parametric knowledge or, conversely, when it overly relies on irrelevant context. This is particularly valuable in retrieval-augmented generation systems (Lewis et al., 2020; Guu et al., 2020; Borgeaud et al., 2022; Izacard et al., 2023), where errors often arise not from bad retrieval but from the model ignoring retrieved information (Petroni et al., 2020; Li et al., 2023). Our contributions are as follows: Using the AttriWiki pipeline involving Wikipedia passages, entity selection, and controlled prompt variations, we create a dataset, in which each completion has a clear, verifiable knowledge source (parametric or contextual). During generation, we record hidden states at key token positions, producing attribution features without manual labels or conflicting sources. We then train lightweight probing classifiers on these representations to distinguish contextual from parametric retrieval and test their generalisation on external QA datasets such as WebQuestions and SQuAD (Berant et al., 2013; Rajpurkar et al., 2016). and compare to existing attribution probing work (Tighidet et al., 2024). We show that contributive attribution is linearly decodable from LLM hidden states, particularly from the middle to upper layers. Attribution mismatches, in which a model answers from the wrong knowledge source, significantly increase error rates, especially in misleading contexts. Yet, correct attribution alone does not guarantee factual correctness. AttriWiki and all code are publicly available.111An anonymised repository is available under https://anonymous.4open.science/r/_AttriWiki_. 2 Related Work Knowledge in LLMs. Understanding hallucinations ultimately requires understanding how large language models acquire, store, and retrieve knowledge. Factual knowledge in LLMs is parametric and emergent: no single parameter encodes a fact, but distributed activation patterns collectively give rise to factual recall (Dai et al., 2022; Yao et al., 2024; Wang and Xu, 2025). Early probing work, such as LAMA (Petroni et al., 2019) showed that factual recall in language models is brittle, varying substantially across languages, paraphrases, and prompt formulations (Kassner et al., 2021; Elazar et al., 2021a; Jiang et al., 2020). In addition, recent works formalise knowledge in models by mapping philosophical notions of knowledge onto model behaviour: Under the âtrue-beliefâ framework (Fierro et al., 2024) for instance, a fact is known by a model if it is reproduced consistently, which we use to make a practical distinction between parametric knowledge and context-derived information. Attribution and Knowledge Conflicts. Once parametric and contextual knowledge are distinguished, a central challenge is detecting their usage. Knowledge conflicts arise when parametric beliefs contradict retrieved or provided context, a situation common in retrieval-augmented settings (Petroni et al., 2020; Li et al., 2023). Attribution methods aim to trace generated content back to its source. Retrieval-augmented generation approaches (Lewis et al., 2020; Izacard et al., 2023) expose evidence to the model, but cannot guarantee its use, and models frequently hallucinate citations even when retrieval is correct (Zuccon et al., 2023). Model-based attribution methods target contributive attribution, either through architectural modifications that encode source identifiers (Khalifa et al., 2024) or probing approaches that infer whether outputs derive from memory or context via knowledge conflicts (Tighidet et al., 2024). However, the former requires substantial architectural changes, while the latter relies on adversarial prompts that conflate provenance with knowledge resolution. Recent mechanistic work has begun to clarify how conflicting knowledge signals are processed. Zhao et al. (2025) trace entity-level knowledge flows through the model, showing that parametric and contextual knowledge are routed through largely distinct attention circuits and coexist as superposed signals rather than competing via direct suppression. Their entity-conditioned probes analyse how competing signals evolve across a forward pass, making their setup better suited to study conflict resolution dynamics rather than genuine source detection in the absence of conflict. Nonetheless, this perspective provides a mechanistic foundation for attribution: because both knowledge sources are routed through largely distinct circuits and persist as superposed signals, their relative contribution might remain detectable even without an explicit conflict. Probing LLM Representations. Probing hidden states offers a lightweight approach for studying knowledge usage directly from model internals. This line of work builds on a long tradition of analysing how linguistic and factual information is encoded across transformer layers (Conneau et al., 2018; Tenney et al., 2019; Hewitt and Manning, 2019). Prior studies show that layers specialise in surface, syntactic, and semantic features (Jawahar et al., 2019), that factual associations emerge in neuron-level keyâvalue structures (Geva et al., 2021), and that hidden states reflect latent capabilities such as lying (Azaria and Mitchell, 2023), causal reasoning (Rohekar et al., 2024), and instruction following (Heo et al., 2025). These findings suggest that signals relevant to attribution and provenance are also present in model activations. Probing for factual knowledge already has been explored by a series of works (see Youssef et al., 2023 for an overview), but recent work applies probing specifically to hallucination detection. MIND (Su et al., 2024) trains classifiers on final-layer activations, outperforming SelfCheckGPT (Manakul et al., 2023) and GPT-4o critics on several benchmarks. However, it provides limited interpretability regarding why a response is classified as hallucinated. Overall, prior work offers strong taxonomies and attribution techniques, but a key gap remains: existing methods lack scalable, controlled datasets that cleanly isolate parametric versus contextual knowledge use. This motivates approaches that explicitly construct such contrasts and probe hidden states to reveal the true origin of model outputs. 3 A Self-Supervised Data Pipeline 3.1 Motivation Although recent work examines the use of parametricâcontextual knowledge, existing datasets do not isolate retrieval modes. Many studies expose both sources simultaneously through conflict settings and infer attribution post hoc from the modelâs answer (Tighidet et al., 2024; Xu et al., 2024; Zhao et al., 2025), conflating provenance with downstream decision-making and conflict resolution. Therefore, checking whether a fact is present in a modelâs parametric memory is necessary to isolate contextual vs. parametric examples. Prior work verifies parametric knowledge via data recency (Zhang et al., 2024) or by querying the model (Cheng et al., 2024). However, these approaches neither explicitly detect attribution nor prevent the simultaneous availability of parametric and contextual knowledge when analysing their interactions. To this end, we introduce AttriWiki, a self-supervised data pipeline that generates paired examples which force models to draw information either from context or from stored knowledge, established through extensive knowledge testing. By ensuring that only a single knowledge source is available at generation time, AttriWiki enables isolated attribution analysis without manual labelling. Figure 1: Overview of the data generation pipeline, showing entity selection, knowledge testing, prompt construction, and hidden-state extraction. 3.2 Data and Models Data. Wikipedia provides dense factual text that mid-sized open language models (approximately 7â8B parameters) only partially memorise, allowing both retrieval modes to occur naturally. We sample 20k pages using the MediaWiki API,222See https://en.wikipedia.org/w/api.php (last accessed May 20, 2026). The Wikipedia text is licensed under C BY-SA 3.0/4.0, and all use complies with the terms of that license. discarding snippets shorter than 300 characters. Data Generation Model. We use GPT-4o-mini (gpt-4o-mini; OpenAI, 2024) as an instruction-following model to paraphrase passages, remove or retain entities, and rewrite completion sentences during dataset construction (see next section). Target Models. Attribution probing is performed on open-weight language models from which we extract hidden states during generation. We use Llama-3.1-8B (Dubey et al., 2024), Mistral-7B (Jiang et al., 2023), and Qwen2.5-7B (Qwen et al., 2025) as target models. 3.3 Method An overview of the data generation pipeline is shown in FigureË1. Because attribution concerns the source of specific factual content, we focus on named entities as attribution anchors, identified using spaCyâs transformer-based NER (Honnibal et al., 2020). Entities that overlap, are synonymous, or closely resemble the passage title are found using a RoBERTa (Liu et al., 2019) cross-encoder (CE) similarity model.333cross-encoder/stsb-roberta-base. The CE assigns pairwise semantic similarity scores, and entities above a threshold of 0.60.6 are considered synonyms (we refer to this procedure as synonym matching). In this case, synonyms are excluded.. The threshold is determined heuristically and used consistently throughout this work. From the remaining entity candidates, GPT-4o-mini selects up to three representative, non-numeric, non-temporal entities per passage (see SectionËC.1.1). Each entityâpassage pair undergoes knowledge testing to determine whether the target model can recall the entity without the passage, serving as an operational proxy for parametric memory. We use three prompt formats to test entity knowledge (see SectionËC.1.2). If any test elicits the correct entity (exact or synonym match), it is labelled known; otherwise, unknown. This high-recall labelling avoids false unknown labels, preferring unanswerable and thus unusable over multi-source samples. We then construct contextual and parametric variants for each passage. For known entities, all mentions are removed using GPT-4o-mini (SectionËC.1.3), thereby forcing parametric retrieval; for unknown entities, the name remains visible, while a different entity is removed to prevent lexical bias this may introduce. GPT-4o-mini appends a short, natural completion cue such as âThe [description] is called âŠ,â yielding paired prompts that differ only in the presence of the entity (see SectionËC.1.4). Each prompt is submitted to the target model to produce a completion. During greedy decoding, we extract hidden representations at two key token positions: (i) the first generated token (FTG), which reflects the modelâs initial response intent allowing us to compare responses without specific entities (Su et al., 2024), and (i) the last token of the entity (LTE), capturing the fully integrated entity semantics (Tighidet et al., 2024).444We omit first-token-entity as it coincides with first-token-generation in almost all cases. Entity spans are located by exact or synonym matching, and the corresponding hidden states are serialised for attribution classification. This automated pipeline produces large volumes of data, with each example having a verifiable knowledge source. AttriWiki Ratio (P:C) Syn Size Llama-3.1-8B â 5:3 10.3% 27,244 Mistral-7B-v0.1 â 3:2 10.7% 26,853 Qwen2.5-7B â 1:1 12.0% 24,336 Table 1: AttriWiki dataset composition for each model. Ratio (P:C) denotes the approximate parametric-to-contextual proportion of examples. Syn denotes the synonym match rate, and Size denotes the total number of examples. Results. TableË1 summarises the AttriWiki dataset for Llama-3.1-8B, Mistral-7B-v0.1, and Qwen2.5-7B. Llama and Mistral show parametric-dominant distributions (â 5:3 and 3:2), while Qwen is nearly balanced (â 1:1). Dataset sizes are similar across models. Synonym match rates (percentage of completions that matched semantically but not verbatim) are highest for Qwen (12.0%). Overall, the pipeline behaves consistently. Figure 2: Per-layer PCA of decoder hidden states at the first generated token (FTG). Qwen has 28 layers (vs. 32 in Llama and Mistral). In the mid to upper layers, contextual and parametric activations increasingly diverge, indicating greater separability. Plots for all layers are given in SectionËB.1. Latent Space Analysis. To gain insight into how contextual and parametric information is represented internally, we perform a per-layer 2D principal component analysis (PCA; Pearson, 1901) on decoder hidden states extracted at the first token of generation (FTG). As shown in FigureË2 (refer to AppendixËB for all the layers), representations from early layers largely overlap, while a clearer separation between contextual and parametric activations emerges from the middle layers onward in all three models. This pattern suggests that attribution-related information becomes increasingly linearly accessible in higher layers, indicating that contextual and parametric signals are likely encoded in separable directions of the hidden state space. 3.4 Evaluating Bias To verify that AttriWiki does not introduce lexical shortcuts between classes as an unintended effect of our knowledge-proxy,555For instance, generic descriptors such as âa countryâ may occur more frequently in contexts associated with widely known facts, inadvertently biasing a classifier toward predicting parametric knowledge based on surface-level phrasing rather than attribution. we train text-only classifiers on the passages. A balanced bag-of-words logistic regression and a logistic regression model trained on DeBERTaV3666microsoft/deberta-v3-large. embeddings (He et al., 2023) are evaluated under five-fold cross-validation (for specifics see AppendixËA). Model Classifier F1F_1 Llama-3.1-8B BoW .652±.008.652±.008 Embedding .663±.002.663±.002 Mistral-7B-v.1 BoW .660±.005.660±.005 Embedding .667±.005.667±.005 Qwen2.5-7B BoW .675±.004.675±.004 Embedding .648±.006.648±.006 Table 2: Bias detection performance using BoW and embedding-based classifiers on AttriWiki with Macro-F1F_1 (mean ± cross-validation standard deviation). Results. Results are shown in TableË2. Bag-of-words and embedding-based classifiers achieve F1F_1 scores of 0.65â0.67, indicating the presence of lexical cues, though these alone are insufficient for reliable performance. BoW on the data for all models shows similar performance and nearly identical unigram profiles: contextual examples favour entity-specific terms (e.g., named, born in, towns), whereas parametric examples lean toward broader encyclopedic vocabulary (e.g., countries, nationalities). This suggests that our knowledge proxy introduces a mild and stable lexical bias. 4 Attribution Classification To assess whether hidden states in AttriWiki truly encode the source of a modelâs knowledge, we train three probing classifiers of increasing complexity. Each model is trained on 64% of AttriWiki, validated on 16%, and tested on the remaining 20%, then evaluated for out-of-domain generalisation. Results are reported as macro-F1F_1. Consult SectionËA.2 for more details. 4.1 Classifiers Since the data for Llama and Mistral shows a mild class imbalance (â 3:2 parametricâcontextual), we compensate through loss weights. To prevent information leakage, we further enforce title-disjoint trainâtest splits, ensuring that no samples originating from the same article appear in both splits. Final-Layer Logistic Regression (Final-LR). As a minimal baseline, we train a single logistic unit on the normalised hidden vector LââHh_L ^H from the final transformer layer: p=Ïâ(â€âL+b),p=Ï(w h_L+b), (1) where ââHw ^H and bââb are the classifier weights and bias, Ïâ(â )Ï(·) is the sigmoid function, and training uses an â2 _2-regularised logistic loss with inverse-frequency class weighting. Layer-Weighted Logistic Regression (Layer-LR). To exploit depth information, we learn a set of unconstrained parameters ââL Ξ ^L, which are transformed via a softmax to obtain aggregation weights =softmaxâ() α=softmax( Ξ) over transformer layers. These weights are used to form a blended representation ÂŻ=ââ=1Lαâââ. h= _ =1^L _ h_ . (2) A logistic head then predicts the probability of parametric retrieval from ÂŻ h. All parameters (,,b)( Ξ,w,b) are trained jointly using binary cross-entropy with logits and positive-class reweighting to account for label imbalance, optimised with AdamW (Kingma and Ba, 2015). Hyperparameters and training details are reported in SectionËA.2. Layer-Weighted MLP (Layer-MLP). We extend the previous model with sparsemax-normalised aggregation weights (Martins and Astudillo, 2016), encouraging focus on a small number of salient layers, motivated by the tendency of softmax-aggregated MLP probes to learn diffuse layer weights. The linear classifier is replaced with a compact two-layer MLP operating on the aggregated representation ÂŻ h: p=Ïâ(2â€âGELUâ(1âÂŻ)+b),p=Ï\! (w_2 GELU(W_1 h)+b ), (3) We use GELU (Hendrycks and Gimpel, 2016), where 1ââmĂHW_1 ^mĂ H projects the hidden representation of dimension H into a bottleneck of size m, and 2ââmw_2 ^m denotes the output-layer weights. Hyperparameters and training details are reported in SectionËA.2. 4.2 Out-of-distribution Datasets All classifiers are first trained on AttriWiki and then evaluated on 2,0002,000 samples from two external QA datasets. The best classifier is additionally compared against the probe of Tighidet et al. (2024) on conflicting scenarios. All samples include a single demonstration line to enforce output format (see SectionËC.2). ParaRel (Elazar et al., 2021b). To situate our approach relative to Tighidet et al. (2024), we evaluate on prompts derived from ParaRel, a structured dataset of factual (subject, relation, object) triples sourced from Wikidata. Each triple is converted into a probing prompt consisting of a counter-factual statement contradicting the modelâs parametric knowledge, followed by a query over the same subjectârelation pair. Labels are assigned based on whether the modelâs response matches the planted counter-object (CK) or the true parametric object (PK), other answers are ignored. Relations are grouped into eight semantic categories. WebQuestions (Berant et al., 2013). Each factoid question lacks supporting context (e.g., âWho developed general relativity?â â Albert Einstein); hence, answers must come from the modelâs parametric memory. Items are wrapped in a one-shot question-answer prompt for compatibility with decoder-only models. SQuAD (Rajpurkar et al., 2016). Passages contain answers verbatim, making the task predominantly contextual. To filter memorised facts, we briefly query each question without its paragraph; if the model still answers correctly, the item is recorded as prior knowledge and used for a separate experiment (see below). The remaining examples form the contextual test set. SQuAD Variants. We derive three controlled variants: (a) an irrelevant context split pairing each known question with an irrelevant, randomly sampled paragraph to force parametric retrieval; (b) an answer-string-decoy version where the paragraph contains the correct answer verbatim in an unrelated context, testing reliance on surface overlap; and (c) a dual-source version retaining both context and prior-knowledge cases to study retrieval preference. Together, these variants test robustness against degenerate heuristics, e.g., predicting attribution solely from the presence of context. 4.3 Results All reported macro-F1F_1 values are accompanied by bootstrap standard deviations estimated from B=1,000B=1,000 samples with replacement. Model Final-LR Layer-LR Layer-MLP Token = FTG (First Token of Generation) Llama-3.1-8B .841±.005.841±.005 .919±.004.919±.004 .922±.004.922±.004 Mistral-7B-v0.1 .844±.005.844±.005 .922±.004.922±.004 .933±.004.933±.004 Qwen2.5-7B .835±.005.835±.005 .912±.004.912±.004 .926±.004.926±.004 Token = LTE (Last Token of Entity) Llama-3.1-8B .854±.005.854±.005 .948±.003.948±.003 .955±.003.955±.003 Mistral-7B-v0.1 .904±.004.904±.004 .958±.003.958±.003 .961±.003.961±.003 Qwen2.5-7B .867±.005.867±.005 .935±.004.935±.004 .949±.006.949±.006 Table 3: Macro-F1F_1 (mean ± bootstrap std, B=1000B=1000) on AttriWiki-test (nâ5,000nâ 5,000). Attribution Performance. TableË3 reports attribution classification on the held-out AttriWiki split. We compare Final-LR, Layer-LR, and Layer-MLP. Layer aggregation improves F1F_1 by 7â11 p over the final-layer baseline, while the MLP yields only marginal gains (â 1 p). Token choice matters: the last token of the entity (LTE) performs best (0.95â0.96), while the first token of generation (FTG) remains slightly lower but above 0.90. Mistral consistently outperforms Llama and Qwen by one point, though trends are consistent across models, with Macro-F1F_1 closely matching accuracy. Model SQuAD WebQ Llama-3.1-8B .977±.003.977±.003 .955±.004.955±.004 Mistral-7B-v0.1 .918±.005.918±.005 .996±.001.996±.001 Qwen2.5-7B .996±.005.996±.005 .991±.002.991±.002 Table 4: Generalisation accuracy (mean ± bootstrap std, B=1000B=1000) of Layer-LR (FTG) on out-of-domain QA datasets. WebQ = WebQuestions. Generalisation. To assess out-of-distribution generalisation, we evaluate the classifier on two out-of-domain QA benchmarks: SQuAD (contextual) and WebQuestions (parametric). As shown in TableË4, it generalises near perfectly without fine-tuning, achieving 0.91â0.99 accuracy across Llama, Mistral, and Qwen on both benchmarks. Similar results are observed for the Layer-MLP in the appendix, in TableË12. We report only on first-token generation (FTG), because it allows us to analyse incorrect responses in later sections. Figure 3: Layer-aggregation weights learned by Layer-LR for first-token generation and last-token entity representations. Curves show learned weights across transformer layers, smoothed with a Gaussian kernel for visual clarity. The x-axis is restricted to layers 10â24, as earlier and later layers receive negligible weight. Which Layers Does the Classifier Rely on? As shown in FigureË3, the classifier concentrates weight in the lower-middle transformer layers across models when using the last token of the entity. For first-token generation, peak contributions occur at layers 16 (Llama-3.1-8B), 19 (Mistral-7B-v0.1), and 21 (Qwen-2.5-7B), while last-token entity representations peak earlier at layers 14 (Llama, Mistral) and 18 (Qwen). Notably, Qwen assigns weight to relatively deeper layers given its shallower overall depth (28 versus 32). This indicates that using the last token of the entity might make the probe focused on entity-specific, lexical information, while the first-token generation hidden states push the probe to focus on potentially high-level information processed in later layers. Comparison to Tighidet et al. (2024). TableË5 compares our AttriWiki-trained probe Layer-LR on FTG against the best-performing layer of Tighidet et al. (2024) evaluated on ParaRel under the same exact same scheme as Tighidet et al. (2024). The AttriWiki probe consistently outperforms across all three models on the never seen before data, with gains ranging from 33 p (Qwen2.5-7B) to 1313 p (Llama-3.1-8B). Unlike Tighidet et al. (2024), our probe requires no conflicting context during training, extracts hidden states at the first generated token rather than the last input token, and learns to aggregate across all layers jointly rather than selecting a single layer post-hoc. We include the per-category results in AppendixËB. Model Probe Macro-F1F_1 Mistral-7B-v0.1 Tighidet et al. (L10) .745±.014.745±.014 AttriWiki .806±.012.806±.012 Llama-3.1-8B Tighidet et al. (L11) .640±.019.640±.019 AttriWiki .771±.016.771±.016 Qwen2.5-7B Tighidet et al. (L22) .757±.007.757±.007 AttriWiki .793±.006.793±.006 Table 5: Comparison of Tighidet et al. and AttriWiki probe (Layer-LR on FTG) across models with Macro-F1F_1(mean ± bootstrap std, B=1000B=1000). Layer number in parentheses indicates the best-performing layer. 4.4 Ablation For the linear probe, ablation is performed on Layer-LR FTG because of its robustness and it does not rely on entity mentions. Degenerate Attribution. We evaluate whether attribution probes rely on degenerate heuristics rather than genuine provenance signals. Firstly, we test the answer-string decoy variant of SQuAD, where the gold answer appears verbatim in an unrelated contextâperformance drops for all classifiers. Nevertheless, Layer-LR remains well above chance (0.724).777Semantic irrelevance is not explicitly verified; we use this ablation more to uncover potential model shortcuts. In contrast, the MLP collapses to near-random performance (0.541), indicating that the MLP overfits to lexical shortcuts such as entity repetition, whereas Layer-LR captures a more robust attribution signal (see TableË11 in AppendixËB). Secondly, in the irrelevant-context settingâwhere known facts are paired with unrelated passagesâ92% of answers are still attributed as parametric, ruling out a trivial heuristic that equates the mere presence of context with contextual attribution. Thirdly, when both parametric memory and context suffice to answer a question, models overwhelmingly default to the contextual channel (87.1%). This rules out the explanation that the probe merely detects whether a fact is stored in the modelâs parameters, rather than the knowledge source. Error and Attribution Mismatch. To analyse the relationship between attribution and answer correctness, we consider two complementary conditions: (i) present parametric knowledge paired with irrelevant context, and (i) absent parametric knowledge paired with relevant context. We use SQuAD with random contexts in case of parametric knowledge and original context otherwise, recording answer correctness. Source alignment (parametric when required, or contextual when required) and correctness form a 2Ă22Ă2 contingency table, analysed using Fisherâs exact test (Fisher, 2018), with relative risk as the effect size. Attribution mismatches strongly correlate with errors (pâȘ0.001p 0.001), with asymmetric effects: relying on misleading context increases errors by up to â 70% when parametric knowledge is required, whereas defaulting to parametric memory in contextual settings increases errors by only â 30%. 5 Discussion Our findings indicate that contributive attribution, whether a completion is driven by contextual evidence or parametric memory, is a representation-level property that is accessible from hidden states, offering a diagnostic lens on knowledge channel usage that complements hallucination taxonomies distinguishing faithfulness from factuality errors (Huang et al., 2025; Qi et al., 2024; Ye et al., 2024). Our mismatch analysis motivates this: using the wrong knowledge source substantially increases error risk, particularly when misleading context overrides parametric knowledge. Yet many errors persist despite correct source alignment, indicating that attribution is one diagnostic axis among others rather than a complete error detector. Moreover, attribution is learnable even without explicit knowledge conflicts, as evidenced by our strong out-of-distribution performance including on unseen conflict settings. This challenges conflict-based approaches (Tighidet et al., 2024; Xu et al., 2024), which conflate attribution with conflict resolution by design. Where prior mechanistic work shows that contextual and parametric signals coexist as superposed activations routed through largely distinct circuits (Zhao et al., 2025), our results go further: their relative contribution is quantitatively detectable at generation time without ever exposing contradictory information, suggesting conflict-free training is not merely sufficient but actually preferable. Two further results confirm prior work. The limited gains from an MLP head and its degradation under the answer-string decoy suggest that added model complexity encourages shortcut learning, while linear probes more reliably predict provenanceâconsistent with the knowledge source being another property encoded linearly in hidden representations (Park et al., 2024; Jiang et al., 2024; Merullo et al., 2025). The preference of the probe for upper-middle layers aligns with prior findings that mid-layer activations are most informative for knowledge attribution (Tighidet et al., 2024), and with broader evidence that later layers encode decision-relevant abstractions beyond surface form (Jawahar et al., 2019; Tenney et al., 2019). 6 Conclusions We study contributive attribution in large language models using a self-supervised, conflict-free pipeline (AttriWiki). Linear probes trained on AttriWiki achieve up to 0.96 F1F_1 on Llama, Mistral, and Qwen models and generalise strongly to SQuAD, WebQuestions and unseen conflict settings. Attribution mismatches increase error risk by 30â70%, particularly under misleading contexts, although correct attribution alone does not guarantee correctness. Together, these findings establish attribution as a meaningful diagnostic signal. Future work should explore how attribution signals can be integrated into downstream systems. In retrieval-augmented generation (RAG; Guu et al., 2020; Borgeaud et al., 2022), attribution probes could provide online signals to detect ignored evidence and trigger selective verification or re-retrieval. In LLM-powered chatbots, exposing attribution may also elicit more critical user engagement by clarifying whether responses rely on retrieved evidence or prior knowledge, similar to Deng et al. (2023). Other works have identified potential interventions through steering model activiations along an identifier direction (Basu et al., 2025; Bi et al., 2025). Limitations Data. AttriWikiâs controlled Wikipedia domain may limit generalisation; testing attribution signals on legal text, dialogue, and noisy web data is essential. Extension to multilingual settings also poses challenges to AttriWiki since knowledge varies across languages (Kassner et al., 2021). Additionally, shortcut learning poses a risk, inputs might be lexically associated with parametric knowledge as demonstrated in SectionË3.4. Furthermore, distribution shift remains a concern: ParaRel covers Wikidata subjects not present in AttriWiki, and per-category results in AppendixËB show that probe performance varies greatly across categories, suggesting sensitivity to unseen domains. Attribution Classification. The classifier would benefit from moving beyond token-level attribution to sequence-level approaches. While our current method performs well on factual questions, it struggles in more natural settings where no clear entity span exists. Aggregation strategies, such as attention-based span scoring, could provide more accurate attribution in these cases. Another limitation is model dependence: each LLM requires a new classifier, and retraining the model often necessitates retraining the classifier and regenerating data if the modelâs knowledge evolves. Ablation. Although informative, error analysis remains challenging. Phenomena such as contradictions, exaggerations, and refusal to answer are often grouped under hallucinations, yet our focus here is on forced factual retrieval. Assessing answer correctness is equally difficult. Minor variations, such as one-letter acronym slips (e.g., GDP â GPD), can bypass the cross-encoder (synonym matching), while valid polarity with low lexical overlap may also lead to misclassifications. For example, the answer âsmallerâ correctly responds to âWere the houses bigger or smaller?â but fails against the gold label âsmaller houses.â Ethics Considerations This work only uses publicly available datasets and does not involve human subjects or personal data. One API-based LLM is used solely for data generation. We do not identify significant ethical risks beyond those common to interpretability research. Acknowledgments This work is supported by the Dutch National Science Foundation (NWO Vici VI.C.212.053) and KPMG NL. We would also like to express our gratitude to Ivan Titov for his valuable guidance and feedback throughout this work. References A. Azaria and T. M. Mitchell (2023) The internal state of an LLM knows when itâs lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), p. 967â976. External Links: Document, Link Cited by: §2. S. Basu, V. Morariu, Z. Wang, R. Rossi, C. Zhao, S. Feizi, and V. Manjunatha (2025) On mechanistic circuits for extractive question-answering. arXiv preprint arXiv:2502.08059. Cited by: §6. J. Berant, A. Chou, R. Frostig, and P. Liang (2013) Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, D. Yarowsky, T. Baldwin, A. Korhonen, K. Livescu, and S. Bethard (Eds.), Seattle, Washington, USA, p. 1533â1544. External Links: Link Cited by: §1, §4.2. B. Bi, S. Liu, Y. Wang, Y. Xu, J. Fang, L. Mei, and X. Cheng (2025) Parameters vs. context: fine-grained control of knowledge reliance in language models. arXiv preprint arXiv:2503.15888. Cited by: §6. S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J. Lespiau, B. Damoc, A. Clark, D. De Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. Rae, E. Elsen, and L. Sifre (2022) Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, p. 2206â2240. External Links: Link Cited by: §1, §6. L. Cecco (2024) Air canada ordered to pay customer who was misled by airlineâs chatbot. The Guardian. Note: https://w.theguardian.com/world/2024/feb/16/air-canada-chatbot-lawsuit External Links: Link Cited by: §1. S. Cheng, L. Pan, X. Yin, X. Wang, and W. Y. Wang (2024) Understanding the interplay between parametric and contextual knowledge for large language models. CoRR abs/2410.08414. External Links: Document, Link, 2410.08414 Cited by: §3.1. A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni (2018) What you can cram into a single \$&!#* vector: probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, I. Gurevych and Y. Miyao (Eds.), p. 2126â2136. External Links: Document, Link Cited by: §2. D. Dai, L. Dong, Y. Hao, Z. Sui, B. Chang, and F. Wei (2022) Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 8493â8502. Cited by: §2. Y. Deng, L. Liao, L. Chen, H. Wang, W. Lei, and T. Chua (2023) Prompting and evaluating large language models for proactive dialogues: clarification, target-guided, and non-collaboration. In Findings of the Association for Computational Linguistics: EMNLP 2023, p. 10602â10621. Cited by: §6. A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. RoziĂšre, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. M. Kloumann, I. Misra, I. Evtimov, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, and et al. (2024) The llama 3 herd of models. CoRR abs/2407.21783. External Links: Document, Link, 2407.21783 Cited by: §3.2. Y. Elazar, N. Kassner, S. Ravfogel, A. Ravichander, E. H. Hovy, H. SchĂŒtze, and Y. Goldberg (2021a) Measuring and improving consistency in pretrained language models. Trans. Assoc. Comput. Linguistics 9, p. 1012â1031. External Links: Document, Link Cited by: §2. Y. Elazar, N. Kassner, S. Ravfogel, A. Ravichander, E. Hovy, H. SchĂŒtze, and Y. Goldberg (2021b) Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics 9, p. 1012â1031. External Links: Document, Link Cited by: §4.2. C. Fierro, R. Dhar, F. Stamatiou, N. Garneau, and A. SĂžgaard (2024) Defining knowledge: bridging epistemology and large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), p. 16096â16111. External Links: Document, Link Cited by: §2. R. A. Fisher (2018) The logic of inductive inference. Journal of the Royal Statistical Society 98 (1), p. 39â54. External Links: Document, ISSN 0952-8385, Link, https://academic.oup.com/jrsssa/article-pdf/98/1/39/49707080/jrsssa_98_1_39.pdf Cited by: §4.4. R. Friel and A. Sanyal (2023) Chainpoll: a high efficacy method for llm hallucination detection. External Links: Link, 2310.18344 Cited by: §1. M. Geva, R. Schuster, J. Berant, and O. Levy (2021) Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, p. 5484â5495. External Links: Document, Link Cited by: §2. K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang (2020) Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine LearningAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtualThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025Proceedings of the First International OpenKG Workshop: Large Knowledge-Enhanced Models co-locacted with The International Joint Conference on Artificial Intelligence (IJCAI 2024), Jeju Island, South Korea, August 3, 2024Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023The Thirty-ninth Annual Conference on Neural Information Processing SystemsProceedings of the 57th Annual Meeting of the Association for Computational LinguisticsProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, H. D. I, A. Singh, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, H. Lin, N. Zhang, T. Wu, M. Wang, G. Qi, H. Wang, H. Chen, H. Bouamor, J. Pino, K. Bali, A. Korhonen, D. Traum, L. MĂ rquez, A. Rogers, J. L. Boyd-Graber, and N. Okazaki (Eds.), Proceedings of Machine Learning ResearchCEUR Workshop Proceedings, Vol. 1193818, p. 3929â3938. External Links: Link Cited by: §1, §6. P. He, J. Gao, and W. Chen (2023) DeBERTaV3: improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §3.4. D. Hendrycks and K. Gimpel (2016) Gaussian Error Linear Units (GELUs). External Links: http://arxiv.org/abs/1606.08415v3 Cited by: §4.1. J. Heo, C. Heinze-Deml, O. Elachqar, K. H. R. Chan, S. Y. Ren, A. C. Miller, U. Nallasamy, and J. Narain (2025) Do llms "know" internally when they follow instructions?. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §2. J. Hewitt and C. D. Manning (2019) A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, p. 4129â4138. External Links: Document, Link Cited by: §2. M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd (2020) spaCy: Industrial-strength Natural Language Processing in Python. External Links: Document Cited by: §3.3. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu (2025) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43 (2), p. 1â55. External Links: Document, ISSN 1558-2868, Link Cited by: §1, §1, §5. G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave (2023) Atlas: few-shot learning with retrieval augmented language models. J. Mach. Learn. Res. 24, p. 251:1â251:43. External Links: Link Cited by: §1, §2. G. Jawahar, B. Sagot, and D. Seddah (2019) What does BERT learn about the structure of language?. Florence, Italy, p. 3651â3657. External Links: Document, Link Cited by: §2, §5. A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de Las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed (2023) Mistral 7b. CoRR abs/2310.06825. External Links: Document, Link, 2310.06825 Cited by: §3.2. Y. Jiang, G. Rajendran, P. K. Ravikumar, B. Aragam, and V. Veitch (2024) On the origins of linear representations in large language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, p. 21879â21911. Cited by: §5. Z. Jiang, F. F. Xu, J. Araki, and G. Neubig (2020) How can we know what language models know?. Transactions of the Association for Computational Linguistics 8, p. 423â438. External Links: Document, Link Cited by: §2. N. Kassner, P. Dufter, and H. SchĂŒtze (2021) Multilingual LAMA: investigating knowledge in multilingual pretrained language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, P. Merlo, J. Tiedemann, and R. Tsarfaty (Eds.), Online, p. 3250â3258. External Links: Document, Link Cited by: §2, Data.. M. Khalifa, D. Wadden, E. Strubell, H. Lee, L. Wang, I. Beltagy, and H. Peng (2024) Source-aware training enables knowledge attribution in language models. External Links: Link, 2404.01019 Cited by: §2. D. P. Kingma and J. Ba (2015) Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Y. Bengio and Y. LeCun (Eds.), External Links: Link Cited by: §4.1. L. Kuhn, Y. Gal, and S. Farquhar (2023) Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §1. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. KĂŒttler, M. Lewis, W. Yih, T. RocktĂ€schel, S. Riedel, and D. Kiela (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. External Links: Link Cited by: §1, §2. D. Li, A. S. Rawat, M. Zaheer, X. Wang, M. Lukasik, A. Veit, F. Yu, and S. Kumar (2023) Large language models with controllable working memory. In Findings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 1774â1793. External Links: Document, Link Cited by: §1, §2. L. R. Lifshitz and R. Hung (2024) Business Law Today, American Bar Association. Note: External Links: Link Cited by: §1. Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov (2019) RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs/1907.11692. External Links: Link, 1907.11692 Cited by: §3.3. P. Manakul, A. Liusie, and M. J. F. Gales (2023) SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), p. 9004â9017. External Links: Document, Link Cited by: §1, §2. A. F. T. Martins and R. F. Astudillo (2016) From softmax to sparsemax: A sparse model of attention and multi-label classification. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, M. Balcan and K. Q. Weinberger (Eds.), JMLR Workshop and Conference Proceedings, Vol. 48, p. 1614â1623. External Links: Link Cited by: §4.1. J. Merullo, N. Smith, S. Wiegreffe, and Y. Elazar (2025) On linear representations and pretraining data frequency in language models. In International Conference on Learning Representations, Vol. 2025, p. 71963â71987. Cited by: §5. OpenAI (2024) OpenAI. Note: Accessed 2025-06-20 External Links: Link Cited by: §3.2. K. Park, Y. J. Choe, and V. Veitch (2024) The linear representation hypothesis and the geometry of large language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, p. 39643â39666. Cited by: §5. K. Pearson (1901) LIII. on lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 2 (11), p. 559â572. External Links: Document Cited by: §3.3. F. Petroni, P. Lewis, A. Piktus, T. RocktĂ€schel, Y. Wu, A. H. Miller, and S. Riedel (2020) How context affects language modelsâ factual predictions. In Conference on Automated Knowledge Base Construction, AKBC 2020, Virtual, June 22-24, 2020, D. Das, H. Hajishirzi, A. McCallum, and S. Singh (Eds.), External Links: Document, Link Cited by: §1, §2. F. Petroni, T. RocktĂ€schel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. H. Miller (2019) Language models as knowledge bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, K. Inui, J. Jiang, V. Ng, and X. Wan (Eds.), p. 2463â2473. External Links: Document, Link Cited by: §2. S. Qi, Y. He, and Z. Yuan (2024) Can we catch the elephant? a survey of the evolvement of hallucination evaluation on natural language generation. External Links: Link, 2404.12041 Cited by: §1, §1, §5. Qwen, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025) Qwen2.5 technical report. External Links: Link, 2412.15115 Cited by: §3.2. P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang (2016) SQuAD: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, J. Su, X. Carreras, and K. Duh (Eds.), p. 2383â2392. External Links: Document, Link Cited by: §1, §4.2. S. S. Ravi, B. Mielczarek, A. Kannappan, D. Kiela, and R. Qian (2024) Lynx: an open source hallucination evaluation model. External Links: Link, 2407.08488 Cited by: §1. R. Y. Rohekar, Y. Gurwicz, S. Yu, and V. Lal (2024) Causal world representation in the gpt model. External Links: Link, 2412.07446 Cited by: §2. A. Shelmanov, E. Fadeeva, A. Tsvigun, I. Tsvigun, Z. Xie, I. Kiselev, N. Daheim, C. Zhang, A. Vazhentsev, M. Sachan, et al. (2025) A head to predict and a head to question: pre-trained uncertainty quantification heads for hallucination detection in llm outputs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 35700â35719. Cited by: §1. A. Simhi, I. Itzhak, F. Barez, G. Stanovsky, and Y. Belinkov (2025) Trust me, iâm wrong: high-certainty hallucinations in llms. arXiv e-prints, p. arXivâ2502. Cited by: §1. W. Su, C. Wang, Q. Ai, Y. Hu, Z. Wu, Y. Zhou, and Y. Liu (2024) Unsupervised real-time hallucination detection based on the internal states of large language models. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), p. 14379â14391. External Links: Document, Link Cited by: §2, §3.3. I. Tenney, P. Xia, B. Chen, A. Wang, A. Poliak, R. T. McCoy, N. Kim, B. V. Durme, S. R. Bowman, D. Das, and E. Pavlick (2019) What do you learn from context? probing for sentence structure in contextualized word representations. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, External Links: Link Cited by: §2, §5. Z. Tighidet, A. Mogini, J. Mei, B. Piwowarski, and P. Gallinari (2024) Probing language models on their knowledge source. CoRR abs/2410.05817. External Links: Document, Link, 2410.05817 Cited by: §B.2, §B.2, Table 13, Table 13, Table 14, Table 14, Table 15, Table 15, §1, §2, §3.1, §3.3, §4.2, §4.2, §4.3, §4.3, Table 5, Table 5, Table 5, Table 5, §5, §5. C. Tomani, K. Chaudhuri, I. Evtimov, D. Cremers, and M. Ibrahim (2024) Uncertainty-based abstention in llms improves safety and reduces hallucinations. arXiv preprint arXiv:2404.10960. Cited by: §1. Z. Wang and C. Xu (2025) Functional abstraction of knowledge recall in large language models. arXiv preprint arXiv:2504.14496. Cited by: §2. T. Worledge, J. H. Shen, N. Meister, C. Winston, and C. Guestrin (2024) Unifying corroborative and contributive attributions in large language models. In IEEE Conference on Secure and Trustworthy Machine Learning, SaTML 2024, Toronto, ON, Canada, April 9-11, 2024, p. 665â683. External Links: Document, Link Cited by: §1. R. Xu, Z. Qi, Z. Guo, C. Wang, H. Wang, Y. Zhang, and W. Xu (2024) Knowledge conflicts for llms: a survey. External Links: Link, 2403.08319 Cited by: §3.1, §5. Y. Yao, N. Zhang, Z. Xi, M. Wang, Z. Xu, S. Deng, and H. Chen (2024) Knowledge circuits in pretrained transformers. Advances in neural information processing systems 37, p. 118571â118602. Cited by: §2. H. Ye, T. Liu, A. Zhang, W. Hua, and W. Jia (2024) Cognitive mirage: A review of hallucinations in large language models. p. 14â36. External Links: Link Cited by: §1, §5. P. Youssef, O. A. Koras, M. Li, J. Schlötterer, and C. Seifert (2023) Give me the facts! A survey on factual knowledge probing in pre-trained language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Findings of ACL, p. 15588â15605. Cited by: §2. H. Zhang, Y. Zhang, X. Li, W. Shi, H. Xu, H. Liu, Y. Wang, L. Shang, Q. Liu, Y. Liu, and R. Tang (2024) Evaluating the external and parametric knowledge fusion of large language models. External Links: Link, 2405.19010 Cited by: §3.1. J. Zhao, Y. Yang, X. Hu, J. Tong, Y. Lu, W. Wu, T. Gui, Q. Zhang, and X. Huang (2025) Understanding parametric and contextual knowledge reconciliation within large language models. External Links: Link Cited by: §2, §3.1, §5. G. Zuccon, B. Koopman, and R. Shaik (2023) ChatGPT hallucinates when attributing answers. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region, SIGIR-AP 2023, Beijing, China, November 26-28, 2023, Q. Ai, Y. Liu, A. Moffat, X. Huang, T. Sakai, and J. Zobel (Eds.), p. 46â51. External Links: Document, Link Cited by: §2. Appendix A Hyperparameters This appendix documents all hyperparameters used in our experiments. We separate configurations for bias evaluation models and attribution classifiers. A.1 Bias evaluation We evaluate bias using two complementary classifiers: a sparse BoW model and an embedding-based model. The former serves as a lightweight lexical baseline, while the latter tests whether bias signals persist in contextualised representations. The corresponding hyperparameters given in TablesË6 and 7, respectively. Component Configuration Text representation TFâIDF ngram_range = (1,2) max_features = 5000 Classifier Logistic regression class_weight = balanced max_iter = 1000 random_state = 42 Table 6: Hyperparameters for the TFâIDF logistic regression bias classifier. Component Configuration Text representation Transformer embeddings model = microsoft/deberta-v3-large max_length = 128 padding = max_length truncation = true repr. = [CLS] dtype = float16 eval mode, no gradients Classifier Logistic regression max_iter = 1000 random_state = 42 Table 7: Hyperparameters for the embedding-based bias classifier. A.2 Attribution classifiers Attribution classifiers are trained to probe internal model representations. Hyperparameters for Layer-LR and Layer-MLP were selected via grid search over dropout p (0,0.1,0.2\0,0.1,0.2\), weight decay λ (0,5Ă10â4,10â3,2Ă10â3\0,5\!Ă\!10^-4,10^-3,2\!Ă\!10^-3\), learning rate η (5Ă10â4,10â3,2Ă10â3\5\!Ă\!10^-4,10^-3,2\!Ă\!10^-3\), and (for Layer-MLP) bottleneck size m (64,128\64,128\), using a title-aware 15% validation split. Models were trained for 10 epochs; the selected configuration maximises validation macro-F1F_1 with early-stopping (tie-breaker: validation accuracy). Selected values are reported in TablesË8, 9 and 10. Component Configuration Input representation Last transformer layer (Xâ[:,â1,:]X[:,-1,:]) Preprocessing StandardScaler (default) Classifier Logistic regression class_weight = balanced solver = lbfgs penalty = â2 _2 max_iter = 1000 Table 8: Hyperparameters for Final-LR (last-layer logistic regression). Component Configuration Input representation Layer stack XââNĂLĂHX ^NĂ LĂ H weighted aggregation Layer weighting Softmax( Ξ), ââαâ=1 _ _ =1 Normalization â2 _2 per layer Dropout p=0.1p=0.1 Classifier head Linear(Hâ1Hâ 1) Loss BCEWithLogitsLoss pos_weight =#âneg/#âpos=\#neg/\#pos Optimizer AdamW learning rate η=2Ă10â3η=2\!Ă\!10^-3 weight decay λ=10â3λ=10^-3 Training batch size =64=64 max epochs =100=100 early stopping (macro-F1F_1, patience =3=3) title-aware validation split (0.150.15, seed =42=42) Table 9: Hyperparameters for Layer-LR (layer-weighted linear classifier). Component Configuration Input representation Layer stack XââNĂLĂHX ^NĂ LĂ H weighted aggregation Layer weighting Sparsemax( Ξ) Normalization â2 _2 per layer MLP head Linear(HâmHâ m), m=64m=64 GELU, Dropout(p=0.1p=0.1) Linear(mâ1mâ 1) Loss BCEWithLogitsLoss pos_weight =#âneg/#âpos=\#neg/\#pos Optimizer AdamW learning rate η=10â3η=10^-3 weight decay λ=10â3λ=10^-3 Training batch size =64=64 max epochs =100=100 early stopping (macro-F1F_1, patience =3=3) title-aware validation split (0.150.15, seed =42=42) Table 10: Hyperparameters for Layer-MLP (layer-weighted MLP classifier). Appendix B Additional Results B.1 Latent Space Analysis Here we report the full latent space plots from FigureË2 for all models, with FigureË4 for Llama-3.1-8B, FigureË5 for Mistral-7B-v0.1, and FigureË6 for Qwen2.5-7B. Figure 4: Per-layer PCA of hidden states for Llama-3.1-8B. Figure 5: Per-layer PCA of hidden states for Mistral-7B-v0.1. Figure 6: Per-layer PCA of hidden states for Qwen2.5-7B. B.2 Out-of-distribution and Generalisation Results This section reports additional experimental results that support the main findings in the paper, including ablation studies and out-of-domain generalisation performance. We omit bootstrap deviations for visual clarity. TableË11 shows the accuracy for the answer-decoy variant for SQuAD, where we can see that the MLP variant of the probe almost collapses to random accuracy, suggesting potential overfitting. TableË12 shows the results for the MLP probe on SQuAD and WebQuestions. Classifier Accuracy Layer-LR .724 Layer-MLP .541 Table 11: Answer-string decoy ablation on SQuAD using Llama-3.1-8B. While performance drops for all classifiers, the linear probe remains well above chance, whereas the MLP collapses to near-random accuracy, indicating reliance on entity repetition. Model SQuAD WebQ Llama-3.1-8B .972 .959 Mistral-7B-v0.1 .958 .948 Qwen-7B .998 .980 Table 12: Generalisation accuracy of Layer-MLP (FTG) on out-of-domain QA datasets. WebQ = WebQuestions. Comparison to Tighidet et al. (2024). In the following TablesË13, 14 and 15 we compare the per-group macro-F1F_1 of Tighidet et al. (2024)âs best-layer probe against the AttriWiki Layer-LR probe on FTG across all three models. Relation groups well-represented in encyclopedic Wikipedia text, such as Geographic and Media, transfer well, with AttriWiki consistently matching or exceeding Tighidet et al.âs performance. The Corporate-products-employment group is a consistent outlier. This reflects the fact that AttriWikiâs training data contains essentially no product-ownership facts. These results confirm that domain shift does affect the AttriWiki probe, but only in a targeted and predictable way: performance degrades specifically for themes absent from Wikipedia-style factual text. Relation group n Tighidet et al. (L10) AttriWiki Corporate 100 .333 ± .023 .369 ± .035 Geographic 920 .780 ± .014 .837 ± .012 Media 16 .329 ± .057 .800 ± .104 Table 13: Per-group macro-F1F_1 (mean ± bootstrap std, B=1000B=1000) for Mistral 7B. Tighidet et al. uses the best overall layer (L10). Relation group n Tighidet et al. (L11) AttriWiki Corporate 52 .411 ± .061 .541 ± .070 Geographic 530 .670 ± .021 .763 ± .019 Media 50 .566 ± .075 .878 ± .046 Occupy-position 16 .328 ± .058 1.000 ± .000 Play-instrument 20 .589 ± .113 .897 ± .070 Table 14: Per-group macro-F1F_1 (mean ± bootstrap std, B=1000B=1000) for Llama 3.1 8B. Tighidet et al. uses the best overall layer (L11). Relation group n Tighidet et al. (L22) AttriWiki Corporate 640 .762 ± .017 .633 ± .019 Geographic 2928 .754 ± .008 .812 ± .007 Hierarchy 46 .686 ± .073 .768 ± .062 Media 178 .798 ± .029 .886 ± .024 Occupy-position 136 .624 ± .042 .837 ± .031 Play-instrument 20 .329 ± .053 .303 ± .054 Religion 208 .794 ± .029 .840 ± .025 Table 15: Per-group macro-F1F_1 (mean ± bootstrap std, B=1000B=1000) for Qwen2.5 7B. Tighidet et al. uses the best overall layer (L22). Appendix C Prompts C.1 AttriWiki C.1.1 Entity selection. This prompt was designed to select up to three entities to prevent longer passages containing more entities from being overrepresented in the data. Given a Wikipedia passage and a list of mentioned entities, you are tasked to narrow down the list of entities to a maximum of three using the following criteria. 1. Remove entities that refer to a number, e.g., âfirstâ, âlastâ, âtwelveâ, âfourthâ. 2. Remove entities that refer to a very specific date, e.g., a range, months, days etc. Just a year is acceptable. 3. Remove entities that refer to any quantity. 4. The resulting entities should contain both well-known examples and lesser-known examples. The string of the remaining entities must be exactly the same. The passage: text The entities: entities Maximum of three selected entities: C.1.2 Knowledge testing. Given a Wikipedia passage and an entity, we present examples of the three knowledge tests. 1. Alice: I canât remember exactly who was the king of England in 1265 during the Battle of Evesham. I canât remember. Bob: Actually, I know. It was King Henry I. 2. Q: Who was the king of England during the Battle of Evesham in the 13th century? A: King Henry I 3. The Battle of Evesham ( 4 August 1265 ) was one of the two main battles of 13th century England âs Second Barons â War . It marked the defeat of Simon de Montfort, Earl of Leicester , and the rebellious barons by Prince Edward - later King Edward I - who led the forces of his father , King Henry I. 1. Dialogue You will receive a Wikipedia passage of an arbitrary topic and an entity that is mentioned somewhere within the passage. You will create a dialogue between Alice and Bob. Alice canât think of the name of [entity]. She describes it perfectly using the Wikipedia passage. Bob is all-knowing, and tells Alice the name of the entity. âą Alice is not allowed to say the name of the entity or part of the entity. âą Alice must only use information provided in the Wikipedia passage to describe said entity. âą Bob must say the exact name of the entity. Here is an example. Wikipedia passage: < The Battle of Evesham ( 4 August 1265 ) was one of the two main battles of 13th century England âs Second Barons â War . It marked the defeat of Simon de Montfort , Earl of Leicester , and the rebellious barons by Prince Edward - later King Edward I - who led the forces of his father , King Henry I . It took place on 4 August 1265 , near the town of Evesham , Worcestershire . > Entity: < King Henry I > Alice-Bob conversation: < Alice: I canât remember exactly who was the king of England in 1265 during the Battle of Evesham. I canât remember. Bob: Actually, I know. It was King Henry I. > Now it is your turn: Wikipedia passage: < passage > Entity: < entity > Alice-Bob conversation (ending with >): < 2. QA-style You will receive a Wikipedia passage of an arbitrary topic and an entity that is mentioned somewhere within the text. Like so: Wikipedia passage: < The Battle of Evesham ( 4 August 1265 ) was one of the two main battles of 13th century England âs Second Barons â War. It marked the defeat of Simon de Montfort, Earl of Leicester, and the rebellious barons by Prince Edward - later King Edward I - who led the forces of his father, King Henry I. It took place on 4 August 1265, near the town of Evesham, Worcestershire. > Entity: < King Henry I > Question-Answer pair: < Q: Who was the king of England during the Battle of Evesham in the 13th century? A: King Henry I > Follow the same pattern using a new Wikipedia passage. You generate a question (Q) to which the answer is the entity (A). Now it is your turn: Wikipedia passage: < passage > Entity: < entity > Question-Answer (ending with >): < Note that the trailing entity mentions are removed using an exact string match in a later phase. 3. Truncated passage. The truncated passage is an exact string match, removing the first mention of the entity in the Wikipedia passage and removing all text succeeding and including the entity. C.1.3 Entity removal. Remove every explicit mention of the entity and any variant (abbreviation, nickname, unambiguous pronoun) from the passage while preserving factual content and fluency. One illustrative example: Entity: United Kingdom Original passage: Winston Churchill was a British statesman, soldier, and writer who served as Prime Minister of the United Kingdom from 1940 to 1945 and again from 1951 to 1955. He led Britain to victory in the Second World War. Among the British public, he is widely considered the greatest Briton of all time. He was born to an aristocratic family in Oxfordshire, England. Rewritten passage: Winston Churchill was a statesman, soldier, and writer who served as Prime Minister from 1940 to 1945 and again from 1951 to 1955. He led the nation to victory in the Second World War. Among the public, he is widely considered one of the greatest leaders of all time. He was born to an aristocratic family in Oxfordshire. Constraints: âą Do not insert placeholder tokens such as "_____" or "[ENTITY]". âą Forbidden filler words: another, other, elsewhere, someplace, something, someone, notable, major, region, and global. âą If deleting the entity breaks a sentence, repair the grammar so the passage would pass professional copy-edit. âą Keep all dates, numbers, and named entities that do not refer to the target entity. âą If the entity acts as an adjective (e.g., âUK policyâ), rewrite only that phrase so the head noun remains (e.g., âthe policyâ). âą Meta-instruction: Do not mimic specific wording from the example; use whatever phrasing fits the original passage. Now your turn Entity: entity Original passage: < passage > Rewritten passage (end with >): < C.1.4 Appending completion sentence. Instruction: You are given a Wikipedia passage that contains an entity. Then you are given the entityâs name. Produce one additional sentence that naturally appends to the passage, reaffirming the entityâs identity. You may use one of the following formats (or a similarly natural variant): âą âThe [description] is called [entity].â âą âThe [description] is named [entity].â âą âLocals called it [entity].â âą âThey refer to it as [entity].â IMPORTANT: 1. You must use the exact entity name as providedâno alterations, changes in capitalization, or partial usage. 2. Your output should only be that single appended sentence. 3. The sentence should be generic and do a good job thoroughly introducing the entity. It should naturally lead up to naming the entity, so that if the entity were removed, a model would be likely to complete the sentence with that entity. One-Shot Example Passage: < Frankenstein is a gothic novel by Mary Shelley that was first published in 1818. The story follows a young scientist who creates a sapient creature through an unorthodox experiment, and it is often hailed as the first true work of science fiction. > Entity: < Mary Shelley > Output: < The author of the novel Frankenstein is named Mary Shelley. > Notice how: âą The entity is introduced exactly as given. âą The sentence flows naturally from the passage. Now it is your turn. Entity: < entity > Original Passage: < passage > Output (ending with >): < C.2 SQuAD and WebQuestions One-shot prompts for SQuAD and WebQ. Context: Frank Herbert was an American scienceâfiction author best known for his novel Dune. Example Q: Who wrote Dune? A: Frank Herbert Context: ⊠Q: ⊠A: ⊠Example Q: Who wrote âDuneâ? A: Frank Herbert Q: ⊠A: âŠ