Paper deep dive
Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
Sanket Badhe, Deep Shah, Nehal Kathrotia
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/21/2026, 2:00:59 AM
Summary
This paper presents a structured taxonomy and analysis of Long-Tail Knowledge (LTK) in Large Language Models (LLMs). It identifies LTK as low-frequency, domain-specific, cultural, and temporal knowledge that is poorly retained due to power-law distributions in training data. The authors analyze mechanisms of knowledge loss across pre-training optimization, representational constraints, post-training alignment, and inference dynamics. They also review technical interventions (e.g., RAG, model editing) and discuss sociotechnical implications including fairness, accountability, and bias.
Entities (10)
Relation Signals (8)
Large Language Models â suffersfrom â long-tail knowledge
confidence 95% · While scaling has improved average-case performance, persistent failures on low-frequency, domain-specific, cultural, and temporal knowledge remain poorly characterized.
long-tail knowledge â manifestsas â Hallucination
confidence 94% · In practice, failures on long-tail queries often manifest as hallucinated responses.
Pre-training Optimization â causes â Gradient Dilution
confidence 92% · The statistical nature of the loss function prioritizes high-frequency patterns. This leads to gradient dilution for rare facts
Retrieval-Augmented Generation â mitigates â long-tail knowledge
confidence 90% · Proposed mitigation strategies, such as retrieval-augmented generation... are typically introduced
Tokenization â causes â Tokenization-Induced Sparsity
confidence 89% · The discrete input layer imposes a structural bottleneck known as Tokenization-Induced Sparsity.
Large Language Models â exhibits â Global North Bias
confidence 88% · LLMs exhibit a cultural concept popularity bias, in which entities and events salient to the Global North are recalled with high accuracy
Model Editing â mitigates â long-tail knowledge
confidence 87% · Proposed mitigation strategies, such as retrieval-augmented generation and model editing
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most appearing infrequently. While scaling has improved average-case performance, persistent failures on low-frequency, domain-specific, cultural, and temporal knowledge remain poorly characterized. This paper develops a structured taxonomy and analysis of long-Tail Knowledge in large language models, synthesizing prior work across technical and sociotechnical perspectives. We introduce a structured analytical framework that synthesizes prior work across four complementary axes: how long-Tail Knowledge is defined, the mechanisms by which it is lost or distorted during training and inference, the technical interventions proposed to mitigate these failures, and the implications of these failures for fairness, accountability, transparency, and user trust. We further examine how existing evaluation practices obscure tail behavior and complicate accountability for rare but consequential failures. The paper concludes by identifying open challenges related to privacy, sustainability, and governance that constrain long-Tail Knowledge representation. Taken together, this paper provides a unifying conceptual framework for understanding how long-Tail Knowledge is defined, lost, evaluated, and manifested in deployed language model systems.
Tags
Links
- Source: https://arxiv.org/abs/2602.16201v1
- Canonical: https://arxiv.org/abs/2602.16201v1
Trouble viewing inline? Open PDF directly â
Full Text
100,340 characters extracted from source content.
Expand or collapse full text
Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications SANKET BADHE, Google, USA DEEP SHAH, Google, USA NEHAL KATHROTIA, Google, USA Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most appearing infrequently. While scaling has improved average-case performance, persistent failures on low-frequency, domain-specific, cultural, and temporal knowledge remain poorly characterized. This paper develops a structured taxonomy and analysis of long-Tail Knowledge in large language models, synthesizing prior work across technical and sociotechnical perspectives. We introduce a structured analytical framework that synthesizes prior work across four complementary axes: how long-Tail Knowledge is defined, the mechanisms by which it is lost or distorted during training and inference, the technical interventions proposed to mitigate these failures, and the implications of these failures for fairness, accountability, transparency, and user trust. We further examine how existing evaluation practices obscure tail behavior and complicate accountability for rare but consequential failures. The paper concludes by identifying open challenges related to privacy, sustainability, and governance that constrain long-Tail Knowledge representation. Taken together, this paper provides a unifying conceptual framework for understanding how long-Tail Knowledge is defined, lost, evaluated, and manifested in deployed language model systems. CCS Concepts:âą Computing methodologiesâ Machine learning approaches; Additional Key Words and Phrases: Long-Tail Knowledge, large language models, data sparsity, knowledge representation, evaluation practices, sociotechnical implications, Cultural concept popularity, Low-resource languages, Geopolitical bias, evaluation benchmarks ACM Reference Format: Sanket Badhe, Deep Shah, and Nehal Kathrotia. 2018. Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym âX). ACM, New York, NY, USA, 22 pages. https://doi.org/X.X 1 Introduction Large Language Models (LLMs) are increasingly used as interfaces for accessing and synthesizing information across a wide range of applications. Unlike traditional information retrieval systems that return documents based on explicit keyword matching, these models generate responses directly from internal parametric representations learned during training. This shift places greater importance on the statistical properties of the data used during pre-training. While scaling laws suggest that increasing model size and data volume yields predictable performance improvements [55], Authorsâ Contact Information: Sanket Badhe, sanketbadhe@google.com, Google, Mountain View, California, USA; Deep Shah, shahdeep@google.com, Google, Mountain View, California, USA; Nehal Kathrotia, nehalk@google.com, Google, Mountain View, California, USA. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM1 arXiv:2602.16201v1 [cs.CL] 18 Feb 2026 2Badhe et al. recent empirical work indicates that these gains are unevenly distributed across different types of knowledge. Improve- ments tend to concentrate on high-frequency facts and linguistic patterns, while knowledge that appears infrequently in training corpora exhibits weaker and less reliable retention [53]. The distribution of information in web-scale text corpora approximately follows a Zipfian or heavy-tailed power law, in which a small fraction of facts occur very frequently while the majority appear rarely [20]. We use the term Long-Tail Knowledge (LTK) to refer to this broad class of low-frequency information. This includes rare entities, niche domain expertise, low-resource languages, and historical or cultural facts with limited digital presence. There is no single agreed-upon definition of LTK in the literature. Prior work approaches this phenomenon from different perspectives, including frequency-based operationalizations in natural language processing and concerns about representational coverage and knowledge gaps in fairness and accountability research [10,82]. This work adopts an inclusive scope in order to synthesize these perspectives into a unified analytical framework. Across recent empirical studies, model performance is observed to degrade non-linearly as the frequency of the target knowledge decreases. Learning rare facts often requires disproportionately larger amounts of data compared to frequent ones [53]. At the same time, standard training objectives prioritize minimizing loss on high-frequency tokens, which biases learning toward common patterns and reduces gradient signal for infrequent knowledge [96]. In practice, failures on long-tail queries often manifest as hallucinated responses. Rather than expressing uncertainty, models frequently generate fluent but factually incorrect statements when queried about rare or weakly represented topics [6,48]. These behaviors have been documented across domains, including settings involving languages and histories that are sparsely represented in training data [59]. Research on LTK is distributed across multiple research communities. Analyses of memorization and training dynamics appear primarily in machine learning and systems venues [114]. Work examining representational coverage and downstream impacts is more common in ethics and fairness research [120]. Proposed mitigation strategies, such as retrieval-augmented generation and model editing, are typically introduced in natural language processing conferences [68,84]. Differences in terminology, benchmarks, and evaluation protocols across these communities make it difficult to compare findings or assess the scope of LTK limitations in a unified manner. This paper addresses this fragmentation by introducing a structured conceptual framework that organizes existing work on LTK across definitions, mechanisms, evaluation practices, and sociotechnical implications.: âą RQ1 (Taxonomy): What categories of LTK are defined or operationalized across different domains and tasks? âąRQ2 (Mechanisms): What training dynamics, architectural constraints, or inference-time processes are reported to contribute to the loss or fragile retention of rare knowledge? âąRQ3 (Interventions): What technical strategies have been proposed to improve LTK performance, and how are their tradeoffs characterized? âąRQ4 (Implications): What impacts of LTK gaps are reported in downstream applications, user interactions, or representational coverage? Existing work on rare or underrepresented knowledge is dispersed across domains, tasks, and evaluation paradigms, often using incompatible definitions and measurement practices. This fragmentation makes it difficult to compare findings, identify recurring failure modes, or assess the real-world significance of reported results. As a consequence, LTK loss is frequently treated as an isolated technical limitation rather than a systemic property of contemporary model development and evaluation pipelines. We address a foundational gap in the literature by synthesizing fragmented Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications3 Linguistic & Dialectical Low-resource languages, di- alects, code-switching gaps Ă Cultural & Geopolitical Global North bias, concept popularity, Western-centric norms  Specialized & Domain Rare diseases, local statutes, archaic technical traditions 3 Temporal & Dynamic New tail (post-cutoff), Forgot- ten past (historical sparsity) Ă Pre-training Optimization Gradient Dilution· Objective Mis- match· Capacity Allocation Bias ÂĄ Representational Limits Tokenization-Induced Spar- sity· Distributed Repre- sentation Interference > Post-Training Alignment Alignment and Instruction- Tuning Compression 8 Inference Dynamics Decoding-Time Probability Trunca- tion· Contextual Retrieval Failure Ă Data-Centric Importance weighting, Syn- thetic data, Curriculum learning Ă” Capacity & Architecture Mixture-of-Experts (MoE), Neuro-symbolic memories ĂĄ Retrieval Methods RAG, Adaptive retrieval, Knowledge graphs _ Model Editing ROME, MEMIT, Causal tracing Human-in-the-Loop Domain experts, Red teaming, RLHF € Epistemic Visibility Exclusion of minority perspectives 6 Unequal Reliability High error rates for non-English speakers . Accountability Blame for stochastic tail failures u Trust & Calibration Overconfidence in hallucinations ± Feedback Loops Model collapse, era- sure of rare knowledge Âč Transparency CoT rationalizes outputs; mis- leading post-hoc reasoning « TAXONOMYMECHANISMSINTERVENTIONSIMPLICATIONS Fig. 1. Overview of Long-Tail Knowledge in LLMs: Taxonomy, Mechanisms, Interventions, and Implications. empirical evidence on LTK in LLMs. This paper is the first to systematically synthesize these findings across definitions, mechanisms, mitigation strategies, and sociotechnical implicaxtions. By organizing technical analyses of training dynamics alongside studies of downstream behavior and deployment contexts, this paper provides a structured account of how statistical sparsity in training data relates to observed limitations in model reliability, coverage, and consistency. This synthesis is necessary to support more rigorous evaluation practices, to clarify the limits of current mitigation strategies, and to inform responsible deployment decisions in settings where rare or specialized knowledge carries outsized social impact. Manuscript submitted to ACM 4Badhe et al. 2 The Taxonomy of the Long-Tail To address the LTK problem effectively, we must first categorize the types of information that systematically reside in regions of low empirical support within the training distribution. Across web-scale corpora, this sparsity is commonly measured through token, entity, or document frequency, and it follows a heavy-tailed distribution. Frequency-based rarity is therefore a necessary but insufficient condition for LTK. This section treats frequency as a shared underlying axis and introduces a taxonomy that categorizes LTK according to the properties that shape how sparsity manifests in practice. Our analysis of the literature reveals four primary ontological dimensions of sparsity: linguistic, cultural, domain- specific, and temporal. These dimensions are not mutually exclusive and frequently interact. They collectively subsume the dominant ways in which LTK is operationalized and evaluated in empirical studies, including cases of compositional rarity and procedural or tacit knowledge that are weakly represented in declarative text corpora. 2.1 Linguistic and Dialectical Sparsity The first dimension of the long tail is linguistic. While scaling laws hold for high-resource languages such as English, model performance degrades sharply for low-resource languages and non-standard dialects [98]. This disparity is partly driven by tokenization artifacts. Tokenizers optimized for English frequently fragment words from low-resource languages into long sequences of subword units, diluting semantic coherence and weakening statistical learning signals [95]. Research indicates that LLMs exhibit severe performance degradation on low-resource African languages, such as Amharic and Sepedi, where data scarcity is compounded by morphological complexity that standard tokenizers fail to capture [2]. Recent work highlights a distinct safety and reliability gap in this linguistic tail. Shen et al. [106] identify two failure modes for low-resource languages: a harmfulness curse, in which models are more likely to generate unsafe content, and a relevance curse, in which instruction-following and task adherence collapse. Linguistic sparsity also extends to dialectal variation within high-resource languages. The ReDial benchmark demonstrates that LLMs exhibit brittleness when processing African American Vernacular English and often fail on reasoning tasks that they solve reliably in standardized English [73,74]. Code-switching presents a further challenge. Models frequently default to the dominant matrix language or fail to preserve syntactic consistency across language boundaries, reflecting compositional rarity rather than absence of component knowledge [41,125]. Beyond dialects, The Script-Gap illustrates a critical safety blind spot in LLM-based systems: models exhibit consistent performance degradation when processing romanized messages compared to native scriptsâa common practice in digital communication for Indian languagesâeven when the models appear to correctly infer the userâs semantic intent. [60]. 2.2 Cultural and Geopolitical Peripheries The second dimension concerns cultural grounding and geopolitical representation. LLMs exhibit a cultural concept popularity bias, in which entities and events salient to the Global North are recalled with high accuracy, while analogous concepts from underrepresented regions are omitted or hallucinated [51]. The CPopQA benchmark shows that models accurately rank holidays in the United States but perform substantially worse for culturally significant events in regions with lower digital visibility, including Sub-Saharan Africa and Southeast Asia [42, 51]. This imbalance induces a Western-centric internal world model. Studies using benchmarks such as CANDLE and CAMeL demonstrate that LLMs struggle to reason about non-Western entities and frequently conflate distinct indigenous Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications5 histories or project Western norms onto local contexts [88,89]. When evaluating subjective global opinions, research indicates that LLMsâespecially those fine-tuned with human feedbackâsystematically favor perspectives from Western, Educated, Industrialized, Rich, and Democratic (WEIRD) populations. Rather than being broadly representative, these models exhibit a substantial shift toward the views of liberal, educated, and wealthy demographics, while the opinions of many other indigenious groups remain significantly underrepresented. [29,102]. The challenge is not merely the absence of data but the Cultural Alignment problem, where the modelâs interpretative frame is incapable of processing the nuances of pluralistic value systems without explicit, resource-intensive adaptation [123]. 2.3 Specialized and Domain-Specific Tails The third dimension encompasses specialized and domain-specific knowledge that appears infrequently in general- purpose training corpora. This category includes both declarative facts and procedural or tacit knowledge, such as diagnostic workflows, legal procedures, and technical practices that are sparsely documented. In the medical domain, LLMs perform well on common conditions but exhibit severe deficits when reasoning about rare diseases [5]. The MIMIC-RD study reports that open-source models fail to predict approximately 60% of observed rare diseases even when relevant phenotypic information is explicitly provided in the prompt [5]. A parallel pattern is observed in the legal domain. While LLMs may demonstrate strong performance on standardized bar exam questions, they frequently hallucinate when queried about local statutes, procedural rules of lower courts, or infrequently cited case law [25,26]. This contrast highlights a persistent gap between transferable reasoning skills, which generalize from high-frequency patterns, and the retrieval or application of low-frequency, context- specific knowledge [53]. Domain-specific tails also include archaic craftsmanship and specialized technical traditions, Scholars are increasingly applying LLMs to digitize hidden technical knowledge and archaic craftsmanship preserved in handwritten manuscripts. While recent work demonstrates that LLMs achieve near-human levels of accuracy when used for post-correction of transcriptions, evaluation remains fragile due to systematic temporal biases [45,54]. For instance, Levchenko identifies an over-historicization phenomenon where models insert archaic characters from incorrect historical periods rather than preserving original period-specific orthography [67]. 2.4 Temporal and Dynamic Knowledge The final dimension is temporal. LTK is not purely static but evolves with time and model training cutoffs. We distinguish between two temporal tails: the New Tail and the Forgotten Past. The New Tail comprises events and facts that emerge after a modelâs training data cutoff. In these cases, models exhibit temporal misalignment, often hallucinating outdated information or failing to recognize recent changes [56, 65, 79]. The Forgotten Past refers to historical knowledge with low contemporary digital presence. LLMs trained primarily on modern web text exhibit a recency bias, leading to degraded performance on historical facts that are infrequently discussed in current discourse [27,47]. This degradation disproportionately affects regions and communities whose historical records were digitized late or unevenly, reinforcing existing gaps in representational coverage [59]. 3 Mechanisms of Knowledge Loss The failure of LLMs to retain and retrieve LTK is not the result of a single technical bottleneck. It is a compound failure that cascades through every stage of the model pipeline. To structure our analysis of these failures, we categorize the mechanisms of knowledge loss into four distinct architectural levels: Pre-training Optimization, Representational Constraints, Post-Training Alignment, and Inference Dynamics. Manuscript submitted to ACM 6Badhe et al. Summary of Architectural Failure Modes: âąLevel 1: Pre-training Optimization. The statistical nature of the loss function prioritizes high-frequency patterns. This leads to gradient dilution for rare facts and a capacity allocation bias that favors generic heuristics over specific details. âąLevel 2: Representational Constraints. The discrete nature of tokenization and the phenomenon of superposi- tion create structural barriers. These barriers prevent the coherent encoding of rare entities and cause distributed interference in the parameter space. âąLevel 3: Post-Training Alignment. Reinforcement learning interventions introduce an alignment tax. This encourages models to hedge or abstain from answering tail queries to minimize the risk of hallucination or unsafety. âąLevel 4: Inference Dynamics. Standard decoding strategies truncate the low-probability tail of the distribution. This systematically filters out correct but rare tokens during generation. 3.1 Level 1: Pre-training Optimization and Objectives The primary driver of tail loss is the optimization landscape of the pre-training phase. We identify three specific mechanisms that degrade the retention of low-frequency data. 3.1.1 Gradient Dilution. In standard autoregressive training, the model minimizes the negative log-likelihood of the next token. This objective function is inherently frequency-dependent. For a tail factíwith probabilityí(í)in the training corpus, the expected gradient updateE[âL]is proportional toí(í). Consequently, high-frequency concepts generate consistent and high-magnitude gradient signals that stabilize the weights of the model. In contrast, long-tail facts produce sparse and high-variance updates. These updates are frequently overwritten by the noise of subsequent updates from dominant data [96]. This phenomenon is known as Gradient Starvation. It results in a model that minimizes loss on the head of the distribution while leaving tail features under-learned [22, 53]. 3.1.2 Objective Mismatch and the Hallucination Nexus. A fundamental misalignment exists between the Maximum Likelihood Estimation objective and the goal of factual retention. The objective incentivizes the model to minimize perplexity rather than maximize factual accuracy. For a rare fact, the model can often achieve a lower loss by predicting a high-probability generic continuation rather than the low-probability true entity [82]. This creates a pressure to learn heuristic associations. For example, the model may predict that a person born in 18th century France is a peasant rather than a specific individual. This statistical correlation drives the Hallucination Nexus. There is a strong inverse correlation between the frequency of a fact in the training data and the hallucination rate of the model. The low signal-to-noise ratio in tail activations triggers the generation of high-frequency priors that are statistically likely but factually incorrect [6, 49, 127]. 3.1.3 Capacity Allocation Bias. Transformers exhibit a Simplicity Bias where the modelâs inductive bias and optimization objectives preferentially favor low-sensitivity functions or simple rules over the memorization of high-complexity sparse data. Bhattamishra et al. [11] demonstrate that this bias leads Transformers to converge to simple hypotheses, which aids generalization but can impair the learning of complex tail information. In the context of LLMs, Barron & White [9] show that limited capacity forces models to prioritize rule extraction over factual memorization, whereas larger models utilize their parameter budget to store factual mappings at the expense of rule-based extrapolation. Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications7 Tirumala et al. [114] further note that while models memorize the head (nouns and numbers) first, larger scales actually mitigate the loss of LTK rather than compressing it into noisy representations. 3.2 Level 2: Representational Constraints Even when gradients are sufficient, structural constraints in the architecture of the Transformer can prevent the accurate storage of tail knowledge. 3.2.1 Tokenization-Induced Sparsity. The discrete input layer imposes a structural bottleneck known as Tokenization- Induced Sparsity. Standard Byte-Pair Encoding or Unigram tokenizers are optimized to minimize the sequence length of the training corpus. This effectively compresses common English words into single tokens. Conversely, rare entities and low-resource language terms are fragmented into long sequences of sub-word units [95,104]. Wang et al. [117] demonstrate that this fragmentation often leads to incorrect tokenization that is misaligned with human comprehension, hindering the modelâs ability to understand input precisely and causing nonsensical responses. While the attention mechanism can theoretically learn the compositionality of these fragments, Kudo [62] identifies that treating these varied sequences as completely different inputs creates a spurious ambiguity that can degrade the modelâs robustness and increase error rates. 3.2.2 Distributed Representation Interference. LLMs operate in a regime of superposition where the model represents more features than it has dimensions by storing them non-orthogonally. While this allows for high capacity, it introduces Distributed Representation Interference [77]. Frequent patterns occupy the principal directions of the representation space. This forces rare facts to be encoded in the noisy residuals [30]. Recent theoretical work by Liu et al. [78] demonstrates that in this strong superposition regime, the interference penalty is disproportionately borne by low- frequency features, which are squeezed into representations with high overlaps. These features are easily overwritten by the crosstalk from high-frequency features that share the same polysemantic neurons [12]. 3.3 Level 3: Post-Training Alignment The fine-tuning phase introduces distinct failure modes that are not present in the base model. 3.3.1 Alignment and Instruction-Tuning Compression. Post-training interventions such as Reinforcement Learning from Human Feedback introduce an Alignment Tax on tail knowledge. RLHF typically optimizes for helpfulness and safety using a reward model trained on a small and high-quality dataset that is heavily biased towards the head of the distribution [18,92,122]. This process induces a form of mode collapse. The model learns to hedge or abstain from answering obscure questions to avoid the penalty of being incorrect. Alternatively, it converges to a generic response style that suppresses specific but rare details [7,61]. Empirical studies show that heavy instruction tuning can degrade the calibration of the model on long-tail facts. The model over-generalizes safety refusals to obscure but harmless queries [76, 94]. 3.4 Level 4: Inference Dynamics Finally, even if knowledge is retained in the weights, it may be inaccessible during the generation phase. 3.4.1 Decoding-Time Probability Truncation. Standard decoding strategies often prune tail knowledge during inference. Nucleus sampling (top-p) and top-k sampling explicitly truncate the tail of the probability distribution to ensure coherence and diversity [43]. However, the correct token for a long-tail fact often resides in this truncated region of Manuscript submitted to ACM 8Badhe et al. low probability. By strictly filtering for high-probability tokens, these decoding algorithms systematically silence the correct but rare answers [119]. This forces the model to select a more common but incorrect alternative from the head of the distribution [83]. 3.4.2 Contextual Retrieval Failure. Rare facts often have weak attention keys that are easily overshadowed by stronger context signals. This leads to Contextual Retrieval Failure. Research indicates that activating these fragile memories requires prompts of high specificity that mimic the exact context seen during training [82]. When the prompt deviates even slightly, the internal retrieval mechanism fails to attend to the relevant parameter subspace. This leads to a silently known fact that is inaccessible during generation [68, 80]. 4 Survey of Mitigation Strategies We now examine the strategies proposed to address LTK in LLMs. This section categorizes interventions by their point of operation within the model lifecycle: data curation, architectural design, inference-time retrieval, parameter editing, and human oversight. We document the operational mechanisms of each approach alongside their empirically observed trade-offs and limitations. 4.1 Data-Centric Interventions The most direct approach involves modifying the training data distribution to amplify the signal of rare concepts. One prominent strategy is importance weighting. This method adjusts the loss function to penalize errors on tail examples more heavily than on head examples. Techniques such as GradTail dynamically increase gradient weights for rare tokens which forces the model to allocate more capacity to learning them [22,96]. While effective in controlled settings, this approach risks catastrophic forgetting of the head as the optimization landscape becomes distorted. Another data-centric method is the generation of synthetic data to populate the tail. Frameworks such as LLM- AutoDA leverage stronger models to generate diverse and high-quality training examples for under-represented classes. This effectively flattens the Zipfian distribution artificially [118]. Similarly, curriculum learning strategies demonstrate that training on high-quality synthetic textbooks can improve coverage of niche domains compared to raw web scrapes [38]. However, reliance on synthetic data introduces the risk of model collapse. Recursive training on generated data leads to a loss of variance and the erasure of the true tail over time [107]. 4.2 Capacity and Architectural Allocation Architectural interventions seek to scale the memory capacity of the model without incurring prohibitive computational costs. The Mixture-of-Experts paradigm replaces dense layers with a set of specialized expert networks that are activated sparsely via a routing mechanism. Models like Mixtral 8x7B demonstrate that this approach allows for a massive increase in total parametersâproviding 47B total parameters while only activating 13B per tokenâthereby creating vast potential memory slots for long-tail facts while maintaining efficient inference [31,50]. However, recent analysis reveals that routing mechanisms often collapse to a few super experts that handle the majority of tokens. This leaves the tail-specific experts under-utilized and undertrained [71, 112]. Other architectural proposals include the integration of explicit memory layers or k-nearest neighbor components directly into the transformer blocks. These neuro-symbolic hybrids attempt to decouple factual storage from reasoning which allows the model to query an internal key-value store for rare facts [58]. While promising, these methods often Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications9 suffer from high latency and the challenge of keeping the memory store synchronized with the evolving weights of the main network [82]. 4.3 Retrieval and Externalization-Based Methods Retrieval-Augmented Generation externalizes LTK by moving the burden of memorization from the weights of the model to a searchable index. By retrieving relevant documents at inference time, RAG allows the model to access a vast and updateable corpus of tail facts without retraining [39,69]. This approach is particularly effective for new tail knowledge that emerges after the training cutoff [116]. However, RAG has significant limitations. Its effectiveness depends entirely on the quality of the retrieval system which itself may suffer from bias towards head documents [70]. Furthermore, standard RAG pipelines apply retrieval indiscriminately. This often hurts performance on queries where the internal parametric memory of the model is sufficient. To address this, adaptive retrieval methods use metrics such as Generative Expected Calibration Error to detect when a query falls into the long tail and trigger retrieval only when necessary [70,82]. Despite these refinements, RAG remains limited by the context window bottleneck and the ability of the model to reason over conflicting retrieved information [33, 105, 121]. 4.4 Model Editing and Localized Updates For specific errors in the tail, model editing techniques offer a way to surgically modify parameters. Methods such as ROME and MEMIT use causal tracing to identify the specific neurons encoding a fact and update them via a closed-form solution [84, 85]. This allows for the correction of outdated or incorrect tail facts without the cost of full fine-tuning. While powerful for single-point edits, these methods face significant scalability challenges. Research indicates that sequential edits can accumulate ripple effects where updates to one fact inadvertently damage the representations of related entities or degrade the general reasoning capabilities of the model [23,37]. Furthermore, lifelong editing eventually leads to a degradation of the internal coherence of the model which renders it unusable after a certain number of updates [44, 124]. 4.5 Human-in-the-Loop and Institutional Interventions Sociotechnical interventions focus on integrating human oversight to verify and correct tail outputs. Human-in-the-Loop frameworks deploy domain experts to audit model outputs in high-stakes fields such as law and medicine. This process reduce hallucinations and increase safety before they reach the end user [4]. Reinforcement Learning from Human Feedback (RLHF) attempts to align the model with human preferences, suffer from algorithmic bias due to optimization regularization â a formal mechanism by which minority preferences are effectively disregarded as preference collapse while dominant human preferences get amplified during alignment [18,122]. Barman et al. [35] argue that this reliance on potentially homogeneous evaluator groups can lead to sycophancy, where the model mirrors the one-sided opinions or ideologies of its tutors, often at the expense of scientific adequacy or nuanced tail knowledge. Institutional interventions also include the development of red teaming protocols specifically designed to probe for long-tail failures. However, these human-centric approaches are inherently unscalable. They rely on the availability of experts who possess the rare knowledge in question. This creates a bottleneck that prevents broad coverage of the long tail [32]. To mitigate these constraints, some suggest shifting toward pluralistic feedback panels that represent a wider spectrum of epistemic standpoints, aiming to achieve interactive objectivity rather than simple preference averaging [35]. Manuscript submitted to ACM 10Badhe et al. 5 Sociotechnical Implications The technical issues discussed in the previous section lead to specific problems when LLMs are used in real-world systems. This section looks at the evidence to show how failures in LTK create systemic risks. We organize this analysis into seven parts, ranging from whose knowledge is included to how we audit these systems. 5.1 Epistemic Visibility and Knowledge Inclusion Studies show that LLMs filter information in a way that favors dominant ideas. Prior work on computational mediation argues that large-scale automated systems do not merely retrieve facts but actively shape epistemic authority by privileging dominant representations [34]. Recent analyses of LLMs show that this effect persists in generative systems, where scale and fluency create an appearance of objectivity while systematically omitting minority and low-frequency perspectives [10,97]. This exclusion can be measured. Kandpal et al. [53] show that the ability of a model to answer a question depends directly on how many times that fact appears in the training data. This makes information with low digital presence almost invisible to the model. This creates a cycle where the model boosts already popular ideas while pushing down local knowledge that does not have enough data to be memorized. 5.2 Unequal Reliability Across User Populations The reliability of these models depends heavily on who is using them. Research shows a digital divide where speakers of languages with less data face more errors. A study by the Stanford HAI policy group found that LLMs trained mostly on English data have higher error rates and are less safe when used in languages like Swahili or Burmese [98]. This gap also extends to cultural and regional knowledge. Myung et al. [87] utilized the BLEnD benchmark to evaluate everyday knowledge across diverse cultures and found a massive disparity: while models achieved high accuracy on questions about United States culture, performance dropped for Ethiopian culture questions prompted in Amharic. Similarly, users in the Global South seeking information about local entities often receive answers grounded in Western perspectives or pure hallucination, as demonstrated by benchmarks like CAMeL which reveal significant alignment failures for Arab cultural contexts compared to Western ones [3, 51]. 5.3 Accountability Ambiguity in AI-Assisted Decision-Making Long-tail failures complicate the attribution of responsibility in AI-assisted workflows. When a traditional system fails due to a clear logic error, the fault is often assignable. However, failures rooted in missing knowledge occupy a gray zone. Dahl et al. [25] document how legal professionals have faced sanctions for submitting hallucinated citations generated by LLMs. These failures often arise not because the model breaks in a traditional sense but because it confidently fills knowledge gaps with statistically plausible but non-existent cases [81]. This creates an accountability vacuum where developers may claim the model is working as intended by minimizing perplexity while users are unable to verify the obscure information that led to the error [36,99]. The result is a system where responsibility for tail failures is difficult to locate as the error stems from the probabilistic nature of the model itself rather than a specific bug [32]. 5.4 Trust Calibration and Overconfidence Effects A critical sociotechnical risk is the misalignment between model confidence and factual accuracy in the tail. LLMs do not exhibit epistemic humility and often present hallucinations with the same rhetorical fluency as verified facts. Zhang et al. [126] propose the Log-Linear Law of Hallucination which predicts that error rates increase linearly as Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications11 knowledge popularity decreases yet model confidence does not degrade proportionally. This calibration failure leads to user over-reliance. Studies on faithfulness show that users frequently accept coherent but incorrect answers in specialized domains because the tone of the model mimics authoritative expert discourse [75,115]. This is particularly dangerous in long-tail scenarios where users turn to the model precisely because they lack the knowledge to verify the answer themselves [49, 82]. 5.5 Feedback Loops and Knowledge Marginalization Over Time The integration of LLM outputs into the web creates recursive feedback loops that can degrade the quality of public knowledge. Shumailov et al. [108] define model collapse as a degenerative process where models trained on synthetic data progressively lose variance and converge on the mode of the distribution. This process disproportionately erases the tails of the distribution. As synthetic content floods the web, rare facts are replaced by generic approximations. This makes them even less likely to be learned by future iterations of the model. Empirical work suggests that even with verification, the accumulation of synthetic data can lead to the permanent loss of low-frequency concepts from the digital record and reinforces the hegemony of the head over time [1]. 5.6 Limits of Transparency and Explainability for Rare Knowledge Explainability mechanisms often fail to faithfully reflect internal model states. Turpin et al. [115] demonstrate that Chain-of-Thought (CoT) reasoning can be systematically unfaithful: models may generate plausible reasoning chains that rationalize an output rather than reveal the causal basis of the decision. While this limitation applies broadly, it is particularly consequential for long-tail knowledge. In head regimes, incorrect reasoning can often be detected through redundancy, consensus, or external verification. In contrast, long-tail failures lack such safeguards, making post-hoc explanations especially misleading when the model has no robust internal representation of the queried knowledge. Prior work shows that explanation faithfulness and uncertainty signaling are limited, providing little assurance that a correct-looking explanation corresponds to genuine knowledge rather than fragile inference [52, 64]. 6 Critical Analysis of Evaluation and Accountability The structural invisibility of LTK is reinforced by the current paradigm of model evaluation. While benchmarks serve as the primary mechanism for establishing trust and verifying capabilities, our analysis reveals that they often systematically exclude the long tail, thereby producing an incomplete picture of model reliability. This section examines how these measurement gaps distort accountability claims and complicate the safe deployment of LLMs in high-stakes domains. 6.1 Scope Limitations of Benchmark-Based Evaluation Standard benchmarks such as MMLU (Massive Multitask Language Understanding) and BIG-bench have become the de facto standards for assessing model quality. However, empirical analysis suggests that these suites are heavily biased towards head knowledge. The construction of these benchmarks often relies on scraping questions from standardized human exams or Wikipedia-derived trivia, sources that naturally reflect the dominant information distribution of the web [40,110]. Consequently, facts that reside in the statistical tail such as local histories, indigenous botanical knowledge, or niche technical specifications are functionally absent from the evaluation set [28, 72]. Manuscript submitted to ACM 12Badhe et al. 6.2 Metric Aggregation and the Visibility of Tail Failures The prevailing practice of reporting model performance via single aggregate metrics actively conceals long-tail failures. Statistical aggregation operates as a smoothing function that drowns out the signal of rare errors. Oakden-Rayner et al. [91] formally define this phenomenon as hidden stratification, demonstrating that a model can achieve state-of-the-art aggregate performance while consistently failing on clinically meaningful minority subsets. For instance, in medical imaging, a high overall accuracy can mask the modelâs consistent inability to detect a rare but aggressive cancer subtype. To counteract this, researchers have proposed slice-based evaluation, which partitions test sets by entity frequency, dialect, or demographic attribute. Studies utilizing these granular metrics reveal that worst-group accuracy often diverges sharply from average performance [8,16]. However, standard industry leaderboards rarely report these disaggregated statistics, privileging a monolithic view of performance that is calibrated to the statistical majority [13]. 6.3 Accountability Claims Under Partial Measurement The gap between what is measured and what is claimed creates a crisis of accountability. When developers claim a model is safe or reliable based on head-heavy benchmarks, they are making a universal claim supported only by partial evidence. This disconnect complicates the assignment of responsibility when failures occur in the tail. Researcher characterize this as the Fallacy of AI Functionality, where internal validity such as benchmark performance is mistakenly equated with external validity i.e. real-world reliability [99]. Empirical audits show that the opacity of closed model evaluations exacerbates this issue. Without access to the specific prompts and test sets used by developers, independent auditors cannot verify whether a claimed capability extends to the long tail or is strictly limited to the tested head examples [36]. This lack of auditability allows developers to disclaim liability for unforeseen errors, even when those errors are predictable consequences of frequency-based learning dynamics [14, 100]. 6.4 Evaluation Practices in Regulated and High-Stakes Domains In domains such as law and medicine, the stakes of long-tail failure are non-negotiable. However, current evaluation practices often fail to align with professional standards of care. In the legal field, models are often evaluated on their ability to pass the Uniform Bar Exam (UBE). While GPT-4 has demonstrated a 90th percentile score on the UBE, this metric assesses general legal reasoning, not the retrieval of specific, low-frequency case law [57]. Deployment audits reveal that these same models frequently hallucinate citations when dealing with obscure or local jurisdictions, a failure mode not captured by the UBE but critical for malpractice liability [26]. Similarly, in healthcare, the focus on USMLE (United States Medical Licensing Examination) performance prioritizes textbook knowledge over clinical judgment in rare scenarios. Systematic reviews of medical LLMs find that while models excel at answering multiple-choice questions, they struggle with assessments that simulate the ambiguity and sparsity of real-world rare disease diagnosis [21,90,109]. This mismatch indicates that regulatory approval pathways based on standard benchmarks may inadvertently authorize systems that are unsafe for long-tail patient populations [120]. 7 Open Challenges and Future Directions The transition of LLMs from experimental artifacts to core information infrastructure necessitates a rigorous accounting of their limitations. Our paper indicates that the LTK problem is not merely a technical error term to be minimized but Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications13 a structural feature of learning from heavy-tailed distributions. This section outlines critical unresolved challenges and emerging future directions. 7.1 Measuring and Evaluating Long-Tail Knowledge in Practice A primary obstacle to progress is the lack of a stable operational definition for LTK. Current definitions largely rely on frequency counts in pre-training corpora such as Common Crawl, which fail to capture domain-specific importance or informational density [53]. What constitutes a rare fact in a general web corpus may be foundational in specialized domains such as materials science or contract law. When benchmarks rely on aggregate metrics they compress heterogeneous failure modes into a single number, obscuring performance on the tail. Raji et al. [99] describe this as the fallacy of AI functionality, where aggregate success is mistaken for universal competence. Evaluation validity is further compromised by static benchmarks. As models scale, overlap between test sets and training data increases, blurring the line between generalization and memorization [46,101]. HELM-style evaluations demonstrate that performance is highly sensitive to prompt formulation, with variance most pronounced for rare knowledge lacking robust internal representations [72,86]. Consequently, the field lacks standardized methods to certify long-tail competence without costly manual audits [36]. These findings suggest a need for tail-aware evaluation practices that explicitly stratify performance by frequency, domain criticality, or representational importance. Future work must reconsider whether benchmark-centric evaluation is sufficient for systems deployed in settings where tail failures dominate risk. 7.2 Accountability Under Knowledge Uncertainty The probabilistic nature of long-tail failures complicates existing accountability frameworks. Unlike conventional software errors, LLM failures in the tail are stochastic and context-dependent. A model may retrieve a rare fact correctly in one context but hallucinate in another following minor prompt perturbations [115]. This instability diffuses responsibility across data curation, model architecture, deployment choices, and user prompting strategies. Failures that occur only in sparse regions of the input space may evade regulatory scrutiny despite being structurally predictable [19, 24]. A key future direction is the development of accountability frameworks that explicitly recognize knowledge uncer- tainty. Rather than attributing errors solely post hoc, governance approaches may need to incorporate documented coverage limits, scoped use policies, or uncertainty-aware deployment standards. 7.3 Privacy and Sustainability Constraints on Long-Tail Learning Efforts to improve long-tail performance face structural constraints related to privacy and sustainability. Long-tail examples are disproportionately vulnerable to memorization and extraction attacks. Carlini et al. [17] show that rare training examples are most susceptible to verbatim leakage. Techniques such as up-sampling or aggressive deduplication can improve rare fact retention but simultaneously increase the risk of exposing sensitive personally identifiable information [15, 66]. Scaling laws imply that capturing increasingly rare knowledge requires exponentially larger models and datasets. The environmental cost of training large dense models is well documented [93,111]. Although sparse architectures such as Mixture-of-Experts offer efficiency gains, the marginal cost of learning the next rarest fact continues to rise. This raises fundamental questions about the economic and ecological viability of encoding the full breadth of human knowledge into parametric models [10, 113]. Manuscript submitted to ACM 14Badhe et al. Future research must grapple with whether long-tail robustness should be pursued through parametric memoriza- tion at all, or whether responsibility for rare knowledge should increasingly shift to external, governed knowledge infrastructures. 7.4 Long-Tail Knowledge as a Dynamic Sociotechnical Phenomenon The long tail is not static. It evolves through interaction between models, users, and information ecosystems. Deployment creates feedback loops that reshape knowledge distributions. Large-scale use of model-generated content reduces informational diversity, a process that can disproportionately erodes rare knowledge [107]. If future models are trained on AI-generated data, omissions and distortions of rare facts may become permanent. Reliance on retrieval-augmented systems introduces additional dependencies. The availability of LTK becomes contingent on external search indices, which are themselves shaped by commercial incentives and algorithmic bias [63]. Changes in indexing or access policies can instantaneously alter a modelâs effective knowledge base. This shifts the locus of inquiry from model parameters to the broader sociotechnical ecosystem in which models operate [103]. These dynamics suggest that preserving LTK is not solely a modeling problem but an ongoing governance challenge. Future work must treat long-tail robustness as a property of evolving sociotechnical systems rather than a static optimization target. 8 Conclusion This paper studied the emerging body of research on LTK in LLMs, with the goal of clarifying how rare, sparse, and underrepresented information is defined, evaluated, and affected across the model lifecycle. Rather than treating long-tail failures as isolated anomalies, the studied literature consistently indicates that such failures arise from systematic interactions between data distributions, training objectives, representational capacity, and inference-time constraints. As LLMs are increasingly deployed as general-purpose knowledge systems, these limitations have become both technically consequential and sociotechnically salient. We organized prior work along four complementary dimensions. First, we synthesized existing definitions of LTK into a taxonomy that captures linguistic, cultural, domain-specific, and temporal forms of sparsity, extending beyond purely frequency-based characterizations. Second, we reviewed mechanisms through which LTK is lost or weakly retained, including gradient dilution, representational interference, tokenization effects, and post-training compression. Third, we surveyed technical interventions proposed to mitigate long-tail failures, such as data-centric rebalancing, retrieval-augmented generation, architectural modularization, and model editing, highlighting their empirical scope and known tradeoffs. Finally, we examined sociotechnical implications, documenting how LTK gaps manifest as unequal system reliability, accountability ambiguity, trust miscalibration, and long-term feedback loops that further marginalize already underrepresented knowledge. A key takeaway from this paper is that LTK failures cannot be understood or addressed in isolation at a single layer of the system. Improvements in model scale or architecture do not fully compensate for sparsity in training data, while inference-time augmentation alone does not resolve deeper issues of representation and evaluation. Conversely, many sociotechnical risks associated with epistemic exclusion and uneven system performance are directly linked to technical design choices that prioritize average-case accuracy over coverage of the tail. This interdependence suggests that assessments of LLM reliability, fairness, and accountability must explicitly account for long-tail behavior rather than treating it as residual error. Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications15 This paper contributes a structured synthesis that maps how disparate strands of research relate to one another and where key gaps remain. By clarifying definitions, mechanisms, evaluation practices, and implications within a unified analytical framework, we aim to support more systematic empirical studies and more transparent assessments of LLM capabilities and limitations. As LLMs continue to mediate access to information across languages, domains, and communities, understanding what these systems systematically fail to know is as important as measuring what they perform well on. LTK will remain a central challenge for both technical robustness and sociotechnical responsibility, and future work will benefit from evaluation and deployment practices that make these limitations visible rather than implicit. 9 Generative AI Usage Statement All study design, literature review, synthesis, and writing were conducted by the authors. Generative AI tools (Gemini) were used only for grammar checking and proofreading during the final polishing of the manuscript. No generative AI system was used to generate content, interpret prior work, or draw conclusions. The authors reviewed and approved all final text and remain fully responsible for the content of the paper. References [1] Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard Baraniuk. 2024. Self-Consuming Generative Models Go MAD. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=ShjMHfmPs0 [2]Tuka Alhanai, Adam Kasumovic, Mohammad Ghassemi, Aven Zitzelberger, Jessica Lundin, and Guillaume Chabot-Couture. 2025. Enhancing LLM Performance for Low-Resource African Languages. In Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI). [3] Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024. Investigating Cultural Alignment of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 12404â12422. doi:10.18653/v1/2024.acl-long.671 [4]Maryam Amirizaniani, Elias Martin, Tanya Roosta, Aman Chadha, and Chirag Shah. 2024. AuditLLM: A Tool for Auditing Large Language Models Using Multiprobe Approach. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (Boise, ID, USA) (CIKM â24). Association for Computing Machinery, New York, NY, USA, 5174â5179. doi:10.1145/3627673.3679222 [5]Anonymous. 2025. MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?. In Machine Learning for Health 2025. https://openreview.net/forum?id=CrZyNUf WHp [6]Amos Azaria and Tom Mitchell. 2023. The Internal State of an LLM Knows When Itâs Lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 967â976. doi:10.18653/v1/2023.findings-emnlp.68 [7] Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, and Deep Ganguli. 2022. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862 [cs.CL] https://arxiv.org/abs/2204.05862 [8]Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W. Duncan Wadsworth, and Hanna Wallach. 2021. Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and Tradeoffs. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES â21). Association for Computing Machinery, New York, NY, USA, 368â378. doi:10.1145/3461702.3462610 [9] Joshua Barron and Devin White. 2025. Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers. In Tiny Titans: The next wave of On-Device Learning for Foundational Models (TTODLer-FM). https://openreview.net/forum?id=skWiP3K67u [10] Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT). [11]Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom. 2023. Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean Functions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 5767â5791. doi:10.18653/v1/2023.acl-long.317 [12]Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html. [13] Su Lin Blodgett, Solon Barocas, Hal DaumĂ© I, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of âBiasâ in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 5454â5476. doi:10.18653/v1/2020.acl-main.485 Manuscript submitted to ACM 16Badhe et al. [14]Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, and Russ Altman. 2021. On the Opportunities and Risks of Foundation Models. ArXiv (2021). https://crfm.stanford.edu/assets/report.pdf [15]Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian TramĂšr. 2022. What Does it Mean for a Language Model to Preserve Privacy?. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT â22). Association for Computing Machinery, New York, NY, USA, 2280â2292. doi:10.1145/3531146.3534642 [16]Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research, Vol. 81), Sorelle A. Friedler and Christo Wilson (Eds.). PMLR, 77â91. https://proceedings.mlr.press/v81/buolamwini18a.html [17] Nicholas Carlini, Florian TramĂšr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ălfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting Training Data from Large Language Models. In 30th USENIX Security Symposium (USENIX Security 21). USENIX Association, 2633â2650. https://w.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting [18]Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, JĂ©rĂ©my Scheurer, Javier Rando, Rachel Freedman, Tomek Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphael Segerie, Micah Carroll, Andi Peng, Phillip J.K. Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. 2023. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Transactions on Machine Learning Research (2023). https://openreview.net/forum?id=bx24KpJ4Eb Survey Certification, Featured Certification. [19] Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, JĂ©rĂ©my Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, and Dylan Hadfield-Menell. 2024. Black-Box Access is Insufficient for Rigorous AI Audits. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT â24). Association for Computing Machinery, New York, NY, USA, 2254â2272. doi:10.1145/3630106.3659037 [20]Kent Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023. Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 7312â7327. doi:10.18653/v1/2023.emnlp-main.453 [21] Xuanzhong Chen, Xiaohao Mao, Qihan Guo, Lun Wang, Shuyang Zhang, and Ting Chen. 2024. RareBench: Can LLMs Serve as Rare Diseases Specialists?. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain) (KDD â24). Association for Computing Machinery, New York, NY, USA, 4850â4861. doi:10.1145/3637528.3671576 [22]Zhao Chen, Vincent Casser, Henrik Kretzschmar, and Dragomir Anguelov. 2022. GradTail: Learning Long-Tailed Data Using Gradient-based Sample Weighting. arXiv:2201.05938 [cs.LG] https://arxiv.org/abs/2201.05938 [23] Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024. Evaluating the Ripple Effects of Knowledge Editing in Language Models. Transactions of the Association for Computational Linguistics 12 (2024), 283â298. doi:10.1162/tacl_a_00644 [24] A. Feder Cooper, Emanuel Moss, Benjamin Laufer, and Helen Nissenbaum. 2022. Accountability in an Algorithmic Society: Relationality, Responsibility, and Robustness in Machine Learning. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT â22). Association for Computing Machinery, New York, NY, USA, 864â876. doi:10.1145/3531146.3533150 [25]Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis 16, 1 (Jan. 2024), 64â93. doi:10.1093/jla/laae003 [26]Fatemeh Dehghani, Roya Dehghani, Yazdan Naderzadeh Ardebili, and Shahryar Rahnamayan. 2025. Large Language Models in Legal Systems: A Survey. Humanities and Social Sciences Communications 12, 1 (2025), 1977. doi:10.1057/s41599-025-05924-3 [27]Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. Time-Aware Language Models as Temporal Knowledge Bases. Transactions of the Association for Computational Linguistics 10 (2022), 257â273. doi:10.1162/tacl_a_00459 [28]Jesse Dodge, Maarten Sap, Ana MarasoviÄ, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 1286â1305. doi:10.18653/v1/2021.emnlp-main.98 [29]Esin DURMUS, Karina Nguyen, Thomas Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2024. Towards Measuring the Representation of Subjective Global Opinions in Language Models. In First Conference on Language Modeling. https://openreview.net/forum?id=zl16jLb91v [30]Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy Models of Superposition. Transformer Circuits Thread (2022). [31]William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res. 23, 1, Article 120 (Jan. 2022), 39 pages. [32]Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications17 Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv:2209.07858 [cs.CL] https://arxiv.org/abs/2209.07858 [33]Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https://arxiv.org/abs/2312.10997 [34]Tarleton Gillespie. 2018. Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions That Shape Social Media. Yale University Press, New Haven, CT. [35]Kristian GonzĂĄlez Barman, Simon Lohse, and Henk W. de Regt. 2025. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values, Whose Perspectives? Philosophy & Technology 38, 35 (2025). doi:10.1007/s13347-025-00861-0 [36]Gabriel Grill. 2024. Constructing Capabilities: The Politics of Testing Infrastructures for Generative AI. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT â24). Association for Computing Machinery, New York, NY, USA, 1838â1849. doi:10.1145/3630106.3659009 [37]Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng. 2024. Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 16801â16819. doi:10.18653/v1/2024.emnlp-main.934 [38]Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio CĂ©sar Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, SĂ©bastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks Are All You Need. arXiv:2306.11644 [cs.CL] https://arxiv.org/abs/2306.11644 [39] Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning (ICMLâ20). JMLR.org, Article 368, 10 pages. [40]Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations. https://openreview.net/forum?id=d7KBjmI3GmQ [41]Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation. arXiv:2302.09210 [cs.CL] https://arxiv.org/abs/2302.09210 [42]Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Ca- bello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders SĂžgaard. 2022. Challenges and Strategies in Cross-Cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 6997â7013. doi:10.18653/v1/2022.acl-long.482 [43] Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations. https://openreview.net/forum?id=rygGQyrFvH [44]Chenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2024. WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge Editing. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 3476â3503. doi:10.18653/v1/2024.findings-acl.207 [45] Mark Humphries, Lianne C. Leddy, Quinn Downton, Meredith Legace, John McConnell, Isabella Murray, and Elizabeth Spence. 2024. Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents. arXiv:2411.03340 [cs.CV] https://arxiv.org/abs/2411.03340 [46]Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg. 2023. Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks. In The 2023 Conference on Empirical Methods in Natural Language Processing. https://openreview. net/forum?id=hA8h2KtSv2 [47] Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun KIM, Stanley Jungkyu Choi, and Minjoon Seo. 2022. Towards Continual Knowledge Learning of Language Models. In International Conference on Learning Representations. https://openreview.net/forum?id= vfsRB5MImo9 [48] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 55, 12, Article 248 (March 2023), 38 pages. doi:10.1145/3571730 [49]Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 55, 12, Article 248 (March 2023), 38 pages. doi:10.1145/3571730 [50] Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, LĂ©lio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, ThĂ©ophile Gervet, Thibaut Lavril, Thomas Wang, TimothĂ©e Lacroix, and William El Sayed. 2024. Mixtral of Experts. arXiv:2401.04088 [cs.LG] https://arxiv.org/abs/2401.04088 [51]Ming Jiang and Mansi Joshi. 2024. CPopQA: Ranking Cultural Concept Popularity by LLMs. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 615â630. doi:10.18653/v1/2024.naacl-short.52 Manuscript submitted to ACM 18Badhe et al. [52]Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language Models (Mostly) Know What They Know. arXiv:2207.05221 [cs.CL] https://arxiv.org/abs/2207.05221 [53]Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICMLâ23). JMLR.org, Article 641, 12 pages. [54]Jenna Kanerva, Cassandra Ledins, Siiri KĂ€pyaho, and Filip Ginter. 2025. OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches. In Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025), Ć pela Arhar Holdt, Nikolai Ilinykh, Barbara Scalvini, Micaella Bruton, Iben Nyholm Debess, and Crina Madalina Tudor (Eds.). University of Tartu Library, Estonia, Tallinn, Estonia, 38â47. https://aclanthology.org/2025.resourceful-1.8/ [55]Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models. arXiv:2001.08361 [cs.LG] https://arxiv.org/abs/2001.08361 [56] Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2023. REALTIME QA: whatâs the answer right now?. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS â23). Curran Associates Inc., Red Hook, NY, USA, Article 2130, 19 pages. [57]Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo. 2024. GPT-4 Passes the Bar Exam. SSRN Electronic Journal (2024). [58]Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through Memorization: Nearest Neighbor Language Models. In International Conference on Learning Representations. https://openreview.net/forum?id=HklBjCEKvH [59] Saurabh Khanna and Xinxu Li. 2025. Invisible Languages of the LLM Universe. arXiv:2510.11557 [cs.CL] https://arxiv.org/abs/2510.11557 [60] Manurag Khullar, Utkarsh Desai, Poorva Malviya, Aman Dalmia, and Zheyuan Ryan Shi. 2025. Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Roman Scripts in a Real World Setting. arXiv:2512.10780 [cs.CL] https://arxiv.org/abs/2512.10780 [61]Hannah Rose Kirk, Ishan Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2024. Understanding the Effects of RLHF on LLM Generalization and Diversity. In International Conference on Learning Representations (ICLR). [62]Taku Kudo. 2018. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Iryna Gurevych and Yusuke Miyao (Eds.). Association for Computational Linguistics, Melbourne, Australia, 66â75. doi:10.18653/v1/P18-1007 [63]Juhi Kulshrestha, Motahhare Eslami, Johnnatan Messias, Muhammad Bilal Zafar, Saptarshi Ghosh, Krishna P. Gummadi, and Karrie Karahalios. 2019. Search bias quantification: investigating political bias in social media and web search. Inf. Retr. 22, 1â2 (April 2019), 188â227. doi:10.1007/s10791- 018-9341-2 [64] Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamil Ì e LukoĆĄi Ì ut Ì e, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez. 2023. Measuring Faithfulness in Chain-of-Thought Reasoning. arXiv:2307.13702 [cs.AI] https://arxiv.org/abs/2307.13702 [65]Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam LiĆĄka, Tayfun Terzi, Mai Gimenez, Cyprien de Masson dâAutume, Tomas Kocisky, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2021. Mind the gap: assessing temporal generalization in neural language models. In Proceedings of the 35th International Conference on Neural Information Processing Systems (NIPS â21). Curran Associates Inc., Red Hook, NY, USA, Article 2247, 16 pages. [66]Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022. Deduplicating Training Data Makes Language Models Better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 8424â8445. doi:10.18653/v1/2022.acl-long.577 [67]Maria Levchenko. 2025.Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities. arXiv:2510.06743 [cs.CV] https://arxiv.org/abs/2510.06743 [68]Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich KĂŒttler, Mike Lewis, Wen-tau Yih, Tim RocktĂ€schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS â20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages. [69]Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich KĂŒttler, Mike Lewis, Wen-tau Yih, Tim RocktĂ€schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS â20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages. [70]Dongyang Li, Junbing Yan, Taolin Zhang, Chengyu Wang, Xiaofeng He, Longtao Huang, Hui Xueâ, and Jun Huang. 2024. On the Role of Long-tail Knowledge in Retrieval Augmented Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications19 Thailand, 120â126. doi:10.18653/v1/2024.acl-short.12 [71]Xinjie Li and Huijuan Xu. 2023. MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video Recognition. Proceedings of the AAAI Conference on Artificial Intelligence 37, 2 (Jun. 2023), 1451â1459. doi:10.1609/aaai.v37i2.25230 [72] Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023. Holistic Evaluation of Language Models. Transactions on Machine Learning Research (2023). https://openreview.net/forum?id=iO4LZibEqW Featured Certification, Expert Certification, Outstanding Certification. [73]Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, and Furu Wei. 2025. Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 6317â6342. doi:10.18653/v1/2025.acl-long.317 [74]Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, and Furu Wei. 2025. Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 6317â6342. doi:10.18653/v1/2025.acl-long.317 [75]Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 3214â3252. doi:10.18653/v1/2022.acl-long.229 [76]Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, and Jianmeng Liu. 2024. Mitigating the Alignment Tax of RLHF. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 580â606. doi:10.18653/v1/2024.emnlp-main.35 [77]Yizhou Liu, Ziming Liu, and Jeff Gore. 2025. Superposition Yields Robust Neural Scaling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=knPz7gtjPW [78]Yizhou Liu, Ziming Liu, and Jeff Gore. 2025. Superposition Yields Robust Neural Scaling. arXiv:2505.10465 [cs.LG] https://arxiv.org/abs/2505.10465 [79]Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A. Smith. 2022. Time Waits for No One! Analysis and Challenges of Temporal Misalignment. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.). Association for Computational Linguistics, Seattle, United States, 5944â5958. doi:10.18653/v1/2022.naacl-main.435 [80] Seiji Maekawa, Hayate Iso, Sairam Gurajada, and Nikita Bhutani. 2024. Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 5506â5521. doi:10.18653/v1/2024.naacl-long.308 [81]Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho. 2025.Hallucination- Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies 22, 2 (2025), 216â242. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/jels.12413 doi:10.1111/jels.12413 [82]Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Proceedings of the 61st Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 9802â9822. doi:10.18653/v1/2023.acl-long.546 [83]Clara Meister, Ryan Cotterell, and Tim Vieira. 2020. If beam search is the answer, what was the question?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 2173â2185. doi:10.18653/v1/2020.emnlp-main.170 [84] Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022. Locating and Editing Factual Associations in GPT. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). https://openreview.net/forum?id=- h6WAS6eE4 [85] Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-Editing Memory in a Transformer. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=MkbcAHIYgyS [86]Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2022. Reframing Instructional Prompts to GPTkâs Language. In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 589â612. doi:10.18653/v1/2022.findings-acl.50 [87]Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, Victor Gutierrez Basulto, Yazmin Ibanez-Garcia, Hwaran Lee, Shamsuddeen Hassan Muhammad, Kiwoong Park, Anar Sabuhi Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, Jose Camacho-Collados, and Alice Oh. 2024. BLEnD: A Manuscript submitted to ACM 20Badhe et al. Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=nrEqH502eC [88]Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2024. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 16366â16393. doi:10.18653/v1/2024.acl-long.862 [89]Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. 2023. Extracting Cultural Commonsense Knowledge at Scale. In Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (W â23). Association for Computing Machinery, New York, NY, USA, 1907â1917. doi:10.1145/3543507.3583535 [90] Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. 2023. Capabilities of GPT-4 on Medical Challenge Problems. arXiv:2303.13375 [cs.CL] https://arxiv.org/abs/2303.13375 [91]Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Re. 2020. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In Proceedings of the ACM Conference on Health, Inference, and Learning (Toronto, Ontario, Canada) (CHIL â20). Association for Computing Machinery, New York, NY, USA, 151â159. doi:10.1145/3368555.3384468 [92] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, et al.2022. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems (NeurIPS). [93]David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 [cs.LG] https://arxiv.org/abs/2104.10350 [94] Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction Tuning with GPT-4. arXiv:2304.03277 [cs.CL] https://arxiv.org/abs/2304.03277 [95] Aleksandar Petrov, Emanuele La Malfa, Philip H.S. Torr, and Adel Bibi. 2023. Language model tokenizers introduce unfairness between languages. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS â23). Curran Associates Inc., Red Hook, NY, USA, Article 1608, 28 pages. [96]Mohammad Pezeshki, SĂ©kou-Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie. 2021. Gradient Starvation: A Learning Proclivity in Neural Networks. In Advances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (Eds.). https://openreview.net/forum?id=aExAsh1UHZo [97] Alistair Plum, Anne-Marie Lutgen, Christoph Purschke, and Achim Rettinger. 2025. Identity-Aware Large Language Models require Cultural Reasoning. arXiv:2510.18510 [cs.CL] https://arxiv.org/abs/2510.18510 [98]Stanford HAI Policy. 2025. Mind the (Language) Gap: Mapping the Challenges of LLM Development in Low-Resource Language Contexts. Stanford Institute for Human-Centered AI White Paper (2025). [99] Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022. The Fallacy of AI Functionality. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT â22). Association for Computing Machinery, New York, NY, USA, 959â972. doi:10.1145/3531146.3533158 [100] Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain) (FAT* â20). Association for Computing Machinery, New York, NY, USA, 33â44. doi:10.1145/3351095.3372873 [101] Oscar Sainz, Jon Campos, Iker GarcĂa-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 10776â10787. doi:10.18653/v1/2023.findings-emnlp.722 [102]Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect?. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICMLâ23). JMLR.org, Article 1244, 34 pages. [103]Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (Atlanta, GA, USA) (FAT* â19). Association for Computing Machinery, New York, NY, USA, 59â68. doi:10.1145/3287560.3287598 [104]Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural Machine Translation of Rare Words with Subword Units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Katrin Erk and Noah A. Smith (Eds.). Association for Computational Linguistics, Berlin, Germany, 1715â1725. doi:10.18653/v1/P16-1162 [105] Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 9248â9274. doi:10.18653/v1/2023.findings-emnlp.620 [106] Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024. The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 2668â2680. doi:10.18653/v1/2024.findings-acl.156 Manuscript submitted to ACM Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications21 [107]Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2024. The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493 [cs.LG] https://arxiv.org/abs/2305.17493 [108]Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024. AI models collapse when trained on recursively generated data. Nature (2024). [109]Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large Language Models Encode Clinical Knowledge. Nature (2023). [110]Aarohi Srivastava, Abhinav Rastogi, and Abhishek Rao. 2023. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research (2023). https://openreview.net/forum?id=uyTL5Bvosj Featured Certification. [111] Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and LluĂs MĂ rquez (Eds.). Association for Computational Linguistics, Florence, Italy, 3645â3650. doi:10.18653/v1/P19-1355 [112]Zunhai Su, Qingyuan Li, Hao Zhang, Weihao Ye, Qibo Xue, YuLei Qian, Yuchen Xie, Ngai Wong, and Kehong Yuan. 2025. Unveiling Super Experts in Mixture-of-Experts Large Language Models. arXiv:2507.23279 [cs.CL] https://arxiv.org/abs/2507.23279 [113] Neil C. Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F. Manso. 2022.The Computational Limits of Deep Learning. arXiv:2007.05558 [cs.LG] https://arxiv.org/abs/2007.05558 [114]Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. Memorization without overfitting: analyzing the training dynamics of large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS â22). Curran Associates Inc., Red Hook, NY, USA, Article 2773, 17 pages. [115]Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Language models donât always say what they think: unfaithful explanations in chain-of-thought prompting. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS â23). Curran Associates Inc., Red Hook, NY, USA, Article 3275, 14 pages. [116]Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2024. FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 13697â13720. doi:10.18653/v1/2024.findings-acl.813 [117] Dixuan Wang, Yanda Li, Junyuan Jiang, Zepeng Ding, Ziqin Luo, Guochao Jiang, Jiaqing Liang, and Deqing Yang. 2025. Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization. arXiv:2405.17067 [cs.CL] https://arxiv.org/abs/2405.17067 [118]Pengkun Wang, Zhe Zhao, HaiBin Wen, Fanfu Wang, Binwu Wang, Qingfu Zhang, and Yang Wang. 2024. LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed Problems. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=VpuOuZOVhP [119] Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, and Preslav Nakov. 2024. Factuality of Large Language Models: A Survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 19519â19529. doi:10.18653/v1/2024.emnlp-main.1088 [120]Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022. Taxonomy of Risks posed by Language Models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT â22). Association for Computing Machinery, New York, NY, USA, 214â229. doi:10.1145/3531146.3533088 [121]Orion Weller, Michael Boratko, Iftekhar Naim, and Jinhyuk Lee. 2025.On the Theoretical Limitations of Embedding-Based Retrieval. arXiv:2508.21038 [cs.IR] https://arxiv.org/abs/2508.21038 [122] Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J. Su. 2025. On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization. arXiv:2405.16455 [stat.ML] https://arxiv.org/abs/2405.16455 [123]Shaoyang Xu, Yongqi Leng, Linhao Yu, and Deyi Xiong. 2025. Self-Pluralising Culture Alignment for Large Language Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Luis Chiruzzo, Alan Ritter, and Lu Wang (Eds.). Association for Computational Linguistics, Albuquerque, New Mexico, 6859â6877. doi:10.18653/v1/2025.naacl-long.350 [124]Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023. Editing Large Language Models: Problems, Methods, and Opportunities. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 10222â10240. doi:10.18653/v1/2023.emnlp- main.632 [125] Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Winata, and Alham Fikri Aji. 2023. Multilingual Large Language Models Are Not (Yet) Code-Switchers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 12567â12582. doi:10.18653/v1/2023.emnlp-main.774 [126]Yuji Zhang, Sha Li, Cheng Qian, Jiateng Liu, Pengfei Yu, Chi Han, Yi R. Fung, Kathleen McKeown, ChengXiang Zhai, Manling Li, and Heng Ji. 2025. The Law of Knowledge Overshadowing: Towards Understanding, Predicting and Preventing LLM Hallucination. In Findings of the Association for Manuscript submitted to ACM 22Badhe et al. Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 23340â23358. doi:10.18653/v1/2025.findings-acl.1199 [127]Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2025. Sirenâs Song in the AI Ocean: A Survey on Hallucination in Large Language Models. Computational Linguistics (09 2025), 1â46. arXiv:https://direct.mit.edu/coli/article-pdf/doi/10.1162/COLI.a.16/2535477/coli.a.16.pdf doi:10.1162/COLI.a.16 Received 12 January 2026; revised 12 March 2009; accepted 5 June 2009 Manuscript submitted to ACM