Paper deep dive
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
Patrik Reizinger, Wieland Brendel
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/22/2026, 2:30:17 AM
Summary
The paper introduces HALLMARK, a benchmark for evaluating LLM-based citation verifiers, identifying false-positive rate (FPR) as the primary deployment bottleneck. It analyzes three failure modes: agentic lookups inflating FPR, venue-realistic base rates making FPR critical, and LLMs over-flagging papers past their training cutoff.
Entities (8)
Relation Signals (5)
HALLMARK â contains â 2,526 BibTeX entries
confidence 98% · HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types
HALLMARK â evaluates â LLM Citation Verifiers
confidence 95% · HALLMARK makes it concrete through three failure modes... evaluate a DOI-lookup baseline, frontier LLMs...
False Positive Rate â determines â Deployability
confidence 92% · the false-positive rate, not recall, decides whether a verifier is deployable.
GPTZero â detected â Hallucinated Citations
confidence 90% · GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set.
LLMs â overflag â Papers past training cutoff
confidence 85% · most LLMs over-flag papers published past their training cutoff
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set. Rule- and LLM-based verifiers are emerging, but no shared benchmark compares them and gives detailed failure diagnostics. We close that gap with HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split. On it we evaluate a DOI-lookup baseline, frontier LLMs zero-shot, tool-augmented agents, and our own rule-based, co-designed verifier bibtex-updater. Across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deployable. HALLMARK makes it concrete through three failure modes: agentic lookups buy recall but inflate false positives; at a venue-realistic base rate, the order-of-magnitude spread in false-positive rates (FPRs) -- not recall -- governs whether a verifier's flags are mostly true catches or mostly noise; and most LLMs over-flag papers published past their training cutoff, where only the two latest-cutoff models hold their false-positive rate near in-distribution levels (a signal we report as descriptive, since it is confounded with possible recall of those entries). Thus FPR is the deployment bottleneck, but an undetected fabrication remains the costlier error for the scientific record.
Tags
Links
- Source: https://arxiv.org/abs/2607.18360v1
- Canonical: https://arxiv.org/abs/2607.18360v1
Trouble viewing inline? Open PDF directly â
Full Text
238,688 characters extracted from source content.
Expand or collapse full text
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers Anonymous Author(s) Abstract Large Language Models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025âs accepted set (Ansari, 2026). Rule- and LLM-based verifiers are emerging, but no shared benchmark compares them and gives detailed failure diagnostics. We close that gap with Hallmark (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split. On it we evaluate a Digital Object Identifier (DOI)-lookup baseline, frontier LLMs zero-shot, tool-augmented agents, and our own rule-based, co-designed verifier bibtex-updater. Across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deployable. Hallmark makes it concrete through three failure modes: agentic lookups buy recall but inflate false positives; at a venue-realistic base rate, the order-of-magnitude spread in false-positive rates (FPRs)ânot recallâgoverns whether a verifierâs flags are mostly true catches or mostly noise; and most LLMs over-flag papers published past their training cutoff, where only the two latest-cutoff models hold their false-positive rate near in-distribution levels (a signal we report as descriptive, since it is confounded with possible recall of those entries). Thus FPR is the deployment bottleneck, but an undetected fabrication remains the costlier error for the scientific record. 1 Introduction Hallucinated citationsâplausible but fabricated references generated by language modelsâhave become a concrete threat to scientific publishing. Shortly after NeurIPS 2025, GPTZero audited the 4,841 accepted papers and found 53 with fabricated citations that had passed peer review [Shmatko et al., 2026, Ansari, 2026]. Independent audits spanning millions of citations report hundreds of affected papers and a sharp rise since 2021 [Xu et al., 2026, Sakai et al., 2026a, Bienz et al., 2026], and neither authors nor reviewers reliably catch the fabrications. Tools to catch them have proliferatedâfrom DOI resolvers to multi-database cross-referencing systemsâbut each is evaluated on its own ad-hoc data under its own conditions, so no shared benchmark compares them. We cannot yet say which tool catches which failure, where each one breaks, or what coverage it leaves open. Hallmark is, to our knowledge, among the first benchmarks to evaluate citation-verification tools under a single protocolâone taxonomy, diagnostic sub-tests, controlled tier difficulty, and contamination-resistant splits (a held-out set drawn so its entries are unlikely to sit in any evaluated toolâs training data)âalongside concurrent work such as the human-validated CiteAudit [Shi et al., 2026] (full comparison in §ËC.6). Source Papers DBLP 2021â2023 513 valid entries Generation Perturbation, LLM, real-world, adversarial Benchmark 2,526 entries 6 sub-tests each Tool Evaluation 13 tools: 1 DB, 12 LLMs Metrics DR, FPR, TW-F1, ECE, MCC Benchmark PipelineHallucination Taxonomy fabricated_doi nonexistent_venue placeholder_authors future_date Tier 1 4 types chimeric_title wrong_venue author_mismatch preprint_as_pub. hybrid_fabrication Tier 2 5 types near_miss_title plausible_fabrication Tier 3 2 typesDifficulty 6 Sub-tests per entry â doi_resolves â title_exists â authors_match â venue_correct â year_correct â fields_complete Figure 1: Overview of Hallmark. Top: The benchmark pipeline: real papers sourced from DBLP are transformed via perturbation, LLM generation, and real-world collection into 2,526 annotated BibTeX entries, each with six diagnostic sub-tests. Thirteen full-coverage verification tools (1 citation-database, 12 zero-shot LLMs) are evaluated using tier-weighted metrics. Bottom left: The three-tier hallucination taxonomy with 11 main types, ordered by the verification effort required to detect them. Bottom right: Each entry undergoes six binary sub-tests that reveal why a tool detects (or misses) a hallucination. LLMs achieve 48â91% detection rates with a pronounced recallâprecision tradeoff, while API tools are limited to â€27% detection. Scope. Hallmark targets citation metadataâdetecting fabricated, inconsistent, or nonexistent bibliographic recordsânot claim-level hallucination, which needs full-text analysis and is an orthogonal problem. We make four contributions: 1. A taxonomy of 14 hallucination types across three difficulty tiers, from non-resolving DOIs to fully plausible fabrications (§Ë3.2). 2. A benchmark of 2,526 BibTeX entries, each with six diagnostic sub-tests that localize why a verifier succeeds or fails, in the spirit of HumanEval [Chen et al., 2021] (§Ë3.3). 3. An evaluation protocol on prevalence-invariant metricsâdetection rate (DR), false positive rate (FPR), and Matthews Correlation Coefficient (MCC)âwith tier-weighting and calibration, so rankings survive the class-ratio shifts between splits (§Ë4). 4. Open infrastructure: the benchmark, pre-computed baselines, and Croissant metadata for drop-in use (§Ë5). We validate Hallmark on thirteen full-coverage systemsâto our knowledge the broadest cohort run under one protocolâand find that a single variable organizes the results: the false-positive rate, not recall, decides whether a verifier can be deployed, even though catching every fabrication is what ultimately builds trust. A DOI-only baseline catches 27% of hallucinations, and LLMs queried directly, without retrieval tools, catch 48â91%, but along a pronounced recallâprecision spectrum. Tab.Ë5 previews three failure modes that follow from it, each measured in §Ë6. Agentic lookups buy recall but inflate false positives; at a venue-realistic âŒ2% 2\% hallucination rate precision falls, and the order-of-magnitude spread in false-positive rates decides whose flags are worth reading. Most LLMs also over-flag papers published past their training cutoff: on a 448-entry 2024â2025 supplement, 8 of 12 degrade sharply, two over-flag only moderately (Gemini 2.5 Pro, GPT-5.4), and only the two latest-cutoff models hold. We report that last resistance as a descriptive signal, because cutoff recency and provider pipeline are confounded and we cannot rule out training-data recall of these entries (§Ë6). 2 Related work Citation hallucination in the LLM era. LLM hallucination [Ji et al., 2023, Huang et al., 2025] manifests distinctly in citations [Alkaissi and McFarlane, 2023, Agrawal et al., 2024, Walters and Wilder, 2023]; venue-level audits document scaleâGhostCite [Xu et al., 2026] on 56K papers, HalluCitation [Sakai et al., 2026a] on ACL/NAACL/EMNLP, the HPC-venue study of Bienz et al. [2026]âmeasuring prevalence, not detection-tool performance. Hallmark addresses the complementary tool-evaluation question on a controlled, tiered, sub-test-decomposed taxonomy. Detection toolsâHaRC [HaRC Contributors, 2024], verify-citations [verify-citations Contributors, 2025], RefChecker [Hu et al., 2024], CheckIfExist [Abbonato, 2026]âreport incompatible metrics on disjoint test data, preventing direct comparison; CheckIfExist (a CrossRef/Semantic-Scholar/OpenAlex cascade verifier) ships without dataset or metrics, and RefChecker requires full-text manuscript context outside Hallmarkâs metadata scope. Concurrent work [Rao and Callison-Burch, 2026, Rao et al., 2026, Sakai et al., 2026b] evaluates LLM-as-citation-generator accuracy or releases lightweight checkers; Hallmark is complementary, benchmarking standalone detection tools. The closest concurrent effort is CiteAudit [Shi et al., 2026], which likewise benchmarks LLMs and commercial verifiers on a human-validated dataset via a multi-agent pipeline. Hallmark differs in what it isolates: diagnostic sub-tests, a 14-type error taxonomy, and a temporal train/test split that controls for training-cutoff contamination, so it reports which metadata failures each tool catches rather than one aggregate score. CiteAudit leads on human validation: its dataset is human-annotated, whereas Hallmarkâs real-world and adversarial entries receive manual author review but no multi-rater human inter-annotator agreement: our released reliability check is an automated LLM-rater proxy (Fleissâ Îș=0.24Îș=0.24), with human Inter-Annotator Agreement (IAA) left to future work (§ËH.4 and 19). Adjacent concurrent detection work [Rao et al., 2026] targets citation-URL liveness on DRBench (the stale-vs-hallucinated distinction), a scope orthogonal to Hallmarkâs BibTeX-metadata cross-consistency checks. §ËC.6 details boundary cases of audit-vs-benchmark categorization. Hallucination detection benchmarks. General-purpose hallucination benchmarks [Li et al., 2023, Ravichander et al., 2025] target factual claims, not metadata; adjacent workâFactScore [Min et al., 2023], RARR [Gao et al., 2023], SciFact [Wadden et al., 2020]âdoes not evaluate BibTeX integrity (DOI/title/author/venue cross-consistency against external databases). Concurrent generator-side benchmarks such as HalluHard [Fan et al., 2026] measure whether LLMs produce claims grounded by retrievable inline citations in multi-turn dialogues; Hallmark is complementary, evaluating whether downstream verifiers can flag fabricated citation metadata once produced. Benchmark design principles. Hallmark synthesizes design principles from established benchmarks: multi-criteria sub-tests from HumanEval [Chen et al., 2021], temporal segmentation from SWE-bench [Jimenez et al., 2024] and LiveCodeBench [Jain et al., 2025], multi-difficulty challenge sets from Dynabench [Kiela et al., 2021], and continuous, expandable sample pools from ONEBench [Ghosh et al., 2025] that aggregate per-entry measurements while resisting contamination and leaderboard rot. We complement these with the evaluation discipline of PostTrainBench [Rank et al., 2026]: pin model versions, timestamp every result, and never assume the newest or largest model is the strongest, since test-set contamination tends to scale with capability. 3 The Hallmark benchmark 3.1 What is a hallucination? The question seems almost trivial. But looking into the details exposes a lot of nuance: telling what is not a hallucination is easy (a bit-by-bit match with a trusted database entry), but the reverse is not. Different spellings, handling of diacritics, name conventions (e.g., how Asian names get represented in Latin alphabets), hyphenation rules, and changing author lists across (preprint) versions complicate the picture and do not afford a clear-cut decision. With these caveats in mind, the main goal of Hallmark is to analyze potential failure modes without being too pedantic about unconditionally enforcing bit-by-bit matches where humans can also err, even with the best of intentions. This does not mean tool designers should not strive for perfection, but they need to acknowledge these nuancesâespecially the ones that have a large effect on deployment. 3.2 Hallucination taxonomy We define 14 citation hallucination typesâ11 empirically grounded and 3 stress-test typesâorganized into three difficulty tiers based on the verification effort required (Tab.Ë1). Tier 1 (Easy) hallucinations are detectable by a single API lookup: a fabricated DOI that does not resolve, a nonexistent venue name, placeholder author names, or a publication date in the future. Tier 2 (Medium) hallucinations require cross-referencing multiple metadata fields: e.g., a chimeric title pairs real authors with a fabricated title, wrong venue assigns a paper to the wrong conference, and hybrid fabrication uses a real DOI whose resolved record does not match the BibTeX metadata (Tab.Ë1). Tier 3 (Hard) hallucinations require deep verification or semantic reasoning: near-miss titles differ by one or two words from a real paper; plausible fabrications are entirely invented but realistic; and arXiv version mismatches cite a preprint with wrong venue and shifted year. The taxonomy is grounded in real incidents: we derived it from hallucinated citations in the NeurIPS 2025 incident and related audits [Xu et al., 2026, Sakai et al., 2026a], then stress-tested it with adversarial brainstorming of failure modes existing tools might miss. Of the 108 real-world hallucinated citations in the benchmark, the 72 we analyzed in detail map cleanly to a single type 79% of the time, with the rest assigned to their highest applicable tier (§ËA.2). Three types we cannot yet ground in any documented incidentâmerged_citation, partial_author_list, arxiv_version_mismatchâlive in a separate stress_test split, evaluated apart from the 11 empirical types and revised as evidence accumulates. Full per-type perturbation rules, LLM generation prompts, and the real-world collection procedure are in §ËA.3 and B.2.3. Table 1: The Hallmark hallucination taxonomy: 11 empirically-grounded types plus 3 stress-test types (âł ) across 3 difficulty tiers. Each type has a characteristic sub-test failure pattern. Sub-tests: DOI resolves, Title exists, Authors match, Venue real, Fields complete, X cross-DB agreement. â = expected pass, â = expected fail, ? = not applicable or varies by entry. Tier Type Description D T A V F X Easy fabricated_doi DOI does not resolve â â â â â â nonexistent_venue Invented conference/journal ? â â â â â placeholder_authors Generic/fake author names ? â â â â â future_date Year in the future ? â â â â â Medium chimeric_title Real authors + fake title â â â ? â â wrong_venue Correct paper, wrong venue â â â â â â author_mismatch Correct title, wrong authors â â â â â â preprint_as_published arXiv cited as venue paper â â â â â â hybrid_fabrication Real DOI + fake metadata ? â â â â â merged_citationâł Metadata from 2+ papers ? â â â â â partial_author_listâ âł Subset of real authors â â â â â â Hard near_miss_title Title off by 1â2 words â â â â â â plausible_fabrication Entirely fabricated, realistic ? â â â ? â arxiv_version_mismatchâł Wrong version claims ? â â â â â â Classified when <<50% authors present without âet al.â indicator. âł motivated types; zero real-world instances in current dataset. Verification decision tree. The tiers reflect a natural workflow: Tier 1 requires single-field lookups (DOI resolution, venue existence); Tier 2 requires cross-referencing (do authors match the resolved DOI record?); Tier 3 demands semantic reasoning (the title almost matches a real paper). A tool that only performs DOI lookups catches Tier 1 but misses Tier 2â3 entirely. Worked example. An illustrative near_miss_title pair (Tier 3)âa real entry and its perturbed counterpart, all other fields unchanged: ⏠title = Attention is All you Need % title = Attention is All you Require % The perturbed entry passes the D (DOI), A (authors), V (venue), and F (fields) sub-tests; only T (title exists) and X (cross-DB agreement) fail (Tab.Ë1), so a DOI-only lookup accepts it, and only title-aware verification catches it. 3.3 Dataset construction The dataset contains two classes of entries: valid references scraped from DBLP and hallucinated references generated through controlled perturbation. Valid entries. We scraped BibTeX records from DBLP [dblp Team, 2025] for papers published at major ML venues (NeurIPS, ICML, ICLR, AAAI, CVPR) between 2021 and 2023. Each entry was verified by confirming DOI resolution, title existence in at least two databases, and author-venue consistency. We retained 1,036 valid entries across all splits. Hallucinated entries. We generated hallucinated entries using four methods: (1) Systematic perturbation: modifying specific fields of valid entries to produce targeted hallucination types (e.g., replacing a DOI with a non-resolving one for fabricated_doi, swapping author lists between papers for author_mismatch). (2) LLM generation: prompting language models to generate plausible but fictional references for types requiring coherent fabrication (plausible_fabrication, chimeric_title). (3) Adversarial crafting: manually constructing entries designed to evade specific detection strategies. (4) Real-world collection: harvesting actual hallucinated citations from published papers identified in audits. The three sources play complementary roles: perturbation entries are controlled diagnostic tests that isolate specific verification capabilities, while LLM-generated and real-world entries supply ecological validity and empirical grounding; stratifying by generation method lets us assess each independently. An evaluation-only supplement of 341 authentic ChatGPT-generated citations (172 valid / 169 hallucinated), hand-verified by Walters and Wilder [2023] across 42 multidisciplinary topics, extends the ecological-validity axis with hallucinations we did not construct (§ËC.5). Labels are assigned deterministically by the generation pipeline; LLM-generated entries are also filtered against bibliographic databases, and real-world and adversarial entries receive manual review (§ËA.3). A ground-truth audit re-resolving entries against live databases recovered 52 real papers that earlier labeling had wrongly marked hallucinatedâ2.5% of the 2,072 public labelsâand the released corpus carries the corrected labels. Most types on dev_public and test_public carry â„30â„ 30 instances, enough for meaningful per-type comparison; test_hidden spreads 244 hallucinated entries over 14 types and does not clear that floor, so its per-type intervals are wider (statistical power in §ËB.1). Quality control and design choices. Every entry undergoes automated validation (field completeness, BibTeX well-formedness, sub-test consistency); see §ËA.3 for the validation rules and acceptance criteria. We strip the url field from all entries to prevent trivial shortcuts, and include canary strings for contamination detection (unique fixed tokens embedded in dataset metadata; if a model outputs them verbatim, it has memorized the split). No evaluated model emitted a canary token in any released prediction record. We use uniform type distribution to maximize per-type power; reweighting utilities support prevalence-adjusted evaluation. The API supports generation-method stratification for assessing tool performance on perturbation vs. LLM-generated vs. real-world entries. Table 2: Dataset statistics by split. Tier distribution refers to hallucinated entries only. dev_public and test_public cover all 14 hallucination types with nâ„30nâ„ 30 for most types (one type on dev_public and six on test_public fall just below, the smallest at n=23n=23); test_hidden covers all 14 types but is too small for an nâ„30nâ„ 30 floor; the stress-test split provides additional evaluation depth for three types (the single valid entry is the contamination canary). The bottom block lists four extension splits, evaluation sets in their own right that probe regimes the core does not sample: the temporal pair probes post-cutoff behavior on 2024â26 papers (§Ë6), test_crossdomain covers PubMed/bioRxiv and non-ML CS venues (§Ë7), and the ChatGPT-citation supplement carries authentic, hand-verified ChatGPT output spanning the humanities, social sciences, and natural sciences (§ËC.5). Each samples a different regime at a different prevalence and is scored separately, so the 2,526-entry total and the dev/test/hidden partition remain the contamination-controlled ML-venue core (construction in §ËA.3). Split Valid Halluc. Total Tier 1 Tier 2 Tier 3 Types dev_public 513 606 1,119 149 280 177 14 test_public 312 519 831 130 238 151 14 stress_test 1 121 122 â 85 36 3 test_hidden 210 244 454 52 115 77 14 Total 1,036 1,490 2,526 331 718 441 14 Extension splits (evaluation-only) temporal probe 30 30 60 10 10 10 9 temporal supplement (2024â25) 300 148 448 58 52 38 14 test_crossdomain 200 300 500 80 140 80 14 ChatGPT citations (WaltersâWilder) 172 169 341 â 17 152 4 3.4 Sub-test design and temporal segmentation Each entry includes six sub-tests (values: True/False/N/A) that decompose citation validity into independently verifiable dimensions: (1) DOI resolves, (2) Title exists in bibliographic databases, (3) Authors match the identified paper, (4) Venue real and correctly attributed, (5) Fields complete, (6) Cross-DB agreement. Each hallucination type has a characteristic failure signature (Tab.Ë1), enabling diagnostic analysis of why a tool succeeds or fails. Following LiveCodeBench [Jain et al., 2025], we tag entries with three temporal segments (pre-2023, 2023â2024, 2025+) so the benchmark can measure how verifiers degrade on recent papers (§Ë6). We read that degradation through the cross-regime ranking, which is prompt-invariant, rather than through absolute post-cutoff FPR, which a prompt-sensitivity sweep moves by 1010â3737 percentage points (p) on wording alone (§ËH.1). 4 Evaluation protocol 4.1 Metrics Two metrics carry the comparison, and both are prevalence-independent, so they stay comparable across splits whose class ratios differ. DR is recall on the hallucinated class; FPR is the fraction of valid entries wrongly flagged. The two trade off against different deployment costs. A missed fabrication that reaches print is usually the worse errorâcostlier than a false alarm a reviewer can dismissâbut which error dominates depends on the regime: prevalence (the fraction of entries that are hallucinated), the cost ratio cFN/cFPc_FN/c_FP, and reviewer capacity (§Ë6). We report the false-negative rate =1âDR=1-DR for completeness. Four further metrics summarize a tool in a single number. F1-Hallucination is the harmonic mean of precision and recall on the hallucinated class; it is prevalence-sensitive, so we compare it only within a split. Tier-weighted F1 (TW-F1) weights each hallucinated entry by its tier (1/2/3), rewarding the detection of harder hallucinations; the ranking holds across uniform, linear, and quadratic weightings (§ËB.1). MCC reads all four confusion-matrix cells and stays comparable when prevalence shifts (dev 54.2%, test 62.5% hallucinated). Expected Calibration Error (ECE) [Naeini et al., 2015] measures how well a toolâs confidence tracks its accuracy; as a rough guide, below 0.050.05 is excellent and above 0.20.2 unreliable. Formal definitions are in §ËB.1; per-tier and per-type breakdowns (§Ë5.3 and 5.4) expose category-specific strengths. 5 Experiments 5.1 Evaluated tools We evaluate thirteen independent full-coverage tools on dev_public: a DOI-only lower bound; GPT-5.1 [OpenAI, 2025]; a later-cutoff GPT-5.4 control (Aug 2025 cutoff) for the temporal story of §Ë6; and ten open- and closed-weight LLMs via OpenRouter: DeepSeek-R1, DeepSeek-V3.2 [DeepSeek-AI et al., 2024], Qwen3-235B and Qwen3-VL-235B-Instruct [Yang et al., 2025], Mistral Large [Mistral AI, 2024], Gemini 2.5 Flash and Pro [Google DeepMind, 2025], Llama 4 Maverick, Claude Sonnet 4.6, and Claude Opus 4.7 (model IDs in §ËB.2.4).111Code, dataset, and pre-computed baseline JSONs at https://anonymous.4open.science/r/hallmark/ (anonymised for review). All LLMs use the same zero-shot prompt; DeepSeek-R1 adds chain-of-thought (âŒ25 25s/entry vs. âŒ5 5s). On top of these, four agentic variants (§ËB.2.4) let the LLM issue up to five database lookups per entry before it commits to a verdict; prompted to cross-reference and decide, the model tends to flag an entry as soon as one lookup returns no matchâthe disposition behind the inflated false-positive rate we report in §Ë5.2. HaRC [HaRC Contributors, 2024] and verify-citations [verify-citations Contributors, 2025] are excluded from Tab.Ë3: even with a Semantic Scholar key, throttling reduces their coverage to below 7% on dev_public (§ËB.2.2). The DOI-only baseline verifies each entryâs DOI against the DOI.org resolver; our headline numbers omit the optional pre-screening layer (DOI format, year bounds, author heuristics), analyzed separately in §Ë6. We label bibtex-updater co-designed because its development overlapped the taxonomy: the typed sub-tests partly mirror the toolâs verification stages, so it may score better here than on a novel hallucination distribution. We therefore read its row as an upper-bound reference, excluded from ranking (§ËG.2). 5.2 Main results Tab.Ë3 reports dev_public. Cost asymmetry across model classes (chain-of-thought vs. single-pass; the agentic tool-call cap) shapes what is feasible to evaluate at scale; see §Ë7 (Compute and token budget) before reading the cross-class Pareto comparisons. Scoring conventions. Small gaps are point-estimate orderingsârankings by the single measured value, with no test that the gap is statistically real: for the Sonnet 4.6 / Opus 4.7 F1 score (F1) gap of 0.3 p, no paired test is available on dev_public, since both rows are summary-only (§ËB.1), and their coverage is not recoverable (endpoint drift, §ËE.2). Abstentionsâentries a tool declines to label either wayâscore as committed-valid in the DR/FPR/F1 triple, with coverage and an aggressive re-flagging stance in §ËE.2. Predictions were collected 2026-05-05; hosted-LLM rows are dated snapshots subject to endpoint drift (§Ë7). Table 3: Main results on dev_public. Bold marks the best point estimate among independent full-coverage tools (see the Scoring conventions paragraph, §Ë5.2). The shaded co-designed block was developed alongside the benchmark taxonomy and is a reference upper bound excluded from ranking; no co-designed cell is bolded (§ËG.2). Decimal cells drop the leading â0.â (except the signed Î column); the zero-shot block is sorted by FPR ascending. Cov. is the fraction of entries a tool commits to; abstentions score as committed-valid. ân/aâ marks the two Anthropic rows, whose dev_public records are summary-only (no stored per-entry predictions), so coverage cannot be computed, and endpoint drift rules out re-measuring it (§ËE.2). Î is the cross-split shift test_public â- dev_public. Scoring conventions and snapshot/drift caveats: §Ë5.2, §ËB.2.4, E.2 and 7. FPR cells are shaded: green marks the low-FPR frontier (FPR â€.13â€.13), red marks over-flagging operating points (FPR â„.41â„.41); the gray co-designed block is excluded from shading as from ranking. Performance Calibr. Coverage Robustness Tool DR â FPR â F1 â MCC â TW-F1 â ECE â Cov. â Î â Citation-database tools DOI-only .268 .185 .373 .099 .329 .143 1.00 +0.094+0.094 Zero-shot LLMs (sorted by FPR) Gemini 2.5 Pro .476 .050 .627 .473 .609 .297 .97 +0.009+0.009 Claude Opus 4.7 .752 .072 .830 .683 .851 .112 n/a â0.005 +-0.005 Gemini 2.5 Flash .500 .100 .631 .429 .628 .265 .99 +0.006+0.006 Claude Sonnet 4.6 .781 .127 .827 .652 .834 .066 n/a â0.002 +-0.002 Llama 4 Maverick .614 .146 .707 .476 .709 .176 1.00 +0.020+0.020 GPT-5.4 (zero-shot) .767 .228 .783 .538 .807 .202 1.00 â0.004 +-0.004 Mistral Large .716 .250 .742 .465 .765 .229 .99 +0.032+0.032 GPT-5.1 (zero-shot) .837 .411 .766 .442 .822 .190 1.00 +0.069+0.069 Qwen3-235B .860 .533 .744 .358 .821 .279 1.00 +0.082+0.082 Qwen3-VL-235B .860 .551 .740 .342 .818 .286 1.00 +0.077+0.077 DeepSeek-R1 .896 .623 .739 .324 .825 .238 .98 â0.303 +-0.303 DeepSeek-V3.2 .911 .702 .727 .268 .821 .316 1.00 +0.026+0.026 Agentic (tool-use; up to 5 tool calls per entry) GPT-5.1 + CrossRef/OpenAlex/arXiv .967 .478 .816 .558 .892 .175 1.00 +0.080+0.080 GPT-5.1 + bibtex-updater (agentic; tool optional) .980 .470 .824 .584 .900 .125 1.00 â0.114 +-0.114 Sonnet 4.6 + bibtex-updater (agentic; tool optional) .990 .431 .841 .630 .913 .118 1.00 â0.088 +-0.088 Co-designed (reference upper bound; see §ËG.2) bibtex-updater (v1.2.0) .865 .092 .890 .771 .908 .383 .82 +0.024+0.024 GPT-5.1 + bibtex-updater (always-call; output in prompt) .843 .144 .856 .698 .872 .078 1.00 +0.112+0.112 Three findings emerge from the independent tools (Tab.Ë3). âą LLMs span a wide recallâprecision spectrum (Fig.Ë2), from ultra-conservative (Gemini 2.5 Pro) to aggressive (DeepSeek-V3.2). The two later-cutoff Anthropic models anchor the precision endâOpus 4.7 and Sonnet 4.6 hold the lowest FPR and, with GPT-5.1, the best calibration in the cohortâbut they are statistically indistinguishable from each other: the F1 gap is 0.3 p, well below the per-type Minimum Detectable Effect (MDE), and no paired test is available on dev_public. We read them as a joint low-FPR frontier rather than ranking one above the other; the point estimate even reverses under scrutiny, since of the 27 recovered real dev papers Sonnet flags 19 as hallucinated to Opusâs 8 (§ËG.2), so any Sonnet edge comes from exactly the over-flagging the benchmark is built to catch (i.e., a higher FPR). Both clear GPT-5.1 and the recall-aggressive open-weight cohort, and GPT-5.4 (cutoff Aug 2025) sits between them at lower FPR and comparable recall, consistent with §Ë6. âą A capability gap remains. Even the highest-recall independent model misses âŒ9% 9\% of hallucinations, and the misses concentrate on subtle types: for GPT-5.1, author_mismatch and near_miss_title are weakest, both demanding exact bibliographic knowledge (Tab.Ë15). âą Agentic lookups give diminishing returns (failure mode i). Five tool calls push GPT-5.1âs recall past bibtex-updaterâs, but its false-positive rate rises to roughly five times bibtex-updaterâs, because the prompted modelâcross-referencing against CrossRef/OpenAlex/arXivâtends to flag an entry as soon as one of them returns no match. The inflation is a property of the harnessâits naive, single-stage use of the lookupsânot of the base model or of tool access itself: swapping Sonnet 4.6 for GPT-5.1 reproduces the profile within â€3.5â€3.5 p on every metric, and the effect is sharpest on a well-calibrated baseâSonnetâs FPR jumps âŒ3.4Ă 3.4Ă over its zero-shot baseline, while GPT-5.1âs already-high zero-shot FPR barely moves. This any-no-match behavior is not a hard-coded rule but an emergent disposition of the cross-reference-and-decide prompt, and a poor one for a precision-bound verifier: because CrossRef, OpenAlex, and arXiv have partial, non-overlapping coverage, a real paper is routinely missing from one of them, so single-source absence gets read as fabrication. The principled alternative is the policy bibtex-updater already usesâflag only on consensus absence across sources (or on positive metadata inconsistency) and route single-source absence to uncertain/human review (§ËE.2)âand the two-stage cascade (§Ë5.5) shows that tool use under this policy avoids the inflation entirely: it reaches DR 0.9960.996 at FPR 0.1080.108 on dev_public (Tab.Ë4), recovering the recall that motivated the any-no-match rule at roughly a quarter of the single-stage harnessesâ false-positive rate. A deterministic re-aggregation over three of the harnessâs four databases confirms the lever: holding sources and matcher fixed, switching from any-no-match flagging (flag if any single database returns no match) to consensus flagging (flag only when all queried databases agree the entry is absent) cuts FPR from 0.730.73 to 0.050.05 (a âŒ15Ă 15Ă reduction) at a recall cost on metadata-corruption types, which pairing consensus with contradiction checksâbibtex-updaterâs designârecovers (Appx.ËD). So failure mode (i) characterizes this common retrieval-augmented pattern, not every agentic design. bibtex-updaterâs low false-positive rate is structural: it comes from conservative matching, not from abstention. It abstains on the âŒ20% 20\% of entries it cannot back with a record; on the entries it commits to, its FPR is essentially unchanged while its detection rate increases (§ËE.2); abstention raises recall and leaves precision unchanged. Because its low FPR does not depend on abstention, bibtex-updater serves as the precision-anchored reference against which the LLMsâ abstention-driven gains are measured, and the reason we keep it out of the ranking (§ËG.2). Because DR and FPR are prevalence-independent, we compare on them, and on F1 only within a split. Takeaway. Both retrieval-augmented harnesses raise FPR over their zero-shot base because the prompted model tends to flag an entry as soon as a single database returns no match: tool budgets trade precision for recall, and the penalty is partly dev_public-distribution-specific (§Ë6). Among independent tools, Opus 4.7 and Sonnet 4.6 form the low-FPR, best-calibrated frontier; the recall-aggressive open-weight cohort detects more hallucinations but flags far more valid entries. These dispositions are measured under one shared prompt: the ranking survives a paraphrase ablation (§ËH.1), but we did not tune prompts per model, so individual operating points may shift under targeted prompt engineering. Figure 2: DRâFPR Pareto frontier on dev_public. Each point is a (tool, configuration) pair. The dotted line traces the Pareto front (DRâ , FPRâ ). Independent zero-shot LLMs occupy the precision-end of the front (Sonnet 4.6, Opus 4.7, Gemini 2.5 Pro); the recall-end is occupied by recall-aggressive open-weight models (DeepSeek-V3.2, Qwen3-VL-235B). Agentic harnesses sit above the LLM zero-shot points on DR but to the right on FPR. bibtex-updater (v1.2.0) flags conservatively and sits at the precision corner (DR 0.865, FPR 0.092; Tab.Ë32): its low FPR is structural, arising from conservative matching rather than from the abstention it adds on unverifiable entries. 5.3 Per-tier analysis DOI-only detection concentrates in Tier 1; the LLMs hold up across tiers and degrade gracefully (Fig.Ë6). The recallâprecision tradeoff persists at every difficulty: aggressive models reach their high Tier 1 detection by flagging indiscriminatelyâtheir FPR is highârather than by sharper discrimination, and stay far apart from the conservative models on Tier 3. (Tier 3 aggregates fold in the stress-test type arxiv_version_mismatch, so per-type sums in Tab.Ë15 do not average exactly to the per-tier numbers.) 5.4 Per-type analysis Per-type detection rates (heatmap in Fig.Ë4; full table in Tab.Ë15) show a clear cohort pattern. GPT-5.1 reaches 74% on 9 of 11 main types and is perfect on nonexistent_venue and placeholder_authors. Its two weak types, author_mismatch and near_miss_title, both turn on exact bibliographic knowledge and map to sub-tests A (authors match) and T (title exists) respectively (Tab.Ë1). At nâ30nâ 30 per type the minimum detectable effect is 20â26 p (§ËB.1), so within-type cell rankings are directional, not significant. 5.5 Stage-2 diagnosis cascade We compose bibtex-updater (Stage 1) with a Claude Sonnet 4.6 diagnoser (Stage 2; up to five tool calls per entry) into a two-stage cascade. Stage 1 emits verified, a typed hallucinated verdict, or uncertain; only the uncertain bucket is forwarded to Stage 2. All cascade results inherit bibtex-updaterâs co-design caveat (§ËG.2). We report two stances: conservative keeps the entries Stage 2 still cannot resolve (its residual uncertain verdicts) as abstentions; aggressive flags any entry not affirmatively verified by either stage, assigning it hallucinated at confidence 0.550.55. Aggressive promotion raises Tier-3 F1 by one to two points and pushes recall toward 1.0, at a few points of added FPR and a âŒ1.5 1.5 p Area Under the ROC Curve (AUROC) drop from collapsing confidences to the fixed 0.550.55 (Tab.Ë4). Table 4: Cascade results, conservative vs aggressive scoring of residual uncertain. Stage 2 is Claude Sonnet 4.6 via OpenRouter on a dated snapshot (2026-05-31), later than the main-run zero-shot rows (2026-05-05) and subject to the same endpoint drift (§Ë7). Decimal cells drop the leading â0.â; âââ marks splits with no scored valid entries (stress_test is evaluated on its 121 hallucinated entries; the splitâs single valid entry is the contamination canary). T3-F1 is F1 on Tier-3 (hard) hallucinations. Stage 1 is bibtex-updater v1.2.0, matching the co-designed row in Tab.Ë3. Split Mode DR â FPR â F1 â TW-F1 â T3-F1 â AUROC â stress_test cons. .951 â .975 .972 .955 â (n=121)(n=121) agg. .959 â .979 .976 .957 â test_public cons. .990 .112 .957 .973 .834 .951 (n=831)(n=831) agg. .992 .160 .950 .971 .845 .936 dev_public cons. .996 .108 .947 .970 .800 .952 (n=1119)(n=1119) agg. .997 .148 .939 .968 .821 .938 Takeaway. The cascade is the high-recall, low-FPR configuration: Stage 2 resolves most of Stage 1âs uncertain bucket instead of leaving it unverified, lifting recall well above standalone bibtex-updaterâs at a comparable false-positive rate (Tab.Ë4). Choose conservative for the lowest FPR; aggressive adds a small hard-tier recall boost when the FPR budget allows. 6 Diagnosing the three failure modes Table 5: Three failure modes that bound LLM-verifier deployment. Overview results on Hallmark; with full tables referenced. â1-in-Nâ is the precision at a venue-realistic âŒ2% 2\% hallucination rate (one true hallucination per N flags). Failure mode What breaks Evidence Implication (i) Agentic lookups, diminishing returns (§Ë5.2) The prompted model tends to flag an entry as soon as any one database returns no match, so partial database coverage becomes a false positive. A 5-call budget lifts recall past the conservative rule-based reference (DR .97â.99 vs. .87) but at âŒ5Ă 5Ă its false-positive rate (.43â.48 vs. .09); the rise comes from the harness (any-no-match flagging), not the base model. More tool calls buy recall, not precision. (i) Base-rate precision drop (§Ë6) At a venue-realistic âŒ2% 2\% base rate, precision is governed by FPR (Bayesâ rule), not recall. FPR spans .05â.70 across verifiers, a âŒ7Ă 7Ă precision gap at a 2%2\% base rate: low-FPR tools reach 1-in-6 to 1-in-9 flags, high-FPR open-weight models 1-in-35 to 1-in-39. FPR, not recall, is the deployment-decisive lever. (i) Post-cutoff calibration breakdown (§Ë6) Most LLMs over-flag papers past their training cutoff: âflag everything unfamiliarâ. On 2024â2025 papers, 8 of 12 LLMs degrade sharply (FPR .59â.89), two over-flag only moderately, and only the two latest-cutoff models hold (Sonnet 4.6 .12, Opus 4.7 .07), confounded with possible recall. Verifier trust expires with the training cutoff. Tab.Ë5 collects the three failure modes that bound LLM-verifier deployment: Failure mode (i), agentic FPR inflation, is the agentic block of Tab.Ë3 (§Ë5.2); failure mode (i), the base-rate precision drop, is the Positive Predictive Value (PPV) analysis below (§Ë6); failure mode (i), post-cutoff calibration breakdown, is the Temporal robustness analysis below (Tab.Ë25). Pre-screening, calibration, and deployment PPV. That precisionâthe positive predictive value (PPV), the fraction of flagged entries that are genuine hallucinationsâfalls under false positives at low prevalence follows directly from Bayesâ rule (§ËE.1) and is no discovery of ours; what Hallmark measures is how far apart the verifiers sit on the false-positive rate that governs it. False-positive rates range an order of magnitude across the cohortâfrom 0.0500.050 to 0.7020.702âand at a venue-realistic âŒ2% 2\% base rate that becomes a âŒ7Ă 7Ă gap in precision: the best verifiers catch one true hallucination per 6â9 flags, the most aggressive fewer than one in 35, even at â„87%â„87\% detection (per-tool PPV in Tab.Ë21, sweep in Tab.Ë22). Calibration tracks the same axis: Sonnet 4.6, Opus 4.7, and GPT-5.1 are the best-calibrated independent tools (ECE 0.0660.066, 0.1120.112, 0.1900.190), while the over-flagging open-weight models reach 0.180.18â0.320.32. The co-designed bibtex-updater joins the low-FPR tier from the other directionâconservative flagging rather than abstentionâreaching one-in-six precision too (§ËE.2). The optional pre-screening layer adds âŒ5 5 p overall (âŒ18 18 p on Tier 1); we report numbers with and without it. Every abstention-excluded metric is paired with its coverage and an aggressive re-flagging strategy, so deferring hard cases does not mask weak performance in precision. Key takeaway. At venue-realistic prevalence, rank verifiers by false-positive rate and calibration, not recall: a few points of FPR decide whether the flagged entries are worth reviewer attention or are dominated by false alarms. Published findings pin the base rate only loosely: incident studies report affected-paper counts rather than rates (GPTZero: 53 of 4,841 NeurIPS 2025 papers [Shmatko et al., 2026]), and even the implied paper-level rate sits well above the entry-level rate the precision calculation needs, so prevalence should be treated as an estimate (for an ablation, see Tab.Ë22). FPR reduction, not recall maximization, is the deployment-decisive lever: prevalence moves deployability while leaving the ranking fixed. Sweeping the base rate over 11â5%5\% leaves the PPV ordering essentially fixed (Tab.Ë22), because prevalence cancels from pairwise comparisons and FPR governs the order. What prevalence moves is the absolute precision: at a 2%2\% base rate even the best verifier reaches only âŒ18% 18\% PPV. The order is also operating-point-robust: confidences are quantized enough that threshold tuning buys almost nothing (default-to-best-F1 gap â€0.3â€0.3 p for all but Gemini 2.5 Flash), so the fixed-0.50.5 point is near-optimal (§ËH.2). When to prefer recall. The operating point is set by three quantities: prevalence, the cost ratio cFN/cFPc_FN/c_FP of a missed fabrication versus a false alarm, and reviewer capacity. FPR-ranking holds in the reviewer-bound venue-audit regimeâlow prevalence, finite reviewer attention, a false alarm cheap to dismissâwhere precision is the limiting factor. When a missed fabrication is far costlier than triaging extra flags, or when a later human-review stage filters out the false alarms the tool raises, recall plus human triage dominates and a high-recall tool (an agentic harness, or DeepSeek-V3.2) is preferable despite its FPR. The two takeawaysâârank on FPRâ and âa missed hallucination is the worse errorââare not in tension; they name different aspects of the (prevalence,cFN/cFP,capacity)(prevalence,c_FN/c_FP,capacity) space, which Tab.Ë6 resolves into tool choices. Table 6: Regime-conditional deployment guidance. The right verifier depends on prevalence, the cost of a missed fabrication relative to a false alarm (cFN/cFPc_FN/c_FP), and reviewer capacity; per-tool numbers in §ËE.3. Regime Pick Why Reviewer-bound venue auditâlow prevalence, a false alarm cheap to dismiss Lowest-FPR, best-calibrated verifier (bibtex-updater, Opus 4.7, Sonnet 4.6) At low prevalence false alarms outnumber true detections even for the best verifier (precision âŒ18% 18\% at a 2%2\% base rate), so a few points of FPR decide usability; add the Stage-2 cascade when recall must rise at the same FPR. High prevalenceâtriage queues, pre-submission self-check High-recall verifier (an agentic harness or DeepSeek-V3.2) Precision improves as the base rate rises, so catching more matters and the extra false alarms are affordable. Costly miss, or downstream human review absorbs false alarms High-recall verifier ++ human triage When a fabrication in print far outweighs an extra flag, recall plus a human filter is a better choice. Costâaccuracy tradeoff. LLM verification is also two orders of magnitude more expensive than the rule-based tools, which run locally for a fraction of a cent per entry; single-call LLMs cost cents, and agentic variants add another 22â5Ă5Ă without proportional gains (Fig.Ë5). With the base-rate PPV drop above, the regime where LLM verification pays off is narrow: high-prevalence settings (triage queues, pre-publication self-check), or pipelines where a downstream human-review stage filters out the false positives. Synthetic vs. real-world representativeness. We can detect no surface-feature artifact separating synthetic from real hallucinations, but we cannot yet claim the two are equivalent. For nine dev_public types with both real and synthetic entries, per-type KolmogorovâSmirnov tests on title length, author count, and year fail to reject the null for 6/9 (p>0.05p>0.05); the three divergent types differ mainly in author count (real entries average 2â3 authors, synthetic 3â4). This is not evidence of equivalence: failing to reject a null is not accepting it, the real set has only âŒ5 5â2020 entries per type (so the KS test is badly underpowered), and these three surface features are exactly the ones our shortcut analysis (§ËC.4.2) shows carry little signal (a logistic regression trained on those features beats the majority-class baselineâalways predicting the more frequent labelâby only 4.64.6 p accuracy, below the 55 p margin we treat as evidence of an exploitable shortcut). The honest reading is narrow: no artifact strong enough to clear the leakage threshold appears, but distributional and semantic equivalence remains untested at these sample sizes. That is the central open validity question, which would need an equivalence test on a much larger real set, not a null-hypothesis test that cannot tell âsimilarâ from âinsufficient dataâ. Tool detection is comparable across strata (GPT-5.1: 0.660.66 on LLM-generated vs. 0.850.85 on perturbations, with Sonnet 4.6 the same direction), consistent with LLM-generated entries being intrinsically harder rather than with self-recognition. This comparability supports perturbations as controlled diagnostics and motivates expanding the 108-entry real-world set. Cross-split robustness. The held-out test_public (831 entries) tells a complementary story to dev-only metrics (the Î column of Tab.Ë3; full results in §ËC.1). âą Calibrated LLMs hold; recall-aggressive LLMs drift. Ten of twelve LLMs gain or hold F1, and FPR shifts stratify along the same axis: the precision-end cohort (Sonnet 4.6, Opus 4.7, GPT-5.4, both Geminis) moves within âŒ1 1 p, while higher-FPR models drift up (GPT-5.1 +6.9+6.9 p, the Qwen variants âŒ+8 +8 p). DeepSeek-R1âs â30.3-30.3 p is a routing artifactâit abstains on 21.7%21.7\% of test_publicânot a precision gain. All shifts sit within bootstrap-CI width at nvalid=312n_valid=312. âą bibtex-updater is cross-split stable. Its FPR rises only +2.4+2.4 p (0.092â0.1150.092â0.115) and DR and F1 hold: conservative matching flags so sparingly that test_publicâs harder valid pool does not inflate its false positives, and its flat riskâcoverage curveâthe error rate as a function of the fraction of entries it commits to, traced by varying the abstention thresholdâconfirms the low FPR is structural, arising from conservative matching rather than from abstention. âą Both agentic bibtex-updater harnesses lose part of their dev-side FPR inflation on test_public. Sonnet+bibtex-updater FPR drops â8.8-8.8 p and GPT-5.1+bibtex-updater â11.4-11.4 p cross-split, shrinking each harnessâs dev-side multiplier, but the harness-driven precision cost persists in both. Takeaway. bibtex-updater is the most cross-split-stable tool: its false-positive rate barely moves and its F1 holds, because the low FPR is intrinsic to conservative matching and does not depend on deferring hard casesâwithin the ML-venue citation regime these splits sample; a first cross-domain probe (§ËG.2) shows detection transferring but FPR rising to 0.375, so the precision claim is regime-bound. This stability also carries the co-design caveat: the benchmarkâs taxonomy was informed by bibtex-updaterâs detection capabilities, so read its cross-split stability as an upper bound, not a fair head-to-head with independent tools. Temporal robustness (failure mode i). A separate 448-entry supplement disjoint from the 2021â2023 corpus (300 valid 2024â2025 DBLP entries, 148 hallucinated; §ËF.1) shows LLM verifiers degrade sharply beyond their training cutoff. The earlier-cutoff models default to âflag everything unfamiliarâ: F1 falls to â0.55â0.55 as detection inflates, with GPT-5.1âs FPR alone rising 41.1%â75.9%41.1\%â75.9\%. DeepSeek-R1 is the degenerate case, routing nearly every entry to uncertain (DR ==F1 =0=0, FPR =0.856=0.856); Tab.Ë25 gives the full panel. We read this through the ranking, not the magnitude. A prompt-sensitivity ablation moves the same modelâs absolute FPR by 1010â3737 p on wording alone (GPT-5.1 0.580â0.2120.580â0.212; Sonnet 4.6 0.121â0.0150.121â0.015; §ËH.1), yet the cross-regime ranking barely moves (mean pairwise Spearman rank correlation Ï=0.90Ï=0.90). So the temporal claim rests on the prompt-invariant ranking rather than on any single FPR value. The pattern extends past 2024â2025: a 60-entry probe on 2026 arXiv submissions reproduces the FPR multiplier (r=â0.82r=-0.82 between baseline aggressiveness and relative post-cutoff degradation). Retrieval does not improve it necessarily: databases also lag on recent papers.222S2 / Crossref / OpenAlex exhibit weeks-to-months ingestion lag for new arXiv preprints, contributing additively to the post-cutoff FPR rise, while the arXiv API itself covers preprints immediately but not venue-published versions; details and empirical confirmation in § F.2. Mitigation and controls. A cutoff-aware prompting variant (§ËF.2) recovers most of GPT-5.1âs post-cutoff FPR (72.6%â0.0%72.6\%â0.0\% on the entries it still commits to) but inflates its abstention uniformly (pre-cutoff uncertain rises to 52.7%52.7\%), whereas on Sonnet 4.6 the same addendum abstains selectively (post-cutoff uncertain 48.9%48.9\% against 8.7%8.7\% pre-cutoff) while halving its committed-entry FPR; the failure is epistemic miscalibration, not structural blindness, and the addendumâs effect is model-dependent. The later-cutoff GPT-5.4 control shows that recency helps but does not cure over-flagging: on these 2024â2025 papers, inside its Aug 2025 window, its FPR is 41.3%41.3\% (§ËF.3), well below the earlier-cutoff over-flagging cluster yet far above the latest-cutoff models. Later-cutoff resistance. Only the two latest-cutoff Anthropic models hold their false-positive rate near in-distribution levels (Sonnet 4.6 12.0%12.0\%, Opus 4.7 7.3%7.3\%); Gemini 2.5 Pro and GPT-5.4 over-flag only moderately, and the remaining eight degrade sharply to 59.559.5â88.7%88.7\% (the âover-flagging clusterâ; Tab.Ë25). Their devâ FPR barely moves (â0.7-0.7 and +0.1+0.1 p). We report this as a descriptive pattern, not a calibration result, because the design cannot cleanly separate two explanations: the valid pool is scraped from DBLP and these are the latest-cutoff models in the cohort, so a low FPR on a valid entry is indistinguishable from training-data recall of that same DBLP record; the risk of such contaminationâbenchmark entries appearing in a modelâs training corpusâgrows with model capability, since stronger models train on larger, more recent corpora [Rank et al., 2026]. Two controls push against, but do not eliminate, the memorization reading (§ËF.4). A recall probe finds Opus 4.7 accepts 84%84\% of valid 2024â2025 papers it cannot recall, consistent with calibration rather than memorization; the recall measure is model-derivedâwe count only papers the model itself reports it cannot recallâso contamination is reduced but not excluded. Sonnet 4.6âs lower FPR is partly recall-driven. A third-provider late-cutoff control (DeepSeek-V4-Pro, complete n=300n=300 run at full coverage) holds its post-cutoff FPR at 0.360.36 where GPT-5.1 reaches 0.930.93 on the same entries, consistent with the effect spanning providers without establishing it.333The control supports the cross-provider reading but does not establish it: DeepSeek-V4-Pro matches rather than exceeds the Anthropic cutoffs, it is a single third-party model, and no recall probe was run on it, so contamination is not separable (§ F.4). The resistance points away from pure memorization for Opus 4.7, but with N=2N=2 holding models from one provider and a single third-party control, it remains a preliminary signal that stops short of a causal attribution. Takeaway. The two latest-cutoff models keep their post-cutoff FPR low while earlier-cutoff models degrade sharply, and the one third-provider late-cutoff control we ran (DeepSeek-V4-Pro) shows the same low FPR. For Opus 4.7, a recall probe points to calibration rather than memorization. We read this as a preliminary descriptive signal, not a causal finding: it rests on a single third-party control, the Anthropic runs have not yet been reproduced on Anthropicâs native API (only through the OpenRouter mirror), and the recall probe relies on each modelâs own report of which papers it cannot recall. Per-type failure modes. API tools miss venue-level hallucinations (preprint_as_published, wrong_venue) because bibliographic APIs do not distinguish venue publication from arXiv availability. GPT-5.1 handles wrong_venue (85%85\%) better than preprint_as_published (74%74\%), and struggles most with author_mismatch (45%45\%) and near_miss_title (58%58\%). Prepending bibtex-updaterâs output to GPT-5.1 cuts calibration error sharply (ECE 0.190â0.0780.190â0.078) at no recall cost: the strongest option when calibration matters more than raw recall (§ËG.2). Tier concentration. Benchmark optimization rewards easy subtasks [Hardt, 2026]. On per-tool, per-tier detection rates, API tools concentrate their wins on Tier 1 (Tier 1/3 ratio >30Ă>30Ă, Gini 0.650.65), while LLMs perform near-uniformly (Gini <0.05<0.05). We therefore report Tier 3 F1 alongside the aggregate (§Ë5.3 and 6) to keep the hard regime visible. 7 Limitations Dataset scale and coverage. The benchmark is small relative to the full diversity of citation hallucinations. At 2,526 entries, most types on dev_public and test_public clear â„30â„ 30 instancesâenough for meaningful per-type comparisonâbut a few fall below it, and test_hidden does not reach the floor at all (§ËB.1). Prevalence differs across splits (54.2â62.5%), so we emphasize the prevalence-independent metrics (DR, FPR). We cover English-language BibTeX from five ML venues (2021â2023); the taxonomy and sub-tests are designed to be domain-agnostic but validated mainly in that regime. A 500-entry cross-domain split (PubMed/bioRxiv plus non-ML CS venues), together with a recency-matched, canonically-resolved rebuild, separates domain from post-cutoff recency out of regime (§ËG.2). The separation reframes the result: the released splitâs apparent cross-domain FPR rise is mostly a recency artifactâthe LLMs flag 2026 dates as impossibleâand once entries are pre-cutoff the residual domain effect is negative or small for the calibrated verifiers, large only for a single recall-aggressive model that already over-flags in its own regime (Tab.Ë35). For bibtex-updater, canonical metadata leaves the FPR at its in-domain level while coverage falls: out of domain the cost is coverage, not precision. Two failure modes stay specific to biomedical dataâprovider safety filters block âŒ3% 3\% of valid biomedical citations for the Anthropic models, distinct from any verdict errorâwhile humanities and non-English settings remain future work. âŒ38% 38\% of valid entries lack DOIsâmany ML papers are never formally publishedâwhich inflates DOI-based FPR on those entries, an effect per-type metrics expose. Syntheticâreal gap. Most hallucinated entries are perturbation-generated, and the real-world anchor is thin. The benchmark holds 108 real-world entries (plus 280 LLM-generated), type-skewed (55% plausible_fabrication), with 15 compound-failure entries that involved judgment calls without human inter-annotator agreement. We report an automated three-rater reliability proxy (Fleissâ Îș=0.24Îș=0.24, fair inter-rater agreement; §ËH.4) that corroborates the relabel audit but does not substitute for human IAA, which remains future work. Expanding real-world coverage from retraction databases is a priority. Temporal fragility and the calibration question. LLM verifiers carry an implicit training-cutoff bias that we can only partly disentangle. §Ë6 measures the post-cutoff FPR rise and the later-cutoff resistance; a recall probe and a non-Anthropic late-cutoff control point away from pure memorization for Opus 4.7 (and only partly for Sonnet 4.6), but the recall measure is model-derived and residual confoundsâsystem-prompt handling, default temperature, native-vs-OpenRouter routingâremain. We therefore read the resistance as a descriptive signal, not a causal attribution. Controls and residual gaps. Three controls probe the calibration-versus-contamination question, none decisive (§ËF.4). (i) A recall probe on valid 2024â2025 papers is consistent with calibration over memorization for Opus 4.7 (it accepts 84%84\% of papers it cannot recall), while Sonnet 4.6âs lower FPR is partly recall-driven. (i) A third-provider late-cutoff control (DeepSeek-V4-Pro) also holds its post-cutoff FPR low relative to the over-flagging cohort (0.360.36 vs. 0.930.93 for GPT-5.1 on the same n=300n=300 subsample, at full coverage), though its cutoff matches rather than exceeds the Anthropic pairâs (§ËF.4). (i) GPT-5.1âs run-to-run variance is small (F1 std 0.0020.002 over three runs), so single-run rankings seem to be stable to sampling noise, at least for some models. The open gaps: the cross-provider control is a single third-party model, a native-API replication of the Anthropic runs is outstanding, and contamination is reduced but not fully excluded even for Opus 4.7. Compute and token budget. Compute caps the coverage of reasoning-mode and extended-thinking variants. Reasoning models cost far more per entry (DeepSeek-R1 âŒ25 25 s vs. âŒ5 5 s), so we did not run extended-thinking variants on the full benchmark (Fig.Ë7 maps where their JSON-output contract holds), the agentic harness is capped at five tool calls per entry, and HaRC/verify-citations are excluded for rate-limit reasons (§ËB.2.2). Regime conditionality and co-design. The âFPR decidesâ framing is calibrated to the reviewer-bound venue-audit regime: low prevalence, finite reviewer attention, a false alarm cheap to dismiss. The PPV ranking is itself prevalence-invariantâsweeping prevalence over 11â5%5\% does not reorder the tools (Tab.Ë22)âso what changes across regimes is whether any verifier reaches a usable absolute precision and the cost asymmetry cFN/cFPc_FN/c_FP; under high prevalence or a high cost of a missed fabrication, recall and human triage regain primacy. Separately, bibtex-updater and the taxonomy were developed in parallel, so the taxonomy reflects fundamental verification steps but the construct-overfitting risk is non-zero; we report it as a co-designed reference (§ËG.2). Whether rankings on synthetic hallucinations predict performance on real errors remains open [Hardt, 2026]; the 108 real-world entries are an initial signal, and the codebase tracks the number of times the dev split has been evaluated against during development, an adaptive-data-analysis hygiene measure that bounds the overfitting introduced by repeatedly reusing the same held-out set [Dwork et al., 2015]. Endpoint drift and reproducibility. Several ablations and the two Anthropic riskâcoverage curves were collected against the OpenRouter API on a dated snapshot that does not reproduce the main-run aggregates. The OpenRouter Anthropic endpoint drifts over time, and the per-entry Sonnet 4.6 / Opus 4.7 dev_public predictions behind the main run were summary-only and never persisted, so their coverage cells are not recoverable. Concretely, a later snapshot roughly doubles both modelsâ FPR (Opus 4.7 0.072â0.1620.072â0.162, Sonnet 4.6 0.127â0.1650.127â0.165) and diverges from the published aggregates by more than 1313 p. We report the internally consistent main-run snapshot rather than splice a drifted operating point onto pinned metrics, mark those two coverage cells ân/aâ (drift caveat in §ËE.2), and read the Anthropic temporal story through the ranking, which drift leaves largely intact because it shifts every modelâs absolute FPR in the same direction while preserving their order (the cross-regime ranking holds at Spearman Ï=0.90Ï=0.90 under prompt perturbation, §ËH.1), rather than through the absolute FPR; within-run deltas (prompt-variant, field-leave-one-out (LOO), rater-agreement) difference conditions measured against the same endpoint on the same day and so survive the drift. The released corpus is pinned at tag v1.2.0, carrying the ground-truth audit of §ËA.3. 8 Conclusion Citation integrity underwrites scientific trust, and the audit chain is only as strong as its weakest verifier. The NeurIPS 2025 incident and concurrent audits [Xu et al., 2026, Sakai et al., 2026a, Bienz et al., 2026] make citation hallucination a deployment problem at scale, yet verification tools have multiplied without a shared way to measure which verifier catches which failure. Hallmark is built for that question: a typed, tier-stratified taxonomy with diagnostic sub-tests and contamination-resistant splits localizes where a verifier breaks rather than only scoring it. It surfaces three failure modes that bound LLM-verifier deployment (Tab.Ë5). (i) Agentic lookups inflate FPR. A five-call budget pushes recall past the conservative rule-based reference, but at âŒ5Ă 5Ă its false-positive rate, because the prompted model tends to flag an entry as soon as any one database returns no match (Tab.Ë3). (i) FPR decides deployability at realistic prevalence. At audit-regime base rates (âŒ2% 2\%), low-FPR verifiers catch one true hallucination per 6â9 flags, while high-FPR open-weight models fall to one per 35 or worse (§ËE.1). (i) Most LLM verifiers degrade sharply past their training cutoff. On 2024â2025 papers, 8 of 12 LLMs over-flag sharply; GPT-5.4 and Gemini 2.5 Pro over-flag only moderately, and only the two latest-cutoff models hold their FPR near in-distribution levels. Because cutoff recency and provider pipeline are confounded, we report this as a descriptive signal, not a causal finding (§Ë6). No tool dominates across regimes; the right verifier depends on the deployment. bibtex-updater is cheapest, precision-anchored, and the most cross-split-stable, its low FPR structural (conservative matching, independent of abstention); Opus 4.7 and Sonnet 4.6 form the best-calibrated low-FPR frontier among independent tools, while Gemini 2.5 Pro reaches an even lower FPR but at much worse calibration and recall (Tabs.Ë6 and E.3). In practice: for a cheap pre-submission self-check, run bibtex-updater first (no LLM inference cost, lowest-FPR tier); when recall matters and the FPR budget allows, the two-stage cascade lifts detection to âŒ0.99 0.99 at FPR â0.11â0.11 (§Ë5.5). For deployment, false-positive rate and calibration decide usability, not recall; for trust, a verifier that lets no fabrication through is what we ultimately want. These two goals pull against each other, and that tension should guide the next generation of hallucination detectors. References Abbonato [2026] Diletta Abbonato. CheckIfExist: Detecting citation hallucinations in the era of AI-generated content. arXiv preprint arXiv:2602.15871, 2026. Agrawal et al. [2024] Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Kalai. Do language models know when theyâre hallucinating references? In EACL, pages 912â928. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.findings-eacl.62. URL https://aclanthology.org/2024.findings-eacl.62. Alkaissi and McFarlane [2023] Hussam Alkaissi and Samy I. McFarlane. Artificial hallucinations in ChatGPT: Implications in scientific writing. Cureus, 15(2), 2023. doi: 10.7759/cureus.35179. URL https://doi.org/10.7759/cureus.35179. Ansari [2026] Samar Ansari. Compound deception in elite peer review: A failure mode taxonomy of 100 fabricated citations at NeurIPS 2025. arXiv preprint arXiv:2602.05930, 2026. Bienz et al. [2026] Amanda Bienz, Carl Pearson, and Simon Garcia de Gonzalo. The case of the mysterious citations. arXiv preprint arXiv:2602.05867, 2026. Chen et al. [2021] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. dblp Team [2025] dblp Team. dblp computer science bibliography â monthly snapshot XML release of october 2025, 2025. URL https://dblp.org. DeepSeek-AI et al. [2024] DeepSeek-AI et al. DeepSeek-V3 technical report. arXiv preprint arXiv:2412.19437, 2024. Dwork et al. [2015] Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth. The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248):636â638, 2015. doi: 10.1126/science.a9375. URL https://doi.org/10.1126/science.a9375. Efron and Tibshirani [1994] Bradley Efron and Robert J. Tibshirani. An Introduction to the Bootstrap. Chapman & Hall/CRC, 1994. El-Yaniv and Wiener [2010] Ran El-Yaniv and Yair Wiener. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11:1605â1641, 2010. URL https://w.jmlr.org/papers/v11/el-yaniv10a.html. Fan et al. [2026] Dongyang Fan, Sebastien Delsad, Nicolas Flammarion, and Maksym Andriushchenko. HalluHard: A hard multi-turn hallucination benchmark. arXiv:2602.01031, 2026. URL https://arxiv.org/abs/2602.01031. Fleiss [1971] Joseph L. Fleiss. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5):378â382, 1971. doi: 10.1037/h0031619. Gao et al. [2023] Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. RARR: Researching and revising what language models say, using language models. In Annual Meeting of the Association for Computational Linguistics, pages 16477â16508. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.acl-long.910. URL https://doi.org/10.18653/v1/2023.acl-long.910. Gebru et al. [2021] Timnit Gebru, Jamie Morgenstern, Brenda Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal DaumĂ© I, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86â92, 2021. doi: 10.1145/3458723. URL https://doi.org/10.1145/3458723. Geifman and El-Yaniv [2017] Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, pages 4878â4887, 2017. URL https://proceedings.neurips.c/paper/2017/hash/4a8423d5e91fda00b7e46540e2b0cf1-Abstract.html. Ghosh et al. [2025] Adhiraj Ghosh, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao, Samuel Albanie, and Matthias Bethge. Onebench to test them all: Sample-level benchmarking over open-ended capabilities. In ACL, pages 32445â32481. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.acl-long.1560. URL https://aclanthology.org/2025.acl-long.1560/. Google DeepMind [2025] Google DeepMind. Gemini 2.5: Our most intelligent AI model, 2025. URL https://deepmind.google/technologies/gemini/. HaRC Contributors [2024] HaRC Contributors. HaRC: Hallucinated reference checker, 2024. URL https://pypi.org/project/harcx/. Hardt [2026] Moritz Hardt. The Emerging Science of Machine Learning Benchmarks. Princeton University Press, 2026. Forthcoming; manuscript available at https://mlbenchmarks.org/. Holland et al. [2020] Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. The dataset nutrition label. In Data Protection and Privacy, pages 1â25. Hart Publishing, 2020. doi: 10.5040/9781509932771.ch-001. URL https://doi.org/10.5040/9781509932771.ch-001. Hu et al. [2024] Xiangkun Hu, Dongyu Ru, Lin Qiu, Qipeng Guo, Tianhang Zhang, Yang Xu, Yun Luo, Pengfei Liu, Yue Zhang, and Zheng Zhang. RefChecker: Reference-based fine-grained hallucination checker and benchmark for large language models. arXiv preprint arXiv:2405.14486, 2024. URL https://github.com/amazon-science/RefChecker. Huang et al. [2025] Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1â55, 2025. doi: 10.1145/3703155. URL https://doi.org/10.1145/3703155. Jain et al. [2025] Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In ICLR, 2025. URL https://openreview.net/forum?id=chfJJYC3iL. Ji et al. [2023] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1â38, 2023. doi: 10.1145/3571730. URL https://doi.org/10.1145/3571730. Jimenez et al. [2024] Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues? In Proceedings of ICLR, 2024. URL https://openreview.net/forum?id=VTF8yNQM66. Kiela et al. [2021] Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al. Dynabench: Rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110â4124. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.naacl-main.324. URL https://doi.org/10.18653/v1/2021.naacl-main.324. Landis and Koch [1977] J. Richard Landis and Gary G. Koch. The measurement of observer agreement for categorical data. Biometrics, 33(1):159â174, 1977. doi: 10.2307/2529310. Li et al. [2023] Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. Halueval: A large-scale hallucination evaluation benchmark for large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6449â6464. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.397. URL https://doi.org/10.18653/v1/2023.emnlp-main.397. Min et al. [2023] Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Conference on Empirical Methods in Natural Language Processing, pages 12076â12100. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.741. URL https://doi.org/10.18653/v1/2023.emnlp-main.741. Mistral AI [2024] Mistral AI. Large enough: Mistral Large 2, 2024. URL https://mistral.ai/news/mistral-large-2407/. Naeini et al. [2015] Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using Bayesian binning. In Proceedings of AAAI, 2015. OpenAI [2025] OpenAI. GPT-5.1 Instant and GPT-5.1 Thinking system card addendum, 2025. URL https://openai.com/index/gpt-5-system-card-addendum-gpt-5-1/. Rank et al. [2026] Ben Rank, Hardik Bhatnagar, Ameya Prabhu, Shira Eisenberg, Karina Nguyen, Matthias Bethge, and Maksym Andriushchenko. PostTrainBench: Can LLM agents automate LLM post-training? arXiv preprint arXiv:2603.08640, 2026. URL https://arxiv.org/abs/2603.08640. Rao and Callison-Burch [2026] Delip Rao and Chris Callison-Burch. BibTeX citation hallucinations in scientific publishing agents: Evaluation and mitigation. arXiv preprint arXiv:2604.03159, 2026. Rao et al. [2026] Delip Rao, Eric Wong, and Chris Callison-Burch. Detecting and correcting reference hallucinations in commercial LLMs and deep research agents. arXiv preprint arXiv:2604.03173, 2026. Ravichander et al. [2025] Abhilasha Ravichander, Shrusti Ghela, David Wadden, and Yejin Choi. HALoGEN: Fantastic LLM hallucinations and where to find them. arXiv preprint arXiv:2501.08292, pages 1402â1425, 2025. doi: 10.18653/v1/2025.acl-long.71. URL https://doi.org/10.18653/v1/2025.acl-long.71. Reizinger [2025] Patrik Reizinger. bibtex-updater: Automated BibTeX verification and updating, 2025. URL https://github.com/rpatrik96/bibtexupdater. Sakai et al. [2026a] Yusuke Sakai, Hidetaka Kamigaito, and Taro Watanabe. HalluCitation matters: Revealing the impact of hallucinated references with 300 hallucinated papers in ACL conferences. arXiv preprint arXiv:2601.18724, 2026a. Sakai et al. [2026b] Yusuke Sakai, Hidetaka Kamigaito, and Taro Watanabe. HalluCiteChecker: A lightweight toolkit for hallucinated citation detection and verification in the era of AI scientists. arXiv preprint arXiv:2604.26835, 2026b. Shi et al. [2026] Kaiwen Shi, Weixiang Sun, Zheyuan Zhang, Lichao Sun, Nitesh V. Chawla, and Yanfang Ye. CiteAudit: You cited it, but did you read it? a benchmark for verifying scientific references in the LLM era. arXiv preprint arXiv:2602.23452, 2026. URL https://arxiv.org/abs/2602.23452. Shmatko et al. [2026] Nazar Shmatko, Alex Adam, and Paul Esau. GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers, 2026. URL https://gptzero.me/news/neurips/. GPTZero analysis of 4,841 accepted NeurIPS 2025 papers, published January 21, 2026. verify-citations Contributors [2025] verify-citations Contributors. verify-citations: Automated citation verification tool, 2025. URL https://pypi.org/project/verify-citations/. Wadden et al. [2020] David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. In Conference on Empirical Methods in Natural Language Processing, pages 7534â7550. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.emnlp-main.609. URL https://doi.org/10.18653/v1/2020.emnlp-main.609. Walters and Wilder [2023] William H. Walters and Esther Isabelle Wilder. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13:14045, 2023. doi: 10.1038/s41598-023-41032-5. Xu et al. [2026] Zuyao Xu, Yuqi Qiu, Lu Sun, Fasheng Miao, Fubin Wu, Xinyi Wang, Xiang Li, et al. GhostCite: A large-scale analysis of citation validity in the age of large language models. arXiv preprint arXiv:2602.06718, 2026. Yang et al. [2025] An Yang et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Appendix Appendix roadmap. The appendix follows the paperâs three failure modes. Appx.ËA covers the dataset and taxonomy and Appx.ËB the evaluation protocol, statistics, and reproducibility setup; Appx.ËC collects the core cross-split and per-type results together with the validity checks. Each failure mode then has a dedicated home: agentic aggregation (mode i) in Appx.ËD, base-rate precision and deployment (mode i) in §ËE.1, and temporal fragility with its calibration controls (mode i) in Appx.ËF. Appx.ËG documents the co-designed bibtex-updater, Appx.ËH reports the ranking-invariance ablations, and a list of abbreviations closes the appendix. Hallmark at a glance three failure modes of LLM citation verifiers (i) Agentic aggregation âŒĂ 15Ă FPR, any-no-match vs. consensus over the same sources (.73 vs. .05) Five-call harnesses reach DR .97â.99, but at FPR .43â.48: the prompted model flags as soon as a single database returns no match. The lever is a deterministic re-aggregation over three of the harnessâs four databases; the any-vs-consensus ordering is the finding, not the level. Details in Appx.ËD. (i) Base-rate precision 17.6% best PPV at a venue-realistic 2% base rate (Opus 4.7) At low prevalence, FPR governs precision (Bayesâ rule), not recall: roughly one flag in six from the best verifier is a true hallucination. Details in §ËE.1. (i) Temporal fragility .07 vs. .89 post-cutoff FPR on the 448-entry 2024â2025 supplement: Opus 4.7 vs. the cohort maximum 8 of 12 LLMs over-flag 2024â2025 papers (âflag everything unfamiliarâ); only the two latest-cutoff models hold (Opus 4.7 .073, Sonnet 4.6 .120). For papers from the past 12 months, no tested tool is reliable without cutoff-aware prompting. Details in Appx.ËF. Benchmark 2,526 entries â · 14 hallucination types â · 3 difficulty tiers â · 6 sub-tests per entry (Tab.Ë2, §Ë3.4) Reference bibtex-updater v1.2.0, co-designed and excluded from ranking (§ËG.2): DR .865 â · FPR .092 â · F1 .890 â · abstains on the âŒ18% 18\% of entries it cannot back with a record (coverage .82); runs locally for a fraction of a cent per entry, ⌠2â3 orders of magnitude cheaper than LLM verifiers (Tab.Ë3, §ËE.3). Deployment regime-conditional guidance: Tab.Ë6, §ËE.3. Decimal cells drop the leading â0.â Appendix A Dataset and taxonomy A.1 Full taxonomy details The taxonomy spans 14 hallucination types across three lookup-difficulty tiers plus a stress-test bucket. Tier 1 covers field-syntax violations detectable without external lookup (DOI format, year bounds, placeholder strings). Tier 2 requires a single external lookup to confirm a field mismatch (venue, author, preprint-vs-published). Tier 3 requires multi-source verification or near-duplicate disambiguation (title near-misses, fully plausible fabrications). The stress-test types (merged_citation, partial_author_list, arxiv_version_mismatch) are theoretically motivated compound failure modes evaluated on a separate split (§ËA.3). Tab.Ë7 provides BibTeX examples for each type, with the hallucinated field highlighted in red. Table 7: Full taxonomy with example BibTeX snippets illustrating each hallucination type. Red text indicates the hallucinated field. Tier Type Example (hallucinated field in red) 1 fabricated_doi doi = 10.9999/nips2024.1847 nonexistent_venue booktitle = Intl. Conf. on Advanced AI Systems placeholder_authors author = John Doe and Jane Smith future_date year = 2030 2 chimeric_title Real authors, title = A Novel Approach... (nonexistent) wrong_venue Real paper, booktitle = ICML (actually NeurIPS) author_mismatch Real title, author = Wrong Author List preprint_as_published arXiv paper, booktitle = NeurIPS (never published) hybrid_fabrication Valid DOI resolves, but title = ... doesnât match 3 near_miss_title title = Attention Is All You Want (vs. âNeedâ) plausible_fabrication Entirely fabricated, all fields realistic but nonexistent Stress merged_citation Authors from paper A, title from paper B, venue from C partial_author_list Real paper, author = First and Last (middle dropped) arxiv_version_mismatch arXiv preprint cited with wrong venue and shifted year A.2 Real-world incident mapping We mapped 72 real-world hallucinated citations from three documented incident studies to our taxonomy. Tab.Ë8 shows the mapping results. Table 8: Mapping of 72 real-world hallucinated citations to Hallmark taxonomy types. Citations were sourced from GPTZeroâs NeurIPS 2025 analysis, GhostCite, and HalluCitation. Taxonomy type Tier Count plausible_fabrication 3 40 fabricated_doi 1 10 chimeric_title 2 8 near_miss_title 3 5 wrong_venue 2 4 author_mismatch 2 3 hybrid_fabrication 2 2 Total mapped 72 Most real-world citations map to a single taxonomy type: of the 72 in this table, 57 (79%) map directly, while the remaining 15 exhibit compound failure modes and take their highest-tier applicable type. The benchmark contains 108 real-world entries in total across all splits; this table covers the 72 drawn from three documented incident studies, and the remaining 36 come from additional documented incidents. plausible_fabrication dominates (55%), reflecting the predominant LLM failure mode: models generate coherent but entirely fictional references rather than subtly corrupting real ones. Three taxonomy typesâmerged_citation, partial_author_list, and arxiv_version_mismatchâhave zero real-world instances, which motivates their designation as stress-test types (§Ë3.2). A.3 Construction details DBLP scraping. Valid entries were scraped from the DBLP API (dblp.org/search/publ/api) using venue-specific queries for NeurIPS, ICML, ICLR, AAAI, and CVPR. We retrieved BibTeX records, verified DOI resolution via CrossRef, and confirmed title existence in Semantic Scholar. Entries failing any verification step were excluded. Perturbation pipeline. Systematic perturbations follow deterministic rules per hallucination type: âą fabricated_doi: Replace DOI with a non-resolving DOI using one of 20 non-existent prefixes and four suffix styles (path, identifier, year-indexed, and conference-indexed) to avoid template-detectable patterns. âą nonexistent_venue: Replace venue with an LLM-generated plausible but nonexistent conference name. âą placeholder_authors: Replace author list with common placeholder names. âą future_date: Set year to current year + 5. âą chimeric_title: Keep authors from paper A, replace title with LLM-generated plausible title. âą wrong_venue: Keep all fields but swap venue with a different real venue. âą author_mismatch: Keep title and venue, replace authors with those from a different paper. âą preprint_as_published: Take an arXiv-only paper and add a fabricated venue field. âą hybrid_fabrication: Keep a valid DOI but replace title and authors with fabricated metadata. âą near_miss_title: Modify 1â2 words via six strategies: synonym substitution (POS-safe pairs), plural/singular flipping, British/American spelling swap, abbreviation expansion/contraction (e.g., âRLâ â âReinforcement Learningâ), hyphenation toggling (e.g., âself-supervisedâ â âself supervisedâ), and article removal. âą plausible_fabrication: Template-based combinatorial generation of a complete, realistic but nonexistent entry (random author combinations, plausible titles, real venues). A separate set of 113 LLM-generated entries (Tab.Ë16) provides ecological validity but uses a different generation method. âą merged_citation: Combine metadata from 2â3 real papers into one entry (e.g., authors from paper A, title from paper B, venue from paper C). âą partial_author_list: Take a real paper and drop one or more middle co-authors, keeping only the first and last. âą arxiv_version_mismatch: Cite an arXiv preprint with a reassigned venue and shifted publication year. Classified as a stress-test type since no real-world instances have been observed. LLM-generated entries. We prompted GPT-5.1 to generate plausible but fictional citations for types requiring coherent fabrication (plausible_fabrication, chimeric_title, fabricated_doi). Each prompt requested a BibTeX entry with realistic metadata for a specified ML venue and year range. Generated entries were verified against CrossRef and DBLP: entries with title similarity â„85%â„ 85\% (token-sort ratio) and author Jaccard similarity â„0.5â„ 0.5 to any real paper were flagged as potential duplicates of real work and excluded. This filtering removed approximately 12% of generated entries, documenting an LLM recall failure rate where the model reproduces real papers rather than fabricating new ones. The remaining entries were assigned sub-test labels based on their hallucination typeâs expected failure pattern. Real-world collection. We harvested 72 hallucinated citations from three documented incident studies: GPTZeroâs NeurIPS 2025 analysis, GhostCite [Xu et al., 2026], and HalluCitation [Sakai et al., 2026a]. Each entry was mapped to the closest taxonomy type based on its failure mode (e.g., a citation with a non-resolving DOI was classified as fabricated_doi; a citation to a non-existent paper with plausible metadata was classified as plausible_fabrication). The real-world sample is type-skewed: 55% are plausible_fabrication, reflecting the predominant LLM failure mode in practice. Types without real-world examples (merged_citation, partial_author_list, arxiv_version_mismatch) are relegated to a separate stress-test split, evaluated independently from the main taxonomy. Adversarial crafting. We manually constructed entries designed to evade specific detection strategies. These include entries with DOIs that resolve to unrelated papers (hybrid_fabrication), entries combining metadata from multiple real papers (merged_citation), and entries with plausible but non-existent venues chosen to be close to real venue names. Adversarial entries stress-test tool robustness beyond what template-based perturbation achieves. Quality control. Every generated entry passes through automated validation: (1) BibTeX well-formedness check (all required fields present, valid syntax), (2) sub-test label consistency (sub-test ground truth matches the hallucination typeâs expected failure pattern), (3) cross-validation with the valid entry pool to prevent accidental duplicates. Extension splits. Four extension splits are released alongside the main splits, evaluation-only and outside the dev/test/hidden partition (Tab.Ë2). The temporal probe (60 entries: 30 valid / 30 hallucinated) is the small pre-supplement probe of post-cutoff behavior. Its valid half pairs 15 papers from 2024 with 15 from 2026, the latter scraped from arXiv ML categories via the arXiv API (scripts/probe_temporal_robustness.py); its hallucinated half (18 from 2024, 12 from 2026) combines 21 perturbation and 9 adversarial entries. The 448-entry temporal supplement (§ËF.1) validates the probeâs findings at scale; the probe additionally covers 2026 arXiv submissions. The cross-domain split (test_crossdomain, 500 entries: 200 valid / 300 hallucinated) probes transfer outside the ML-venue regime the main splits sample: 299 biomedical entries (156 PubMed, 143 bioRxiv) and 201 from non-ML CS venues (OSDI, CCS, CHI, USENIX Security, KDD, SIGGRAPH, and others). Valid entries are scraped; hallucinated entries follow the perturbation pipeline across the taxonomy, including the stress types, with 80/140/80 entries in Tiers 1/2/3; pre-screening uses reference_year 2026 for this split. bibtex-updaterâs evaluation on this split is in §ËG.2. The ChatGPT-citation supplement (341 entries: 172 valid / 169 hallucinated) converts the hand-coded corpus of Walters and Wilder [2023] into authentic, multidisciplinary ChatGPT hallucinations; its construction and results are in §ËC.5. A.4 Datasheet for Hallmark Following Gebru et al. [2021] and Holland et al. [2020], we provide a datasheet for the Hallmark dataset. Motivation. Hallmark was created to provide a standardized benchmark for evaluating citation hallucination detection tools, motivated by the NeurIPS 2025 incident and subsequent audits. Composition. The public release contains 2,072 BibTeX entries: 826 valid entries scraped from DBLP and arXiv and 1,246 hallucinated entries generated through perturbation, LLM generation, adversarial crafting, and real-world collection across 14 hallucination types. A held-out test_hidden split contains 454 additional entries (per-entry labels withheld; aggregate label counts in Tab.Ë2), bringing the grand total to 2,526. Split counts: dev_public 1,119 (513 valid / 606 hallucinated), test_public 831 (312 / 519), stress_test 122 (1 / 121). The dev/test/hidden split is deterministic: stratified by hallucination type and tier with seed=8042 (recorded in metadata.json), so the partition is regenerable from the released corpus. Each entry includes 6 binary sub-test labels. Four extension splits are released alongside the 2,526-entry benchmark, evaluation-only and outside the dev/test/hidden partition: a 60-entry temporal probe, a 448-entry temporal supplement (2024â2025), a 500-entry cross-domain split, and a 341-entry supplement of authentic ChatGPT-generated citations converted from Walters and Wilder [2023] (Tab.Ë2; the ChatGPT supplementâs construction in §ËC.5). Collection process. Valid entries were scraped from the DBLP API and verified against CrossRef and Semantic Scholar. Hallucinated entries were generated using the methods described in §Ë3.3 and §ËA.3. Preprocessing/Cleaning/Labeling. BibTeX records were normalized to a consistent field ordering. Unicode characters were preserved. Entries were split into dev/test/hidden sets by stratified sampling across hallucination types and tiers (deterministic seed: see Composition). Labels were assigned deterministically by the generation pipeline; valid entries carry ground-truth sub-test labels derived from database cross-checks. The released corpus is post-relabel: a systematic ground-truth audit (scripts/relabel_ground_truth.py, commits 32fe9d6 and 1475a2f) re-resolved entries against live databases and corrected real papers that had been mislabeled HALLUCINATED; the per-entry flip log is released at results/reviewer_experiments/relabel_flips.json. Recommended uses. Hallmark is intended for evaluating and comparing citation verification tools, ablating individual verification sub-tests, and cross-tool ranking via tier-weighted F1 and Plackett-Luce infrastructure. Non-recommended uses. Hallmark should not be used to train hallucination generators, to produce convincing fake citations, or for any purpose that violates the MIT license. Distribution. The dataset is distributed under the MIT license (https://opensource.org/licenses/MIT) via GitHub at https://anonymous.4open.science/r/hallmark/ (anonymised for review). The content release reported throughout this paper is pinned to repository tag v1.2.0, which fixes the byte-level split contentsâthe core splits are byte-identical to the post-relabel v1.1.1, and v1.2.0 adds the evaluation-only extension artifactsâand a manifest.json carrying sha256 digests pins the precomputed baseline-result files under data/v1.0/baseline_results/. A Croissant 1.0 metadata record is included in the repository root. The test_hidden split is not publicly distributed. Maintenance. The benchmark is maintained by the authors and accepts community contributions through pull requests validated by automated checks for BibTeX well-formedness, sub-test consistency, and duplicate detection. Two version axes are tracked separately. Each entry carries a schema_version field (1.0) recording the record structure; the schema is stable and the data/v1.0/ directory name reflects it. The content releaseâwhich corrections update as point releasesâis identified by the version field in metadata.json (1.2.0; the ground-truth relabel landed in 1.1.1) and the matching repository tag v1.2.0. The co-designed reference tool bibtex-updater is pinned to git tag v1.2.0 throughout: both the standalone results (§ËG.2) and the cascade Stage 1 (Tab.Ë4). The cascade Stage 2 is a dated snapshot (Sonnet 4.6 via OpenRouter, 2026-05-31), as LLM endpoints drift (§Ë7). The released per-entry bibtex-updater verdicts (results/relabel_delta/btu_v1_2_0/bibtexupdater_dev,test_public_per_entry.jsonl alongside data/v1.0/baseline_results/bibtexupdater_dev,test_public.json) let a reader recompute the bibtex-updater aggregates (DR 0.865, FPR 0.092) offline, without re-installing the tool or re-querying external databases. Robustness ablations collected against third-party inference endpoints carry a snapshot date and an endpoint-drift caveat (§Ë7), so they are reproducible as dated within-run deltas rather than as absolute re-runs. Ethical and legal considerations. Hallmark contains only bibliographic metadata; it includes no personally identifiable information beyond author names already present in published bibliographic records. No human subjects are involved and no IRB approval is required. Appendix B Evaluation protocol and reproducibility B.1 Metrics and statistical analysis Tier-weighted F1. For entries ei\e_i\ with tier weights wiâ1,2,3w_iâ\1,2,3\, predictions y^i y_i, and labels yiy_i: TW-Precision =âiwiâ â[y^i=yi=H]âiwiâ â[y^i=H], = _iw_i·1[ y_i=y_i=H] _iw_i·1[ y_i=H], TW-Recall =âiwiâ â[y^i=yi=H]âiwiâ â[yi=H], = _iw_i·1[ y_i=y_i=H] _iw_i·1[y_i=H], (1) where H denotes the hallucinated class, and TW-F1 is their harmonic mean. Valid entries carry no difficulty tier, so false positives contribute uniform weight 1.0 in the TW-Precision denominator. Because false positives carry uniform weight while true positives are tier-weighted, TW-F1 can be inflated by aggressive over-flagging; MCC provides a complementary prevalence-corrected view. Expected Calibration Error. We partition predictions into B=10B=10 equal-width confidence bins and compute: ECE=âb=1B|Sb|Nâ|accâ(Sb)âconfâ(Sb)|,ECE= _b=1^B |S_b|N |acc(S_b)-conf(S_b) |, (2) where SbS_b is the set of predictions in bin b, accâ(Sb)acc(S_b) is the fraction of correct predictions, and confâ(Sb)conf(S_b) is the mean confidence. Adaptive (equal-mass) binning is available in the codebase. Baseline integration. Hallmark provides a baseline registry that supports discovery, availability checking, and dispatch for all integrated tools. New baselines register via a decorator pattern, specifying their dependencies and whether they require API keys. A weekly CI workflow re-runs the DOI-only baseline in full and a 50-entry stratified verify-citations sample on the dev split, and validates the pre-computed rate-limited baselines by checksum. The stress_test split contains 122 entries: 121 hallucinated plus a single valid canary entry planted for contamination detection; with one valid entry FPR is not meaningful, so we report only Detection Rate. The statistical procedures below underlie every numerical claim in the paper. Bootstrap CIs accompany the aggregate metrics in Tab.Ë3 and the per-type breakdown in Tab.Ë15; paired bootstrap p-values support pairwise tool comparisons in §Ë5.2; the per-type power analysis (Tab.Ë9) bounds which type-level differences in Tab.Ë15 are statistically distinguishable. Bootstrap confidence intervals. The aggregate metrics in Tab.Ë3 are accompanied by 95% confidence intervals computed via stratified bootstrap with 10,000 resamples (seed=42), wherever a toolâs full-coverage per-entry predictions are available. Stratification by hallucination type ensures that the resampled datasets preserve the original type distribution, preventing bootstrap bias from underrepresented types. For reference, GPT-5.1âs dev_public 95% CIs are: DR [0.812, 0.861]; F1 [0.747, 0.785] (point estimates 0.837 / 0.766); full intervals for every full-coverage baseline and metric are available in the released evaluation artifacts. Three tools are summary-only on dev_publicâDOI-only and both Anthropic models (Claude Sonnet 4.6, Claude Opus 4.7)âbecause their released dev_public aggregates were reconstructed from confusion-matrix counts rather than stored per-entry predictions; we report point estimates only for these three and do not attach a CI or a paired test on that split. bibtex-updaterâs full-coverage per-entry verdicts are released on both splits and recompute its point estimates exactly; as a co-designed tool it is reported separately from the independent-tool comparison (Tab.Ë32), so we give its point estimates without a paired bootstrap. For the two Anthropic models we re-ran the full 1,119-entry dev_public split to recover per-entry predictions, but the OpenRouter Anthropic endpoint had drifted since the original run: a controlled replay on shared inputs shows only 90%/75% Opus/Sonnet label agreement, so the re-run diverges from the published aggregates by more than 13 p and we do not adopt its intervals. We release the drifted re-run as a reproducibility artifact rather than substitute it for the published point estimates. Significance testing. We use paired bootstrap tests [Efron and Tibshirani, 1994] to compare tools pairwise. For each pair of tools (A,B)(A,B), we resample entries with replacement and compute Î=F1AâF1B =F1_A-F1_B for each resample. The p-value is the fraction of resamples where Îâ€0 †0. The only headline pair for which both tools carry full-coverage per-entry predictions on a shared split is Sonnet 4.6 vs. Opus 4.7 on test_public: Sonnetâs F1 (0.866, CI [0.847, 0.884]) exceeds Opusâs (0.846, CI [0.828, 0.863]) by ++0.020, one-sided p=0.024p=0.024, two-sided p=0.049p=0.049, Cohenâs h=0.057h=0.057: borderline-significant at α=0.05α=0.05 with a negligible effect size; a different split or seed could plausibly flip the direction. The Sonnet/Opus comparison on dev_public involves the summary-only Anthropic dev_public cells, and the bibtex-updater-vs-Sonnet F1 comparison sits across the co-design boundary, so we report those as point-estimate differences without a paired p-value. Per-type power analysis. Per-type counts cap the resolution of type-level comparisons. Most types carry âŒ30 30 instances on dev_public and test_public; a few sit below 30 because the ground-truth audit reassigned some real papers out of their hallucination-type buckets (see Tab.Ë2 caption). For detection rate, a type with 30 hallucinated entries and 90% true detection rate has a 95% Clopper-Pearson confidence interval of [0.74,0.98][0.74,0.98], compared to [0.55,0.998][0.55,0.998] at n=10n=10. Type-level power is modest, so only large per-type gaps are detectable. Tab.Ë9 reports the minimum detectable effect (MDE) at 80% power (α=0.05α=0.05, two-sided z-test) for each hallucination type in the dev_public split. A type with n=30n=30 has an MDE of 25.6p, so tool differences smaller than this cannot be reliably detected; larger types (e.g., plausible_fabrication, n=82n=82) reach an MDE of 15.4p. These values bound which per-type comparisons are statistically meaningful. Table 9: Per-type minimum detectable effect (MDE) at 80% power on dev_public. Assumes baseline detection rate of 50% (worst-case MDE). Tier Type n MDE (p) 1 fabricated_doi 39 22.4 nonexistent_venue 32 24.7 placeholder_authors 34 24.0 future_date 31 25.1 2 chimeric_title 35 23.7 wrong_venue 35 23.7 author_mismatch 34 24.0 preprint_as_pub. 31 25.1 hybrid_fabrication 46 20.6 3 near_miss_title 35 23.7 plausible_fabrication 82 15.4 Tier weight sensitivity. We evaluate tier-weighted F1 under five weighting schemes: uniform 1,1,1\1,1,1\ (equivalent to standard macro-F1 across tiers), linear 1,2,3\1,2,3\ (default), quadratic 1,4,9\1,4,9\, log logâĄ(2),logâĄ(3),logâĄ(4)\ (2), (3), (4)\, and inverse difficulty (weights proportional to 1âmean DR1-mean DR across tools). Tab.Ë10 reports TW-F1 for the three reference tools across all five schemes. The relative ordering of tools (bibtex-updater >> GPT-5.1 >> DOI-only) is preserved under all weighting schemes, confirming that our conclusions are robust to the specific choice of tier weights. bibtex-updaterâs narrow post-relabel sensitivity range (0.890â0.921) reflects its consistently strong per-tier performance; DOI-only shows wider variation (0.255â0.392) because its detection is concentrated in Tier 1. Table 10: Tier-weighted F1 under five weighting schemes (dev_public). The bibtex-updater column is post-relabel and its default (linear) value matches Tab.Ë3; the GPT-5.1 and DOI-only columns are the pre-relabel snapshot (their per-entry predictions were not retained for a post-relabel sweep), whose default-scheme TW-F1 shifts to 0.8220.822 and 0.3290.329 post-relabel (Tab.Ë3). The tool ordering bibtex-updater >> GPT-5.1 >> DOI-only is preserved under every scheme. Weighting scheme DOI-only GPT-5.1 bibtex-updater Uniform 1,1,1\1,1,1\ 0.392 0.875 0.890 Linear 1,2,3\1,2,3\ (default) 0.310 0.860 0.908 Quadratic 1,4,9\1,4,9\ 0.255 0.850 0.921 Log logâĄ2,logâĄ3,logâĄ4\ 2, 3, 4\ 0.337 0.864 0.902 Inverse difficulty 0.342 0.865 0.891 Range (max â- min) 0.137 0.025 0.031 Takeaway. Read the paperâs numbers at the resolution the statistics support: per-type cells carry an MDE of 15â26 p (Tab.Ë9), and the one headline pair with a full paired testâSonnet 4.6 vs. Opus 4.7 on test_publicâseparates by 2.02.0 p F1 at a borderline two-sided p=0.049p=0.049 with a negligible effect size. Aggregate rankings are the stable quantity: the tool ordering survives all five tier-weighting schemes (Tab.Ë10), whereas small per-type and per-pair gaps are point estimates. B.2 Evaluation implementation All baselines are implemented as Python wrappers conforming to the Hallmark baseline interface. Each wrapper: (1) converts Hallmark BenchmarkEntry objects to the toolâs expected input format, (2) invokes the tool, (3) maps the toolâs output to a Hallmark Prediction with label, confidence, and reason (Fig.Ë3). Citation BibTeX entry Pre-screening DOI format year bounds name heuristics Tool call API / LLM Metrics DR, FPR, TW-F1, ECE Override direct verdict passflagshared layer Figure 3: Evaluation pipeline. Each citation entry passes through a shared pre-screening layer (DOI format, year bounds, name heuristics) that may emit a final verdict directly; otherwise the entry is forwarded to the tool. Metrics aggregate over all entries, including pre-screening overrides. The shared layer means every toolâs reported FPR includes pre-screening false positives; see Tab.Ë3. B.2.1 Pre-screening layer specification The pre-screening layer implements three checks that run before external tool invocation: 1. DOI format validation: Checks that DOI strings match the expected format (10.X/...) and that the DOI prefix corresponds to a known registrant. 2. Year bounds checking: Flags entries with publication years in the future or before 1900.444Genuinely valid pre-1900 references exist, of course (classical mathematics and physics citations); the bound is a corpus-tuned heuristicâno entry pool in the benchmark carries a publication year before 2021âso deployments verifying older literature should relax or drop this check. 3. Author name heuristics: Detects common placeholder patterns (âJohn Doe,â âA. Author,â single-word author names, repeated names). Pre-screening results are tagged with [Pre-screening override] in the reason string to maintain transparency about which detections come from the pre-screening layer vs. the external tool. B.2.2 HaRC and verify-citations: rate-limit disclaimer We attempted HaRC [HaRC Contributors, 2024] and verify-citations [verify-citations Contributors, 2025] as full-coverage baselines but Semantic Scholar API rate limits cap their effective coverage at <<7% even with paid keys. Re-running HaRC with an authenticated S2 key (1 RPS, 5-hour budget) leaves 0/1,119 entries actually checked by harcx: âŒ90 90 s per-entry latency (DBLP/Google Scholar fallbacks that bypass S2) exceeds the 30-min per-batch timeout. verify-citations exhibits the same throughput pathology (6.3% coverage, 71/1,119 entries processed). At this coverage level, neither toolâs reported metrics reflect its intrinsic verification capability: they reflect the shared pre-screening layer (§ËB.2.1) applied to the entries that happened to complete before timeout. Both tools are therefore excluded from Tab.Ë3. For completeness, HaRCâs pre-screening-only metrics on the 20/1,119 entries it did process (1.8% coverage) are DR = 0.209, FPR = 0.045, F1 = 0.335; for verify-citations we do not report per-metric numbers, since at 71/1,119 entries (6.3% coverage) they would characterize the timeout survivors, not the tool. B.2.3 Baseline wrappers and prompt template LLM prompt template. The LLM-based baselines (GPT-5.1, Claude, and OpenRouter models) use a zero-shot prompt with no few-shot examples. The exact template, reproduced verbatim from hallmark/baselines/llm_verifier.py (VERIFICATION_PROMPT; bibtex is the substituted entry), is: ⏠You are a citation verification expert. Analyze the following BibTeX entry and determine if it is a VALID real publication or a HALLUCINATED (fabricated) citation. BibTeX entry: âbibtex bibtex â Consider: 1. Is the title plausible and does it match known work by these authors? 2. Are the authors real researchers in this field? 3. Is the venue (journal/conference) real? 4. Does the year make sense? 5. If a DOI is present, does it look properly formatted? When the entry is HALLUCINATED, classify the hallucination mode using exactly one of: âfabricated_doiâ, ânonexistent_venueâ, âplaceholder_authorsâ, âfuture_dateâ, âchimeric_titleâ, âwrong_venueâ, âswapped_authorsâ, âpreprint_as_publishedâ, âhybrid_fabricationâ, ânear_miss_titleâ, âplausible_fabricationâ, âmerged_citationâ, âpartial_author_listâ, âarxiv_version_mismatchâ. Brief definitions: - âfabricated_doiâ: DOI does not resolve / is invented. - ânonexistent_venueâ: venue/journal does not exist. - âplaceholder_authorsâ: authors are placeholders ("Author1", "et al." alone, etc.). - âfuture_dateâ: year is in the future relative to publication. - âchimeric_titleâ: title combines fragments from multiple real works. - âwrong_venueâ: real paper but cited at wrong venue. - âswapped_authorsâ: authors swapped or mismatched against the real paper. - âpreprint_as_publishedâ: arXiv preprint cited as published in a venue. - âhybrid_fabricationâ: real DOI but other metadata (authors/title) doesnât match the DOI target. - ânear_miss_titleâ: title differs from a real paper by small but meaningful edits. - âplausible_fabricationâ: entirely fabricated yet plausible-sounding paper. - âmerged_citationâ: metadata combined from two real papers. - âpartial_author_listâ: real paper but author list is incomplete. - âarxiv_version_mismatchâ: arXiv version cited as a different version (or as published). Respond with JSON only: "label": "VALID" or "HALLUCINATED" or "UNCERTAIN", "confidence": 0.0 to 1.0, "predicted_hallucination_type": "<one of the 14 types above, or null>", "reason": "brief explanation" âpredicted_hallucination_typeâ MUST be null when label is VALID or UNCERTAIN. B.2.4 Reproducibility: LLM experimental setup Model roster. The evaluation pipeline accepts any OpenRouter-, OpenAI-, or Anthropic-compatible model ID as a drop-in kwarg (model=...). Tab.Ë11 lists every zero-shot baseline reported in the paper, the API endpoint used for the predictions in data/v1.0/baseline_results/ and results/temporal_supplement/, and the corresponding training cutoff (cross-referenced against Tab.Ë24). The two Anthropic models are reachable through both the native API (baselines llm_anthropic and llm_anthropic_opus_4_7) and the OpenRouter mirror; we use the OpenRouter mirror for the released predictions. Agentic baselines (llm_agentic_openai, llm_agentic_btu_openai) use gpt-5.1; the reported Anthropic agentic-bibtex-updater run (llm_agentic_btu_sonnet_4_6) uses the OpenRouter mirror anthropic/claude-sonnet-4.6, matching the zero-shot Sonnet endpoint. The registry also retains legacy native-API agentic variants (llm_agentic_anthropic, llm_agentic_btu_anthropic) hard-coded to claude-sonnet-4-5-20250929; these are not used for any number reported in the paper. OpenRouter is accessed via the OpenAI SDK with base_url=https://openrouter.ai/api/v1. All zero-shot, DOI-only, and bibtex-updater main-run predictions in data/v1.0/baseline_results/ were collected on 2026-05-05 via the endpoints above; hosted-LLM rows are dated snapshots subject to endpoint drift (§Ë7). Table 11: Zero-shot LLM baseline roster. Endpoint indicates the API used for the released predictions; native = vendor SDK, OR = OpenRouter mirror. Training cutoffs cross-reference Tab.Ë24; ââ€â marks inferred upper bounds. Run date is the collection date of the released predictions; the two Anthropic rows are the non-reproducing snapshot discussed in §Ë7. gpt-5.1, claude-sonnet-4.6, and claude-opus-4.7 are floating aliases (only gpt-5.4-2026-03-05 is a dated-pinned ID), so every row is reproducible only as a dated snapshot. Model Endpoint / model ID Provider Cutoff Run date GPT-5.1 native: gpt-5.1 OpenAI Sep 2024 2026-05-05 GPT-5.4 native: gpt-5.4-2026-03-05 OpenAI Aug 2025 2026-05-05 Claude Sonnet 4.6 native: claude-sonnet-4-6; OR: anthropic/claude-sonnet-4.6 Anthropic †Aug 2025 2026-05-05â Claude Opus 4.7 native: claude-opus-4-7; OR: anthropic/claude-opus-4.7 Anthropic †Oct 2025 2026-05-05â DeepSeek-R1 OR: deepseek/deepseek-r1 DeepSeek Jul 2024 2026-05-05 DeepSeek-V3.2 OR: deepseek/deepseek-v3.2 DeepSeek †Oct 2024 2026-05-05 Qwen3-235B-A22B-2507 OR: qwen/qwen3-235b-a22b-2507 Alibaba Jun 2025 2026-05-05 Qwen3-VL-235B-A22B-Instruct OR: qwen/qwen3-vl-235b-a22b-instruct Alibaba †Jun 2025 2026-05-05 Mistral Large 2512 OR: mistralai/mistral-large-2512 Mistral †mid-2025 2026-05-05 Gemini 2.5 Flash OR: google/gemini-2.5-flash Google Jan 2025 2026-05-05 Gemini 2.5 Pro OR: google/gemini-2.5-pro Google Jan 2025 2026-05-05 Llama 4 Maverick OR: meta-llama/llama-4-maverick Meta †Aug 2024 2026-05-05 â Non-reproducing snapshot: the OpenRouter Anthropic endpoint has since drifted (§Ë7, §ËB.1). Decoding configuration. Zero-shot calls request temperature=0.0 on every endpoint that accepts it. The one forced exception is gpt-5.5, which the code sets to temperature=1.0; gpt-5.1 and gpt-5.4 are sent temperature=0.0, but the GPT-5 endpoint samples stochastically regardless of the requested value (E3 in §ËF.4 measures the resulting run-to-run variance). The completion-token budget is max_completion_tokens=1024 on OpenAI- and OpenRouter-compatible endpointsâsized to absorb hidden thinking tokens emitted by reasoning modelsâand max_tokens=256 on the Anthropic native endpoint. This single 1024-(non-Anthropic)/256-(Anthropic-native) split applies to every main-table, temporal-supplement, and cutoff-aware run; the thinking-budget smoke test (§ËF.5) deliberately sweeps around this budget to locate the saturation boundary. We pass seed=42 on OpenAI-SDK-compatible endpoints; the Anthropic Messages API does not expose an equivalent. OpenRouter is left on its default dynamic provider routing (no provider pin), so an open-weight slug may be served by different underlying providers (e.g. DeepInfra, Together, Fireworks) on different runs; the OpenRouter rows are therefore reproducible only as dated snapshots. The zero-shot prompt (§ËB.2.3) is sent as a single user-role message with no system prompt. We do not use JSON mode; responses are parsed with a tolerant extractor that handles bare JSON and fenced code blocks, with an UNCERTAIN@0.5 fallback on parse failure (throughout, label@c denotes the label assigned with confidence c). Agentic calls use a 1024-token completion budget to accommodate tool-call turns. Client-side retry is max_retries=5 for zero-shot and 3 for agentic baselines, with a 120 s per-request timeout. Agentic baselines: tools and iteration. We evaluate four agentic variants: OpenAI (gpt-5.1) and Anthropic (OpenRouter anthropic/claude-sonnet-4.6) backends, each with either a multi-source tool suite or bibtex-updater-as-tool. The OpenAI agentic and the Anthropic agentic-bibtex-updater runs are the variants reported in the paper; numbers come from data/v1.0/baseline_results/llm_agentic_btu_sonnet_4_6_dev_public.json and the corresponding OpenAI checkpoint. Tools are exposed via the providerâs native function-calling interface with tool_choice="auto": the model may answer without any tool call, and we track the parametric-vs-tool split per prediction. Up to MAX_TOOL_CALLS=5 invocations per entry, plus one final verdict turn. If the cap is exceeded without a verdict, the entry is assigned UNCERTAIN@0.5. The multi-source suite comprises 4 tools, each returning authors, title, venue, year, doi with fields truncated to 500 characters: âą resolve_doi(doi): CrossRef DOI resolution (api.crossref.org/works/doi). âą search_crossref(query, limit=5): CrossRef bibliographic search (max 20 results). âą search_openalex(query, limit=5): OpenAlex polite pool (max 25). âą search_arxiv(query, limit=5): arXiv Atom API (max 20). All HTTP requests use a HALLMARK/1.0 User-Agent and 10â15 s timeouts; tool exceptions are retried with exponential backoff (up to 3 retries, base delay 1 s). The bibtex-updater-as-tool variant exposes a single tool, verify_with_bibtex_updater(bibtex), that invokes the bibtex-check CLI with --rate-limit 120 --academic-only, returning status, confidence, mismatched_fields, api_sources, errors. Subprocess timeout: 180 s. External-database access. The three database-backed endpointsâDOI-only, standalone bibtex-updater, and the agentic suiteâquery five external services: CrossRef (api.crossref.org, polite pool, no key), OpenAlex (polite pool, identified by a mailto contact), Semantic Scholar (the graph/v1 paper-lookup API, accessed with an S2_API_KEY), the arXiv Atom API, and DBLP. Semantic Scholar resolves a citation by matching its title, authors, venue, and identifiers against the S2 corpus and returning the canonical recordâor noneâwhich is the evidence bibtex-updater weighs to confirm or flag an entry. Semantic Scholar requires an API key for reliable programmatic access: the free key moves requests off the throttled shared public pool onto a dedicated per-key rate limit. The key affects only throughput and reliability: an unauthenticated re-run hits the shared limit (frequent HTTP 429) and resolves the corpus more slowly, but queries the same records and returns the same matches, so it does not change the reported verdicts. The main-run numbers were queried on 2026-05-04/05. Agentic tool results are replayable offline from the released sha256-keyed tool-cache (mechanics under Tool-call caching below), so the agentic baselines reproduce without live calls. The DOI-only and standalone-bibtex-updater endpoints have no such cache, so their numbers are dated snapshots of the corresponding databases. System prompts. The multi-source agentic system prompt: ⏠You are a citation verification expert with access to bibliographic lookup tools. Your task: determine whether a BibTeX entry is a VALID real publication or a HALLUCINATED (fabricated) citation. Strategy: 1. Inspect the entry for obvious red flags (fake DOI prefix, future year, placeholder authors). 2. Use tools to cross-reference: resolve the DOI, or search by title/author. 3. After gathering evidence (or after finding sufficient signal), emit your verdict. When you are ready to give your final answer, output ONLY valid JSON -- no prose, no markdown fences: "label": "VALID" or "HALLUCINATED", "confidence": 0.0 to 1.0, "reason": "concise explanation citing the evidence you found" Do NOT output the JSON until you have used enough tools or determined that parametric knowledge is sufficient. Listing 1: System prompt (multi-source) The bibtex-updater-as-tool agentic system prompt: ⏠You are a citation verification expert with access to a specialized tool, âverify_with_bibtex_updaterâ, which cross-references a BibTeX entry against CrossRef, DBLP, and Semantic Scholar and returns a structured verdict. Strategy: 1. For almost every entry, call âverify_with_bibtex_updaterâ once with the exact BibTeX string you were given. 2. Interpret the returned âstatusâ field: statuses like âverifiedâ, âurl_verifiedâ, or âpublished_version_existsâ suggest VALID; statuses like ânot_foundâ, âtitle_mismatchâ, âauthor_mismatchâ, âhallucinatedâ, âfuture_dateâ, âdoi_not_foundâ, or âvenue_mismatchâ suggest HALLUCINATED. 3. If the tool returns âapi_errorâ or you suspect the tool is wrong (e.g. it reports âverifiedâ but the entry still looks suspicious on inspection), apply your own judgment -- the tool is evidence, not an oracle. 4. If the first call is unambiguous, do NOT call again. Extra calls waste budget. When you are ready to give your final answer, output ONLY valid JSON -- no prose, no markdown fences: "label": "VALID" or "HALLUCINATED", "confidence": 0.0 to 1.0, "reason": "concise explanation referencing the tool status and any disagreement" Listing 2: System prompt (bibtex-updater-as-tool) Tool-call caching. Agentic tool results are cached in SQLite keyed on sha256(tool_name::canonical_args_json). TTL is disabled by default; cache hits bypass the network. This is a reproducibility artifact, transparent to the LLM. Failure handling. After three consecutive API errors, the baseline aborts and marks remaining entries as UNCERTAIN@0.5 with reason tag [Error fallback]. These entries retain Cov. = 1.00 but contribute zero to detection rate. Appendix C Core results and validity C.1 Full test_public results Tab.Ë12 reports all tools with both dev_public and test_public evaluations: the DOI-only baseline, twelve zero-shot LLM verifiers, three agentic harnesses, bibtex-updater, and the always-call evidence-injection co-designed variant. Table 12: Full test_public (831 entries) results vs. dev_public (1,119 entries). Zero-shot LLMs ordered by |ÎâFPR|| |. Precision-end LLMs (Sonnet 4.6, Opus 4.7, GPT-5.4, Gemini 2.5 Pro/Flash) move within ± 2.5 p; recall-aggressive open-weight models drift +2.6+2.6 to +8.2+8.2 p. DeepSeek-R1 is the lone zero-shot outlier moving in the opposite direction (â30.3-30.3 p): the chain-of-thought verifier abstains heavily on test_public (180/831 UNCERTAIN, 21.7%, vs. 18/1119 on dev). Both bibtex-updater-agentic harnesses move in the opposite direction (Sonnet+bibtex-updater â8.8-8.8 p, GPT-5.1+bibtex-updater â11.4-11.4 p), partly reversing the dev-side harness FPR penalty. bibtex-updater is cross-split stable (+2.4+2.4 p), as its abstention policy refuses to flag entries it cannot back with a record. In the Î column, green marks cross-split-stable tools (|ÎâFPR|â€1| |†1 p) and red upward drift (â„+6.9â„+6.9 p); DeepSeek-R1âs â30.3-30.3 p is an abstention artifact and is left unshaded. dev_public test_public Tool DR FPR F1 DR FPR F1 Î Citation-database baseline DOI-only .268 .185 .373 .387 .279 .498 +0.094+0.094 Zero-shot LLMs Claude Sonnet 4.6 .781 .127 .827 .821 .125 .866 â0.002 +-0.002 GPT-5.4 .767 .228 .783 .780 .224 .815 â0.004 +-0.004 Claude Opus 4.7 .752 .072 .830 .763 .067 .846 â0.005 +-0.005 Gemini 2.5 Flash .500 .100 .631 .505 .106 .644 +0.006+0.006 Gemini 2.5 Pro .476 .050 .627 .458 .059 .613 +0.009+0.009 Llama 4 Maverick .614 .146 .707 .631 .167 .729 +0.020+0.020 DeepSeek-V3.2 .911 .702 .727 .911 .728 .776 +0.026+0.026 Mistral Large .716 .250 .742 .688 .282 .741 +0.032+0.032 GPT-5.1 .837 .411 .766 .852 .481 .796 +0.069+0.069 Qwen3-VL-235B .860 .551 .740 .927 .628 .804 +0.077+0.077 Qwen3-235B .860 .533 .744 .909 .615 .798 +0.082+0.082 DeepSeek-R1 .896 .623 .739 .809 .319 .812 â0.303 +-0.303 Agentic harnesses (tool-use; up to 5 tool calls per entry) GPT-5.1 + CrossRef/OpenAlex/arXiv .967 .478 .816 .942 .558 .827 +0.080+0.080 GPT-5.1 + bibtex-updater (agentic; tool optional) .980 .470 .824 .960 .356 .883 â0.114 +-0.114 Sonnet 4.6 + bibtex-updater (agentic; tool optional) .990 .431 .841 .990 .343 .902 â0.088 +-0.088 Co-designed (reference upper bound; §ËG.2) bibtex-updater (v1.2.0) .865 .092 .890 .877 .115 .901 +0.024+0.024 GPT-5.1 + bibtex-updater (always-call; output in prompt) .843 .144 .856 .855 .256 .851 +0.112+0.112 Takeaway. The cross-split pattern stratifies cleanly: precision-end LLMs (Opus 4.7 â0.5-0.5 p, GPT-5.4 â0.4-0.4 p, Gemini Flash +0.6+0.6 p, Gemini 2.5 Pro +0.9+0.9 p, Sonnet 4.6 â0.2-0.2 p) barely move; recall-aggressive open-weight models drift more (+2.6+2.6 to +8.2+8.2 p). bibtex-updater joins the stable cluster (+2.4+2.4 p), and its F1 lead over Sonnet 4.6 narrows from 6.36.3 p on dev_public to 3.53.5 p on test_public but does not reverse: the abstention policy removes the cross-split FPR blow-up the recall-only configuration showed. C.2 LLM baseline agreement and calibration The aggregate and per-type numbers for every LLM baseline are in Tabs.Ë3 and 15. Here we add three diagnostics that the released per-entry dev_public predictions support for the open-weight pair Qwen3-235B and DeepSeek-V3.2, which characterize how these verifiers fail rather than how often. Generation-method stratification. Tab.Ë13 stratifies detection rate by the entryâs generation method, probing whether an LLM verifier has an edge on LLM-generated entries. Table 13: Detection rate stratified by generation method on dev_public. LLM-generated entries are the hardest source for both models, the opposite of what a generator-familiarity advantage would predict. Generation method n Qwen3-235B DeepSeek-V3.2 Adversarial 60 0.917 1.000 Perturbation 410 0.866 0.937 Real-world 46 0.978 0.935 LLM-generated 90 0.730 0.722 Scraped (FPR) 513 0.533 0.702 For hallucinated entries the metric is detection rate; for scraped (valid) entries it is false positive rate over the full valid pool. Values are regenerated on the relabeled dev_public split; the relabel reclassified 23 recovered real papers out of the LLM-generated hallucinated bucket (n 113â 90). GPT-5.1âs per-source detection rate is omitted because only evaluation-level (not entry-level) GPT-5.1 predictions are available on dev_public. LLM-generated hallucinations remain the hardest source for both models (DR 0.722â0.730) relative to perturbation-based entries (DR 0.866â0.937), consistent with the GPT-5.1 finding in §ËB.1. Were verifiers systematically attuned to LLM-generated text, those entries would be easier to catch; the reverse holds here. The generator is GPT-5.1 (§ËA.3), so for these two models the stratification probes cross-model familiarity rather than strict self-recognition; the same-generator case is the GPT-5.1 stratification in Tab.Ë16. Adversarial entries are near-perfectly detected (0.917â1.000), so current adversarial strategies do not fool LLM verifiers. Pairwise agreement. Tab.Ë14 reports Cohenâs Îș for all 28 pairs of the eight zero-shot baselines with stored per-entry dev_public predictions (scripts/compute_pairwise_kappa.py); the remaining four baselines persisted only aggregate metrics. Table 14: Pairwise agreement between LLM baselines on dev_public (Cohenâs Îș; n=1,119n=1,119; abstentions scored as committed-valid, matching §Ë5.2). The two Anthropic columns use the later OpenRouter snapshot and carry its drift caveat (§ËE.2). Agreement clusters by disposition: the precision pair (OpusâSonnet, Îș=0.75Îș=0.75) and the recall-aggressive cohort (R1âV3.2 0.500.50, R1âQwen3 0.530.53) agree internally, while cross-disposition pairs fall as low as Îș=0.19Îș=0.19 (V3.2âFlash). Cells are shaded in proportion to Îș (darker == higher agreement). Opus 4.7 Sonnet 4.6 R1 V3.2 Flash Mistral Qwen3 Sonnet 4.6 .746 DeepSeek-R1 .317 .309 DeepSeek-V3.2 .243 .229 .497 Gemini 2.5 Flash .409 .448 .220 .187 Mistral Large .480 .513 .240 .298 .541 Qwen3-235B .370 .384 .529 .454 .311 .295 GPT-5.4 .576 .630 .379 .314 .512 .605 .500 The matrix orders itself by disposition: within-cohort Îș runs two to three times the cross-cohort values, so verifier errors are correlated within cohort and ensembling verifiers of the same disposition adds little independent evidence (cf. the voter ensemble in §ËH.2). The moderate Cohenâs Îș (0.454) between DeepSeek-V3.2 and Qwen3-235B indicates that despite similar overall detection rates, the two models disagree on a substantial fraction of entries. Of the 1,119 entries where both models have predictions, they agree and are both correct on 620, agree but are both wrong on 274 (predominantly false positives on valid entries), and disagree on 225. Thirty hallucinated entries are missed by both models, a âhard coreâ that may require a different verification strategy such as API-based cross-referencing. Confidence calibration. A verifierâs confidence is useful only if it separates correct verdicts from wrong ones. For Qwen3-235B and DeepSeek-V3.2 it does not: both assign a median confidence of 0.950.95 whether the verdict is right or wrong, and their mean confidence on correct predictions exceeds that on incorrect ones by only 0.0080.008 and 0.0010.001 respectively, so no confidence threshold can route their unreliable flags to human review. GPT-5.1âs confidences retain some discriminative signal (ECE 0.1900.190, Tab.Ë3), though far from calibrated. For deployment pipelines that triage flags by confidence, the two open-weight modelsâ scores are effectively uninformative. C.3 Full per-type results Fig.Ë4 visualizes per-type detection rates as a heatmap; Tab.Ë15 reports the underlying detection rate for every hallucination type and baseline. HaRC and verify-citations are excluded due to <7%<7\% effective coverage (§ËB.2.2). Figure 4: Per-type detection rates across all full-coverage tools. HaRC and verify-citations are excluded due to <7%<7\% effective coverage (§ËB.2.2). author_mismatch (shown in the heatmap under its enum value âswapped authorsâ) and near_miss_title remain the hardest types across the cohort. Table 15: Per-type detection rate on dev_public for full-coverage independent tools. HaRC and verify-citations are excluded due to <7%<7\% effective coverage (§ËB.2.2). Most main types have âŒ30 30 instances; correcting the ground-truth mislabels leaves one type (hybrid_fabrication, n=26n=26) below 30 (Tab.Ë2 caption). Stress-test types are evaluated in the separate stress_test split. The DOI-only, Sonnet 4.6, and Opus 4.7 per-type cells are reproduced from an earlier evaluation run, as their per-entry predictions were not retained on dev_public; the ten remaining columns are scored on the released labels. Column headers abbreviated for space: S4.6 = Sonnet 4.6, O4.7 = Opus 4.7, G51 = GPT-5.1, G54 = GPT-5.4, R1 = DeepSeek-R1, V3 = DeepSeek-V3.2, Q3 = Qwen3-235B, ML = Mistral Large, GF = Gemini 2.5 Flash, L4 = Llama 4 Mav., GP = Gemini 2.5 Pro, QV = Qwen3-VL. The two red-shaded types are the cohort-wide hard types, the only per-type gaps that exceed the 20â26 p MDE (§ËB.1). Numeric cells are shaded in proportion to the detection rate (darker == higher). Tier Type DOI G51 G54 S4.6 O4.7 R1 V3 Q3 ML GF L4 GP QV 1 fabricated_doi 0.590 1.000 0.974 0.821 0.436 1.000 1.000 1.000 0.974 0.943 1.000 0.703 1.000 nonexistent_venue 0.125 1.000 1.000 1.000 0.975 1.000 1.000 1.000 0.974 0.513 0.949 0.571 1.000 placeholder_authors 0.235 1.000 0.927 0.956 1.000 1.000 1.000 0.902 0.927 0.463 0.610 0.805 0.927 future_date 1.000 1.000 1.000 1.000 0.100 1.000 1.000 1.000 1.000 1.000 1.000 0.857 1.000 2 chimeric_title 0.143 0.894 0.872 0.907 0.977 0.867 0.979 0.830 0.809 0.500 0.596 0.614 0.915 wrong_venue 0.200 0.851 0.766 0.761 0.891 0.957 0.915 0.915 0.689 0.587 0.617 0.422 0.851 author_mismatch 0.206 0.448 0.433 0.612 0.701 0.657 0.716 0.522 0.348 0.106 0.194 0.106 0.522 preprint_as_pub. 0.129 0.742 0.581 0.742 0.806 0.900 0.839 0.903 0.452 0.419 0.548 0.100 0.806 hybrid_fabrication 0.261 1.000 0.962 0.800 0.822 1.000 1.000 0.885 0.923 0.692 0.692 0.846 0.923 3 near_miss_title 0.086 0.577 0.462 0.366 0.463 0.731 0.788 0.673 0.558 0.220 0.269 0.196 0.712 plausible_fabrication 0.036 0.934 0.895 0.845 0.810 0.974 0.961 0.987 0.842 0.640 0.842 0.680 0.974 DOI-only with pre-screening (§ËB.2.1) achieves non-zero detection across all types; without pre-screening it detects only fabricated_doi entries. GPT-5.1âs two weakest typesâauthor_mismatch (0.448) and near_miss_title (0.577)ârequire exact bibliographic recall that even large LLMs lack (parenthesized values are per-type detection rates). Takeaway. Read the per-type heatmap as a failure-mode diagnostic, not a tool ranking: at nâ30nâ30 per cell, MDE is 20â26 p, so only the cohort-wide hard types (author_mismatch, near_miss_title) carry enough power for tool-vs-tool comparison. C.4 Validity checks C.4.1 Core subset analysis Scaling the hallucinated pool does not distort the evaluation signals. The pool grew out of 114 seed hallucinations (71 in dev, 43 in test) through generation across all main hallucination types, and to check the scaling we re-evaluated all baselines on the core subset containing only the 114 seed hallucinations alongside the valid entries. This check was performed on an earlier corpus snapshot: 816 hallucinated (453 dev / 363 test) and 720 valid public entries at the time; it predates the final corpus expansion and the ground-truth relabel, and the released corpus counts are in Tab.Ë2. On that snapshot, aggregate metrics differ by less than 2% between the core subset and the full pool across all baselines: detection rate changes by at most 0.015, F1 by at most 0.02, and tier-weighted F1 by at most 0.018. Per-type rankings are preserved: no baseline changes rank on any type. We conclude that the scaled entries are consistent with the seed distribution and do not introduce systematic bias. C.4.2 Shortcut analysis Metadata features alone barely predict hallucination labels, so the dataset offers no shortcut a model could exploit instead of reading the content. To test this, we train a logistic regression classifier on eight entry-level featuresâpresence of DOI, field count, author count, title length (characters), title word count, year (numeric), BibTeX entry type, and presence of a booktitle fieldâand evaluate via 5-fold cross-validation on the dev_public split. The logistic regression achieves a cross-validated accuracy of 58.8% (majority-class baseline: 54.2%), a margin of 4.6p. Since this margin is below 5p, metadata features provide negligible predictive signal beyond class prevalence. This shows that the metadata features themselves do not leak class information; which content fields actually drive detection is quantified separately by the field-leave-one-out ablation (§ËH.3), where dropping the title raises FPR by up to 35.535.5 p. Among the metadata features, has_doi carries the most weight, consistent with the DOI-only baselineâs non-zero detection rate, but no single feature pushes cross-validated accuracy meaningfully above the class-prevalence baseline, confirming that the perturbation pipeline does not introduce systematic metadata artifacts. Detection rate by generation method. To assess whether GPT-5.1 exhibits self-recognition bias on LLM-generated entries, we stratify detection rate by the entryâs generation method (Tab.Ë16). LLM-generated hallucinations are harder for GPT-5.1 to detect (DR = 0.656) than perturbation-based entries (DR = 0.846), the opposite of what a self-recognition advantage would predict, though this stratification cannot separate intrinsic difficulty from a recognition effect it might mask. This is expected: LLM-generated entries are designed for ecological validity and lack the systematic structural patterns that perturbation introduces. Perturbation-generated entries dominate the hallucinated pool (410410 vs. 9090 LLM-generated on dev_public, Tab.Ë16) and are detected âŒ19 19 p more easily, so the public leaderboard is weighted toward this easier class. The harder, content-driven probes of construct validity are the LLM-generated subset and the field-leave-one-out ablation (§ËH.3); read absolute detection rates with this class mix in mind. Adversarial entries, despite being crafted to evade detection, are detected without error (DR = 1.000), suggesting that current adversarial strategies do not fool LLM verifiers. Table 16: GPT-5.1 detection rate on dev_public stratified by generation method. DR is computed over hallucinated entries, FPR over valid entries; a method with no entries of a given polarity shows âââ. The ground-truth audit places 4 perturbation-origin and 23 LLM-generated entries in the valid pool, so those two methods carry a (small-n) per-method FPR alongside the scraped pool (n=486n=486, FPR 0.405); the four per-method FPRs reconcile to the headline dev_public FPR of 0.411 over all 513 valid entries. The âAdversarialâ generation method in the data files covers hybrid_fabrication and plausible_fabrication entries that use template-based perturbation (real DOI + fabricated metadata, or combinatorial assembly of plausible fields), distinguishing them from âPerturbationâ (simple single-field modifications) and âLLM-generatedâ (entries produced by language models). Generation method n (hall.) n (valid) DR FPR Adversarial 60 â 1.000 â Perturbation 410 4 0.846 1.000 Real-world 46 â 0.891 â LLM-generated 90 23 0.656 0.435 Scraped (valid) â 486 â 0.405 Cross-split validation. We evaluate GPT-5.1 on test_public (831 entries, 62.5% hallucinated) to verify that findings generalize beyond dev_public. Detection rate is consistent: DR = 0.852 (test) vs. 0.837 (dev). Per-tier patterns are preserved: Tier 1 DR = 1.000, Tier 2 DR = 0.754, Tier 3 DR = 0.831 on dev_public. FPR is elevated on both splits (dev: 0.411, test: 0.481), reflecting GPT-5.1âs tendency to flag valid entries with non-canonical metadata (uncommon venues, name variants). bibtex-updaterâs FPR is by contrast cross-split stable (dev 0.092 â test 0.115), because its abstention policy declines to flag entries it cannot back with a record rather than guessing on test_publicâs harder valid pool. MCC = 0.397 (test) provides a prevalence-invariant comparison point. Takeaway. The validity checks hold: scaling the hallucinated pool from its 114 seeds shifts aggregate metrics by at most 2 p with no rank changes (§ËC.4.1), and a metadata-only classifier exceeds the majority baseline by only 4.6 pâbelow the 5 p leakage barâso metadata offers no shortcut around reading the content (§ËC.4.2). The class mix is the caveat to carry: perturbation entries dominate the pool and are âŒ19 19 p easier for GPT-5.1 than the LLM-generated subset (DR 0.8460.846 vs. 0.6560.656), a direction that also runs opposite to a self-recognition advantage. C.5 External validity: authentic ChatGPT citations (the WaltersâWilder supplement) The stratification in Tab.Ë16 rests on Hallmarkâs own LLM-generated entries, so it cannot rule out that our generation pipeline leaves fingerprints a verifier could exploit; an external check needs authentic hallucinations we did not construct. We therefore convert the hand-coded citation corpus of Walters and Wilder [2023] into a fourth evaluation-only extension split, supplement_chatgpt_citations (Tab.Ë2). Walters and Wilder prompted ChatGPT-3.5 and GPT-4 to write 84 short literature reviews across 42 topics spanning the humanities, social sciences, and natural sciences, then hand-verified every one of the 636 resulting citations, reporting fabrication rates of 55% (GPT-3.5) and 18% (GPT-4). The corpus extends the benchmark along the two axes it is thinnest on: the hallucinations are authentic LLM output with expert labels rather than perturbations (the dev_public pool behind Tab.Ë16 has 90 LLM-generated hallucinated entries), and the subject matter is multidisciplinary rather than ML, making it the natural companion to the cross-domain split of §Ë7. Conversion. We keep the supported publication typeâjournal articlesâbecause books, chapters, and websites (â32%â32\% of the source corpus) are structurally unresolvable for the database-backed verifiers evaluated here. A field mismatch on a real work is a hallucination and is mapped to the matching taxonomy type rather than left valid (Tab.Ë17); when several fields are wrong, the most identity-defining field wins (title >> author >> venue). Thirty-four real, findable articles whose only substantive error is a wrong non-future year and/or a volume/page slip are excluded because the taxonomy has no matching type; formatting-only deviations (capitalization, initials-vs-names, missing pages) are likewise not treated as hallucinations. The result is 341 entries (172 valid / 169 hallucinated) that pass the same validation pipeline as the core corpus; the converter, its per-case audit trail, and the raw tool outputs are released with the benchmark code. Table 17: Mapping the WaltersâWilder source coding onto the Hallmark taxonomy. A substantive field error on a real work maps to the corresponding hallucination type; only wrong-year/volume/page-only errors (34 entries) have no matching type and are excluded. Source coding Hallmark type Tier n work itself fabricated plausible_fabrication 3 139 real work, wrong title near_miss_title 3 13 real work, wrong authorship swapped_authors 2 12 real work, wrong/invented journal wrong_venue 2 5 real work, no substantive error valid â 172 bibtex-updater results. On the full supplement (coverage 341/341), bibtex-updater reaches DR 0.929 (95% CI [0.905,0.953][0.905,0.953]), FPR 0.076 ([0.041,0.116][0.041,0.116]), F1 0.926, and tier-weighted F1 0.959, close to its dev_public profile (Tab.Ë32) on data from disciplines it was never tuned on. The per-type gradient replicates the main-split pattern: plausible_fabrication 1.000 (n=139n=139) and wrong_venue 1.000 (n=5n=5), against near_miss_title 0.769 (n=13n=13) and swapped_authors 0.250 (n=12n=12). The tool is in effect a fabrication detector: an invented work is simply not found in any backing database, whereas a corrupted citation of a real work resolves by title and the tool then accepts the record without re-checking authorship, the same abstention-driven weakness as on the core splits (§ËE.2). Stratified by generator, GPT-3.5 output is easier (DR 0.990, FPR 0.053) than GPT-4 output (DR 0.841, FPR 0.078): GPT-4 fabricates less and its errors are subtler, mostly author swaps on real papers. Across subject fields, detection is stable (0.85â0.96) while FPR is highest in the natural sciences (0.123 vs. 0.000 humanities, 0.056 social sciences). Converter audit. Because a conversion artifact is indistinguishable from tool behavior in the aggregate numbers, we audited every false positive and every miss case by case. The audit caught one APA-parser bugâdropped final authors on âA, B, & Câ references and deleted ellipsis truncation, which bibtex-updater reads as silent author-list truncationâthat had inflated DR and FPR simultaneously: spurious flags on valid entries and wrong-reason detections on swapped_authors. All numbers above use the fixed converter; the released per-case trail documents both scoring passes. The residual errors are genuine tool behavior: the 13 remaining false positives are metadata sensitivity on real works (strict subtitle/punctuation matching, online-first vs. print years), and all 12 misses are author or title corruptions the tool resolves to the real record and abstains on. An aggregate score on a converted external corpus is thus only as trustworthy as a per-case audit of its false positives and misses. Zero-shot LLM verifiers. We run four zero-shot LLM verifiers spanning the conservativeâaggressive spectrum of Tab.Ë25 on the supplement, under the default prompt and the main-table protocol (§ËB.2.4); every cited work predates all four training cutoffs, so the supplement isolates parametric verificationâthe model verifying a citation against what its weights absorbed in pretraining, without retrievalâon non-ML literature with no post-cutoff confound. Tab.Ë18 reports the results next to each modelâs dev_public baseline. Three patterns emerge. First, detection of authentic fabrications separates the models most: GPT-5.1 (0.964) and Sonnet 4.6 (0.961) approach bibtex-updaterâs 1.000 on plausible_fabrication, while Gemini 2.5 Flash recognizes only 0.345 of them, its overall DR falling from 0.500 to 0.290 even as its FPR stays low (0.029). The drop is consistent with thinner parametric coverage of the multidisciplinary literature, but with four models and no direct probe of what each has memorized, we cannot separate missing coverage from a conservative disposition that declines to flag what it cannot confirm. Second, the FPR ordering from dev_public is preserved and widens: the spread grows from 0.100â0.533 on ML venues to 0.029â0.860 here, with the aggressive Qwen3-235B flagging 86% of valid multidisciplinary papers; its seemingly strong swapped_authors cell (0.917) is a byproduct of that indiscriminate flagging. Third, the subtle-corruption weakness spans the standalone verifier class: every calibrated standalone verifier scores at or below 0.25 on swapped_authors (bibtex-updater 0.250, GPT-5.1 0.250, Sonnet 4.6 0.182, Gemini Flash 0.000), so the benchmarkâs Tier-2 finding replicates on authentic data for parametric and database-backed verification alike; the agentic combination below is the one configuration that improves on it. Stratified by generator, GPT-4 output is harder than GPT-3.5 output for every verifier (e.g. Sonnet 4.6 DR 0.94 vs. 0.58; bibtex-updater 0.99 vs. 0.84): the better generator fabricates less and corrupts more subtly. Sonnet 4.6 is the only zero-shot model to abstain (15 entries) and posts the best LLM F1 (0.896) with an FPR (0.058) below its ML-venue baseline, consistent with its calibrated profile in the temporal analysis (§ËF.1). Table 18: Zero-shot LLM verifiers vs. bibtex-updater on the WaltersâWilder supplement (N=341N=341: 172 valid / 169 hallucinated). DR/FPR on dev_public (2021â2023 ML venues) alongside the supplement (authentic, multidisciplinary, pre-cutoff for all models). Abstentions are excluded from DR/FPR (§ËB.1); Cov. is the committed fraction (Sonnet 4.6 abstains on 15 entries, the agentic harness on 22; all other verifiers commit on all 341). Fab. is the detection rate on plausible_fabrication (n=139n=139), Swap. on swapped_authors (n=12n=12). The bottom block is tool-involving: the standalone co-designed tool and the agentic btu-only harness of Tab.Ë3, whose dev columns are that tableâs agentic (tool-optional) row. Green marks supplement FPRs at or below the dev_public value for calibrated verifiers; red the aggressive outlier. âIndiscriminate flagging (FPR 0.860), not discrimination. DR FPR Supplement Model dev suppl. dev suppl. F1 Fab. Swap. Cov. Gemini 2.5 Flash 0.500 0.290 0.100 0.029 0.439 0.345 0.000 1.00 Claude Sonnet 4.6 0.781 0.865 0.127 0.058 0.896 0.961 0.182 0.96 GPT-5.1 0.837 0.846 0.411 0.285 0.792 0.964 0.250 1.00 Qwen3-235B 0.860 0.976 0.533 0.860 0.685 1.000 0.917â 1.00 bibtex-updater 0.865 0.929 0.092 0.076 0.926 1.000 0.250 1.00 GPT-5.1 + btu (agentic) 0.980 0.975 0.470 0.045 0.967 1.000 0.571 0.94 GPT-5.1 as a bibtex-updater dispatcher. The agentic btu-only harness of Tab.Ë3âGPT-5.1 with bibtex-updater as its only tool, free to decide when to call it and whether to trust itâdominates both of its components on the supplement: DR 0.975 (95% CI [0.952,0.994][0.952,0.994]), FPR 0.045 ([0.013,0.079][0.013,0.079]), F1 0.967, tier-weighted F1 0.984, committing on 319 of 341 entries. The combination inherits the toolâs perfect fabrication detection (1.000) and improves on both weaknesses the components share: near_miss_title rises to 1.000 (tool alone 0.769, model alone 0.385) and swapped_authors to 0.571 (4 of 7 committed; both components alone 0.250): the dispatcher treats the toolâs mismatch statuses as evidence where the standalone wrapper abstains, and routes genuinely ambiguous cases to UNCERTAIN. Notably, the harness FPR inflation of failure mode (i) does not reproduce here: the same configuration that posts FPR 0.470 on dev_public posts 0.045 on the supplement. The valid half of this corpus consists of real, database-indexed journal articles that the tool verifies cleanly, so the dispatcher rarely faces the ambiguous partial-match output that produces dev-side over-flagging, consistent with §Ë5.2âs reading that mode (i) characterizes harness behavior on borderline metadata, not the dispatch pattern itself. Protocol note: this runâs tool calls used authenticated OpenAlex and Semantic Scholar access, which the standalone bibtex-updater run above predates; authentication changes service level (rate limits, latency), while the records consulted are identical. Takeaway. On authentic, multidisciplinary ChatGPT citations, two of the benchmarkâs central findings replicate on data we did not construct: subtle corruptions of real works remain the hard class for every calibrated standalone verifier (swapped_authors DR â€0.25â€0.25, parametric and database-backed alike), and FPRâthe deployment-decisive metricâspans 0.029â0.860 across the LLM cohort, preserving and widening the dev_public ordering. Detection of authentic fabrications spans 0.345â0.964 across LLMs (vs. 1.000 for the database-backed tool), a spread consistent with differences in parametric coverage. The compositional exception is the agentic dispatcher: GPT-5.1 orchestrating bibtex-updater reaches DR 0.975 / FPR 0.045 / F1 0.967 and lifts swapped_authors to 0.571, without the mode-(i) FPR inflation it shows on ML venues. Ecological validity cuts both ways: synthetic perturbations neither overstate the difficulty of authentic fabrications nor manufacture the subtle-corruption weakness. Caveat: n=12n=12 for swapped_authors, and the supplement is articles-only. C.6 Comparison and additional figures C.6.1 Comparison with related citation efforts Tab.Ë19 contrasts Hallmark with concurrent citation-audit efforts. Prior work focuses on auditing published papers; Hallmark is, to our knowledge, among the first evaluation benchmarks for citation-verification tools, alongside concurrent work such as CiteAudit [Shi et al., 2026]. Table 19: Comparison of Hallmark with related citation analysis efforts. CiteAudit [Shi et al., 2026] is a concurrent, human-validated effort that evaluates citation-verification tools (LLMs and commercial tools); âââ marks attributes we could not verify from the available description. aHallmarkâs real-world and adversarial entries receive manual review; perturbation and LLM-generated entries are verified algorithmically (§Ë3.3). Entries Types Sub-tests Tool eval. Human-valid. Open Citation audits GhostCite [Xu et al., 2026] 56K papers â â â â â HalluCitation [Sakai et al., 2026a] âŒ300 300 papers 3 â â â â Mysterious Citations [Bienz et al., 2026] HPC venues â â â â â Citation verification benchmarks CiteAudit [Shi et al., 2026] â â â â â â Hallmark (ours) 2,526 14 6 â partiala â C.6.2 Additional figures Figure 5: Costâaccuracy tradeoff. All sixteen evaluated tools with recorded throughput are shown; rate-limited tools (HaRC, verify-citations) are plotted at their sub-7%-coverage operating points (§ËB.2.2). Rate-limited tools are impractical for venue-scale deployment. Per-entry cost is the dominant feasibility constraint at low prevalence (âŒ2% 2\%): even highly accurate verifiers misallocate reviewer effort because precision is bottlenecked by base rate, not capability. Figure 6: Detection rate by difficulty tier across all full-coverage tools. DOI-only concentrates in Tier 1; every LLM achieves broad coverage with graceful degradation as difficulty rises. Appendix D Failure mode (i): agentic aggregation Tab.Ë5 names failure mode (i)âthe false-positive inflation of agentic, retrieval-augmented verifiers (Tab.Ë3). Here we isolate its mechanism: the inflation is a property of the aggregation rule that turns multi-database lookups into a verdict, with the base model held fixed, and Tab.Ë20 quantifies the fix. Cross-source aggregation on the agentic databases. Failure mode (i) attributes the agentic harnessâs inflated FPR to its emergent any-no-match disposition; we isolate that aggregation rule directly. Holding the databases and the matcher fixed, we query three of the harnessâs four databasesâCrossRef, OpenAlex, and arXivâfor every dev_public entry, count a source as confirming when it returns a record whose title matches the entry (â„0.85â„ 0.85 similarity), and score three rules over identical evidence. Any-no-match (flag if any source misses) reaches FPR 0.7290.729 at DR 0.8350.835; consensus (flag only if all three miss) cuts FPR to 0.0490.049âroughly 15Ă15Ă lowerâat DR 0.2510.251; majority (flag if fewer than half confirm) falls between, at FPR 0.2010.201 / DR 0.4170.417. The FPR reduction is therefore a property of the aggregation rule, not the model or the sources. The recall cost is expected and instructive: title-presence consensus catches only non-existence hallucinations and misses the metadata-corruption types (near_miss_title, chimeric_title, wrong_venue), where a real paperâs title still matchesâso the principled rule is consensus-absence paired with positive-contradiction checks, exactly bibtex-updaterâs design (DR 0.8650.865 / FPR 0.0920.092). Two caveats bound the reading: Semantic Scholar, the harnessâs fourth backend, is excluded because its API key returned HTTP 403 on every endpoint during our evaluationâwhereupon the harness silently falls back to a rate-limited keyless poolâso we aggregate over the other three;555This backend fragility is itself an instance of failure mode (i). Semantic Scholarâs free API tier has been curtailed since 2024 and its public support repository was archived in January 2025 (https://github.com/allenai/s2-folks); separately, OpenAlex began requiring an API key in early 2026, so the harnessâs keyless OpenAlex path no longer functions either. A verifier that reads such a failed lookup as âno record foundâ converts backend unavailability into a false positive. and the deterministic title matcher is a proxy for the harnessâs per-source judgement, so the absolute any-no-match FPR need not equal the LLM harnessâs (and the proxy DR 0.8350.835 likewise understates the harnessâs 0.970.97â0.990.99); the any-vs-consensus ordering is what transfers, not the level. Table 20: Cross-source aggregation on the agentic harnessâs databases (dev_public; deterministic title-match proxy over CrossRef/OpenAlex/arXiv at â„0.85â„ 0.85 similarity). Holding databases and matcher fixed, the aggregation rule alone moves FPR by âŒ15Ă 15Ă. bibtex-updater pairs consensus-absence with positive-contradiction checks, recovering the recall that title-only consensus loses. Read the any-vs-consensus ordering, not the absolute level (the proxy differs from the LLM harnessâs own operating point; Tab.Ë3). Shading marks the lever: red is the any-no-match FPR and green the consensus FPR on identical evidence (âŒ15Ă 15Ă apart); the gray row is the co-designed reference, as in Tab.Ë3. Aggregation rule DR â FPR â Any-no-match (flag if any source misses) 0.835 0.729 Majority (flag if fewer than half confirm) 0.417 0.201 Consensus (flag only if all three miss) 0.251 0.049 bibtex-updater (consensus-absence ++ contradiction) 0.865 0.092 Takeaway. Failure mode (i) is an aggregation failure: on identical lookups with the base model held fixed, flagging on any single-source miss gives FPR 0.7290.729, whereas requiring consensus absence lowers it âŒ15Ă 15Ă to 0.0490.049, at a recall cost that pairing consensus with positive-contradiction checks (bibtex-updater) recovers (DR 0.8650.865 / FPR 0.0920.092). The figures are a deterministic title-match proxy over three of the harnessâs four databases, so the any-vs-consensus ordering is the finding; the absolute level is proxy-dependent. Appendix E Failure mode (i): base-rate precision and deployment E.1 Deployment PPV analysis At venue-realistic prevalence (11â5%5\%), positive predictive value drops sharply. Tab.Ë21 reports PPV =DRâ prevDRâ prev+FPRâ (1âprev)= DR·prevDR·prev+FPR·(1-prev), where prev is the entry-level hallucination prevalence, for the full-coverage tools and the co-designed bibtex-updater (whose DR/FPR use the headline committed-VALID convention; its abstention is reported separately in §ËE.2). Table 21: Positive predictive value at venue-realistic prevalence. The Anthropic rows use the pinned dated snapshot; under endpoint drift a re-runner should expect higher absolute FPR, hence lower PPV (§Ë7). In the PPV2% column, green marks roughly one true catch per six flags (PPV â„16%â„16\%) and red under one per twenty (PPV <5%<5\%); the gray row is the co-designed reference. Tool DR FPR PPV2% PPV5% Opus 4.7 0.752 0.072 17.6% 35.5% Gemini 2.5 Pro 0.476 0.050 16.3% 33.4% bibtex-updater (v1.2.0) 0.865 0.092 16.2% 33.1% Sonnet 4.6 0.781 0.127 11.2% 24.5% Gemini 2.5 Flash 0.500 0.100 9.3% 20.8% Mistral Large 0.716 0.250 5.5% 13.1% GPT-5.1 0.837 0.411 4.0% 9.7% Qwen3-235B 0.860 0.533 3.2% 7.8% DeepSeek-R1 0.896 0.623 2.9% 7.0% DOI-only 0.268 0.185 2.9% 7.1% DeepSeek-V3.2 0.911 0.702 2.6% 6.4% At a 2%2\% base rate, GPT-5.1 achieves only 4.0% PPV: roughly 24 of every 25 flagged entries are false alarms. Aggressive models (DeepSeek-R1/V3.2) fall below 3% despite â„89%â„89\% DR. Gemini 2.5 Flash more than doubles GPT-5.1âs PPV (9.3% vs. 4.0%) through its much lower FPR, despite a lower detection rate: the precision-bottleneck logic favors the conservative model at deployment-realistic prevalence. Takeaway. At a 2%2\% base rate, low-FPR tools deliver multiples of the PPV of high-recall ones: Opus 4.7 hits 17.6% and Sonnet 4.6 11.2%, while DeepSeek-R1/V3.2 fall below 3% despite â„89%â„89\% DR. Pick the verifier by FPR, not recall. Prevalence sweep. To separate the ranking of tools (which is prevalence-invariant because DR and FPR are) from deployability (absolute PPV, which is not), Tab.Ë22 sweeps the same closed-form PPVâ(prev)=DRâ prevDRâ prev+FPRâ (1âprev)PPV(prev)= DR·prevDR·prev+FPR·(1-prev) (3) over prevalence values 0.005,0.01,0.02,0.05,0.10,0.20,0.30\0.005,0.01,0.02,0.05,0.10,0.20,0.30\ for four representative tools spanning the FPR spectrum, using the exact DR/FPR from Tab.Ë3. Table 22: PPV vs. prevalence for four representative tools (Eq.Ë3; DR/FPR from Tab.Ë3). (a) The ranking is essentially prevalence-invariant: Opus 4.7 >> Sonnet 4.6 >> GPT-5.1 >> DeepSeek-V3.2 at every prevalence, because PPV is monotone in DR/FPR and the FPR gaps dominate. (b) Absolute PPV (deployability) changes sharply with prevalence: the same tool moves from near-useless at venue-realistic prevalence to usable as prevalence rises (Opus 4.7: 5.0% at a 0.5%0.5\% base rate to 81.7% at 30%30\%). Shading in the 2% column mirrors Tab.Ë21; at other prevalences the ranking is unchanged and only deployability moves. Tool DR FPR .5% 1% 2% 5% 10% 20% 30% Claude Opus 4.7 0.752 0.072 5.0% 9.5% 17.6% 35.5% 53.7% 72.3% 81.7% Claude Sonnet 4.6 0.781 0.127 3.0% 5.8% 11.2% 24.5% 40.6% 60.6% 72.5% GPT-5.1 0.837 0.411 1.0% 2.0% 4.0% 9.7% 18.5% 33.7% 46.6% DeepSeek-V3.2 0.911 0.702 0.6% 1.3% 2.6% 6.4% 12.6% 24.5% 35.7% The sweep makes the deployment lesson concrete: at venue-realistic 11â2%2\% even the best tool flags four-to-five false alarms per true catch, but the relative ordering a benchmark establishes transfers unchanged to any deployment prevalence: only the operating point on the PPV curve moves. E.2 Selective prediction and the Coverage column The Coverage column of Tab.Ë3 reports, for each tool, the fraction of entries on which it commits to a VALID/HALLUCINATED verdict rather than abstaining (UNCERTAIN). We report it because abstention is a usable deployment actionâdefer to a human reviewerâand because it is the lens under which the co-designed bibtex-updater reads cleanly: it abstains on the 1818â21%21\% of entries it cannot verify against a backing record (coverage 0.820.82 on dev, 0.790.79 on test). Its low headline FPR (0.0920.092) does not come from that abstention: on the entries it commits to, its selective FPR (0.0990.099) is essentially unchanged from full coverage, so abstention raises recall and leaves precision unchanged. Where the LLM verifiers lower their false-positive rate as they abstainâprecision-via-abstentionâbibtex-updaterâs riskâcoverage curve is flat: its precision is structuralâit comes from conservative matching, not from abstention (its selective FPR 0.0990.099 â its full-coverage 0.0920.092)âso selective prediction adds little. We use âstructuralâ throughout in this operational sense; it does not mean invariant, as the FPR still rises modestly across splits (0.092â0.1150.092â 0.115; Tab.Ë12). No peer citation benchmark reports coverage; the selective-prediction framing is standard elsewhere [El-Yaniv and Wiener, 2010, Geifman and El-Yaniv, 2017]. Dual scoring. We score each tool three ways (Tab.Ë23). The selective stance scores only the committed (non-abstained) entries; the conservative stance scores the full split with every abstention forced to VALID; the aggressive stance scores the full split with every abstention forced to a flag (HALLUCINATED@0.55, i.e. scored as a flag with confidence 0.550.55, just above the 0.50.5 decision threshold). The conservativeâaggressive gap is the cost of the abstention, and it is non-trivial exactly where coverage is below one: bibtex-updater on dev_public (coverage 0.8220.822) earns selective DR/FPR/F1 of 0.9870.987/0.0990.099/0.9450.945 on the 920920 entries it commits to, while over the full split conservative scoring (abstainâ ) gives 0.7390.739/0.0900.090/0.8150.815 and aggressive scoring (abstainâ ) gives 0.9900.990/0.1810.181/0.9240.924; DeepSeek-R1 on test_public (coverage 0.7830.783) pays +16.2+16.2 p between its conservative and aggressive stances. The selective FPR (0.0990.099) essentially matches the full-coverage headline FPR (0.0920.092): abstention does not add precision for bibtex-updater; it raises recall (DR 0.865â0.9870.865â 0.987) at near-unchanged FPR. Reporting Coverage with the aggressive number closes the obvious loopholeâabstaining on everything would otherwise give perfect committed-precisionâso a tool cannot hide a weak verdict behind UNCERTAIN. Riskâcoverage. For each tool with graded confidences we abstain on the least-confident entries first (smallest |Pâ(hallucinated)â0.5||P(hallucinated)-0.5|) and score FPR on the rest, tracing a riskâcoverage curve (data in the released coverage_reporting.json; we summarize each curve by its FPR at 90%90\% coverage in Tab.Ë23b). Abstaining on the least-confident decile yields a real FPR reduction for the better-calibrated LLM verifiers: the high-AUROC verifiers (Sonnet 4.6, Opus 4.7; AUROC 0.910.91â0.930.93) sit at the favorable corner with FPR â0.10â0.10â0.120.12 at 90%90\% coverage, while the recall-aggressive open-weight models stay above 0.50.5. bibtex-updater is the opposite case: its riskâcoverage curve is flatâFPR is essentially constant (0.0920.092 at full coverage, 0.0990.099 on the committed subset)âso selective prediction adds no precision. Where the LLM verifiers lower their false-positive rate as they abstain, bibtex-updaterâs precision is structural; it is the precision-anchored reference against which the LLMsâ abstention gains are read. Three caveats. (i) The two Anthropic dev_public coverage cells are not recoverable (endpoint drift, summary-only predictions; §Ë7)âan OpenRouter re-run lands at FPR 0.1620.162 / 0.1650.165 against the published 0.0720.072 / 0.1270.127, the span of the driftâso they read ân/aâ in Tab.Ë3 and their riskâcoverage curves are appendix-only: the curve shape and AUROC are drift-immune deterministic re-scores, but the operating point is not. (i) LLM abstention is prompt-dependent (§ËH.1: Sonnetâs UNCERTAIN rate moves 2.0%â8.7%2.0\%â 8.7\% by wording), so coverage is read as âcoverage under the default promptâ, not an intrinsic constant. (i) bibtex-updaterâs abstention and an LLMâs occupy the same column but differ in mechanism: the tool abstains because the databases hold no matching record (a data-coverage gap), whereas an LLM abstains because it will not commit (an epistemic gap); under the route-to-human framing both are the same useful action, which is why the column is shared. bibtex-updater abstains on 199/1119199/1119 entries on dev_public (coverage 0.8220.822) and 171/831171/831 on test_public (coverage 0.7940.794). Table 23: Coverage and selective prediction on dev_public (offline re-score; cells match the released aggregates to 5Ă10â45Ă 10^-4). (a) per-tool coverage with conservative (abstainâ ) and aggressive (abstainâ ) DR/FPR/F1 over the full split; the gap is the abstentionâs cost. For bibtex-updater the selective DR/FPR/F1 on the 920920 committed entries is 0.9870.987/0.0990.099/0.9450.945 (dev) and 0.9890.989/0.1060.106/0.9560.956 (test, n=660n=660): its selective FPR essentially matches its full-coverage headline FPR (0.0920.092), so its abstention raises recall, not precision. On test_public (coverage 0.7940.794) its conservative and aggressive stances are 0.7170.717/0.0960.096/0.8080.808 and 0.9920.992/0.1860.186/0.9430.943. (b) FPR at 90%90\% coverage from the riskâcoverage curve. â Anthropic curves are appendix-only with a drift caveat (§Ë7); â- marks tools the threshold sweep did not cover (curves are in the released JSON). In (b), green marks FPR â€.12â€.12 at 90%90\% coverage (the favorable riskâcoverage corner) and red FPR >.5>.5; the gray row in (a) is the co-designed reference. (a) Coverage + dual scoring Tool Cov. Cons. DR/FPR/F1 Aggr. DR/FPR/F1 bibtex-updater (v1.2.0) 0.822 0.739 / 0.090 / 0.815 0.990 / 0.181 / 0.924 Gemini 2.5 Pro 0.967 0.476 / 0.050 / 0.627 0.497 / 0.074 / 0.637 Gemini 2.5 Flash 0.988 0.500 / 0.100 / 0.631 0.508 / 0.105 / 0.636 Mistral Large 0.989 0.716 / 0.251 / 0.742 0.721 / 0.253 / 0.745 DeepSeek-R1 0.984 0.896 / 0.623 / 0.739 0.898 / 0.628 / 0.739 Opus 4.7 â n/a (drift) 0.752 / 0.072 / 0.830 n/a Sonnet 4.6 â n/a (drift) 0.781 / 0.127 / 0.827 n/a (b) FPR at 90%90\% coverage (riskâcoverage summary) Tool AUROC FPR@90% F1@90% Sonnet 4.6 â 0.928 0.120 0.914 Opus 4.7 â 0.906 0.095 0.909 GPT-5.4 0.834 0.222 0.816 Mistral Large 0.744 0.241 0.736 Gemini 2.5 Flash 0.739 0.097 0.578 DeepSeek-R1 0.741 0.581 0.752 DeepSeek-V3.2 0.609 0.699 0.711 Takeaway. The precision-ceiling finding survives the obvious confounds: rankings are invariant to prompt wording (§ËH.1) and decision threshold (§ËH.2), detection leans on title/author content rather than a format tell (§ËH.3), and independent raters reproduce the over-flagging the benchmark measures (§ËH.4). Coverage makes the abstention lever explicit; the better-calibrated LLMs trade a small coverage cut for a real FPR reduction (precision-via-abstention), whereas the co-designed bibtex-updaterâs low FPR (0.0920.092) is structural: its riskâcoverage curve is flat, so abstaining on the 1818â21%21\% of entries it cannot verify recovers recall (DR 0.865â0.9870.865â 0.987) at near-unchanged FPR rather than adding precision. E.3 Regime-conditional deployment guidance At venue-realistic prevalence (âŒ2% 2\%), the precision-bottleneck logic favors low-FPR tools regardless of recall: PPV per flagged entry is âŒ18% 18\% for Opus 4.7, âŒ16% 16\% for the co-designed bibtex-updater (precision-competitive with Opus 4.7 because it flags conservatively, so its low FPR is structural), âŒ11% 11\% for Sonnet 4.6, âŒ4% 4\% for GPT-5.1, and <3%<3\% for high-recall open-weight models (Tab.Ë21). âą Pre-2024, recall-prioritized triage: a high-recall verifier such as an agentic harness (DR â0.97â0.97â0.990.99) or DeepSeek-V3.2 (DR = 0.911), accepting the higher FPR where reviewer capacity absorbs false positives. Pre-2024, precision-prioritized: bibtex-updater (DR = 0.865, FPR = 0.092, PPV âŒ16% 16\% at a 2%2\% base rate, âŒ2 2â33 orders of magnitude cheaper than LLM verifiers; its F1 lead over Sonnet 4.6 narrows but does not reverse on test_public, §Ë6), Opus 4.7 zero-shot (FPR = 0.072), or Sonnet 4.6 (FPR = 0.127). âą Near or beyond training cutoffs: Sonnet 4.6 zero-shot for the best calibrated FPR among LLMs tested. Agentic harnesses match higher detection but inflate FPR over their zero-shot bases (GPT-5.1 +6 p; Sonnet 4.6 +30 p, a 3.4Ă3.4Ă multiplier from a low base); deploy only where reviewer capacity absorbs false positives. âą Past 12 months: no tool we tested is reliable without cutoff-aware prompting (§ËF.2); the addendum recovers FPR on GPT-5.1 and Qwen3-235B at the cost of indiscriminate abstention, barely moves Gemini 2.5 Flash, and abstains selectively only on Sonnet 4.6, so it requires per-provider tuning. Appendix F Failure mode (i): temporal fragility and calibration Training-data cutoffs of evaluated LLMs. The temporal evaluation assumes papers from calendar years 2024â2025 post-date most evaluated modelsâ training data. Tab.Ë24 compiles vendor-disclosed training cutoffs (or the most recent reliable upper bound when no disclosure is available) for every LLM used as a baseline. Three observations frame the temporal results. (i) Every evaluated model has a disclosed or inferred cutoff no later than late 2025, so the 448-entry temporal supplement (drawn from 2024â2025 DBLP proceedings) is partially or fully out-of-distribution for all of them. (i) The two closed-weight models with earliest cutoffs (GPT-5.1, Sep 2024) and the conservative open-weight Gemini 2.5 Flash (Jan 2025) are hardest-hit by the post-cutoff FPR rise, consistent with their reliance on parametric confirmation. (i) DeepSeek models do not publish an official cutoff in their technical reports, so we report the OpenRouter-surfaced metadata and flag it as a provenance gap that future versions of Hallmark should track via release-date probes. Table 24: Training-data cutoffs of LLM baselines evaluated in Hallmark. âDisclosedâ indicates a vendor-published cutoff in an official model card, API documentation, system card, or technical report; âInferredâ indicates an upper bound derived from release date and predecessor disclosures. Anthropic distinguishes âtraining data cutoffâ from âreliable knowledge cutoffâ; we report the latter for Claude variants. Rows above the midrule have cutoffs at or before mid-2025; rows below have post-Aug 2025 cutoffs (the âlater-cutoffâ cohort referenced in §Ë6). Model Training cutoff Source Status DeepSeek-R1 Jul 2024 OpenRouter model metadata Disclosed (via host) Llama 4 Maverick †Aug 2024 Meta release notes; release Apr 2025 Inferred GPT-5.1 Sep 2024 (GPT-5 family) OpenAI system card (GPT-5) Disclosed (family) DeepSeek-V3.2 †Oct 2024 Release date; V3 tech report Inferred Gemini 2.5 Flash Jan 2025 Vertex AI model documentation Disclosed Gemini 2.5 Pro Jan 2025 Vertex AI model documentation Disclosed Qwen3-235B-A22B-2507 Jun 2025 OpenRouter model metadata Disclosed (via host) Qwen3-VL-235B-A22B-Instruct †Jun 2025 OpenRouter metadata; Qwen3 family lag Inferred Mistral Large 2512 †mid-2025 Release date; predecessor cutoff Inferred GPT-5.4 Aug 2025 OpenAI release (gpt-5.4-2026-03-05) Disclosed Claude Sonnet 4.6 †Aug 2025 Release Feb 2026; Anthropic family lag Inferred Claude Opus 4.7 †Oct 2025 Release Apr 2026; Anthropic family lag Inferred 2024â2025 DBLP papers are out-of-distribution to varying degrees across the cohort: late-2025 entries are post-cutoff for every model, while early-2024 entries are post-cutoff only for the earliest-cutoff models (GPT-5.1: Sep 2024; DeepSeek-R1: Jul 2024). Across the nine non-Anthropic LLMs tested on the supplement the temporal hit splits into four regimes: four models with healthy baselines degrade sharply (GPT-5.1, Gemini 2.5 Flash, Mistral Large, Llama 4 Maverick; Î ++35â62 p); three with already-elevated baselines stay near-ceiling or worsen (Qwen3-235B, DeepSeek-V3.2, Qwen3-VL-235B at FPR 0.89, a new cohort maximum); DeepSeek-R1 routes nearly every entry to UNCERTAIN; and Gemini 2.5 Pro over-flags only moderately (FPR 0.25). Pro-tier post-training, not data recency, drives that last case: Gemini 2.5 Flashâsame vendor, comparable cutoffâdegrades sharply (FPR 0.595). Training-data recency alone does not rescue parametric verification. The latest-cutoff non-Anthropic model tested (Qwen3-235B: Jun 2025) still ends in the near-ceiling group, so the gap between âseen during trainingâ and âreliably recalledâ matters as much as the cutoff. The two later-cutoff Anthropic models (Sonnet 4.6, Opus 4.7) sharpen the point: their inferred cutoffs (Aug 2025, Oct 2025) are not radically later than Qwen3-235Bâs, yet their FPR shifts on the supplement are an order of magnitude smaller. Cutoff recency is therefore necessary but insufficient; pipeline-level differences (Reinforcement Learning from Human Feedback (RLHF) for abstention, âI donât knowâ calibration) appear to dominate the temporal-robustness story. Baseline temporal consistency. For bibtex-updater, detection rates are consistent across temporal segments (± 2%); DOI-only improves slightly on newer entries because recent papers more consistently include DOIs. F.1 Temporal supplement (2024â2025) To validate the 60-entry temporal probe at scale (probe composition in §ËA.3), we construct a 448-entry temporal supplement: 300 valid entries scraped from DBLP across six ML venues (NeurIPS, ICML, ICLR, AAAI, CVPR, ECCV) for publication years 2024â2025, plus 148 hallucinated entries generated via the standard perturbation pipeline (58/52/38 across Tiers 1/2/3, spanning all 14 types; Tab.Ë2). This provides 10Ă10Ă more valid entries than the probe (300 vs. 30), yielding substantially tighter confidence intervals on FPR estimates. The 448-entry set (300 valid / 148 hallucinated, full coverage) is the canonical temporal supplement reported here; every released prediction file resolves to the same 448 keys. A later 858-entry superset of the same 2024â2025 DBLP pool backs the extended reviewer-experiment robustness checks (recall probe, late-cutoff control, GPT-5.1 run-to-run variance) reported separately in the released artifacts (results/reviewer_experiments/FINDINGS.md), and is presented alongsideânot in place ofâthe canonical 448-entry table. Tab.Ë25 compares each LLM baselineâs performance on the original benchmark (2021â2023 valid entries) against the temporal supplement (2024â2025 valid entries). The pattern is unambiguous: every modelâs FPR increases dramatically on recent papers, while detection rate inflates because models flag nearly everything as hallucinated. Table 25: Temporal supplement: LLM baseline performance on 2021â2023 benchmark entries (dev_public, n=1,119n=1,119) vs. 2024â2025 supplement entries (n=448n=448). Î and Î report the absolute shift. The top block (Gemini Flash through DeepSeek-R1) covers the six LLMs with training cutoffs at or before mid-2025, sorted by baseline FPR ascending; DeepSeek-R1 is placed last because its 24â25 behavior is anomalous (see â ). Below the midrule are five models evaluated on a later dated snapshot (§ËB.2.4): the two later-cutoff Anthropic models (Sonnet 4.6, Opus 4.7; documented or inferable post-Aug 2025 cutoffs, Appx.ËF) do not show the post-cutoff FPR rise, while Gemini 2.5 Pro over-flags only moderately and Llama 4 Maverick and Qwen3-VL-235B degrade like the top block. â DeepSeek-R1 classified nearly all entries as UNCERTAIN; see footnote below table. In the FPR24â25 column, red marks the over-flagging cluster (FPR â„.59â„.59) and green the two models whose post-cutoff FPR holds; Gemini 2.5 Pro (.250) sits between and is unshaded. Model DR21â23 DR24â25 Î FPR21â23 FPR24â25 Î F124â25 MCC24â25 Gemini Flash 0.500 0.844 +0.344 0.100 0.595 +0.495 0.555 0.251 GPT-5.1 0.837 0.958 +0.121 0.411 0.759 +0.348 0.546 0.246 Mistral Large 0.716 0.952 +0.236 0.250 0.793 +0.543 0.537 0.208 Qwen3-235B 0.860 0.986 +0.126 0.533 0.809 +0.276 0.541 0.245 DeepSeek-V3.2 0.911 0.980 +0.069 0.702 0.759 +0.057 0.558 0.278 DeepSeek-R1â 0.896 0.000 â-0.896 0.623 0.856 +0.233 0.000 0.000 Evaluated on a later dated snapshot (§ËB.2.4): Claude Sonnet 4.6 0.781 0.818 +0.037 0.127 0.120 â-0.007 0.793 0.688 Claude Opus 4.7 0.752 0.716 â-0.036 0.072 0.073 +0.001 0.768 0.669 Gemini 2.5 Pro 0.476 0.626 +0.150 0.050 0.250 +0.200 0.588 0.366 Llama 4 Maverick 0.614 0.939 +0.325 0.146 0.763 +0.617 0.539 0.216 Qwen3-VL-235B 0.860 1.000 +0.140 0.551 0.887 +0.336 0.527 0.201 â DeepSeek-R1 classified nearly all 2024â2025 entries as UNCERTAIN, yielding DR = 0 (no hallucinations confidently detected) and FPR = 0.856 (the few non-uncertain predictions were false positives). Its chain-of-thought reasoning explicitly identifies uncertainty about post-training papers but defaults to over-flagging. The supplement confirms three patterns from the probe, now with substantially greater statistical power, and surfaces a fourth via the later-evaluated models. First, FPR rises sharply on recent papers for most models: GPT-5.1âs FPR climbs from 41.1% to 75.9% (1.8Ă1.8Ă); three conservative-baseline models (Gemini Flash ++49.5 p, Mistral Large ++54.3 p, Llama 4 Maverick ++61.7 p) post the largest absolute jumps, while GPT-5.1âs increase is more modest at ++34.8 p. Second, the detection-rate increases are spurious: the four largest-jump models (Gemini Flash, Mistral Large, Llama 4 Maverick, GPT-5.1) achieve DR â„0.84â„0.84 on the supplement, but only because they flag nearly everything; F1 falls to â0.55â0.55 (from 0.61â0.77 on dev_public) and MCC to â0.25â0.25, confirming near-random predictions. Third, already-aggressive models stay near-ceiling: DeepSeek-V3.2 (baseline FPR 70.2%) shows only a 5.7 p increase but ends at 75.9%; Qwen3-VL-235B is the most extreme, climbing 55.1% â 88.7%, a new cohort maximum. Fourth, a vendor-/pipeline-correlated subset resists: Claude Sonnet 4.6 (FPR 12.0%) and Opus 4.7 (7.3%) hold; Gemini 2.5 Pro (FPR 25.0%) and a later-cutoff GPT-5.4 (FPR 41.3%) over-flag only moderately: pipeline differences appear to dominate over training recency for the residual gap. The 60-entry probe, which additionally covered 2026 arXiv submissions, showed the same FPR-multiplier pattern with a strong negative correlation (r=â0.82r=-0.82) between baseline aggressiveness and relative degradation, evidence that the failure mode generalizes beyond 2024â2025 DBLP. These results establish that LLM citation verification leans heavily on parametric knowledge with a graded temporal boundary near the training cutoff: most models fall into the âflag everything unfamiliarâ failure mode for post-cutoff papers. But the partial-to-full resistance from Gemini 2.5 Pro, GPT-5.4, and the two Anthropic models shows this is not a hard structural limit of parametric verification: training-pipeline interventions (post-training calibration, abstention RLHF) can shift the trade-off without retrieval augmentation, even at comparable training cutoffs. We caution against reading the latest-cutoff modelsâ low post-cutoff FPR as evidence of better calibration: the 2024â2025 supplement is drawn from DBLP proceedings, so for a model whose training data plausibly ingested those same DBLP records, a correct valid verdict is indistinguishable from training-data recall (contamination) rather than principled abstention or verification. Their resistance is therefore not attributed to calibration; disentangling recall from calibration would require post-cutoff valid entries the models provably never saw, which the current supplement cannot guarantee. Takeaway. The post-cutoff failure is graded, not binary: 8 of the 12 evaluated LLMs over-flag on 2024â2025 papers (FPR â0.60â0.60â0.890.89 for the seven near-ceiling models plus Gemini Flash at 0.595), while Sonnet 4.6 (0.120) and Opus 4.7 (0.073) hold and Gemini 2.5 Pro over-flags only moderately (0.250). The 12th LLM, the later-cutoff GPT-5.4, is evaluated in the separate temporal-cutoff probe (§ËF.3) and over-flags only moderately (supplement FPR 0.41). Training-pipeline differences (abstention RLHF, calibration) outweigh recency; donât trust a model just because its cutoff is recent. Caveat: n=2n=2 Anthropic models. F.2 Robustness check: cutoff-aware prompting The temporal results above establish H1: LLMs do not know their own cutoff; post-cutoff FPR rises sharply because the model treats unfamiliar papers as suspicious. A separate hypothesis remains untested: H2, that LLMs can route post-cutoff citations to UNCERTAIN if explicitly reminded of the cutoff. H1 and H2 are orthogonal: H1 failing does not imply H2 must fail (the model may have latent metacognitive capacity that the default prompt does not elicit), and H2 failing would upgrade H1 from âepistemic miscalibrationâ to âstructural blindness to temporal uncertainty.â Prompt addendum. Each cutoff-aware variant appends the following text to the default verification prompt, verbatim: ⏠Note: your training data has a knowledge cutoff. If the citation could post-date your training data, or if you cannot recall the paper with confidence, respond with UNCERTAIN rather than HALLUCINATED or VALID. Do not guess. No other element of the prompt, temperature (T=0T=0), seed (42), or token budget (max_completion_tokens=1024, the non-Anthropic value of §ËB.2.4; Sonnet 4.6 runs through the OpenRouter mirror, which uses the same budget) is changed, so any observed shift in predictions is attributable to the addendum alone. Subset of models. We run the cutoff-aware variant on GPT-5.1, Gemini 2.5 Flash, Qwen3-235B, and Claude Sonnet 4.6. DeepSeek-R1 and DeepSeek-V3.2 are omitted because they already saturate at UNCERTAIN (or near-UNCERTAIN) on 2024â2025 entries under the default prompt (Tab.Ë25, footnote â ): there is zero headroom for the addendum to shift their behavior, and any measured âimprovementâ would be dominated by noise in the few non-UNCERTAIN predictions. Mistral Large is omitted because its FPR dynamics closely track Qwen3-235B. The Sonnet 4.6 sweep was added later (2026-07-02, through the OpenRouter mirror) with both a cutoff-aware and a same-snapshot default pass on the canonical 448-entry supplement and the 150-entry pre-cutoff sample, so its comparison is fully within-run and unaffected by the endpoint drift of §Ë7. The four-model subset spans the full conservativeâaggressive FPR spectrum observed in Tab.Ë25: Sonnet 4.6 anchors the calibrated end, Gemini Flash and GPT-5.1 populate the conservative-to-middle range, and Qwen3-235B the aggressive end. Metrics. We report four per-segment metrics (detection rate, FPR, UNCERTAIN rate, coverage), stratified by each modelâs training-data cutoff (Tab.Ë24) into pre-cutoff and post-cutoff entries. The comparison is computed from cached predictions via hallmark.evaluation.temporal.compare_prompt_variants, so the ablation is reproducible from the released prediction dumps without re-running the API calls. The default-prompt FPR baseline here (FPRdef_def in Tab.Ë26) is this ablationâs own snapshot, so its absolute levels differ by a few points from the main temporal supplement (Tab.Ë25; e.g. Gemini 2.5 Flash 57.5%57.5\% here vs. 59.5%59.5\% there, GPT-5.1 72.6%72.6\% vs. 75.9%75.9\%, Sonnet 4.6 18.6%18.6\% vs. 12.0%12.0\% on its 2026-07-02 snapshot); read the Î within this experiment, as elsewhere for drift-sensitive endpoints. Table 26: Cutoff-aware ablation results. Post-cutoff: 2024â2025 temporal supplement (N=448N=448 unique entries, all post-cutoff for the four tested models). Pre-cutoff: 150-entry stratified sample of dev_public (2021â2023, pre-cutoff for all four models; 75 valid/75 hallucinated). Default pre-cutoff UNCERTAIN rates are taken from the full dev_public evaluation for the top three models; the Sonnet 4.6 rows (below the midrule) were collected on a later dated snapshot (2026-07-02) with a same-snapshot default baseline on both pools, so its comparison is fully within-run (its default supplement FPR reads 18.6 here against 12.0 in Tab.Ë25, the span of the endpoint drift). Î columns report cutoff-aware minus default. Red marks the mitigationâs cost: pre-cutoff abstention inflated by the addendum, so the FPR recovery is not a free deployment win. Post-cutoff (N=448N=448) Pre-cutoff (N=150N=150) Model FPRdef_def FPRCA_CA Î UNCCA_CA UNCdef_def UNCCA_CA GPT-5.1 72.6 0.0 -72.6 94.4 0.0 52.7 Gemini 2.5 Flash 57.5 47.0 -10.5 19.2 1.2 0.7 Qwen3-235B 78.5 0.0 -78.5 99.8 0.2 95.3 Claude Sonnet 4.6 18.6 9.7 -8.9 48.9 3.3 8.7 Outcome classification (per model). The four tested models exhibit qualitatively different outcomes, so a single aggregate classification would obscure the signal. âą GPT-5.1: FPR drop =72.6=72.6 p, pre-cutoff UNCERTAIN=52.7=52.7%, post-cutoff UNCERTAIN=94.4=94.4% â partial metacognition. âą Gemini 2.5 Flash: FPR drop =10.5=10.5 p, pre-cutoff UNCERTAIN=0.7=0.7%, post-cutoff UNCERTAIN=19.2=19.2% â strong H1 confirmation. âą Qwen3-235B: FPR drop =78.5=78.5 p, pre-cutoff UNCERTAIN=95.3=95.3%, post-cutoff UNCERTAIN=99.8=99.8% â partial metacognition. âą Claude Sonnet 4.6: FPR drop =8.9=8.9 p (18.6%â9.7%18.6\%â 9.7\% on committed entries), pre-cutoff UNCERTAIN=8.7=8.7%, post-cutoff UNCERTAIN=48.9=48.9% â selective abstention, the closest to H2 in the cohort. Sonnet 4.6 comes closest to H2 holds: it is the only model whose abstention inflates selectivelyâpost-cutoff UNCERTAIN rises to 48.9%48.9\% while pre-cutoff abstention stays below 9%9\%âand its committed-entry FPR halves (18.6%â9.7%18.6\%â 9.7\%). The selectivity still routes roughly half the post-cutoff entries to a human, so abstention remains the mitigationâs price. The other three respond in qualitatively different ways: GPT-5.1 and Qwen3-235B lose discrimination (post-cutoff UNCERTAIN above 94%, with pre-cutoff abstention inflating by 52 p and 95 p respectively), while Gemini 2.5 Flash is barely moved by the reminder (FPR drop only 10.5 p while pre-cutoff UNCERTAIN stays near zero). The practical implication stands: prompt-level temporal-awareness mitigations are not portable across modelsâthe four responses differ qualitativelyâand cannot be relied on as a substitute for retrieval-augmented verification. Takeaway. The cutoff-aware addendum confirms epistemic miscalibration over structural blindness: on GPT-5.1 it drops post-cutoff FPR from 72.6% to 0.0% but inflates pre-cutoff UNCERTAIN to 52.7%, trading discrimination for abstention; Sonnet 4.6 is the only model that abstains selectively (post-cutoff UNCERTAIN 48.9% against 8.7% pre-cutoff) while halving its committed-entry FPR (18.6%â9.7%18.6\%â 9.7\%). Prompt-level mitigation works yet isnât a free deployment win. Caveat: n=4n=4 models tested. F.3 Later-cutoff robustness check: GPT-5.4 on the temporal supplement H1 predicts that the FPR rise on post-cutoff entries is a direct consequence of the training-data cutoff: moving the cutoff forward should shift which entries the model recognizes and, in turn, drop the FPR on entries that cross the new cutoff line. We probe this with a single later-cutoff model: GPT-5.4 (gpt-5.4-2026-03-05), released on 5 March 2026 with an August 31, 2025 training cutoff, roughly 11 months later than GPT-5.1âs September 2024 cutoff. We run it zero-shot on the same 448-entry 2024â2025 temporal supplement used in Tab.Ë26 and compare entry-by-entry against the existing GPT-5.1 default-prompt predictions. GPT-5.4 is reported here as a single-model temporal-cutoff probe; its zero-shot dev_public performance is also included as a baseline row in Tab.Ë3. The GPT-5.4 per-entry predictions on all 448 supplement entries are released, so the probeâs aggregate FPR (0.413) and DR (0.899) are reproducible offline without further API calls. Table 27: Stratified GPT-5.1 vs. GPT-5.4 on the 2024â2025 temporal supplement (N=448N=448). Stratum is the entryâs publication year. 2024 & earlier (236 entries: 161 valid, 75 hallucinated) is post-cutoff for GPT-5.1 but pre-cutoff for GPT-5.4. 2025 (191 entries: 139 valid, 52 hallucinated) is post-cutoff for both, although some venues (ICLR/ICML/AAAI/CVPR/ACL 2025) predate August 2025. Future (21 entries: all hallucinated future_date types) is post-cutoff for both and contains no valid entries, so FPR is undefined. GPT-5.1 cells are computed from its released per-entry predictions with abstentions excluded (9/448), so the all-448 FPR reads 74.2%; scoring its five valid-entry abstentions as flagged reproduces the 75.9% of Tab.Ë25. The red-shaded cell is the residual: even on pre-cutoff entries GPT-5.4 over-flags 28%28\% of valid papers, so a later cutoff alone does not close the gap. FPR (%) DR (%) Stratum N GPT-5.1 GPT-5.4 Î GPT-5.1 GPT-5.4 2024 & earlier 236 53.1 28.0 â25.2-25.2 92.0 85.3 2025 191 99.3 56.8 â42.4-42.4 100.0 92.3 Future (2026+) 21 â â â 100.0 100.0 All 448 74.2 41.3 â32.9-32.9 95.9 89.9 Two features of the result replicate the paperâs temporal narrative across a model generation: âą FPR drops monotonically with cutoff distance. The 2024-and-earlier stratum shifts from post-cutoff (GPT-5.1) to pre-cutoff (GPT-5.4) and FPR drops by 2525 p. The 2025 stratum is post-cutoff for both models but is closer to the GPT-5.4 cutoff; FPR drops by 4242 p, consistent with a graded notion of âunfamiliarâ rather than a binary cutoff. âą GPT-5.4 still over-flags. On pre-cutoff 2024 entries, GPT-5.4âs FPR is 28%28\%, far from zero. Training cutoff alone does not explain the full FPR burden: long-tail recognition, paraphrased titles, and pre-print/published-version mismatches contribute residual false positives even for papers the modelâs training data should have covered. This tempers a naĂŻve âjust use a newer modelâ mitigation. âą UNCERTAIN rate does not rise. Unlike the cutoff-aware prompt ablation (Tab.Ë26), which inflated abstention without shifting discrimination, moving to a later-cutoff model reduces FPR and drops UNCERTAIN from 2.0%2.0\% to 0.0%0.0\%. The gain comes from familiarity with the newer entries rather than from added caution. We interpret this as evidence that H1 is a genuine temporal phenomenon rather than an artifact of GPT-5.1 specifically, and that the practical mitigation path is retrieval-augmented verification rather than prompt-engineered abstention, consistent with the agentic-baseline finding in Tab.Ë3. Takeaway. H1 replicates across model generations: GPT-5.4âs later cutoff drops FPR by 2525â4242 p stratum-wise (74.2%â41.3%74.2\%â41.3\% overall) and UNCERTAIN falls to 0%0\%: familiarity, not caution. Yet pre-cutoff FPR is still 28%28\%, so âdeploy a newer modelâ does not close the gap; pair it with retrieval-augmented verification. F.4 Calibration-vs-memorization controls Three targeted experiments probe the mechanism behind the temporal-robustness results, using the temporal supplement and a dev_public subsample. Each is reported with its own N and complements Tab.Ë25. All three were run on 2026-05-28 through the Hallmark Python API: GPT-5.1 via the native OpenAI endpoint (gpt-5.1), and the Anthropic and DeepSeek models via OpenRouter (anthropic/claude-sonnet-4.6, anthropic/claude-opus-4.7, deepseek/deepseek-v4-pro). (E1) Recall probe: calibration vs. memorization. A low post-cutoff FPR on a real paper is ambiguous: the model may be calibrated (declining to flag the unfamiliar) or merely recalling a paper its training data contained. This is the contamination confound noted in §Ë6: the supplementâs valid entries are scraped from DBLP, and a model whose training data ingested those same DBLP records can accept them from memory rather than verification. On 150150 valid 20242024â20252025 papers we run, per model, the standard verifier (FPR) and a recall probe that supplies only the title and year and asks for the author list and venue with no lookup (ârecalledâ == predicted-author last-name Jaccard â„0.5â„ 0.5 vs. truth). If a low FPR were memorization-driven, accepted-as-valid verdicts would concentrate on recalled papers. Table 28: Recall probe on 150150 valid 2024â2025 papers. Pâ(validâŁrecalled)P( valid ) vs. Pâ(validâŁnot recalled)P( valid recalled) separates memorization from calibration. Model Recall Verify FPR Pâ(vâŁrec)P( v ) Pâ(vâŁrecÂŻ)P( v rec) GPT-5.1 0%0\% 92%92\% â 8%8\% Claude Sonnet 4.6 10%10\% 24%24\% 87%87\% 50%50\% Claude Opus 4.7 29%29\% 13%13\% 91%91\% %84\% Opus 4.7 accepts 84%84\% of valid papers it cannot recall, so its low FPR reflects genuine calibration rather than memorization; Sonnet 4.6 shows a large recall-conditioned gap (87%87\% vs. 50%50\%), so its resistance is partly memorization-assisted; GPT-5.1 recalls none and flags 92%92\%. This partially resolves the contamination confound: it is largely ruled out for Opus 4.7 and only partly for Sonnet 4.6. Caveats: title-given recall is a lower bound on memorization; abstention and âcannot recallâ are not perfectly separable; N=150N=150, one sample per entry (the GPT-5 endpoint samples stochastically at the requested T=0T=0; see E3). (E2) A third-provider late-cutoff control. As a check that the FPR resistance is not specific to the Anthropic pipeline, we evaluate DeepSeek-V4-Proâa third-provider modelâon a fixed stratified n=300n=300 subsample of the supplement, alongside GPT-5.1 and Sonnet 4.6 re-run on the same entries. This is weaker than the intended control: DeepSeek-V4-Pro matches the Anthropic cutoff rather than exceeding it, and the reading below rests on the cross-model ordering, not its absolute FPR. On the complete run its FPR is 0.360.36 at full coverage (two residual uncertain verdicts; DR 0.870.87), versus 0.930.93 for GPT-5.1 and 0.290.29 for Sonnet 4.6 on the same entries, and its year breakdown (0.290.29 on 2024 entries, 0.450.45 on 2025) shows the same graded recency pattern as the cohort. The absolute levels here sit above Tab.Ë25 (e.g. Sonnet 0.29 vs. 0.120) because this control ran on a later endpoint snapshot and an n=300n=300 subsample; we read the within-run cross-model ordering, not the level (§Ë7). A non-Anthropic model thus also resists the post-cutoff FPR rise, consistent with the effect tracking training recency across providers. Caveats: a single third-party control; provider-reported cutoff; contamination is not separable for this model (no recall probe run on it). (E3) Run-to-run variance at the GPT-5 endpoint. We send gpt-5.1 temperature=0.0, but the GPT-5 endpoint samples stochastically regardless of the requested value, so each baseline run is one stochastic draw. Three independent GPT-5.1 runs on a fixed n=150n=150 dev_public subsample give F1 =0.815±0.002=0.815± 0.002, FPR =0.528±0.009=0.528± 0.009, DR =0.965±0.000=0.965± 0.000 (mean ± sample std; 1/1501/150 label flips). The F1 run-to-run std (0.0020.002) is far below the Sonnet 4.6â-GPT-5.1 F1 gap (0.0690.069, â30Ăâ30Ă) and the Sonnet 4.6â-Opus 4.7 gap (0.0160.016, â7Ăâ7Ă), so single-run rankings are stable to sampling noise. Caveat: N=3N=3 runs; n=150n=150 subsample. Takeaway. The recall probe separates the two readings of a low post-cutoff FPR: Opus 4.7 accepts 84%84\% of valid 2024â2025 papers it cannot recall, so its resistance is largely calibration; Sonnet 4.6âs recall-conditioned gap (87%87\% vs. 50%50\%) makes its resistance partly memorization-assisted; GPT-5.1 recalls none and flags 92%92\%. A third-provider control (DeepSeek-V4-Pro) resists alongside themâread the within-run ordering, as its 23%23\% UNCERTAIN/parse failures deflate the levelâand GPT-5.1âs run-to-run F1 std (0.0020.002) sits âŒ7Ă 7Ă below even the smallest headline F1 gap (0.0160.016), so single-run rankings are stable to sampling noise. Caveats: title-given recall lower-bounds memorization; n=150n=150â300300 per probe. F.5 Thinking-budget regime boundary Motivation. The main-table baselines all share a fixed max_completion_tokens budget (1024 tokens on the OpenAI/OpenRouter endpoints, including the chain-of-thought DeepSeek-R1 baseline; 256 on the Anthropic native endpoint; see §ËB.2.4). At this budget, the structured JSON-output contract holds across the entire cohort. Three families of post-snapshot thinking-tier models break this contractâGemini 3.1 Pro/Flash-Lite, DeepSeek-V4, and Qwen3.5âemitting chain-of-thought scratchpad before the verdict and exhausting the budget mid-reasoning. During the codebase additions for the Q2 2026 frontier cohort we observed concrete failure modes that motivated their exclusion from the main table: âą google/gemini-3.1-pro-preview: unbounded thinking tokens overrun the 1024-token max_completion_tokens budget, causing JSON parse failures across nearly all entries; google/gemini-2.5-pro (the non-thinking GA tier) is the closest reliable substitute and is included in the main table. âą google/gemini-3.1-flash-lite-preview: supports the full thinking range (minimalâ ); same unbounded-thinking risk as gemini-3.1-pro-preview. âą qwen/qwen3.5-397b-a17b: âŒ40% 40\% empty responses (thinking mode consumed the budget before the JSON object closed). âą qwen/qwen3.5-122b-a10b: parses successfully but emits 500+ reasoning tokens per reply, inflating cost âŒ5Ă 5Ă versus the matched non-thinking baseline. âą deepseek/deepseek-v4-pro and deepseek-v4-flash: thinking models with reasoning_effort levels high/xhigh supported; same budget-exhaustion risk as the V3-R1 chain-of-thought baseline at 1024 tokens. Smoke-test design. To quantify the regime boundary rather than assert it, we run a stratified n=100n=100 subsample of dev_public (proportional across the 11 main hallucination types, â„5â„ 5 entries per type, fixed seed) on three post-snapshot models under two budget regimes: Table 29: Smoke-test cohort and budget regimes. Regime A is a deliberately tight probe at or below the main-table non-thinking budget (1024 tokens on the OpenAI/OpenRouter endpoints; GPT-5.5 is additionally probed at 256 to expose its saturation boundary, Tab.Ë30); Regime B grants thinking-tier headroom equivalent to âŒ8Ă 8Ă that budget. Parse-failure rate (PF) is the primary diagnostic: a model that exceeds PF â„0.30â„ 0.30 under Regime A is excluded from the main table on methodological grounds (the JSON contract does not hold), even if Regime B recovers competitive accuracy. Model Tier Cutoff Regime A Regime B GPT-5.5 non-reasoning May 2026 256 tok 1024 tok Gemini 3.1 Pro thinking â„Q1 2026 2048 tok (thinking_budget 1024) 8192 tok (thinking_budget 4096) DeepSeek-V4-Pro thinking â„Q1 2026 4096 tok (reasoning_effort low) 8192 tok (reasoning_effort high) Reported metrics. Per (model, regime) cell we report DR, FPR, F1, MCC, parse-failure rate (PF) (fraction of entries whose response is missing or non-JSON), mean and p95 output tokens, and $/entry. Cost is computed at the OpenRouter listed rate as of 2026-05-04. Results. Tab.Ë30 reports parse-failure rate, saturation ratio p95/capp_95/cap, and full classification metrics for all nine (model, regime) cells; Tab.Ë31 gives the per-type detection-rate breakdown. Table 30: Thinking-budget regime boundary smoke test on stratified n=100n=100 subsample of dev_public. PF = parse-failure rate; p95/capp_95/cap = saturation ratio (1.00 = the 95th-percentile output hits the budget cap). DR / FPR / F1 / MCC computed treating UNCERTAIN and parse-failure entries as label=VALID. $ is the OpenRouter list-price cost for the cell at May 2026 rates. Red-shaded saturation cells mark the regime where the cap is biting (p95/capâ„0.95p_95/capâ„ 0.95), matching the shaded band of Fig.Ë7. Model Budget n PF p95/capp_95/cap DR â FPR â F1 â MCC â mean tok p95p_95 tok $ GPT-5.5 256 100 0.66 1.000 0.100 0.000 0.182 0.180 249 256 $0.87 1024 100 0.12 1.000 0.557 0.000 0.716 0.523 510 1024 $1.65 4096 100 0.01 0.339 0.643 0.000 0.783 0.592 602 1389 $1.93 Gemini 3.1 Pro 2048 + 1024 reas. 100 0.00 0.998 0.657 0.033 0.786 0.573 834 2044 $1.34 8192 + 4096 reas. 100 0.01 0.969 0.629 0.033 0.765 0.548 1183 7937 $1.87 16384 + 8192 reas. 100 0.01 0.964 0.643 0.033 0.776 0.560 1688 15793 $2.62 DeepSeek-V4-Pro 4096, effort=low 100 0.06 1.000 0.600 0.000 0.750 0.557 1282 4096 $0.16 8192, effort=high 100 0.00 0.366 0.700 0.033 0.817 0.611 1240 3002 $0.16 16384, effort=high 100 0.00 0.211 0.643 0.000 0.783 0.592 1355 3463 $0.17 Table 31: Per-type detection rate across smoke-test cells (stratified n=100n=100 subsample, â„5â„ 5 per type). Each cell reports the fraction of hallucinated entries of that type the model labeled HALLUCINATED; UNCERTAIN and parse failures count as non-detections. Per-type n is small (â„5â„ 5); 95% binomial CI width is roughly ±35± 35 p at n=5n=5, so within-row differences across budgets are noise unless they exceed that band. Useful for relative shape (which types each model finds easiest/hardest at each budget), not for fine absolute comparisons. GPT-5.5 Gemini 3.1 Pro DeepSeek-V4-Pro Type Tier n A B C A B C A B C fabricated_doi 1 5 0.20 0.80 0.80 0.80 0.60 0.60 1.00 1.00 0.80 nonexistent_venue 1 5 0.00 1.00 1.00 0.60 0.80 0.60 1.00 1.00 1.00 placeholder_authors 1 5 0.00 1.00 1.00 1.00 1.00 1.00 0.80 1.00 1.00 future_date 1 5 0.60 1.00 0.80 1.00 1.00 1.00 1.00 1.00 0.80 chimeric_title 2 5 0.20 0.40 0.40 1.00 1.00 1.00 0.40 0.60 0.40 wrong_venue 2 5 0.00 0.60 1.00 0.40 0.40 0.60 0.80 0.80 0.60 swapped_authors 2 5 0.00 0.20 0.60 0.40 0.40 0.40 0.40 0.60 0.40 preprint_as_pub. 2 5 0.00 0.40 0.40 0.40 0.40 0.40 0.60 0.60 0.80 hybrid_fabrication 2 5 0.20 0.60 0.80 0.80 0.80 0.80 0.60 0.80 0.80 near_miss_title 3 5 0.00 0.20 0.40 0.40 0.20 0.40 0.40 0.40 0.40 plausible_fabrication 3 5 0.00 0.00 0.20 0.60 0.40 0.40 0.00 0.40 0.20 merged_citation S 5 0.20 0.80 0.80 0.80 0.80 0.80 0.40 0.60 0.60 partial_author_list S 5 0.00 0.00 0.00 0.20 0.40 0.20 0.20 0.00 0.20 arxiv_version_mismatch S 5 0.00 0.80 0.80 0.80 0.60 0.80 0.80 1.00 1.00 Three archetypes of regime-boundary behavior. The smoke test reveals three qualitatively distinct responses to budget headroom: âą Clean ceiling (GPT-5.5). Saturation 1.00 â 1.00 â 0.34 across the three regimes; parse-failure rate falls 66% â 12% â 1%; F1 climbs 0.18 â 0.72 â 0.78. The model has an implicit reasoning trace that consumes the entire max_completion_tokens budget at 256, and its capability is fully recovered at 4k. This is the textbook regime-boundary behavior. âą Non-converging tail (Gemini 3.1 Pro). Saturation 0.998 â 0.969 â 0.964 from 2k to 16k; mean output grows monotonically (834 â 1183 â 1688) and the worst-case reasoning trace tracks the cap rather than converging to a fixed length. F1 is essentially flat (0.79 // 0.77 // 0.78) and parse failure stays at â€1%†1\%, so the cap is not biting in absolute terms; but the model has no native budget self-regulation we can detect at n=100n=100. This is a stronger negative result than we expected: for some thinking-tier models, structured-output verification is not budget-bounded at any practical budget, and a fair fixed-budget protocol cannot reliably bracket their reasoning trace. âą Clean recovery (DeepSeek-V4-Pro). Saturation 1.00 â 0.37 â 0.21; F1 0.75 â 0.82 â 0.78. Recovers cleanly at 8k under reasoning_effort=high, and 16k adds nothing. Best-cell F1 (0.82) is competitive with the main-table cohort. Figure 7: Thinking-budget regime boundary. F1 vs. saturation ratio p95/capp_95/cap for each (model, budget) cell, where p95p_95 is the 95th-percentile output token count and capcap is the configured max_completion_tokens; saturation â1â 1 means the reasoning trace is hitting the budget ceiling and the structured output is likely truncated. Arrows trace each modelâs trajectory from low to high budget; marker size is proportional to parse-failure rate; the shaded right band (sat>0.95sat>0.95) flags the regime where the cap is biting. Three archetypes emerge. GPT-5.5 (blue) escapes the saturated band once the budget exceeds 1k tokens. DeepSeek-V4-Pro (green) recoversâsaturation drops well below 1.0 once the budget exceeds its typical reasoning-trace length, freeing the JSON contract to complete reliablyâand plateaus. Gemini 3.1 Pro (red) stays inside the saturated band at every tested budget; we detect no native budget self-regulation up to 16k tokens. Reading the table. None of the three models would have replaced a main-table baseline in expectation: their best-cell F1 (GPT-5.5: 0.78; Gemini 3.1 Pro: 0.79; DeepSeek-V4-Pro: 0.82) sits in the existing cohortâs mid-range, comfortably under the cohort-leading independent F1 (Opus 4.7âs 0.830). The protocol-budget concern is empirically resolved for two of three: GPT-5.5 and DeepSeek-V4-Pro reach a clean budget ceiling and fit the fixed-budget protocol once we know what budget to allocate. The third (Gemini 3.1 Pro) does not, which is a methodological caveat for future versions of Hallmark: a fair fixed-budget benchmark must either (i) allocate the highest budget needed by any one model (inflating cost âŒ10Ă 10Ă across the cohort), or (i) use a budget at which the JSON contract holds for all listed models and exclude models whose tail does not converge within it. We choose (i) and report the boundary explicitly. Cost and wall-clock. The full nine-cell smoke test cost $10.77 at OpenRouter list prices (May 2026) and consumed ⌠303 wall-clock minutes (most of which was DeepSeek-V4-Proâs chain-of-thought latency). Scaling to full dev_public (n=1,119n=1,119) at the same price ratios costs âŒ$â120 120, well within a typical per-paper evaluation budget. Wall-clock per cell is ⌠10â65 s/entry depending on tier and budget; cells parallelize across providers without contention. Takeaway. The main cohort is fixed-budget (1024 tokens on the OpenAI/OpenRouter endpoints these smoke-test models use) by construction. Thinking-tier models break the JSON contract at that budget, but two of three (GPT-5.5, DeepSeek-V4-Pro) recover cleanly when the budget bites; one (Gemini 3.1 Pro) does not converge even at 16k tokens: a methodological boundary, not a capability gap. None would have replaced an existing baseline; Hallmark will track thinking-tier as a separate cohort once budgets are calibrated. Appendix G bibtex-updater: the co-designed reference tool This section collects everything on bibtex-updater, the co-designed reference tool: its implementation (§ËG.1), its benchmark evaluation and the co-design-bias controls (§ËG.2), and a tool-augmented LLM baseline (§ËG.2). G.1 Implementation We additionally describe bibtex-updater, a co-developed open-source citation verification tool designed for deployment by venues, reviewers, and authors: not to optimize benchmark scores, but to provide reliable, automated checking that integrates into existing publication workflows. G.1.1 Design goals Three practical requirements drive the design: 1. Zero human effort. Verification must be fully automated: no manual review, no prompt engineering, no LLM inference costs. This rules out approaches requiring human-in-the-loop validation or expensive API calls to language models. 2. Workflow integration. The tool must integrate into existing pipelines: CI/CD (GitHub Actions), pre-commit hooks, Overleaf builds, and one-off command-line checks. A tool that requires a separate platform or manual invocation will not be adopted. 3. Graceful degradation. When APIs are unavailable or rate-limited, the tool should return partial results rather than fail silently. Venues processing hundreds of submissions cannot tolerate flaky infrastructure. G.1.2 Verification pipeline bibtex-updater implements a multi-stage pipeline that processes each BibTeX entry through increasingly expensive checks: Pre-API validation (zero cost). Before any network calls, the tool checks for syntactic red flags: DOIs that fail to resolve (HEAD request to doi.org), future publication years, implausible dates (<1800<1800), and malformed fields. These cheap checks catch Tier 1 hallucinations without API overhead. Multi-source lookup. The tool queries Crossref, DBLP, and Semantic Scholar using title and first-author search, and resolves arXiv IDs against the arXiv Atom API for preprint-specific consistency checks. Each source returns candidate records that are scored using a weighted combination of fuzzy title matching (70%, token-sort ratio) and author Jaccard similarity (30%). The best-scoring candidate across all sources is selected for field-by-field comparison. Post-match analysis. Once a candidate is identified, the tool compares DOI, title, authors, year, and venue against the input entry. Venue comparison uses alias-aware matching for 17 major ML/AI venues (e.g., NeurIPS/NIPS, ICML, ICLR, CVPR), so that common name variations do not trigger false positives. A dedicated preprint detection stage queries Semantic Scholar to identify entries that claim venue publication when only an arXiv preprint exists. Status assignment. Each entry receives one of eight status codes: verified, not_found, hallucinated (match score <0.50<0.50), or specific mismatch types (title_mismatch, author_mismatch, year_mismatch, venue_mismatch, partial_match). The HALLMARK wrapper maps these to binary labels and confidence scores for benchmark evaluation. G.1.3 Deployment modes The tool supports three deployment scenarios: CI/CD integration (--strict flag exits with nonzero code on detection, gating submission workflows), pre-commit hooks (validates .bib files on every commit), and batch processing (concurrent workers with rate limiting for venue-scale use). For venue-scale deployment, we recommend a two-stage workflow: run bibtex-updater to flag suspicious citations, then manual review of high-confidence flags (confidence â„0.7â„ 0.7); at venue-realistic prevalence, expect roughly one true hallucination per six flags (§ËE.1). At NeurIPS scale (approximately 10,000 submissions, approximately 50 references each), this requires under 6 hours with 8 workers. The tool requires no GPU or LLM API keys and is MIT-licensed; code is publicly available at https://anonymous.4open.science/r/hallmark/ (anonymised for review). G.2 Evaluation and co-design bias bibtex-updater [Reizinger, 2025] is a multi-database cross-referencing tool co-developed alongside this benchmark. Its results appear in Tab.Ë3 under the âCo-designed (reference upper bound)â section, separated from independent tools by a rule. We include it in the main table to give a complete picture of the recallâprecision spectrum: bibtex-updater anchors the precision end by flagging conservatively (DR 0.865, FPR 0.092), not by topping detection rate, and it abstains on the âŒ18 18â21%21\% of entries it cannot verify against a backing record (coverage 0.820.82 dev / 0.790.79 test), an abstention that recovers recall; the low FPR comes from conservative matching, not from abstaining. The explicit upper-bound label and this appendix keep the comparison transparent. The co-design relationship means the benchmark taxonomy and sub-test structure were informed by the toolâs detection capabilities, creating a potential circularity that could inflate its apparent performance; readers should treat its numbers as an informative upper bound, not a fair head-to-head comparison with independent tools. Table 32: Results for bibtex-updater v1.2.0 (co-designed alongside this benchmark) on dev_public. Reported separately from independent tools in Tab.Ë3 due to co-design bias concerns. The tool abstains when no backing record is found; its DR/FPR/F1 score abstentions as committed-VALID over the full split (the headline convention), while Cov. reports the 0.820.82 fraction it commits to. Its selective and conservative/aggressive split is in §ËE.2. Tool DR â FPR â F1 â MCC â TW-F1 â ECE â Cov. bibtex-updater 0.865 0.092 0.890 0.771 0.908 0.383 0.82 Per-type breakdown for bibtex-updater. Tab.Ë33 shows per-type detection rates for bibtex-updater. It detects perfectly (a per-type detection rate of 1.000; parenthesized values throughout this passage) on fabricated_doi, placeholder_authors, future_date, chimeric_title, hybrid_fabrication, plausible_fabrication, and merged_citation, and stays strong on author_mismatch (0.985) and near_miss_title (0.923). The weakest types are precisely those where the tool abstains rather than guessesâpartial_author_list (0.219), nonexistent_venue (0.667), preprint_as_pub. (0.677), wrong_venue (0.681), and arxiv_version_mismatch (0.714)âwhere a could-not-verify verdict is scored as VALID in this forced-binary detection rate (the abstention behavior of §ËE.2). These cells are scored from the toolâs per-entry verdicts on the released labels, the same verdicts behind the aggregate in Tab.Ë32. Table 33: Per-type detection rates for bibtex-updater (v1.2.0) on dev_public. Forced-binary detection rate on the released labels, scored from the toolâs per-entry verdicts. Types on which the tool abstains (could-not-verify) score lower, since abstentions count as VALID in the detection rate (§ËE.2). Tier Type bibtex-updater 1 fabricated_doi 1.000 nonexistent_venue 0.667 placeholder_authors 1.000 future_date 1.000 2 chimeric_title 1.000 wrong_venue 0.681 author_mismatch 0.985 preprint_as_pub. 0.677 hybrid_fabrication 1.000 3 near_miss_title 0.923 plausible_fabrication 1.000 Stress merged_citation 1.000 arxiv_version_mismatch 0.714 partial_author_list 0.219 Readers should interpret bibtex-updaterâs results with caution: (1) its multi-database verification strategy was developed concurrently with the benchmark design; (2) the sub-test structure mirrors the toolâs internal verification pipeline; (3) the pre-screening layer was designed to complement its known blind spots. We encourage independent tool developers to submit evaluations via Hallmarkâs contribution framework to establish unbiased baselines. One observation bounds the circularity concern from the other side. The ground-truth audit recovered 52 real papers (27 dev_public, 25 test_public) that earlier labeling had wrongly marked HALLUCINATED; on those entries bibtex-updaterâs database cross-check returned a VALID verdict that diverged from the then-current (buggy) labels rather than tracking them. A tool whose apparent accuracy were merely an artifact of co-design would mirror the benchmarkâs labels, including their errors: the divergence is evidence that its detections rest on external database evidence, not on the label distribution. The tool abstains on entries it cannot back with a record, reporting dev_public DR 0.865 / FPR 0.092 / F1 0.890 (Tab.Ë32). Its VALID verdicts on the recovered papers show its judgments track external evidence, not the label distribution; the relabel removed an artifact that had flattered the more aggressive LLM verifiers more than the rule-based tool (§Ë6). Cross-domain probe. The released test_crossdomain split (500 entries: 200 valid / 300 hallucinated; composition in §ËA.3) evaluates bibtex-updater v1.2.0 outside the ML-venue regime the main splits sample. Detection transfers: DR reaches 0.890, against 0.865 on dev_public. The released split reports FPR 0.375 (Tab.Ë34), but two confounds inflate that number. First, the splitâs valid biomedical entries are all dated 2026, post-cutoff for every LLM (the recency confound we take up below); bibtex-updater queries live databases and carries no cutoff, so this does not affect its own verdicts. Second, and specific to the rule-based tool, our biomedical scrape carried metadata noiseâVancouver-style author initials, epub-versus-print year ambiguity, preprint venue stringsâthat its exact-match verification reads as contradictions and flags, manufacturing false positives that reflect the citation record rather than the toolâs judgment. A recency-matched, canonically-resolved rebuild removes both: 152 valid biomedical entries from 2021â2023, each field replaced by the CrossRef record its DOI resolves to (canonical author, year, venue), plus the same three hallucination tiers. On that clean split bibtex-updaterâs headline FPR is 0.112, close to its in-domain 0.092, though it commits to only 29% of the valid biomedical entries, abstaining on the other 71% (against 18% in-domain). The canonical metadata converts the released splitâs mis-flags into honest abstentions: out of domain the tool abstains rather than over-flags, because its ML-tuned assumptions (full author names, one canonical year, exact venue strings) cannot confirm biomedical records against its backing databases. The cost of leaving the regime is coverage, not precision (Tab.Ë34). Table 34: bibtex-updater v1.2.0 in and out of regime. dev_public (ML venues, 2021â2023) vs. the released test_crossdomain split vs. a recency-matched, canonically-resolved biomedical rebuild (matched: 152 valid entries from 2021â2023, each field resolved to its DOIâs CrossRef record). Abstentions scored committed-VALID (the headline convention); Cov. is the committed fraction over all entries. The released splitâs FPR (red) is inflated by post-cutoff recency and scrape-metadata noise; canonical resolution returns the FPR to its in-domain level (green) while coverage falls to 0.582 (the tool abstains on 71% of the valid biomedical entries). Split DR â FPR â F1 â Cov. dev_public 0.865 0.092 0.890 0.82 test_crossdomain 0.890 0.375 0.832 0.736 test_crossdomain_matched 0.777 0.112 0.847 0.582 Recovering coverage with a Stage-2 diagnoser. bibtex-updaterâs out-of-domain abstention is a coverage gap that a second stage can fill. We run the paperâs two-stage cascade on the matched split: Stage 1 bibtex-updater decides the 263 entries it can confirm or refute and defers the 189 it cannot to Stage 2, an agentic Sonnet 4.6 diagnoser (up to five tool calls per entry).666Stage 1 is served from the persisted standalone verdicts; re-running bibtex-check live hangs on arXiv DOI HEAD checks against the fabricated future-dated DOIs, so the live cascade is impractical here. Stage 2 is a dated OpenRouter snapshot (§ B.2.4). Stage 2 recovers the coverage: the cascade commits on all 152 valid biomedical entries (valid-entry coverage 0.29â1.000.29â 1.00) at FPR 0.189 and DR 0.975 (F1 0.941; the aggressive stance is nearly identical at FPR 0.211, since Stage 2 leaves almost nothing uncertain). Closing the coverage gap costs little precisionâheadline FPR 0.112 with bibtex-updater abstaining, 0.189 with the cascade committing on everythingâand stays well below the recall-aggressive modelsâ over-flagging (Qwen3-235B 0.974): out of domain, the abstention is recoverable. The LLM cohort: domain versus recency. The zero-shot LLM verifiers of Tab.Ë3 were unevaluated cross-domain; we run eleven of them on both the released 2026 biomedical entries and the recency-matched rebuild, which turns the released splitâs confound into a clean 2Ă22Ă 2 over domain and recency (Tab.Ë35).777DeepSeek-R1 is excluded: at the 120 s client timeout used for every model its reasoning exceeds the budget on roughly 40% of biomedical entries, so a matched-split FPR would rest on a coverage-biased subset; on the entries it does answer its FPR is close to its in-domain rate. The released splitâs high biomedical FPR is a post-cutoff artifact: 89% of the biomedical false positives there give the future date as the reason (â2026 is in the future,â âbeyond my training cutoffâ), and on the recency-matched entriesâsame domain, pre-cutoffâthat share falls to 2%, with the remaining false positives turning into genuine recall failures (an unresolvable DOI, an unfamiliar venue). The residual pure-domain effect (matched FPR minus dev_public FPR, recency held fixed) is model-dependent: negative or small for the calibrated verifiers (GPT-5.1 â0.31-0.31, GPT-5.4 â0.12-0.12, Sonnet 4.6 â0.03-0.03, Gemini 2.5 Pro â0.03-0.03, Llama 4 Maverick â0.09-0.09), and large only for a recall-aggressive model (Qwen3-235B +0.44+0.44), which flags legitimate bioRxiv preprints cited in article form as fabrications. Biomedical citations are not inherently harder to verify: the cross-domain FPR rise was recency, and what domain adds is small except where a model already over-flags in its own regime. Table 35: Domain Ă recency decomposition of LLM false positives. Each cell is FPR on valid entries. Columns cross domain (in = ML dev_public; out = biomedical) with recency (pre-cutoff 2021â2023; post-cutoff: in-domain = the 2024â25 temporal supplement, out-domain = the released 2026 biomedical entries). Îdom _dom is the pure domain effect (out-pre â- in-pre, recency held fixed). Two readings: (i) the out-domain post-cutoff column is â1.0â1.0 for every model, reflecting the post-cutoff date heuristic (89% of those false positives cite the future date) rather than domain transfer; (i) once recency is matched, Îdom _dom is negative or small for calibrated verifiers and large only for the recall-aggressive Qwen3-235B. Sorted by out-domain pre-cutoff FPR; DeepSeek-R1 omitted (see text). pre-cutoff post-cutoff Model in-dom. out-dom. Îdom _dom in-dom. out-dom. Gemini 2.5 Flash 0.100 0.013 â0.09-0.09 0.595 0.991 Gemini 2.5 Pro 0.050 0.020 â0.03-0.03 0.250 1.000 Llama 4 Maverick 0.146 0.059 â0.09-0.09 0.763 1.000 Claude Sonnet 4.6 0.127 0.099 â0.03-0.03 0.120 0.957 GPT-5.1 0.411 0.105 â0.31-0.31 0.759 1.000 GPT-5.4 0.228 0.112 â0.12-0.12 â 1.000 Claude Opus 4.7 0.072 0.125 +0.05+0.05 0.073 0.948 Mistral Large 0.250 0.178 â0.07-0.07 0.793 1.000 DeepSeek-V3.2 0.702 0.408 â0.29-0.29 0.759 1.000 Qwen3-VL-235B 0.551 0.553 +0.00+0.00 0.887 1.000 Qwen3-235B 0.533 0.974 +0.44+0.44 0.809 1.000 Provider content filtering on biomedical citations. Both Anthropic models we query through OpenRouter (Claude Opus 4.7 and Sonnet 4.6) return an empty, content-filtered response (finish_reason: content_filter) on the same 5 of 152 valid biomedical entries (3.3%) in the recency-matched split: all virology and immunology papers on SARS-CoV-2 inhibitor resistance, ACE2 receptor-binding and antibody-escape mutations, and HIV envelope neutralization, i.e. dual-useâadjacent titles the provider filter blocks. The refusal is deterministic: it reproduces at completion budgets from 512 to 4096 tokens, so it is not output truncation, and a trivial control prompt to the same endpoint returns normally, so it is not an outage or a credit fault. These entries fall to the verifierâs error path and score as UNCERTAIN, i.e. committed-VALID under our convention, so they do not inflate false-positive rates; the failure is one of availability, not accuracy: a small but real fraction of legitimate biomedical citations that these models decline to evaluate at all. The filter never fires on the ML-venue core, so the failure mode is invisible in-domain and surfaces only out of domain: a deployment consideration specific to safety-filtered providers, orthogonal to the verdict-quality question the rest of this section studies. Takeaway. The cross-domain FPR rise is mostly a recency artifact, not a domain effect. For the LLMs, 89% of the released biomedical false positives are post-cutoff date heuristics that fall to 2% once the entries are pre-cutoff; the residual domain effect is negative or small for calibrated verifiers and large only for a model that already over-flags in its own regime. For bibtex-updater, canonical metadata shows the released 0.375 was inflated by scrape noise: its true out-of-domain behavior is a coverage dropâit abstains on 71% of valid biomedical citationsâat an in-domain-level FPR (0.112). A Stage-2 Sonnet diagnoser recovers that coverage (0.29â1.000.29â 1.00 on valid entries) at FPR 0.189: the out-of-domain abstention is a recoverable gap. Precision is a property of the toolâregime pair, and the pair that matters is domain and recency together. Tool-augmented LLM baseline. To test whether combining LLM reasoning with API-backed verification is greater than the sum of its parts, we augment GPT-5.1 with bibtex-updaterâs structured output. For each entry, we first run bibtex-check to obtain the verification status, mismatched fields, confidence score, and APIs consulted, then inject this evidence into an augmented prompt that instructs the model to use the tool findings as evidence while applying its own judgment. Tab.Ë36 compares the augmented model against its components. Table 36: Tool-augmented GPT-5.1 vs. standalone components on dev_public. The augmented model improves precision and calibration over GPT-5.1 alone but does not match bibtex-updaterâs detection rate. Baseline DR â FPR â F1 â MCC â TW-F1 â ECE â GPT-5.1 (standalone) 0.837 0.411 0.766 0.442 0.822 0.190 bibtex-updater (standalone, v1.2.0) 0.865 0.092 0.890 0.771 0.908 0.383 GPT-5.1 + bibtex-updater 0.843 0.144 0.856 0.698 0.872 0.078 The augmented model improves over standalone GPT-5.1 on every headline metric: detection rate rises slightly (84.3% vs. 83.7%) while FPR drops sharply (0.144 vs. 0.411) and ECE improves to 0.078, the best calibration of any baseline. However, it still falls short of bibtex-updaterâs standalone detection rate (84.3% vs. 86.5%, where the tool abstains rather than guessing on unbacked entries). Tab.Ë37 reveals why. On the two types where bibtex-updater excels and GPT-5.1 strugglesâauthor_mismatch and near_miss_titleâthe augmented model degrades: author_mismatch drops to 47.8%, below even standalone GPT-5.1 (63.2%) and far below bibtex-updaterâs 98.5%, and near_miss_title falls from 62.9% to 50.0%. The LLM overrides tool-detected metadata mismatches, treating them as potential API artifacts rather than genuine hallucination signals. Where the augmented model does improve is on types requiring semantic judgment: chimeric_title rises to 97.9% (from 92.3%) and plausible_fabrication to 88.2% (from 82.1%), suggesting the tool evidence helps the LLM confirm its existing suspicions rather than revise its judgment. Table 37: Per-type detection rates for tool-augmented GPT-5.1 vs. standalone components. The augmented model fails to transfer bibtex-updaterâs strength on metadata-based types. The bibtex-updater (v1.2.0) and Augmented columns are scored from per-entry verdicts on the released labels; the GPT-5.1 standalone column is reproduced from an earlier evaluation run, as its per-entry predictions were not retained. Red marks the two cells where evidence injection falls below both standalone components, the same two hard types shaded in Tab.Ë15. Tier Type GPT-5.1 bibtex- updater Augmented 1 fabricated_doi 0.973 1.000 0.923 nonexistent_venue 0.970 0.667 0.795 placeholder_authors 0.941 1.000 1.000 future_date 1.000 1.000 1.000 2 chimeric_title 0.923 1.000 0.979 wrong_venue 0.833 0.681 0.681 author_mismatch 0.632 0.985 0.478 preprint_as_pub. 0.840 0.677 0.677 hybrid_fabrication 0.673 1.000 1.000 3 near_miss_title 0.629 0.923 0.500 plausible_fabrication 0.821 1.000 0.882 These results demonstrate that naĂŻve evidence injectionâpresenting tool output as LLM contextâis insufficient to realize the complementarity between parametric and retrieval-based verification. The LLM treats tool evidence as advisory rather than authoritative, defaulting to its own judgment when conflicts arise. More structured integration strategiesâsuch as forcing acceptance of tool-detected field mismatches, weighted ensembling, or agentic tool-use where the LLM iteratively queries APIsâmay better combine the complementary strengths. Takeaway. Read bibtex-updaterâs dev_public lead as a precision-oriented reference, not a recall ceiling: its low FPR is structuralâit flags conservativelyâand comes from that conservative matching rather than from the abstention it adds on unverifiable entries. The tool is cross-split stableâFPR rises only +2.4+2.4 p (0.092â0.1150.092â 0.115) and F1 holds (0.890â0.9010.890â 0.901), because abstaining on entries it cannot back with a record keeps it from guessing on test_publicâs harder valid poolâso its F1 lead over Sonnet 4.6 narrows from 6.36.3 p on dev_public to 3.53.5 p on test_public but does not reverse. We still recommend independent submissions, not co-developed tools, to set fair baselines. Appendix H Robustness ablations We report four ablations that probe whether the precision-ceiling finding is an artifact of a single design choiceâa prompt, a threshold, an input field, or a single annotatorâalongside the selective-prediction analysis of §ËE.2. We read all four as robustness evidence for the ranking, not as tuning for best numbers: the tool ranking is what transfers across regimes (§Ë6), so we test its invariance to phrasing (§ËH.1), operating point (§ËH.2), input format (§ËH.3), and rater (§ËH.4). The temporal-mechanism probes (the recall and late-cutoff controls, cutoff-aware prompting, and the thinking-budget boundary) live with failure mode (i) in Appx.ËF; the bootstrap CIs are in §ËB.1. The prompt-variant, field-LOO, and rater runs use a fresh dated OpenRouter snapshot (2026-05-31) that does not reproduce the main-run absolute aggregates; we therefore read each as a within-run difference, robust to endpoint drift (§Ë7). H.1 Prompt-sensitivity sweep We sweep four prompt variantsâdefault (the paperâs prompt), notaxo (taxonomy removed), uncertain (abstention explicitly encouraged), and terse (compressed instructions)âover four models on a stratified n=150n=150 dev_public sample (81 hallucinated / 69 valid), at temperature 0 and seed 42 (Tab.Ë38); GPT-5.1 runs through its OpenAI-direct endpoint, the other three through OpenRouter. The finding is two-part. First, the model ranking is prompt-invariant: mean pairwise Spearman Ï=0.90Ï=0.90 for F1, DR, and FPR across the four variants (Sonnet 4.6 >> GPT-5.1 >> DeepSeek-V3.2 >> Gemini 2.5 Flash on F1 at default), and the sole departure from Ï=1.0Ï=1.0 is a 0.0010.001 F1 near-tie between GPT-5.1 (0.7940.794) and Sonnet 4.6 (0.7930.793) under the uncertain variant: a tie, not a reordering. This is the result that defends the single-prompt design of §Ë5.1. Second, the absolute FPR is wording-sensitive: GPT-5.1 shows the largest swing: FPR 0.5800.580 at default down to 0.2120.212 under terse (â36.8-36.8 p), with its UNCERTAIN rate reaching 19.3%19.3\% under uncertain (from 0%0\%). The abstention-encouraging uncertain variant drops Sonnet 4.6âs FPR from 0.1210.121 to 0.0150.015 (â10.6-10.6 p) and DeepSeek-V3.2âs from 0.8990.899 to 0.6000.600 (â29.9-29.9 p), while Gemini 2.5 Flash moves only â4.3-4.3 p. Sonnetâs UNCERTAIN rate rises from 2.0%2.0\% (notaxo/terse) to 8.7%8.7\% (uncertain), a âŒ7 7 p coverage swing from wording alone. Pooled over all models and non-default variants, the verdict-flip rate against the default is 17.4%17.4\% (all entries) and 13.6%13.6\% when UNCERTAIN-involving flips are excluded, with GPT-5.1 the most prompt-sensitive model (25.8%25.8\% mean flip). The practical consequence is a scope bound on the post-cutoff FPR magnitude: because a single wording change can move FPR by 1010â3737 p, we treat the absolute post-cutoff false-alarm rate (§Ë6) as prompt-conditional and rest the temporal claim on the cross-regime ranking, which this sweep shows is prompt-invariant. Table 38: Prompt-sensitivity sweep (n=150n=150 dev_public; temperature 0; seed 42; OpenRouter models on a fresh 2026-05-31 snapshot, GPT-5.1 via the OpenAI-direct endpoint). DR, FPR, F1, ECE, and UNCERTAIN-rate per model and prompt variant. Model F1/DR ranking is near-invariant to the variant (mean pairwise Spearman Ï=0.90Ï=0.90); the absolute FPR is not (the uncertain and terse variants drop it sharply). Model Variant DR â FPR â F1 â ECE â UNC. Sonnet 4.6 default 0.909 0.121 0.903 0.087 4.7% notaxo 0.744 0.101 0.811 0.098 2.0% uncertain 0.667 0.015 0.793 0.083 8.7% terse 0.759 0.074 0.833 0.113 2.0% GPT-5.1 default 0.914 0.580 0.759 0.233 0.0% notaxo 0.802 0.420 0.743 0.225 0.0% uncertain 0.800 0.250 0.794 0.145 19.3% terse 0.720 0.212 0.755 0.170 6.0% DeepSeek-V3.2 default 0.975 0.899 0.712 0.372 0.0% notaxo 0.938 0.768 0.724 0.335 0.0% uncertain 0.872 0.600 0.735 0.282 4.7% terse 0.889 0.667 0.724 0.315 0.0% Gemini 2.5 Flash default 0.469 0.130 0.594 0.316 0.0% notaxo 0.370 0.087 0.513 0.324 0.0% uncertain 0.383 0.087 0.525 0.314 0.0% terse 0.321 0.101 0.456 0.414 0.0% Mean pairwise Spearman Ï F1/DR: 0.900.90 FPR: 0.900.90 Takeaway. Prompt wording moves the absolute numbers and leaves the ranking alone: a single variant shifts FPR by 1010â3737 p (GPT-5.1: 0.580â0.2120.580â 0.212 under terse) and Sonnetâs coverage by âŒ7 7 p, whereas the model ordering holds at mean pairwise Spearman Ï=0.90Ï=0.90, with the sole departure a 0.0010.001 F1 near-tie. This is why the temporal claims rest on cross-regime rankings and every absolute post-cutoff FPR reads as prompt-conditional (§Ë6). H.2 Threshold and aggregation ablation We re-score the stored per-entry confidences on dev_public (no API calls) to test two operating-point choices: the decision threshold and the field-aggregation rule (Tab.Ë39). Threshold. Confidences are effectively quantized, so threshold tuning buys almost nothing: the gap between the default 0.50.5 threshold and the best-F1 threshold is below 0.350.35 p for seven of eight tools (Opus 4.7 0.220.22 p, Sonnet 4.6 0.180.18 p, GPT-5.4 0.340.34 p, the rest â€0.07â€0.07 p). The lone exception is Gemini 2.5 Flash (8.228.22 p), and it buys F1 only by trading FPR up to 0.3430.343. AUROC orders the verifiers as expected (Sonnet 0.9280.928, Opus 0.9060.906, GPT-5.4 0.8340.834, then the recall-aggressive open-weight models 0.610.61â0.740.74). For the seven quantized-confidence tools the fixed-0.50.5 operating point of Tab.Ë3 is therefore near-optimalâbest-F1 is chosen in-sample on dev_public, so this lower-bounds the true headroomâwhile Gemini 2.5 Flash is the exception noted above. The ranking is threshold-robust regardless, in contrast to peer benchmarks that hard-code an unjustified match threshold. Aggregation. The benchmarkâs detector flags an entry when any one of the cross-database sub-tests fails (the âany-missâ ruleâthe benchmarkâs internal detector, distinct from the agentic harnessâs any-no-match rule in Appx.ËD). Sweeping stricter quorums trades DR for FPR monotonically: on the structured sub-tests, any-miss gives DR 0.9500.950 / FPR 0.0000.000, while requiring two, three, or all four misses reduces DR to 0.2690.269, 0.0480.048, 0.0020.002; on the field-level resolver the same sweep drops DR from 0.7590.759 (any-miss) to 0.1280.128 (unanimous). A noisy-voter ensemble makes the same point with real verifiers: treating the eight zero-shot LLMs with stored per-entry predictions (Tab.Ë14) as eight independent voters, flagging whenever any single voter flags reproduces the any-no-match profile (DR 0.9920.992 / FPR 0.8710.871), simple majority (â„5/8â„ 5/8) gives DR 0.8190.819 / FPR 0.1770.177 / F1 0.8320.832, and supermajority and unanimity buy lower FPR at steep DR cost. The benchmark deploys the any-miss rule (F1 0.9750.975); requiring cross-database agreement instead nudges F1 to 0.9850.985 at a small false-positive cost (FPR 0.000â0.0350.000â 0.035), so any-miss sits within 11 p of the familyâs F1 maximum while holding FPR at zero. Table 39: Threshold and aggregation ablation on dev_public (offline re-score of stored confidences; no API calls). Top: per-tool AUROC, default-0.50.5 F1, and the F1 gap to the best thresholdânear zero for all but Gemini 2.5 Flash, confirming the fixed operating point is near-optimal. Bottom: the field-aggregation sweep; the benchmark uses the any-miss rule, which sits within 11 p of the familyâs F1 maximum while holding FPR at zero, and stricter quorums trade DR for FPR monotonically. (a) Threshold: default 0.50.5 vs. best-F1 (8 tools) Tool AUROC F1@0.5 best-F1 gap (p) Claude Sonnet 4.6 0.928 0.891 0.893 0.18 Claude Opus 4.7 0.906 0.889 0.891 0.22 GPT-5.4 0.834 0.783 0.786 0.34 Mistral Large 0.744 0.742 0.742 0.00 DeepSeek-R1 0.741 0.739 0.739 0.07 Gemini 2.5 Flash 0.739 0.631 0.713 8.22 Qwen3-235B 0.695 0.744 0.744 0.00 DeepSeek-V3.2 0.609 0.727 0.727 0.00 (b) Aggregation rule (cross-DB sub-tests / noisy 8-voter ensemble) Rule DR FPR F1 MCC structured: any-miss (kâ„1kâ„1/4) 0.950 0.000 0.975 0.948 structured: cross-DB-agreement only 1.000 0.035 0.985 0.968 structured: majority (kâ„3kâ„3/4) 0.048 0.000 0.091 0.150 field-level: any-miss (kâ„1kâ„1/4) 0.759 0.139 0.809 0.619 field-level: two-fields (kâ„2kâ„2/4) 0.365 0.031 0.525 0.407 ensemble: majority (â„5/8â„ 5/8) 0.819 0.177 0.832 0.640 ensemble: supermajority (â„7/8â„ 7/8) 0.634 0.057 0.754 0.596 H.3 Input-format and field leave-one-out We test whether detection rests on a surface format tell or on field content by varying the input representation (full BibTeX vs. structured fields) and dropping one field at a time (title / authors / venue / year / DOI), on the same n=150n=150 stratified sample with two models (Tab.Ë40). This is the dynamic counterpart to the static format-tell audit (§ËC.4.2). Two results. Title carries most of the signal. For Gemini 2.5 Flash, dropping the title raises FPR by +35.5+35.5 p over the structured baseline (0.145â0.5000.145â 0.500) and dropping authors by +18.8+18.8 p, while venue, year, and DOI move FPR little; DeepSeek-V3.2âalready near-saturated at FPR 0.810.81 structuredâshows the same direction at smaller magnitude (+12.9+12.9 p for title). Format matters less than content: switching between full BibTeX and structured fields shifts DR/FPR by at most 7.47.4/14.514.5 p. Detection therefore leans on title and author semantics rather than a surface artifact, reinforcing the construct-validity argument that the result is not driven by a single co-designed field or by formatting (§ËG.2 and C.4.2). Table 40: Input-format and field leave-one-out on dev_public (n=150n=150; temperature 0; seed 42; fresh 2026-05-31 snapshot). Î is measured against each modelâs structured-field baseline. Dropping the title or authors spikes FPR; dropping venue/year/DOI does not: detection leans on title/author content, not a format tell. (DOI LOO affects only the 7171 DOI-bearing entries.) Model Condition DR â FPR â F1 â Î Gemini 2.5 Flash full 0.543 0.101 0.667 â structured 0.617 0.145 0.709 (base) â-title 0.789 0.500 0.709 +0.355+0.355 â-authors 0.709 0.333 0.709 +0.188+0.188 â-venue 0.519 0.101 0.646 â0.043-0.043 â-year 0.593 0.275 0.649 +0.130+0.130 â-DOI 0.469 0.145 0.589 0.000 +0.000 DeepSeek-V3.2 full 1.000 0.957 0.711 â structured 0.975 0.812 0.731 (base) â-title 0.988 0.940 0.712 +0.129+0.129 â-authors 0.987 0.896 0.719 +0.084+0.084 â-venue 0.926 0.826 0.704 +0.014+0.014 â-year 0.926 0.754 0.721 â0.058-0.058 â-DOI 0.963 0.841 0.719 +0.029+0.029 H.4 Multi-rater reliability proxy This is an automated multi-rater reliability proxy, not human inter-annotator agreement. Three independent LLM raters (Sonnet 4.6, DeepSeek-V3.2, Gemini 2.5 Pro) label 132 blinded entriesâ80 real-world-incident hallucinations and the 52 relabel-recovered real papersâand we report their agreement (UNCERTAIN mapped to HALLUCINATED, binary; Tab.Ë41). It does not substitute for human IAA, which remains future work (§Ë7); it serves to test whether independent raters reproduce the relabel. Overall agreement is only fair: Fleissâ Îș=0.24Îș=0.24 [Fleiss, 1971] on the âfairâ band of the LandisâKoch scale [Landis and Koch, 1977], with pairwise Cohenâs Îș ranging from 0.450.45 (SonnetâDeepSeek, moderate) down to 0.030.03 (DeepSeekâGemini Pro, slight). Against the corrected ground-truth labelsâthe database-backed labels after the relabel audit (§ËA.3)âDeepSeek-V3.2 agrees best (Îș=0.716Îș=0.716, accuracy 0.8710.871) and Gemini 2.5 Pro worst (Îș=0.026Îș=0.026); the majority vote reaches accuracy 0.7200.720 (Îș=0.340Îș=0.340). The diagnostic split is the point. Majority-vote accuracy is 0.9750.975 on the 80 real-world hallucinations but only 0.3270.327 on the 52 relabel-recovered real papers: the LLM raters confidently agree on genuine fabrications yet systematically over-flag the recovered real papers, the same over-flagging failure mode Hallmark is built to measure. Independent raters would have repeated the original labeling error, which both corroborates that the relabel audit corrected a real bias and motivates the precision-via-abstention contribution (§ËE.2). Table 41: Multi-rater reliability proxy (not human IAA) on 132 blinded dev_public entries (80 real-world-incident hallucinations, 52 relabel-recovered real papers), 3 independent LLM raters, binary (UNCERTAINâHALLUCINATED). The diagnostic split does the work: the raters agree on genuine fabrications but over-flag the recovered real papers, exactly the failure mode the benchmark measures; the red-shaded cell marks that diagnostic. Agreement statistic Value Fleissâ Îș (3 raters, binary) 0.2380.238 (fair) Cohenâs Îș SonnetâDeepSeek 0.4540.454 (moderate) Cohenâs Îș SonnetâGemini Pro 0.2660.266 (fair) Cohenâs Îș DeepSeekâGemini Pro 0.0290.029 (slight) Rater Îș vs. ground truth acc. vs. ground truth DeepSeek-V3.2 0.716 0.871 Sonnet 4.6 0.360 0.727 Gemini 2.5 Pro 0.026 0.576 Majority vote 0.340 0.720 Majority-vote accuracy by pool real-world hallucinations (n=80n=80) 0.975 relabel-recovered real papers (n=52n=52) 0.327 Takeaway. Independent raters reproduce the failure mode the benchmark measures: majority-vote accuracy is 0.9750.975 on the 80 real-world hallucinations yet 0.3270.327 on the 52 relabel-recovered real papers, so three independent LLM raters would have repeated the labeling error the ground-truth audit corrected. With Fleissâ Îș=0.24Îș=0.24 overall, LLM raters are no substitute for database-backed ground truth, nor for the human IAA that remains future work (§Ë7). Acronyms AUROC Area Under the ROC Curve DOI Digital Object Identifier DR detection rate ECE Expected Calibration Error F1 F1 score FPR false positive rate IAA Inter-Annotator Agreement LLM Large Language Model LOO leave-one-out MCC Matthews Correlation Coefficient MDE Minimum Detectable Effect PF parse-failure rate PPV Positive Predictive Value RLHF Reinforcement Learning from Human Feedback TW-F1 tier-weighted F1