Paper deep dive
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2607.23514v1
- Canonical: https://arxiv.org/abs/2607.23514v1
Trouble viewing inline? Open PDF directly →
Full Text
59,901 characters extracted from source content.
Expand or collapse full text
Novel Claim or Déjà Vu? Rethinking “Contamination-Free” Dynamic Evaluation for Multimodal Automated Fact-Checking Haorui He Department of Interactive Media, Hong Kong Baptist University School of Computing and Data Science, The University of Hong Kong harryhe@connect.hku.hk Xinwen Chen Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University xinwwwc@gmail.com Dacheng Wen Department of Interactive Media, Hong Kong Baptist University School of Computing and Data Science, The University of Hong Kong wdacheng@connect.hku.hk Reynold Cheng School of Computing and Data Science, The University of Hong Kong ckcheng@cs.hku.hk Francis C. M. Lau School of Computing and Data Science, The University of Hong Kong fcmlau@cs.hku.hk Yupeng Li ∗ Department of Interactive Media, Hong Kong Baptist University ivanypli@gmail.com Abstract Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most ex- isting static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM’s internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to- date information. To address this, emerging dynamic benchmarks collect claims published after LLMs’ knowledge cut-off dates, assum- ing they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the- art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlight- ing three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09%–29.30% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Con- tamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distort- ing system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation. CCS Concepts • Information systems→Multimedia information systems;• Computing methodologies→ Natural language processing. Keywords Multimodal Automated Fact-Checking; Dynamic Evaluation ∗ Corresponding author. This work was done while Haorui He was under the supervi- sion of Yupeng Li. This work is licensed under a Creative Commons Attribution 4.0 International License. M ’26, Rio de Janeiro, Brazil. © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2213-4/2026/11 https://doi.org/10.1145/3767308.3836298 1 Introduction The rapid and widespread dissemination of (multimodal) misinfor- mation poses a pressing societal challenge that cannot be addressed by professional human fact-checkers alone [1,8,16,19,31], which motivates automated countermeasures. State-of-the-art (SOTA) multimodal automated fact-checking (MAFC) systems, e.g., DE- FAME [3], leverage powerful multimodal large language models (LLMs) to retrieve and analyze textual and/or visual evidence from online sources (e.g., the Web) to verify (check-worthy) claims. 1 The evaluation of MAFC systems relies on benchmarks built from claims drawn from fact-checking articles published by profes- sional agencies. However, these benchmarks, such as AVeriTeC [27], are static and are not updated after their initial construction. Over time, their outdated coverage of claims from past events introduces a risk of contamination: the LLM backbones powering the MAFC systems have seen these events and related context during pretrain- ing, enabling them to verify claims from parametric knowledge alone and to sidestep the core MAFC challenge of retrieving and reasoning over external evidence [31,34]. 2 As illustrated in Fig. 1, such contamination can artificially inflate estimated MAFC performance relative to true capability on genuinely novel claims arising from timely news events that fall outside the LLMs’ training data [31,34]. To address this, recent benchmarks such as VERITAS [25] and XFACTA [34] adopt a “dynamic” eval- uation paradigm. They continuously aggregate claims published after models’ knowledge cut-off dates and regularly refresh the dataset (e.g., monthly or quarterly), aiming to keep the evaluation data unseen and support reliable assessment of MAFC systems. However, four critical issues remain unresolved. First, the extent of contamination in existing static MAFC benchmarks is un- known. Prior studies have not directly quantified the contamina- tion risks in these benchmarks, making it difficult to assess the reliability of their evaluation results. Second, it is unclear how effec- tively dynamic benchmarks eliminate contamination risks. Claims published after LLMs’ knowledge cut-off dates may still be verifi- able using information available before the cut-off, or may revisit previously fact-checked topics [28]. Thus, some seemingly “novel” 1 Check-worthy claims hold significant public interest and affect public behavior [15]. 2 Table 4 provides examples of such contaminated claims. arXiv:2607.23514v1 [cs.CL] 26 Jul 2026 M ’26, November 10–14, 2026, Rio de Janeiro, Brazil.He et al., Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, and Yupeng Li MAFC BenchmarksReal-World Claims Pretraining Corpora Contamination I’ve never seen them! Real-World Performance Benchmark Performance Social Platforms Contamination Detection How does it affect evaluation? RQ3 RQ4 What is the real performance? RQ2 What are the reasons? RQ1 To what extent? MessagingApps WebsitesNews Outlets Post-Cut-Off Claims Pre-Cut-Off Claims Contaminated (Déjà Vu) Claims Uncontaminated (Novel) Claims Figure 1: An overview of contamination in MAFC and our studied research questions (RQs). claims may in fact be déjà vu to LLMs. As a result, dynamic bench- marks could still overestimate MAFC performance on truly unseen claims, undermining the goal of uncontaminated evaluation. These contamination risks in both static and dynamic benchmarks raise a third issue: How does such contamination distort MAFC evalua- tion? While Rothermel et al. [25]observed performance declines on claims published after LLMs’ knowledge cut-off dates, empirical evidence based on grouping claims by these dates remains mixed: Quelle and Bovet[24]report no abrupt drop, and Fontana et al. [7] find declines only for real claims. Moreover, temporal performance shifts may stem from factors unrelated to contamination, such as distributional shifts [16] or changes in claim complexity [29,32]. Consequently, without directly isolating and controlling for con- tamination, it is unclear to what extent existing benchmark scores may be artificially inflated. Finally, given the contamination risks in existing benchmarks and their potential to distort evaluation, the true MAFC capabilities of SOTA LLMs remain unclear. Without benchmarks that rigorously exclude claims verifiable via pre-cut-off knowledge, we cannot assess models’ generalization to genuinely unseen claims from fast-evolving, real-world media environments. To systematically investigate these issues, this work conducts empirical studies to address the following research questions (RQs). •RQ1: To what extent are existing static and dynamic MAFC benchmarks contaminated? To address RQ1, we adopt an established evidence sufficiency evaluation pipeline proposed in AVeriTeC [27] for detecting potential contamination. Specifically, we measure the sufficiency of evidence extracted directly from an LLM’s internal knowledge against the oracle evidence used by human fact-checkers for the same claim, flagging high-sufficiency claims as potentially contaminated. We evaluate six SOTA LLMs, including both proprietary models (e.g., GPT-5.2 and Gemini-3.0- Pro) and open-source models (e.g., DeepSeek-V3.2 and Qwen3.5- 122B-A10B). We then compare the SOTA static MAFC benchmark, AVeriTeC 3 , with ClaimReview2025Q4, our newly constructed benchmark comprising claims published in Q4 2025 to simulate post-cut-off evaluation. As shown in Sec. 4.1, dynamic evaluation reduces but does not eliminate contamination: 17.09%–29.30% of claims remain potentially contaminated. • RQ2: How does contamination arise in dynamic bench- marks? To address RQ2, we conduct a case study of claims whose oracle fact-checking articles were published after LLMs’ knowl- edge cut-off dates, yet whose LLM-generated articles still show high similarity to those oracle articles. As shown in Sec. 4.2, such contamination persists because many newly published claims can be verified using public knowledge available before the cut-off, either directly or by synthesizing multiple known facts. • RQ3: How does contamination affect MAFC performance evaluation? For RQ3, we compare Accuracy and Macro-F1 of the six LLMs on contaminated vs. uncontaminated claim subsets. As shown in Sec. 4.3, contamination can significantly inflate MAFC performance by as much as 11.34 percentage points in Macro- F1 and can even alter model rankings. Further analysis reveals that contamination enables the LLMs to achieve higher retrieval precision, bypassing exploration of the evidence space. •RQ4: How do SOTA LLMs perform under contamination- controlled MAFC evaluation? Building on the findings from RQ1–RQ3, we re-evaluate all models on a unified contamination- controlled set containing only claims that are uncontaminated for all six models under all three similarity metrics. As detailed in Sec. 4.4, DeepSeek-V3.2 achieves the highest Macro-F1 score. However, all models score below 56% Macro-F1, underscoring substantial room for improvement in MAFC solutions. Our code and appendix are available at https://trustworthycomp. github.io/Rethink-MAFC-Eval/. 2 Related Work Automated fact-checking (AFC) systems verify check-worthy claims by retrieving relevant evidence and predicting a verdict, such as Supported, Refuted, or Not Enough Evidence, often accompanied by an explanatory justification. 4 The majority of AFC systems are text-based [8,10,11,17,26,27,37,41]. However, multimodal misinformation, which combines text with other modalities such as images or videos, has been shown to be even more mislead- ing [4,9,14,20,23,39]. Thus, recent research has increasingly focused on MAFC systems, typically built on multimodal LLMs. For example, RAGAR [14] leverages LLMs to generate textual de- scriptions for images to enable multimodal understanding, while DEFAME [3] employs LLMs to dynamically select tools for extract- ing and reasoning over both textual and visual evidence. Most existing systems are evaluated on static benchmarks, such as AVeriTeC [27], MOCHEG [35], VERITE [22], and AVerImaTeC [6]. As discussed in Sec. 1, these benchmarks risk contamination, which can compromise the reliability of evaluation results. To address 3 According to the manual evaluation by Akhtar et al. [1], 28.68% of claims in AVeriTeC either contain multimodal content or require multimodal reasoning to be fact-checked. 4 Definitions of these verdict categories are provided in Appendix A. Novel Claim or Déjà Vu? Rethinking “Contamination-Free” Dynamic Evaluation for Multimodal Automated Fact-CheckingMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. this, the emerging dynamic evaluation paradigm continuously in- troduces unseen, time-stamped data to enhance robustness and miti- gate such contamination risks [13,33,40]. For example, LiveBench [33] automatically sources new questions across various domains, such as reasoning, coding, math, and writing, from recent public bench- marks and updates its evaluation data on a monthly basis. For MAFC, two dynamic benchmarks exist: XFACTA [34], which is no longer actively maintained (last updated in August 2025), and VERITAS [25], the SOTA benchmark that aggregates real-world claims from 108 professional fact-checking agencies and updates quarterly through a fully automated pipeline. Nonetheless, there is still limited understanding of the extent of contamina- tion in existing static and dynamic MAFC benchmarks and its impact on evaluation outcomes. In this work, we leverage the evidence sufficiency evaluation pipeline from AVeriTeC for detecting potential contamination. Our analysis (Sec. 4) reveals that dynamic MAFC benchmarks remain susceptible to contamina- tion, which can significantly distort evaluation results. To mitigate this, we filter out potentially contaminated samples to provide a contamination-controlled assessment of SOTA LLMs. Closest to our work, Yoon et al. [38]examine contamination in MAFC with systems using LLM-based query expansion. However, their study is limited to systems reliant on human-specified work- flows and only evaluates existing static benchmarks. In contrast, we study SOTA MAFC systems that operate as fully autonomous agents and analyze emerging dynamic benchmarks, providing a more forward-looking assessment of contamination risks in MAFC. 3 Contamination Detection for MAFC To address the RQs motivating this work, we first define contamina- tion for the MAFC task, then introduce our contamination detection pipeline, and finally describe the benchmarks we use. 5 3.1 Definition of Contamination in MAFC MAFC systems predict the veracity label of a claim푐. This task unfolds in two sequential stages: (1) evidence retrieval, where the system gathers a set of multimodal evidence itemsE=푒 1 , . . . ,푒 푘 from external sources (e.g., the open web), each item potentially containing textual and/or visual information, to support or refute 푐; and (2) claim verification, where the system aggregates and reasons overEto assign a veracity label to푐. Prior studies [10,26, 42] demonstrate that MAFC performance is largely determined by the evidence retrieval stage. In light of this, we consider a claim as contaminated for a target LLM when the LLM possesses sufficient parametric knowledge matching the evidence used by human fact- checkers. While parametric knowledge can help verify claims about past events, MAFC systems are primarily intended for emerging events with little or no public evidence [25,34]. Thus, this definition captures whether the model can bypass the core MAFC challenge of retrieving evidence from external, in-the-wild sources. 3.2 Contamination Detection Pipeline Based on the definition, the next question becomes: how can we evaluate the sufficiency of the evidence in LLMs’ internal knowledge with respect to the oracle evidence used by humans? To address this, 5 Example prompts used in this section are provided in Appendix B. we adopt the established evidence sufficiency evaluation pipeline proposed in AVeriTeC [27] for potential contamination detection. This choice is aligned with the definition: we measure contamina- tion risk by the extent to which the LLMs can generate sufficient oracle evidence from their parametric knowledge. As shown in Fig. 2, this pipeline comprises the following steps. Step 1: Fact-checking Article Generation. Given a claim푐 in a MAFC benchmark, we prompt a target LLM푚to generate a fact-checking article ̃ 푎 for푐. This setup is intended to elicit the internal parametric knowledge available to the model. Step 2: Evidence Extraction. After obtaining the LLM-generated article ̃ 푎for claim푐and its corresponding human-written oracle fact-checking article푎, crawled from fact-checking agencies, we employ GPT-4o-Mini to extract evidence items from both the or- acle article푎and the LLM-generated article ̃ 푎, yielding the sets E=푒 1 , . . . ,푒 푛 and ̃ E= ̃ 푒 1 , . . . , ̃ 푒 푛 ′ , respectively. 6 Each evidence item is a claim-relevant textual statement extracted from a source article. Although textual in form, it may encode interpretations or reasoning derived from multimodal content (e.g., images, audio, or video). This extraction step reduces stylistic mismatches between human-written and LLM-generated articles while preserving the core content and removing extraneous elements such as cookie notices, advertisements, and other UI noise. Applying the same extraction to both articles ensures a fair comparison. Step 3: Similarity Matching. After extracting the evidence, we require a method to assess the similarity betweenE(oracle) and ̃ E(LLM-generated). Here, we follow the Hungarian-matching protocol from AVeriTeC [27], whereEplays the role of the human- annotated reference evidence, while ̃ Eserves as the “retrieved” evidence derived purely from the target LLM’s internal knowledge. Formally, for a generated evidence set ̃ E and an oracle evidence set E, we compute the claim-level contamination score as follows: 푠( ̃ E,E)= 1 |E| max ∑︁ ̃ 푒 푖 ∈ ̃ E ∑︁ 푒 푗 ∈E 푓( ̃ 푒 푖 ,푒 푗 )푋( ̃ 푒 푖 ,푒 푗 ),(1) where푓is a pairwise similarity function and푋( ̃ 푒 푖 ,푒 푗 ) ∈ 0,1 encodes a one-to-one assignment obtained via the Hungarian al- gorithm. Normalization by|E|prevents artificial inflation from over-generation while penalizing omissions of oracle evidence. The resulting score quantifies the degree of similarity between gener- ated evidence and oracle evidence. Using this pipeline, we compute a claim-level contamination score푠( ̃ E,E)for each claim푐with respect to a target LLM푚. A claim푐is flagged as potentially con- taminated for푚if its score satisfies푠 ≥ 휏, where휏is a predefined threshold. To obtain a benchmark-level contamination score푆for LLM 푚, we average the claim-level scores across all claims: 푆= 1 |C| ∑︁ 푐∈C 푠( ̃ E 푐 ,E 푐 ),(2) whereC denotes the set of claims in the benchmark. 3.3 Evaluated MAFC Benchmarks To apply this pipeline in practice, we select two complementary MAFC benchmarks: one established static benchmark and one newly constructed dynamic benchmark. 6 Appendix C validates that evidence extraction is robust to different LLMs. M ’26, November 10–14, 2026, Rio de Janeiro, Brazil.He et al., Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, and Yupeng Li > Threshold Yes No Web Crawling Fact-Checking Article Generation Claim from Static/Dynamic Benchmark LLM-Generated Fact-Checking Article Evidence Extraction Similarity Matching Oracle Fact-Checking Article Contaminated (Déjà Vu) Claim Uncontaminated (Novel) Claim LLM-Generated Evidence Set Oracle Evidence Set Contamination Score Figure 2: An overview of contamination detection via an evidence sufficiency evaluation pipeline. Static Benchmark. To evaluate the contamination level of static MAFC benchmarks, we use the SOTA AVeriTeC [27] benchmark, which consists of claims from 50 fact-checking agencies, each anno- tated by professional fact-checkers. This benchmark addresses key limitations of previous MAFC benchmarks, such as temporal evi- dence leakage and evidence insufficiency, and is widely adopted by SOTA systems [3,10,37]. For our experiments, we use the AVeriTeC development set, which contains 500 claims. The most recent claim in this benchmark is dated October 31, 2020, posing a significant risk of contamination, as the knowledge cut-off dates for SOTA LLMs are typically no earlier than 2023 (see Table 1). To ensure data quality, we first deduplicate the claims, resulting in 491 unique claims, then crawl their corresponding oracle fact-checking articles using the source URLs provided in AVeriTeC. Dynamic Benchmark. To empirically evaluate post-cut-off claims, we adopt the data collection protocol of VERITAS and ClaimRe- view2024+ [3,25] to source up-to-date fact-checks from the Claim- Review project, which aggregates structured articles from fact- checking agencies worldwide. 7 Given that the knowledge cut-off date of the SOTA LLM GPT-5.2 is August 2025, we collect only claims published in Q4 2025 (Oct.–Dec. 2025) and denote the re- sulting benchmark as ClaimReview2025Q4. 8 To maintain compara- bility with AVeriTeC, we restrict our collection to English claims. For source credibility, we retain only claims fact-checked by the agencies that are signatories of the International Fact-Checking Net- work (IFCN). 9 Then, we deduplicate repeated claims from different agencies. Finally, the raw claim text provided by ClaimReview may be unsuitable for MAFC evaluation. It may reveal the verdict (e.g., “a fake video...”, “a manipulated image...”), consist of incomplete sentences, or contain unnecessary meta-information such as “A viral social media post claims...”, instead of directly expressing a fac- tual proposition. Thus, we leverage GPT-4o-Mini to filter out these malformed claims. After this processing, the resulting ClaimRe- view2025Q4 benchmark contains 901 precise, self-contained claims. 7 As XFACTA is no longer maintained and VERITAS is not open-source at the time of writing, we curate our own dynamic benchmark to support this evaluation using a similar claim source and curation pipeline. Please note that our main contribution is the empirical analysis of contamination, rather than introducing a new benchmark. The curated benchmark is solely for enabling our analysis. 8 We adapt collection code from MisinfoMe [18] and CimpleKG [5]. 9 Full list available at https://ifcncodeofprinciples.poynter.org/signatories. As with AVeriTeC, we crawl the oracle fact-checking articles for each claim from their source URLs. 4 Experiments This section details our empirical evaluation, designed to address the four research questions (RQs) posed in Sec. 1. 4.1 RQ1: The Extent of Contamination To address RQ1, we apply the evidence sufficiency evaluation pipeline to evaluate the contamination levels in existing benchmarks. Experimental Setup. As described in Sec. 3, we quantify con- tamination by measuring the pairwise similarity scores between evidence items in LLM-generated fact-checking articles and those in human-written oracle articles for each claim, on both static and dynamic benchmarks (i.e., AVeriTeC and ClaimReview2025Q4). To compute the pairwise scores, we consider both textual and semantic similarities. For textual similarity, we follow the default configuration in AVeriTeC [27], using METEOR 10 as the pairwise similarity function푓and setting휏=25%. For semantic similarity, we employ two SOTA text representation models: Gemma-Emb- 0.3B 11 and Qwen3-Emb-0.6B 12 and instantiate푓as the cosine simi- larity between the embeddings of each pair of evidence items. To determine the contamination threshold for our semantic metrics, we randomly sampled 180 claims (20% of ClaimReview2025Q4) for human evaluation. Two annotators with postgraduate-level exper- tise in journalism and/or fact-checking, both proficient in English, independently assessed each claim for contamination. They agreed on 97.8% of the claims; the remaining disagreements were resolved through discussion. We used these adjudicated labels as ground truth and performed a threshold-sweep analysis over휏 ∈ [0,1]. 13 At approximately휏=0.6, contamination labels exhibited the highest agreement with human annotations: 75.6% for Gemma-Emb-0.3B and 82.2% for Qwen3-Emb-0.6B. Thus, we selected휏=0.6 as an em- pirically grounded contamination threshold for semantic metrics. 10 https://w.nltk.org/_modules/nltk/translate/meteor_score.html 11 https://huggingface.co/google/embeddinggemma-300m 12 https://huggingface.co/Qwen/Qwen3-Embedding-0.6B 13 Please see Appendix D for the complete threshold-sweep analysis over 휏 ∈ [0, 1]. Novel Claim or Déjà Vu? Rethinking “Contamination-Free” Dynamic Evaluation for Multimodal Automated Fact-CheckingMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. To benchmark contamination levels in SOTA LLMs, we evaluate both proprietary models, including OpenAI’s GPT-5.2, GPT-4o- Mini, Google’s Gemini-3.0-Pro, and Gemini-3.0-Flash, and two open- source models (DeepSeek’s DeepSeek-V3.2 and Alibaba’s Qwen3.5- 122B-A10B). Table 1 provides an overview of these LLMs. All models use a temperature of 0.01 to encourage deterministic behavior. Table 1: Overview of the evaluated LLMs. LLMsOpen-Source Knowledge Cut-Off GPT-5.2×Aug. 2025 GPT-4o-Mini×Oct. 2023 Gemini-3.0-Pro×Jan. 2025 Gemini-3.0-Flash×Jan. 2025 DeepSeek-V3.2✓Unknown Qwen3.5-122B-A10B✓Unknown Results. Table 2 shows the average textual and semantic contami- nation scores between LLM-generated fact-checking articles and or- acle articles, calculated using three metrics across six LLMs. Results are averaged over at least three repeated trials. The results reveal several key findings. Finding 1.1: The dynamic benchmark can effectively alleviate contamination. The dynamic benchmark (ClaimReview2025Q4) consistently exhibits lower contamination levels than the static benchmark (AVeriTeC) across all six models. Under METEOR, the contamination scores decrease by 2.92–4.95 percentage points; under the two semantic metrics, the reductions are 6.20–11.16 percentage points. Finding 1.2: Strong contami- nation signals are widespread across LLMs on AVeriTeC; on ClaimReview2025Q4, peak scores are concentrated in par- ticular models. On AVeriTeC, nearly all models exhibit average contamination scores above the contamination thresholds, with only two exceptions under the strict METEOR metric, which re- quires exact textual matches. GPT-4o-Mini shows the strongest contamination level under METEOR (26.80%), while Gemini-3.0- Pro achieves the highest scores under both semantic metrics (67.00% on Gemma-Emb-0.3B and 65.78% on Qwen3-Emb-0.6B). In contrast, on ClaimReview2025Q4, Qwen3.5-122B-A10B attains the highest scores across all three metrics: 23.16% (METEOR), 60.38% (Gemma- Emb-0.3B), and 58.48% (Qwen3-Emb-0.6B). Table 3 further reports the proportion of contaminated claims in ClaimReview2025Q4 identified by the criterion푠( ̃ E,E) ≥ 휏for each metric and each model. The results provide a more fine-grained view of contamination in the dynamic benchmark, yielding the following findings. Finding 1.3: A non-trivial portion of claims in the dynamic benchmark still faces contamination risk across all models. When considering the intersection of the three metrics (i.e., exceeding the thresholds for all three), the proportion of potentially contaminated claims ranges from 17.09% to 29.30%, depending on the model. This indicates that a dynamic evaluation protocol substantially reduces, but does not fully eliminate, the risk of contamination. Finding 1.4: Semantic metrics identify a markedly larger contaminated subset than lexical match- ing. For every model, the proportions under Gemma-Emb-0.3B and Qwen3-Emb-0.6B are consistently much higher than those under METEOR, indicating that contamination often appears in semanti- cally similar reformulations rather than near-verbatim reuse alone. This finding further suggests that relying only on lexical over- lap can underestimate the practical contamination risk in MAFC benchmarks. Finding 1.5: Qwen3.5-122B-A10B is the most con- taminated model on the dynamic benchmark. It has the high- est contaminated-claim ratio across all three metrics: 35.18% (ME- TEOR), 55.94% (Gemma-Emb-0.3B), and 50.94% (Qwen3-Emb-0.6B). Notably, since the knowledge cut-off date for Qwen3.5-122B-A10B is unknown (see Table 1), the results may indicate that the model could include online content as recent as Q4 2025. To further exam- ine the extent of contamination, Fig. 3 visualizes the claim-level con- tamination score distribution for Qwen3.5-122B-A10B on AVeriTeC and ClaimReview2025Q4. This visualization reveals Finding 1.6: A subset of claims in the dynamic benchmark exhibits ex- tremely high contamination risk. All three distributions retain a visible right tail above the contamination thresholds, with some samples showing exceptionally high scores (e.g., over 80%). 4.2 RQ2: The Cause of Contamination Given that a subset of claims in the dynamic benchmark exhibits strong contamination despite their source articles being published after the LLMs’ knowledge cut-off dates, we investigate the under- lying reasons for this unexpected contamination. Experimental Setup. We present a qualitative case study of four claims from ClaimReview2025Q4, all published in Q4 2025. De- spite their recent publication, these claims exceed contamination thresholds for all three similarity metrics across all six models. Results. The four representative examples presented in Table 4 highlight two main sources of contamination in the dynamic bench- mark, summarized as follows. Finding 2.1: Contamination arises from claims that directly reference facts available before the LLMs’ knowledge cut-off dates. Even when fact-checking ar- ticles are published post-cut-off, claims about historical events, established policies, or other pre-existing facts can often be veri- fied using information already present in the LLM’s training data. For example, claim 1 centers on EMTALA, a law enacted in 1986 that is widely documented in sources like Wikipedia. 14 Similarly, claim 2 concerns Scarborough’s canceled fireworks display in 2022, as widely reported by outlets such as the BBC, 15 and claim 3 dis- cusses the free TV licence policy, established in 2020 and repeatedly covered thereafter. 16 As a result, direct or contextual knowledge of these facts was likely incorporated into the pretraining data of LLMs, bypassing the need for external evidence retrieval that MAFC systems are designed to perform. Finding 2.2: Contamina- tion can also result from synthesizing multiple pre-cut-off facts. Some claims may be novel in their formulation (never having appeared previously), but their veracity can be determined by syn- thesizing multiple pieces of knowledge available before the cut-off. For instance, claim 4 involves Barron Trump’s eligibility for the U.S. Senate in 2028. While this specific claim is new, it can be refuted by combining two well-established facts: Barron Trump’s birth year (2006) and the age requirement for senators (30 years, set in 1787). 14 https://en.wikipedia.org/wiki/Emergency_Medical_Treatment_and_Active_ Labor_Act 15 https://w.bbc.com/news/uk-england-york-north-yorkshire-64139048 16 https://w.gov.uk/free-discount-tv-licence M ’26, November 10–14, 2026, Rio de Janeiro, Brazil.He et al., Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, and Yupeng Li Table 2: Average contamination scores between LLM-generated fact-checking articles and oracle articles from two benchmarks for six models across three metrics. The highest score in each column is bolded; the second-highest is underlined. LLMs METEORGemma-Emb-0.3BQwen3-Emb-0.6B AVeriTeCClaimReview2025Q4AVeriTeCClaimReview2025Q4AVeriTeCClaimReview2025Q4 GPT-5.223.78%19.15%65.28%54.12%63.82%53.47% GPT-4o-Mini26.80%21.93% 63.70%55.28%61.67%53.11% Gemini-3.0-Pro26.58%21.63%67.00%57.82%65.78%56.89% Gemini-3.0-Flash25.42%20.90%65.30%57.12%64.00%55.72% DeepSeek-V3.224.65%20.45%64.88%57.05%63.59%55.65% Qwen3.5-122B-A10B26.08%23.16%66.58% 60.38%64.87%58.48% 0.00.20.40.60.81.0 Score 2.0 4.0 6.0 Frequency AVeriTeC ClaimReview2025Q4 (a) METEOR 0.00.20.40.60.81.0 Score 2.0 4.0 6.0 Frequency AVeriTeC ClaimReview2025Q4 (b) Gemma-Emb-0.3B 0.00.20.40.60.81.0 Score 2.0 4.0 6.0 Frequency AVeriTeC ClaimReview2025Q4 (c) Qwen3-Emb-0.6B Figure 3: Contamination score distributions of Qwen3.5- 122B-A10B on AVeriTeC and ClaimReview2025Q4. Table 3: Proportion of potentially contaminated claims in ClaimReview2025Q4 for each LLM under different metrics. The intersection column indicates claims identified as con- taminated by all three metrics. LLMsMETEOR Gemma-Emb-0.3B Qwen3-Emb-0.6B Intersection GPT-5.222.31%38.62%38.51%17.09% GPT-4o-Mini32.52%42.62%36.40%22.97% Gemini-3.0-Pro28.41%49.06%46.95%22.86% Gemini-3.0-Flash25.97%44.51%40.73%20.09% DeepSeek-V3.225.19%45.95%41.62%20.31% Qwen3.5-122B-A10B 35.18%55.94%50.94%29.30% This case study shows that many contamination signals stem from claims that, although newly published, are verifiable using public information available before the cut-off. 4.3 RQ3: The Impact of Contamination Given that MAFC benchmarks risk contamination, we evaluate how such contamination affects MAFC performance. Experimental Setup. We instantiate the SOTA MAFC framework DEFAME [3], using each LLM selected in RQ1 as the backbone. DEFAME employs the standard ReAct-style [36] agentic paradigm, iteratively reasoning and selecting actions, making it a suitable de- fault for our evaluation. In our experiments, LLMs within DEFAME access the open web by invoking Google Web Search and/or Image Search tools via the Serper API 17 , which returns the top three most relevant web pages per query, including text, images, diagrams, etc. For each LLM, we define a model-specific contaminated subset by taking the intersection of claims that exceed contamina- tion thresholds across all three similarity metrics. The remaining claims constitute the uncontaminated subset. Unlike prior bench- marks that group claims by publication date and thus confound contamination with temporal distribution shifts, both subsets are drawn from the same sources and time frame. 18 We then compare Accuracy and Macro-F1 between these two subsets. Results. Table 5 compares MAFC performance on the contam- inated and uncontaminated subsets for each model. The results reveal four key findings. Finding 3.1: Contamination can signif- icantly inflate evaluation results. Both Accuracy and Macro-F1 17 https://serper.dev 18 Appendix E validates that they are highly comparable in topical and stylistic features. Novel Claim or Déjà Vu? Rethinking “Contamination-Free” Dynamic Evaluation for Multimodal Automated Fact-CheckingMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. Table 4: Contaminated claims from ClaimReview2025Q4. Content is condensed for brevity while preserving original meaning. Contaminated oracle fact-check snippets are highlighted in purple; matched LLM-generated snippets are shown in green. ClaimOracle Fact-ChecksLLM-Generated Fact-Checks Former U.S. President Ronald Reagan signed the Emergency Medical Treat- ment and Active Labor Act, which re- quired emergency rooms to provide emergency health care to anyone, re- gardless of immigration status, in 1986. (Claim 1, published on Nov. 1, 2025) Reagan did sign the Emergency Medical Treat- ment and Active Labor Act in 1986. The act re- quired “participating hospitals” that offer emer- gency services to provide care regardless of pa- tients’ ability to pay or immigration status. The Emergency Medical Treatment and Active Labor Act (EMTALA) was signed into law by Pres- ident Ronald Reagan on April 7, 1986. The enforce- ment and legal interpretation of the law mandate that emergency rooms provide screening and sta- bilization to all individuals, including those regard- less of their immigration status. A town in England canceled a New Year’s Eve fireworks show so a walrus could sleep. (Claim 2, published on Dec. 25, 2025) A real BBC News article confirmed the basics of the story as shared on social media: In 2022, the English town of Scarborough canceled a New Year’s Eve fireworks display at the last minute over fears it could cause distress to a walrus. In December 2022, the seaside town of Scarbor- ough in North Yorkshire, England, officially can- celed its New Year’s Eve fireworks display to pro- tect the welfare of a visiting Arctic walrus. UK residents over the age of 60 can claim a free lifetime exemption from paying the TV licence. (Claim 3, published on Nov. 28, 2025) From 2020, the BBC decided to pay for free TV licences for the households of anyone over the age of 75 receiving Pension Credit. There is no such blanket lifetime exemption for the over 60s. In the United Kingdom, the age threshold for a free TV licence is 75, not 60. Further, eligibility for those aged 75 and over is not automatic; it is contingent upon receiving specific state benefits. Barron Trump announced a 2028 Senate run. (Claim 4, published on Nov. 22, 2025) The U.S. Constitution requires senators to be at least 30 years old. Barron Trump was born on March 20, 2006, making him 19 years old. So in 2028 he would still not meet the constitutional minimum age of 30 for qualification to serve in the U.S. Senate. Barron Trump was born on March 20, 2006. By the end of 2028, he would be 22 years old. He will not turn 30 until 2036; therefore, he is constitutionally ineligible to run for or serve in the U.S. Senate during the 2028 election cycle because the consti- tutional minimum age requirement is 30 years. Table 5: MAFC performance on the contaminated and uncon- taminated subsets of ClaimReview2025Q4. Red superscripts indicate drops in performance when moving from the con- taminated to the uncontaminated subset;∗marks statistically significant drops (푝< 0.05, bootstrap test). LLMs ContaminatedUncontaminated AccuracyMacro-F1AccuracyMacro-F1 GPT-5.261.04%57.68%52.88% -8.16% 52.81% -4.87% GPT-4o-Mini65.22%56.41%55.62% -9.60% ∗ 46.82% -9.59% ∗ Gemini-3.0-Pro71.36%61.07%60.86% -10.50% ∗ 52.50% -8.57% ∗ Gemini-3.0-Flash74.59% 61.22%66.81% -7.78% ∗ 49.88% -11.34% ∗ DeepSeek-V3.257.92%54.79%52.92% -5.00% 52.57% -2.22% Qwen3.5-122B-A10B74.62%57.48%63.74% -10.88% ∗ 47.16% -10.31% ∗ decrease for every model when moving from the contaminated subset to the uncontaminated subset, with statistically significant drops observed for Qwen3.5-122B-A10B, Gemini-3.0-Pro, Gemini- 3.0-Flash, and GPT-4o-Mini. Finding 3.2: The magnitude of this inflation varies substantially across models. Qwen3.5-122B- A10B exhibits the largest Accuracy drop, decreasing by 10.88 per- centage points, and also shows a substantial Macro-F1 decline of 10.31 percentage points. Gemini-3.0-Pro, Gemini-3.0-Flash, and GPT-4o-Mini also exhibit substantial declines, with Accuracy/Macro- F1 drops of (10.50/8.57), (7.78/11.34), and (9.60/9.59) percentage points, respectively, while GPT-5.2 and DeepSeek-V3.2 show no statistically significant degradation. Finding 3.3: The most con- taminated model also suffers the largest Accuracy decline. In Table 2, Qwen3.5-122B-A10B exhibits the strongest contamination levels on ClaimReview2025Q4 across all three metrics. Consistently, this model also demonstrates the largest Accuracy drop when com- paring contaminated to uncontaminated claims. This alignment is consistent with a positive association between the extent of con- tamination and the degree of benchmark score inflation, which also provides converging evidence that our contamination detection pipeline captures a meaningful evaluation artifact. Finding 3.4: Contamination can distort MAFC performance rankings of LLMs. On the contaminated subset, Gemini-3.0-Flash achieves the highest Macro-F1 score (61.22%); however, on the uncontaminated subset, the best Macro-F1 shifts to GPT-5.2 (52.81%). This indicates that benchmark contamination can obscure models’ relative capa- bilities when fact-checking genuinely unseen claims, motivating a contamination-controlled evaluation. Table 6: Trajectory-level reformulation statistics on contam- inated (Con.) and uncontaminated (Uncon.) subsets. Values report the average number of transitions per claim.∗indi- cates a statistically significant increase relative to the corre- sponding Con. counterpart (푝< 0.05; bootstrap test). ModelType Spe. Gen. Exp. Rep. Qwen3.5-122B-A10B Con.0.53 0.470.89 0.04 Uncon. 0.600.45 1.24 ∗ 0.03 Gemini-3.0-Pro Con.0.800.491.220.11 Uncon. 0.94 0.56 1.54 ∗ 0.12 Analysis. To further understand why contamination changes downstream MAFC performance, we follow prior work [2, 12, 21] M ’26, November 10–14, 2026, Rio de Janeiro, Brazil.He et al., Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, and Yupeng Li and analyze trajectory-level reformulations during the fact- checking process. Specifically, for each fact-checking trajectory, we follow Ning et al. [21]and use GPT-5-Nano (temperature 0.01, with three repeated trials) to label every adjacent search query pair(푞 푘 → 푞 푘+1 )as one of four types: Specialization (Spe.) (nar- rowing the query by adding constraints), Generalization (Gen.) (broadening the query by relaxing constraints), Exploration (Exp.) (pivoting to a different facet within the same topic), or Repetition (Rep.) (issuing an identical or near-duplicate reformulation). 19 We then compute the average number of each reformulation type for both contaminated and uncontaminated subsets, focusing on the top two models with the largest Accuracy declines as shown in Table 5, i.e., Qwen3.5-122B-A10B and Gemini-3.0-Pro. As shown in Table 6, the average counts of Specialization, Generalization, and Repetition remain statistically indistinguishable across contami- nated and uncontaminated subsets. However, contaminated trajec- tories consistently devote significantly fewer steps to Exploration, which suggests Finding 3.5: Contamination enables models to retrieve relevant evidence more precisely, bypassing com- prehensive exploration of the evidence space. When facing contaminated claims, LLMs retrieve relevant evidence efficiently, without thoroughly considering alternative queries or being dis- tracted by irrelevant details. As a result, contamination may inflate accuracy scores by shortcutting evidence retrieval. 4.4RQ4: Contamination-Controlled Evaluation Motivated by the above findings, which highlight the significant impact of contamination on evaluation results, we further assess the LLMs in a stricter contamination-controlled setting. Experimental Setup. Using the same DEFAME framework and six LLM backbones, we obtain a unified contamination-controlled evaluation set by retaining only claims that are uncontaminated for all six models under METEOR, Gemma-Emb-0.3B, and Qwen3- Emb-0.6B. Table 7 reports the verdict category distributions and sizes of the three metric-specific contamination-controlled subsets whose intersection defines this final evaluation set. Each model is evaluated on this unified contamination-controlled set. Table 7: Verdict category distributions of contamination- controlled claims under each similarity metric. CategoriesMETEOR Gemma-Emb-0.3B Qwen3-Emb-0.6B Intersection Supported50 (12.92%)23 (10.55%)33 (13.47%)19 (12.34%) Refuted264 (68.22%)157 (72.02%)167 (68.16%)105 (68.18%) Not Enough Evidence73 (18.86%)38 (17.43%)45 (18.37%)30 (19.48%) Total387 (100.00%)218 (100.00%)245 (100.00%)154 (100.00%) Results. The statistics in Table 7 and Table 8 lead to three main ob- servations. Finding 4.1: The evaluation set exhibits noticeable label imbalance. As shown in Table 7, all three contamination- controlled subsets are dominated by Refuted claims. Due to this imbalance, we use Macro-F1 as our primary metric. Future dy- namic benchmark construction should more carefully control label 19 Ning et al. [21]provide examples of these categories and validate the robustness of their method to different LLMs. We follow their default setup to use GPT-5-Nano. Table 8: MAFC performance of LLMs on the unified contamination-controlled evaluation set. LLMsAccuracy Macro-F1 GPT-5.251.95%51.66% GPT-4o-Mini49.35%39.14% Gemini-3.0-Pro59.09%50.88% Gemini-3.0-Flash63.64%48.75% DeepSeek-V3.255.19%55.93% Qwen3.5-122B-A10B62.34% 47.04% distributions. In this regard, mechanisms such as the claim recti- fication step employed in VERITAS [25] may provide a practical framework for rebalancing claim labels. Finding 4.2: DeepSeek- V3.2 achieves the strongest overall MAFC performance in the contamination-controlled setting. Specifically, it attains the highest Macro-F1 (55.93%) on the unified contamination-controlled set. The smaller GPT-4o-Mini exhibits the lowest MAFC perfor- mance, highlighting the impact of model size on fact-checking capa- bility. Finding 4.3: Significant room for improvement remains in MAFC. All models achieve a Macro-F1 below 56%, indicating that current approaches remain far from saturated. 5 Conclusion This work re-examines the assumption that dynamic evaluation is inherently contamination-free for multimodal automated fact- checking (MAFC). To this end, we investigated four research ques- tions: (RQ1) the extent to which existing static and dynamic MAFC benchmarks are contaminated, (RQ2) how contamination arises in dynamic benchmarks, (RQ3) how such contamination affects MAFC evaluation, and (RQ4) how SOTA LLMs perform under contamination-controlled settings. Our empirical study yields 16 findings, with three key results highlighted below. First, dynamic benchmarks reduce contamination, but they do not fully eliminate it. Second, contamination in dynamic benchmarks arises because even newly published claims are often verifiable using public knowl- edge available before the cut-off. Third, contamination significantly inflates evaluation outcomes and can even change model rankings, plausibly because it allows LLMs to shortcut evidence retrieval rather than broadly exploring the evidence space. These results suggest that contamination can obscure MAFC performance on genuinely unseen claims, prompting us to construct contamination- controlled subsets to reassess SOTA LLMs, further highlighting the substantial room for future improvement. Beyond evaluation, our contamination detection pipeline can also serve as a data processing tool for decontaminating MAFC benchmarks. Our work lays the foundation for trustworthy evaluation of MAFC systems in dynamic media environments. We encourage fu- ture research to move beyond timestamp-based filtering by adopting more robust controls for claim novelty, advancing MAFC system development, detecting subtle latent contamination beyond observ- able evidence overlap (for which our similarity-based estimates may be conservative), and expanding empirical studies beyond English to enhance digital media literacy and support integrated information ecosystems for societal well-being. Novel Claim or Déjà Vu? Rethinking “Contamination-Free” Dynamic Evaluation for Multimodal Automated Fact-CheckingMM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. Acknowledgments This work was supported by National Natural Science Founda- tion of China (No. 62202402), Guangdong and Hong Kong Univer- sities “1+1+1” Joint Research Collaboration Scheme, Project No. 2025A0505000001, the Early Career Scheme (ECS) from the Re- search Grants Council of HKSAR (HKBU 22202423), the General Research Fund (GRF) from the Research Grants Council of HKSAR (HKBU 12203425), a grant from the Germany/Hong Kong Joint Research Scheme sponsored by the Research Grants Council of HK- SAR and the German Academic Exchange Service of Germany (No. G-HKBU208/25), the Initiation Grant for Faculty Niche Research Areas 2023/24 (No. RC-FNRA-IG/23-24/COMM/01), Research Clus- ter Matching Scheme (No. RCMS/24-25/01) of Hong Kong Baptist University, the grants from the Research Grants Council of HKSAR (HKU 17202325), the University of Hong Kong (Project 2409100399), the HKU Faculty Exchange Award 2024 (Faculty of Engineering), and Startup Grant (Tier 1) for New Academics AY2020/21 of Hong Kong Baptist University. References [1] Mubashara Akhtar, Michael Schlichtkrull, Zhijiang Guo, Oana Cocarascu, Elena Simperl, and Andreas Vlachos. 2023. Multimodal Automated Fact-Checking: A Survey. In Findings of EMNLP. [2] Paolo Boldi, Francesco Bonchi, Carlos Castillo, and Sebastiano Vigna. 2011. Query Reformulation Mining: Models, Patterns, and Applications. Information Retrieval 14, 3 (2011), 257–289. [3] Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. 2025. DEFAME: Dynamic Evidence-Based Fact-Checking with Multimodal Experts. In Proc. of ICML. [4]Yuyan Bu, Qiang Sheng, Juan Cao, Peng Qi, Danding Wang, and Jintao Li. 2024. FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspec- tive of Creative Process. In Proc. of M. [5] Grégoire Burel, Martino Mensio, Youri Peskine, Raphael Troncy, Paolo Papotti, and Harith Alani. 2024. CimpleKG: A Continuously Updated Knowledge Graph on Misinformation, Factors and Fact-Checks. In Proc. of ISWC. [6] Rui Cao, Zifeng Ding, Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2025. AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web. In Proc. of NeurIPS. [7]Nicoló Fontana, Francesco Corso, Enrico Zuccolotto, and Francesco Pierri. 2025. Evaluating Open-Source Large Language Models for Automated Fact-Checking. arXiv preprint arXiv:2503.05565 (2025). [8]Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking. Transactions of the Association for Computational Linguistics 10 (2022), 178–206. [9]Michael Hameleers, Thomas E Powell, Toni GLA Van Der Meer, and Lieke Bos. 2020. A Picture Paints a Thousand Lies? The Effects and Mechanisms of Mul- timodal Disinformation and Rebuttals Disseminated via Social Media. Political Communication 37, 2 (2020), 281–301. [10]Haorui He, Yupeng Li, Dacheng Wen, Yang Chen, Reynold Cheng, Donglong Chen, and Francis Lau. 2026. Debating Truth: Debate-Driven Claim Verification with Multiple Large Language Model Agents. In Proc. of W. [11]Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen, Reynold Cheng, and Francis Lau. 2026. Fact2Fiction: Targeted Poisoning Attack to Agentic Fact- Checking System. In Proc. of AAAI. [12]Jeff Huang and Efthimis N Efthimiadis. 2009. Analyzing and Evaluating Query Reformulation Strategies in Web Search Logs. In Proc. of CIKM. [13]Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. Live- CodeBench: Holistic and Contamination-Free Evaluation of Large Language Models for Code. In Proc. of ICLR. [14] Mohammed Abdul Khaliq, Paul Yu-Chun Chang, Mingyang Ma, Bernhard Pflugfelder, and Filip Miletić. 2024. RAGAR, Your Falsehood Radar: RAG- Augmented Reasoning for Political Fact-Checking Using Multimodal Large Lan- guage Models. In Proc. of FEVER. [15] Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2021. Toward Automated Fact-Checking: Developing an Annotation Schema and Bench- mark for Consistent Automated Claim Detection. Digital Threats: Research and Practice 2, 2 (2021), 1–16. [16]Yupeng Li, Haorui He, Jin Bai, and Dacheng Wen. 2024. MCFEND: A Multi-Source Benchmark Dataset for Chinese Fake News Detection. In Proc. of W. [17]Yifeng Luo, Yupeng Li, Dacheng Wen, and Liang Lan. 2024. Message Injection Attack on Rumor Detection under the Black-Box Evasion Setting Using Large Language Model. In Proc. of W. [18] Martino Mensio and Harith Alani. 2019. MisinfoMe: Who’s Interacting with Misinformation?. In Proc. of ISWC. [19]Qiong Nan, Juan Cao, Yongchun Zhu, Yanyan Wang, and Jintao Li. 2021. MD- FEND: Multi-Domain Fake News Detection. In Proc. of CIKM. [20]Eryn J Newman, Maryanne Garry, Daniel M Bernstein, Justin Kantner, and D Stephen Lindsay. 2012. Nonprobative Photographs (or Words) Inflate Truthi- ness. Psychonomic Bulletin & Review 19, 5 (2012), 969–974. [21]Jingjie Ning, João Coelho, Yibo Kong, Yunfan Long, Bruno Martins, João Ma- galhães, Jamie Callan, and Chenyan Xiong. 2026. Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests. In Proc. of SIGIR. [22]Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis C Petrantonakis. 2024. VERITE: A Robust Benchmark for Multimodal Misinformation Detection Accounting for Unimodal Bias. International Journal of Multimedia Information Retrieval 13, 1 (2024), 4. [23]Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, and Tat-Seng Chua. 2023. FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms. In Proc. of AAAI. [24] Dorian Quelle and Alexandre Bovet. 2024. The Perils and Promises of Fact- Checking with Large Language Models. Frontiers in Artificial Intelligence 7 (2024). [25]Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, and Anna Rohrbach. 2026. VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact- Checking. In Proc. of ACL. [26] Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Auto- mated Verification of Textual Claims (AVeriTeC) Shared Task. In Proc. of FEVER. [27]Michael Schlichtkrull, Zhijiang Guo, and Andreas Vlachos. 2024. AVeriTeC: A Dataset for Real-World Claim Verification with Evidence from the Web. In Proc. of NeurIPS. [28] Qiang Sheng, Juan Cao, Xueyao Zhang, Xirong Li, and Lei Zhong. 2021. Ar- ticle Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked Claims. In Proc. of ACL. [29]HLE Team. 2025. Humanity’s Last Exam. arXiv preprint arXiv:2501.14249 (2025). [30]James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: A Large-Scale Dataset for Fact Extraction and VERification. In Proc. of NAACL. [31]Andreas Vlachos and Sebastian Riedel. 2014. Fact-Checking: Task Definition and Dataset Construction. In Proc. of ACL. [32] Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. 2025. BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. arXiv preprint arXiv:2504.12516 (2025). [33]Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Sid- dhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al. 2025. LiveBench: A Challenging, Contamination-Limited LLM Benchmark. In Proc. of ICLR. [34]Yuzhuo Xiao, Zeyu Han, Yuhan Wang, and Huaizu Jiang. 2025. XFACTA: Con- temporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs. arXiv preprint arXiv:2508.09999 (2025). [35]Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang. 2023. End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models. In Proc. of SIGIR. [36]Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. In Proc. of ICLR. [37]Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2024. HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims. In Proc. of FEVER. [38]Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2025. Hypotheti- cal Documents or Knowledge Leakage? Rethinking LLM-Based Query Expansion. In Findings of ACL. [39]Fanrui Zhang, Dian Li, Qiang Zhang, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha, et al.2025. Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning. In Proc. of NeurIPS. [40] Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang, Bowen Li, Chengxing Xie, Junhao Wang, Maoquan Wang, Yufan Huang, Shengyu Fu, et al.2025. SWE-Bench Goes Live!. In Proc. of NeurIPS. [41] Xueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng, Lei Zhong, and Kai Shu. 2021. Mining Dual Emotion for Fake News Detection. In Proc. of W. [42]Liwen Zheng, Chaozhuo Li, Xi Zhang, Yu-Ming Shang, Feiran Huang, and Haoran Jia. 2024. Evidence Retrieval Is Almost All You Need for Fact Verification. In Findings of ACL. M ’26, November 10–14, 2026, Rio de Janeiro, Brazil.He et al., Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, and Yupeng Li 0.00.20.40.60.81.0 Threshold 0.2 0.4 0.6 0.8 1.0 Accuracy METEOR Gemma-Emb-0.3B Qwen3-Emb-0.6B (a) Threshold-sweep analysis. Contaminated Uncontaminated (b) t-SNE visualization. Figure 4: Supporting analyses for threshold selection and subset comparability. A Definitions of the Verdict Categories Following FEVER [30], the ClaimReview2025Q4 benchmark uses three verdict categories for claims: (i) Supported: The claim is supported by the arguments and evidence presented. (i) Refuted: The claim is contradicted by the arguments and evidence presented. (i) Not Enough Evidence: The presented evidence is not enough to support or refute the claim. It applies when the evidence either explicitly indicates that relevant evidence cannot be found or leaves certain aspects of the claim neither supported nor refuted. B Example Prompts This section presents example prompts. Evidence Generation. We use the following prompt to generate a fact-checking article based on the claim. Instructions Write a fact-checking article to verify the claim. Claim [CLAIM] Evidence Extraction. We use the following prompt to extract evidence from both the oracle article and the LLM-generated article. Instructions You are a precise evidence extraction expert. Claim [CLAIM] Task Extract evidence sentences from the generated text that are directly related to the core factual content of the claim. Rules 1. Only extract content that appears in the generated text. 2. Evidence must address the main factual assertion(s) made in the claim. 3. Do not infer, summarize, or add information. 4. Do not extract sentences that merely restate the claim. 5. Avoid duplication. Fact-Checking Article [FACT-CHECKING ARTICLE] Output Output strictly in JSON format: "Reason": "concise extraction reasoning", "Evidences": ["Evidence": "..."] CImpact of LLM Choice on Evidence Extraction Table 9: Average contamination scores across different ex- traction models on ClaimReview2025Q4. MetricExtractorGPT-5.2 Gemini-3.0-Pro Qwen3.5-122B-A10B METEOR GPT-4o-Mini19.15%21.63%23.16% GPT-5-Nano18.01%20.95%20.86% Gemma-Emb-0.3B GPT-4o-Mini54.12%57.82%60.38% GPT-5-Nano53.30%58.18%58.01% Qwen3-Emb-0.6B GPT-4o-Mini53.47%56.89%58.48% GPT-5-Nano53.57%57.76%56.38% We assess the robustness of our contamination detection pipeline to different LLM extractors. As shown in Table 9, contamination scores remain consistent when using either GPT-4o-Mini (default) or GPT-5-Nano for evidence extraction. D Threshold Selection To select the contamination threshold for the semantic similarity metrics, we conducted a threshold-sweep analysis on 180 randomly sampled claims (20% of ClaimReview2025Q4). Two postgraduate- level annotators with expertise in journalism and/or fact-checking independently labeled each claim as contaminated or uncontami- nated. After adjudicating disagreements, we treated the resulting labels as ground truth and evaluated threshold values over휏 ∈ [0,1]. Fig. 4a reports the classification accuracy of the contamination la- bels against these human annotations across threshold values. For the semantic metrics, both curves peak in the range of휏=0.6–0.62, where the accuracy reaches 75.6% for Gemma-Emb-0.3B and 82.2% for Qwen3-Emb-0.6B at휏=0.6. We also performed the same sweep for METEOR, whose peak occurs close to the default AVeriTeC threshold. Therefore, we use휏=0.6 for the semantic metrics and retain the default threshold of 25% for METEOR in all experiments. E Impact of Non-Contamination Factors A key concern is that performance differences between contam- inated and uncontaminated subsets may reflect topical or stylis- tic variation rather than contamination itself. Unlike prior work, both subsets here are drawn from the same ClaimReview2025Q4 sources and time frame, minimizing the likelihood of such con- founding factors. To empirically verify this, we embed each claim using Qwen3-Emb-0.6B (768 dimensions), which captures topical and stylistic features such as language complexity and ambiguity, and compute centroid-based cosine similarities: 88.2% within the contaminated group, 89.9% within the uncontaminated group, and 87.9% across groups. The high cross-group similarity indicates that the two groups are highly comparable. Fig. 4b further visualizes the two groups (downsampled to the same size of 154 claims each) with t-SNE; the points are largely mixed.