Paper deep dive
RCMN: Understanding Misleadingness in Influential Public Discourse
Peiling Yi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 2:55:24 AM
Summary
The paper introduces Reader-Centric Misleadingness Understanding (RCMN), a framework and dataset for analyzing misleadingness in influential public discourse. RCMN operationalizes misleadingness through five dimensions: misleading mechanism, likely reader interpretation, evidence-warranted interpretation, emotional arousal, and communicative intent. The authors construct an evidence-grounded dataset derived from Fact-Check Insights and benchmark five generative foundation models (Qwen3-VL-8B, DeepSeek-V4-Flash, Gemma-4-12B, GPT-5.6 Sol, Claude Fable 5) to evaluate their ability to recover reader-centric cues from lightweight representations. Findings indicate that misleadingness extends beyond fabrication to include omission and exaggeration, and while models can recover reader interpretations from limited data, identifying specific misleading mechanisms requires richer contextual grounding.
Entities (12)
Relation Signals (10)
RCMN â affiliatedwith â Kingston University London
confidence 95% ¡ Peiling Yi Affiliation: Kingston University London... we introduce Reader-Centric Misleadingness Understanding (RCMN)
RCMN â createdby â Peiling Yi
confidence 95% ¡ we introduce Reader-Centric Misleadingness Understanding (RCMN)... Peiling Yi Affiliation
RCMN â usesdatasource â Fact-Check Insights
confidence 95% ¡ The RCMN dataset is seeded from Fact Check Insights
misleadingness â associatedwith â Communicative Intent
confidence 90% ¡ misleadingness... is frequently associated with... distortive communicative intent
misleadingness â associatedwith â emotional arousal
confidence 90% ¡ misleadingness... is frequently associated with heightened emotional arousal
RCMN â evaluates â Gemma-4-12B
confidence 90% ¡ benchmark three recent open-source models (...and Gemma-4-12B)
RCMN â evaluates â GPT 5.6 Sol
confidence 90% ¡ alongside two closed-source models (GPT-5.6 Sol...)
RCMN â evaluates â Claude-fable-5
confidence 90% ¡ alongside two closed-source models (...and Claude Fable 5)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated. Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers. To address this gap, we introduce Reader-Centric Misleadingness Understanding (RCMN), a framework that operationalises misleadingness through five dimensions: misleading mechanism, likely reader interpretation, evidence-warranted interpretation, emotional arousal, and communicative intent. Based on this framework, we construct an evidence-grounded dataset of influential public discourse. Empirical findings show that misleadingness is diverse and extends well beyond fabrication, with unsupported inference, exaggeration, and omission among the prevalent mechanisms, and is frequently associated with heightened emotional arousal and distortive communicative intent. Moreover, we investigate whether lightweight claim-and-context representations retain sufficient cues for understanding reader-centric misleadingness without access to richer contextual, evidential, and multimodal information. Evaluation across five recent generative foundation models shows that reader-level interpretations can often be recovered from such limited representations, whereas identifying how misleadingness is produced remains considerably more challenging. These findings highlight the potential of lightweight representations for scalable misleadingness analysis, while reliable understanding of misleading mechanisms continues to require richer contextual and evidential grounding.
Tags
Links
- Source: https://arxiv.org/abs/2608.27358v1
- Canonical: https://arxiv.org/abs/2608.27358v1
Trouble viewing inline? Open PDF directly â
Full Text
59,717 characters extracted from source content.
Expand or collapse full text
RCMN: Understanding Misleadingness in Influential Public Discourse Peiling Yi Affiliation: School of Computer Science and Mathematics Affiliation: Faculty of Engineering, Computing and the Environment Affiliation: Kingston University London, United Kingdom Email: p.yi@kingston.ac.uk Abstract Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated. Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers. To address this gap, we introduce Reader-Centric Misleadingness Understanding (RCMN), a framework that operationalises misleadingness through five dimensions: misleading mechanism, likely reader interpretation, evidence-warranted interpretation, emotional arousal, and communicative intent. Based on this framework, we construct an evidence-grounded dataset of influential public discourse. Empirical findings show that misleadingness is diverse and extends well beyond fabrication, with unsupported inference, exaggeration, and omission among the prevalent mechanisms, and is frequently associated with heightened emotional arousal and distortive communicative intent. Moreover, we investigate whether lightweight claim-and-context representations retain sufficient cues for understanding reader-centric misleadingness without access to richer contextual, evidential, and multimodal information. Evaluation across five recent generative foundation models shows that reader-level interpretations can often be recovered from such limited representations, whereas identifying how misleadingness is produced remains considerably more challenging. These findings highlight the potential of lightweight representations for scalable misleadingness analysis, while reliable understanding of misleading mechanisms continues to require richer contextual and evidential grounding. 1 Introduction Influential public discourse refers to publicly circulated communication that has substantial visibility, prominence, or relevance to public debate and can shape how audiences understand issues of societal concernDruckman (2001). The statement plays a central role in shaping how people understand political, social, economic, and other public-interest issuesMcCombs and Shaw (1972); Entman and others (1993); Scheufele and Tewksbury (2007). Such discourse does more than communicate isolated facts: it can frame which aspects of an issue receive attention, establish causal explanations, attribute responsibility, emphasise particular risks or consequences, and influence how audiences interpret subsequent informationChong and Druckman (2010). As these messages are increasingly amplified and recirculated through news platforms and social media, their influence may extend well beyond the original communication, contributing to broader public narratives and potentially affecting attitudes, trust, and decision-makingVosoughi et al. (2018); Lazer et al. (2018). Given the societal importance of such discourse, its potential to mislead is particularly consequential Ecker et al. (2022), posing a persistent challenge to public knowledge and representative democracy. Crucially, grounded information can lead to a distorted understanding when it is selectively presented, framed, or stripped of relevant context. As illustrated in Figure 1, At the factual-verification level, the post reports numerically accurate changes in the Supplemental Poverty Measure (SPM). At the misleadingness-understanding level, the same message encourages readers to interpret these changes as evidence that Trumpâs policies reduced poverty whereas Bidenâs policies caused it to rise. This interpretation is promoted through omission and selective presentation. The comparison gives insufficient attention to the COVID-19 pandemic, temporary economic-relief programmes, the choice of poverty measure, and other relevant socioeconomic factors . Its competitive political framing also produces moderate emotional arousal by encouraging blame-and-credit attribution and serves a persuasive communicative intent rather than a purely informative one. Figure 1: From fact verification to misleadingness understanding. All claims, numerical values, and contextual explanations presented in the figure are adapted from the FactCheck.org analysis by Gore (2023). However, computationally understanding misleadingness remains particularly challenging in Influential Online Public Discourse. 1) Beyond Veracity-Centred Modelling. Existing approaches predominantly focus on verifying the factual accuracy of individual claims, while online discourse often derives its persuasive or misleading effect from how claims are framed, combined, contextualised, and presented to audiencesBudak et al. (2024); Pasquetto et al. (2024). 2) Difficult Operationalisation of Reader-Centric Misleadingness. Capturing misleadingness in online discourse requires characterising mechanisms of distortion, likely reader interpretations, emotional arousal, and communicative intent. These dimensions depend on substantial contextual judgement and domain knowledge, making reader-centric characterisation considerably more resource-intensive and complex than conventional claim-level assessment.Gabriel et al. (2022); Ni et al. (2024); Modzelewski et al. (2026). 3) Expensive Contextual and Multimodal Reasoning. Assessing misleadingness may require recovering rapidly evolving background context and reasoning across text, images, video, or audio to identify omission, emphasis, exaggeration, or recontextualisation. Retrieving and processing such distributed evidence increases computational and latency costs, limiting scalable and real-time analysis Chakraborty et al. (2023); Xie et al. (2025). To address these challenges, we propose Reader-Centric Misleadingness Understanding (RCMN), which conceptualises misleadingness in terms of the understanding a message is likely to induce in readers and comprises three complementary components. 1) RCMN Taxonomy. To move beyond veracity-centred modelling, we characterise misleadingness along five complementary dimensions: misleading mechanisms, likely reader interpretations, Evidence-Warranted Interpretation, emotional arousal, and communicative intent. 2) RCMN Dataset for Online Influential Discourse. To operationalise these reader-centric dimensions in real-world discourse, we construct RCMN from publicly circulated messages that attracted professional fact-checking attention and are often associated with public figures, organisations, and significant political or social events. By integrating original source materials, contextual information, and expert fact-checking evidence, the dataset provides structured representations of reader-centric misleadingness for fine-grained analysis and model supervision. 3) RCMN Benchmark. To reduce the cost of contextual and multimodal reasoning at inference time, we benchmark three recent open-source models (Qwen3-VL-8B, DeepSeek-V4-Flash, and Gemma-4-12B) alongside two closed-source models (GPT-5.6 Sol and Claude Fable 5), examining whether they can recover reader-centric misleadingness cues from substantially lower-cost claim-and-context representations without requiring full evidential retrieval or complete multimodal processing. The empirical findings reveal four main patterns. First, misleadingness in influential public discourse extends well beyond outright fabrication: most cases arise through unsupported inference, exaggeration, omission, miscontextualisation, or misattribution, demonstrating that factual verification alone is insufficient to capture how messages may mislead readers. Second, misleading cases are more frequently associated with high emotional arousal and distortive communicative intent, although arousal alone is not sufficient to indicate misleadingness. Third, in non-misleading cases, the interpretation encouraged by the message generally aligns closely with the evidence-warranted interpretation, supporting interpretive divergence as a useful basis for operationalising misleadingness. Finally, benchmarking five recent large language and multimodal models reveals a clear asymmetry in recoverability: likely reader interpretations and broad affective and communicative cues can often be recovered from lightweight claim-and-context representations, whereas identifying the precise mechanism of misleadingness remains substantially more difficult and dependent on richer contextual and evidential information, particularly when misleadingness arises from omitted or displaced information. Our contribution: RCMN reframes misleadingness as a structured reader-centric language-understanding problem that captures how misleadingness is produced, what interpretation a message encourages relative to the available evidence, and which affective and communicative signals shape that interpretation. By unifying these dimensions within an evidence-grounded taxonomy, dataset, and benchmark, RCMN provides a systematic framework for studying misleading communication beyond claim-level veracity and for evaluating how effectively five recent generative foundation models can recover reader-centric misleadingness cues from limited contextual information. 2 Related work In this section, we review related work from two perspectives: benchmark datasets and methodological approaches, tracing the progression from factual verification towards reader-centric understanding of misleadingness. Table 1 summarises representative datasets and their expanding task scope. Table 1: Progression of representative datasets from fact verification to reader-centric misleadingness understanding. Dataset Year Modality Primary task Labels or outputs I. Fact verification and fake-news detection LIAR Wang (2017) 2017 Text Claim-veracity classification Six-level truthfulness labels: pants-fire, false, barely true, half true, mostly true, and true. FEVER Thorne et al. (2018) 2018 Claims and textual evidence Evidence-based fact verification Supported, refuted, and not enough information. MultiFC Augenstein et al. (2019) 2019 Text, evidence, and metadata Multi-domain claim verification Heterogeneous veracity ratings collected from multiple fact-checking organisations. FakeNewsNet Shu et al. (2020) 2020 News, social context, and metadata Fake-news detection Real and fake labels with user-engagement and propagation information. Fakeddit Nakamura et al. (2020) 2020 Text, image, metadata, and comments Multimodal fake-news classification Binary, three-way, and six-way fine-grained labels. Factify / Factify 2 Mishra et al. (2022); Suryavardan et al. (2023) 2022â23 Claims, images, and evidence Multimodal fact verification Support, refute, insufficient-evidence, and related verification categories. MuMiN Nielsen and McConville (2022) 2022 Claims, images, articles, and social graphs Multilingual misinformation detection Claim-veracity labels linked to multilingual social-network data. MOCHEG Yao et al. (2023) 2023 Claims, images, and web evidence Multimodal fact-checking and explanation Veracity labels, multimodal evidence, and natural-language explanations. VeriTaS Rothermel et al. (2026) 2026 Textual and audiovisual claims Dynamic multimodal fact-checking Standardised verdict dimensions, retrieved original media, and textual justifications. I. Out-of-context and multimodal misalignment detection NewsCLIPpings Luo et al. (2021) 2021 Image and caption Out-of-context misinformation detection Pristine and semantically mismatched imageâcaption pairs. VERITE Papadopoulos et al. (2024) 2024 Image, caption, and external context Real-world out-of-context detection Truthful, miscaptioned, and out-of-context imageâcaption pairs. 5Pils-OOC Tonglet et al. (2025) 2025 Image, caption, and retrieved evidence Out-of-context detection and image contextualisation Accurate/out-of-context captions; predicted original image context and caption veracity. I. Towards misleadingness understanding M4FC Geng et al. (2025) 2025 Images, multilingual claims, and metadata Multitask multimodal fact-checking Fake image: authentic / manipulated-or-fake; location verification: whether candidate location is consistent; verdict: true / false. M-Misleading Li et al. (2026) 2026 News image, headline, and article Misleading-omission detection and correction misleading / non-misleading based specifically on misleading omission IV. Reader-centric misleadingness understanding RCMN (Ours) 2026 Claims, contextual metadata, multimodal sources/links, fact-check metadata/links, and structured evidence Reader-centric misleadingness understanding Mechanism: fabrication/alteration, miscontextualisation, omission/selective presentation, misattribution, exaggeration/quantitative distortion, unsupported inference; Arousal: low, moderate, high; Intent: informative, persuasive, distortive; Likely reader interpretation; Evidence-warranted interpretation. 2.1 Datasets and Task Evolution Early misinformation benchmarks largely operationalised the problem through veracity or factuality prediction. LIAR Wang (2017) introduced fine-grained veracity labels, FEVER Thorne et al. (2018) classified claims as supported, refuted, or lacking sufficient evidence, and MultiFC Augenstein et al. (2019) extended evidence-based verification to naturally occurring claims collected across multiple fact-checking organisations. FakeNewsNet Shu et al. (2020) broadened this setting by incorporating news content, social context, and spatiotemporal information. Later resources extended misinformation detection and fact-checking to multimodal settings. Fakeddit Nakamura et al. (2020) provides large-scale textâimage examples for multimodal fake-news detection, while Factify Mishra et al. (2022), Factify 2Suryavardan et al. (2023), MuMiN Nielsen and McConville (2022), MOCHEG Yao et al. (2023) and VeriTaS Rothermel et al. (2026) broaden the task space through multimodal fact verification, multilingual and social-context modelling, evidence use, and explanation generation. A further line of work addresses out-of-context misinformation, where authentic media become misleading through reuse or mismatched contextualisation. NewsCLIPpings Luo et al. (2021) and VERITE Papadopoulos et al. (2024) benchmark the detection of out-of-context or miscaptioned imageâtext pairs, while COVE Tonglet et al. (2025) explicitly reconstructs aspects of an imageâs original context before assessing caption veracity. These studies importantly demonstrate that misleadingness can arise without media fabrication; These studies importantly demonstrate that misleadingness can arise without media fabrication; Recent work has begun to move beyond conventional fact verification towards modelling how context, communicative intent, and selective presentation can shape misleading interpretations. M4FC Geng et al. (2025) further incorporates claimant-intent prediction, image contextualisation, location verification, and verdict prediction. M-Misleading Li et al. (2026) compares preview-supported and article-supported interpretations to identify omission-based cases in which factually compatible news previews nevertheless induce misleading interpretations. However, these approaches still capture only part of the broader phenomenon of reader-level misleadingness. 2.2 Methods Reader-centric misleadingness detection shifts the modelling objective beyond factual correctness towards understanding how content, context, communicative intent, emotion, and reader interpretation interact Yi and Zubiaga (2026). Existing methods provide several foundations for this direction. Human-centred misinformation research models cognitive, emotional, and behavioural responses to misleading content, including inferred writer intent and potential reader actions Gabriel et al. (2022), while emotion-aware approaches represent reader perception, emotion categories and intensity, and semantic emotion roles Oberländer et al. (2020). More recently, interpretation-aware methods compare the understanding induced by limited content with that supported by fuller evidence to detect omission-based interpretation drift Li et al. (2026). Explainable fact-checking methods further combine multimodal evidence retrieval, veracity prediction, and natural-language explanation generation Yao et al. (2023). However, these methodological strands remain fragmented Yi and Zubiaga (2026). 3 RCMN taxonomy In the study, Misleadingness refers to the extent to which the interpretation encouraged by a message diverges from the interpretation warranted by relevant evidence and context, regardless of whether its individual statements are literally true or false Rogers et al. (2017); Reboul (2021). Drawing on insights from psychology, linguistics, media theory, and communication studies, Yi and Zubiaga (2026) identify three core dimensions that shape misleadingness: emotional arousal, communicative intent, and context. Information may therefore mislead not only through factual inaccuracy, but also by evoking strong emotions, activating moral responses, signalling persuasive or distortive intent, or presenting otherwise accurate information without sufficient context. Building on existing studies, we characterise RCMN along five complementary dimensions. 3.1 Misleading Mechanism This dimension captures the mechanisms or manipulation techniques through which misleading content is produced van der Linden et al. (2026). ⢠Fabrication or alteration: Content is invented, synthetically generated, or technically modified in a way that changes its meaning or evidential value. ⢠Miscontextualisation: Authentic content is presented in an incorrect temporal, spatial, event, or discourse context. ⢠Omission or selective presentation: Relevant contextual, qualifying, or contradictory information is omitted, or evidence is selectively presented in a way that favours a particular interpretation. ⢠Misattribution: Content, a quotation, statement, or action is incorrectly attributed to a person, organisation, publication, or account. ⢠Exaggeration or quantitative distortion: The scale, frequency, certainty, severity, or numerical magnitude of a phenomenon is overstated or otherwise distorted. ⢠Unsupported inference: The message encourages a causal, evaluative, predictive, or generalised conclusion that is not sufficiently supported by the available evidence and context, even when the underlying statements may be factually accurate. ⢠Not misleading: No misleading mechanism is identified when the message is evaluated against the relevant evidence and context. 3.2 Likely Reader Interpretation & Evidence-Warranted Interpretation These two dimensions allow us to measure interpretive divergence: the gap between what a message encourages readers to infer and what the available evidence supports. This divergence provides a basis for assessing the degree of misleadingness. Likely Reader Interpretation captures the broader inference or conclusion that a message encourages readers to draw, including interpretations shaped by framing, selective presentation, implication, or omitted context Gabriel et al. (2022). Evidence-Warranted Interpretation captures the interpretation justified by the fact-check analysis and relevant contextual evidence Atanasova et al. (2020). 3.3 Emotional Arousal This dimension captures the intensity of the emotional response that the content is likely to evoke LĂźhring et al. (2024). ⢠Low: calm, neutral, or minimally emotional presentation. ⢠Moderate: noticeable but restrained emotional emphasis. ⢠High: strongly alerting, urgent, alarming, or emotionally charged presentation. 3.4 Communicative Intent This dimension captures the communicative goal that the message appears to serve Da et al. (2021). ⢠Informative communication: Primarily aims to provide, report, describe, or explain information to the reader. ⢠Persuasive communication: Primarily aims to influence readersâ beliefs, evaluations, or attitudes toward a person, event, issue, or position. ⢠Distortive communication: primarily steer readers toward an interpretation that is not adequately warranted by the available evidence, through selective, exaggerated, decontextualised, fabricated, or otherwise misleading presentation. 4 RCMN Datasets Figure 2: An example instance from the RCMN dataset Figure 2 illustrates the structure of the reader-centric dataset through a representative instance. The dataset is organised into five complementary groups, each serving a distinct purpose. 1) Recovered original source content records the multimodal content presented to readers, preserving the observable textual and visual information from which an interpretation may be formed. 2) Evidence & provenance records the source, provenance, and reference evidence needed to establish what is supported and to assess whether the presented content is misleading. 3) Recovered context and evidence captures relevant contextual information that is absent or not explicit in the presented content but may materially affect its interpretation. 4) Reader-centric annotations provide evidence-grounded labels across the five RCMN dimensions, capturing how misleadingness is produced, interpreted, and communicated. 5) Reader engagement signals provide complementary evidence of how audiences actually respond to and interact with the content. 4.1 Data Source The RCMN dataset is seeded from Fact Check Insights Fact-Check Insights (2026), which aggregates claims drawn from real-world information environments and investigated by independent fact-checking organisations. Rather than treating these records simply as fact-checking examples, we use them as an entry point for identifying and reconstructing potentially misleading public communication. This source is particularly suitable for our study for three reasons: 1) Fact-checkers tend to investigate claims that have already circulated publicly and attracted social or public attention, providing a useful proxy for potentially influential online discourse in which misleading communication may shape reader understanding. 2) Its structured metadata and provenance information facilitate tracing claims back to their original communications and recovering the contextual and evidential information needed to analyse how misleadingness is produced and what interpretations it may encourage. 3) Its coverage across multiple fact-checking organisations exposes RCMN to diverse topics, sources, communicative settings, and forms of misleadingness, supporting the construction of a broad reader-centric taxonomy beyond binary factual-veracity judgements. However, Fact-Check Insights exhibit substantial heterogeneity, multilinguality, and sparsity across the full schema. The records span the period from 1970 to 2025 and comprise 260,863 entries across 57 fields, There is no record contain claim source and complete information for all fields. The dataset also demonstrates considerable linguistic and organisational diversity: English accounts for only 7% of the records, while Filipino is the most represented language. Furthermore, the data include contributions from 1,007 fact-checking organisations and contain 24,869 distinct claim-rating labels, reflecting significant variation in annotation conventions across organisations and languages. Therefore, we consider Fact-Check Insights only as candidate identification and evidence recovery; its fact-check ratings are not treated as the target labels of RCMN. 4.2 Data construct Table 2: Selected core claim fields from Fact-Check Insights. â denotes fields under itemReviewed, and # denotes fields under reviewRating. Field name Purpose id Unique identifier for each claim review record. claimReviewed Original textual claim being fact-checked. datePublished Publication date of the fact-checking article. url URL of the fact-checking page. author.name Fact-checking organisation or author. *.author.name Original claim maker. *.author.@type Type of claim maker, such as person or organisation. *.name Title of the reviewed item. *.datePublished Publication date of the original claim. #.alternateName Veracity label assigned by the fact-checker. #.ratingExplanation Explanation or rationale for the rating. Fact-Check Insights provides only a limited set of structured fields for constructing RCMN, but each record links to a corresponding fact-check article that often preserves or references additional information about the original communication, its context, and the evidence used in the verification process. The fact-check article is therefore used as a recovery gateway to reconstruct the source message and retrieve the contextual and evidential information required for RCMN analysis and annotation. For each Fact-Check Insights record, eleven base fields are retained, as shown in Table 2, to support record linkage and source recovery. The corresponding fact-check article is then inspected, and its outbound links are followed to recover the original post, article, image, video, advertisement, or archived copy. When the original source is unavailable, the message is reconstructed from screenshots, quotations, embedded media, or descriptions preserved in the fact-check article, with recovery level, content representation, provenance, and confidence recorded accordingly. 4.3 Annotation The annotation pipeline is guided by four principles: evidence grounding, separation of factual veracity from misleadingness, multi-level annotation, and auditability. Accordingly, each annotation and explanation must be supported by explicit evidence and recorded in a form that enables subsequent human verification and adjudication. Importantly, annotations are not derived directly from fact-check verdicts such as False, Mostly False, or Half True; The focus is instead on the communicative process through which a message may encourage a misleading interpretation. The pipeline consists of six stages, as illustrated in Figure 3. S1) Evidence acquisition. The complete fact-check article is used as the primary evidence recovery. Relevant information about the reviewed communication, its surrounding context, omitted or distorted information, and how the message was presented is recovered from the article. S2) Source reference. The original or recovered source,such as a social-media post, image, video, quotation, or advertisement, is consulted when available. It serves as supplementary reference evidence for understanding the speaker, wording, media, platform, communication setting, and presentation, but its availability is not required for every record. S3) Evidence recovery. GPT-5.6 Sol, configured with high reasoning effort, was used to extract and structure evidence relevant to reader-centric misleadingness from the fact-check article and available source information. When an article addressed multiple claims, only evidence relevant to the reviewed claim was retained. The recovered evidence was subsequently checked against the source material to remove unsupported, fabricated, or incorrectly attributed information. The recovered evidence was subsequently manually verified against the source material to remove unsupported, fabricated, or incorrectly attributed information. S4) Initial annotation. Apply the RCMN annotation scheme to each instance based on the recovered evidence and context. The annotation considers the relationship between (i) what the communication explicitly states or shows, (i) the interpretation encouraged by its wording, framing, and presentation, and (i) the interpretation warranted by the available evidence and context. The model assigns initial labels across the RCMN dimensions and generates an evidence-warranted interpretation as a standardised reference for assessing divergence between the communicated and evidence-supported interpretations. S5) Human verification. Human annotators verify each model-proposed annotation against the recovered source content, evidence, and context. They check whether the misleading mechanism is evidence-supported, whether the likely reader interpretation follows from the presented content, whether the evidence-warranted interpretation reflects the fuller evidence, and whether the arousal and communicative-intent labels are justified by observable cues. The recovered evidence is also checked for incorrect attribution, unsupported additions, or missing material context. Any identified discrepancies are revised, while ambiguous or disputed cases are escalated for further review. S6) Adjudication. Difficult or ambiguous cases are reviewed by a human adjudicator, who resolves disagreements or uncertainty and determines the final gold annotation. Figure 3: RCMN dataset annotation pipeline Evidential grounding and annotation quality control. Sufficient evidence was available for 2,075 of the 2,216 instances (93.6%), whereas only 101 cases (4.6%) were judged to have insufficient evidence. Moreover, 2,117 instances (95.5%) contain direct two-sided evidence, allowing the annotation to be compared with evidence-supported context rather than inferred from the fact-check verdict alone. Annotation reliability was further strengthened through repeated review: 774 instances (34.9%) received a second check and 323 (14.6%) received a third check. 4.4 Empirical Findings Table 3: Statistics of the deduplicated dataset covering 2019â2025. Percentages are calculated over all 2,216 instances unless otherwise stated. Category Count Share Dataset composition Unique instances 2,216 100.0% Publication period 2019â2025 double check 774 34.9%⥠Three check 323 14.6%⥠More 1 0.05%⥠Annotation status Sufficient evidence 2,075 93.6% Article confirms claim / non-misleading 63 2.84% Insufficient evidence 101 4.6% Source actor type Politician / candidate 973 43.9%§ Political / public organisation 92 4.2%§ Media / journalist / commentator 57 2.6%§ Social-media / anonymous source 688 31.0%§ Other named public / online actor 406 18.3%§ Major discourse-domain indicators Elections / campaigns 392 17.7%Âś Economy / employment 345 15.6%Âś Health / public health 318 14.4%Âś International affairs / conflict 292 13.2%Âś Government / public policy 217 9.8%Âś Crime / public safety 183 8.3%Âś Immigration / border 146 6.6%Âś Social / cultural issues 141 6.4%Âś Climate / environment 73 3.3%Âś Fact-checking organisations PolitiFact 1,039 46.9% FactCheck.org 540 24.4% FactRakers 276 12.5% Washington Post 226 10.2% Other organisations 135 6.1% Primary misleading mechanism Unsupported inference 509 24.8%â Exaggeration / quantitative distortion 468 22.8%â Omission / selective presentation 361 17.6%â Fabrication / alteration 330 16.1%â Miscontextualisation 230 11.2%â Misattribution 154 7.5%â Emotional arousal Low 305 13.8% Moderate 634 28.6% High 1,228 55.4% Not assessable 49 2.2% Communicative intent Informative 55 2.5% Persuasive 699 31.5% Distortive 1,406 63.4% Not assessable 56 2.5% Evidence coverage Direct two-sided evidence 2,117 95.5% No structured evidence 99 4.5% Original-source evidence items 2,130 â Corrective fact-check evidence items 2,404 â â Percentages are calculated over the 2,052 instances with an assigned mechanism. ⥠Percentages are calculated over the 1,098 duplicate groups. Table 3 summarises the dataset composition, annotation outcomes, and evidence coverage. More importantly, the dataset reveals several distinctive characteristics of misleadingness. F1: Misleadingness represented in the RCMN dataset is strongly embedded in public-facing online discourse. Politicians and candidates constitute the largest source group (43.9%), while elections and campaigns (17.7%), the economy and employment (15.6%), public health (14.4%), and international affairs and conflict (13.2%) are the most prominent discourse domains. F2: Misleading mechanisms are diversitiy. The diverse distribution of misleading mechanisms shows that determining whether an individual claim is factually true is insufficient for identifying misleadingness. Most cases do not rely on outright fabrication; instead, they mislead through unsupported inference. F3: Misleadingness is strongly associated with the way information is communicated. More than half of the instances exhibit high emotional arousal, while distortive communicative intent constitutes the largest intent category. Together, these patterns suggest that misleadingness is often encouraged by how information is presented and framed. F4: Distortive Intent signal misleading All non-misleading cases, 75% were labelled as persuasive and 23.7% as informative; 0% were assigned distortive intent. By contrast, 73.1% of the misleading cases were labelled as distortive, whereas only 0.3% were informative. These results suggest that distortive, rather than persuasive, intent may serve as a strong indicator of misleadingness. F5: Reader-likely interpretation vs evidence-warranted interpretation. A semantic comparison provides further validation. Among cases labelled as non-misleading, the encouraged interpretation shows strong semantic alignment with the evidence-warranted interpretation in nearly all cases, with only one case exhibiting weaker or partial alignment. In contrast, misleading cases exhibit greater divergence between the encouraged and evidence-warranted interpretations. This pattern supports the misleadingness operationalisation: when a message is labelled as non-misleading, the interpretation it encourages is generally consistent with that warranted by the available evidence. Non-misleadingâIencouragedâIwarranted Non-misleading\; \;I_encouragedâ I_warranted (1) F6: Misleading communication is associated with higher emotional arousal. We observe a clear association between misleadingness and emotional arousal. Among messages with an identified misleading mechanism, 58.3% exhibit high arousal, compared with only 12.7% of non-misleading messages. Conversely, 28.6% of non-misleading messages exhibit low arousal, compared with 13.7% of misleading messages. The difference in arousal distributions is statistically significant (Ď2â(2)=51.80Ď^2(2)=51.80, p<0.001p<0.001), indicating that misleading communication is more frequently associated with emotionally intense presentation. However, the association is small (CramĂŠrâs V=0.157V=0.157), showing that emotional arousal is associated with, but does not, by itself, determine misleadingness. 5 Benchmarks The RCMN benchmark investigate "To what extent can claim and contextual information preserve sufficient cues for understanding multimodal misleadingness without access to the richer contextual, evidential, and multimodal information used to establish reference annotations?" 5.1 Task Definition In the study, we formally formulate: Xlimited= X_limited=\ claim,person,location,time,event, ,\,person,\,location,\,time,\,event, source setting,context, setting,\,context\, (2) Xfull= X_full=\ original source,multimodal content, source,\,multimodal content, fact-check article,fact-check results, -check article,\,fact-check results, evidence,source setting,context ,\,source setting,\,context\ (3) XfullX_full is used to establish the gold annotation YâĄXfull=Ym,Yr,Ye,Yi,Y\X_full\ \=\Y_m,Y_r,Y_e,Y_i\, where YmY_m denotes the misleading mechanism, YrY_r the likely reader interpretation, YeY_e the emotional-arousal label, and YiY_i the communicative intent. At inference time, the evaluated model does not receive the complete evidence used to construct the annotation. Instead, it is provided with a lower-cost representation XlimitedX_limited, a sample shown in Table 4. The model is therefore required to estimate Y^=fâĄ(Xlimited) Y=f(X_limited) The objective is for Y^âY Yâ Y. Table 4: Example of the input used in the RCMN benchmark. Input Claim/Message: Credit card debt is above $1 trillion for the first time ever. Person/Organisation: Jim Justice. Location: United States. Time: 14 August 2023. Event: Release of Q2 2023 New York Federal Reserve credit-card balance data. Source setting: X social-media post. Context: West Virginia Gov. Jim Justice posted the debt figure while criticising Bidenâs economic policies, connecting it to âBidenomics,â the âradical left,â and families relying on credit cards. The post linked to a CNBC story about the debt milestone. 5.2 Models & Settings To support reproducibility and comparison across model families, we evaluate three open-weight models: Qwen3-VL-8B-Instruct Bai et al. (2025), DeepSeek-V4-FlashXu et al. (2026), and Gemma-4-12BTeam et al. (2026), together with GPT-5.6 SolOpenAI (2026) and Claude-fable-5. All models are evaluated under a common zero-shot benchmark protocol using the same 2,216 instances, identical non-verification input fields, task instructions, output schema, and evaluation criteria. Model-specific inference settings are standardised where possible: the open-weight models use 4-bit NF4 quantisation, greedy decoding, and a maximum of 350 generated tokens, while the API models use their respective structured-output interfaces with medium reasoning or thinking settings. Although Qwen3-VL-8B-Instruct and gemma-4-12B are multimodal models, no image input is provided in this benchmark; all models are evaluated using the same limited claim-and-context representation. 5.3 Evaluation Classification. Misleading mechanism, emotional arousal, and communicative intent are formulated as multi-class classification tasks. We use Macro-F1 as the primary evaluation metric because it gives equal weight to each class and is therefore less affected by class imbalance. We additionally report class-wise F1 scores to examine performance across individual categories. Likely Reader Interpretation. We evaluate generated likely reader interpretations at both lexical and semantic levels. Lexical similarity is measured using ROUGE-L. Because semantically equivalent interpretations may differ substantially in wording, we additionally conduct a meaning-level semantic-equivalence evaluation. Each generated interpretation is compared with the reference annotation and classified as fully equivalent, partially equivalent, or non-equivalent. Full equivalence indicates that the same central reader takeaway is preserved; partial equivalence indicates that the main interpretation is retained but an important qualifier, scope, causal relation, or implication is missing or altered; and non-equivalence indicates a substantially different or contradictory interpretation. Empty or malformed reference interpretations are excluded from semantic evaluation. We further report a semantic adequacy score, Ssem=Nfull+0.5âNpartialNfull+Npartial+Nnon,S_sem= N_full+0.5N_partialN_full+N_partial+N_non, which assigns scores of 1, 0.5, and 0 to fully equivalent, partially equivalent, and non-equivalent outputs, respectively. 5.4 Results Table 5 presents the class-level classification results, while Table 6 reports the generation performance for likely reader interpretation. R1: Strong semantic recovery despite limited lexical overlap. The generation results reveal a clear distinction between lexical similarity and semantic recovery. Although ROUGE-L scores are relatively modest (0.228 - 0.387), 84 - 97% of generated interpretations are fully semantically equivalent to the reference, with only <=2% classified as non-equivalent. Claude Fable 5 provides the clearest example: despite achieving the lowest ROUGE-L score (0.228), it obtains the highest full-equivalence rate (97%) and a semantic adequacy score of 0.98. This indicates that models can recover the intended reader interpretation from limited claim-and-context information even when their wording differs substantially from the reference. The results also demonstrate that lexical-overlap metrics alone can underestimate the quality of reader-interpretation generation. R2:Recovery of Emotional and Communicative Cues. GPT-5.6 Sol achieves the strongest overall performance for both emotional arousal (Macro-F1 = 0.643) and communicative intent (Macro-F1 = 0.607). Its performance is particularly strong for high arousal (F1 = 0.871), persuasive intent (F1 = 0.966), and distortive intent (F1 = 0.959). These results suggest that broad emotional and communicative cues can often be recovered from limited claim-and-context information, even without conducting full evidential verification or processing the original multimodal content. R3: Challenging on Misleading-mechanism Classification. Misleading-mechanism classification remains challenging across most models, with substantially lower Macro-F1 than emotional arousal for four of the five models. Claude Fable 5 achieves the strongest overall mechanism performance (Macro-F1 = 0.520). The not misleading class is particularly difficult for most models, with F1 scores ranging from 0.031 to 0.184 for DeepSeek-v4-flash, Qwen3-VL-8B, GPT-5.6 Sol, and Gemma-4-12B, while Claude Fable 5 performs considerably better (0.393). This pattern may reflect the difficulty of making reliable judgments about misleading mechanisms from limited information, particularly when the available cues appear potentially suspicious. R4: Intermediate Recoverability of Unsupported Inference. Unsupported inference presents an interesting intermediate case. Models can sometimes recognise that a claim makes a stronger causal or interpretive conclusion than is directly supported by the supplied context, yet reliable identification still depends on knowing what conclusions the underlying evidence actually warrants. Taken together, these differences suggest that misleading mechanisms vary in their recoverability from low-cost textual cues. Mechanisms with more explicit lexical, numerical, or presentational signals appear more accessible, whereas mechanisms defined by missing, displaced, or externally verifiable information depend much more strongly on additional evidence. Table 5: Per-class and Macro-F1 performance across reader-centric misleadingness dimensions. All models are evaluated on the same 2216 shared instances. Bold indicates the best result in each row. Dimension Class GPT-5.6 Sol Qwen3-VL-8B Deepseek-v4-flash Gemma-4-12B Claude Fable 5 Mechanism Fabrication / alteration 0.521 0.147 0.587 0.237 0.634 Miscontextualisation 0.350 0.123 0.523 0.284 0.508 Omission / selective presentation 0.276 0.275 0.394 0.090 0.463 Misattribution 0.199 0.149 0.173 0.080 0.414 Exaggeration / quantitative distortion 0.586 0.341 0.480 0.173 0.647 Unsupported inference 0.567 0.113 0.102 0.433 0.578 Not misleading 0.179 0.052 0.031 0.184 0.393 Macro-F1 0.383 0.170 0.333 0.206 0.520 Arousal Low 0.288 0.597 0.342 0.487 0.472 Moderate 0.772 0.550 0.590 0.430 0.555 High 0.871 0.613 0.824 0.547 0.729 Macro-F1 0.643 0.502 0.550 0.488 0.586 Intent Informative 0.525 0.147 0.231 0.213 0.306 Persuasive 0.966 0.533 0.6139 0.433 0.501 Distortive 0.959 0.603 0.749 0.421 0.824 Macro-F1 0.607 0.320 0.405 0.263 0.408 Mean Macro-F1 across dimensions 0.544 0.331 0.429 0.319 0.505 Table 6: Evaluation of generated likely reader interpretations. ROUGE-L measures lexical overlap, while semantic-equivalence evaluation measures meaning-level agreement with the reference interpretation. Model ROUGE-L â Full Equiv. â Partial Equiv. Non-Equiv. â Semantic Adequacy â GPT-5.6 Sol 0.387 91% 8% 1% 0.95 Qwen 0.289 84% 14% 2% 0.91 DeepSeek 0.256 93% 5% 2% 0.95 Gemma 0.314 91% 7% 1% 0.95 Claude Fable 5 0.228 97% 3% <1% 0.98 5.5 Discussion Returning to the benchmark question, our results show that claim-and-context information preserves substantial but incomplete cues for understanding multimodal misleadingness. Across model families, likely reader interpretations, emotional arousal, and communicative intent can often be recovered without access to the richer evidential and multimodal information used to establish the reference annotations. In contrast, identifying the precise mechanism through which misleadingness arises remains considerably less reliable, particularly when it depends on omitted information, displaced context, or evidence external to the message. These results reveal an important distinction between contextual signals of misleadingness and context-grounded misleadingness understanding. Lightweight claim-and-context representations can preserve useful interpretive, affective, and communicative signals, but they are not sufficient for reliably determining whether and how a message misleads. Such judgements often require richer contextual and evidential information about what is absent, how the message relates to its original setting, and what interpretation is warranted by the available evidence. Claim-and-context representations are therefore best viewed as a low-cost basis for preliminary analysis and targeted retrieval, rather than a replacement for full contextual and multimodal reasoning. 6 Limitations This study has several limitations that should be considered when interpreting the findings. Limited representation of non-misleading communication. The dataset is naturally biased towards disputed or potentially misleading claims. As a result, only a small proportion of RCMN instances are labelled as not misleading. This imbalance limits the benchmarkâs ability to evaluate how reliably models distinguish misleading communication from ordinary, evidence-compatible communication, and may partly contribute to the low F1 scores observed for the not misleading class. Future work should incorporate a larger and more diverse set of non-misleading controls. AI-assisted annotation and interpretive subjectivity. The annotation process combines AI-assisted evidence recovery and initial annotation with human verification and adjudication. This design improves scalability and auditability, but reader-centric dimensions such as likely interpretation, emotional arousal, and communicative intent remain partly interpretive. Human verification reduces unsupported AI inference, but it cannot eliminate disagreement about how different readers may understand the same communication. The resulting labels should therefore be interpreted as evidence-grounded reference annotations rather than deterministic representations of every possible reader response. Reconstruction rather than complete multimodal preservation. The original post, image, video, or other media is not recoverable for every instance. In such cases, the communication is reconstructed from the fact-check article and associated provenance information, with the original source used as supplementary reference when available. 7 Conclusion This work introduces RCMN, a reader-centric framework, evidence-grounded dataset, and benchmark for understanding misleadingness in influential public discourse beyond claim-level factuality. RCMN captures how messages become misleading, what interpretations they encourage relative to available evidence, and the affective and communicative signals that shape those interpretations. Our findings show that reader-centric misleadingness contains both readily observable communicative cues and deeper evidence-dependent mechanisms. This distinction points to a promising direction for future research: developing adaptive models that use lightweight contextual signals for initial assessment while selectively retrieving richer contextual, evidential, or multimodal information when deeper verification is required. Such approaches could support more scalable and reliable understanding of how and why influential public discourse may mislead. 8 Ethics Statement RCMN is constructed from publicly circulated discourse and professionally reviewed source material. The annotation process combines AI-assisted evidence recovery and initial annotation with human verification and adjudication. Because the dataset may contain sensitive or politically salient public communication, the current study focuses on research use and does not infer private attributes or intentions beyond the communicative signals defined in the annotation framework. References Atanasova et al. (2020) P. Atanasova, J. G. Simonsen, C. Lioma, and I. Augenstein Generating fact checking explanations. In Proceedings of the 58th annual meeting of the association for computational linguistics, p. 7352â7364. Cited by: §3.2. Augenstein et al. (2019) I. Augenstein, C. Lioma, D. Wang, L. C. Lima, C. Hansen, C. Hansen, and J. G. Simonsen MultiFC: a real-world multi-domain dataset for evidence-based fact checking of claims. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), p. 4685â4697. Cited by: §2.1, Table 1. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §5.2. Budak et al. (2024) C. Budak, B. Nyhan, D. M. Rothschild, E. Thorson, and D. J. Watts Misunderstanding the harms of online misinformation. Nature 630 (8015), p. 45â53. Cited by: §1. Chakraborty et al. (2023) M. Chakraborty, K. Pahwa, A. Rani, S. Chatterjee, D. Dalal, H. Dave, P. Gurumurthy, A. Mahor, S. Mukherjee, A. Pakala, et al. Factify3m: a benchmark for multimodal fact verification with explainability through 5w question-answering. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 15282â15322. Cited by: §1. Chong and Druckman (2010) D. Chong and J. N. Druckman Dynamic public opinion: communication effects over time. American Political Science Review 104 (4), p. 663â680. Cited by: §1. Da et al. (2021) J. Da, M. Forbes, R. Zellers, A. Zheng, J. D. Hwang, A. Bosselut, and Y. Choi Edited media understanding frames: reasoning about the intent and implications of visual misinformation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 2026â2039. Cited by: §3.4. Druckman (2001) J. N. Druckman On the limits of framing effects: who can frame?. The journal of politics 63 (4), p. 1041â1066. Cited by: §1. Ecker et al. (2022) U. K. Ecker, S. Lewandowsky, J. Cook, P. Schmid, L. K. Fazio, N. Brashier, P. Kendeou, E. K. Vraga, and M. A. Amazeen The psychological drivers of misinformation belief and its resistance to correction. Nature reviews psychology 1 (1), p. 13â29. Cited by: §1. Entman et al. (1993) R. M. Entman et al. Framing: towards clarification of a fractured paradigm. McQuailâs reader in mass communication theory 390, p. 397. Cited by: §1. Fact-Check Insights (2026) Fact-Check InsightsGuide to the data(Website) External Links: Link Cited by: §4.1. Gabriel et al. (2022) S. Gabriel, S. Hallinan, M. Sap, P. Nguyen, F. Roesner, E. Choi, and Y. Choi Misinfo reaction frames: reasoning about readersâ reactions to news headlines. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3108â3127. Cited by: §1, §2.2, §3.2. Geng et al. (2025) J. Geng, J. Tonglet, and I. Gurevych M4FC: a multimodal, multilingual, multicultural, multitask real-world fact-checking dataset. arXiv preprint arXiv:2510.23508. Cited by: §2.1, Table 1. Gore (2023) D. Gore Trumpâs misleading poverty rate comparison. Note: FactCheck.org External Links: Link Cited by: Figure 1. Lazer et al. (2018) D. M. Lazer, M. A. Baum, Y. Benkler, A. J. Berinsky, K. M. Greenhill, F. Menczer, M. J. Metzger, B. Nyhan, G. Pennycook, D. Rothschild, et al. The science of fake news. Science 359 (6380), p. 1094â1096. Cited by: §1. Li et al. (2026) F. Li, J. Wu, T. Fu, D. Li, H. Wan, W. Zhou, and M. Kan Whatâs left unsaid? detecting and correcting misleading omissions in multimodal news previews. arXiv preprint arXiv:2601.05563. Cited by: §2.1, §2.2, Table 1. LĂźhring et al. (2024) J. LĂźhring, A. Shetty, C. Koschmieder, D. Garcia, A. Waldherr, and H. Metzler Emotions in misinformation studies: distinguishing affective state from emotional response and misinformation recognition from acceptance. Cognitive research: principles and implications 9 (1), p. 82. Cited by: §3.3. Luo et al. (2021) G. Luo, T. Darrell, and A. Rohrbach Newsclippings: automatic generation of out-of-context multimodal media. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 6801â6817. Cited by: §2.1, Table 1. McCombs and Shaw (1972) M. E. McCombs and D. L. Shaw The agenda-setting function of mass media. Public opinion quarterly 36 (2), p. 176â187. Cited by: §1. Mishra et al. (2022) S. Mishra, S. Suryavardan, A. Bhaskar, P. Chopra, A. N. Reganti, P. Patwa, A. Das, T. Chakraborty, A. P. Sheth, A. Ekbal, et al. FACTIFY: a multi-modal fact verification dataset.. In DE-FACTIFY@ AAAI, p. np. Cited by: §2.1, Table 1. Modzelewski et al. (2026) A. Modzelewski, W. Sosnowski, E. Papadopulos, E. Sartori, T. Labruna, G. Da San Martino, and A. Wierzbicki MALicious intent dataset and inoculating llms for enhanced disinformation detection. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3125â3148. Cited by: §1. Nakamura et al. (2020) K. Nakamura, S. Levy, and W. Y. Wang Fakeddit: a new multimodal benchmark dataset for fine-grained fake news detection. In Proceedings of the twelfth language resources and evaluation conference, p. 6149â6157. Cited by: §2.1, Table 1. Ni et al. (2024) J. Ni, M. Shi, D. Stammbach, M. Sachan, E. Ash, and M. Leippold Afacta: assisting the annotation of factual claim detection with reliable llm annotators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1890â1912. Cited by: §1. Nielsen and McConville (2022) D. S. Nielsen and R. McConville Mumin: a large-scale multilingual multimodal fact-checked misinformation social network dataset. In Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval, p. 3141â3153. Cited by: §2.1, Table 1. Oberländer et al. (2020) L. A. M. Oberländer, E. Kim, and R. Klinger GoodNewsEveryone: a corpus of news headlines annotated with emotions, semantic roles, and reader perception. In Proceedings of the Twelfth Language Resources and Evaluation Conference, p. 1554â1566. Cited by: §2.2. OpenAI (2026) OpenAI GPT-5.6 sol model. Note: OpenAI API DocumentationAccessed August 2026 Cited by: §5.2. Papadopoulos et al. (2024) S. Papadopoulos, C. Koutlis, S. Papadopoulos, and P. C. Petrantonakis Verite: a robust benchmark for multimodal misinformation detection accounting for unimodal bias. International Journal of Multimedia Information Retrieval 13 (1), p. 4. Cited by: §2.1, Table 1. Pasquetto et al. (2024) I. V. Pasquetto, G. Lim, and S. Bradshaw Misinformed about misinformation: on the polarizing discourse on misinformation and its consequences for the field. Harvard Kennedy School Misinformation Review 5 (5), p. 1â8. Cited by: §1. Reboul (2021) A. Reboul Truthfully misleading: truth, informativity, and manipulation in linguistic communication. Frontiers in Communication 6, p. 646820. Cited by: §3. Rogers et al. (2017) T. Rogers, R. Zeckhauser, F. Gino, M. I. Norton, and M. E. Schweitzer Artful paltering: the risks and rewards of using truthful statements to mislead others.. Journal of personality and social psychology 112 (3), p. 456. Cited by: §3. Rothermel et al. (2026) M. Rothermel, M. Kornmann, M. Rohrbach, and A. Rohrbach VeriTaS: the first dynamic benchmark for multimodal automated fact-checking. arXiv preprint arXiv:2601.08611. Cited by: §2.1, Table 1. Scheufele and Tewksbury (2007) D. A. Scheufele and D. Tewksbury Framing, agenda setting, and priming: the evolution of three media effects models. Journal of communication 57 (1), p. 9â20. Cited by: §1. Shu et al. (2020) K. Shu, D. Mahudeswaran, S. Wang, D. Lee, and H. Liu Fakenewsnet: a data repository with news content, social context, and spatiotemporal information for studying fake news on social media. Big data 8 (3), p. 171â188. Cited by: §2.1, Table 1. Suryavardan et al. (2023) S. Suryavardan, S. Mishra, P. Patwa, M. Chakraborty, A. Rani, A. Reganti, A. Chadha, A. Das, A. Sheth, M. Chinnakotla, et al. Factify 2: a multimodal fake news and satire news dataset. arXiv preprint arXiv:2304.03897. Cited by: §2.1, Table 1. Team et al. (2026) G. Team, S. E. Abd, V. Aggarwal, R. Algayres, A. Andreev, O. Bachem, I. Ballantyne, C. Brick, V. CÄrbune, M. Casbon, et al. Gemma 4 technical report. arXiv preprint arXiv:2607.02770. Cited by: §5.2. Thorne et al. (2018) J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal FEVER: a large-scale dataset for fact extraction and verification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), p. 809â819. Cited by: §2.1, Table 1. Tonglet et al. (2025) J. Tonglet, G. Thiem, and I. Gurevych COVE: COntext and VEracity prediction for out-of-context images. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 2029â2049. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2.1, Table 1. van der Linden et al. (2026) S. van der Linden, D. Louison-Lavoy, N. Blazer, N. S. Noble, and J. Roozenbeek Prebunking misinformation techniques in social media feeds: results from an instagram field study. Harvard Kennedy School Misinformation Review. Cited by: §3.1. Vosoughi et al. (2018) S. Vosoughi, D. Roy, and S. Aral The spread of true and false news online. science 359 (6380), p. 1146â1151. Cited by: §1. Wang (2017) W. Y. Wang âLiar, liar pants on fireâ: a new benchmark dataset for fake news detection. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 2: short papers), p. 422â426. Cited by: §2.1, Table 1. Xie et al. (2025) Z. Xie, R. Xing, Y. Wang, J. Geng, H. Iqbal, D. Sahnan, I. Gurevych, and P. Nakov FIRE: fact-checking with iterative retrieval and verification. In Findings of the Association for Computational Linguistics: NAACL 2025, p. 2901â2914. Cited by: §1. Xu et al. (2026) A. Xu, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Ling, et al. Deepseek-v4: towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348. Cited by: §5.2. Yao et al. (2023) B. M. Yao, A. Shah, L. Sun, J. Cho, and L. Huang End-to-end multimodal fact-checking and explanation generation: a challenging dataset and models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, p. 2733â2743. Cited by: §2.1, §2.2, Table 1. Yi and Zubiaga (2026) P. Yi and A. Zubiaga From fact verification to understanding misleadingness: a survey and roadmap on reader-centric multimodal misinformation detection. Cited by: §2.2, §2.2, §3.