Paper deep dive
The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation
Zacharias Chrysidis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 11:54:49 AM
Summary
The paper introduces CONVEX, a large-scale dataset of over 150,000 multimodal misinformation posts (images and videos) from X's Community Notes. The study analyzes the evolution, virality, and detectability of three categories: miscaptioned, edited, and AI-generated content. Key findings reveal that AI-generated content achieves disproportionate virality through passive engagement (likes) rather than active discourse (replies) and reaches community consensus more rapidly once flagged, despite slower initial reporting. Additionally, the research demonstrates a decline in the efficacy of specialized detectors and Vision-Language Models (VLMs) over time as generative models evolve.
Entities (8)
Relation Signals (4)
CONVEX â derivedfrom â Community Notes
confidence 100% ¡ we use the publicly available Community Notes corpus to construct CONVEX
DALL¡E 3 â isatypeof â Generative AI
confidence 100% ¡ In the image subset, volume increases correspond with the public release of DALL¡E 3
Generative AI Models â impacts â Detection Efficacy
confidence 95% ¡ evaluation of specialized detectors and vision-language models reveals a consistent decline in performance over time as generative models evolve
AI-generated content â hashighviralityvia â passive engagement
confidence 90% ¡ AI-generated content achieves disproportionate virality, its spread is driven primarily by passive engagement
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As generative AI advances, the distinction between authentic and synthetic media is increasingly blurred, challenging the integrity of online information. In this study, we present CONVEX, a large-scale dataset of multimodal misinformation involving miscaptioned, edited, and AI-generated visual content, comprising over 150K multimodal posts with associated notes and engagement metrics from X's Community Notes. We analyze how multimodal misinformation evolves in terms of virality, engagement, and consensus dynamics, with a focus on synthetic media. Our results show that while AI-generated content achieves disproportionate virality, its spread is driven primarily by passive engagement rather than active discourse. Despite slower initial reporting, AI-generated content reaches community consensus more quickly once flagged. Moreover, our evaluation of specialized detectors and vision-language models reveals a consistent decline in performance over time in distinguishing synthetic from authentic images as generative models evolve. These findings highlight the need for continuous monitoring and adaptive strategies in the rapidly evolving digital information environment.
Tags
Links
- Source: https://arxiv.org/abs/2604.15372v1
- Canonical: https://arxiv.org/abs/2604.15372v1
Trouble viewing inline? Open PDF directly â
Full Text
45,233 characters extracted from source content.
Expand or collapse full text
The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation * Zacharias Chrysidis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos Centre for Research and Technology Hellas, Greece zchrysid, stefpapad, papadop@iti.gr Abstract As generative AI advances, the distinction between au- thentic and synthetic media is increasingly blurred, chal- lenging the integrity of online information. In this study, we present CONVEX, a large-scale dataset of multimodal misinformation involving miscaptioned, edited, and AI- generated visual content, comprising over 150K multimodal posts with associated notes and engagement metrics from Xâs Community Notes. We analyze how multimodal misin- formation evolves in terms of virality, engagement, and con- sensus dynamics, with a focus on synthetic media. Our re- sults show that while AI-generated content achieves dispro- portionate virality, its spread is driven primarily by passive engagement rather than active discourse. Despite slower initial reporting, AI-generated content reaches community consensus more quickly once flagged. Moreover, our eval- uation of specialized detectors and vision-language models reveals a consistent decline in performance over time in dis- tinguishing synthetic from authentic images as generative models evolve. These findings highlight the need for con- tinuous monitoring and adaptive strategies in the rapidly evolving digital information environment. 1. Introduction Amid the rapid evolution of media and communication technologies, the scale and velocity at which misinforma- tion spreads have made it a major societal concern with po- tential negative impacts on democratic processes [4], vul- nerable groups [13], and public health [11], among other domains. While early research largely focused on textual claims, misleading information increasingly incorporates multimodal content [1], including images and videos, which tend to be perceived as more persuasive when used to sup- port misleading claims or narratives [37]. Generative AI amplifies these concerns by enabling the creation of realistic * Accepted at the 3rd Workshop on New Trends in AI-Generated Media and Security (AIMS), held in conjunction with CVPR 2026. Black River, Jamaica, now. #Melissa #Jamaica USER COMMUNITY NOTE This image is AI-generated. It was posted on Joemar Sombero's Facebook page. His intro reads "Real disasters ⢠AI visuals â˘": facebook.com/weekendtraveler195 This image does not match the location of the hospital: mapcarta.com/W374514947/Map EXTERNAL SOURCES Figure 1. Example of an AI-generated image post on X, with the Community Note and external sources used to verify it. synthetic images, videos, and text at scale, raising questions about how AI-generated media may transform the produc- tion and spread of misinformation online [2, 12]. Addressing misinformation at scale remains challenging for social media platforms. Mitigation approaches rely on professional fact-checkers or automated detection systems [3]. While expert verification can provide high-quality as- sessments, it struggles to scale to the massive volume of online content [38]. In turn, automated methods remain constrained by biases in training data, limited generaliza- tion and trustworthiness, and the rapidly evolving strategies used to produce misleading information [10, 15]. Conse- quently, platforms are increasingly exploring community- based moderation, where users collaboratively contribute context and evaluate potentially misleading content [7]. One prominent example is Xâs Community Notes, a crowdsourced fact-checking system introduced in 2021 that allows contributors to write contextual notes on posts they consider misleading, while other users rate their helpfulness [39]. Figure 1 illustrates a community note, in which an AI- generated image is challenged through comparison with ex- ternal satellite imagery, revealing inconsistencies between arXiv:2604.15372v1 [cs.CR] 15 Apr 2026 the depicted building and the stated location. Since its in- troduction, Community Notes has expanded globally and produced millions of notes, providing a large-scale resource of community-annotated content [31]. Prior research has utilized this data to analyze consensus dynamics [38], the timeliness of community responses [35], and the longitudi- nal impact of the system on user engagement [8]. However, the evolving landscape of multimodal misinformation and AI-generated media remains unexplored. In this work, we leverage Community Notes to construct a large-scale dataset of multimodal misinformation â cat- egorized into miscaptioned, edited, and AI-generated im- ages and videos â and analyze their prevalence, engagement dynamics, and community-driven oversight. Specifically, we introduce CONVEX, the âCommunity Notes for Visual Misinformation on Xâ dataset, a collection of over 150,000 note-post pairs with crowdsourced annotations, associated media, and engagement statistics. Notably, our data collec- tion and annotation pipeline is designed to integrate future releases of Community Notes for continuous monitoring. We conduct a longitudinal data analysis of how different misinformation categories evolve in terms of virality, en- gagement, and consensus dynamics. Our analysis indicates that AI-generated content volume correlates with the evo- lution and availability of generative models. We find that generated media achieves disproportionate virality primar- ily through passive engagement (e.g., favorites), contrast- ing the more discursive patterns (e.g., replies) of miscap- tioned content. Furthermore, while AI-generated media is initially slower to be reported, it currently exhibits signifi- cantly higher community consensus once identified. Finally, motivated by the use of AI tools by Community Notes contributors, we construct a real-world benchmark of authentic versus AI-generated images. Our evaluation of specialized Synthetic Image Detectors (SIDs) and Vision- Language Models (VLMs) reveals consistent and signifi- cant decline in detection efficacy over time as generative models continue to evolve. We make the codebase, dataset, and appendix publicly available for reproducibility 1 . 2. Related Work 2.1. Multimodal and AI-Generated Misinformation Research on misinformation has increasingly recognized the importance of multimodal content, including images and videos that shape how users perceive and engage with claims online [1, 19]. Recent studies show that manipulated or misleading visuals can be especially persuasive when presented as evidence of real-world events [16, 37]. More- over, multimodal misinformation has become a substantial and evolving phenomenon, with empirical analyses indicat- ing that contextual manipulations remain the most common 1 https://github.com/zachos99/convex-dataset form of visual misinformation even as AI-generated con- tent has risen rapidly in recent years [12, 36]. At the same time, the emergence of generative AI systems has increased concerns about the scalable production of realistic synthetic media and its implications for fact-checking and online in- formation integrity [2, 30]. However, prior work primarily relies on synthetic data [27, 32, 33] or fact-checking sites [34, 40, 42], which lack insight into how multimodal misinformation circulates âin the wildâ. Moreover, existing social media datasets [5, 6, 18] are limited to a small number of events. This lack of scale and diversity limits their ability to generalize to the rapidly evolving misinformation landscape. 2.2. Community-based Fact-Checking Community-based fact-checking systems leverage collec- tive intelligence to complement expert content moderation [3, 29]. A leading example is Xâs Community Notes, which has provided a valuable resource for studying real-world misinformation dynamics. Early studies describe the transi- tion to Community Notes and the systemâs evolving design and infrastructure [31]. During the pilot period, internal tests report that users exposed to notes were 25â34% less likely to like or repost misleading tweets [39]. However, an analysis after the platform-wide rollout find no significant reduction in engagement, likely because notes appear too late in the postâs lifecycle [8]. Other research highlights lim- its to scalability and consensus formation: note production is highly concentrated (the top 10% of contributors produce about 58% of notes), while only around 11.5% of proposed notes ultimately reach publication consensus [35]. Another line of work examine political asymmetries in note language and behavior [24], the relationship be- tween community-based moderation and professional fact- checking [7], and patterns of source credibility and bias in the outlets cited within notes [20]. Recently, several studies have explored how LLMs could assist the production, rank- ing, or summarization of Community Notes [9, 25]. How- ever, most research on Community Notes focuses on textual misinformation, while the evolution of multimodal misin- formation over time remains largely unexplored. 2.3. Detection of AI-Generated Media Detecting AI-generated images has become an active re- search area as generative models rapidly improve in re- alism.Recent surveys describe a broad landscape of detection methods, spanning artifact-based, frequency- domain, representation-based, and multimodal reasoning approaches [26, 28]. These works also highlight the dif- ficulty of achieving robust detection across diverse gener- ative models and real-world distribution shifts [21]. Re- cent specialized detectors, such as SPAI [22], RINE [23] and B-Free [14], aim to improve robustness through spec- 07-202310-202301-202404-202407-202410-202401-202504-202507-202510-2025 0 100 200 300 400 DALL¡E 3 Midjourney v6 DALL¡E 3 free-tier GPT-4o Images Gemini 2.5 Flash Image Gemini 3 Pro Image Miscaptioned Edited AI-generated Miscaptioned Edited AI-generated (a) Image Set 10-202301-202404-202407-202410-202401-202504-202507-202510-2025 0 100 200 300 400 500 600 700 800 Veo Veo 2 Veo 3 Sora Sora 2 Veo 3.1 Miscaptioned Edited AI-generated Miscaptioned Edited AI-generated (b) Video Set Figure 2. Monthly Community Notes volume for both modalities and releases of popular generative AI tools. tral learning, intermediate visual representations, and bias- suppressing training. In parallel, VLMs have emerged as general-purpose systems for visual understanding and rea- soning [41], with recent work exploring their potential as reasoning-based tools for identifying synthetic media [17]. Despite these advancements, previous studies largely rely on curated synthetic-image benchmarks rather than âin the wildâ images. Our work addresses this gap by evaluat- ing detection models on images drawn from Community- annotated posts on X. 3. Dataset Construction 3.1. Data Collection To construct CONVEX for longitudinal analysis of multi- modal misinformation, we use the publicly available Com- munity Notes corpus 2 . We collect all notes and metadata from January 2021 to January 2026 and retain entries la- beled as âmisinformed or potentially misleadingâ, resulting in 1,806,168 notes. To isolate multimodal misinformation, we identify notes that reference images or videos. We employ keyword- based filtering (e.g., âphotoâ, âgraphicâ for images; âclipâ, âfootageâ for videos; see Appendix 10 for full list) to iden- tify multimodal entries. We then retrieve the associated X posts using the twikit library 3 and collect the cor- responding media files, author metadata, and engagement metrics (favorites, retweets, views).Inaccessible posts (deleted or suspended) are excluded.The final corpus comprises 66,135 image-related and 86,131 video-related noteâpost pairs. 3.2. Data Annotation We classify note-post pairs into three misinformation cat- egories: Miscaptioned (authentic media presented in mis- 2 https://x.com/i/communitynotes/download-data 3 https://twikit.readthedocs.io leading contexts), Edited (digitally altered media), and AI- generated (synthetic media). To annotate the dataset at scale, we adopt a hybrid weakly supervised approach: ⢠Keyword-based: For each category, we define a key- word list (e.g., âout-of-contextâ, âreused photoâ for mis- captioned; âphotoshoppedâ, âdigitally alteredâ for edited; âAI-generatedâ, âsynthetic imageâ for AI-generated) and apply them to the Community Note. ⢠VLM-based: We use Gemma 3 in a zero-shot setting to classify each entry using the post text, associated media, and Community Note text. ⢠Classification: When keyword-based and VLM-based labels disagree, we rerun the VLM with the keyword- derived label provided as additional context. Final labels are assigned via majority voting over keyword-based la- bels and the two VLM predictions. See Appendix 11 for the full methodology and prompts. In the image subset (66K entries), 60.2% are classified as miscaptioned, 23.3% as edited and 16.3% as AI-generated. In the video subset (86K entries), 75.6% are classified as miscaptioned, 12.8% as AI-generated, and 9.3% as edited. Few ambiguous cases were excluded from the analysis. Al- though this weakly supervised procedure may introduce la- beling noise, it enables consistent annotation at scale which is necessary for longitudinal analysis. To assess label quality, we manually examine a strat- ified random sample of 600 note-post pairs (100 per category-modality combination), with weights reflecting the 2023â2025 distribution. Agreement with manual labels was 91%, 95%, and 92% for AI-generated, edited, and mis- captioned images, versus 87%, 88%, and 91% for videos. 3.3. Continuous Monitoring To ensure continuous tracking and long-term analysis, our data collection and annotation pipeline is designed to oper- ate on successive Community Notes releases. This enables monitoring shifts in the multimodal misinformation land- 05-202309-202301-202405-202409-202401-202505-202509-2025 0.5 1.0 1.5 2.0 2.5 3.0 Virality Share Miscaptioned Edited AI-generated (a) Image Set 09-202301-202405-202409-202401-202505-202509-2025 0.5 1.0 1.5 2.0 2.5 3.0 Virality Share Miscaptioned Edited AI-generated (b) Video Set Figure 3. Virality Share across both modalities. scape over time rather than relying on static snapshots. 4. Evolution of Multimodal Misinformation We examine monthly Community Notes volume for image- and video-related entries beginning in May 2023 and September 2023, respectively, following the platformâs of- ficial media annotation rollout. Figure 2 shows the evolu- tion of notes categorized as Miscaptioned, Edited and AI- generated across both modalities. Miscaptioned content re- mains the dominant category for both images and videos, reflecting the low barrier to entry for repurposing authentic visual media. While edited content volume is relatively sta- ble, AI-generated visual content displays a steady upward trajectory, with accelerated growth in recent periods â most notably within the video subset. These surges often align with major generative model releases. In the image subset, volume increases correspond with the public release of DALL-E 3 (August 2024) and the integration of GPT-4o image generation into free tiers (April 2025), or more recently, Gemini 3 Pro Image (âNano Banana Proâ). Similarly, video notes spiked following the public releases of Sora 2 and Veo 3.1. These patterns sug- gest that expanded access to high-fidelity generative tools correlates with the volume of AI-generated visual content. 5. Attention Dynamics 5.1. Virality Share To assess how different types of misinformation capture public attention, we analyze engagement metrics, includ- ing retweets (rt), favorites (f ), and replies (rp). We define an aggregate interaction score: A = rt + f + rp(1) To account for the heavy-tailed nature of social media inter- actions, we define a post as viral if its score A exceeds the 99th percentile of the monthly distribution for its modality. To compare the viral potential of each misinformation category while accounting for its baseline prevalence, we define a Virality Share metric: V (c,m) = P(c| viral,m) P(c| m) (2) where c denotes the misinformation category and m the month. This metric measures whether a given category is over- or under-represented among viral posts relative to its overall frequency. V â 1 indicate proportional representa- tion, V > 1 indicate over-representation, and V < 1 indi- cate under-representation. AI-generated content exhibits the highest average V in both modalities, reaching 1.56 in the image subset and 1.25 in the video subset. Edited content remains close to propor- tional representation (1.02 image, 1.13 video), while mis- captioned posts are slightly under-represented among vi- ral posts (0.91 image, 0.97 video). These results indicate that AI-generated content is disproportionately represented among highly engaging posts. As shown in Figure 3, AI-generated content is con- sistently over-represented among viral posts across most months in both modalities, with pronounced spikes in mid- 2024 and mid-2025. While month-to-month volatility is expected given the small number of posts in the top 1%, the overall pattern indicates persistent over-representation rather than isolated outliers. In contrast, miscaptioned con- tent remains close to proportional representation (V â 1), and edited content fluctuates around the baseline. 5.2. Engagement Dynamics Beyond virality, we examine user engagement; the ratio of active (retweets and replies) versus passive (favorites) en- 05-202309-202301-202405-202409-202401-202505-202509-2025 â0.3 â0.2 â0.1 0.0 0.1 0.2 Engagement Index Miscaptioned Edited AI-generated (a) Image Set 09-202301-202405-202409-202401-202505-202509-2025 â0.3 â0.2 â0.1 0.0 0.1 0.2 0.3 Engagement Index Miscaptioned Edited AI-generated (b) Video Set Figure 4. Monthly Engagement Index for both modalities. gagement. To smooth variance in the highly skewed en- gagement metrics on X, we apply a log transformation to each signal s: l = log(1 + s)(3) We then standardize these values within each month using a z-score: z = lâ Îź m Ď m (4) where Îź m and Ď m denote the monthly mean and standard deviation. This normalization places all engagement signals on a comparable, zero-centered scale. Using the standard- ized signals, we define an Engagement index (E) as: E = z rt + z rp â 2z f (5) The factor of two assigns equal total weight to active and passive interactions. Higher values indicate posts that trig- ger discursive participationâwhere users are more likely to retweet a post or reply than to simply âlikeâ, while lower values reflect passive approval. Figure 4 illustrates monthly median E values across mis- information categories. In both modalities miscaptioned content consistently shows higher median E values than edited and AI-generated content. In contrast, AI-generated posts exhibit lower values, reflecting attention driven pri- marily by passive engagement rather than active discussion. These results suggest that miscaptioned content frequently sparks public discussion or controversy, while AI-generated content tends to propagate through passive engagement. 6. Consensus Dynamics Community Notes begin in a âNeeds More Ratingsâ (NMR) state and transition to either âHelpfulâ or âNot Helpfulâ once sufficient cross-partisan agreement is reached, indicat- ing community consensus. Only notes that exit NMR are displayed under the flagged post. Using the âNote Status Historyâ file, we compute consensus-related metrics at both the tweet and note level. At the note level, we compute the (1) Notes / Tweets: the average number of notes per tweet, and (2) Helpful (%): the share of notes rated Helpful. At the tweet level, we measure (3) Consensus Probability: the fraction of tweets receiving at least one note transition from NMR to consensus, (4) First Note: the median time (in hours) between tweet creation and its first associated note, and (5) Notes to Consensus: the average number of notes required before the first consensus note appears. Table 1 summarizes the results. Across both modali- ties, AI-generated content receives fewer notes per tweet and takes longer to receive its first note (11.3h vs 9.1h in video). However, once annotated, AI-related posts are more likely to reach consensus. In the video subset, AI-generated tweets exhibit a consensus probability of 0.362 compared to 0.295 and 0.299 for other categories, and require fewer notes on average before consensus (1.26 vs 1.40). More- over, AI-generated content shows the highest Helpful share (81.4% in video; 69.4% in image). These patterns indi- cate a distinct annotation dynamic for AI-generated visual content: it attracts fewer and slower annotations, yet once examined, raters converge more quickly and with higher agreement. In contrast, miscaptioned content receives more immediate attention but exhibits lower consensus rates and a smaller share of helpful outcomes. One explanation is that current generative outputs con- tain identifiable structural artifacts that facilitate objective verification once attention is drawn. Furthermore, the fre- quent citation of external verification tools, a phenomenon which we analyze in Section 7, provides a standardized evi- dence base that accelerates rater agreement. Combined with the engagement results in Section 5, this suggests that while synthetic misinformation exhibits higher virality through passive engagement, it undergoes swift collective correction once subjected to crowdsourced scrutiny. Table 1. Consensus metrics across misinformation categories for image and video subsets. ImageVideo Type Notes / Tweet Helpful (%) Consensus Probability First Note (h) Notes to Consensus Notes / Tweet Helpful (%) Consensus Probability First Note (h) Notes to Consensus Miscaptioned1.8966.80.2849.411.471.8571.10.2959.061.41 Edited1.8570.50.3048.471.421.8371.20.2999.761.40 AI-Generated1.7869.40.30511.691.411.6181.40.36211.301.26 Table 2. AI-related mentions in Community Notes. Values repre- sent the percentage of notes containing general references to AI or to specific a model. ImageVideo TypeModelGeneralAnyModelGeneralAny Miscaptioned0.941.582.480.981.722.68 Edited1.351.472.800.921.632.53 AI-generated9.4756.7560.5812.2547.0753.04 7. AI References in Community Notes To quantify how AI is used and referenced in modera- tion discussions, we examine general references to AI gen- eration (e.g., âAI generatedâ, âcreated by artificial intelli- genceâ), which attribute content creation to AI systems, and mentions of specific AI models or tools (e.g., ChatGPT, Grok, Gemini, Midjourney, Sora). See Appendix 12 for fur- ther methodological details. AI-related references are far more prevalent in Commu- nity Notes than in user posts; appearing inâ 11% of notes compared to less than 1% of posts. As shown in Table 2, general references to AI concentrate strongly in notes ad- dressing AI-generated misinformation. In the image subset, 60.6% of notes attached to AI-generated content contain at least one AI-related reference (either specific AI model or general references to AI) compared to 2â3% for miscap- tioned and edited content. Similar patterns appear in the video subset, where over half of AI-related notes reference AI explicitly. However, a substantial fraction of notes ad- dressing synthetic media, â40% in the image subset, do not contain explicit AI-generation phrases. This suggests that identifying synthetic media often relies on contextual reasoning or visual forensics beyond simple keyword cues, which supports the adopted hybrid labeling approach de- scribed in Section 3. Explicit references to specific AI systems are less fre- quent than general references to AI but still present across the dataset. The most frequently mentioned systems are general-purpose assistants such as Grok (46.8% in im- age set; 39.4% in video set), ChatGPT (16% in image set; 12.1% in video set) and Gemini (14.5% in image set; 4.6% in video set), which appear across both modal- ities. Modality-specific tools are also referenced: in the image subset, mentions often include Midjourney (11.6%), while in the video subset references are dominated by Sora (37.2%) and Veo (4.1%). Notably, a substantial fraction of model mentions appear inside URLs rather than descrip- tive text. For example, notes often include links to Grok- generated replies embedded on X or shared conversations with systems such as Grok or ChatGPT. This suggests that community contributors are increasingly citing AI-assisted analyses as supporting evidence to justify the synthetic na- ture of a post. 8. Evaluation of Detection Systems As AI systems are increasingly cited within the explanatory context of Community Notes, we examine how reliably cur- rent automated methods distinguish AI-generated from au- thentic images in the wild. 8.1. Experimental Setup From CONVEX, we select all images assigned to the AI- generated class from 2023 to 2025, yielding 10,866 im- ages. Earlier cases (2020â2022) and a small number of en- tries from early 2026 are excluded due to their very limited count. To construct a balanced evaluation set, we sample an equal number of images from the Miscaptioned classâreal images presented in misleading contextsâto serve as the au- thentic ground truth. To control for temporal shifts in image quality and model capabilities, sampling is matched by year to the AI-generated set (1,524 images for 2023; 3,822 for 2024; 5,520 for 2025). The resulting test set contains ap- proximately 21.7K images, balanced between AI-generated and authentic images, and is used exclusively to evaluate off-the-shelf models without further training. We evaluate three Synthetic Image Detectors (SIDs), SPAI, RINE, and BFree, and three VLMs, Gemma 3 27B, Grok 4.1 Fast, and GPT-5-mini. We evaluate all VLMs in a zero-shot setting and employ SIDs using their original pre- trained weights without additional fine-tuning. For SPAI, we use a probability threshold of 0.5. For RINE, predic- tions at or above MODERATE EVIDENCE are classified as AI-generated. BFree outputs a single logit score per image; positive values are classified as AI-generated and negative values as non-AI. For VLMs, we prompt the model to re- turn a single label (AI or real) based on the image alone. See Appendix 13 for details and prompts. 202320242025 40% 50% 60% 70% 80% TPR 202320242025 10% 15% 20% 25% 30% FPR 202320242025 55% 60% 65% 70% 75% 80% Accuracy SPAIRINEBFREEGemma-3GrokGPT-5-Mini Figure 5. Temporal performance of AI-image detection systems from 2023 to 2025. Panels show TPR, FPR, and Accuracy across six- month evaluation periods. TPR declines consistently across models, while FPR remains relatively stable, suggesting that performance degradation is driven by reduced sensitivity to newer AI-generated images. Table 3. Overall benchmark performance on the image dataset. Bold denotes best values and underlined the second-best. ModelTPRâ FPRâ Precisionâ F1â Accuracyâ SPAI41.48 24.6662.9049.9958.34 RINE52.75 30.5463.7557.7361.03 BFree52.72 27.8165.6558.4762.42 Gemma 3 27B 70.21 31.0069.5169.8669.60 Grok 4.1 Fast 52.28 21.1671.3560.3565.51 GPT-5-mini55.00 15.0678.6464.7369.91 We treat the AI-generated class as the positive class in all evaluations and report True Positive Rate (TPR), False Positive Rate (FPR), Precision, F1-score, and Accuracy. 8.2. Overall Results Table 3 summarizes overall detection performance across all evaluated models. VLMs consistently outperform the specialized SIDs across all reported metrics.GPT-5- mini achieves the highest overall accuracy (69.91%), while BFree is the strongest among the specialized detectors (ac- curacy 62.42%, F1 58.47%). This indicates that state-of- the-art general-purpose VLMs can match or exceed the per- formance of dedicated SIDs in an âin the wildâ setting. The primary differences between models appear in the trade-off between TPR and FPR. Gemma 3 achieves the highest TPR (70.21%) and F1-score (69.86%), at the cost of a high FPR (31%). In contrast, GPT-5-mini exhibits the opposite be- havior: it achieves the lowest FPR (15.06%) and the high- est precision (78.64%), while maintaining moderate TPR (55%). Grok 4.1 Fast occupies a middle ground, with bal- anced TPR and FPR, while SIDs generally exhibit lower TPR with only moderate reductions in FPR. 8.3. Temporal Degradation To examine how detection performance evolves over time, we evaluate each model across the 2023â2025 period using six-month bins. As Figure 5 shows, across all models, TPR declines substantially over time, with drops ranging from roughly 16% to 36% between early 2023 and late 2025. The strongest degradation is observed for the specialized SIDs: RINE declines from 74.69% in early 2023 to 39.34% in late 2025, while BFree decreases from 70.66% to 41.73% over the same period. VLMs also exhibit declining TPR, though to a slightly lesser extent. For example, Gemma 3 drops from 82.15% to 62.22%, and Grok declines from 61.12% to 45.67%. These TPR reductions also translate in decrease in overall accuracy across the same period. For instance, RINE drops from 74.16% in early 2023 to 54.04% in late 2025, while Gemma 3 drops from 75.44% to 65.38%. In contrast, FPR remains relatively stable over time, fluctuat- ing within a narrow range. This indicates that the perfor- mance degradation is primarily driven by reduced sensitiv- ity to AI-generated images rather than by increased misclas- sification of authentic images. This performance decline reflects a distribution shift caused by the rapid evolution of generative models. SIDs trained on earlier generations of synthetic data rely on visual cues that become less reliable as AI-generated images be- come more realistic and stylistically diverse, while VLMs, despite broader pretraining and stronger reasoning, are con- strained by their training data cutoffs. Additionally, given the training of VLMs on large-scale web data, it is pos- sible that some evaluation imagesâor visually similar in- stancesâwere present in their training corpora, which may partially influence performance. These results highlight that relying solely on static AI-image detectors is insufficient in rapidly evolving misinformation environments. 8.4. Qualitative Analysis Figure 6 shows four examples of AI-generated images and corresponding predictions by SIDs and VLMs. In (a), all models correctly identify the image as synthetic, which may indicate the presence of noticeable artifacts alongside se- mantic inconsistencies that VLMs can potentially capture. SPAI1.00â RINEVery Strongâ BFree2.197 â Gemma 3AIâ Grok 4.1AIâ GPT 5AIâ SPAI0.0273â RINEWeakâ BFree-2.4059â Gemma 3AIâ Grok 4.1AIâ GPT 5AIâ SPAI0.998â RINEVery Strongâ BFree5.821â Gemma 3Realâ Grok 4.1Realâ GPT 5Realâ SPAI3.10e-06â RINEWeakâ BFree-0.352â Gemma 3AIâ Grok 4.1Realâ GPT 5Realâ (a)(b) (c)(d) Figure 6. Examples of AI-generated images with corresponding predictions from SIDs and VLMs. Notably, the image is relatively older, posted in 2023, and it is possible that some models have been trained on it or similar images. In (b), all three SIDs fail while VLMs make the correct prediction. The image is relatively clean and lacks strong low-level artifacts, leading to weak SID sig- nals, while its unusual or less plausible semantic content ânamely, the pope with a rainbow flagâ enables VLMs to classify it as AI-generated. The opposite pattern appears in (c), where SIDs correctly detect the image as synthetic, possibly due to subtle artifact patterns, whereas VLMs in- correctly predict it as real; if the supposed identity of the depicted figure (former President Biden) is not recognized, the scene remains semantically plausible and does not raise obvious concerns. Finally, (d) presents a case where most models fail, potentially suggesting that the image resem- bles a degraded or historical photograph, thereby obscuring artifact-based cues. Taken together, these examples sug- gest that SIDs and VLMs may rely on different and po- tentially complementary signalsâlow-level artifacts versus high-level semanticsâwhich may account for their differ- ing predictions across cases. 9. Conclusion In this study, we present CONVEX, a large-scale dataset of multimodal misinformation, miscaptioned, edited, and AI- generated images and videos, collected from Xâs Commu- nity Notes. We leverage this data to conduct a longitudinal analysis of how misinformation evolves in terms of virality, engagement, consensus dynamics, and detectability. Our results show that AI-generated visual content is un- dergoing a rise in volume as generative models evolve. While it achieves disproportionate virality, this spread is driven primarily by passive engagement (e.g., favorites) rather than the active discource typical of miscaptioned media. Furthermore, despite slower initial reporting, AI- generated visuals reach community consensus more quickly than other categories. This suggests that synthetic content currently possesses recognizable artifacts or standardized cues that facilitate collective verificationâa process increas- ingly aided by the integration of AI-detection tools within the crowd-sourced annotation process. To explore the reliability of specialized synthetic image detectors and VLMs, we create an evaluation benchmark of authentic vs. AI-generated images. Our evaluation un- covers a significant decline in True Positive Rate over time. This highlights the vulnerability of static detection systems, which struggle to generalize as generative models become more realistic and closer to authentic imagery. The dissemination of AI-generated content and misin- formation represents a dynamic challenge within a rapidly shifting digital environment. Due to the quick evolution of generative capabilities, this landscape necessitates a human- in-the-loop approach that combines community monitoring with automated fact-checking. Such efforts require constant improvement through the iterative re-training of detection models to keep pace with advancements in generative AI. While this study offers a snapshot of the current state of the field, our proposed pipeline is designed for long-term analysis. It can be used to continuously monitor the evolu- tion of synthetic media as new Community Notes data be- comes available, providing researchers and platforms with real-time insights into multimodal misinformation. Acknowledgments This work received funding by the Horizon Europe projects AI-CODE (grant agreement no. 101135437) and AI4Trust (101070190). References [1] Mubashara Akhtar, Michael Schlichtkrull, Zhijiang Guo, Oana Cocarascu, Elena Simperl, and Andreas Vlachos. Mul- timodal automated fact-checking: A survey. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5430â5448, 2023. 1, 2 [2] Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, et al. Factuality challenges in the era of large lan- guage models and opportunities for fact-checking. Nature Machine Intelligence, 6(8):852â863, 2024. 1, 2 [3] Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty, David Corney, Emilio Ferrara, Iryna Gurevych, Scott Hale, Eduard Hovy, Heng Ji, Irene Larraz, et al. Community mod- eration and the new epistemology of fact checking on social media. arXiv preprint arXiv:2505.20067, 2025. 1, 2 [4] W Lance Bennett and Steven Livingston. The disinformation order: Disruptive communication and the decline of demo- cratic institutions. European journal of communication, 33 (2):122â139, 2018. 1 [5] Christina Boididou, Katerina Andreadou, Symeon Pa- padopoulos, Duc Tien Dang Nguyen, Giulia Boato, Michael Riegler, Yiannis Kompatsiaris, et al. Verifying multimedia use at mediaeval 2015. In MediaEval 2015. CEUR-WS, 2015. 2 [6] Christina Boididou, Stuart E Middleton, Zhiwei Jin, Symeon Papadopoulos, Duc-Tien Dang-Nguyen, Giulia Boato, and Yiannis Kompatsiaris. Verifying information with multime- dia content on twitter: a comparative study of automated approaches. Multimedia tools and applications, 77:15545â 15571, 2018. 2 [7] Nadav Borenstein, Greta Warren, Desmond Elliott, and Is- abelle Augenstein. Can community notes replace profes- sional fact-checkers?In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 535â552, 2025. 1, 2 [8] Yuwei Chuai, Haoye Tian, Nicolas Pr Ě ollochs, and Gabriele Lenzini. Did the roll-out of community notes reduce en- gagement with misinformation on x/twitter? Proceedings of the ACM on human-computer interaction, 8(CSCW2):1â52, 2024. 2 [9] Soham De, Michiel A Bakker, Jay Baxter, and Martin Saveski. Supernotes: Driving consensus in crowd-sourced fact-checking. In Proceedings of the ACM on Web Confer- ence 2025, pages 3751â3761, 2025. 2 [10] Jingyi Deng, Chenhao Lin, Zhengyu Zhao, Shuai Liu, Zhe Peng, Qian Wang, and Chao Shen. A survey of defenses against ai-generated visual media: Detection, disruption, and authentication. ACM Computing Surveys, 58(5):1â35, 2025. 1 [11] Israel Junior Borges Do Nascimento, Ana Beatriz Pizarro, Jussara M Almeida, Natasha Azzopardi-Muscat, Mar- cos Andr Ě e Gonc ̧alves, Maria Bj Ě orklund, and David Novillo- Ortiz. Infodemics and health misinformation: a systematic review of reviews. Bulletin of the World Health Organiza- tion, 100(9):544, 2022. 1 [12] Nicholas Dufour, Arkanath Pathak, Pouya Samangouei, Nikki Hariri, Shashi Deshetti, Andrew Dudfield, Christo- pher Guess, Pablo Hern Ě andez Escayola, Bobby Tran, Mevan Babakar, et al. Ammeba: A large-scale survey and dataset of media-based misinformation in-the-wild. arXiv preprint arXiv:2405.11697, 1(8), 2024. 1, 2 [13] Jos Ě e Gamir-R Ě Äąos, Raquel Tarullo, Miguel Ib Ě a Ě nez-Cuquerella, et al. Multimodal disinformation about otherness on the in- ternet. the spread of racist, xenophobic and islamophobic fake news in 2020. An ` alisi, pages 49â64, 2021. 1 [14] Fabrizio Guillaro, Giada Zingarini, Ben Usman, Avneesh Sud, Davide Cozzolino, and Luisa Verdoliva. A bias-free training paradigm for more general ai-generated image de- tection. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18685â18694, 2025. 2 [15] Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. A survey on automated fact-checking. Transactions of the association for computational linguistics, 10:178â206, 2022. 1 [16] Michael Hameleers, Thomas E Powell, Toni GLA Van Der Meer, and Lieke Bos. A picture paints a thousand lies? the effects and mechanisms of multimodal disinformation and rebuttals disseminated via social media. Political com- munication, 37(2):281â301, 2020. 2 [17] Shan Jia, Reilin Lyu, Kangran Zhao, Yize Chen, Zhiyuan Yan, Yan Ju, Chuanbo Hu, Xin Li, Baoyuan Wu, and Siwei Lyu. Can chatgpt detect deepfakes? a study of using mul- timodal large language models for media forensics. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4324â4333, 2024. 3 [18] Zhiwei Jin, Juan Cao, Yongdong Zhang, Jianshe Zhou, and Qi Tian. Novel visual and statistical image features for mi- croblogs news verification. IEEE transactions on multime- dia, 19(3):598â608, 2016. 2 [19] Zhiwei Jin, Juan Cao, Han Guo, Yongdong Zhang, and Jiebo Luo. Multimodal fusion with recurrent neural networks for rumor detection on microblogs. In Proceedings of the 25th ACM international conference on Multimedia, pages 795â 816, 2017. 2 [20] Uku Kangur, Roshni Chakraborty, and Rajesh Sharma. Who checks the checkers? exploring source credibility in twitterâs community notes. Journal of Computational Social Science, 9(1):24, 2026. 2 [21] Dimitrios Karageogiou, Quentin Bammey, Valentin Por- cellini, Bertrand Goupil, Denis Teyssou, and Symeon Pa- padopoulos. Evolution of detection performance throughout the online lifespan of synthetic images. In European Con- ference on Computer Vision, pages 400â417. Springer, 2024. 2 [22] Dimitrios Karageorgiou, Symeon Papadopoulos, Ioannis Kompatsiaris, and Efstratios Gavves.Any-resolution ai- generated image detection by spectral learning. In Proceed- ings of the Computer Vision and Pattern Recognition Con- ference, pages 18706â18717, 2025. 2 [23] Christos Koutlis and Symeon Papadopoulos. Leveraging rep- resentations from intermediate encoder-blocks for synthetic image detection. In European Conference on computer vi- sion, pages 394â411. Springer, 2024. 2 [24] Simon Fox Kuuse, Uku Kangur, Roshni Chakraborty, and Rajesh Sharma. Crowdsourced fact-checking or biased com- mentary? analyzing political bias in twitterâs community notes. In Companion Proceedings of the ACM on Web Con- ference 2025, pages 2661â2669, 2025. 2 [25] Haiwen Li, Soham De, Manon Revel, Andreas Haupt, Brad Miller, Keith Coleman, Jay Baxter, Martin Saveski, and Michiel A Bakker. Scaling human judgment in community notes with llms. arXiv preprint arXiv:2506.24118, 2025. 2 [26] Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, and Shu Hu. Detecting multimedia generated by large ai models: A survey. arXiv preprint arXiv:2402.00045, 2024. 2 [27] Grace Luo, Trevor Darrell, and Anna Rohrbach. Newsclip- pings: Automatic generation of out-of-context multimodal media. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6801â6817, 2021. 2 [28] Arpan Mahara and Naphtali Rishe.Methods and trends in detecting ai-generated images: A comprehensive review. Computer Science Review, 60:100908, 2026. 2 [29] Cameron Martel, Jennifer Allen, Gordon Pennycook, and David G Rand. Crowds can effectively identify misinfor- mation at scale. Perspectives on Psychological Science, 19 (2):477â488, 2024. 2 [30] Yisroel Mirsky and Wenke Lee. The creation and detection of deepfakes: A survey. ACM computing surveys (CSUR), 54(1):1â41, 2021. 2 [31] Saeedeh Mohammadi, Narges Chinichian, Hannah Doyal, Kristina Skutilova, Hao Cui, Michele dâErrico, Siobhan Grayson, and Taha Yasseri. From birdwatch to community notes, from twitter to x: four years of community-based con- tent moderation. arXiv preprint arXiv:2510.09585, 2025. 2 [32] Eric M Ě uller-Budack, Jonas Theiner, Sebastian Diering, Max- imilian Idahl, and Ralph Ewerth. Multimodal analytics for real-world news using measures of cross-modal entity con- sistency. In Proceedings of the 2020 international confer- ence on multimedia retrieval, pages 16â25, 2020. 2 [33] Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis Petrantonakis. Synthetic mis- informers: Generating and combating multimodal misinfor- mation. In Proceedings of the 2nd ACM International Work- shop on Multimedia AI against Disinformation, pages 36â44, 2023. 2 [34] Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis C Petrantonakis. Verite: a ro- bust benchmark for multimodal misinformation detection ac- counting for unimodal bias. International Journal of Multi- media Information Retrieval, 13(1):4, 2024. 2 [35] Olesya Razuvayevskaya, Adel Tayebi, Ulrikke Dybdal Sørensen, Kalina Bontcheva, and Richard Rogers. Timeli- ness, consensus, and composition of the crowd: Community notes on x. arXiv preprint arXiv:2510.12559, 2025. 2 [36] Jonathan Tonglet, Gabriel Thiem, and Iryna Gurevych. Cove: Context and veracity prediction for out-of-context im- ages. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computa- tional Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 2029â2049, 2025. 2 [37] Teresa Weikmann and Sophie Lecheler. Visual disinforma- tion in a digital age: A literature synthesis and research agenda. New Media & Society, 25(12):3696â3713, 2023. 1, 2 [38] Valerie Wirtschafter and Sharanya Majumder. Future chal- lenges for online, crowdsourced content moderation: evi- dence from twitterâs community notes. Journal of Online Trust and Safety, 2(1), 2023. 1, 2 [39] Stefan Wojcik, Sophie Hilgard, Nick Judd, Delia Mocanu, Stephen Ragain, MB Hunzaker, Keith Coleman, and Jay Baxter. Birdwatch: Crowd wisdom and bridging algorithms can inform understanding and reduce the spread of misinfor- mation. arXiv preprint arXiv:2210.15723, 2022. 1, 2 [40] Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang. End-to-end multimodal fact-checking and explanation generation: A challenging dataset and mod- els. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2733â2743, 2023. 2 [41] Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen.A survey on multimodal large language models. National Science Review, 11(12): nwae403, 2024. 3 [42] Dimitrina Zlatkova, Preslav Nakov, and Ivan Koychev. Fact- checking meets fauxtography: Verifying claims about im- ages. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter- national Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2099â2108, 2019. 2