Paper deep dive
AI-generated Images Challenge Visual Trust in High-risk Scenarios
Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/28/2026, 3:44:08 AM
Summary
The paper introduces SafeIMG, a safety-oriented benchmark for evaluating AI-generated image detection in high-risk public and individual safety scenarios. Using GPT Image 2 to generate 12 specific risk categories, the study finds that current specialized detectors and vision-language models (VLMs) significantly underperform compared to human evaluators (81.7% accuracy). Models struggle with commonsense and physical inconsistencies, covering only 29.8% of human-annotated anomalies, highlighting a critical gap in reliable visual trust for safety-critical evidence.
Entities (12)
Relation Signals (11)
SafeIMG → covers → Public Safety
confidence 95% · SafeIMG integrates... public safety and individual safety.
SafeIMG → covers → Individual Safety
confidence 95% · SafeIMG integrates... public safety and individual safety.
SafeIMG → uses → GPT Image 2
confidence 95% · SafeIMG... generated using GPT Image 2
SafeIMG → evaluates → LNP
confidence 90% · evaluate specialized synthetic-image detectors... LNP
SafeIMG → evaluates → Gemini 3.1 Pro
confidence 90% · Comparison of representative models... Gemini 3.1 Pro
SafeIMG → evaluates → Claude Opus 4.6
confidence 90% · Comparison of representative models... Claude Opus 4.6
SafeIMG → evaluates → Doubao Seed 2.0 Pro
confidence 90% · Comparison of representative models... Doubao Seed 2.0 Pro
SafeIMG → evaluates → CNNSpot
confidence 90% · evaluate specialized synthetic-image detectors... CNNSpot
SafeIMG → evaluates → Vision-Language Models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2. Unlike benchmarks centred on generic imagery and image-level labels, SafeIMG evaluates not only whether detectors recognise synthetic images, but also whether their decisions reflect human-identified anomalies. To this end, SafeIMG provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies. We evaluate specialized synthetic-image detectors and vision-language models (VLMs), and find that neither provides reliable detection. The strongest VLM identifies only 49.5% of generated images, whereas the best specialised detector identifies 33.1%, compared with 81.7% accuracy for human evaluators. Model explanations cover only 29.8\% of human-annotated anomalies and predominantly capture local defects in text, faces and hands. Their coverage falls to 15.0% for commonsense conflicts and 12.0% for physical inconsistencies, while detection performance deteriorates further after dissemination-induced image degradation. These findings show that current detectors lack the accuracy, explanatory alignment and robustness needed to evaluate AI-generated images reliably across public- and individual-safety settings.
Tags
Links
- Source: https://arxiv.org/abs/2607.22745v1
- Canonical: https://arxiv.org/abs/2607.22745v1
Trouble viewing inline? Open PDF directly →
Full Text
102,181 characters extracted from source content.
Expand or collapse full text
AI-generated Images Challenge Visual Trust in High-risk Scenarios Yi-Zhi Wang 1,2 , Yichen Xiao 1,2 , Linan Yue 1,2 , Weibo Gao 3 , Yichao Du 4 , Pengfei Fang 1,2 , Shimin Di 1,2 , and Min-Ling Zhang 1,2 1 School of Computer Science and Engineering, Southeast University, 2 Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, 3 The Hong Kong Polytechnic University, 4 School of Artificial Intelligence, Wuhan University Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2. Unlike benchmarks centred on generic imagery and image-level labels, SafeIMG evaluates not only whether detectors recognise synthetic images, but also whether their decisions reflect human-identified anomalies. To this end, SafeIMG provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies. We evaluate specialized synthetic-image detectors and vision-language models (VLMs), and find that neither provides reliable detection. The strongest VLM identifies only 49.5% of generated images, whereas the best specialised detector identifies 33.1%, compared with 81.7% accuracy for human evaluators. Model explanations cover only 29.8% of human-annotated anomalies and predominantly capture local defects in text, faces and hands. Their coverage falls to 15.0% for commonsense conflicts and 12.0% for physical inconsistencies, while detection performance deteriorates further after dissemination-induced image degradation. These findings show that current detectors lack the accuracy, explanatory alignment and robustness needed to evaluate AI-generated images reliably across public- and individual-safety settings. Projects: https://safeimg.github.io/ Code Repository: https://github.com/Snowstorm1492/SafeIMG Datasets: https://huggingface.co/datasets/Snowstorm1492/SafeIMG P e r s o n - I 1 P e r s o n - I 2 P e r s o n - I 3 P e r s o n - I 4 1 P - c i l b u P 2 P - c i l b u P 3 P - c i l b u P 4 P - c i l b u P 5 P - c i l b u P 6 P - c i l b u P P u b l i c - P 7 P u b l i c - P 8 LNP FreDect Gemini 3.1 Pro Claude Opus 4.6 Doubao Seed 2.0 Pro Doubao Seed 2.0 Mini 53.1 35.6 24.4 12.2 22.2 2.2 21.4 20.9 34.1 53.8 45.1 0 40.7 26.6 40.4 78.7 38.3 1.1 40.8 37.6 42.4 25.9 41.2 1.2 54.4 43.4 38.1 30.4 33.6 13.3 63.2 48.1 52.9 42.3 20.2 9.6 57.0 54.7 47.7 29.1 41.9 18.6 64.8 46.2 38.5 29.7 42.9 7.7 60.0 31.6 41.1 33.0 37.9 12.6 54.5 38.4 22.2 16.2 41.4 14.1 57.4 52.1 45.7 17.0 18.1 8.5 42.7 40.4 39.3 46.1 33.7 14.6 SafeIMG (a) Category-level image detection per- formance on SafeIMG. e s n e s n o m m o C 51.8% HIT 15.0% MISS 85.0% Text 21.5% HIT 58.7% MISS 41.3% Other 10.7% HIT 28.6% MISS 71.4% Faces 10.5% HIT 56.2% MISS 43.8% Physics 2.7% HIT 12.0% MISS 88.0% Hands 1.9% HIT 64.7% MISS 35.3% Light 0.9% HIT 25.0% MISS 75.0% (b) Distribution of human anno- tated artifact types and hit rates. I1(9) I2(33) I3(16) I4(13) P1(12) P2(6)P5(20) P6(10) P7(10) P8(2) Overall 2030405060 Gemini 3.1 Pro Preview Claude Sonnet 4.6 Doubao Seed 2.0 Pro Qwen3.6 Plus Qwen3.6 Flash GPT-5.5 (c) Comparison of representative models on difficult subset. Figure 1: Overview of SafeIMG and key empirical findings. The figure highlights the challenges of detecting AI- generated images in high-risk scenarios, where current models exhibit uneven capability and limited understanding of human-perceived artifacts. Corresponding author(s): Linan Yue, Email: lnyue@seu.edu.cn arXiv:2607.22745v1 [cs.CV] 23 Jul 2026 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 1. Introduction Image generation models have advanced rapidly in recent years (Ho et al., 2020, Song et al., 2021, Lipman et al., 2023, Esser et al., 2024, Chen et al., 2024a, Yan et al., 2025, Chen et al., 2025). Modern text-to-image systems can synthesize realistic images, follow detailed natural-language instructions, and compose complex visual scenes with fine-grained control over objects, layout, style, and text (Zhang et al., 2023, Tuo et al., 2023, Esser et al., 2024). These capabilities are no longer limited to professional design software or closed research systems. Public products and API services now allow non-specialist users to generate large volumes of realistic images at low cost by writing natural-language prompts. This progress changes the security implications of synthetic images. Earlier AI-generated images were often discussed as aesthetic or creative artifacts, where realism and visual quality were the main concerns (Heusel et al., 2017, Karras et al., 2019, Dhariwal and Nichol, 2021). Recent models can increasingly produce images with evidentiary attributes. These images appear to document real-world facts, events, identities, transactions, or records (Chesney and Citron, 2019, Vaccari and Chadwick, 2020, Nightingale and Farid, 2022). For example, a generated image may resemble a news photograph from a disaster scene, a bank- transfer screenshot, an identity credential, a legal document, or a private communication record. In such cases, the image is not merely consumed as visual content. It may shape public judgment, financial trust, identity verification, legal interpretation, or personal reputation. As synthetic images enter these high-risk contexts, the social assumption that “seeing is believing” becomes increasingly fragile (Chesney and Citron, 2019, Vaccari and Chadwick, 2020, Nightingale and Farid, 2022). Detecting AI-generated images has therefore become an important problem in multimodal safety. In the upper part of Figure 2, existing studies have proposed benchmarks and detection methods for synthetic-image recognition (Wang et al., 2020, Ojha et al., 2023, Sha et al., 2022, Zhu et al., 2023). One line of work focuses on general natural images, emphasizing large-scale data construction, cross-generator generalization, and robustness to common degradations (Wang et al., 2020, Zhu et al., 2023, Park and Owens, 2024). Another line examines more specific domains, including scientific images, artistic images, originality judgment, and vision-language models (VLMs) for forged-content recognition (Hu et al., 2026, Li and Stamp, 2025, Zhu et al., 2025, Chen et al., 2024b). These studies provide important foundations for understanding synthetic-image detection. However, existing benchmarks remain insufficient for evaluating AI-generated images in safety-critical evidentiary scenarios. First, many classic benchmarks were constructed with earlier generators, such as BigGAN (Brock et al., 2019), ADM (Dhariwal and Nichol, 2021), GLIDE (Nichol et al., 2021), VQ-Diffusion (Gu et al., 2022), or early versions of Stable Diffusion (Rombach et al., 2022). Although these generators were representative at the time, they do not capture the realism of recent image generation systems. Detectors that perform well on earlier synthetic images may therefore fail to generalize to images produced by current models (Ojha et al., 2023, Wang et al., 2023, Yan et al., 2024, Park and Owens, 2024). Second, even recent benchmarks (Zhu et al., 2023, Park and Owens, 2024) primarily contain general-purpose or domain-specific imagery without an explicit focus on public and individual safety. It therefore remains unclear how reliably current detectors identify synthetic images across safety-sensitive contexts. Third, current evaluations often rely on overall detection accuracy, which provides limited diagnostic insight (Zhu et al., 2023, Chen et al., 2024b, Yang et al., 2025). In safety-oriented detection, image types may differ substantially in risk and difficulty. A detector may perform well on public-safety images but fail on transaction screenshots. It may detect news-like photographs but fail after cropping, re-encoding, screenshots, or social-media transformations. Without a fine-grained taxonomy of risk domains and evidence types, it is difficult to identify where detectors are reliable and where they remain unsafe. A further limitation concerns the evidence used to make authenticity judgements. Some generated images 2 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios · P1-Natural Disasters P2-Unrest & Attack P3-Traffic Accidents P4-Fire & Explosion P5-Pollution & HazMat P6-Infrastructure Damage P7-Crowd Safety P8-Public Health I1-Personal Emergencies I2-Transaction Proofs I3-Communication Records I4-Identity Endorsement CommonsenseFontsOtherFacesPhysicsHandsLighting Undersized car door Disproportionate fingersWrong characterAsymmetric eyes Incorrect time Inconsistent road width Physics Opposite shadows Category distribution of SafeIMG Human-annotated category distribution of SafeIMG AIGCDetectBenchmark GenImage Single object & Lacks realism Detecting AI-generated Artwork RealHD Irrelevance to security Realistic and detailed scenes Story-rich and controversial contents Fine-grained annotations and explanations Generated by GPT Image 2 → More challenging Existing benchmarks Our benchmark Figure 2: Overview and distinguishing characteristics of SafeIMG. SafeIMG organises 12 safety-oriented scenarios into public- and individual-safety domains and uses structured prompts to generate realistic, context-rich images with GPT Image 2. Human annotators localise suspicious regions and provide artefact categories and textual explanations, covering both local visual defects and higher-level commonsense or physical inconsistencies. The comparison with existing benchmarks highlights SafeIMG’s emphasis on safety relevance, complex visual content and fine-grained diagnostic annotations. contain visible local artefacts, such as malformed text, faces or hands. These cues may become less dependable as generation models improve. Visual realism also does not necessarily imply real-world plausibility. An image may appear photorealistic while containing implausible spatial relations, broken causal structure, commonsense conflicts or violations of physical constraints. Detecting these problems requires more than recognising low-level generator traces. It requires assessing whether the depicted content is coherent with the physical, logical and social context it claims to represent. Image-level labels cannot determine whether a detector has identified such anomalies or reached a correct decision for an unrelated reason. To address these gaps, we introduce SafeIMG, a safety-oriented benchmark for AI-generated image detection. As shown in Figure 2, SafeIMG integrates three components: risk-oriented scenario design, structured image generation and fine-grained artefact annotation. Its scenario taxonomy comprises 12 categories across two domains, public safety and individual safety. The public-safety domain covers disasters, violence, transport accidents, hazardous incidents, infrastructure failures, crowd safety and public-health events. The individual-safety domain includes personal emergencies, transaction records, private communications, fabricated scene evidence and identity endorsement. Rather than treating these categories as generic visual classes, SafeIMG defines them according to their evidentiary functions and potential safety consequences. For each category, a risk-aware scenario space specifies the relevant content, media forms and safety concerns. These specifications guide the construction of contextually grounded prompts, which are used to generate synthetic images with GPT Image 2. Human annotators subsequently localise suspicious regions, assign artefact categories and provide textual explanations for the identified anomalies. The annotation schema covers both explicit local defects and higher-level inconsistencies involving commonsense, scene logic and physical plausibility. Together, these components support evaluation of not only whether a detector recognises synthetic content, but also where and why it considers an image suspicious. 3 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios To thoroughly investigate the detection capabilities and characteristics of existing methods, we evaluate two major categories of AI-generated image detectors: specialized forensic models designed for synthetic-image detection (e.g., CNNSpot (Wang et al., 2020), FreDect (Frank et al., 2020), LNP (Liu et al., 2022)), and VLMs with image understanding and forgery-recognition capabilities (e.g., GPT series (OpenAI, 2025, 2026), Claude series (Anthropic, 2026), Qwen Series (Bai et al., 2025, Qwen Team, 2026)). Across these evaluations, three consistent patterns emerge from the results. First, automatic systems remain markedly less sensitive than humans. The strongest VLM recognises 49.5% of generated images, and the best specialised detector recognises 33.1%, whereas human evaluators achieve 81.7% accuracy. Performance also varies substantially across the 12 safety scenarios, indicating that aggregate scores conceal important category-specific failures. Second, model explanations cover only 29.8% of human-annotated anomalies and focus predominantly on local cues involving text, faces and image texture. They are substantially less sensitive to commonsense conflicts and physical inconsistencies that require scene-level reasoning. Finally, robustness is also limited, with one representative model losing more than 90% of its initial detection accuracy under severe dissemination-induced degradation. These findings show that safety-oriented image detection requires evaluation beyond aggregate accuracy, including scenario-specific reliability, explanatory validity and robustness under realistic propagation conditions. Our contributions are summarized as follows: • We formulate safety-oriented AI-generated image detection as an evaluation setting centred on public- and individual-safety contexts. This setting shifts the focus from generic authenticity classification towards risk-sensitive assessment of potentially misleading visual content. • We introduce SafeIMG, a benchmark covering 12 safety-sensitive scenarios generated using GPT Image 2. Its risk-oriented taxonomy and scenario spaces specify the evidentiary function, relevant content, media forms and potential consequences of each category. •We provide fine-grained human annotations that include suspicious-region localisation, artefact categories and textual explanations. The annotations cover both local visual defects and higher-level commonsense or physical inconsistencies, enabling diagnostic evaluation and human–AI explanation alignment analysis. •We systematically evaluate multiple types of detection methods and find that current models more readily recognize explicit local defects but struggle to identify commonsense conflicts and physical anomalies, revealing a key capability gap in high-risk image detection. 2. SafeIMG Benchmark SafeIMG is a safety-oriented benchmark of AI-generated images covering 12 scenarios across public- and individual-safety domains. The final benchmark contains two complementary subsets. The annotated subset comprises images with human-identified anomalies, each accompanied by suspicious-region localisation, a fine-grained artefact label and a textual explanation. Its annotations cover both conventional local defects and higher-level inconsistencies involving commonsense, scene logic and physical constraints. The extremely difficult subset contains highly realistic images for which annotators cannot identify a defensible visual anomaly. This design enables SafeIMG to evaluate both evidence-grounded detection and detection when no readily observable human cue is available. As shown in Figure 3, SafeIMG is constructed through five stages. First, the Risk-Oriented Image Taxonomy organises the benchmark into public- and individual-safety domains covering 12 high-risk scenarios. Second, Scenario Space Specification defines the relevant content and potential safety consequences of each category. Third, concrete scenario instances are converted into structured prompts. Fourth, GPT Image 2 generates a candidate pool of synthetic images. Finally, human screening removes invalid generations and separates the retained images into the annotated and extremely difficult subsets. Images entering the annotated subset are further labelled with suspicious regions, artefact categories and natural-language explanations. 4 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Physics: One end of the caution tape is broken and appears to be floating unnaturally. Fonts: The back of the police uniform displays garbled characters. Other: The chair in the image has an abnormal form. Other: The clothing has a strange shape, seemingly featuring two hoods. On the afternoon of May 11, 2024 (When), the riverside promenade along the Pudong waterfront in Shanghai (Where) was hit by a storm surge (What). The river water overflowed the wave-break platform and submerged benches and the bases of lampposts. The warning tape was tilted by the wind. Staff members persuaded two citizens (Who) to leave. The turbid water surface was strewn with branches, plastic road barriers, and broken wooden planks. In the distance, tall buildings were reflected in the yellowish floodwater... Physics: One end of the caution tape is broken and appears to be floating unnaturally. Fonts: The back of the police uniform displays garbled characters. Other: The chair in the image has an abnormal form. Other: The clothing has a strange shape, seemingly featuring two hoods. On the afternoon of May 11, 2024 (When), the riverside promenade along the Pudong waterfront in Shanghai (Where) was hit by a storm surge (What). The river water overflowed the wave-break platform and submerged benches and the bases of lampposts. The warning tape was tilted by the wind. Staff members persuaded two citizens (Who) to leave. The turbid water surface was strewn with branches, plastic road barriers, and broken wooden planks. In the distance, tall buildings were reflected in the yellowish floodwater... "Images of natural disasters may be used to prove the occurrence of a disaster, the extent of the affected areas... If such images are fabricated or mistakenly perceived as authentic, they may exaggerate the severity of the disaster, incite public panic..." Content List Media Form List Risk Summary 퐀 1 퐀 1 퐀 1 flood, earthquake, typhoon, tsunami, torrential rain,... documentary photography, news coverage footage Stage 2: Scenario Space SpecificationStage 3: Prompt Construction Stage 1: Risk-Oriented Image Taxonomy Public Safety Individual Safety P1 P8 I1I2 I3I4 P2P3P4 P5P6P7 Stage 5: Human AnnotationStage 4: Image Generation i p Prompt GPT Image 2 "Images of natural disasters may be used to prove the occurrence of a disaster, the extent of the affected areas... If such images are fabricated or mistakenly perceived as authentic, they may exaggerate the severity of the disaster, incite public panic..." Content List Media Form List Risk Summary flood, earthquake, typhoon, tsunami, torrential rain,... documentary photography, news coverage footage Stage 2: Scenario Space SpecificationStage 3: Prompt Construction Stage 1: Risk-Oriented Image Taxonomy Public Safety Individual Safety P1 P8 I1I2 I3I4 P2P3P4 P5P6P7 Stage 5: Human Annotation Stage 4: Image Generation Prompt GPT Image 2 On the afternoon of May 11, 2024... AI-Generated Image p i G(·) x i ε 1 M 1 R 1 Figure 3: Construction pipeline of the SafeIMG benchmark for high-risk visual evidence scenarios. SafeIMG is built through five stages. (1) The taxonomy in stage 1 defines 12 high-risk image categories under public or individual safety. (2) For each category, the scenario space specifies the content listℰ, media form listℳ, and risk summaryR; (3) complete prompts are then constructed before (4) generating images with GPT Image 2. (5) Finally, human annotators localize suspicious regions, assign artifact labels, and provide textual explanations for fine-grained evaluation. 2.1. Risk-Oriented Image Taxonomy Existing AI-generated image detection benchmarks are mostly organized based on general visual semantics, such as objects, styles, or generator sources. Such categorization is suitable for evaluating general image authenticity detection, but it is insufficient for covering safety issues in high-risk evidence scenarios. In real- world dissemination, the risk of an image depends not only on its visual content, but also on its evidentiary use. Based on this consideration, to address the lack of safety-risk coverage and evidentiary image scenarios in existing benchmarks, SafeIMG constructs a risk-oriented image taxonomy. The data are organized according to the possible function of an image and its potential safety impact. Specifically, as shown in Table 1, SafeIMG divides high-risk images into two major groups: public safety and individual safety: The public safety group contains 8 categorie: P1 Natural Disasters; P2 Social Unrest, Public Violence, and Attack Incidents; P3 Transportation Accidents; P4 Fires, Explosions, and Energy Facility Accidents; P5 Pollution and Hazardous Material Accidents; P6 Building and Civil Infrastructure Accidents; P7 Crowd Gathering and Venue Safety Accidents; and P8 Public Health and Biosecurity Events. These images are often used to support claims about whether an event occurred, as well as the corresponding degree of damage or risk level. If such images are fabricated or misclassified as real, they may mislead the public’s judgment of the authenticity and severity of the event, cause confusion or panic, and affect public opinion and social order. The individual safety group contains 4 categories: I1 Personal Accidents and Emergencies; I2 Private Receipts and Transaction Records; I3 Personal Chat and Communication Records; and I4 Fabricated Scene Evidence and Identity Endorsement. These categories focus on visual evidence that may be used in private communication, financial disputes, or personal identity claims. Compared with public safety scenarios, individual safety images are more directly related to the life and property safety of individuals. If such images are circulated or accepted as evidence, or even misused as tools for fraud, they may cause harm to personal property, reputation, privacy, and social trust. 5 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Table 1: Taxonomy of SafeIMG image categories and their evidentiary function. ID CategoryEvidentiary Function Public Safety Group P1 Natural DisastersSupports claims about disaster impact and emergency response. P2 Social Unrest, Public Violence, and Attack Incidents Shapes claims about public order, security risks, and social stability. P3 Transportation AccidentsSupports claims about accident occurrence, damage, and responsibility. P4 Fires, Explosions, and Energy Facility AccidentsSupports risk assessment of hazardous incidents and emergency handling. P5 Pollution and Hazardous Material AccidentsSupports claims that may trigger public panic or regulatory action. P6 Building and Civil Infrastructure AccidentsSupports claims about infrastructure safety, damage, and liability. P7 Crowd Gathering and Venue Safety AccidentsSupports judgments about crowd control, evacuation, and event safety. P8 Public Health and Biosecurity EventsSupports claims affecting public health decisions and social trust. Individual Safety Group I1 Personal Accidents and EmergenciesSupports claims about personal danger, rescue needs, or liability. I2 Private Receipts and Transaction RecordsSupports claims about payments, refunds, reimbursement, or disputes. I3 Personal Chat and Communication RecordsSupports claims about private commitments, misconduct, or disputes. I4 Fabricated Scene Evidence and Identity EndorsementSupports claims of presence, status, or endorsement. Overall, the taxonomy of SafeIMG does not attempt to exhaust harmful forms of AI-generated images. Instead, it selects representative high-risk scenarios to provide a basis for benchmark construction. By organizing categories around evidentiary function, SafeIMG shifts the evaluation focus from general visual realism detection to risk-sensitive image authenticity detection. This taxonomy serves as the basis for subsequent scenario space specification, prompt construction, image generation, and human annotation. 2.2. Scenario Space Specification In Section 2.1, we define the risk categories. However, the content and risk focus vary greatly across categories. The taxonomy above only specifies which high-risk categories should be covered by the dataset, but it does not directly determine what should be generated for each category. Directly generating images from category names may lead to two problems. First, scenes within the same category may become repetitive and visually homogeneous. Second, the risk focus of different categories may remain unclear, causing generated images to present only generic visual content without a clear risk orientation. Therefore, to ensure the quality and diversity of synthetic images, we specify a generation scope for each category, namely a Scenario Space, to guide prompt construction under different categories. Specifically, the Scenario Space of each category contains three components: (1) a risk summaryR c ; (2) a content listℰ c ; and (3) a media form listℳ c . The risk summary explains the core reason why images in this category may cause harm when fabricated or misused. The content list specifies which concrete events, objects, or materials can be generated under this category. The media form list specifies the visual forms in which such content commonly appears. Through this step, SafeIMG converts abstract risk categories into operable generation scopes. 풮 c = (R c , e, m) ∣ e ∈ ℰ c , m ∈ ℳ c .(1) For public safety images, the risk summary indicates the potential harm of images in the corresponding category. For example, the risk of P1 Natural Disasters lies in the possibility of exaggerating disaster severity, creating panic, or misleading rescue decisions. The content list mainly corresponds to different public events or accident scenes. For example, the content list of P1 includes 11 specific disaster types, such as floods, earthquakes, typhoons, and heavy rain. The media form list of public safety images mainly includes documentary photography and news-report images. For individual safety images, the risk summary indicates the specific personal-rights scenarios in which such images may be involved. For example, images in I2 Private Receipts and Transaction Records may be used in 6 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios telecom fraud. The content list may include either concrete events or various types of personal evidentiary materials. For example, Private Receipts and Transaction Records mainly include payment screenshots, refund records, transfer vouchers, and order pages. The media forms of individual safety images are more diverse, mainly including documentary photography, electronic-device screenshots, and photos taken of mobile phone screens. Through Scenario Space Specification, SafeIMG establishes an intermediate layer between taxonomy and prompt construction. The taxonomy determines which risk categories the dataset covers. The scenario space specifies the generation scope: what content can be generated under each category, what media forms can be used, and what risk focus should be highlighted. Subsequent prompt construction can then build complete prompts based on this scenario space. Scenario Space Specification can reduce repetitive generation within the same category and improve the relevance of synthetic images to high-risk scenarios. This stage does not directly generate final prompts. Instead, it specifies the generation scope of each category and lays the foundation for subsequent sample-level prompt construction. 2.3. Prompt Construction In the previous section, we define the scenario space풮 c for each risk category, specifying the content scope, media forms, and risk focus of that category. This section further constructs complete prompts based on the scenario space of each category. Specifically, we synthesize complete prompts in batches for each category. For each sample under a specific category, we first randomly select a concrete scenario instance from the scenario space of that category, and then expand it into a complete generation instruction using universal prompt expansion rules and category-specific constraints. Formally, this process can be written as: p i =Ψ(U, s i , A c ),s i ∈ 풮 c .(2) Here,p i denotes the final prompt.Ψ(⋅)denotes the prompt constructor,Udenotes the universal prompt expansion rules,s i denotes a concrete scenario instance, andA c is the category-specific constraints, specificity: The prompt expansion rulesUare shared by all categories and control the overall narrative structure of the prompt. We require the model to use the 5W+1H structure commonly used in news narration, namely Who, What, When, Where, Why, and How. This structure ensures that the prompt has basic scenario completeness, so that the generated result contains a relatively clear event background and detailed information. For public safety images, 5W+1H helps supplement the location, time, cause, affected subjects, and on-site process of the event. For individual safety images, this narrative structure also makes the background more detailed and the evidentiary chain more reliable. In this way, 5W+1H serves not only as a descriptive template, but also as a constraint on scenario completeness, making the generated images more concrete and credible. The concrete scenario instances i is selected from the scenario space of a specific category. It usually determines the basic content framework of the sample, including the main event or material type, the adopted media form, and the main risk orientation. For example, in the Natural Disasters category, a scenario instance can be “a news-report image showing severe urban road flooding caused by heavy rain, with high dissemination potential.” The category-specific constraintsA c supplements the special requirements of different categories in prompt construction. These constraints do not change the main content of the scenario instance, but they specify forms or details that are important for credibility. Different categories have different category-specific constraints. For example, for categories under the public safety group, we require the image to contain elements that indicate the time of the event. For I2 and I3 under the individual safety group, we require the model to explicitly avoid obvious placeholders such as “x” or “1234” in common textual elements, including text, ID numbers, names, amounts, timestamps, and institutional identifiers. Instead, the generated content 7 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios should be more natural, random, and contextually appropriate. The complete category-specific constraint settings are provided in the appendix. In practice, we combineU,s i , andA c to convert abstract categories into implementable sample-level prompts. Through this process, SafeIMG improves prompt specificity, scenario diversity, and risk relevance while maintaining category consistency, thereby laying the foundation for subsequent image generation. 2.4. Image Generation After constructing the prompts, SafeIMG uses GPT Image 2 to generate synthetic images for the defined safety scenarios. Given a complete prompt p i , the generation process is written as x i = G(p i ),(3) where G denotes the image-generation model and x i is the resulting image. To preserve the provenance of each image, we associatex i with its safety categoryc i , scenario instances i , generation prompt p i and media form m i . We define the resulting sample record as z i = (x i , c i , s i , p i , m i ).(4) The complete pool of generated candidates is then 풟 gen = z i N gen i=1 ,(5) whereN gen is the number of images produced before human screening. At this stage,풟 gen contains all generated candidates, including valid images, extremely difficult images and invalid generations. Each sample can be traced to its category, scenario specification and generation instruction. 2.5. Human Annotation Most AI-generated image detection datasets provide only image-level authenticity labels. Such labels indicate whether an image is synthetic but do not reveal the evidence supporting that judgement. SafeIMG therefore introduces evidence-level human annotation to record where an image appears suspicious, what type of anomaly is present and why the anomaly is implausible. A central feature of our annotation design is its scope. SafeIMG does not restrict annotation to familiar local defects, such as malformed text, faces, hands or lighting. Annotators are also instructed to examine higher-level inconsistencies that require an understanding of the depicted scene and the real world. These include implausible spatial relations, inconsistent event logic, violations of commonsense and conflicts with physical constraints. This distinction enables SafeIMG to assess whether detection models move beyond superficial generation artefacts towards evaluating the overall plausibility of visual content. The annotation protocol comprises two stages: sample triage and evidence-level annotation. Stage 1: Sample triage. Annotators first screen every sample z i ∈ 풟 gen and assign a triage label: q i =τ(z i ) ∈ ann, hard, invalid ,(6) whereτ(⋅) denotes the human-screening procedure. Samples are labelledinvalidwhen they substantially deviate from the intended category, omit the principal subject, contain obvious placeholder text or exhibit severe generation failures. These samples are excluded from subsequent annotation and evaluation. The discarded set is defined as: 풟 invalid = z i ∈ 풟 gen ∣ q i = invalid .(7) 8 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Samples are labelledhardwhen they are visually convincing and annotators cannot identify a defensible anomaly or suspicious region. These samples form the extremely difficult subset: 풟 hard = (z i ,∅) ∣ z i ∈ 풟 gen , q i = hard ,(8) where ∅ indicates that no defensible region-level anomaly was identified by the annotators. The remaining valid images contain at least one identifiable anomaly and are assigned the labelann. These images proceed to evidence-level annotation. This triage distinguishes poor-quality generation failures from highly realistic images without readily observable defects. It also ensures that the extremely difficult samples remain part of SafeIMG rather than being incorrectly discarded as annotation failures. Stage 2: Evidence-level annotation. For each sample assigned to the annotated branch, annotators examine both local appearance defects and the consistency of the overall scene. The annotation taxonomy contains two complementary groups. Local artefacts include abnormalities involving text, faces, hands, lighting and other directly observable visual details. High-level anomalies include commonsense conflicts, implausible object interactions, inconsistent event logic and violations of physical constraints. Each identified anomaly is documented using three components. Annotators first draw a bounding box around the relevant object or suspicious region. They then assign a fine-grained artefact label describing the anomaly type. Finally, they provide a concise natural-language explanation of why the selected content appears abnormal or incompatible with the depicted scene. For high-level anomalies, the explanation specifies the violated commonsense expectation, logical relation or physical constraint. An image may contain multiple independently annotated anomalies. The annotations associated with image x i are represented as 풜 i = a ij K i j=1 ,a ij = (b ij ,ℓ ij , e ij ) ,(9) whereK i is the number of annotated anomalies. Here,b ij denotes the bounding box,ℓ ij the artefact category and e ij the textual explanation. Finally, the annotated subset is defined as: 풟 ann = (z i , 풜 i ) ∣ z i ∈ 풟 gen , q i = ann, K i ≥ 1 .(10) This subset contains images with human-identified local or high-level anomalies and their corresponding evidence-level annotations. The final SafeIMG benchmark is the disjoint union of these two subsets: 풟 SafeIMG = 풟 ann ⊔ 풟 hard ,(11) where ⊔ denotes a disjoint union. The two subsets support complementary evaluation objectives.풟 ann enables evidence-level analysis of local artefacts and commonsense conflicts. By contrast,풟 hard evaluates detectors on highly realistic images for which human observers cannot specify a reliable visual anomaly. Together, they allow SafeIMG to assess both explainable detection and detection under the absence of readily observable human evidence. 3. Empirical Studies 3.1. Experimental Setup This section introduces the experimental setup on SafeIMG, including the evaluation data, evaluated models, and evaluation metrics. Our goal is to systematically evaluate the detection capability of existing methods in 9 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 0255075100 Images within category (%) P1 n=113 P2 n=104 P3 n=86 P4 n=91 P5 n=95 P6 n=99 P7 n=94 P8 n=89 I1 n=90 I2 n=91 I3 n=94 I4 n=85 101 12 (10.6%) 98 6 (5.8%) 83 3 (3.5%) 87 4 (4.4%) 75 20 (21.1%) 89 10 (10.1%) 84 10 (10.6%) 87 2 (2.2%) 81 9 (10.0%) 58 33 (36.3%) 78 16 (17.0%) 72 13 (15.3%) AnnotatedHard (a) Image composition by safety category. CommonsenseTextOtherFacePhysicsHandLighting P1 P2 P3 P4 P5 P6 P7 P8 I1 I2 I3 I4 Overall 41%23%18%5%8%4%2% 46%21%9%13%6%4%1% 66%19%5%6%4%0%1% 39%25%18%11%4%2%2% 63%19%8%3%3%3%0% 47%25%12%10%5%0%1% 60%16%6%15%1%2%0% 42%19%16%12%5%5%1% 49%23%18%2%2%4%2% 54%15%31%0%0%0%0% 76%9%1%13%0%0%0% 27%37%30%1%0%5%1% 50%21%13%8%4%3%1% Within-row share (%) 0>0–10>10–25>25–50>50–80 (b) Anomaly-label composition by safety category. 020406080 Explanation length (tokens) Overall n=2,652 P1 n=264 P2 n=311 P3 n=252 P4 n=270 P5 n=180 P6 n=241 P7 n=282 P8 n=258 I1 n=215 I2 n=91 I3 n=135 I4 n=153 Mean ± s.d.MeanMedianP95 to maximum All groups: 0% of explanations were ≥100 tokens. (c) Explanation length by safety category. 020406080 Explanation length (tokens) Commonsense n=1,325 Text n=559 Other n=353 Face n=223 Physics n=94 Hand n=72 Lighting n=26 Mean ± s.d.MeanMedianP95 to maximum All groups: 0% of explanations were ≥100 tokens. (d) Explanation length by anomaly type. Figure 4: SafeIMG evaluation data. (a) Category composition of the 1,131 synthetic images. Bars show the percentages assigned to the annotated (pale gold) and hard (green) subsets. Values give the corresponding image counts and hard-image rates. (b) Within-category distributions of the seven region-level anomaly labels. Cell values are rounded percentages, unrounded rows sum to 100%. (c,d) Explanation-length summaries by safety category and anomaly type, respectively. Rectangles show the mean±s.d., circles mark the means and vertical lines mark the medians. Horizontal segments span the 95th percentile to the maximum. P1–P8 denote public-safety categories and I1–I4 denote individual-safety categories. high-risk evidentiary image scenarios, and to further analyze whether the detection rationales produced by models are consistent with the visual issues annotated by humans. Evaluation data. SafeIMG comprises 1,131 synthetic images generated with GPT Image 2 and spanning 12 safety-sensitive categories: eight public-safety categories and four individual-safety categories (Fig. 4a). Human triage assigned 993 images (87.8%) to the annotated subset풟 ann and 138 (12.2%) to the hard subset 풟 hard . The hard-image fraction ranged from 2.2% in P8 to 36.3% in I2, revealing pronounced category-level variation in the visibility of local artefacts. The annotated subset yielded 2,652 region-level annotations, equivalent to 2.67 regions per annotated image and 2.34 regions per benchmark image. Category-level means ranged from 1.57 (I2) to 3.36 (P7) regions per annotated image and from 1.00 to 3.00 regions per benchmark image. Each annotation includes a bounding box, one of seven anomaly labels and a concise English explanation. Commonsense violations were most frequent (50%), followed by text artefacts (21%) and other local defects (13%). Faces (8%), physical constraints (4%), hands (3%) and lighting (1%) accounted for the remaining annotations (Fig. 4b). Explanations contained 15.50 tokens on average (median, 13; s.d., 8.58; 95th percentile, 31; maximum, 73). Category- and anomaly-specific distributions are summarised in Figs. 4c and 4d, respectively. We use풟 ann to quantify artefact prevalence, model coverage of human-identified anomalies and human–AI 10 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios explanation alignment. We evaluate풟 hard separately to test whether models recognise synthetic images when no defensible local anomaly can be identified. To estimate class-specific prediction bias, we further construct풟 real , a balanced reference set of 600 genuine images. The set contains 50 images in each of the 12 safety categories and spans source material dated from 2000 to 2026. All public-safety images (P1–P8) were sourced from Wikimedia Commons. For the individual-safety categories, I1 came from MEDIC (Alam et al., 2023), I2 from WildReceipt (Sun et al., 2021), and I3 from RICO (Deka et al., 2017). I4 were acquired through internet. It includes 348 landscape, 190 portrait and 62 square images, with a mean resolution of 1.287 megapixels. Evaluated Models. We evaluate two types of AI-generated image detection methods. The first type is general-purpose vision-language models (VLMs), including multimodal foundation models from GPT, Gemini, Claude, Qwen, Doubao, and others. These models possess strong visual understanding and natural-language explanation capabilities, and thus can be directly used for authenticity judgment and rationale generation. For these models, we use a unified prompt requiring the model to output an image label, a confidence score, and several independent visual rationales. The output label is restricted to eitherrealorai_generated; the confidence score is represented by discrete levels; and each rationale is required to be an independently checkable visual observation. The second type is specialized AI-generated image detection models. These methods usually directly output whether an image is AI-generated, or produce a detection score related to the probability of AI generation. For detectors that provide continuous scores, we convert them into binary labels according to their default thresholds or officially recommended settings, while preserving the original scores for subsequent analysis. Unlike general-purpose VLMs, specialized detectors usually do not provide natural-language explanations, and are therefore mainly used for image-level detection performance comparison. In addition to evaluating these two types of models, we conducted a human evaluation. We invited three human evaluators to independently perform the authenticity assessment without consulting one another. These evaluators were independent of the annotators who conducted the evidence-level annotation described in Section 2.5. We calculated the evaluation metric separately for each evaluator and report the arithmetic mean across the three evaluators. Evaluation metrics. SafeIMG evaluates detection models at both the image and evidence levels. Unless otherwise specified, generated-image detection is evaluated on the complete benchmark풟 SafeIMG . Fine- grained artefact and explanation analyses are restricted to풟 ann , because풟 hard does not contain region-level annotations. We additionally report performance on풟 hard to assess detection when no readily observable human cue is available. For a model M and a generated-image set 풟, we define the AI detection rate by 1 ∣풟∣ ∑ x i ∈풟 I [ ˆ y M (x i ) = ai_generated] ,(12) where ˆ y M (x i )denotes the predicted label andI[⋅]is the indicator function. Because every image in풟is synthetic, the AI rate is equivalent to recall for the AI-generated class. It measures generated-image sensitivity and should not be interpreted as complete binary-classification accuracy. 3.2. Overall Detection Performance We first report the overall image-level detection performance of different models on SafeIMG. This section provides an integrated overview of the detection capability of existing methods, including general-purpose VLMs, specialized detectors, and human evaluators, within the context of high-risk evidentiary image scenarios. The overall results, detailed in Table 2, reveal that current models and methods face clear limitations when identifying risk-oriented images synthesized by frontier generative models. 11 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Table 2: Detection accuracy (%) of VLMs, specialized detectors, and human evaluators on SafeIMG. Overall is weighted by category sample size. Within each model group, the highest value in each column is highlighted in blue and the second-highest value is highlighted in pink. ModelP1P2P3P4P5P6P7P8I1I2I3I4Overall VLM Base Models Claude Opus 4.630.4 42.3 29.1 29.7 33.0 16.2 17.046.112.253.878.725.932.6 Claude Sonnet 4.620.5 30.8 29.1 22.0 24.5 11.1 12.8 22.5 4.442.946.820.021.7 GPT 5.5 (reason effort = high) 6.2 13.5 23.3 12.1 12.6 4.04.34.51.1 11.0 8.50.08.9 GPT 5.4 (reason effort = high) 0.01.00.01.11.10.00.00.00.04.41.10.00.5 GPT 5.5 (no think)17.7 22.1 32.6 18.7 26.3 13.1 6.4 21.3 1.15.57.42.415.3 GPT 5.4 (no think)3.5 17.3 15.1 5.56.34.04.34.50.03.34.30.06.2 Doubao Seed 2.0 Pro38.1 52.947.7 38.541.122.2 45.7 39.3 24.4 34.1 40.442.438.2 Doubao Seed 2.0 Mini54.463.257.064.860.054.557.442.753.121.4 40.740.849.5 Gemini 3.1 Pro43.448.154.746.231.638.452.140.435.620.9 26.6 37.639.5 Qwen3.6 Flash24.3 21.4 32.6 22.0 16.0 16.2 14.9 13.5 4.4 21.1 19.4 10.818.1 Qwen3.6 Plus23.9 27.9 37.2 28.6 26.3 20.2 20.2 16.9 4.4 24.2 27.7 23.524.1 Specialized AI Detection Models CNNSpot9.76.78.117.612.611.1 1.12.22.24.40.00.07.3 FreDect33.620.241.942.937.941.418.133.722.245.138.341.233.1 LNP13.39.618.67.712.614.18.514.62.20.01.11.29.1 Human Evaluation Human83.281.180.282.477.578.882.386.993.370.791.171.881.7 At a high level, the best-performing general-purpose VLM achieves a detection accuracy of 49.5%, while the best-performing specialized image detector reaches only 33.1%. In stark contrast, human evaluators attain a significantly higher average accuracy of 81.7%, underscoring the substantial gap between current automated systems and human-level forensic judgment. Different VLMs exhibit substantial performance differences, with some models showing particularly poor discriminative ability for AI-generated images. This indicates that general visual understanding ability does not naturally translate into reliable AI-generated image detection capability. Furthermore, stronger model capability or more intensive reasoning settings do not always lead to consistent improvements. For instance, within the GPT series, the no-think setting occasionally outperforms the high-reasoning setting on several public safety categories. This suggests that VLM detection capability is highly model-dependent and cannot be simply predicted by scale or reasoning intensity. Existing specialized detectors show limited performance on SafeIMG due to a clear distribution shift problem. In particular, early methods like CNNSpot and LNP, which learned statistical artifacts , achieve only 7.3% and 9.1% accuracy, respectively, while FreDect performs relatively better at 33.1%. This indicates that frequency-domain cues retain some value, but relying on low-level traces is insufficient for images produced by newer generators like GPT Image 2. The performance of these detectors is also highly imbalanced across categories, suggesting that different evidentiary scenarios present distinct visual cues that no single detector can comprehensively cover. From a category-level perspective, the difficulty is not uniform. Text-intensive or interface-like images, such as private receipts (I2) and chat records (I3), are easier for some VLMs to identify due to rich textual and layout cues. In contrast, categories resembling real photographs or news scenes, such as personal accidents (I1), infrastructure accidents (P6), and crowd gatherings (P7), are more challenging for both VLMs and specialized detectors. This pattern is echoed in human evaluation, where complex backgrounds, occlusions, 12 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 0%10%20%30%40%50% AI-image accuracy 55% 60% 65% 70% 75% 80% 95% 100% Real-image accuracy Claude Opus 4.6 (32.6, 98.5) Claude Sonnet 4.6 (21.7, 98.7) GPT-5.5 (high) (8.9, 99.7) GPT-5.4 (high) (0.5, 100.0) GPT-5.5 (no think) (15.3, 100.0) GPT-5.4 (no think) (6.2, 100.0) Doubao Seed 2.0 Pro (38.2, 65.5) Doubao Seed 2.0 Mini (49.5, 60.5) Gemini 3.1 Pro (39.5, 78.7) Qwen3 Flash (18.1, 97.3) Qwen3 Plus (24.1, 94.8) Gemini (Google) Doubao (ByteDance) Claude (Anthropic) GPT (OpenAI) Qwen (Alibaba) Figure 5: Trade-off between AI-image accuracy and real-image accuracy across VLMs on SafeIMG. Points above the diagonal indicate higher AI than real accuracy, and vice versa. and crowds (e.g., in P6, P7, and P8) also increase the difficulty of manual recognition. Overall, the differences in performance across scenarios such as public safety images, private receipts, chat records, and identity endorsement further demonstrate that SafeIMG is not merely a test of identifying generic visual defects. Instead, it poses a comprehensive challenge involving scenario understanding, text/interface consistency, commonsense reasoning, and forensic sensitivity. The results validate the necessity of SafeIMG as a safety-oriented benchmark, highlighting that current automated systems are insufficient for high-stakes environments and that manual review, while more accurate, is also prone to error and should be complemented by automatic detection, cross-modal consistency verification, and source-credibility analysis. 3.3. Trade-off between AI and Real Image Accuracy Beyond overall detection accuracy, we further examine how different VLMs balance the identification of AI-generated images in풟 SafeIMG against the correct classification of real images in풟 real . Figure 5 plots each model’s accuracy on AI-generated images against its accuracy on real images, revealing a clear trade-off that separates the models into two distinct groups. The first group, comprising GPT-5.4, GPT-5.5, Claude Opus 4.6, Claude Sonnet 4.6, Qwen3 Flash, and Qwen3 Plus, exhibits a strong real-image bias. These models achieve extremely high real-image accuracy ranging from 94.8% to 100.0%, but their AI-image accuracy remains disappointingly low, with GPT-5.4 (reason effort = high) dropping to only 0.5%. This indicates that such models are heavily conservative in making "AI-generated" judgments, tending to classify almost all inputs as real to avoid false positives. While this behavior is desirable in applications where misclassifying real images as fake carries high costs, it renders these models nearly ineffective for AI-generated image detection in safety-critical evidentiary scenarios. The second group, including Doubao Seed 2.0 Mini, Doubao Seed 2.0 Pro, and Gemini 3.1 Pro, adopts a more balanced strategy. Doubao Seed 2.0 Mini achieves the highest AI-image accuracy among all VLMs 13 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 0x1x2x3x Perturbation Strength 0 5 10 15 20 25 30 35 40 AI-Image Accuracy (%) Claude Opus 4.6 Claude Sonnet 4.6 Gemini 3.1 Pro Preview Qwen 3.6 Plus (a) Other (170) Commonsense (166) Geometry & Structure (160) Hands (125) Physics (34) Text (284) Face (239) Texture & Detail (237) Lighting (207) 10.5% 10.2% 9.9% 7.7% 2.1% 17.5% 14.7% 14.6% 12.8% (b) Figure 6: (a). AI-image detection accuracy of representative VLMs under network propagation perturbations. (b). Distribution of model-generated rationales that are not supported by human annotations. at 49.5%, while maintaining a moderate real-image accuracy of 60.5%. Similarly, Doubao Seed 2.0 Pro and Gemini 3.1 Pro achieve AI-image accuracies of 38.2% and 39.5%, with real-image accuracies of 65.5% and 78.7%, respectively. These models demonstrate that improving AI-image detection does not necessarily require sacrificing real-image performance entirely, though a clear trade-off remains. Notably, no model in our evaluation achieves superior performance on both dimensions simultaneously. The GPT, Claude, and Qwen series prioritize real-image fidelity at the expense of detection sensitivity, while the Doubao and Gemini series lean toward higher detection rates at the cost of more frequent false alarms on real images. This divergence suggests that current VLMs have not yet converged on an optimal strategy for evidentiary image forensics, and the choice of model should depend on the specific risk tolerance of the deployment scenario. Overall, the above analysis reinforces our finding that general visual understanding ability does not naturally translate into reliable AI-generated image detection. The pronounced trade-off between AI and real image accuracy highlights a fundamental challenge for current VLMs, which struggle to distinguish highly realistic synthetic content from genuine evidence, particularly when high precision on real images is required. 3.4. Impact of Network Propagation on Detection Accuracy Beyond evaluating models on pristine AI-generated images, we further investigate how common network propagation perturbations affect detection performance in real-world social media scenarios. Figure 6a presents the detection accuracy of four VLMs under increasing levels of JPEG compression and image degradation, denoted as 0x, 1x, 2x, and 3x, where higher levels indicate more severe quality degradation. Notably, all models exhibit a consistent downward trend in detection accuracy as perturbation intensity increases, confirming that network propagation poses a serious challenge to AI-generated image detection. However, the rate of decline varies substantially across models, revealing distinct robustness characteristics. Gemini 3.1 Pro demonstrates the strongest overall robustness among the evaluated models, maintaining the highest accuracy across all perturbation levels. Its absolute performance remains superior throughout, though the rate of decline is moderate. Claude Opus 4.6 exhibits the steepest relative drop among all models, falling from 32.6% at 0x to merely 2.9% at 3x, a reduction of over 90% relative to its starting point. This sharp degradation implies that its detection capability relies heavily on high-frequency image details that are easily destroyed by compression. Qwen3 Plus, by comparison, shows the most gradual relative decline, decreasing from 24.1% to 11.6%, retaining nearly half of its initial accuracy even under severe perturbation. 14 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Claude Sonnet 4.6 performs the poorest across all settings, with accuracy plummeting to only 0.4% at 3x, indicating that certain VLMs are almost entirely dependent on low-level visual artifacts, making them highly vulnerable to common image transformations encountered during social media sharing. These results carry important practical implications. In real-world evidentiary scenarios, images are rarely preserved in their original quality, as they are frequently compressed, resized, and re-encoded during online dissemination. The substantial performance drop observed across all models underscores that current VLMs lack the robustness required for reliable deployment in forensic applications. Furthermore, the pronounced variation in degradation sensitivity across models suggests that robustness to network propagation should be considered a critical evaluation dimension in future benchmark development, alongside pristine-image detection accuracy. 3.5. Human-AI Alignment in Artifact Explanations In the previous sections, we analyzed the performance of detection models on SafeIMG images. We next examine whether model-generated rationales align with the evidence identified by human annotators. Because this analysis requires fine-grained reference annotations, all experiments in this section are conducted on풟 ann . We use a language model to compare the semantic content of human annotations with model- generated rationales and quantify their agreement. A match does not require identical wording, it is recorded when a model rationale identifies the same problem or anomaly described in a human annotation. AI- generated rationales evaluated in this experiment are produced by claude opus 4.6. And we use GPT-5.5 as the judge model for the Human–AI alignment analysis. First, from the perspective of whether human-annotated artifacts are covered by the model, the alignment between the model and humans is still limited. For each image, we calculate the proportion of human- annotated artifacts that are hit by the model’s reasons, and then average this score over the whole dataset. As shown in Table 3, Human Covered by AI is generally low across categories, with an average of only 29.8%. This indicates that even when the model can make image-level authenticity judgments to some extent, its explanations still fail to sufficiently cover the key artifacts annotated by humans. Across categories, I4 and P4 have relatively higher Human Covered by AI scores, reaching 48.6% and 49.9%, respectively. This suggests that in fabricated evidence and identity endorsement scenarios, as well as fire and explosion accidents scenarios, the model is relatively more likely to capture the anomalous cues that humans focus on. In contrast, I1, I2, and P5 have much lower coverage rates, at 19.1%, 21.4%, and 21.5%, respectively. This indicates larger discrepancies between model explanations and human annotations in personal emergencies, private receipts and transaction records, and pollution or hazardous material incidents. We further analyze why the consistency between model outputs and human annotations is low by examining which types of human-annotated artifacts are most often missed by the model. As shown in Figure 1b, the hit rate varies substantially across artifact types. The model achieves relatively higher coverage on Hands, Text, and Faces, with hit rates of 64.71%, 58.67%, and 56.25%, respectively. These issues usually have clear local visual manifestations. In contrast, the hit rates for Commonsense and Physics are much lower, at only 15.01% and 12.00%, respectively. This shows that the model is more likely to capture local and explicit visual artifacts, while current models remain limited in recognizing problems that require higher-level understanding, such as scene logic, commonsense conflicts, and physical inconsistencies. In other words, the model’s explanatory ability is still biased toward visible local defects and remains weaker in judging the overall plausibility of a scene. In addition, we analyze the reasons proposed by the model that are not supported by human annotations. These reasons are not necessarily incorrect judgments or hallucinations. On the contrary, they may include key artifacts overlooked by humans, reflecting differences between humans and AI in what they focus on 15 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Table 3: Statistics of Human Covered by AI and AI Supported by Human across categories. MetricP1 P2 P3 P4 P5 P6 P7 P8 I1 I2 I3 I4 Overall Human Covered by AI 34.1 27.3 26.7 49.9 21.5 34.4 28.6 22.9 19.1 21.4 30.1 48.6 29.8 AI Supported by Human 18.5 17.6 11.4 21.2 8.6 17.9 18.1 16.7 6.9 5.8 9.3 24.8 14.3 when judging images. Specifically, because the model’s reasons do not come with fixed category labels, and because its focus does not fully match the human annotation taxonomy, we do not directly apply the human artifact labels to these reasons. Instead, we use a language model to conduct semantic summarization and keyword extraction, and finally group them into several thematic categories. As shown in Figure 6b, the model reasons not supported by human annotations are mainly concentrated in themes such as Text, Face, Texture&Detail, Lighting, and Commonsense. Among them, Text accounts for the largest share, at 17.51%; Face and Texture&Detail account for 14.73% and 14.61%, respectively; Lighting accounts for 12.76%; and Commonsense accounts for only 10.23%. These results show that when explaining images, the model tends to actively inspect cues such as text, human features, and texture details. This is consistent with our earlier findings: the model tends to focus on explicit and local texture or detail-level features in images, while there remains substantial room for improvement in higher-level logical issues that are truly important in human annotations, such as physical anomalies. 3.6. Detection Accuracy on Challenging AI-Generated Images In this section, we further evaluate representative VLMs on풟 hard . This subset contains highly realistic synthetic images for which human annotators could not identify a defensible visual anomaly. It therefore provides a challenging test of model accuracy when readily observable generation artefacts are absent. As shown in Figure 1c, most models achieve substantially lower detection accuracy on풟 hard than on the complete SafeIMG benchmark. The accuracy of Doubao Seed 2.0 Pro decreases from 38.2% to 22.5%, corresponding to a relative reduction of 41.1%. Similarly, the accuracy of Qwen3.6 Flash falls from 18.1% to 10.1%, a relative reduction of 44.2%. These declines suggest that current detection capabilities depend partly on explicit visual cues that are unavailable in the hard subset. The subset therefore provides a more discriminative evaluation of highly realistic synthetic-image detection. Detection accuracy also varies considerably across safety categories. No model maintains consistently strong performance across all scenarios. Gemini 3.1 Pro Preview performs relatively well on I1, I4 and several public-safety categories, whereas Claude Sonnet 4.6 achieves higher accuracy on I2 and I3. These fragmented strengths indicate that current detection capabilities remain category-dependent. Performance gains within one type of visual content do not consistently transfer to other public- or individual-safety scenarios. Overall, the results reveal limited generalisability when obvious generation artefacts are unavailable, with model judgements remaining sensitive to image content and scenario characteristics. 3.7. Case Study We further analyze the differences between human and VLM judgments through representative cases. As shown in Figure 7, the four examples present the ground-truth label, human voting results, Claude’s judgment, and its detection rationales. The upper-left example is an AI-generated corporate group photo. Most human evaluators misclassified it as real, while Claude correctly identified it as AI-generated based on overly smooth faces, overly regular composition, and uniform lighting. The upper-right example is a real urban smoke scene. Human evaluators 16 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Human voting: AI-Generated (1 Real / 2 AI_Generated) Difficulty: HARD [Other]: Misshapen water bottle Human voting: Real (3 Real / 0 AI_Generated) Difficulty: NORMAL Model judgment: AI_Generated (confidence: 5) Model judgment: Real (confidence: 4) Several faces share an unusually similar smoothness and idealized appearance, with very uniform skin textures lacking natural variation → AI-generated faces. The overall composition has a stock-photo-like perfection with overly uniform lighting, symmetrical positioning, and generic corporate setting → AI-generated corporate imagery. The orchid plant on the left has leaves and flowers that appear somewhat simplified and lack fine botanical detail typical. The windows on the midground blue and white buildings, as well as those further back, are highly irregular in size, shape, and alignment, lacking logical architectural structure. Cables or wires crisscross over the foreground buildings in chaotic patterns, without clear or realistic attachment points. Structures on the roofs appear as blobby, undefined shapes rather than distinct mechanical or structural elements. Model judgment: AI_Generated (confidence: 4) Human voting: Real (2 Real / 1 AI_Generated) Difficulty: HARD Label: AI-GeneratedLabel: Real Image Human voting: Real (3 Real / 0 AI_Generated) Difficulty: HARD Label: AI-Generated Model judgment: Real (confidence: 4) The image shows standard UI elements including status bar and typical chat bubble formatting consistent with the real WeChat app. Label: AI-Generated Lighting is consistent throughout the scene - shelf LED lighting, overhead fixtures, and display backlight all create coherent shadows and highlights on the subjects. The showroom environment has realistic depth, with proper perspective on shelves, products, chairs, and background elements. Clothing wrinkles and fabric textures appear natural and physically consistent with body poses and gravity. The profile avatars show small cityscape and silhouette images → low-resolution thumbnails typical of real messaging apps. The text content follows a realistic conversational pattern of a job recruitment scenario, with specific address details and natural Chinese language formatting including proper punctuation. Figure 7: Representative cases comparing human evaluations with VLMs. The examples illustrate agreements and disagreements between human and model judgments, showing that humans may be misled by realistic synthetic images, while the model may either mistake imperfections in real images for synthetic artefacts or overlook localized anomalies in AI-generated images. correctly classified it as real, but Claude misclassified it as AI-generated with high confidence. Its rationale focuses on irregular building windows, chaotic cables, and blurry rooftop structures, suggesting that VLMs may mistake low-quality details or complex backgrounds for synthetic artifacts. The lower-left example is an AI-generated chat screenshot. Both humans and Claude misclassified it as real, indicating that text-based images can be highly deceptive when the interface layout, textual content, and conversational logic appear natural. The lower-right example is an AI-generated showroom scene. Humans correctly identified it based on key artifacts such as a misshapen water bottle, whereas Claude misclassified it as real because the overall lighting, perspective, and textures appeared consistent. Overall, these cases show that high-risk images challenge both humans and VLMs. Humans can be misled by highly realistic synthetic images, while VLMs may either over-rely on superficial local cues and produce false alarms, or overlook key artifacts when the overall scene appears coherent. 17 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 4. Related Work High-quality benchmarks are an important foundation for advancing methods for detecting AI-generated images. Early benchmarks were mostly built around GANs or CNN-based generators, focusing on whether detectors could distinguish real and generated images based on low-level statistical artifacts. For example, ForenSynths (Wang et al., 2020) provided an important experimental basis for cross-generator detection and promoted the development of subsequent general-purpose detectors. With the rise of new-generation text-to-image models such as diffusion models, benchmarks have begun to cover more complex generation sources and more realistic image content (Ho et al., 2020, Nichol et al., 2021, Rombach et al., 2022, Saharia et al., 2022). DE-FAKE (Sha et al., 2022) extends the detection task to models such as DALL-E 2 and Stable Diffusion, while considering both real/fake detection and source attribution. GenImage (Zhu et al., 2023) further constructs a million-scale dataset of real–generated image pairs, covering diverse visual content and multiple generators, and introduces evaluation protocols under cross-generator and image-degradation settings. These general-purpose benchmarks provide a large-scale data foundation for AI-generated image detection. Recent benchmarks have also paid increasing attention to the generalization challenges caused by shifts in image distributions. Comprehensive evaluation suites such as AIGCDetectBenchmark (Zhong et al., 2023) have been widely used to compare different detectors under multiple generators, categories, and post-processing conditions. AI-GenBench (Pellegrini et al., 2025) introduces a temporally evolving evaluation framework to simulate the continuous evolution of generative models from GANs to diffusion models and future models. Community Forensics (Park and Owens, 2024) samples images from a large number of open community models, emphasizing the importance of generator diversity in training data for open-world generalization. Together, these benchmarks show that AI-generated image detection should not be evaluated only on a static and closed set of generators, but must continuously confront new models, new post-processing pipelines, and new content distributions. Another line of development focuses on domain-specific or task-specific benchmarks. Scientific images, artistic images, anime images, and generative image editing scenarios have visual structures and risk contexts that differ from those of general natural images (Hu et al., 2026, Li and Stamp, 2025, Zhu et al., 2025, Chen et al., 2024b). For example, SciFigDetect (Hu et al., 2026) focuses on AI-generated scientific figures, emphasizing that scientific images often contain structured layouts, dense text, and alignment with academic semantics. Detecting AI-generated Artwork (Li and Stamp, 2025) and AnimeDL-2M (Zhu et al., 2025) focus on AI-generated detection in artworks and anime images, respectively. GIM (Chen et al., 2024b) targets generative image manipulation detection and localization, extending the task boundary from whole-image authenticity classification to local manipulation recognition. These domain-specific benchmarks demonstrate that general natural image detection cannot fully cover all high-value visual scenarios, and that different domains require tailored data construction and evaluation protocols. Overall, existing benchmarks have laid an important foundation for AI-generated image detection. However, they still fall short in sufficiently covering security-critical visual evidence scenarios. On the one hand, classic datasets rely on earlier generators and may not reflect the capability improvements of new-generation models (Brock et al., 2019, Dhariwal and Nichol, 2021, Nichol et al., 2021, Gu et al., 2022, Rombach et al., 2022). On the other hand, existing benchmarks are mostly oriented toward general images or specific tasks such as scientific and artistic content, with limited systematic investigation of high-risk evidentiary images (Vaccari and Chadwick, 2020, Huang et al., 2025). SafeIMG is constructed precisely around this gap, aiming to evaluate the capability boundaries of detectors in security-critical scenarios. 18 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios 5. Conclusion We present SafeIMG, a safety-oriented benchmark for AI-generated image detection across 12 public- and individual-safety scenarios. SafeIMG combines risk-aware image construction with human annotations of local artefacts, commonsense conflicts and physical inconsistencies, together with a challenging extremely difficult subset. Evaluations of vision–language models, specialised detectors and human observers show that current automated systems remain substantially less reliable than humans. Model rationales align poorly with human-identified evidence, particularly for high-level anomalies, and detection performance deteriorates after dissemination-induced image degradation. These findings show that reliable detection in safety-sensitive contexts requires scene-level reasoning, evidence-grounded explanations and propagation robustness beyond generic real-or-synthetic classification. We hope that SafeIMG can provide a foundation for future research on more reliable and interpretable AI-generated image detection methods. References Firoj Alam, Tanvirul Alam, Md Arid Hasan, Abul Hasnat, Muhammad Imran, and Ferda Ofli. Medic: a multi-task learning dataset for disaster image classification. Neural Computing and Applications, 35(3): 2609–2632, 2023. Anthropic. Introducing Claude Opus 4.6.https://w.anthropic.com/news/claude-opus-4-6, 2026. Shuai Bai, Yuxuan Cai, Ruizhe Chen, et al. Qwen3-VL technical report, 2025. URLhttps://arxiv.org/ abs/2511.21631. Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations, 2019. URLhttps://arxiv.org/ abs/1809.11096. Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024a. URL https://arxiv.org/abs/2403.04692. Sixiang Chen, Jinbin Bai, Zhuoran Zhao, Tian Ye, Qingyu Shi, Donghao Zhou, Wenhao Chai, Xin Lin, Jianzong Wu, Chao Tang, Shilin Xu, Tao Zhang, Haobo Yuan, Yikang Zhou, Wei Chow, Linfeng Li, Xiangtai Li, Lei Zhu, and Lu Qi. An empirical study of GPT-4o image generation capabilities, 2025. URL https://arxiv.org/abs/2504.05979. Yirui Chen, Xudong Huang, Quan Zhang, Wei Li, Mingjian Zhu, Qiangyu Yan, Simiao Li, Hanting Chen, Hailin Hu, Jie Yang, Wei Liu, and Jie Hu. GIM: A million-scale benchmark for generative image manipulation detection and localization, 2024b. URL https://arxiv.org/abs/2406.16531. Robert Chesney and Danielle Keats Citron. Deep fakes: A looming challenge for privacy, democracy, and national security. California Law Review, 107(6):1753–1820, 2019. URLhttps://doi.org/10.15779/ Z38RV0D15J. Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. Rico: A mobile app dataset for building data-driven design applications. In Proceedings of the 30th annual ACM symposium on user interface software and technology, pages 845–854, 2017. Prafulla Dhariwal and Alexander Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, volume 34, pages 8780–8794, 2021. 19 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis, 2024. URL https://arxiv.org/abs/2403.03206. Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. Leveraging frequency analysis for deep fake image recognition. In International conference on machine learning, pages 3247–3258. PMLR, 2020. Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, volume 30, 2017. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pages 6840–6851, 2020. You Hu, Chenzhuo Zhao, Changfa Mo, Haotian Liu, and Xiaobai Li. SciFigDetect: A benchmark for AI- generated scientific figure detection, 2026. URL https://arxiv.org/abs/2604.08211. Tai-Ming Huang, Wei-Tung Lin, Kai-Lung Hua, Wen-Huang Cheng, Junichi Yamagishi, and Jun-Cheng Chen. ThinkFake: Reasoning in multimodal large language models for AI-generated image detection, 2025. URL https://arxiv.org/abs/2509.19841. Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019. Meien Li and Mark Stamp. Detecting AI-generated artwork, 2025. URLhttps://arxiv.org/abs/2504. 07078. Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. In International Conference on Learning Representations, 2023. URLhttps:// arxiv.org/abs/2210.02747. Bo Liu, Fan Yang, Xiuli Bi, Bin Xiao, Weisheng Li, and Xinbo Gao. Detecting generated images by real images. In European conference on computer vision, pages 95–110. Springer, 2022. Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models, 2021. URL https://arxiv.org/abs/2112.10741. Sophie J. Nightingale and Hany Farid. AI-synthesized faces are indistinguishable from real faces and more trustworthy. Proceedings of the National Academy of Sciences, 119(8):e2120481119, 2022. doi: 10.1073/pnas.2120481119. Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24480–24489, 2023. OpenAI. GPT-5 system card. https://openai.com/index/gpt-5-system-card/, 2025. 20 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios OpenAI. Introducing GPT-5.5. https://openai.com/index/introducing-gpt-5-5/, 2026. Jeongsoo Park and Andrew Owens. Community forensics: Using thousands of generators to train fake image detectors, 2024. URL https://arxiv.org/abs/2411.04125. Lorenzo Pellegrini, Davide Cozzolino, Serafino Pandolfini, Davide Maltoni, Matteo Ferrara, Luisa Verdoliva, Marco Prati, and Marco Ramilli. AI-GenBench: A new ongoing benchmark for AI-generated image detection, 2025. URL https://arxiv.org/abs/2504.20865. Qwen Team. Qwen3.5-Omni technical report, 2026. URL https://arxiv.org/abs/2604.15804. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, S. Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, pages 36479–36494, 2022. Zeyang Sha, Zheng Li, Ning Yu, and Yang Zhang. DE-FAKE: Detection and attribution of fake images generated by text-to-image generation models, 2022. URL https://arxiv.org/abs/2210.06998. Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021. URL https://arxiv.org/abs/2011.13456. Hongbin Sun, Zhanghui Kuang, Xiaoyu Yue, Chenhao Lin, and Wayne Zhang. Spatial dual-modality graph reasoning for key information extraction. arXiv preprint arXiv:2103.14470, 2021. Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie. AnyText: Multilingual visual text generation and editing, 2023. URL https://arxiv.org/abs/2311.03054. Cristian Vaccari and Andrew Chadwick. Deepfakes and disinformation: Exploring the impact of synthetic po- litical video on deception, uncertainty, and trust in news. Social Media + Society, 6(1):2056305120903408, 2020. doi: 10.1177/2056305120903408. Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. CNN-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8695–8704, 2020. Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. DIRE for diffusion-generated image detection, 2023. URL https://arxiv.org/abs/2303.09295. Shilin Yan, Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, and Weidi Xie. A sanity check for AI-generated image detection, 2024. URL https://arxiv.org/abs/2406.19435. Zhiyuan Yan, Junyan Ye, Weijia Li, Zilong Huang, Shenghai Yuan, Xiangyang He, Kaiqing Lin, Jun He, Conghui He, and Li Yuan. GPT-ImgEval: A comprehensive benchmark for diagnosing GPT4o in image generation, 2025. URL https://arxiv.org/abs/2504.02782. Michael Yang, Shijian Deng, William T. Doan, Kai Wang, Tianyu Yang, Harsh Singh, and Yapeng Tian. Explainable AI-generated image detection rewardbench, 2025. URLhttps://arxiv.org/abs/2511. 12363. 21 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. Nan Zhong, Yiran Xu, Sheng Li, Zhenxing Qian, and Xinpeng Zhang. PatchCraft: Exploring texture patch for efficient AI-generated image detection, 2023. URL https://arxiv.org/abs/2311.12397. Chenyang Zhu, Xing Zhang, Yuyang Sun, Ching-Chun Chang, and Isao Echizen. AnimeDL-2M: Million-scale AI-generated anime image detection and localization in diffusion era, 2025. URLhttps://arxiv.org/ abs/2504.11015. Mingjian Zhu, Hanting Chen, Qiangyu Yan, Xudong Huang, Guanyu Lin, Wei Li, Zhijun Tu, Hailin Hu, Jie Hu, and Yunhe Wang. GenImage: A million-scale benchmark for detecting AI-generated image, 2023. URL https://arxiv.org/abs/2306.08571. 22 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Commonsense Hands Faces Lighting Fonts Physics Other I1I2I3I4Individual Commonsense Hands Faces Lighting Fonts Physics Other P1P2P3P4Public -20%-10%0+10%+20% Commonsense Hands Faces Lighting Fonts Physics Other P5 -20%-10%0+10%+20% P6 -20%-10%0+10%+20% P7 -20%-10%0+10%+20% P8 -20%-10%0+10%+20% Overall -20% -10% 0 +10% +20% Prevalence Difference Figure 8: Prevalence difference of artifact labels between false negatives and all AI-generated images. Individual, Public, and Overall denote individual-safety, public-safety, and full-dataset aggregates. Crosses “×” indicate labels absent in both compared sets. Appendix A. Detection Error Analysis on Base VLMs To analyze the failure patterns of base VLMs, we compare the prevalence of each human-annotated artifact label in false negatives with its prevalence among all AI-generated images. The statistic is computed at the image level: if the same label appears multiple times in one image, it is counted only once. ∆(l) = P(l in false negative∣false negative)− P(l in all AI-generated images∣all AI-generated images) (13) As shown in Figure 8, the prevalence differences are relatively small, with all label-level changes below 5%. This suggests that false negatives are not dominated by any single artifact type; instead, their artifact distribution is broadly similar to that of the overall AI-generated image set. However, category-level results reveal more localized shifts. For example, in personal-information images, “Text” and “Other” become relatively more common among false negatives. In some public-scene categories, labels such as “Physics”, “Other”, and “Text” also show positive shifts. This indicates that detection failures do not follow a single global pattern, but are instead related to category-specific artifacts. In general, this analysis suggests that base VLM detection errors are not strongly biased toward one universal artifact label, although certain artifact types may become more prominent within specific image categories. 23 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios B. Prompt Examples for Image Generation Prompt Example on P1 - Natural Disasters On the afternoon of May 11, 2024, a storm surge strikes a riverside walkway in Los Angeles. River water has overflowed the waterfront platform and submerged benches and the bases of streetlights. Warning tape is blown sideways by the wind. Staff members are persuading two citizens to leave. The muddy water surface carries branches, plastic road barriers, and broken wooden boards. Distant high-rise buildings are reflected in the yellowish floodwater. In the bottom-right corner, there must be a small, natural camera-auto date stamp reading “2024-05-11,” in an orange transparent seven-segment digital style. Realistic documentary photography of storm-surge disaster damage. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P2 - Social Unrest, Public Violence, and Attacks On the afternoon of May 6, 2024, a labor dispute triggers a group protest outside a large commercial plaza in Chaoyang District, Beijing. Dozens of citizens gather at the plaza entrance holding handwritten signs. Some barriers have been pushed over. Police and security personnel set up warning lines and guide bystanders away. Flyers, plastic water bottles, and fallen traffic cones are scattered on the ground. The incident occurs after negotiations break down and emotions escalate. In the bottom-right corner, include a small natural orange transparent seven-segment camera-auto date stamp reading “2024-05-06.” Realistic documentary photography style. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P3 - Transportation Accidents On the morning of May 5, 2024, during the rainy morning rush hour on Beijing’s East Fifth Ring Road, a wet road surface causes a multi-vehicle rear-end collision. Three sedans and a box truck are crushed together. Front ends are deformed, with broken glass and bumpers scattered across the lane. Traffic police close two lanes, firefighters check whether anyone is trapped inside the vehicles, and distant cars queue in slow traffic. In the bottom-right corner, include a small natural orange transparent seven-segment camera-auto date stamp reading “2024-05-05.” Realistic documentary photography of a traffic accident. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P4 - Fires, Explosions, and Energy Facility Accidents On the morning of May 13, 2024, a gas explosion occurs at a ground-floor shop beneath an old residential building in Yuexiu District, Guangzhou. The roll-up door facing the street is deformed, and glass fragments are scattered on the sidewalk. Firefighters inspect the upstairs residents and help evacuate an elderly person. Gas-company repair workers shut off a roadside valve. The accident is suspected to have been caused by a restaurant gas leak ignited by an open flame. In the bottom-right corner, include a small natural orange transparent seven-segment date stamp reading “2024-05-13.” Realistic documentary photography of the aftermath of an urban gas explosion. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. 24 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Prompt Example on P5 - Pollution, Hazardous-Material Leaks, and Nuclear/Radiation Incidents On the morning of May 5, 2024, abnormal blue-green wastewater appears in a drainage canal at an industrial park in Tongzhou District, Beijing. Environmental law-enforcement officers and emergency responders take samples beside the canal. Warning tape is set up along the banks, and nearby residents are persuaded to leave by community staff. Foam floats on the water surface and a pungent odor is present. The incident is suspected to have been caused by illegal industrial wastewater discharge from a company in the park. In the bottom-right corner, include a small natural orange transparent seven-segment date stamp reading “2024-05-05.” Realistic documentary photography of a pollution emergency response. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P6 - Building, Structural, and Civil-Infrastructure Accidents On the morning of May 5, 2024, a partial collapse occurs at an old residential building in Haidian District, Beijing. An entire exterior wall and two floors of balconies have collapsed onto the road below. Concrete blocks, steel bars, and window frames cover the sidewalk. Firefighters set up life-detection equipment beside the debris, while community staff organize residents to evacuate. The accident is suspected to have been caused by long-term structural aging and unauthorized renovation. In the bottom-right corner, include a small natural orange transparent seven-segment date stamp reading “2024-05-05.” Realistic documentary photography of a major building-collapse accident. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P7 - Crowd Gathering and Venue Safety Accidents On the morning of May 5, 2024, a crowd crush occurs at the entrance of a large shopping mall promotion event in Chaoyang District, Beijing. The crowd in front of the glass doors becomes uncontrollably congested. Several sections of metal barriers are crushed down. Shoes, shopping bags, and broken barrier tape are scattered on the ground. Security staff and firefighters are opening an emergency rescue lane, while medical workers treat fallen people outside the warning line. The accident is caused by a limited-product sale, an overly narrow entrance passage, and a sudden surge of visitors. In the bottom-right corner, include a small natural orange transparent seven-segment date stamp reading “2024-05-05.” Realistic documentary photography of a major crowd-crush accident. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. Prompt Example on P8 - Public Health and Biosecurity Events On the afternoon of May 9, 2024, the emergency hall of a large hospital in Pudong, Shanghai becomes crowded due to a concentrated surge of respiratory infection patients. The waiting area is filled with masked patients. Nurses maintain order in front of the triage desk. Temporary fever-clinic signs and mobile testing equipment are clearly visible. The situation is caused by overlapping community-cluster infections and peak medical demand. In the bottom-right corner, include a small natural orange transparent seven-segment date stamp reading “2024-05-09.” Realistic documentary photography of hospital overcrowding during a public-health incident. Do not include “someone”, “x”, or any other placeholder text anywhere in the image; use normal names. 25 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Prompt Example on I1 - Personal Accidents and Emergency Crises A highly realistic documentary-style photograph of a minor traffic accident at an urban intersection at night. A private car is stopped beside a crosswalk with its front bumper damaged. An electric-scooter rider sits by the roadside waiting for an ambulance, while traffic police place reflective cones. A spilled delivery box and helmet lie on the ground. The scene feels tense but non-graphic. All names must be normal. All IDs and numbers must be completely random. Do not include placeholders such as “someone”, “x”, “1234”, “8888”, “6666”, or similar content. Prompt Example on I2 - Receipts and Transaction Proofs A realistic mobile subscription order screenshot. The page shows a monthly enterprise software subscription purchased on 2024.10.31, with order number SA-70593162, subscription period from 2024.11.01 to 2024.11.30, payment amount 299.00, and invoice status “available to request.” All names must be normal. All IDs and numbers must be completely random. Do not include placeholders such as “someone”, “x”, “1234”, “8888”, “6666”, or similar content. The generated image must be a vertical standard canvas, with a complete 9:19.5 vertical mobile screenshot area centered in the image. The area outside the screenshot must remain a solid blank background to fit the canvas size. Do not stretch, compress, or crop the interface content. Prompt Example on I3 - Personal Chats and Communication Records A realistic mobile chat or communication-record screenshot. The page shows a mobile email app interface, with sender richard@rivermail.com, recipient service@besttravel.com, date 2025.03.18, and email subject “Inquiry About Flight Ticket Refund Progress.” The content involves refund number RF-58204179, refund amount 2,860.00, and processing deadline. The conversation content should be limited to 6–10 messages to avoid abnormal layout caused by excessive length. All names must be normal. All IDs and numbers must be completely random. Do not include placeholders such as “someone”, “x”, “1234”, “8888”, “6666”, or similar content. The generated image must be a vertical standard canvas, with a complete 9:19.5 vertical mobile screenshot area centered in the image. The area outside the screenshot must remain a solid blank background. Do not stretch, compress, or crop the interface content. Prompt Example on I4 - Fabricated Scene Proofs and Identity Endorsements A highly realistic business-scene photograph. The person sits at the main executive desk in a high-rise office in Tokyo Bay. Behind them are floor-to-ceiling windows with a sea view and city skyline. On the desk are a business plan, signing pen, laptop, and an opened financing proposal. The overall scene creates the impression of a technology-company founder handling major business matters as identity endorsement. All names must be normal. All IDs and numbers must be completely random. Do not include placeholders such as “someone”, “x”, “1234”, “8888”, “6666”, or similar content. 26 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios C. Prompts for Evaluation Prompt for Evaluation You are a visual forensics assistant. Given one image, decide whether it is a real photo or AI-generated. Then list the visual evidence as separate atomic reasons. Output only strict JSON in this schema: "label":"real|ai_generated","confidence":1,"reasons":["..."]. Label must be one of: real, ai_generated. Confidence must be an integer from 1 to 5, where 5 means highly confident. Reasons must be a list of strings. Please decide whether this image is a real photo or AI-generated. Return strict JSON only, with fields label, confidence, and reasons. Reason-writing rules: Each item in reasons must be one atomic visual claim, independently checkable in the image, and no more than 60 words. Do not combine unrelated artifacts in one item. Do not duplicate the same artifact. If multiple similar artifacts appear, summarize them as one grouped item. Use concrete visual evidence: faces, hands, text, object boundaries, lighting, shadows, reflections, perspective, anatomy, object interactions, repeated patterns, or common-sense consistency. Avoid vague phrases such as looks AI-generated, looks weird, or low quality unless paired with concrete evidence. If no obvious artifact is visible, give 1-3 reasons supporting realism. If artifacts are visible, list the most important ones first. The reasons list will later be compared against human annotations, so write each reason as a clear, self-contained issue description. D. Case Studies on SafeIMG Examples Figure 9: Case study on P1 - Natural Disasters. The broadcast depicts a severe Harbin blizzard in July, an implausible seasonal event, and reports the temperature range in reversed order. A malicious actor could use such fabricated news imagery to exaggerate a disaster, provoke public panic, or distort travel and emergency-response decisions. 27 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Figure 10: Case study on P2 - Social Unrest, Public Violence, and Attack Incidents. The scene lacks a coherent collision history: the vehicle has no visible impact counterpart, a streetlight is unsupported, and the crowd remains panicked despite police control. A malicious actor could present it as a vehicle attack to incite fear, inflame tensions, or trigger an unwarranted security response. Figure 11: Case study on P3 - Transportation Accidents. The report attributes the crash to tire failure although the visible damage is concentrated on the vehicle body, and the truck’s direction of travel conflicts with the road sign. A malicious actor could use the image to fabricate an accident narrative, falsely assign responsibility, or mislead emergency and traffic-management decisions. Figure 12: Case study on P4 - Fires, Explosions, and Energy Facility Accidents. Bystanders stand inside the police cordon, while the elevated firefighter is too far from the fire to intervene, contradicting incident-response logic. A malicious actor could present it as a major urban fire to cause panic, misrepresent responders, or divert emergency resources. 28 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Figure 13: Case study on P5 - Pollution and Hazardous Material Accidents. The overhead sign and roadside arrow direct traffic in opposite directions, an implausible arrangement during a coordinated hazardous-material response. A malicious actor could use the image to fabricate a chemical spill, provoke evacuation or traffic disruption, or trigger unwarranted public and regulatory action. Figure 14: Case study on P6 - Building and Civil Infrastructure Accidents. The sinkhole scene violates road-safety logic: one side remains accessible without a cordon, and a traffic light stands in the roadway. A malicious actor could circulate it as false evidence of infrastructure failure to create panic, disrupt travel, or manipulate liability claims. Figure 15: Case study on P7 - Crowd Gathering and Venue Safety Accidents. The nighttime scene conflicts with the stated time and season in Xinjiang, while the schedule repeats one program as both current and upcoming. A malicious actor could use it to fabricate a venue emergency, exaggerate crowd-control failures, or inflame public disorder. 29 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Figure 16: Case study on P8 - Public Health and Biosecurity Events. The triage setup is operationally implausible: the desk sits outside the tent, the thermometer is too small to function as shown, and the notice repeats its date. A malicious actor could use it to invent an outbreak response, spread health misinformation, or undermine trust in public-health institutions. Figure 17: Case study on I1 - Personal Accidents and Emergencies. The post-fire kitchen is causally inconsistent: power remains on, while heat-sensitive bottles, bags, and food survive the surrounding fire damage. A malicious actor could use it as false evidence of an emergency to solicit money, support insurance fraud, or shift liability. Figure 18: Case study on I2 - Private Receipts and Transaction Records. The banking record violates the interface’s expected chronology, placing earlier transactions above later ones in a supposedly ordered ledger. A malicious actor could use the fabricated record to claim a payment or refund that never occurred, enabling fraud, false reimbursement, or manipulation of a financial dispute. 30 Seeing Is No Longer Believing: AI-generated Images Challenge Visual Trust in High-risk Scenarios Figure 19: Case study on I3 - Personal Chat and Communication Records. The purported bank-message screenshot conflicts with the institution’s real identity and normal messaging conventions: the logo is incorrect, and timestamps appear in an implausible location. A malicious actor could use it to impersonate a bank, fabricate account activity, or lend credibility to phishing and payment scams. Figure 20: Case study on I4 - Fabricated Scene Evidence and Identity Endorsement. The displayed corporate timeline is not chronological, creating a temporal inconsistency that undermines the claimed business setting. A malicious actor could use the image to fabricate affiliation or endorsement, support investment fraud, or damage reputations. 31