Paper deep dive
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, Jianfu Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/23/2026, 1:45:57 AM
Summary
The paper introduces Defake-o3, an explainable AI-generated image (AIGI) detector that moves beyond speculative rationales to provide verifiable evidence. It combines interactive visual search (iterative zooming) with verifier-guided evidence alignment using reinforcement learning. The authors also introduce GroundFake, a dataset for grounded explainable detection, and FakeFrontier, an out-of-distribution benchmark for evaluating evidence quality.
Entities (7)
Relation Signals (6)
Defake-o3 → evaluatedon → FakeFrontier
confidence 95% · Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves
Defake-o3 → trainedon → GroundFake
confidence 95% · To support this objective, we construct GroundFake... To support this objective, we construct GroundFake... supervised fine-tuning
Defake-o3 → uses → interactive visual search
confidence 95% · It combines interactive visual search with verifier-guided evidence alignment
Defake-o3 → uses → Evidence Verifier
confidence 95% · an Evidence Verifier... provides reinforcement learning rewards
GroundFake → contains → localized bounding-box evidence
confidence 90% · with localized bounding-box evidence, human verification based on visual grounding
FakeFrontier → contains → 10 recent generators
confidence 90% · built from real images and outputs of 10 recent generators
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.
Tags
Links
- Source: https://arxiv.org/abs/2608.16259v1
- Canonical: https://arxiv.org/abs/2608.16259v1
Trouble viewing inline? Open PDF directly →
Full Text
116,432 characters extracted from source content.
Expand or collapse full text
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection Bowen Deng ∗ des_wv@sjtu.edu.cn Shanghai Jiao Tong University Shanghai, China Jiahui Zhan jiahuizhan@sjtu.edu.cn Shanghai Jiao Tong University Shanghai, China Yikun Ji da-kun@sjtu.edu.cn Shanghai Jiao Tong University Shanghai, China Haozhen Yan orion810@sjtu.edu.cn Shanghai Jiao Tong University Shanghai, China Jianfu Zhang ∗† c.sis@sjtu.edu.cn Shanghai Jiao Tong University Shanghai, China Abstract The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explain- able and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate specula- tive rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evi- dence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verifica- tion based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 im- proves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence. 1 Introduction The rapid progress of image generation models, including autore- gressive text-to-image systems [10,45,55,66] and diffusion-based approaches [4,7,16,21,36,46,47,51,60], has made synthetic im- ages increasingly photorealistic and harder to distinguish from authentic photographs. Although these advances benefit creative applications, they also create serious risks for public trust by en- abling the spread of visually convincing misinformation at scale. As a result, robust and trustworthy AI-generated image (AIGI) detec- tion has become increasingly important. Traditional AIGI detection methods typically cast the task as binary classification. Built on Convolutional Neural Networks (CNNs) [42] and Vision Transform- ers [15], these data-driven detectors have achieved strong detection performance [9,12,52,57,58,64], yet they remain fundamentally black-box models. In most cases, they output only an authenticity score without revealing the visual evidence behind the decision. As ∗ Also with Ant Group. † Corresponding author. User Input Defake-o3 (Ours) Final Output of Defake-o3 Was this image generated by artificial intelligence? Please provide your verdict and reasoning. The image displays a large stadium, likely Estadio BBVA, with a cityscape and mountains in the background at sunset... I am zooming in on the stadium to check the architectural details and the 'BBVA' logo. "bbox": [150, 400, 780, 700] ... I noticed some text-like features on the right facade of the stadium. I will zoom in to verify if the text is legible or if it consists of pseudo-characters, which is a common sign of AI generation ... "bbox": [620, 505, 735, 560] Upon closer inspection, the characters are completely unrecognizable gibberish, a classic hallmark of AI image generation ... . . . . . . ... The moving vehicles are rendered as stretched, formless blobs rather than actual cars ... I have gathered enough evidence to deliver my verdict. Observing the highway spanning the lower portion of the image...I will now zoom in on the right side of the highway overpassfor a final, closer inspection. "bbox": [400, 750, 950, 920] Global Evidence: The overall image has a ‘painterlyÕ or overly smooth texture common in AI generations, particularly noticeable in the vegetation and distant buildings, which lack photographic noise and sharp definition. Local Evidence: 1. bbox: [620, 505, 735, 560]: The smaller signage on the right side of the stadium's facade consists of garbled, illegible squiggles that imitate text but form no coherent words. 2. bbox: [400, 750, 950, 920]: The vehicles on the highway are significantly distorted. They appear as blurry, oblong blobs without distinct wheels, windows, or headlights... Existing MLLM Methods Verdict: Fake This is a fake image. The stadiumÕs roofline has an overly smooth, melted curvature... The mountain range behind the city has unnaturally sharp, symmetrical peaks and lacks realistic atmospheric haze... Iterative visual exploration Grounded verifiable evidence Reliable predictions Single-pass visual inspection Vague or hallucinated rationales Unreliable predictions Figure 1: Overview of Defake-o3. Compared with existing single-pass MLLM detectors, Defake-o3 iteratively zooms into suspicious regions to verify candidate artifacts and pro- duces a structured verdict supported by global evidence and localized key evidence. a result, their predictions are difficult to verify, which limits their trustworthiness and practical deployment in high-stakes settings. Recent studies have explored Multimodal Large Language Mod- els (MLLMs) [2, 3, 5, 14, 32, 35] for explainable AIGI detection [26, 33,59,63,65], producing natural-language rationales alongside authenticity predictions. However, these rationales are often spec- ulative rather than supported by verifiable evidence. First, they are frequently ungrounded, relying on vague, generic, or hallu- cinated cues instead of precise visual artifacts, with no explicit criterion for valid evidence. Second, existing methods generally perform single-pass, coarse-grained inspection. As modern genera- tors leave increasingly subtle and localized flaws, standard open- source MLLMs, constrained by input resolution and single-pass visual encoding, often fail to examine suspicious regions in suffi- cient detail. Consequently, their explanations may appear plausible arXiv:2608.16259v1 [cs.CV] 17 Aug 2026 Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang while remaining weakly grounded, poorly localized, and difficult to verify. To address these limitations, we propose Defake-o3, an ex- plainable AIGI detector that moves from speculative rationales to verifiable evidence through two complementary components: in- teractive visual search, which determines where and at what granularity to inspect, and verifier-guided evidence alignment, which defines what constitutes valid evidence. Defake-o3 adopts an iterative “thinking-with-images” mechanism [56,61,69], using a zoom-in tool to crop and magnify suspicious regions and uncover subtle artifacts missed in the full image (Fig. 1). To support this ob- jective, we construct GroundFake, a debiased training dataset that controls image format, aspect ratio, semantic category, and aesthetic quality to reduce shortcut bias. It provides localized bounding-box evidence under strict human verification based on visual grounding and artifact specificity, together with corrected reasoning trajecto- ries and valid/invalid evidence labels. These annotations support both supervised fine-tuning and the training of an Evidence Veri- fier, which serves as the reward model [13] during GRPO-based reinforcement learning [48,50,67] to reward grounded evidence and penalize baseless claims. For evaluation on recent generators, we introduce FakeFrontier, an out-of-distribution benchmark con- taining 2,000 real images and 2,000 synthetic images generated by 10 recent models, including Nano Banana [20] and Seedream 4.5 [6]. We also develop an MLLM-based protocol to evaluate both clas- sification performance and the quality and persuasiveness of the produced evidence. In summary, our contributions are threefold: •We propose Defake-o3, an explainable AIGI detector that com- bines interactive visual search with verifier-guided evidence alignment. Iterative zoom-in inspection reveals subtle artifacts, while the Evidence Verifier guides reinforcement learning toward grounded evidence and away from baseless claims. • We introduce GroundFake, a debiased training dataset for grounded explainable AIGI detection. It provides localized evidence verified according to explicit visual-grounding and artifact-specificity cri- teria, together with corrected reasoning trajectories and valid/in- valid evidence labels for supervised fine-tuning and verifier train- ing. •We present FakeFrontier, an out-of-distribution benchmark containing real images and outputs from 10 recent generators, along with an MLLM-based protocol for evaluating both detec- tion performance and evidence quality. 2 Related Work Conventional Black-Box AIGI Detection: A dominant line of AIGI detection formulates the task as binary classification. CNNSpot [57] uses data augmentation to improve cross-model generalization of a standard ResNet detector. DIRE [58] and DRCT [9] exploit recon- struction discrepancies revealed by pre-trained diffusion models. NPR [52] identifies synthetic images through local pixel statistics introduced by upsampling, while AIDE [64] combines frequency and semantic cues in a dual-stream architecture. Recent studies [12] further suggest that detector generalization is also sensitive to dis- tributional bias between real and synthetic data. These methods have established strong classification baselines for AIGI detection, but they remain fundamentally black-box detectors. In practice, they typically output only an authenticity score, without revealing which visual cues support the prediction or providing localized, visually verifiable evidence that humans can inspect. Consequently, their decisions are difficult to interpret and verify, which motivates recent efforts toward explainable AIGI detection. Explainable Detection via MLLMs: To move beyond score-only black-box detectors, recent studies have extended Multimodal Large Language Models (MLLMs) to explainable AIGI detection, while related works also consider broader image or video forgery set- tings. Instruction-tuning-based methods such as FakeVLM [59] and FakeScope [34] construct large-scale visual instruction data so that MLLMs can jointly predict authenticity and generate textual expla- nations. IVY-FAKE [25] further unifies explainable detection across images and videos. To improve spatial grounding, methods such as LEGION [26] and FakeShield [63] augment MLLMs with artifact lo- calization, often by leveraging external models such as SAM [27] to predict masks over suspicious artifact regions. More recently, align- ment strategies based on preference optimization or reinforcement learning have been explored to improve explanation quality. AIGI- Holmes [70] adopts DPO [44] to suppress low-quality explanations, but its outputs remain coarse-grained without explicit localization. So-Fake [23] and FakeXplain [24] further combine explanation generation with bounding-box grounding under RL-style training. However, current methods still mainly optimize the similarity of the explanation with reference annotations, rather than explicitly learn- ing whether a claimed piece of evidence is visually grounded and artifact-specific enough to support an AIGI verdict. As a result, even localized outputs can remain vague, spurious, or hallucinated, and existing resources still provide limited evidence-level supervision for training and evaluating truly verifiable explanations. Dynamic Visual Reasoning and Exploration: Recent advances in MLLM-based visual reasoning have shifted from fixed, text- dominant chain-of-thought toward interactive visual exploration, where models actively decide where to inspect and at what gran- ularity. Early works such as Visual CoT [49] and LLaVA-CoT [62] made reasoning more explicit through region grounding or staged decomposition, but still relied largely on static visual inputs. Later studies treated visual search as a core reasoning primitive: V* [61] introduced LLM-guided search for high-resolution detail ground- ing, Visual-RFT [37] showed that verifiable rewards can improve visual reasoning behavior, and Pixel Reasoner [56], DeepEyes [69], and Mini-o3 [30] further enabled models to zoom, inspect, and it- eratively interact with images during inference. These advances are especially relevant to explainable AIGI detection, where recent generators leave increasingly subtle and localized artifacts that are hard to capture with single-pass visual encoding. However, they mainly improve where to look, rather than what should count as valid evidence for an AIGI verdict. Therefore, interactive visual search alone does not guarantee grounded, artifact-specific, and visually verifiable explanations. 3 GroundFake Dataset Construction To support grounded explainable AIGI detection, we construct GroundFake, a training dataset containing 16,000 images, stepwise reasoning transcripts, and evidence annotations. For fake images, Defake-o3 Fake Images SDXL FLUX Midjourney V5 Caption Image agemI Automated Annotation Evidence Candidates: Global: The image exhibits a hyper-realistic.... Local: 1: bbox: [779,565,861,626], There is a distorted, unrecognizable wooden object.... resembling a hallucinated artifact . 2: bbox: [600,250,1000,550], The staircase railing lacks...; the glass panels seem to... the handrail transition is... . 3: bbox: [0,320,150,500], The leaves of the plant... ...here are some common errors found in AI-generated images: 1...2...3...4...5...You will be presented with an image... You should zoom in on specific areas for a closer look. Finally, please provide the following in a JSON file: 1...2...3...4...5...The following shows the json format: ..... Human Annotation Interface #2 #3 #1 #1 AcceptReject There is a distorted, unrecognizable wooden object....resembling a hallucinated artifact. #2 AcceptReject The staircase railing lacks... the glass panels seem to... the handrail transition is... . #3 AcceptReject The leaves of the plant have an artificial, plastic-like sheen... For Training Evidence Verifier Reasoning Trajectory Correction Global: I will start by observing the entire image... I will zoom in on the staircase area on the right... Local: 1: bbox: [600,200,1000,900], I am examining the staircase structure, focusing on ... - The connections look suspicious. + While the design is minimalist and some connections are simplified, they don't offer definitive proof of forgery. Next, I will look closely at the objects underneath the stairs... 2: bbox: [750,550,900,650], I am inspecting the wooden objects under the stairs... Next, I will move to the left side of the image... 3:bbox: [0,300,300,600], I am looking at the plant and the window frame on the left. - The plant's leaves and branches look unnatural. I have found enough evidence to form a conclusion. + The foliage and the background reflections appear mostly natural and consistent with a high-end render. However, the artifact found under the stairs is sufficient for a conclusion. I am going to present my final results. DALLE3 Input: Reasoning: Global: I will start by observing the entire image...I will zoom in on the staircase area on the right... Local: 1: bbox: [600,200,1000,900], I am examining the staircase structure... Next, I will look closely at the objects underneath the stairs... 2: bbox: [750,550,900,650], I am inspecting the wooden objects under the stairs..... Next, I will move to the left side... 3: bbox: [0,300,300,600], I am looking at the plant and the window frame... I have found enough evidence... VLM Verdict: Fake For SFT Image Collection Real Images Open Images V7 Nano Banana Figure 2: GroundFake construction pipeline. Starting from a balanced, bias-controlled image pool, we generate label-conditioned candidate reasoning trajectories and localized evidence, manually verify localized key evidence for fake images, and rewrite the trajectories into annotation-consistent supervision transcripts. The resulting dataset supports both SFT and Evidence Verifier training. the localized key evidence is manually verified under visual ground- ing and artifact specificity criteria. We further rewrite the reasoning trajectories using the verified evidence for annotation consistency. GroundFake provides two complementary supervision signals for Defake-o3: rewritten trajectories for supervised fine-tuning and valid/invalid evidence labels for Evidence Verifier training. The overall construction pipeline of GroundFake is illustrated in Fig. 2. 3.1 Image Collection and Bias Control We construct a balanced and bias-controlled image pool of 16k im- ages (8k real and 8k fake) for GroundFake. Real images are sampled from Open Images V7 [17], while fake images are collected from recent image generators and supplemented with additional FLUX.1- dev samples [28] to broaden photographic conditions. To reduce dataset-specific shortcuts, we standardize file format and align the aspect-ratio, coarse semantic-category, and aesthetic-quality distri- butions between the real and fake subsets, encouraging the detector to rely on generation artifacts rather than trivial source cues. 3.2 Label-Conditioned Candidate Annotation At this stage, we use Gemini 3 Pro Preview [19] to produce candi- date reasoning trajectories and evidence proposals conditioned on the known image label. The ground-truth label is provided only to elicit class-consistent candidate evidence and localization proposals. The prompt lists several common artifact categories in AI-generated images, including incoherent text, deformed faces or hands, dis- torted object structures, and violations of common sense or physical constraints, as heuristic cues for candidate generation. The model is then instructed to inspect the image region by region, with at- tention to fine-grained details, and to return a structured output containing a reasoning trajectory, a label-consistent conclusion, and supporting evidence. Specifically, the generated output contains three components: • Reasoning Trajectory: A sequential list in which each step records a region of interest (specified by a bounding box) together with the corresponding intermediate reasoning text. •Verdict: A label-consistent conclusion in “real”, “fake”, included to maintain a complete supervision format. •Evidence: The visual and textual basis associated with the con- clusion. For both real and fake images, the model provides one global evidence item that summarizes coarse scene-level coher- ence. For fake images, it further proposes several localized key evidence items, each paired with a bounding box that points to a candidate generation artifact. We treat the Gemini outputs as candidate annotations. Gemini 3 Pro Preview provides useful candidate trajectories and localization proposals, but a non-trivial portion of the localized key evidence remains weakly grounded or hallucinated. We therefore conduct a subsequent human verification step on the localized key evidence for fake images before these annotations are used for trajectory rewriting or Evidence Verifier training. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang 3.3 Human Annotation and Verification We focus human verification on the localized key evidence for fake images, because these artifact claims are the most hallucination- prone and are the annotations later used for verifier supervision. The proposed candidate key evidence may still be weakly grounded or invalid. Typical failure cases include: (1) the evidence text is irrelevant to the content inside the bounding box; (2) the text con- tradicts the actual visual content (e.g., correctly rendered text or normal hands are falsely described as defective); or (3) the claimed cue is not specific enough to AI generation (e.g., “smooth skin” which may also appear in heavily post-processed real photographs). We therefore ask human annotators to judge each localized key evi- dence item, consisting of the original image, the bounding box, and the associated text, with a binary label: Valid or Invalid. Each item is independently reviewed by three annotators under two criteria: (1)Visual Grounding: The textual description accurately matches the visual content within the bounding box. (2)Artifact Specificity: The described flaw is more specific to AI generation than a generic attribute that may also appear in real photographs. We retain the vote ratio across the three annotators as a soft validity label for Evidence Verifier training, and use the retained valid evidence in the subsequent trajectory rewriting step. This yields a higher-quality subset of localized evidence by filtering out unreasonable, weakly grounded, or non-specific indicators of AIGI. 3.4 Reasoning Trajectory Rewriting for Annotation Consistency Simply discarding invalid key evidence would leave the reasoning transcript inconsistent with the retained evidence annotations used for training. We therefore employ Gemini 3 Flash Preview [18] to rewrite the reasoning trajectories. Given the original trajectory and the human-verified valid evidence list, the model rewrites the intermediate reasoning text so that the resulting transcript remains consistent with the verified evidence and no longer depends on discarded hallucinated claims. The rewritten trajectories are then used as annotation-consistent supervision transcripts, providing stable training targets for tool use and final output formatting. 4 Defake-o3 Methodology To address the limitations of single-pass and weakly grounded MLLM detectors, we propose Defake-o3. Defake-o3 combines in- teractive visual search with verifier-guided evidence alignment. At inference time, it iteratively zooms into suspicious regions and then outputs a structured verdict with global evidence and, when available, localized key evidence. Training proceeds in two stages: SFT on human-filtered GroundFake traces to learn tool use and structured evidence generation, followed by GRPO with a hybrid evidence reward that combines rule-based matching and an Evi- dence Verifier trained from human verification labels. An overview is shown in Fig. 3. 4.1 Task Formulation Given an input imageI, the Defake-o3 model휋 휃 performs a multi- turn interaction with the image and then produces a structured output. Interactive Visual Search. At turn푖, the model receives an obser- vation푂 푖 , which is either the full image or a zoomed-in patch from a previous step, and chooses an action퐴 푖 ∈ Zoom In, Final Output. AZoom Inaction predicts a 2D bounding box푏 푖 and returns the corresponding crop as the next observation푂 푖+1 . AFinal Output action terminates the interaction and triggers the final output. Final Structured Output. We denote the final output byY= ( ˆ 푦,푒 푔 ,퐸) , where ˆ 푦 ∈ real, fakeis the verdict,푒 푔 = (∅,푡 푔 )is a mandatory global evidence item summarizing image-level cues, and 퐸=푒 1 , . . .,푒 퐾 is a possibly empty set of localized key evidence items. Each푒= (푏,푡) ∈ 퐸consists of a bounding box푏≠ ∅ and a text description푡of a local artifact. For a real verdict, we set퐸= ∅. For a fake verdict, the model outputs localized key evidence when clear and visually verifiable local artifacts are found. Otherwise, it may keep퐸=∅and return only the global evidence to avoid hallucinated localization. Throughout this paper, we reserve the term verifiable evidence for the localized evidence items in 퐸; the global evidence푒 푔 serves only as a contextual image-level assessment. 4.2 Supervised Fine-Tuning (SFT) In the SFT stage, we bootstrap the base MLLM using the human- filtered and trajectory-corrected traces in GroundFake. Each train- ing sample is cast as a multi-turn interaction sequence in which the model learns to invoke the zoom-in tool, condition on the returned patches, and emit the final structured output defined in Sec. 4.1. We train the model with standard autoregressive next-token prediction over the full trace. This stage mainly initializes tool use, sequential visual context accumulation, and structured evidence generation, while evidence quality is further aligned in the subsequent RL stage. 4.3 Evidence Verifier Training For grounded AIGI detection, exact matching to reference annota- tions, such as spatial overlap with human boxes [23,24], is useful but insufficient as a sole reward signal. Reference annotations can be incomplete, and annotation matching alone does not explicitly penalize invalid or fabricated evidence. To provide a denser and more human-aligned signal, we train an Evidence Verifier푉 휙 un- der the human verification protocol in Sec. 3.3. In GroundFake, each localized evidence item is judged by visual grounding and artifact specificity. Using these annotations, the verifier takes the original imageI, the cropped local regionI 푏 , the bounding box푏, and the evidence text푡as input, and predicts a validity score푠 ∈ [0,1]. We instantiate푉 휙 by appending a linear classification head to an MLLM and optimize it as a binary classifier with soft labels derived from annotator agreement. Specifically, if푚out of푀annotators mark an evidence item as valid, we set its target to푦 soft =푚/푀and train the verifier with a soft-label binary cross-entropy loss: 푠=푉 휙 (I,I 푏 ,푏,푡), L ver =−푦 soft log푠−(1−푦 soft ) log(1−푠). (1) This allows푉 휙 to capture the graded certainty of human judgments and provide a learned validity signal for the subsequent RL stage. Defake-o3 Policy Model (휋 ఏ ) Current Observation Decision Making if Action Is “Final Output” Model Output "verdict":"fake", "evidence": [ “bbox”: null, “text”: “The overall image...", "bbox": [110, 715, 170, 745], "text": "The text on the hull of the white boat is nonsensical gibberish...", "bbox": [610, 720, 830, 880], "text": "The interior of the red boat is distorted; the seats appear as simple, melted shapes..." ] Ground Truth Label: Fake Global Evidence: The overall image... Local Evidence #1 [120, 695, 360, 740]: The text on the side of the white boat is unrecognizable and consists of malformed characters... Local Evidence #2 [780, 840, 830, 900]: The rope on the red boat appears to emerge directly from the smooth surfaceof the hull's bow... Verification Global Image or Local Patch if Action is “Zoom In” New Observation Model-Based Total Reward: 푹 풕풐풕풂풍 = 푹 풄풍풔 + 푹 풆풗풊 GRPO Advantage Estimation: 퐴 = 푅 −푚푒푎푛(푹) 푠푡푑(푹) Rule-Based Feedback & Update Gradient Update Region Matching ... [The] [text] [on] [the] [hull] [of] [the] [white] [boat] ... ... [The] [text] [on] [the] [side] [of] [the] [white] [boat] ... Text Matching IoU BLEU-2 ⊕ 푹 풓풖풍풆 image text bbox image text bbox Evidence Verifier 풔 풆 ퟏ 풔 풆 ퟐ Aggregate Function 푹 풎풐풅풆풍 Evidence Reward: 푹 풆풗풊 = ퟏ−휶 푴 푹 풓풖풍풆 + 휶 푴 푹 풎풐풅풆풍 Classification Reward: 푹 풄풍풔 =ퟏ 풊풇 풗풆풓풅풊풄풕==풍풂풃풆풍 풆풍풔풆 ퟎ Update Current Observation Zoom In Figure 3: Overview of Defake-o3. The model performs interactive visual search by iteratively zooming into suspicious regions and then outputs a structured verdict with global and localized evidence. For every correctly classified fake image, the total reward combines the classification reward with rule-based evidence matching and model-based scoring from the Evidence Verifier. 4.4 Reward Function Design Our RL objective combines task correctness with localized evidence quality. We penalize unsupported evidence strongly and avoid re- warding the model for proposing many weak evidence items. Ac- cordingly, the reward is defined on the verdict and the localized key evidence set 퐸 from Sec. 4.1. 4.4.1 Total Reward (푅 푡표푡푎푙 ). For a final outputY=( ˆ 푦,푒 푔 ,퐸), where 퐸=푒 1 , . . .,푒 퐾 denotes the localized key evidence set, we define 푅 푡표푡푎푙 = ( 푅 푓푎푖푙 ,invalid format, 푅 푐푙푠 + I( ˆ 푦= fake) 푅 푒푣푖 ,otherwise, (2) where invalid format means that at least one tool call or the final structured output fails to parse, and we set푅 푓푎푖푙 =−1. The indicator function I(·) equals 1 when its argument is true and 0 otherwise. The classification reward measures only verdict correctness: 푅 푐푙푠 = I( ˆ 푦=푦 푔푡 ),(3) where푦 푔푡 is the ground-truth label; while the finer-grained shaping of localized evidence quality is handled by 푅 푒푣푖 . 4.4.2 Evidence Reward (푅 푒푣푖 ). We compute the evidence reward on the localized key evidence set퐸. We use a hybrid evidence reward that combines a coarse rule-based matching term with a verifier- based validity term: 푅 푒푣푖 =(1− 훼 푀 )푅 푟푢푙푒 + 훼 푀 푅 푚표푑푒푙 ,(4) where푅 푟푢푙푒 anchors the policy to the available reference annota- tions, while푅 푚표푑푒푙 provides a learned signal for evidence validity under the human verification protocol (see Sec. 4.3). This hybrid form avoids relying entirely on either exact annotation matching or the learned verifier alone. Unless otherwise stated, we set훼 푀 =0.5. 4.4.3 Rule-based Reward (푅 푟푢푙푒 ). Although exact matching is in- sufficient as a sole reward signal, it remains a useful coarse anchor when reference localized evidence is available. For the predicted localized evidence set퐸, we form a union mask푀 푝푟푒푑 from all pre- dicted boxes and concatenate all evidence texts into푇 푝푟푒푑 . Likewise, for the reference localized evidence, we form푀 푟푒푓 and푇 푟푒푓 . If the model incorrectly classifies a real image as fake, we set푅 푟푢푙푒 =0. Otherwise, the rule-based reward is defined as: 푅 푟푢푙푒 = 휆 푖표푢 IoU(푀 푝푟푒푑 ,푀 푟푒푓 )+ 휆 푏푙푒푢 BLEU-2(푇 푝푟푒푑 ,푇 푟푒푓 )(5) where휆 푖표푢 = 휆 푏푙푒푢 =1 by default. We use set-level matching, since reference annotations may be incomplete and multiple valid decompositions of evidence can exist for the same image. The BLEU- 2 term serves as a lightweight lexical anchor. 4.4.4 Model-based Reward (푅 푚표푑푒푙 ). For each localized evidence item푒=(푏,푡) ∈ 퐸, we extract the corresponding cropI 푏 and query the Evidence Verifier푉 휙 , obtaining a score푠 푒 =푉 휙 (I,I 푏 ,푏,푡). We then convert the verifier score into an item-level reward: 푟 푚표푑푒푙,푒 = ( −퐶 invalid ,푠 푒 < 휏, 퐶 valid · 푠 푒 −휏 1−휏 , 푠 푒 ≥ 휏, (6) where휏is the acceptance threshold under the verification protocol in Sec. 4.3. This mapping assigns a fixed penalty to low-scoring evi- dence and only bounded positive reward to evidence whose verifier score exceeds휏. When a real image is incorrectly predicted as fake, the same mapping naturally penalizes the proposed localized evi- dence, since such evidence is unlikely to pass the verifier threshold. Unless otherwise stated, we use 휏= 0.5 and퐶 valid =퐶 invalid = 1. To avoid rewarding the model for proposing many borderline evidence items, we aggregate item-level rewards with cumulative penalties and capped positive rewards. Let E − =푒 | 푟 푚표푑푒푙,푒 < 0, E + =푒 | 푟 푚표푑푒푙,푒 ≥ 0,(7) Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang Table 1: Results on the GroundFake test set. MethodAccF1BLEU-1 BLEU-2 ROUGE-LIoU CNNSpot0.9760.976– NPR0.8380.830– AIDE0.9240.922– FakeVLM0.8570.8540.1050.0450.110– LEGION0.5020.0240.0030.0010.0020.000 FakeShield0.4940.2750.0290.0130.0260.017 Defake-Direct (w/o RL)0.9460.9430.2820.1540.2340.165 Defake-Direct0.9580.9570.3330.1950.2520.262 Defake-CoT (w/o RL)0.9880.9880.3300.1840.2670.222 Defake-CoT0.992 0.9920.3350.1840.2570.280 Defake-o3 (w/o RL)0.9720.9710.3190.1790.2500.241 Defake-o30.992 0.992 0.364 0.2150.281 0.311 Acc/F1: image-level classification performance. BLEU-1/2 and ROUGE-L: evidence-level lexical overlap with the reference annotations. IoU: evidence-level spatial overlap with the reference regions. and let푟 (1) ≥ 푟 (2) denote the largest and second-largest reward of the evidence inE + , respectively. We define: 푅 푚표푑푒푙 = ∑︁ 푒∈E − 푟 푚표푑푒푙,푒 | z Penalty +A(E + ) | z Reward , A(E + )= 0,|E + |= 0, 푟 (1) ,|E + |= 1, 푟 (1) + 훽 푟 (2) , |E + | ≥ 2. (8) This aggregation reflects our preference for a small number of strong, verifiable artifacts rather than many weak ones. In our implementation, we set 훽= 0.5. 4.5 Reinforcement Learning Starting from the SFT-trained weights, we optimize Defake-o3 using GRPO [50,67]. Each sampled rollout contains multiple rounds of Zoom Inactions followed by a final structured output. Through trajectory-level optimization, the model learns to use the zoom- in tool judiciously and to favor a small number of high-quality localized evidence items over many weak or vague ones. 5 Experiments 5.1 Experimental Setup Implementation Details. Unless otherwise stated, all experi- ments are conducted on 8 A100 GPUs. We employ Qwen3-VL- 8B-Instruct [3] as our base model and fine-tune it using LoRA [22], which is applied to all linear layers within the text backbone, with a rank of푟=16 and an alpha of훼=32. During the SFT stage, we train the model using a global batch size of 8 and a learning rate of 1×10 −4 . In the RL stage, the group size is set to 16, and the learning rate is 1×10 −5 . Following DAPO [67], we implement several improvements to the standard GRPO algorithm: (1) remov- ing the KL divergence constraint; (2) adopting token-level loss instead of sequence-level loss; (3) applying a clip higher strategy with휖 low =0.2 and휖 high =0.28; and (4) utilizing overlong reward masking, a mechanism that limits sampling length without sup- pressing the generation of longer reasoning paths. The Evidence Verifier shares the same base model, LoRA configuration, and SFT training settings, with the exception that its linear classification head is fully trainable. Baselines and Ablations. We compare Defake-o3 with traditional binary classification approaches including CNNSpot [57], NPR [52], and AIDE [64]. For a fair comparison, these models are re-trained on our GroundFake training set. For MLLM-based explainable methods, we employ and test FakeVLM [59], LEGION [26], and FakeShield [63]. We evaluate FakeVLM and FakeShield using their officially released weights, while LEGION is reproduced utilizing its official training code and data. To validate the effectiveness of our proposed method, we design two variants of our model for ablation studies: (1) Defake-CoT : relies solely on pure textual chain- of-thought reasoning without zoom-in tool calls; and (2) Defake- Direct: directly outputs the final verdict and evidence without any preliminary reasoning steps. 5.2 Results on GroundFake Test Set Table 1 reports classification (Acc/F1), evidence text overlap (BLEU- 1/2 and ROUGE-L), and evidence region overlap (IoU) of state-of- the-art AIGI detectors and Defake-o3 on the GroundFake test set. RL consistently improves Acc/F1 and IoU for all three variants. En- abled by interactive visual search for verifiable evidence, Defake-o3 achieves the strongest overall explainable performance, match- ing the best Acc/F1 while attaining the highest BLEU-1, BLEU-2, ROUGE-L, and IoU. 5.3 FakeFrontier Benchmark Beyond the in-domain evaluation on GroundFake, we further as- sess out-of-distribution generalization on recent generators using FakeFrontier. FakeFrontier contains 4,000 images, evenly split between real and fake. The real subset is drawn from Open Images V7 [17] and Chameleon [64], while the fake subset is collected from outputs of 10 recent image generators, including Nano Banana [20], GPT-Image-1.5 [40], and Seedream 4.5 [6]. The Open Images V7 images used in FakeFrontier do not overlap with those used in GroundFake. Unlike GroundFake, FakeFrontier does not provide human evidence annotations and is therefore used to evaluate both classification on the latest generators and explanation quality with- out relying on reference boxes or texts. To evaluate explanations under this annotation-free setting, we adopt an MLLM-based pro- tocol on 200 fake images randomly sampled from FakeFrontier. We instantiate the judge with three open-source MLLMs, Qwen3- VL-235B-A22B-Thinking [3], Kimi K2.5 [54], and GLM-4.6V [68], to reduce evaluator-specific bias. Each evidence item is assessed independently rather than as part of a bundled explanation. We con- sider two complementary aspects: evidence quality, which measures whether an evidence item is visually grounded and specific to AI generation, and evidence persuasiveness, which measures whether it can convince the judge to predict “fake”. Table 2 reports the MLLM- based evaluation, and Table 3 reports classification results on the full FakeFrontier benchmark. Quality Evaluation. We first evaluate each evidence item indepen- dently with an MLLM judge. The judge assigns a score from 1 to 10, where 1 denotes evidence that is irrelevant, contradictory, or not specific to AI generation, and 10 denotes evidence that is visually grounded and highly specific to AI generation. We normalize the score to[0,1]and report it as the Quality Score (QS). For each image, we average the normalized scores over its evidence items, Defake-o3 Visual Reasoning with Zoom-in Tool Calls Global: The image depicts a corgi sitting in a gaming chair in a setup that resembles a streamer's room.... I will then zoom in on specific details like the dog, the chair, and the desk accessories. Local #1 [470, 200, 730, 450]: I am examining the dog's head and the headset... The logo on the headset looks like a generic 'H' or similar character that doesn't quite fit a brand. I will check the chair next. Local #2 [680, 130, 850, 250]: I am looking at the chair's headrest and back. The logo 'MXLGR' is visible... I will move to the desk area next. Local #3 [0, 450, 450, 700]: I am examining the desk area, including the keyboard, mouse, and cans... The wires and cables look a bit messy. I will check the monitor and background next. Local #4 [0, 150, 280, 480]: I am looking at the monitor screen and the background posters. The image on the screen is a fantasy scene... I will check the bottom left corner next. Local #5 [0, 970, 150, 1000]: I am examining the timestamp in the bottom left corner. The text '2202S7:80:42:209' looks like a nonsensical timestamp format... I will now synthesize my findings. Input Image Defake-o3 (Ours) Global Evidence: The overall image has a hyper-realistic, overly smooth aesthetic typical of AI generation, particularly in the lighting and the texture of the dog's fur. Local Evidence #1 bbox: [660, 330, 710, 390] The logo on the headset is a nonsensical blue symbol... Local Evidence #2 bbox: [10, 975, 145, 995] The timestamp in the bottom left corner... is a nonsensical and non- functional format... Verdict: Fake Visual Reasoning with Zoom-in Tool Calls Global: The image shows a news anchor in a studio setting... I will then zoom in on specific details like the face, hands, and objects on the desk to check for AI artifacts. Local #1 [400, 220, 600, 450]: I am examining the face and hair. The skin texture appears very smooth... The hair looks well-rendered. I will move to the hands next. Local #2 [150, 750, 850, 850]: I am looking at the hands and the objects on the desk. The hands appear somewhat stiff. There is a small object on the right side of the desk that looks suspicious. I will examine that area closer. Local #3 [750, 760, 850, 840]: The object on the desk is unclear. It looks like a cluster of dark shapes that don't form a recognizable object like a phone or a prop. It seems to be floating or attached in a strange way. I will check the text and background next. Local #4 [30, 60, 980, 150]: The text 'BREAKING NEWS: MARKETS SOAR' is legible and well-rendered. However, the anomaly on the desk remains the most suspicious element. I am ready to provide my verdict. Input Image Defake-o3 (Ours) Global Evidence: The overall image has a hyper-realistic, overly polished look typical of AI generation, particularly in the perfect lighting and the smoothness of the skin texture. Local Evidence #1 bbox: [760, 770, 840, 830] There is a nonsensical, dark object floating on the desk surface near the anchor's right hand. It does not resemble a real object like a phone or a piece of equipment and lacks a clear structure. Verdict: Fake FakeVLM Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeShield Verdict: Fake ...The shadows and highlights on the dog's hind legs appear inconsistent with the lighting on the rest of the dog and the environment... FakeVLM Verdict: FakeThe image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeShield Verdict: Fake...The figure appears to be under a different light source compared to the rest of the scene... The absence of a shadow... suggests a violation of physical laws... #1 #2 #4 #3 #5 #1 #2 #3 #4 Figure 4: Qualitative examples from the GroundFake test set. Defake-o3 first performs coarse scene-level inspection and then uses targeted zoom-in steps to verify candidate artifacts, yielding a small set of localized, visually verifiable evidence items. Compared with FakeVLM [59] and FakeShield [63], its outputs are more directly grounded in explicit, image-specific artifacts. FakeShield masks are omitted for brevity. Table 2: Evaluation with the MLLM-based protocol on the fake-image subset of FakeFrontier. MethodAvg #Evi Qwen3-VL-235B (Prior-Acc = 0.2421) Kimi K2.5 (Prior-Acc = 0.3526)GLM-4.6V (Prior-Acc = 0.0737) QSHit@ImgHit@EviQSHit@ImgHit@EviQSHit@ImgHit@Evi Defake-o33.08040.53780.78390.63620.43340.74710.58950.68120.90450.8249 Defake-CoT2.93970.48410.77550.62390.38110.71190.57040.59800.89610.7584 Defake-Direct3.19100.37210.60470.48610.28770.56450.45040.44310.70850.5554 FakeVLM3.64210.12680.37020.25580.12040.45260.35310.24380.54560.5092 FakeShield2.21610.08750.23120.32960.07630.27470.42180.13930.27470.4255 Avg #Evi: average number of evidence items per image. Prior-Acc: probability that the MLLM evaluator classifies an image as fake without additional information. QS: Quality Score. Hit@Img / Hit@Evi: Image Hit Rate / Evidence Hit Rate. LEGION is excluded from the explainability evaluation, as it classifies nearly all fake images as real. Table 3: Classification results on FakeFrontier benchmark. MethodReal Acc Fake Acc AccF1 Defake-o30.90200.9446 0.9232 0.9246 Defake-CoT0.86750.95320.91020.9137 Defake-Direct0.84850.78260.81570.8088 FakeVLM0.71270.71020.71140.7108 FakeShield0.81350.41070.61270.5139 LEGION0.99300.00450.50040.0090 CNNSpot0.82000.47860.64990.5767 NPR0.86100.11470.48910.1829 AIDE0.74950.50830.62930.5775 and then average the resulting image scores over the evaluation set. Persuasion Evaluation. We next evaluate whether a single ev- idence item is strong enough to support a fake verdict. For each fake image, the judge first predicts from the image alone, which reveals its prior tendency and is reported as Prior-Acc in Table 2. Table 4: Accuracy on other OoD test sets. Test SetDefake-o3 FakeVLM LEGION FakeShield CNNSpotNPRAIDE AIGI-Now 0.91800.85960.49970.83070.77300.6925 0.8473 EvalGEN0.98720.96680.09070.95130.76600.2430 0.7045 MNW0.88710.84170.02650.50720.64350.3775 0.7712 We then provide one evidence item from a detector and ask the judge to assess the same image again. Each evidence item is tested independently. Because this decision can be sensitive to prompting, we use three system prompts and average the results. We report two metrics: (1) Hit@Img, the proportion of fake images for which at least one evidence item leads the judge to output “fake”; and (2) Hit@Evi, the proportion of evidence items that lead the judge to output “fake”. For detectors that output plain text rather than structured JSON, we first use Qwen3-VL-8B-Instruct to split their outputs into independent evidence items. These metrics are comple- mentary: QS measures the visual grounding and artifact specificity of the evidence, Hit@Img measures whether a detector can surface at least one decisive flaw for a fake image, and Hit@Evi measures Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang how often its individual evidence items are actually useful rather than noisy or hallucinated. As shown in Tables 2 and 3, Defake-o3 achieves the highest Acc/F1 on FakeFrontier and also ranks first in QS, Hit@Img, and Hit@Evi under all three judges. Although baseline methods produce more evidence items on average, their much lower Hit@Evi suggests that many of those claims are weak or unpersuasive. 5.4 Results on External OoD Benchmarks Beyond FakeFrontier, we further evaluate classification general- ization on three external OoD benchmarks: AIGI-Now [11], Eval- GEN [12], and MNW [38]. For these external benchmarks, we report classification accuracy only. To maintain a strict OoD setting, we remove samples from generators that overlap with the Ground- Fake training set. As shown in Table 4, Defake-o3 achieves the best accuracy on all three datasets, reaching 0.9180 on AIGI-Now, 0.9872 on EvalGEN, and 0.8871 on MNW. It consistently outper- forms both explainable MLLM baselines and black-box detectors, showing that the advantage of Defake-o3 is not limited to Ground- Fake or FakeFrontier. These results suggest that interactive visual search and verifier-guided evidence alignment improve robustness under broader distribution shifts, without sacrificing classification strength. 5.5 Ablation Studies We conduct ablation studies to assess the effectiveness of interactive visual search and the role of the Evidence Verifier. Effectiveness of Interactive Visual Search: Based on the results in Tables 1–3, Defake-o3 exhibits clear advantages in both classifi- cation performance and explanation quality over the Defake-CoT model (which lacks the zoom-in mechanism) and the Defake-Direct model (which lacks any reasoning process). The Role of the Evidence Verifier: Figure 5 shows that the Evi- dence Verifier assigns high scores to human-accepted evidence and much lower scores to human-rejected evidence or forcibly fabri- cated evidence on random bounding boxes (Random in Fake/Real), indicating that it captures the visual grounding and artifact speci- ficity criteria in our annotation protocol. This learned signal leads to better reward shaping during RL. Compared with the rule-only vari- ant (훼 푀 =0 in Table 5), the model trained with the Evidence Verifier improves IoU on GroundFake from 0.252 to 0.311, while causing only minor changes in lexical overlap with the reference annota- tions. Table 5 further shows that the best results require a balance between rule-based matching and verifier guidance. Using only the verifier (훼 푀 =1) degrades text-explanation performance. Although a smaller verifier weight (훼 푀 =0.25) yields a slightly higher IoU, it noticeably degrades Acc/F1. We therefore use훼 푀 =0.5 as the best overall trade-off. 5.6 Qualitative Results Figure 4 illustrates how Defake-o3 turns exploratory inspection into compact, verifiable evidence. It zooms into candidate regions, retains only verifiable artifacts—such as the malformed headset logo, invalid timestamp, and implausible desk object—and discards Figure 5: Output value distribution of the Evidence Verifier on generated evidence. Table 5: Ablation of reward hyperparameters evaluated on the GroundFake test set. 훼 푀 휆 푖표푢 /휆 푏푙푒푢 AccF1BLEU-1BLEU-2ROUGE-LIoU 100.9840.9840.2750.1500.2220.195 0.7510.9100.9020.3030.1760.2350.260 0.510.992 0.9920.3640.2150.2810.311 0.2510.9620.9600.3300.1940.2700.318 010.9860.9860.3630.2180.2910.252 unsupported suspicions. In contrast, FakeVLM and FakeShield pro- vide coarser evidence that is less directly tied to image-specific artifacts. 6 Conclusion In this paper, we presented Defake-o3, an explainable AIGI detec- tor that moves from speculative rationales to verifiable evidence through interactive visual search and verifier-guided evidence align- ment. We further introduced GroundFake and FakeFrontier for grounded training and out-of-distribution evaluation. Experiments show that Defake-o3 achieves the best overall performance among compared methods, with top or tied-for-top classification results and stronger judge-based evidence quality. Acknowledgments This work was supported in part by the National Natural Sci- ence Foundation of China (Grant Nos. 62302295, 62595733, and 62561160155), the Shanghai Municipal Science and Technology Major Project (Grant No. 2021SHZDZX0102). This work was also supported by Ant Group. In particular, we sincerely thank Zijuan Yu, Yan Hong, and Jun Lan from Ant Group for their valuable help. References [1] Stability AI. 2024. stable-diffusion-3.5-large. https://huggingface.co/stabilityai/ stable-diffusion-3.5-large. Accessed: 2026-04-09. [2]Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736. [3]Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al.2025. Qwen3-VL Technical Report. arXiv preprint arXiv:2511.21631 (2025). [4]James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. 2023. Improving Image Generation with Better Captions. Technical Report. OpenAI. https://cdn.openai. com/papers/dall-e-3.pdf Defake-o3 [5]Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901. [6]ByteDance. 2025. Seedream 4.5. https://seed.bytedance.com/en/seedream4_5. Accessed: 2026-03-31. [7] Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Zhaohui Hou, Shijie Huang, Dengyang Jiang, Xin Jin, Liangchen Li, et al.2025. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer. arXiv preprint arXiv:2511.22699 (2025). [8]Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al.2025. HunyuanImage 3.0 Technical Report. arXiv preprint arXiv:2509.23951 (2025). https://arxiv.org/abs/ 2509.23951 [9]Baoying Chen, Jishen Zeng, Jianquan Yang, and Rui Yang. 2024. DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images. In Forty-first International Conference on Machine Learning. [10] Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020. Generative pretraining from pixels. In International conference on machine learning. PMLR, 1691–1703. [11]Ruoxin Chen, Jiahui Gao, Kaiqing Lin, Keyue Zhang, Yandan Zhao, Is- abel Guan, Taiping Yao, and Shouhong Ding. 2026.AlignGemini: Gen- eralizable AI-Generated Image Detection Through Task-Model Alignment. arXiv:2512.06746 [cs.CV] https://arxiv.org/abs/2512.06746 [12]Ruoxin Chen, Junwei Xi, Zhiyuan Yan, Ke-Yue Zhang, Shuang Wu, Jingyi Xie, Xu Chen, Lei Xu, Isabel Guan, Taiping Yao, et al.2025. Dual Data Align- ment Makes AI-Generated Image Detector Easier Generalizable. arXiv preprint arXiv:2505.14359 (2025). [13]Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017). [14] Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. 2023. InstructBLIP: Towards General-Purpose Vision-Language Models with Instruction Tuning. Advances in neural information processing systems 36 (2023), 49250–49267. [15] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al.2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020). [16]Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al.2024. Scaling rectified flow transformers for high-resolution image synthesis. In Forty- first international conference on machine learning. [17] Google. 2022. Open Images V7. https://storage.googleapis.com/openimages/ web/factsfigures_v7.html. Accessed: 2026-03-31. [18] Google. 2025. Gemini 3 Flash Preview. https://ai.google.dev/gemini-api/docs/ models/gemini-3-flash-preview. Accessed: 2026-08-09. [19]Google. 2025. Gemini 3 Pro Preview. https://ai.google.dev/gemini-api/docs/ models/gemini-3-pro-preview. Accessed: 2026-08-09. [20]Google. 2025. Gemini Image – Nano Banana. https://deepmind.google/models/ gemini-image/. Accessed: 2026-03-31. [21]Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020.Denoising Dif- fusion Probabilistic Models. In Advances in Neural Information Process- ing Systems, Vol. 33. 6840–6851.https://papers.nips.c/paper/2020/hash/ 4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html [22] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al.2022. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 1, 2 (2022), 3. [23]Zhenglin Huang, Tianxiao Li, Xiangtai Li, Haiquan Wen, Yiwei He, Jiangning Zhang, Hao Fei, Xi Yang, Xiaowei Huang, Bei Peng, and Guangliang Cheng. 2025. So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection. arXiv preprint arXiv:2505.18660 (2025). arXiv:2505.18660 [cs.CV] https://arxiv. org/abs/2505.18660 [24]Yikun Ji, Yan Hong, Qi Fan, Jun Lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, and Jianfu Zhang. 2026. FakeXplain: AI-Generated Image Detection via Human- Aligned Grounded Reasoning. In ICLR.https://openreview.net/forum?id= UcpTOa8OnG [25]Changjiang Jiang, Wenhui Dong, Zhonghao Zhang, Chenyang Si, Fengchang Yu, Wei Peng, Xinbin Yuan, Yifei Bi, Ming Zhao, Zian Zhou, and Caifeng Shan. 2025. IVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection. arXiv preprint arXiv:2506.00979 (2025). arXiv:2506.00979 [cs.CV] https://arxiv.org/abs/2506.00979 [26] Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Weijia Li, Peilin Feng, Baichuan Zhou, Bin Wang, Dahua Lin, Linfeng Zhang, and Conghui He. 2025. LEGION: Learning to Ground and Explain for Synthetic Image Detection. arXiv preprint arXiv:2503.15264 (2025). arXiv:2503.15264 [cs.CV] https://arxiv.org/abs/ 2503.15264 [27]Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. 2023. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision. 4015–4026. [28]Black Forest Labs. 2024. black-forest-labs/FLUX.1-dev. https://huggingface.co/ black-forest-labs/FLUX.1-dev. Accessed: 2026-03-31. [29] Black Forest Labs. 2026. FLUX.2. https://bfl.ai/flux2. Accessed: 2026-04-09. [30]Xin Lai, Junyi Li, Wei Li, Tao Liu, Tianjian Li, and Hengshuang Zhao. 2025. Mini-o3: Scaling up reasoning patterns and interaction turns for visual search. arXiv preprint arXiv:2509.07969 (2025). [31]LAION-AI. 2022. LAION-Aesthetics_Predictor V1. https://github.com/LAION- AI/aesthetic-predictor. Accessed: 2026-03-31. [32] Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Boot- strapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. In International conference on machine learning. PMLR, 19730– 19742. [33]Yixuan Li, Xuelin Liu, Xiaoyang Wang, Bu Sung Lee, Shiqi Wang, Anderson Rocha, and Weisi Lin. 2025. FakeBench: Probing Explainable Fake Image Detec- tion via Large Multimodal Models. IEEE Transactions on Information Forensics and Security (2025). [34]Yixuan Li, Yu Tian, Yipo Huang, Wei Lu, Shiqi Wang, Weisi Lin, and An- derson Rocha. 2025. FakeScope: Large Multimodal Expert Model for Trans- parent AI-Generated Image Forensics. arXiv preprint arXiv:2503.24267 (2025). arXiv:2503.24267 [cs.CV] https://arxiv.org/abs/2503.24267 [35]Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual in- struction tuning. Advances in neural information processing systems 36 (2023), 34892–34916. [36] Xingchao Liu, Chengyue Gong, and Qiang Liu. 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 (2022). [37]Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang. 2025. Visual-RFT: Visual Reinforcement Fine-Tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 2034– 2044. [38] Microsoft. 2025. MNW Benchmark Dataset. https://github.com/microsoft/MNW/ tree/main. Accessed: 2026-03-31. [39] Midjourney. 2023. Midjourney. https://w.midjourney.com/home. Accessed: 2026-03-31. [40] OpenAI. 2025. GPT-Image-1.5 Model | OpenAI API. https://developers.openai. com/api/docs/models/gpt-image-1.5. Accessed: 2026-03-31. [41] OpenAI. 2026. GPT Image 1 Model | OpenAI API. https://platform.openai.com/ docs/models/gpt-image-1. Accessed: 2026-04-09. [42]Keiron O’shea and Ryan Nash. 2015. An introduction to convolutional neural networks. arXiv preprint arXiv:1511.08458 (2015). [43]Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving La- tent Diffusion Models for High-Resolution Image Synthesis. arXiv preprint arXiv:2307.01952 (2023). [44] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36 (2023), 53728–53741. [45]Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning. PMLR, 8821–8831. [46] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Mod- els. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10674–10685. doi:10.1109/CVPR52688.2022.01042 [47]Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al.2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35 (2022), 36479–36494. [48]John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017). [49]Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. 2024. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning. Advances in Neural Information Processing Systems 37 (2024), 8612– 8642. [50] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al.2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300 (2024). [51]Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020). Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang [52]Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. 2024. Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 28130–28139. [53]ByteDance Seed Vision Team. 2025. Seedream 3.0 Technical Report. https:// seed.bytedance.com/en/public_papers/seedream-3-0-technical-report. Accessed: 2026-04-09. [54]Kimi Team. 2026. Kimi K2.5: Visual Agentic Intelligence. https://w.kimi.com/ blog/kimi-k2-5. Accessed: 2026-08-08. [55]Aaron Van Den Oord, Oriol Vinyals, et al.2017. Neural discrete representation learning. Advances in neural information processing systems 30 (2017). [56]Haozhe Wang, Alex Su, Weiming Ren, Fangzhen Lin, and Wenhu Chen. 2025. Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven rein- forcement learning. arXiv preprint arXiv:2505.15966 (2025). [57]Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. 2020. CNN-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8695–8704. [58]Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. 2023. DIRE for Diffusion-Generated Image Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 22445– 22455. [59]Siwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang, Zichen Wen, Yize Chen, Jiang Wu, Wenjun Wu, Conghui He, and Weijia Li. 2025. Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation. arXiv preprint arXiv:2503.14905 (2025). arXiv:2503.14905 [cs.CV] https://arxiv. org/abs/2503.14905 [60] Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng- ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al.2025. Qwen-Image Technical Report. arXiv preprint arXiv:2508.02324 (2025). [61]Penghao Wu and Saining Xie. 2024. V*: Guided Visual Search as a Core Mecha- nism in Multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13084–13094. [62]Guowei Xu, Peng Jin, Ziang Wu, Hao Li, Yibing Song, Lichao Sun, and Li Yuan. 2025. LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. In Proceed- ings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, Los Alamitos, CA, USA, 2087–2098. [63]Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, and Jian Zhang. 2025. FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models. In ICLR. https://openreview.net/forum? id=pAQzEY7M03 [64] Shilin Yan, Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, and Weidi Xie. 2024. A Sanity Check for AI-Generated Image Detection. arXiv preprint arXiv:2406.19435 (2024). [65]Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, et al.2024. LOKI: A Compre- hensive Synthetic Data Detection Benchmark Using Large Multimodal Models. arXiv preprint arXiv:2410.09732 (2024). [66]Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al.2022. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789 2, 3 (2022), 5. [67]Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al.2025. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv preprint arXiv:2503.14476 (2025). [68]Z.ai. 2025. GLM-4.6V: Open Source Multimodal Models with Native Tool Use. https://z.ai/blog/glm-4.6v. Accessed: 2026-03-31. [69] Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. 2025. DeepEyes: Incentivizing “Thinking with Images” via Reinforcement Learning. arXiv preprint arXiv:2505.14362 (2025). [70]Ziyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun, Jiayi Ji, Ke Yan, Shouhong Ding, Xiaoshuai Sun, Yunsheng Wu, and Rongrong Ji. 2025. AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models. arXiv preprint arXiv:2507.02664 (2025). arXiv:2507.02664 [cs.CV] https://arxiv.org/abs/2507.02664 Defake-o3 Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection Supplementary Material Figure S1: Comparison of validation loss between Defake-o3 and Defake-CoT in the SFT stage. Figure S2: Average number of tool calls and other metrics during the RL stage of Defake-o3. A Prompts In this section, we summarize all the prompts utilized throughout our study. The full content of each prompt is provided at the end of this supplementary material. •Prompts for Inference: The system prompts used during infer- ence by Defake-o3, Defake-CoT, and Defake-Direct are detailed in Figures S7, S8, and S9, respectively, and the unified user prompt is provided in Figure S10. •Prompts for GroundFake Dataset Construction: The sys- tem and user prompts used to instruct Gemini 3 Pro Preview to generate evidence candidates are shown in Figures S11 and S12, and the prompt for trajectory rewriting is detailed in Figure S13. •Prompts for FakeFrontier Benchmark: The prompt used for extracting descriptions for image generation is shown in Figure S14, the prompt for evidence quality evaluation is detailed in Figure S15, the three system prompts for persuasion evaluation are provided in Figures S16, S17, and S18, and the prompt for parsing baseline explanations is shown in Figure S19. B Robustness Evaluation We evaluate the performance of our model on the FakeFrontier benchmark under various image degradations, including JPEG com- pression (Quality Factor = 70), Gaussian blur (휎=1), and resizing (×0.5). As shown in Table S1, Defake-o3 exhibits strong robustness against these common visual perturbations. C Discussion on Training Dynamics During the SFT stage, we consistently observed that the evaluation loss of the Defake-CoT model was significantly higher than that of Defake-o3 (see Figure S1), despite their reasoning text being identical. This indicates that the absence of magnified visual patches makes the reasoning process substantially more difficult. Furthermore, during the RL training of Defake-o3, we observed that although we did not introduce any explicit reward for tool invocation, the number of zoom-in tool calls gradually increased after an initial decrease (see Figure S2). This suggests that the model, during the learning process, autonomously overcame the inertia of not invoking tools [69] and recognized the crucial role of the zoom-in tool in localizing generative flaws. D Ablation on Rule-based Reward On the GroundFake test set, incorporating the Evidence Verifier into the reward function significantly improves the Intersection over Union (IoU) compared to the model relying solely on the rule- based reward, albeit with a slight decrease in textual metrics. A natural question arises: if we solely use the rule-based reward but increase the weight of the IoU component, can we achieve the same effect? As shown in Table S2, increasing the ratio휆 푖표푢 /휆 푏푙푒푢 under the rule-based reward only setting (훼 푀 =0) can improve the IoU to some extent. However, it still falls short of the spatial accuracy achieved when the Evidence Verifier is introduced (훼 푀 =0.5). This confirms that the Evidence Verifier provides unique signals beyond what simple bounding box matching can offer. E Base Model Performance We report the performance of the pre-trained Qwen3-VL-8B-Instruct base model on the GroundFake dataset without any fine-tuning. During inference, we utilize the same prompt as Defake-Direct. The performances of Defake-Direct and Defake-o3 are also provided in Table S3 for comparison. As can be seen, our training pipeline brings substantial performance improvements over the general-purpose base model. F Inference Speed We present a throughput comparison of the Defake-o3, Defake-CoT, and Defake-Direct methods, alongside three explainable baselines (FakeVLM, LEGION, and FakeShield) on an 8×A100 GPU platform. As shown in Table S4, Defake-o3 trades lower inference efficiency for higher detection performance and explanation quality, largely due to the sequential generation and multi-turn interaction nature of its “thinking-with-images” paradigm. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang Table S1: Robustness evaluation on FakeFrontier. PerturbationDefake-o3FakeVLMLEGIONFakeShield Original0.92320.71140.50040.6127 JPEG Compression 0.89240.68910.50040.5952 Gaussian Blur0.87460.68880.49910.6090 Resize0.91820.66820.50040.5867 Table S2: Ablation of rule-based reward hyperparameters evaluated on the GroundFake test set. 훼 푀 휆 푖표푢 /휆 푏푙푒푢 AccF1BLEU-1BLEU-2ROUGE-LIoU 0.510.992 0.992 0.3640.2150.2810.311 010.9860.9860.3630.2180.2910.252 020.9520.9500.3210.1900.2660.237 040.9740.9730.3410.2020.2860.258 Table S3: Performance of the base model compared to our trained models on GroundFake. ModelAccF1BLEU-1BLEU-2ROUGE-LIoU Qwen3-VL-8B-Instruct0.8240.8230.1080.0510.1590.143 Defake-Direct0.9580.9570.3330.1950.2520.262 Defake-o30.992 0.9920.3640.2150.2810.311 Table S4: Throughput comparison during inference. MethodThroughput (samples/second) Defake-Direct10.668 Defake-CoT8.180 Defake-o32.606 FakeVLM3.008 LEGION7.398 FakeShield0.535 G More Details About GroundFake G.1 Image Collection and Bias Control We collected a balanced training set of 16k images (8k real and 8k fake) and applied several bias control procedures to reduce obvious shortcut differences between real and fake images. Data Sources. Our training objective is to enable the model to iden- tify subtle generative flaws within AI-generated images. However, if the quality of the generated images in the training set is too low and the errors are overly apparent, the resulting model may per- form poorly on high-quality images produced by newer generators. Therefore, we aim to construct a generated image dataset that main- tains a high average quality while presenting a gradient of difficulty. To achieve this, we scraped 6.5k generated images from the Inter- net, covering newer generators including SDXL [43], DALLE3 [4], Midjourney V5 [39], and Nano Banana [20]. For authentic images, we sampled 8k real images from OpenImageV7 [17]. Concurrently, we observed that recent generators typically pro- duce images of high aesthetic quality (e.g., perfect composition and lighting). This tendency creates a stylistic discrepancy between the fake image subset and the real image subset, potentially in- ducing a tendency for the model to overfit to aesthetic features rather than genuine generative artifacts. To counteract this, we generated an additional 1.5k fake images using FLUX.1-dev [28] paired with a specialized custom LoRA [22]. This LoRA enables the model to simulate the diverse and often imperfect shooting con- ditions typical of authentic photographs, such as underexposure and motion blur. During this generation process, we utilized the captions of real images from OpenImageV7 as prompts, which ef- fectively produced fake images that closely mimic the diverse styles of real photographs (see Figure S3). The final dataset comprises 8k generated images and 8k authentic images. All these images have a longer side of at least 1024 pixels. Bias Control Procedures. AIGI detectors can easily rely on dataset- specific shortcuts (non-semantic biases, e.g., aesthetic style) instead of actual generation artifacts. We therefore applied the following procedures: •Format Standardization: All images were converted to high- quality JPG to remove trivial file format cues across different sources. •Aspect Ratio Alignment: Among the collected fake images, square formats are overrepresented. To reduce the shortcut that square images are fake, we cropped fake images to follow the aspect ratio distribution of OpenImageV7. •Category Alignment: We used Qwen3-VL-8B-Thinking [3] as a coarse classifier to group images into four broad categories: human, animal, object, and scene. At the same time, we filtered out clearly non-photorealistic content such as cartoons and digital art. The four categories were kept balanced at a ratio of 1:1:1:1. •Aesthetic Quality Alignment: As mentioned above, recent generators often produce highly polished images, while real photographs span a wider quality range. To reduce the shortcut that very high aesthetic quality implies fake, we matched the aesthetic score distribution of the real subset to that of the fake subset using the LAION-Aesthetic-Predictor [31]. The additional FLUX.1-dev samples further broaden the fake subset with more diverse photographic conditions, including motion blur, noise, and imperfect lighting. G.2 Label-Conditioned Candidate Annotation To guide the MLLM in generating high-quality explanations, we summarized several common AI generation errors that serve as strong, definitive evidence of synthetic origin. These include: (1) unrecognizable or garbled text; (2) missing, redundant, or distorted object components, including deformed faces or hands; (3) repetitive patterns, such as two identical faces or a single person holding two identical cups simultaneously; (4) anomalous lighting, such as areas appearing unnaturally bright without a clear light source; and (5) other violations of common sense or physical laws, such as objects floating in mid-air. We explicitly incorporated this information into the system prompt to make the model more attentive to these specific issues. For each image, we employed Gemini 3 Pro Preview [19] to gen- erate reasoning trajectories and evidence candidates. To improve annotation efficiency, we provided the ground-truth label of the im- age in the prompt. The system prompt used is detailed in Figure S11 and the user prompt is shown in Figure S12. Defake-o3 Figure S3: Examples of images from the GroundFake dataset. It is worth noting that we simultaneously instructed the model to generate Chinese explanatory text for the evidence to facilitate the annotation process, as all our human annotators are native Chinese speakers. G.3 Human Annotation and Verification In this stage, the evidence candidates generated for the fake im- ages in the previous step were further inspected and filtered by human annotators. This verification is necessary because the evi- dence generated by MLLMs frequently exhibits several issues: (1) The descriptive text is entirely irrelevant to the region enclosed by the bounding box; (2) The text contradicts the actual visual con- tent within the bounding box (e.g., text or hands that are rendered correctly are falsely claimed to be deformed); (3) The evidence is insufficient or overly sensitive to be deemed a definitive generative flaw (e.g., claiming “the character’s skin is too smooth,” which fre- quently occurs in real images due to post-processing). Meanwhile, the human-annotated data obtained at this stage was also used to train our Evidence Verifier, which learned how to identify such invalid evidence. During the annotation process, each piece of evidence was marked as either “Valid” or “Invalid”. Annotators were instructed that a piece of evidence should only be marked as “Valid” if the text ac- curately matches the image region within the bounding box AND correctly points out an error that is exclusive to AI-generated im- ages (i.e., not an artifact that could reasonably appear in real pho- tographs due to lighting constraints, simple post-processing, or motion blur). The annotation of each evidence item was strictly independent. Annotators were provided with the image, the bounding box, and the text in both English and Chinese. A total of 6 annotators partic- ipated in this process, with each piece of evidence independently evaluated by 3 of them. To ensure annotation quality, we organized the tasks into batches of approximately 1,000 evidence items. In each batch, we interspersed a certain proportion of evidence from real images, as well as forcibly generated bogus evidence on real images. We mandated that the “Valid” rate for these control items must not exceed 5% within any batch; otherwise, the entire batch was rejected and re-annotated. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang Table S6: Persuasion Evaluation Results under Prompt I. Method Qwen3-VL-235BKimi K2.5GLM-4.6V Hit@ImgHit@EviHit@ImgHit@EviHit@ImgHit@Evi Defake-o30.93470.88250.86430.72920.93970.9527 Defake-CoT0.95480.86840.88440.75210.95480.8974 Defake-Direct0.75380.71650.66830.58900.76880.6803 FakeVLM0.58950.47400.54740.48700.70530.7428 FakeShield0.31660.54200.31160.54650.32660.5646 Table S5: Impact of removing invalid evidence during train- ing. Invalid FilteringRL Training GroundFakeFakeFrontier AccIoUReal AccFake AccAcc ✗0.9840.2010.8360.9160.876 ✓✗0.9720.2410.8830.9040.894 ✓0.992 0.311 0.9020.945 0.923 Excluding the control items from real images used for quality checks, there was an average of 2.397 pieces of evidence to be an- notated per fake image. Under a majority voting rule, 74.99% of the evidence candidates were ultimately marked as “Valid”. Further- more, all 3 annotators reached a unanimous consensus on 89.43% of the evidence items. G.4 Reasoning Trajectory Rewriting for Annotation Consistency In this phase, we utilized Gemini 3 Flash Preview [18] to rewrite portions of the reasoning trajectories. This step ensures that the thought processes within the trajectories remain logically consis- tent with the filtered evidence (i.e., after the ‘Invalid‘ evidence was removed by human annotators). The system prompt used for this task is provided in Figure S13. In the user prompt for this task, we provided the original image, the uncorrected output from Gemini 3 Pro Preview, and the human annotation results presented as a boolean list. G.5 Impact of Human Verification on Classification Accuracy During the human annotation phase, evidence that did not align with human judgment was removed. The primary goal of this pro- cedure was to align the model’s explanations with human cognitive standards. Another intriguing question is: does this filtering process also affect the model’s overall classification accuracy? As shown in Table S5, removing human-annotated invalid evi- dence during the SFT stage understandably results in a higher IoU on the GroundFake test set, because the ground truth references have also been stripped of these invalid elements. Interestingly, this filtering also leads to an increase in overall accuracy on the OoD FakeFrontier benchmark. This improvement stems from a noticeable reduction in the false positive rate on real images. This demonstrates that removing invalid, overly sensitive, or halluci- nated evidence during training fundamentally helps reduce the probability of the model misclassifying real images. H More Details About FakeFrontier H.1 Image Sources and Generation Methods The 2,000 real images in the FakeFrontier benchmark are evenly sourced from OpenImageV7 [17] and Chameleon [64]. The 2,000 fake images are generated by 10 state-of-the-art generators: GPT- Image-1.5 [40], GPT-Image-1 [41], Seedream 4.5 [6], Seedream 3.0 [53], Z-Image-Turbo [7], Qwen-Image [60], Nano Banana [20], FLUX.2-pro [29], HunyuanImage 3.0 [8], and Stable Diffusion 3.5 Large [1]. For each generator, we produced 200 images. Specifi- cally, 100 images were generated using prompts derived from the captions of OpenImagesV7 samples, and the remaining 100 were generated using prompts based on Chameleon captions. We utilized Qwen3-VL-8B-Instruct to generate these captions, instructing the model to preserve the stylistic information (e.g., underexposure, motion blur) of the original real images. This approach ensures that the generated images mimic the diverse aesthetic conditions of real-world photography. The prompt used for generating image captions is detailed in Figure S14. During the image generation process, we employed several stan- dard resolutions: 1344×768, 1216×832, 1024×1024, 832×1216, and 768×1216. For each reference real image, its synthetic counter- part was generated using the standard resolution that most closely matched its aspect ratio. All other generation parameters were set to the models’ default or officially recommended configurations. H.2 MLLM-based Protocol We deployed Qwen3-VL-235B-A22B-Thinking, Kimi K2.5, and GLM- 4.6V locally to evaluate the explanation quality. During inference, all models used their default generation parameters. For the Quality Evaluation, we used the system prompt detailed in Figure S15. For the Persuasion Evaluation, since the susceptibility of Large Language Models to persuasion is highly dependent on the system prompt, we employed three distinct sets of system prompts. The final results presented in the main text are the average scores across these three prompts. Prompt I is detailed in Figure S16. The distinction in Prompt I (detailed in Figure S17) is the inclusion of common AI generation errors as background knowledge. The distinction in Prompt I (detailed in Figure S18) is a warning to the model that the provided evidence might be misleading, forcing it to rely more heavily on its autonomous judgment of the image content. This significantly increases the model’s skepticism; consequently, evidence that still manages to persuade the model under these conditions possesses remarkably high quality. H.3 Detailed Results of Persuasion Evaluation We provide the raw persuasion evaluation results under each of the three individual prompts for reference, as shown in Tables S6, S7, and S8. It can be observed that the difficulty of persuading the MLLM to deliver a “fake” verdict progressively increases from Prompt I to Prompt I. This indicates that evidence capable of per- suading the MLLM under the more stringent conditions of Prompt I is significantly more effective and reasonable. Moreover, across all different prompt configurations, Defake-o3 consistently main- tains a clear advantage over the other baseline methods. Defake-o3 Table S7: Persuasion Evaluation Results under Prompt I. Method Qwen3-VL-235BKimi K2.5GLM-4.6V Hit@ImgHit@EviHit@ImgHit@EviHit@ImgHit@Evi Defake-o30.89450.74550.83920.70800.94970.9462 Defake-CoT0.91460.76070.86930.71970.95980.8923 Defake-Direct0.69350.55430.64820.54490.76880.6787 FakeVLM0.41580.23840.53160.41760.66320.6561 FakeShield0.27640.37410.31160.48070.32660.5692 Table S8: Persuasion Evaluation Results under Prompt I. Method Qwen3-VL-235BKimi K2.5GLM-4.6V Hit@ImgHit@EviHit@ImgHit@EviHit@ImgHit@Evi Defake-o30.52260.28060.53770.33120.82410.5759 Defake-CoT0.45730.24270.38190.23930.77390.4855 Defake-Direct0.36680.18740.37690.21730.58790.3071 FakeVLM0.10530.05490.27890.15460.26840.1286 FakeShield0.10050.07260.20100.23810.17090.1429 H.4 Parsing Explanatory Text Our explanation quality evaluation is conducted at the granularity of individual “evidence” items. Defake-o3 naturally outputs struc- tured JSON with evidence separated into a list. However, baseline methods often output an unformatted block of text. Therefore, their explanations must be parsed and split before evaluation. We utilized Qwen3-VL-8B-Instruct with the prompt detailed in Figure S19 to achieve this. I Details About Baselines Here we briefly introduce each baseline method, along with their training details or the source of their weights used in our experi- ments. • CNNSpot [57] is a standard image classifier trained on real and ProGAN-generated images, with enhanced cross-generator generalization via preprocessing, postprocessing, and data aug- mentation. We trained it on GroundFake using the Adam optimizer with a learning rate of 1× 10 −4 and a batch size of 64 for 15 epochs. •NPR [52] models neighboring pixel relationships induced by generator up-sampling operations and uses these local structural artifacts as a source-invariant representation for generalizable fake image detection. We trained it on GroundFake using the Adam optimizer with a learning rate of 2× 10 −4 and a batch size of 32 for 10 epochs. •AIDE [64] combines CLIP-based global semantic embeddings with DCT-guided high/low-frequency patch sampling and SRM- based noise features to detect AI-generated images using hybrid high-level and low-level cues. We trained it on GroundFake using the AdamW optimizer with a learning rate of 1× 10 −4 and a batch size of 32 for 20 epochs. • FakeVLM [59] is a specialized multimodal large language model trained on the FakeClue dataset that performs synthetic image classification and generates natural-language artifact explana- tions from fine-grained visual clues. We used the pre-trained weights published on their official GitHub repository. •LEGION [26] integrates a global image encoder, an MLLM, a grounding image encoder, and a pixel decoder to jointly per- form synthetic image detection, artifact localization, and textual explanation. We reproduced LEGION on their SynthScars dataset using their official training scripts and default parameters. •FakeShield [63] combines a domain tag-guided explainable forgery detection module with a multimodal forgery localiza- tion module to predict image authenticity, localize manipulated regions, and explain tampering evidence in language. We used the pre-trained weights published on their official GitHub repository. J More Qualitative Results We provide additional qualitative results of Defake-o3 on the Ground- Fake test set, along with corresponding outputs from FakeVLM and FakeShield for a comprehensive comparison of explanation quality. As can be seen from the examples in Figures S4, S5, and S6, Defake-o3 is able to precisely localize specific generative flaws like unrecognizable objects, and easily identify surreal, obviously AI- generated scenes, such as the two penguins taking a selfie with a smartphone in the bottom example of Figure S5. In contrast, the explanations provided by the baseline methods are often extremely vague and unreliable. For instance, in 5 out of the 6 examples, FakeVLM outputs the exact same boilerplate explanation: "The image exhibits underlying characteristic incon- sistencies in its features that suggest it is artificially created." This indicates that FakeVLM struggles to pinpoint explicit flaws and instead makes unreliable judgments based merely on overall im- age texture. Meanwhile, FakeShield always attempts to explain its verdicts from the perspective of lighting and shadows, even when the lighting in the generated images does not appear prob- lematic to a human observer. Furthermore, in the bottom example of Figure S5, FakeShield mistakenly identifies the smartphone as a "mirror", demonstrating that its fine-tuning process may have compromised the model’s pre-existing world knowledge. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang Verdict: Fake Global Evidence: The overall image has a hyper-realistic, overly smooth texture... particularly in the way the light interacts with the floor and the plants. Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeVLM Verdict: Fake The natural light coming from the outside appears to be inconsistent in terms of brightness and color temperature... The shadows on the floor and wall appear to be inconsistent with the direction and intensity of the light... FakeShield Local Evidence #1: The small stools on the right side have nonsensical structures. The legs are thin, floating, or merging into the floor... Local Evidence #2: The object on the coffee table, presumably a plant or a decorative item, has a distorted, melted appearance. It lacks a clear structure... Verdict: Fake Global Evidence: The overall image has a hyper-realistic, overly smooth texture characteristic of AI generation, particularly in the reflections on the car body. Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeVLM Verdict: Fake The shadows and highlights on the car's hood and the background appear inconsistent with the natural lighting in the rest of the image... FakeShield Local Evidence #1: The logo on the hood is a distorted shield emblem with nonsensical internal details that do not resemble a real car manufacturer's logo. Local Evidence #2: The emblem on the front grille is a generic blue and silver circular logo that blends into the grille slats and lacks a coherent design. Figure S4: Additional qualitative comparison of Defake-o3 against FakeVLM and FakeShield on the GroundFake test set (Part 1). The bounding boxes are from the outputs of Defake-o3. Defake-o3 Verdict: Fake Global Evidence: The overall image has a hyper-realistic, overly smooth aesthetic typical of AI generation, particularly in the lighting and the perfect composition. Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeVLM Verdict: Fake The lighting on the bird does not match the rest of the scene. The bird appears to be lit from the front, while the rest of the scene suggests a light source coming from the right side, as indicated by the shadows on the cage... FakeShield Local Evidence #1: The metal hook connecting the cage to the wooden beam is nonsensical; it ... floats without a proper attachment point. Local Evidence #2: The object attached to the left side of the cage is a distorted cluster of texture that resembles seeds or a decorative ornament but lacks a logical structure. Verdict: Fake Global Evidence: The overall image depicts a surreal scenario of penguins taking a selfie with a smartphone, which is biologically impossible... Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The penguins have unnatural proportions, with their heads and beaks appearing disproportionately large compared to their bodies... FakeVLM Verdict: Fake ...The reflection in the mirror does not match the lighting of the actual penguin. The penguin's body is lit from the right side, casting a shadow to the left, while the reflection in the mirror shows a different angle of lighting... FakeShield Local Evidence #1: The penguin is holding a smartphone with a flip-style case or stand that extends like a limb. The hand holding the phone is anatomically nonsensical... Local Evidence #2: The mouths of both penguins are open in a way that defies natural expression. The interior of the mouths appears distorted, with teeth-like structures... Figure S5: Additional qualitative comparison of Defake-o3 against FakeVLM and FakeShield on the GroundFake test set (Part 2). Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang Verdict: Fake Global Evidence: The overall image has a hyper-realistic, overly smooth quality typical of AI generation, particularly in the lighting and the perfect arrangement of the food. Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeVLM Verdict: Real (×) The lighting in the image is consistent, with the shadows falling in a natural direction, indicating a single light source. The colors and textures of the food and tableware are realistic... FakeShield Local Evidence #1: There is a strange, white, dough-like object floating or standing on the black square dish in the center. It lacks a clear texture... Local Evidence #2: The cutlery on the plate is distorted. The handles of the spoons or utensils merge into the plate or each other in a nonsensical way... Verdict: Fake Global Evidence: The overall image has a slightly oversmoothed, 'waxy' texture common in AI generation, particularly in the skin and the way objects blend into the background. Input Image Defake-o3 Label: Fake #2 #1 Verdict: Fake The image exhibits underlying characteristic inconsistencies in its features that suggest it is artificially created. FakeVLM Verdict: Fake The proportions of the person's body are distorted, which is especially noticeable around the waistline... The pot and the person's arm appear to have a different exposure level compared to the surrounding environment... FakeShield Local Evidence #1: There is a long, straight metal rod sticking out of the pot that serves no clear purpose and defies physics, appearing to float... Local Evidence #2: The control panel on the microwave appliance is distorted, with nonsensical, melted-looking buttons and dials instead of functional controls. Figure S6: Additional qualitative comparison of Defake-o3 against FakeVLM and FakeShield on the GroundFake test set (Part 3). Defake-o3 You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will be presented with an image which may be AI generated. Please examine the image carefully, paying particular attention to small details. Finally, please provide the following in a JSON file: 1. Your evidence. This should be a list, with each element containing a piece of evidence you found that proves this image is AI-generated or not, consisting of a bounding box (in the format of [x1, y1, x2, y2]) and a piece of text. The bbox can be null, indicating that the evidence is for the entire image. For example, the overall texture of the image, whether it is oversaturated, etc. If the bbox is not null, it indicates that the area is key evidence. Only clear and decisive factors can be considered key evidence. For example, if a small area looks blurry, but in fact the real image may also look blurry in a small area due to image compression, this cannot be considered as critical evidence for AIGC. All your key evidence should be clear and unambiguous. You should always provide a piece of overall evidence with a null bbox, and zero or more pieces of key evidence with bounding boxes. If a piece of evidence requires multiple objects to be compared with each other (for example, two duplicate faces), the bbox should include all of them for direct comparison. If your verdict is "fake", only if you can't find any key evidence should you provide only a piece of general evidence with a null bbox. While your overall evidence can be slightly longer, your key evidence should be concise, just in one or two sentences. If your verdict is "real", just provide an overall evidence with null bbox; otherwise, for "fake" verdict, you can give extra key evidence to support your verdict. 2. Your final verdict, "fake" or "real". The following shows the json format: ```json "evidence": [ "bbox": null or [int, int, int, int], # note: the bbox for the first evidence is always null! "text": text , ... ], "verdict": "real" or "fake" ``` You should use the image cropping tool to examine each suspicious area of the image one by one. You can only use the tool once per round of conversation. Think step by step, and put your thoughts within the `<think></think>` tag. Place your tool calls and the JSON text block representing the final result outside the `<think></think>` tag. In each round of the conversation, you should first think, then use the tool once, or give the final answer. For example, a conversation might look like this: * <think> The image displays a large stadium, ... I am zooming in on the stadium to check the architectural details. </think> <tool_call>"name": "crop_image", " arguments": "x1": 150, "y1": 400, "x2": 780, "y2": 700</tool_call> * tool: here is the cropped image... * <think> I noticed some text-like features on the right facade of the stadium. I will zoom in to verify if the text is legible or if it consists of pseudo- characters... </think> <tool_call>"name": "crop_image", "arguments": "x1": 600, "y1": 500, "x2": 750, "y2": 600</tool_call> * tool: here is the cropped image... ... * <think> ... I have gathered enough evidence to make my final verdict. </think> ```json ... ``` Figure S7: System Prompt for Defake-o3. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will be presented with an image which may be AI generated. Please examine the image carefully, paying particular attention to small details. Finally, please provide the following in a JSON file: 1. Your evidence. This should be a list, with each element containing a piece of evidence you found that proves this image is AI-generated or not, consisting of a bounding box (in the format of [x1, y1, x2, y2]) and a piece of text. The bbox can be null, indicating that the evidence is for the entire image. For example, the overall texture of the image, whether it is oversaturated, etc. If the bbox is not null, it indicates that the area is key evidence. Only clear and decisive factors can be considered key evidence. For example, if a small area looks blurry, but in fact the real image may also look blurry in a small area due to image compression, this cannot be considered as critical evidence for AIGC. All your key evidence should be clear and unambiguous. You should always provide a piece of overall evidence with a null bbox, and zero or more pieces of key evidence with bounding boxes. If a piece of evidence requires multiple objects to be compared with each other (for example, two duplicate faces), the bbox should include all of them for direct comparison. If your verdict is "fake", only if you can't find any key evidence should you provide only a piece of general evidence with a null bbox. While your overall evidence can be slightly longer, your key evidence should be concise, just in one or two sentences. If your verdict is "real", just provide an overall evidence with null bbox; otherwise, for "fake" verdict, give extra key evidence to support your verdict. 2. Your final verdict, "fake" or "real". The following shows the json format: ```json "evidence": [ "bbox": null or [int, int, int, int], # note: the bbox for the first evidence is always null! "text": text , ... ], "verdict": "real" or "fake" ``` Think step by step, and put your thoughts within the `<think></think>` tag. Place your JSON text block representing the final result outside the `<think></think>` tag. Figure S8: System Prompt for Defake-CoT. Defake-o3 You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will be presented with an image which may be AI generated. Please examine the image carefully, paying particular attention to small details. Finally, please provide the following in a JSON file: 1. Your evidence. This should be a list, with each element containing a piece of evidence you found that proves this image is AI-generated or not, consisting of a bounding box (in the format of [x1, y1, x2, y2]) and a piece of text. The bbox can be null, indicating that the evidence is for the entire image. For example, the overall texture of the image, whether it is oversaturated, etc. If the bbox is not null, it indicates that the area is key evidence. Only clear and decisive factors can be considered key evidence. For example, if a small area looks blurry, but in fact the real image may also look blurry in a small area due to image compression, this cannot be considered as critical evidence for AIGC. All your key evidence should be clear and unambiguous. You should always provide a piece of overall evidence with a null bbox, and zero or more pieces of key evidence with bounding boxes. If a piece of evidence requires multiple objects to be compared with each other (for example, two duplicate faces), the bbox should include all of them for direct comparison. If your verdict is "fake", only if you can't find any key evidence should you provide only a piece of general evidence with a null bbox. While your overall evidence can be slightly longer, your key evidence should be concise, just in one or two sentences. If your verdict is "real", just provide an overall evidence with null bbox; otherwise, for "fake" verdict, give extra key evidence to support your verdict. 2. Your final verdict, "fake" or "real". The following shows the json format: ```json "evidence": [ "bbox": null or [int, int, int, int], # note: the bbox for the first evidence is always null! "text": text , ... ], "verdict": "real" or "fake" ``` You should directly give this JSON text block. Figure S9: System Prompt for Defake-Direct. Here is an image which may be AI generated. Please observe carefully and give your verdict and reasoning. Figure S10: User Prompt for Inference. Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will be presented with an image which may be AI generated. Please examine the image carefully, paying particular attention to small details. You should zoom in on specific areas for a closer look. Finally, please provide the following in a JSON file: 1. Your exploration trajectory. This should be a list, with each element containing the area you explored and your thought. Each area is given as a bounding box in the format [ymin, xmin, ymax, xmax] normalized to 0-1000. If you explored the entire image, the bbox should be null. At the end of each of your thoughts, you should mention which area you plan to explore next, or state that you are going to present your final results. You should explore the image thoroughly, always starting with the whole image (a null bbox), and then zoom in to up to 5 areas. 2. Your evidence. This should also be a list, with each element containing a piece of evidence you found that proves this image is AI-generated or not, also consisting of a bounding box and a piece of text. The bbox can be null, indicating that the evidence is for the entire image. For example, the overall texture of the image, whether it is oversaturated, etc. If the bbox is not null, it indicates that the area is key evidence. Only clear and decisive factors can be considered key evidence. For example, if a small area looks blurry, but in fact the real image may also look blurry in a small area due to image compression, this cannot be considered as critical evidence for AIGC. All your key evidence should be clear and unambiguous. You should always provide a piece of overall evidence with a null bbox, and zero or more pieces of key evidence with bounding boxes. If a piece of evidence requires multiple objects to be compared with each other (for example, two duplicate faces), the bbox should include all of them for direct comparison. If your verdict is "fake", only if you can't find any key evidence should you provide only a piece of general evidence with a null bbox. While your overall evidence can be slightly longer, your key evidence should be concise, just in one or two sentences. If your verdict is "real", just provide an overall evidence with null bbox; otherwise, for "fake" verdict, give extra key evidence to support your verdict. Additional requirement: Please provide a Chinese translation of the text. 3. Your final verdict, "fake" or "real". The following shows the json format: ```json "exploration_trajectory": [ "bbox": null or [int, int, int, int], # note: the bbox for the first exploration step is always null! "thought": text , ... ], "evidence": [ "bbox": null or [int, int, int, int], # note: the bbox for the first evidence is always null! "text": text, "chinese_text": text , ... ], "verdict": "real" or "fake" ``` Figure S11: System Prompt for Evidence Candidate Generation (GroundFake). Here is an image which may be AI generated. Please observe carefully and give your verdict and reasoning. Hint: The correct verdict should be'<label>'. But please behave as if you don't know this hint before you give your final verdict. Figure S12: User Prompt for Evidence Candidate Generation (GroundFake). Defake-o3 Here is a prompt for evaluating AI-generated images with a large language model: --- You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will be presented with an image which may be AI generated. Please examine the image carefully, paying particular attention to small details. You should zoom in on specific areas for a closer look. Finally, please provide the following in a JSON file: 1. Your exploration trajectory. This should be a list, with each element containing the area you explored and your thought. Each area is given as a bounding box in the format [ymin, xmin, ymax, xmax] normalized to 0-1000. If you explored the entire image, the bbox should be null. At the end of each of your thoughts, you should mention which area you plan to explore next, or state that you are going to present your final results. You should explore the image thoroughly, always starting with the whole image (a null bbox), and then zoom in to up to 5 areas. 2. Your evidence. This should also be a list, with each element containing a piece of evidence you found that proves this image is AI-generated or not, also consisting of a bounding box and a piece of text. The bbox can be null, indicating that the evidence is for the entire image. For example, the overall texture of the image, whether it is oversaturated, etc. If the bbox is not null, it indicates that the area is key evidence. Only clear and decisive factors can be considered key evidence. For example, if a small area looks blurry, but in fact the real image may also look blurry in a small area due to image compression, this cannot be considered as critical evidence for AIGC. All your key evidence should be clear and unambiguous. You should always provide a piece of overall evidence with a null bbox, and zero or more pieces of key evidence with bounding boxes. If a piece of evidence requires multiple objects to be compared with each other (for example, two duplicate faces), the bbox should include all of them for direct comparison. If your verdict is "fake", only if you can't find any key evidence should you provide only a piece of general evidence with a null bbox. While your overall evidence can be slightly longer, your key evidence should be concise, just in one or two sentences. If your verdict is "real", just provide an overall evidence with null bbox; otherwise, for "fake" verdict, give extra key evidence to support your verdict. Additional requirement: Please provide a Chinese translation of the text. 3. Your final verdict, "fake" or "real". The following shows the json format: ```json "exploration_trajectory": [ "bbox": null or [int, int, int, int], # note: the bbox for the first exploration step is always null! "thought": text , ... ], "evidence": [ "bbox": null or [int, int, int, int], # note: the bbox for the first evidence is always null! "text": text, "chinese_text": text , ... ], "verdict": "real" or "fake" ``` --- Note that this prompt is not for you. Instead, here is already the output from using this prompt with an image. You will see the image, the analysis result and a bool list. The analysis result was obtained using the above prompt. And the bool list is of the same length of the evidence list, indicating whether each item of the evidence is correct. Your goal is to correct the analysis result based on the bool list. For each item of the bool list, if it is true or null, you should leave the corresponding item of the evidence as it is. If it is false, it means that this is not key evidence of an AI-generated image according to expert evaluation. And you should remove this item from the evidence list. At the same time, you should update the thought in the exploration trajectory related to this evidence item, no longer considering it as decisive evidence. Also, you may or may not need to update the first evidence item (with null bbox) because this is the overall evidence for the entire image. Here are some examples: * Evidence: "The text on the red box is unrecognizable"; Marked as: False; Reason: The text is not clearly visible at the current resolution, regardless of whether it is a real image or an AI-generated image. * Evidence: "The person's hands are deformed"; Marked as: False; Reason: The person's hands look normal from the current angle. By the way, you should keep "verdict" unchanged. You should give the corrected analysis result in the same json format as above. Figure S13: System Prompt for Trajectory Rewriting (GroundFake). Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, and Jianfu Zhang You are an expert photo captioner. Describe the photo in natural language using at most two sentences. Include: 1. A concise description of the image's style or shooting conditions in one sentence. For example, "Professional photography", "Photograph with natural lighting and high depth of field", "Amateur photo from the 2000s, in dim lighting, taken with a flash", "Amateur photo, even lighting, casual composition", "Amateur photo, with slight motion blur, handheld camera". 2. The main subjects, their actions, the setting. Now give your caption directly. Figure S14: Prompt for Image Captioning (FakeFrontier). You will see an image which may be AI generated or not, along with a piece of evidence that attempts to prove that the image was generated by AI. Your task is not to judge the authenticity of the image, but rather to evaluate the quality of the provided evidence. You should determine whether the evidence is sufficient and convincing enough to support the claim that the image is AI-generated. You should rate the quality of the evidence on a scale of 1 (worst) to 10 (best) where 1 means the evidence is weak, irrelevant, or does not match the image at all; and 10 means the evidence is clear, specific, and directly supports the claim that the image is AI-generated. You should explain why you give the rating, and you should give your final rating in a pair of tags like this: <rating>1~10</rating> Figure S15: System Prompt for Quality Evaluation (FakeFrontier). You will see an image which may be AI generated or not, along with a piece of evidence that attempts to prove that the image was generated by AI. Your task is to determine whether the image was generated by AI, referring to this evidence. Provide your verdict and reasons; your final verdict should be given in a **JSON text block** like this: ```json "verdict": "real" or "fake" ``` Figure S16: Persuasion Evaluation Prompt I (FakeFrontier). You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will see an image which may be AI generated or not, along with a piece of evidence that attempts to prove that the image was generated by AI. Your task is to determine whether the image was generated by AI, referring to this evidence. Provide your verdict and reasons; your final verdict should be given in a **JSON text block** like this: ```json "verdict": "real" or "fake" ``` Figure S17: Persuasion Evaluation Prompt I (FakeFrontier). Defake-o3 You are an expert in the field of forged image detection. For reference, here are some common errors found in AI-generated images: 1. Unrecognizable text 2. Missing, redundant, or distorted object components, such as deformed face, deformed limbs, extra or missing limbs, including inconsistency with logos or unique features of well-known brands 3. Repeating patterns, such as two identical faces, one person holding two water cups, or a large number of people with unusually similar appearances 4. Unusual lighting, such as unusually bright areas in the absence of light sources 5. Other obvious anomalies that defy common sense or physics, such as objects floating in mid-air You will see an image which may be AI generated or not, along with a piece of evidence that attempts to prove that the image was generated by AI. Your task is to determine whether the image was generated by AI, referring to this evidence. Note that the evidence may be misleading, and you should make a verdict based on the actual content of the image. Provide your verdict and reasons; your final verdict should be given in a **JSON text block** like this: ```json "verdict": "real" or "fake" ``` Figure S18: Persuasion Evaluation Prompt I (FakeFrontier). You are an assistant that converts AIGC-detection model output text into JSON. Return ONLY a JSON object with this schema: "evidence": [ "text": "reason point 1", "text": "reason point 2" ] Rules: - Split the reasons into concise key points. - Keep each point factual and directly based on the input text. - Do not include bbox or any fields other than evidence[].text. - Use English. Figure S19: Prompt for Evidence Parsing.