Paper deep dive
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
Jianpeng Wang, Haoyu Wang, Baoying Chen, Jishen Zeng, Yiming Qin, Yiqi Yang, Zhongjie Ba
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/1/2026, 1:34:03 AM
Summary
PromptForge-350k is a large-scale dataset and ICL-Net is a novel forgery localization network designed to detect prompt-based AI image edits. The authors introduce an automated mask annotation framework using keypoint alignment and semantic similarity to generate ground-truth masks for 354,258 image pairs. ICL-Net utilizes a triple-stream backbone and intra-image contrastive learning to achieve state-of-the-art performance in identifying manipulated regions.
Entities (5)
Relation Signals (3)
ICL-Net → usestechnique → Intra-image Contrastive Learning
confidence 100% · ICL-Net, an effective forgery localization network featuring a triple-stream backbone and intra-image contrastive learning
ICL-Net → trainedon → PromptForge-350k
confidence 95% · Extensive experiments demonstrate that our method achieves an IoU of 62.5% on PromptForge-350k
PromptForge-350k → containsdatafrom → Nano-Banana
confidence 90% · PromptForge-350k... covering four state-of-the-art prompt-based AI image editing models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid democratization of prompt-based AI image editing has recently exacerbated the risks associated with malicious content fabrication and misinformation. However, forgery localization methods targeting these emerging editing techniques remain significantly under-explored. To bridge this gap, we first introduce a fully automated mask annotating framework that leverages keypoint alignment and semantic space similarity to generate precise ground-truth masks for edited regions. Based on this framework, we construct PromptForge-350k, a large-scale forgery localization dataset covering four state-of-the-art prompt-based AI image editing models, thereby mitigating the data scarcity in this domain. Furthermore, we propose ICL-Net, an effective forgery localization network featuring a triple-stream backbone and intra-image contrastive learning. This design enables the model to capture highly robust and generalizable forensic features. Extensive experiments demonstrate that our method achieves an IoU of 62.5% on PromptForge-350k, outperforming SOTA methods by 5.1%. Additionally, it exhibits strong robustness against common degradations with an IoU drop of less than 1%, and shows promising generalization capabilities on unseen editing models, achieving an average IoU of 41.5%.
Tags
Links
- Source: https://arxiv.org/abs/2603.29386v1
- Canonical: https://arxiv.org/abs/2603.29386v1
Trouble viewing inline? Open PDF directly →
Full Text
42,451 characters extracted from source content.
Expand or collapse full text
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization Jianpeng Wang, Haoyu Wang, Baoying Chen, Jishen Zeng, Yiming Qin, Yiqi Yang, Zhongjie Ba Abstract The rapid democratization of prompt-based AI image editing has recently exacerbated the risks associated with malicious content fabrication and misinformation. However, forgery local- ization methods targeting these emerging editing techniques remain significantly under-explored. To bridge this gap, we first introduce a fully automated mask annotating framework that lever- ages keypoint alignment and semantic space similarity to generate precise ground-truth masks for edited regions. Based on this framework, we construct PromptForge-350k, a large-scale forgery localization dataset covering four state-of-the-art prompt-based AI image editing models, thereby mitigating the data scarcity in this domain. Furthermore, we propose ICL-Net, an effective forgery localization network featuring a triple-stream backbone and intra-image contrastive learning. This design enables the model to capture highly robust and generalizable forensic features. Extensive experiments demonstrate that our method achieves an IoU of 62.5% on PromptForge-350k, out- performing SOTA methods by 5.1%. Additionally, it exhibits strong robustness against common degradations with an IoU drop of less than 1%, and shows promising generalization capabilities on unseen editing models, achieving an average IoU of 41.5%. Keywords: Multimedia Security, Image Forgery Localization, AI Security, AI Image Editing, Media Forensics 1 Introduction The rapid advancement of Generative AI technologies has made AI-powered image editing tools in- creasingly accessible and ubiquitous. By leveraging prompt-based image editing models, such as Flux.Kontext [13] and Nano-Banana [6], users can generate coherent manipulated images by merely providing an input image and an editing instruction. Nevertheless, this lowered barrier to image ma- nipulation has facilitated the widespread dissemination of forged content online, exacerbating risks associated with misinformation and fraud. Consequently, developing forgery localization methods tailored to these emerging image editing techniques has become critical. Despite this urgency, current research on image forgery localization lags significantly behind the rapid evolution of AI-powered image editing technologies. In terms of forgery localization datasets, existing datasets are primarily limited to traditional Photoshop [10, 18, 20] or mask-based AI editing [4, 5, 19], failing to cover emerging prompt-based techniques. Furthermore, since prompt-based models regenerate the entire image without explicit edited region masks, the absence of automated annotation methodologies poses a significant hurdle. Regarding forgery localization methods, existing approaches rely heavily on low-level artifacts such as splicing traces[8], sensor noise[1], and JPEG compression inconsistencies[36]. However, these subtle forensic traces are often disrupted or erased during the generative reconstruction process inherent to prompt-based editing. This necessitates the exploration of localization strategies that look beyond low-level signal patterns. In this work, as illustrated in Fig 1, we first introduce an automated mask annotation framework based on pixel-level alignment and semantic space feature similarity. This approach guarantees the accuracy of edited region mask annotations while eliminating the need for labor-intensive manual labeling. Building upon this framework, we construct PromptForge-350k, a comprehensive forgery localization dataset comprising 354,258 edited image pairs annotated with precise pixel-level edited region masks. This dataset encompasses diverse samples generated by four state-of-the-art open- source or closed-source image editing models (Nano-banana[6], BAGEL[7], Flux.Kontext[13], Step1x- 1 arXiv:2603.29386v1 [cs.CV] 31 Mar 2026 Forgery Localization Dataset Forgery Localization Network Mask Annotating Framework Pixel-Level Alignment Original Image Edited Image Edited Region Mask Semantic-Based Mask Annotation Input Image ICL-Net GT Mask Intra-Image Contrastive Loss Predicted Mask Original Image Edited Image Mask Figure 1: Overview of our work: (1) A fully automated mask annotating framework targeting prompt- based AI image editing. (2) PromptForge-350k, a comprehensive forgery localization dataset. (3) ICL-Net, an effective forgery localization network. edit[15]) across eight distinct editing tasks. Furthermore, we propose ICL-Net, a specialized forgery localization network that exploits Intra-image Contrastive Learning to effectively capture forensic traces in prompt-based AI edited images. Our contributions are summarized as follows: • We introduce a fully automated framework capable of generating precise edited-region masks for prompt-based AI forged images, circumventing the need for manual annotation. • We construct a comprehensive forgery localization dataset, PromptForge-350k, containing auto- matically annotated ground-truth masks, mitigating the data scarcity in this domain. • We propose ICL-Net, an effective forgery localization network driven by intra-image contrastive learning to capture robust and generalizable forensic features. Quantitative results demonstrate the effectiveness of our approach, which achieves an IoU of 62.5% on the PromptForge-350k dataset, surpassing existing SOTA methods by 5.1%. Notably, ICL-Net demonstrates high robustness against common image degradations, exhibiting a negligible IoU drop of less than 1%. It also generalizes well to unseen editing models, achieving an average IoU of 41.5%. Both the dataset and source code will be made publicly available to facilitate reproducibility and future research. 2 Related Works 2.1 AI Image Forgery Datasets Existing AI image forgery datasets can be categorized into two types based on the paradigms of the image editing methods used for their construction. Early mask-based editing methods require users to manually draw a mask over the editing region and provide instructions; the model then regenerates the content within the masked area while preserving the remaining regions. Conversely, the emerging prompt-based editing paradigm eliminates the dependency on explicit edited region masks, allowing 2 users to merely provide editing instructions via natural language. While mature dataset construction methodologies and large-scale datasets exist for mask-based approaches, forgery localization datasets tailored to prompt-based AI editing remain largely unexplored. Mask-based Editing Dataset. MagicBrush [33] utilizes manually annotated masks to perform image editing via DALL-E 2. TGIF [19] employes images and corresponding captions from MSCOCO as inputs for mask-based AI editing models. GRE [26], FakeShield [31], and GIM [5] adopt a similar approach: employing the SAM model to generate object masks and LLMs to synthesize editing instruc- tions. Additionally, BRGen [4] addresses background editing scenarios often overlooked by existing datasets, thereby enhancing dataset diversity. Prompt-based Editing Dataset. Several recent works have attempted to construct datasets tailored to prompt-based image editing. UltraEdit [34] employs a prompt-to-prompt strategy to control SDXL for generating edited images. X-Edit [2] utilizes InstructPix2Pix[3] to generate edited images, obtaining approximate edited region through pixel differencing. Both PicoBanana [22] and X2Edit [17] constructed datasets using state-of-the-art prompt-based editing methods, but lack edited region masks. 2.2 Image Forgery Localization Methods Image editing inevitably introduces pixel-level discrepancies between edited and non-edited regions, which serve as critical forensic cues. To exploit these discrepancies, the community has developed diverse forgery localization methodologies, which can be broadly categorized into three streams: Feature-Centric Approaches. Many works focus on high-frequency or statistical anomalies. MVSS-Net [8], PSCC-Net [16], and HiFiNet [12] incorporate noise-sensitive branches, dual-path struc- tures, and LoG filtering respectively to extract pixel-level artifacts. Similarly, PIM-Net [1] leverage CMOS noise patterns in non-edited regions to identify forged areas. Model Architecture Designs. Recent studies have introduced specialized architectures to ex- tract forgery features. APSC-Net [23] employs multi-scale extractors for holistic analysis. Sparse- ViT [25] performs attention within grouped feature maps, a concept further extended by NFA-ViT [4] via noise-guided amplification, to mitigate semantic interference. Additionally, Mesoscopic Insights [36] combines parallel CNN-Transformer encoders with DCT enhancement to capture mesoscopic features. Advanced Learning Paradigms. Beyond architectural innovations, advanced learning strate- gies have been adopted to enhance forgery localization efficacy. TruFor [11], NCL-IML [35], and FOCAL [29] harness contrastive learning, while CoDE [21] and AdaIFL [14] integrate reinforcement learning and dynamic routing. Furthermore, DiffForensics [32] and InPDiffusion [28] reformulate lo- calization as a mask generation task using conditional diffusion models. 3 Mask Annotating Framework In this section, we introduce the mask annotating framework employed for our PromptForge-350k dataset. Building upon the original and edited images from the X2Edit[17] and PicoBanana[22] datasets, we select samples associated with local editing tasks to undergo the following annotation process: Pixel-Level Alignment. We first perform pixel-level alignment between the original and edited images. This step is essential because certain image editing models (e.g., Nano-Banana and Kon- text) internally apply pixel cropping to input images, which would cause spatial misalignment and compromise the subsequent mask annotation process. Semantic-Based Mask Annotation. Subsequently, we leverage DINO v3 [27] to extract dense semantic features from the aligned images and compute the cosine similarity at corresponding spa- tial positions. This strategy effectively circumvents pixel-level variances introduced by prompt-based methods in non-edited regions, enabling robust mask annotation. 3.1 Pixel-Level Alignment The alignment process comprises three stages: keypoint extraction and matching, affine matrix estimation, transformation and cropping, as illustrated in Fig 2. For implementation details, we provide the pseudo code in Algorithm 1 of the supplementary material. 3 RANSAC Original Image Edited Image ORB BF Matcher Least Squares Affine Trans. Original Img Keypoints Edited Img Keypoints Keypoint Pairs Coarse Affine Matrix Inliers Masks Refined Affine Matrix Cropping Aligned Original Image Aligned Edited Image Affined Original Image Keypoint Extraction and MatchingAffine Matrix Estimation Transformation and Cropping Figure 2: Pipeline of Pixel-Level Alignment: The process involves keypoint extraction and matching, followed by affine matrix estimation. Finally, we apply a coordinate transformation, and the black borders are cropped out. Formally, as defined in Equation 1, we employ a 6-degree-of-freedom 2D affine to model geomet- ric distortions, including scaling, translation, and shearing. The parameters a 1 ,...,a 6 denote the transformation coefficients to be estimated. x ′ y ′ 1 = a 1 a 2 a 3 a 4 a 5 a 6 001 | z transformation matrixA x y 1 (1) The transformation matrix A defines the coordinate mapping between the original and edited images. Therefore, the primary objective of our pixel alignment process is to solve for this matrix A. 3.1.1 Keypoint Extraction and Matching. First, we employ the ORB algorithm [24] to extract 1,000 keypoints from the original and edited images, respectively. Each keypoint is characterized by its spatial coordinate and a 32-byte binary descriptor encoding local features. Keypoint matching between the two groups of keypoints is subsequently performed using a Brute-Force Matcher (BFMatcher) based on the Hamming distance. To ensure match reliability, we apply Lowe’s ratio test, discarding matches where the distance ratio between the nearest and second-nearest matching options exceeds 0.75. 3.1.2 Affine Matrix Estimation. Next, to handle potential noise and outliers in the matched keypoint pairs, we utilize the RANSAC [9] algorithm. This iterative algorithm robustly estimates the model parameters supported by the max- imum number of inliers. Specifically, applying RANSAC to the keypoint pairs yields a coarse trans- formation matrix A and a set of inlier indicators, denoted as idx inliers . This boolean list indicates whether a corresponding keypoint pair is consistent with the coarse matrix A (i.e., is an inlier). Lever- aging idx inliers , we select the inlier keypoint pairs and compute a refined affine transformation matrix, A refined , by solving a least-squares problem on this inlier subset. 3.1.3 Transformation and Cropping. Once A refined is obtained, we warp the original image to align it with the coordinate system of the edited image. Since affine transformations may introduce black borders at the image boundaries, we apply an identical cropping operation to both the transformed original image and the edited image to eliminate such regions. 3.2 Semantic-based Mask Annotation Prompt-based editing models typically induce global pixel-level changes, rendering naive pixel-level differencing ineffective for accurate mask estimation. Consequently, we propose utilizing semantic features from the pre-trained DINO v3 model [27] to localize edited regions based on semantic similarity. 4 Otsu Threshold Aligned Original Image Aligned Edited Image DINO v3 Cosine Similarity Original Img Features Edited Img Features Similarity Map Edited Region Mask Figure 3: Pipeline of Semantic-Based Mask Annotation. We first utilize DINO v3 to extract semantic features, followed by calculating the pixel-wise feature similarity. Finally, the similarity map is bina- rized to output the edited region mask. Specifically, as shown in Fig 3, we extract dense features from both the aligned original and edited images using the pre-trained DINO v3 Large model. We then compute the cosine similarity between features at corresponding spatial locations. To automate mask generation, we employ Otsu’s method on the similarity map to dynamically determine an optimal threshold. Regions with similarity scores below this threshold are classified as edited, while those above are considered non-edited. The detailed mask annotation pipeline is provided in Algorithm 2 of the Supplementary Material. 4 PromptForge-350k Dataset In this section, we first provide additional details regarding our dataset construction process. We then analyze the computational cost of the proposed annotation framework, and finally evaluate the annotation quality through a user study. 4.1 Dataset Construction We applied our proposed annotation pipeline to process image pairs from the X2Edit dataset [17] (com- prising subsets edited by BAGEL [7], Kontext [13], and Step1x [15]) and the PicoBanana dataset [22] (edited by Nano-Banana [6]). Given our focus on local editing, we filtered the datasets based on their “editing task” metadata, excluding global editing tasks such as style transfer, color tone conversion, and viewpoint changes. The final dataset includes eight distinct editing task types: object replace- ment, addition, removal, material conversion, pose change, text editing, portrait editing, and object state conversion. The left panel of Fig 4 illustrates the distribution of these editing tasks. To ensure dataset quality, we discarded samples meeting any of the following criteria: (1) Fewer than 10 keypoints extracted from either image, (2) Fewer than 10 matched keypoint pairs, or (3) An inlier ratio below 60%. Such samples typically correspond to cases where the edited region is excessively large, the image content is overly simple, or the editing operation failed. In total, the constructed dataset comprises 354,258 pairs, including 100,000 pairs from each of the three X2Edit subsets and 54,258 pairs from Picobanana. The dataset was randomly partitioned into training and test sets with a 95:5 split. Image resolutions range from 512×512 to 1024×1024, with edited region masks standardized to 128×128. 4.2 Computational Cost Analysis Dataset construction was performed on a server equipped with an Intel Xeon Platinum 8469C CPU and 8 NVIDIA H20 GPUs, taking approximately 103 hours. As illustrated in the right panel of Fig 4, we profiled the computational cost of each stage. Notably, our proposed pixel alignment method incurs a marginal overhead of 5%, while feature extraction accounts for 72% of the total runtime. To further enhance efficiency, employing semantic feature extraction models with faster inference speeds than DINO v3 Large could effectively accelerate the dataset construction process. 5 8% 4% 72% 5% 3% 9% Dataset Loading and Decoding Image Preprocessing Pixel Alignment Feature Extraction Dataset Writing and Encoding Others 23,639 16,157 50,402 43,321 27,191 45,921 78,101 69,526 Object Replacement Object Addition Object Removal Material conversion Pose Change Text Editing Portrait Editing Object State Conversion Figure 4: Statistics of the proposed dataset. Left: Distribution of editing task categories. Right: Time consumption breakdown of each operation during the dataset construction process. 4.3 Dataset Quality Assessment To assess the quality of the dataset, we compiled a set of 200 samples by randomly selecting 50 samples from each of the four subsets. Three volunteers were recruited to rate the annotation quality by assigning each sample to one of the following three categories: (1) Perfect: The annotation is precise or contains only negligible noise; (2) Minor Error: The annotation is largely accurate, with erroneous areas covering less than 20% of the edited region; and (3) Significnt Error: The annotation contains substantial errors. Discrepancies among the volunteers were resolved via majority voting. The final classification results were 172, 20, and 8, respectively. These results demonstrate that our annotation method yields reliable masks for the edited regions in the vast majority of cases. 5 ICL-Net Input image High-pass Filter Noise Backbone Main Backbone Frozen Backbone Fused Feature Channel Concat Edited Region Non-Edited Region Contrastive Loss Channel Concat Classification Head Seg. Loss Total Loss Index + Predicted Mask GT Mask Figure 5: Architecture of the proposed forgery localization network, ICL-Net. The network features three parallel backbones and is optimized via intra-image contrastive loss and segmentation loss. Unlike traditional Photoshop editing or mask-based AI editing, which operate on specific local regions, prompt-based editing involves a global re-synthesis process where the entire image is regen- erated by the model. Consequently, the forensic traces relied upon by traditional forgery localization methods (such as JPEG compression artifacts, splicing boundaries, upsampling anomalies, and CMOS sensor noise) are significantly diminished or obliterated during this process. This necessitates the development of forgery localization methods that are independent of these specific priors. We posit that in prompt-based editing, non-edited regions structurally and semantically preserve the original image content, whereas edited regions are synthesized entirely based on editing prompts and context. This discrepancy leads to subtle inconsistencies in textural details and semantic distributions between the edited and non-edited regions. Based on this insight, as illustrated in Fig 5, we introduce ICL-Net, a forgery localization network based on Intra-image Contrastive Learning, which guides the network to learn discriminative forgery features. ICL-Net enables precise localization of manipulated regions despite the absence of traditional forensic traces. 6 5.1 Network Structure Triple-Stream Backbone. ICL-Net comprises three parallel backbones with identical structures, all initialized with the same pre-trained SegFormer-B4 [30] model. Specifically, the Noise Backbone focuses on processing the high-frequency components of the input image, aiming to capture fine-grained textural artifacts associated with forgery. The Main Backbone focuses on extracting semantic-level forgery features. To preserve the generalizable knowledge learned from pre-training and mitigate catastrophic forgetting, we incorporate a Frozen Backbone, which shares the same initial parameters as the Main Backbone but keeps its parameters fixed during training. Classification Head. Features processed by the three backbones are concatenated along the channel dimension and fed into a trainable Classification Head. It consists of four groups of convolu- tional and upsampling layers, designed to increase the spatial resolution to match the ground truth mask while reducing the channel dimension to 1, ultimately generating the predicted mask. 5.2 Loss Function Intra-Image Contrastive Loss. Features maps from the Noise and Main backbones are con- catenated along the channel dimension to yield a fused feature. Guided by the ground truth mask, we spatially partition these fused features into two distinct sets: those corresponding to forged (edited) regions and those corresponding to real (non-edited) regions. Subsequently, a contrastive loss is com- puted to maximize the separability between these two feature categories. The formulation is as follows: L contrastive =− 1 n f n f X i=1 log 1 n f P n f j=1 exp(sim(F f [i],F f [j])/τ ) P n r k=1 exp(sim(F f [i],F r [k])/τ ) ! (2) In Equation 2, F f and F r denote the set of feature vectors corresponding to the forgery and real regions, respectively, while n f and n r represent the number of pixels in these regions. To balance computational cost and training efficiency, we randomly sample a subset of pixels if the number of candidates exceeds a predifined threshold during training. Additionally, sim(·) denotes cosine similar- ity, and τ is the temperature parameter. Segmentation Loss. The forgery localization task inherently suffers from severe class imbalance, as forged regions typically occupying less than 20% of the total image area. Relying solely on the Binary Cross-Entropy (BCE) loss commonly used in semantic segmentation can lead to suboptimal convergence. To address this, we employ a hybrid segmentation loss combining Focal Loss and Dice Loss. This combination effectively mitigates class imbalance and compels the network to focus on hard-to-classify classes, thereby enhancing training stability. Finally, the Contrastive Loss and Segmentation Loss are aggregated to form the Total Loss, which guides the network optimization. Let ˆ M and M denote the predicted mask and ground truth mask, respectively: L total = λ 1 L contrastive + λ 2 L Dice ( ˆ M,M ) + λ 3 L Focal ( ˆ M,M )(3) 6 Experiment 6.1 Experimental Setup Parameters. We resize input images to 512× 512 and apply random JPEG compression and cropping for training data augmentation. For the objective function, the loss weights are configured as λ 1 = 1,λ 2 = 4,λ 3 = 20. Training is conducted using the AdamW optimizer with an initial learning rate of 1e −4 . To ensure stable convergence, the learning rate is halved if the validation performance plateaus for three consecutive evaluations (performed every 800 iterations), with a minimum lower bound of 1e −8 . Metrics. We employ pixel-level F1 score, IoU (Intersection over Union), Precision, and Recall as evaluation metrics. Consistent with common practice, we report the F1 score for the forged class to focus specifically on forged regions. This is crucial because, due to significant class imbalance, a random guess would only achieve an F1 score of approximately 15%. 7 Table 1: Performance comparison between ICL-Net and SOTA methods. Our method achieves the best performance on all four subsets. The best and second-best results are highlighted in bold and underlined, respectively. MetricMethod Subset of PromptForge-350K Nano.BAGELKontextStep1xAverage F1 MVSS-Net31.8144.1954.2652.6145.72 TruFor45.3054.5870.2661.5058.61 FOCAL59.2855.4981.1171.0966.74 Mesorch50.1853.1874.9166.1261.10 NFA-ViT54.84 61.7677.1972.3066.52 ICL-Net (ours)65.7570.1584.2680.6275.20 IoU MVSS-Net22.4233.0242.2841.2234.74 TruFor35.4143.7259.6651.8047.65 FOCAL44.2540.25 68.7057.3252.63 Mesorch41.5444.1165.1957.2952.03 NFA-ViT46.0752.1967.8363.4657.39 ICL-Net (ours)51.4355.1874.2368.9762.45 Precision MVSS-Net32.6842.8549.5751.1444.06 TruFor46.8254.5268.6863.1858.30 FOCAL67.9665.8380.9975.2672.51 Mesorch61.2067.3279.6376.7771.23 NFA-ViT64.3869.2780.7177.4072.94 ICL-Net (ours)69.2773.7783.2781.4576.94 Recall MVSS-Net44.3858.1971.8264.5559.74 TruFor53.8063.7877.6666.3065.39 FOCAL53.5349.0181.4768.1263.03 Mesorch49.6650.4474.7063.8459.66 NFA-ViT54.7161.9976.9871.7766.36 ICL-Net (ours)63.5069.3085.7180.0274.63 Comparision Methods. We benchmark our ICL-Net against five state-of-the-art forgery local- ization methods: MVSS-Net (PAMI 2022)[8], TruFor (CVPR 2023)[11], FOCAL (TDSC 2025) [29], Mesorch (AAAI 2025)[36], and NFA-ViT (AAAI 2026)[4]. To ensure a fair comparison, we retrain all baseline methods on our dataset, strictly adhering to the training protocols and hyperparameters detailed in their respective papers. 6.2 Main Results As presented in Tab 1, ICL-Net outperforms all competing methods, achieving an average pixel-level F1 score of 75.20% and an IoU of 62.45%. Performance on Different Subsets. Among the four editing methods, the closed-source Nano- Banana subset poses the greatest challenge. We attribute this to the fact that the network architectures and training data of closed-source models likely diverge significantly from open-source alternatives, resulting in distinct and more elusive forgery traces. Analysis of Competitors. Among the comparison methods, NFA-ViT, Mesorch, and FOCAL exhibit relatively compatitive performance. We reason that while these methods focus on enhancing high-frequency forgery traces, they do not explicitly suppress semantics features. In contrast, MVSS- Net and TruFor employ mechanisms specifically designed to decouple semantic information to avoid interference. While this strategy enhances generalization for traditional splicing or copy-move forgery, it proves detrimental for prompt-based editing localization, where semantic inconsistencies are crucial cues. Qualitative Analysis. Qualitative comparisons are visualized in Fig 6. The first two rows display the original and edited images, respectively. The third row illustrates the overlay of predicted masks 8 Nano-BananaBAGELKontextStep1x Original Edited Mask Figure 6: Subjective evaluation results of ICL-Net. The third row highlights the predicted forged regions in red, the ground-truth edited regions in green, and the overlap between predictions and ground truth in yellow. Table 2: Robustness analysis. We evaluate ICL-Net under varying degrees of JPEG compression and random cropping. PerturbationFactorF1IoUPrecisionRecall ICL-Net (baseline)None75.2062.4576.9474.13 ICL-Net w/ JPEG Compression 9074.8361.9274.7475.68 8074.4561.3774.8974.60 7073.4960.2375.2672.29 6072.2458.8275.7369.58 ICL-Net w/ Random Crop 10%76.8264.1777.4876.60 20%76.4963.6776.9976.45 30%74.6861.3573.8176.22 40%69.2255.1168.8871.06 ICL-Netw/ JPEG+Crop J80 + C20%74.6161.5674.9375.06 (red) and ground truth masks (green), with the intersection (yellow) indicating correct predictions. These visualizations qualitatively validate the superior precision of our proposed method. 6.3 Robustness Analysis To evaluate the ICL-Net’s robustness against real-world degradations, we applied JPEG compression and random cropping to the test set, to simulate artifacts introduced during social media dissemination. As quantitative results in Tab 2 demonstrate, under simultaneous perturbations of JPEG compression (quality factor 80) and 20% cropping, the F1 score and IoU of our network decreased by a mere 0.58% and 0.86%, respectively. This indicates that our method does not heavily rely on fragile forgery traces susceptible to being distorted by image degradations, highlighting its promising potential for real-world social network applications. 6.4 Cross-Model Generalization To assess cross-model generalization, we adopted a Leave-One-Out (LOO) strategy, where one subset was excluded from training and subsequently used for evaluation. As illustrated in Tab 3, perfor- 9 Table 3: Cross-model generalization performance. We exclude one subset from the training set and evaluate the network’s ability to generalize to unseen editing methods. SubsetSettingF1IoUPrecisionRecall In-domain65.7551.4369.2763.50 Nano-Banana Out-of-domain31.0718.9535.2529.49 In-domain70.1555.1873.7767.30 BAGEL Out-of-domain51.0736.3964.6743.63 In-domain84.2674.2383.2785.71 Kontext Out-of-domain68.4453.5375.1063.93 In-domain80.6268.9781.4580.02 Step1x Out-of-domain70.9357.1277.8265.41 In-domain75.2062.4576.9474.13 Average Out-of- domain 55.3741.5063.2150.62 mance declines across all unseen editing methods. A notable trend is that the drop in Recall is more pronounced than in Precision, suggesting that ICL-Net tends to yield conservative predictions when encountering forged images from unseen models. For the three open-source editing models, our approach achieves an IoU ranging from 36.39% to 57.12%, even though they were unseen during the training phase. However, for Nano-Banana, ICL-Net attains an IoU of only 18.95%, significantly lower than the generalization performance observed on the three open-source counterparts. This indicates that the forgery artifacts of closed-source methods differ significantly from those of open-source methods. Consequently, forgery localization networks trained exclusively on open-source models exhibit limitations when generalizing to closed-source methods. 6.5 Ablation Studies We conducted ablation studies to validate the effectiveness of our key design components: the con- trastive loss, the noise backbone and the frozen backbone. The results presented in Tab 4 demonstrate that each component plays a critical role. In particular, the exclusion of the contrastive loss leads to a significant performance drop, with F1 and IoU declining by 15.6% and 17.38%, respectively. This indicates that the contrastive loss is instrumental in guiding the network to differentiate between edited and non-edited regions, thereby ensuring the network learns discriminative features. Table 4: Ablation study: removing the contrastive loss, noise backbone, or the frozen backbone each leads to a noticeable drop in localization performance. OperationF1IoUPrecisionRecall ICL-Net (baseline)75.2062.4576.9474.13 ICL-Net w/o contrastive loss59.6345.0760.4559.55 ICL-Net w/o noise backbone73.2159.8471.5775.43 ICL-Net w/o frozen backbone70.4456.9575.6166.81 7 Limitations and Future Work Limitations of the Annotation Pipeline. As discussed in Section 4.3, our dataset annotation method yields incorrect masks in a small fraction of edge cases. Manual inspection reveals that these mislabeled samples are primarily concentrated in scenarios involving solely object color modifications, such as changing an apple from red to green. This is likely attributed to the DINO v3 model’s relatively 10 low sensitivity to object color, as it prioritizes semantic information. Future work could address this by incorporating additional discriminative cues for joint decision-making or by fine-tuning the DINO v3 model to further enhance annotation accuracy. Generalization Gap. Although our forgery localization method demonstrates satisfactory per- formance, its generalization capability to unseen closed-source models remains limited. Enhancing cross-model generalization, especially for closed-source models, represents a critical direction for fu- ture research. 8 Conclusion In this work, we address the critical yet underexplored challenge of image forgery localization for prompt-based AI editing. We present PromptForge-350k, a large-scale dataset constructed via our fully automated annotating pipeline. Furthermore, we propose ICL-Net, an effective forgery localiza- tion network that incorporates intra-image contrastive learning. Extensive experiments validate the superiority of our approach, which achieves state-of-the-art performance on PromptForge-350k while demonstrating resilience against common post-processing perturbations. We hope that PromptForge- 350k and ICL-Net will serve as a solid foundation to facilitate future research in the growing field of prompt-based AI image forgery localization. References [1] Ningning Bai, Xiaofeng Wang, Ruidong Han, Jianpeng Hou, Yihang Wang, and Shanmin Pang. PIM-Net: Progressive Inconsistency Mining Network for image manipulation localization. Pattern Recognition, 159:111136, March 2025. [2] Valentina Bazyleva, Nicol`o Bonettini, and Gaurav Bharaj. X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2754–2764, 2025. [3] Tim Brooks, Aleksander Holynski, and Alexei A. Efros. InstructPix2Pix: Learning To Follow Image Editing Instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023. [4] Lvpan Cai, Haowei Wang, Jiayi Ji, Yanshu Zhoumen, Shen Chen, Taiping Yao, and Xiaoshuai Sun. Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach, December 2025. [5] Yirui Chen, Xudong Huang, Quan Zhang, Wei Li, Mingjian Zhu, Qiangyu Yan, Simiao Li, Hanting Chen, Hailin Hu, Jie Yang, Wei Liu, and Jie Hu. GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization. The Thirty-Ninth AAAI Conference on Artificial Intelligence, 2025. [6] Gheorghe Comanici, Eric Bieber, and Schaekerman. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities, December 2025. [7] Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Wei- hao Yu, Xiaonan Nie, Ziang Song, Guang Shi, and Haoqi Fan. Emerging Properties in Unified Multimodal Pretraining, July 2025. [8] Chengbo Dong, Xinru Chen, Ruohan Hu, Juan Cao, and Xirong Li. MVSS-Net: Multi-View Multi- Scale Supervised Networks for Image Manipulation Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3):3539–3553, March 2023. [9] Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM, 24(6):381–395, June 1981. 11 [10] Haiying Guan, Mark Kozak, Eric Robertson, Yooyoung Lee, Amy N Yates, Andrew Delgado, Daniel Zhou, Timothee Kheyrkhah, Jeff Smith, and Jonathan Fiscus. Mfc datasets: Large-scale benchmark datasets for media forensic challenge evaluation. In 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW), pages 63–72. IEEE, 2019. [11] Fabrizio Guillaro, Davide Cozzolino, Avneesh Sud, Nicholas Dufour, and Luisa Verdoliva. TruFor: Leveraging all-round clues for trustworthy image forgery detection and localization, May 2023. [12] Xiao Guo, Xiaohong Liu, Zhiyuan Ren, Steven Grosz, Iacopo Masi, and Xiaoming Liu. Hierarchi- cal Fine-Grained Image Forgery Detection and Localization. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3155–3165, Vancouver, BC, Canada, June 2023. IEEE. [13] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas M ̈uller, Dustin Podell, Robin Rombach, Harry Saini, Axel Sauer, and Luke Smith. FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space, June 2025. [14] Yuxi Li, Fuyuan Cheng, Wangbo Yu, Guangshuo Wang, Guibo Luo, and Yuesheng Zhu. AdaIFL: Adaptive Image Forgery Localization via a Dynamic and Importance-Aware Transformer Net- work. In Aleˇs Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G ̈ul Varol, editors, Computer Vision – ECCV 2024, pages 477–493, Cham, 2025. Springer Nature Switzerland. [15] Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, Guopeng Li, Yuang Peng, Quan Sun, Jingwei Wu, Yan Cai, Zheng Ge, Ranchen Ming, Lei Xia, Xianfang Zeng, Yibo Zhu, Binxing Jiao, Xiangyu Zhang, Gang Yu, and Daxin Jiang. Step1X-Edit: A Practical Framework for General Image Editing, July 2025. [16] Xiaohong Liu, Yaojie Liu, Jun Chen, and Xiaoming Liu. PSCC-Net: Progressive Spatio-Channel Correlation Network for Image Manipulation Detection and Localization, August 2022. [17] Jian Ma, Xujie Zhu, Zihao Pan, Qirong Peng, Xu Guo, Chen Chen, and Haonan Lu. X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning, November 2025. [18] Ga ̈el Mahfoudi, Badr Tajini, Florent Retraint, Frederic Morain-Nicolier, Jean Luc Dugelay, and Marc Pic. Defacto: Image and face manipulation dataset. In 2019 27Th european signal processing conference (EUSIPCO), pages 1–5. IEEE, 2019. [19] Hannes Mareen, Dimitrios Karageorgiou, Glenn Van Wallendael, Peter Lambert, and Symeon Papadopoulos. TGIF: Text-Guided Inpainting Forgery Dataset. In 2024 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6, December 2024. [20] Adam Novozamsky, Babak Mahdian, and Stanislav Saic. IMD2020: A Large-Scale Annotated Dataset Tailored for Detecting Manipulated Images. In 2020 IEEE Winter Applications of Com- puter Vision Workshops (WACVW), pages 71–80, Snowmass Village, CO, USA, March 2020. IEEE. [21] Rongxuan Peng, Shunquan Tan, Xianbo Mo, Bin Li, and Jiwu Huang. Employing Reinforcement Learning to Construct a Decision-Making Environment for Image Forgery Localization. IEEE Transactions on Information Forensics and Security, 19:4820–4834, 2024. [22] Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, and Zhe Gan. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing, October 2025. [23] Chenfan Qu, Yiwu Zhong, Chongyu Liu, Guitao Xu, Dezhi Peng, Fengjun Guo, and Lianwen Jin. Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel Methods. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10781–10790, Seattle, WA, USA, June 2024. IEEE. 12 [24] Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ORB: An efficient alternative to SIFT or SURF. In 2011 International Conference on Computer Vision, pages 2564–2571, November 2011. [25] Lei Su, Xiaochen Ma, Xuekang Zhu, Chaoqun Niu, Zeyu Lei, and Ji-Zhe Zhou. Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding Transformer. Proceedings of the AAAI Conference on Artificial Intelligence, 39(7):7024–7032, April 2025. [26] Zhihao Sun, Haipeng Fang, Juan Cao, Xinying Zhao, and Danding Wang. Rethinking Image Editing Detection in the Era of Generative AI Revolution. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 3538–3547, Melbourne VIC Australia, October 2024. ACM. [27] Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ̈el Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timoth ́e Darcet, Th ́eo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Herv ́e J ́egou, Patrick Labatut, and Piotr Bojanowski. DINOv3, August 2025. [28] Kai Wang, Shaozhang Niu, Qixian Hao, and Jiwei Zhang. InpDiffusion: Image Inpainting Lo- calization via Conditional Diffusion Models. Proceedings of the AAAI Conference on Artificial Intelligence, 39(7):7771–7779, April 2025. [29] Haiwei Wu, Yiming Chen, Jiantao Zhou, and Yuanman Li. Rethinking Image Forgery Detection via Soft Contrastive Learning and Unsupervised Clustering, May 2025. [30] Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. Seg- Former: Simple and Efficient Design for Semantic Segmentation with Transformers. In Advances in Neural Information Processing Systems, volume 34, pages 12077–12090. Curran Associates, Inc., 2021. [31] Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, and Jian Zhang. FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models, April 2025. [32] Zeqin Yu, Jiangqun Ni, Yuzhen Lin, Haoyi Deng, and Bin Li. DiffForensics: Leveraging Diffusion Prior to Image Forgery Detection and Localization. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12765–12774, Seattle, WA, USA, June 2024. IEEE. [33] Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. MAGICBRUSH: A Manually Anno- tated Dataset for Instruction-Guided Image Editing. Advances in Neural Information Processing Systems (NeurIPS), 2023. [34] Haozhe Zhao, Xiaojian Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang. UltraEdit: Instruction-based Fine-Grained Image Editing at Scale. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. [35] Jizhe Zhou, Xiaochen Ma, Xia Du, Ahmed Y. Alhammadi, and Wentao Feng. Pre-training-free Image Manipulation Localization through Non-Mutually Exclusive Contrastive Learning. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 22289–22299, Paris, France, October 2023. IEEE. [36] Xuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang, Bo Du, Xiwen Wang, Zeyu Lei, Wentao Feng, Chi-Man Pun, and Ji-Zhe Zhou. Mesoscopic Insights: Orchestrating Multi-Scale & Hy- brid Architecture for Image Manipulation Localization. The Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI), 2025. 13