Paper deep dive
FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking
Muñoz-Haro Javier, Teruel Andres, Tolosana Ruben, DeAlcala Daniel, Vera-Rodriguez Ruben, Morales Aythami, Fierrez Julian
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 8/4/2026, 10:53:52 AM
Summary
The paper introduces FakeIDet3-DB, a comprehensive database of digital manipulations on real, government-issued identity documents to address the domain gap caused by privacy regulations. It proposes PACE, a privacy-aware patch extraction algorithm using Integral Image mapping and Non-Maximum Suppression to generate 5.2M patches from 6.4K images while preventing PII leakage. The study evaluates state-of-the-art forensic models, revealing significant challenges in detecting and localizing both classical and Generative AI-driven attacks.
Entities (14)
Relation Signals (10)
FakeIDet3-DB → contains → Digital Manipulations
confidence 95% · FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations
PACE → uses → Integral Image mapping
confidence 92% · PACE... leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS).
PACE → uses → Non-Maximum Suppression
confidence 92% · PACE... leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS).
PACE → prevents → PII leakage
confidence 90% · PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage
FakeIDet3-DB → complieswith → GDPR
confidence 88% · to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches.
Generative AI → enables → Face-Swapping
confidence 85% · Generative AI-driven manipulations (e.g., face-swapping, inpainting)
Generative AI → enables → Inpainting
confidence 85% · Generative AI-driven manipulations (e.g., face-swapping, inpainting)
Latent Diffusion Models → usedfor → Text Inpainting
confidence 83% · we employ Latent Diffusion Models [10] for text inpainting
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we introduce FakeIDet3-DB, the first comprehensive database of digital manipulations on real, government-issued IDs. FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations (e.g., face-swapping, inpainting) enhanced with advanced image refinement procedures to suppress visual artifacts. In addition, to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches. In order to maximize forensic utility, we formulate privacy-aware patch extraction from a real ID as a geometrically constrained image processing problem. We propose PACE, a Pseudo-Anonymized Contextual patch Extraction algorithm, which leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS). PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage while maximizing semantic density in peri-censorship regions, yielding almost 5.2M patches extracted from more than 6.4K images from real/fake IDs. Furthermore, an extensive evaluation of the proposed FakeIDet3-DB is performed using state-of-the-art models, showcasing they all struggle to detect and locate attacks coming from generative and classic techniques (32.45\% EER in detection and 83.48\% AUC-ROC in localization).
Tags
Links
- Source: https://arxiv.org/abs/2607.26641v1
- Canonical: https://arxiv.org/abs/2607.26641v1
Trouble viewing inline? Open PDF directly →
Full Text
71,449 characters extracted from source content.
Expand or collapse full text
KYC Know Your Customer ID Identity Document AI Artificial Intelligence OVD Optical Variable Device OCR Optical Character Recognition PAI Presentation Attack Instrument PII Personally Identifiable Information BID Brazilian Identity Document PVC Polyvinyl Chloride PAD Presentation Attack Detection IFDL Image Forgery Detection and Localization CNN Convolutional Neural Network GAN Generative Adversarial Network LoRA Low-Rank Adaptation ULD Unconditional Latent Diffusion LDM Latent Diffusion Model SAM-2 Segment Anything Model 2 EER Equal Error Rate GIMP GNU Image Manipulation Program ROI Region Of Interest FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking Javier Muñoz-Haro, Andres Teruel, Ruben Tolosana, Daniel DeAlcala, Ruben Vera-Rodriguez, Aythami Morales, Julian Fierrez All authors are with School of Engineering, Universidad Autónoma de Madrid, Ciudad Universitaria de Cantoblanco, Madrid, Spain.Corresponding author: Javier Muñoz-Haro (javier.munnoz@uam.es). Abstract Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we introduce FakeIDet3-DB, the first comprehensive database of digital manipulations on real, government-issued IDs. FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations (e.g., face-swapping, inpainting) enhanced with advanced image refinement procedures to suppress visual artifacts. In addition, to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches. In order to maximize forensic utility, we formulate privacy-aware patch extraction from a real ID as a geometrically constrained image processing problem. We propose PACE, a Pseudo-Anonymized Contextual patch Extraction algorithm, which leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS). PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage while maximizing semantic density in peri-censorship regions, yielding almost 5.2M patches extracted from more than 6.4K images from real/fake IDs. Furthermore, an extensive evaluation of the proposed FakeIDet3-DB is performed using state-of-the-art models, showcasing they all struggle to detect and locate attacks coming from generative and classic techniques (32.45% EER in detection and 83.48% AUC-ROC in localization). Finally, we establish a privacy-compliant benchmark for both forgery detection and localization, demonstrating the challenges introduced by patch-based processing and real-world ID attacks. FakeIDet3-DB and the corresponding benchmark code are avaible at https://github.com/BiometricsAI/FakeIDet3-DB. Index Terms: FakeIDet3-DB, Benchmark, Digital Attacks, Generative AI, Fake Identity Document, Algorithm. I Introduction Figure 1: An overview of the main contributions of the present article. The left panel illustrates the attack generation pipeline considered in FakeIDet3-DB, integrating both classical methods and GenAI models which undergo a refinement post-processing to produce highly-realistic attacks. Right panel depicts the proposed PACE, an efficient, privacy-aware patch extraction algorithm that leverages Integral Images and a distance-aware greedy NMS to maximize residual sensitive information extraction around redacted sections without compromising any PII from the ID owner. The proliferation of digital services has driven the transition from physical identity verification to remote, image-based authentication ecosystems. In standard Know Your Customer (KYC) protocols, users are required to submit smartphone-captured images of their physical identity document (ID). From an image processing perspective, authenticating these documents is a highly complex task, as the system must reliably analyze intricate spatial structures—such as Optical Variable Devices (e.g., holograms, precise laser engravings)—under unconstrained acquisition conditions. This vulnerability is frequently exploited through injection attacks, where malicious actors bypass the camera sensor to directly feed manipulated ID images into the verification pipeline. Concurrently, the rapid evolution of Generative AI (GenAI) models has fundamentally altered the landscape of digital image manipulation [28]. Rather than simply synthesizing images from scratch, these AI-driven tools can perform targeted, high-fidelity semantic alterations within existing visual content. Applied to ID tampering, techniques such as face-swapping [39], morphing [40], or text inpainting [21] act as localized signal perturbations. To evade detection, these GenAI attacks are often coupled with sophisticated image refinement procedures—such as feathering or illumination correction—that suppress tampering artifacts [41] and seamlessly align the forged regions with the pristine source distribution. Developing robust image forensic models [13, 42] to detect these subtle structural anomalies in IDs is hindered by a critical data acquisition bottleneck. Due to strict privacy regulations, existing ID databases rely almost exclusively on synthetic, laboratory-created templates [20]. These templates lack the rich, high-frequency security patterns with complex and delicate visual forms from legitimately manufactured government IDs (e.g., fine-grained engravings). Consequently, forensic models suffer from a severe domain shift: they achieve high accuracy on in-domain synthetic distributions but fail drastically when attempting to generalize to the topological and structural complexity of real-world manipulations [19, 38]. To overcome these fundamental limitations, we introduce FakeIDet3-DB, the first database of digital manipulations applied to real, government-issued IDs. Adopting a recent privacy-aware framework [23, 24], FakeIDet3-DB strictly complies with data protection regulations by applying irreversible modifications—pitch-black anonymization masks—and extracting patches from non-redacted areas, thus preventing any Personally Identifiable Information (PII) leakage. While anonymization solves the privacy bottleneck, it introduces severe visual artifacts that hinder standard feature extraction. To address this, we formulate patch extraction as a geometrically constrained image processing problem. We propose an optimized, context-aware extraction algorithm that utilizes Integral Image mapping ((1)O(1) local complexity) [18] and distance-driven Non-Maximum Suppression (NMS). This algorithmic approach effectively contours the complex geometries of the anonymization masks, targeting the highly informative peri-censorship regions where residual sensitive data and manipulation artifacts typically reside, while strictly guaranteeing zero censorship contamination. The main contributions of this article are summarized in Fig. 1 and as follows: • FakeIDet3-DB: The most comprehensive ID database to date for digital manipulations. Sourced from 250 real images of 47 distinct government-issued IDs, we generated 8 different attack typologies, resulting in over 6.4K images of real/fake IDs. These range from Classical Manipulations (e.g., copy-move, splicing) to GenAI Manipulations (e.g., face-swapping, morphing, or text inpainting). Crucially, these attacks are enhanced with advanced image refinement procedures to minimize visual artifacts and simulate high-quality, real-world forgeries. • Privacy-Aware Patch Extraction Algorithm: We introduce a Pseudo-Anonymized Contextual patch Extraction algorithm (PACE). By leveraging global spatial optimization rather than naive uniform grids, our algorithm significantly increases the semantic density of the extracted dataset, maximizing the retention of semantically-relevant artifacts located at the boundaries of occluded regions. As a result, we extract over 5.2M patches from pseudo-anonymized ID images at both 64×64 and 128×128 patch sizes, following our previous framework [24]. • Standardized Benchmark: To foster this challenging line of research within the image processing and forensics community, we design and release a reproducible benchmark111https://github.com/BiometricsAI/FakeIDet3-DB for both the detection and spatial localization of fake IDs, providing access to real-world document evaluations without compromising PII. The remainder of this article is structured as follows. Section I reviews current ID databases, competitions and patch extraction methodologies, addressing their primary limitations. Section I introduces FakeIDet3-DB, detailing the image manipulation and refinement processes. Section IV formulates the proposed context-aware patch extraction algorithm under spatial constraints. In Section V, we evaluate state-of-the-art media forensics architectures against FakeIDet3-DB, and finally, Section VI draws the conclusions of this work. I Related Work I-A Fake ID Databases and Competitions Table I depicts the landscape of fake ID detection databases, which has historically relied on proxy representations due to strict privacy constraints regarding real IDs. Early databases focused on physical forgeries, utilizing either laboratory-manufactured replicas, such as the KID34K dataset [27], or digital templates printed on Polyvinyl Chloride (PVC), like the MIDV database family and DLC-2021 [6, 5, 1, 29]. More recently, addressing the rise of digital tampering, several works have introduced entirely synthetic databases (e.g., IDNet [44], providing over 800K generated samples, and DASAC [22]) or digitally manipulated fictitious templates (e.g., FantasyID [20]). While these databases offer large-scale volume and diverse attack typologies, they inherently lack the rich, high-frequency security features of real, government-issued IDs. To tackle the sensitive nature of real IDs, FakeIDet2-db [24] proposed a patch-based framework utilizing pitch-black redacted areas over ID owners’ sensitive data. However, its scope was exclusively limited to physical forgeries, leaving challenging digital manipulations on real IDs unexplored. The fundamental limitation of relying on synthetic or lab-created proxies as “real” IDs is the severe domain shift they introduce. Recent international evaluations have severely exposed this vulnerability. In the Second PAD-ID Competition [38], while the winning solution achieved a 6.36% Equal Error Rate (EER) by leveraging a massive proprietary database, runner-up models evaluated against government-issued real IDs suffered severe performance degradation (EERs of 23.87% and 31.94%) compared to evaluations on lab-made proxies. Similarly, in the DeepID challenge [19], forensic models trained on synthetic data achieved a near-perfect F1 score of 0.99 on synthetic test sets, but their performance plummeted to 0.72 when evaluated on government-issued real IDs provided by industry partners (PXL Vision). This substantial performance degradation underscores the urgent need for public, open-source databases featuring government-issued real IDs to democratize research and bridge the gap between laboratory proxies and real fraud. TABLE I: Main characteristics of public fake ID databases in the literature. We point out whether the databases contain physical or digital attacks. Refined Attacks refer to databases containing attacks that undergone a post-processing or fine-tuning procedure to create high-quality manipulations. Database #Samples #IDs #Devices Attack Types #Attacks Government-Issued Real IDs? Refined Attacks? Public? DLC-2021 (2021) [29] 1,424 (videos) 1,000 2 Physical 4 ✗ N/A ✓ Benalcazar et al. (2023) [4] 3,000 3,000 N/A Digital 1 ✗ ✗ ✓ KID34K (2023) [27] 34,662 92 12 Physical 3 ✗ N/A ✓ IDNet (2024) [44] 837,060 837,060 Synthetic Digital 7 ✗ ✗ ✓ FantasyID (2025) [20] 3,284 362 3 Digital 3 ✗ ✓ ✓ DASAC (2025) [22] 150,833 1,472 3 Digital, Physical 5 ✗ ✗ ✓ FakeIDet2-db (2026) [24] 2,000 47 3 Physical 3 ✓ N/A ✓ FakeIDet3-DB (proposed) 6,436 47 3 Digital 8 ✓ ✓ ✓ I-B Geometrical Constrained Patch Extraction Spatial patch extraction has evolved significantly beyond rigid uniform grids [24]. PatchMatch [3] pioneered stochastic perturbation for rapid spatial mapping. In medical and geospatial domains, Kochkarev et al. [18] introduced probabilistic sampling via distance transforms and integral images to balance rare classes. Additionally, in-context learning frameworks like PatchICL [25] leverage boundary-guided Gumbel-Top-K sampling for selective processing. However, all these patch extraction methods lack the deterministic, zero-overlap topological constraints mandated by privacy and data regulation laws, which is a critical gap addressed by our proposed distance-driven patch extraction framework. Figure 2: Visual results of the digital attacks included in FakeIDet3-DB including different types of manipulations and post-processing techniques. It is noticeable how refined attacks remove clear artifacts left by text inpainting models, seamlessly blending background with foreground text. For face manipulations, we observe that for GenAI face morphing attacks, our refined pipeline contours the morphed subject and faithfully blends it with the background canvas. Note that for content-removal attacks, no post-processing is needed. I FakeIDet3-DB: Attacks Generation As demonstrated in previous section, the persistent lack of government-issued real IDs in research databases introduces a severe domain gap that limits real-world applicability. To overcome this fundamental barrier, we introduce the FakeIDet3-DB database. By leveraging proven privacy-aware frameworks [23, 24], this database securely utilizes real ID data without compromising PII. FakeIDet3-DB includes two main types of digital attacks. Classical Attacks use existing content and classical algorithms (e.g., copy-move, splicing, and landmark-based face morphing). Conversely, GenAI Attacks leverage generative models to alter real content. Both types undergo a refined post-processing pipeline to faithfully embed the modified elements, producing attacks with varying degrees of visual realism. Fig. 2 shows attacks generated under identical conditions with and without our proposed refinement pipeline (i.e., refined vs. cheapfake). I-A Digital Attacks Formulation To generate the digital attacks, we utilize the 250 real IDs introduced in FakeIDet2-db [24]. The synthesis is modeled as a set of image-domain transformations, categorized into classical signal-processing attacks and GenAI-driven manipulations. I-A1 Classical Attacks Classical attacks rely on the spatial redistribution of existing pristine pixels. These manipulations are divided into text-based and face-based transformations. For text-based manipulations, the methodology consists of a spatial mapping of a source text region ℬsrcB_src from a source image IsrcI_src into a target bounding box ℬtgtB_tgt in a target image ItgtI_tgt. To ensure the forgery remains visually pristine, semantic compatibility is strictly enforced (e.g., date fields are exclusively mapped to date fields). We formulate two widely recognized spatial transformations [12]: • Copy-Move (Intra-document): The manipulation is constrained such that Isrc=ItgtI_src=I_tgt. Furthermore, fine-grained attacks are generated by restricting ℬsrcB_src to sub-character topological regions (e.g., duplicating a single digit), introducing high-precision spatial alterations. • Splicing (Inter-document): The manipulation maps content between disjoint image domains, where Isrc≠ItgtI_src≠ I_tgt. Face-based manipulations encompass portrait swapping and face morphing. Portrait swapping functions as a facial splicing attack (inter-document mapping of the portrait region), ensuring gender-matched identities. For face morphing, we utilize open-source frameworks222https://github.com/alyssaq/face_morpher,333https://github.com/Azmarie/Face-Morphing/ based on classical geometric warping. Given a source face, denoted as FsrcF_src and a target face, denoted as FtgtF_tgt, facial landmarks LsrcL_src and LtgtL_tgt are extracted. Delaunay triangulation [40] is used to compute a spatial affine warping function W, aligning both faces to a mean geometry. The morphed face FmorphF_morph is then synthesized via linear alpha blending: Fmorph=λ⋅(Fsrc,Lsrc,Ltgt)+(1−λ)⋅(Ftgt,Ltgt,Lsrc)F_morph=λ·W(F_src,L_src,L_tgt)+(1-λ)·W(F_tgt,L_tgt,L_src) (1) where λ∈[0,1]λ∈[0,1] controls the blending factor. Figure 3: Proposed refined post-processing pipeline for text digital attacks. The GenAI pipeline (top, light-blue background) involves GenAI methods (Φtext _text) to synthesize the new text, while also removes the target text area (Φinpaint _inpaint), creating IbgI_bg. Finally, an alpha-blending operation αt _t from the extracted text mask is applied to embed the new content. The classical pipeline (bottom, light-green background) leverages the existing content (IsrcI_src) along with the target text removed image (IbgI_bg) and embeds the source content crop fakeT_fake into the ItgtI_tgt in a similar fashion as the GenAI pipeline (i.e., alpha-blending operation). I-A2 GenAI Attacks In generative AI-driven digital attacks, the forgery process is conceptualized as a conditional generative function. To ensure high-quality synthesis, we employ Latent Diffusion Models [10] for text inpainting (DiffSTE [17], UDiffText [49], TextDiffuser 2 [9], RS-STE [11]), face morphing (DiffMorpher [48], FreeMorph [7]) and face swapping (REFace [2]). We also rely on Generative Adversarial Networks [14] for face swapping (FaceDancer [31], InsightFace444https://github.com/deepinsight/insightface) and content removal (LaMa-Inpaint [36]). This multi-model strategy endows FakeIDet3-DB with a wide spectrum of visual manipulation qualities and distinct generation fingerprints [43, 34, 26], forcing detection models to learn robust, architecture-agnostic features. For text-based manipulations, the generative process is modeled as an inpainting function Φtext _text. We select a target bounding box ℬtgtB_tgt and create a binary mask MℬM_B along with a contextual text prompt P. Crucially, we account for the structural limitations of current text-inpainting LDMs, which struggle with character-space mismatches. Therefore, we impose a strict character-length constraint between the original legitimate string and the synthesized prompt: ||=|original text||P|=|original text|. For names and surnames, we sample replacements from a Hispanic demographic database. For the ID number, valid modulo-23 sequences are algorithmically generated. To provide sufficient receptive field for font and texture replication, an expanded center crop Ttgt⊂ItgtT_tgt⊂ I_tgt (e.g., 768×768768× 768) is processed: Tfake=Φtext(Ttgt,,Mℬ)T_fake= _text(T_tgt,P,M_B) (2) The newly generated text crop TfakeT_fake is subsequently extracted for downstream integration. For face manipulations, portrait crops are either automatically or manually extracted from the target ItgtI_tgt and source IsrcI_src images, yielding Fsrc⊂IsrcF_src⊂ I_src and Ftgt⊂ItgtF_tgt⊂ I_tgt. The generative model Φface _face maps the source and target characteristics with the operation Ffake=Φface(Fsrc,Ftgt)F_fake= _face(F_src,F_tgt) to generate a realistic synthetic face FfakeF_fake, which is reserved for subsequent seamless blending. Furthermore, we introduce a Content Removal manipulation, which acts as a purely subtractive alteration. Using the dataset’s metadata, a random field is targeted, and its spatial binary mask MℬM_B is generated. Using an inpainting network Φinpaint _inpaint (e.g., LaMa-Inpaint [36]), we apply the operation Iremoved=Φinpaint(Itgt,Mℬ)I_removed= _inpaint(I_tgt,M_B), which removes the foreground semantic information while explicitly emulating the underlying high-frequency security patterns. In both text and face manipulations, synthesizing the raw content (TfakeT_fake or FfakeF_fake) is only the first step. The realistic embedding of these forgered pixels into the complex visual domain of a real ID document is critical to create a faithful attack. The following section details the post-processing pipelines designed to achieve this seamless integration. Figure 4: Proposed refined post-processing pipeline for face digital attacks. Given two images, source (IsrcI_src) and target (ItgtI_tgt), a portrait photo crop is applied to both images to restrict the context to the ID owner’s face. Both images are then forwarded into a generic face manipulation function Φface _face, which can either be classical algorithms or GenAI models. Once the manipulated FfakeF_fake image is generated, a binary mask locating the areas of the new face is produced by Φseg_face _seg\_face. In parallel, the ID owner face is segmented, obtaining a mask that conditions the Φinpaint _inpaint function to remove the portrait photo area, yielding IbgI_bg. Finally, an alpha-blending operation αf _f embeds the new manipulated face in IbgI_bg, creating the final IrefinedI_refined. I-B Post-Processing Refinement Pipeline We describe next how the manipulated content generated in previous section is seamlessly embedded into the digital attacks. In this work, cheapfake refers to low-effort forgeries through a naive spatial substitution strategy between target (ItgtI_tgt) and source (IsrcI_src). The source content (coming from a generative model or existing content) is scaled and placed directly over the target using a static bounding region via hard-pixel replacement without any post-processing. Because this generic masking forces the source content into a rigid spatial constraint rather than respecting the semantic contours, it introduces unnatural boundary transitions. In the following sections, we formalize the post-processing techniques employed to create manipulations with varying degrees of visual realism. The pipeline diverges slightly depending on the manipulation modality (text vs. face). I-B1 Text-based Manipulations To overcome the visual discrepancies originated from cheapfake attacks, we introduce a refined post-processing pipeline (Fig. 3) for high-fidelity manipulations. This framework models text integration as a semantic alpha-blending problem. First, to prevent structural artifacts from the original text, the pipeline reconstructs the background canvas. Let Mℬ∈0,1H×WM_B∈\0,1\^H× W be the binary mask of the target field. We process ItgtI_tgt and MℬM_B through an inpainting network Φinpaint _inpaint (e.g., LaMa-Inpaint [36]) applying the operation Ibg=Φinpaint(Itgt,Mℬ)I_bg= _inpaint(I_tgt,M_B) to generate the underlying document patterns (e.g., security grids, fine-grained engravings, etc). Concurrently, let TfakeT_fake denote the text crop synthesized by the generative model. To isolate the typographic ink from any mismatched generated background, we apply a high-fidelity text segmentation network Φseg_text _seg\_text (e.g., Hi-SAM [45]) operating in a sliding-window patch mode, Mstroke=Φseg_text(Tfake)M_stroke= _seg\_text(T_fake), where Mstroke∈0,1H×WM_stroke∈\0,1\^H× W is a pixel-level text stroke mask. Unlike hard-pixel masking, text strokes require subtle edge softening to blend naturally with the document’s resolution. We apply a Gaussian smoothing kernel GσtG_ _t (e.g., a 3×33× 3 kernel) to generate a soft alpha matte αt∈[0,1]H×W _t∈[0,1]^H× W: αt=Mstroke∗Gσt _t=M_stroke G_ _t (3) The final composition is driven by the Hadamard product (⊙ ), transferring only the feathered typographic ink onto the clean background at the target Region Of Interest (ROI): Irefined=αt⊙Tfake+(−αt)⊙IbgI_refined= _t T_fake+(1- _t) I_bg (4) This integration strictly preserves the authentic security features of IbgI_bg, yielding sophisticated forgeries that significantly increase forgery detection and localization difficulty. I-B2 Face-based Manipulations To mitigate structural artifacts characteristic from cheapfake, we propose a semantic-aware refined pipeline for face manipulations, depicted in Fig. 4. This framework abandons rigid bounding boxes in favor of dynamic semantic segmentation and spatial alpha matting. First, we isolate the original subject to reconstruct the background canvas. We obtain a precise mask of the real portrait photo area Morig∈0,1H×WM_orig∈\0,1\^H× W using a promptable segmentation model Φseg_face _seg\_face (e.g., SAM-2 [30]). To guarantee the complete removal of the original facial contour (including hair) and prevent peripheral artifacts during background reconstruction, we apply a massive morphological dilation operation to MorigM_orig using a large structuring element K (e.g., a 75×7575× 75 kernel), denoted as Merase=Morig⊕KM_erase=M_orig K. The background canvas is then generated as Ibg=Φinpaint(I,Merase)I_bg= _inpaint(I,M_erase). Concurrently, the pipeline semantically segments the generative or classic model’s raw output FfakeF_fake to extract the precise manipulated contours, discarding peripheral mismatched pixels: Mfake=Φseg_face(Ffake)M_fake= _seg\_face(F_fake). To ensure a seamless transition between the synthesized skin/hair and the document background, MfakeM_fake is softened into an alpha matte αf∈[0,1]H×W _f∈[0,1]^H× W via a Gaussian kernel GσfG_ _f (e.g., a 5×55× 5 kernel): αf=Mfake∗Gσf _f=M_fake G_ _f (5) The final content-aware composition is achieved via alpha blending: Irefined=αf⊙Ffake+(−αf)⊙IbgI_refined= _f F_fake+(1- _f) I_bg (6) By replacing rigid templates with semantic isolation and background inpainting, this formulation effectively eliminates sharp boundaries, yielding highly realistic visual forgeries. Moreover, we subject a single authentic ID to multiple simultaneous refined manipulations to simulate a comprehensive identity forgery scenario. In this context, an attacker compromises several fields concurrently—for instance, by replacing textual personal data in conjunction with face morphing manipulations of the target subject. We define such forgeries as multi-attacks. IV FakeIDet3-DB: Proposed Patch Extraction Methodology 1 Input: Image ∈ℝH×W×CI ^H× W× C, Anon. Mask A∈0,1H×WM_A∈\0,1\^H× W, GT Mask GT∈0,1H×WM_GT∈\0,1\^H× W, Patch Size (h,w)(h,w), Evaluation Stride s Output: Extracted Patches ℰE_I, Extracted Ground Truth Masks ℰGTE_GT 2 3ℰ←∅,ℰGT←∅E_I← , _GT← , ←∅C← , ←∅V← // 1. (1)O(1) Validity Checking via Integral Images 4 (x,y)←∑x′≤x,y′≤yA(x′,y′)I(x,y)← _x ≤ x,y ≤ yM_A(x ,y ) ; 5 6for (x,y)∈[0,W−w]×[0,H−h](x,y)∈[0,W-w]×[0,H-h] to with step s do 7 S←(x+w,y+h)−(x,y+h)−(x+w,y)+(x,y)S (x+w,y+h)-I(x,y+h)-I(x+w,y)+I(x,y); 8 if S=0S=0 then ←∪(x,y)V ∪\(x,y)\ ; // Mark as valid coordinate 9 10 11 // 2. Distance-Aware Prioritization 12 ←DistanceTransform(1−A)D (1-M_A); 13 ←Sort(,key←(x+w2,y+h2),order←ascending)C (V,key (x+ w2,y+ h2),order ); 14 // 3. Greedy Spatial Non-Maximum Suppression (NMS) 15 ←H×WO 0^H× W 16for (x,y)∈(x,y) do 17 if ∑[y:y+h,x:x+w]=0 [y:y+h,x:x+w]=0 then [y:y+h,x:x+w]←1O[y:y+h,x:x+w]← 1 ; // Mark area as occupied 18 ℰ←ℰ∪[y:y+h,x:x+w]E_I _I∪\I[y:y+h,x:x+w]\; 19 ℰGT←ℰGT∪GT[y:y+h,x:x+w]E_GT _GT∪\M_GT[y:y+h,x:x+w]\; 20 21 22return ℰ,ℰGTE_I,E_GT; Algorithm 1 Proposed PACE Patch Extraction Sharing complete, uncensored real IDs introduces severe privacy and regulatory risks. To mitigate this vulnerability, patch-based processing [16, 8] enables a secure collaboration paradigm between data-rich ID Holders and technology-rich AI Researchers as introduced in [24]. Under this privacy-aware framework, ID Holders locally censor sensitive information, extract valid patches, and shuffle them to completely dismantle the document’s spatial layout. Furthermore, ID censoring can be applied at different levels. As per [24], sensitive data can be completely occluded (i.e., fully-anonymized ID) or partially occluded (i.e,. pseudo-anonymized ID) leaving small, residual sections of the sensitive data uncovered which do not compromise the owners’ PII. Consequently, AI models can be safely trained (e.g., through self-supervised learning) on these strictly anonymized patches before being deployed back to the ID Holders for secure, local inference. While this paradigm successfully protects privacy, anonymization artifacts (i.e., pitch-black rectangles) severely poison feature extraction. Naive uniform grid algorithms used in [24] fail to address this topological challenge, suffering from both semantic dilution (over-representing uninformative background) and censorship contamination (inadvertently accepting partially redacted cells). Because censored areas mimic digital manipulations, training on contaminated patches heavily degrades attack detection. Conversely, in pseudo-anonymized configurations [24], the most discriminative forensic features—such as residual sensitive data and local manipulation artifacts—reside strictly at the immediate boundaries of these occluded regions as shown in Fig. 5. To overcome these limitations, we propose an optimized, context-aware patch extraction approach (Algorithm 1), summarized visually in Fig. 1 (right panel). Our proposal considers a geometrically constrained greedy strategy designed to precisely target and maximize peri-censorship information while strictly guaranteeing zero overlap with the redacted pixels. Figure 5: Visual example of a pseudo-anonymized ID, where redacted sections make practically impossible to infer the sensitive information from the owner, thus preventing PII leaking. Uncovered sections are denoted as residual sensitive data, which are important for targeted patch extraction. IV-A Proposed Patch Extraction Method: PACE The enforced objective for the proposed PACE algorithm is a constrained spatial packing problem: we must extract the maximum number of valid patches that are as strictly close to the anonymization mask AM_A as possible, without any overlap with the mask itself or between the patches. Naively evaluating every possible patch candidate involves nested loops with a computational complexity of (H⋅W⋅h⋅w)O(H· W· h· w), where H and W are the height and width of the image and h and w are the height and width of the patches. As this is intractable for large-scale datasets, we bypass this bottleneck via a three-stage optimization pipeline, described next. IV-A1 (1)O(1) Validity Checking via Integral Images To instantaneously determine if a proposed patch contains any forbidden (anonymized) pixels, we compute the Summed-Area Table, or Integral Image I [18], of the binary anonymization mask AM_A. The value at any location (x,y)(x,y) in I represents the sum of all pixels above and to the left of that coordinate. Consequently, the sum of pixels S (i.e., the cost) within any rectangular region bound by (x,y)(x,y) and (x+w,y+h)(x+w,y+h) can be computed in (1)O(1) time using exactly four array references: S=(x+w,y+h)+(x,y)−(x,y+h)−(x+w,y)S=I(x+w,y+h)+I(x,y)-I(x,y+h)-I(x+w,y) (7) If the cost is zero (i.e., S=0S=0), then the patch is considered strictly valid (V). This reduces the search complexity for valid coordinates from (h⋅w)O(h· w) per patch to (1)O(1). Furthermore, we introduce the stride parameter s, which allows to skip adjacent coordinates in both x and y axes. Since close coordinates contain highly correlated content, this reduces the number of valid coordinates by a factor of s2s^2. IV-A2 Distance-Aware Prioritization Identifying valid patches is insufficient, as those physically adjacent to the occlusions (i.e., those containing residual sensitive information) must be prioritized. We apply a Distance Transform [33, 32] to the inverted anonymization mask (1−A)(1-M_A) . This operation yields a topographical proximity map ∈ℝH×WD ^H× W, where each scalar value denotes the exact Euclidean or Manhattan distance to the nearest censored boundary. We assign a priority score to each valid candidate coordinate (x,y)∈(x,y) by sampling the Distance Transform at the geometric centroid of the proposed patch: (x+w2,y+h2)D(x+ w2,y+ h2). The candidate set is then strictly sorted in ascending order of these centroid distances. This sorting operation introduces a computational bottleneck when evaluated along all valid coordinates, specially for very high-resolution images. However, thanks to the stride parameter s, the set of valid coordinates is reduced, alleviating heavy computation at a minor drop in semantically extracted information (more details in Section IV-B). Figure 6: Visual examples of different patch extractions regarding ROI Coverage (CRoIC_RoI) and Semantic Density (DsemD_sem) proposed metrics. It can be observed that, as both metrics increase, the patches contain more pixels from sensitive regions, which is related to the DsemD_sem metric, while also covering the most part of residual sensitive data, which relates to the CRoIC_RoI metric. IV-A3 Greedy Spatial Non-Maximum Suppression (NMS) With the candidate list C sorted by boundary proximity, we must select the final patches while preventing spatial redundancy. We implement a Greedy Spatial Non-Maximum Suppression (NMS) utilizing a boolean occupancy grid ∈0,1H×WO∈\0,1\^H× W, initialized to zeros. Iterating through the sorted list, a patch is accepted if and only if its entire spatial grid maps to zeros in O (∑[y:y+h,x:x+w]=0 [y:y+h,x:x+w]=0). Upon acceptance, this grid is updated to ones (occupied), effectively claiming the area. Because the list is ordered by distance, this greedy strategy forces the algorithm to “crystallize” the patches around the perimeter of the anonymized regions first. Once the immediate border is saturated, subsequent valid patches naturally form consecutive outward layers. Finally, for every selected coordinate, the corresponding visual data ℰE_I is extracted from the image I, while the identical spatial footprint ℰGTE_GT is extracted from the Ground Truth manipulation mask GTM_GT. For real ID images, GTM_GT acts as a tensor of zeros. This parallel extraction guarantees perfect spatial alignment between the image domain and the supervisory signal, optimizing the data pipeline for subsequent model training. TABLE I: Spatial Yield and Semantic Density Comparison on 250 Real Pseudo-Anonymized IDs for 128 × 128 patch size. Best result is highlighted in bold. Second best result is underlined. Method ROI Cov. (%) ↑ Sem.Dens. (px/patch) ↑ Time/Img. (s) ↓ Naive Grid [24] 63.41 3,045 0.003 Kochkarev et al. [18] 52.76 3,435 0.061 PatchMatch [3] 45.13 3,765 0.077 PACE (s=1s=1) 73.16 3,804 6.127 PACE (s=2s=2) 73.64 3,805 1.626 PACE (s=8s=8) 73.15 3,753 0.135 IV-B PACE: Spatial Yield and Semantic Density Evaluation To ensure a rigorous comparative analysis under strict privacy constraints, all evaluated patch-extraction algorithms inherently enforce an absolute zero-overlap policy with the anonymization mask via (1)O(1) integral image validation. We contrast PACE against topological adaptations of prominent state-of-the-art samplers. Specifically, the methodology by Kochkarev et al. [18] is implemented as a probabilistic sampler weighted by the normalized inverse distance transform. Concurrently, the stochastic local search inspired by the random search phase of the PatchMatch algorithm [3] is also considered as a competing method for optimal patch selection. This unified framework isolates the geometric efficacy of each sampling strategy, evaluating their capacity to securely patch peri-sensitive regions, where the residual sensitive data lies. We define two metadata-driven metrics to evaluate extraction efficacy: ROI Coverage (CRoIC_RoI) and Semantic Density (DsemD_sem). CRoIC_RoI measures the global percentage of critical forensic pixels successfully retrieved, whereas DsemD_sem evaluates the average amount of informative pixels packed into each patch, serving as a direct indicator of residual sensitive data coverage and purity. Let ℛR denote the set of exposed ROI pixels, defined as the ground-truth ID bounding boxes excluding the anonymization mask. Given a set of N valid extracted patches ℰ=E1,E2,…,ENE=\E_1,E_2,…,E_N\ extracted by an algorithm, the metrics are formalized as: CRoI(%)=1|ℛ|∑n=1N|Ei∩ℛ|×100,Dsem=1N∑n=1N|Ei∩ℛ|C_RoI(\%)= 1|R| _n=1^N|E_i |× 100, D_sem= 1N _n=1^N|E_i | where |En∩ℛ||E_n | represents the exact number of pristine, sensitive pixels captured by the n-th patch. Fig. 6 depicts these metrics visually. On the bottom left, we see a patch extraction where the semantic density (DsemD_sem) and ROI Coverage (CRoIC_RoI) are low, as patches do not overlap with any of the ROI regions and no sensitive pixels are contained in the region the patch encompasses. The top left image shows the case of high CRoIC_RoI and low DsemD_sem, as patches cover the majority of the residual sensitive data, but for each patch, a fewer amount of sensitive pixels are contained. Conversely, in the bottom right image we see that patches contain a great amount of pixels belonging to residual sensitive data, but they do not cover them for the most part. Finally, the top right image showcases an ideal extraction, where patches are aligned with the residual sensitive data pixels, therefore producing purer patches, while also covering practically the whole area of residual sensitive information. Figure 7: Comparison in terms of valid extracted 128×128 patches (green), contaminated patches (orange) vs non-valid patches (red) in pseudo-anonymized configurations between the Naive Uniform Grid Method [24], Kochkarev et al. [18], PatchMatch [3] and our proposed Pseudo-Anonymized Contextual patch Extraction method, PACE. It can be observed that the proposed PACE captures higher patch density in residual sensitive areas without compromising general coverage of the whole ID image compared with the competing methods thanks to a greedy strategy that prioritizes coordinates adjacent to redacted areas. Table I presents the evaluation across 250 real pseudo-anonymized IDs at 128×128 patch resolution, benchmarking our approach against both rigid structural baselines and stochastic state-of-the-art methodologies. Evaluations were carried out in our server, with an Intel(R) Xeon(R) Gold 5420+ CPU using just one core. The limitations of a rigid uniform grid are evident: due to severe phase misalignment against irregular censorship boundaries, the naive grid captures only 63.41% of the available ROI. The introduction of probabilistic sampling (Kochkarev et al. [8]) and stochastic perturbation (PatchMatch [3]) mitigates background dilution, yielding higher Semantic Densities (3,435 and 3,765 px/patch, respectively). However, this randomness causes severe structural fragmentation. By selectively skipping certain regions, these methods introduce significant discontinuities along the perimeter, reducing the ROI coverage below 53%. By operating free of grid constraints while maintaining absolute deterministic packing, our proposed PACE dynamically contours the masks without leaving perimeter gaps. This strategic packing boosts ROI coverage to a leading 73.64% and maximizes Semantic Density (3,805 pristine pixels/patch) at an optimal evaluation stride (s=2s=2). Fig. 7 depicts this phenomena visually, where it can be observed that our proposed patch greedy strategy extract patches that contour the redacted areas, better capturing the residual sensitive data. Conversely, the competing methods either capture contaminated areas (orange cells), or fail to properly occupy semantically dense regions where residual sensitive data is uncovered. We would also like to note that our proposed method performance degrade very slightly as the stride parameter s goes up, showcasing its flexibility for practical applications. V FakeIDet3-DB: Experimental Results FakeIDet3-DB pioneers the inclusion of digital attacks on real, government-issued IDs. In this section we benchmark these attacks across both fake ID detection and localization tasks, utilizing standard metrics in the literature such as EER (%) [37, 38] and pixel-level AUC (%) [15, 47, 46], respectively. For this analysis, we select five state-of-the-art image forensic methods as baselines: TruFor [15] and Re-MTKD [46] (denoted as Re-MTKDAAAI25) target general image forensics and perform both tasks. Additionally, we include its fine-tuned version—the DeepID challenge winner [19]—denoted as Re-MTKDICCV25, to assess how pre-trained domain data affects performance. Due to architectural specifics, SparseViT [35] is evaluated exclusively on localization, while FakeIDet2 [24] serves strictly as a detection baseline given its original focus on physical ID attacks. Our primary objective is to evaluate how challenging the detection of FakeIDet3-DB’s attacks is by state-of-the-art detectors. Section V-A evaluates the state-of-the-art detectors presented using the entire database (train and test sets) without anonymization, feeding the whole ID image. This allows a direct comparison to other public databases in the literature based on digital attacks. Section V-B evaluates the detectors on our proposed patch-based benchmark, adhering to the privacy-aware protocol from [24], using patches extracted from pseudo-anonymized IDs via our proposed PACE patch extractor. All evaluations have been carried out using a single NVIDIA RTX 4090 GPU with weights strictly frozen. V-A Attack Difficulty Assessment We evaluate the whole FakeIDet3-DB (i.e., no train/test split) to effectively demonstrate the quality and diversity of the proposed attacks. Given the wide range of attacks and post-processing methods, we divide the analysis into Classical and GenAI Attacks. Table I shows the performance on Classical Attacks, where we can see that the detectors struggle to detect and locate the proposed classical attacks. Regarding detection, we observe that both TruFor and Re-MTKDICCV25 perform quite on par overall (31.14% EER vs. 32.09% EER). Nevertheless, Re-MTKDICCV25 is better at detecting manipulations involving existing content (i.e., Copy-Move and Splicing) while TruFor excels detecting face morphing manipulations, which indicates that is more robust to face manipulations. While being optimized to detect physical ID attacks, FakeIDet2 performs quite poorly overall when detecting digital manipulations, which suggests that domain adaptation may not be a crucial factor for digital ID attacks detection. This phenomena is exacerbated for Re-MTKDAAAI25, where the performance is near random with 48.85% EER. TABLE I: Performance on classical attacks. Best result is highlighted in bold. Second best result is underlined. Copy-Move Splicing Face Morph. Avg. Model Cheapfake Refined Cheapfake Refined Cheapfake Refined Detection (EER% ↓ ) TruFor[CVPR23] [15] 32.80 43.21 19.60 33.65 22.43 35.20 31.14 Re-MTKD[AAAI25] [46] 50.04 52.21 43.60 48.41 50.49 48.39 48.85 Re-MTKD[ICCV25] [19] 17.20 40.00 10.11 31.19 42.20 51.89 32.09 FakeIDet2[InfFus26] [24] 45.19 47.39 45.18 48.91 48.81 48.81 47.38 Localization (AUC% ↑ ) TruFor[CVPR23] [15] 96.35 88.75 97.25 89.19 84.00 75.58 88.52 Re-MTKD[AAAI25] [46] 71.10 68.39 81.81 68.39 92.45 89.75 78.64 Re-MTKD[ICCV25] [19] 93.32 81.63 95.39 81.12 76.35 61.93 81.62 SparseViT[AAAI25] [35] 83.89 74.74 92.87 78.36 59.87 50.20 73.25 Analyzing the localization task, TruFor outperforms the other detectors by an evident margin across most of the manipulation types with an average of 88.52% AUC, which we believe is motivated by i) the introduction of forensic priors, such as the Noiseprint++ module, which extracts residual forensic information from the original image, and i) the model was pre-trained with data containing copy-move and splicing attacks. Conversely, SparseViT yielded the lowest performance on classical attacks (73.25% AUC). We attribute this to two main factors: the artifact-inducing downsampling of high-resolution IDs to 512×512, and the architecture’s sparse-attention module. By restricting dense interactions between adjacent patch embeddings, sparse attention may dilute the weights in localized tampered regions, severely hindering the accurate localization of small manipulations. TABLE IV: Performance on GenAI Attacks. Best result is highlighted in bold. Second best result is underlined. Text Inp. Face Swap. Face Morph. Cont. Remov. Multi-Attack Avg. Model Cheapfake Refined Cheapfake Refined Cheapfake Refined Refined Refined Detection (EER% ↓ ) TruFor[CVPR23] [15] 40.40 46.00 17.65 15.07 7.02 17.96 50.00 16.07 26.27 Re-MTKD[AAAI25] [46] 37.64 40.54 43.87 41.07 26.49 29.61 52.00 38.38 38.69 Re-MTKD[ICCV25] [19] 3.55 6.67 5.09 4.26 4.02 15.46 50.00 8.44 12.18 FakeIDet2[InfFus26] [24] 53.37 54.80 43.66 51.78 51.81 47.00 48.78 54.81 50.75 Localization (AUC% ↑ ) TruFor[CVPR23] [15] 93.57 86.02 81.22 77.67 84.48 75.04 50.20 79.31 78.44 Re-MTKD[AAAI25] [46] 76.62 69.47 89.70 92.17 98.33 96.37 57.97 85.84 83.30 Re-MTKD[ICCV25] [19] 89.88 71.56 98.29 90.03 91.61 86.35 51.32 68.68 80.96 SparseViT[AAAI25] [35] 84.07 69.70 60.55 72.93 99.91 83.21 51.77 77.15 74.91 Regarding the detection of GenAI attacks, the results are presented in Table IV. The first noticeable aspect is that Re-MTKDICCV25 detects GenAI attacks not only far better than Classical Attacks on average (12.18% EER vs. 32.09% EER), but also far better than the competing methods. For digital attacks with face and text manipulations, we observe that it achieves results under 10% EER consistently, with the exception of refined face morphing (15.46% EER) and content removal attacks (50% EER). Regarding its pre-trained counterpart, Re-MTKDAAAI25, performance drops quite evidently across all types of attacks, regardless the refinement post-processing procedure. TruFor performs better detecting GenAI and Classical Attacks (26.27% EER vs. 31.14% EER), exhibiting mild generalization capabilities as it was not exposed to GenAI manipulations when pre-trained. However, for content removal attacks, we observe that no detector outperforms random guessing. Regarding localization performance, the detectors perform comparably on classical attacks, with TruFor leading the group. Intriguingly, Re-MTKDAAAI25 outperforms its fine-tuned version to IDs by a small margin (83.30% AUC vs. 80.96% AUC), revealing a clear decorrelation between detection and localization performance. This discrepancy can be explained by examining the Re-MTKD architecture [46], which decouples these tasks into two distinct heads: a classification head that leverages bottleneck features from the Cue-Net, and a localization head fed by the decoder. Because these heads are optimized using different objective functions, their respective weight updates are applied independently. Consequently, this separate optimization can cause the decoder to overfit to highly specific, dense localization cues present only in FantasyID, which fail to generalize to FakeIDet3-DB due to the more complex nature of the spatial supervisory signal. In contrast, the classification head extracts more generic, high-level features that remain applicable across domains, explaining why the model remains proficient at detection but fails at localization. TruFor’s performance in GenAI attacks excels in Text Inpainting attacks (93.57% AUC and 86.02% AUC) with respect to the rest of the models, while struggling with face manipulations from generative models (from 75.04% AUC to 84.48% AUC). Furthermore, TruFor localization performance is worse in GenAI than Classic Attacks (78.44% AUC vs. 88.52% AUC), which is expected as its pre-training data encompassed only classic manipulations. Finally SparseViT remains the worst performing model in GenAI attacks (74.91%), although it has the best performance in cheapfake face morphing attacks (99.91% AUC), which can be considered the easiest attack due to the evident manipulation traces left. Additionally, we would like to note that our refinement post-processing produces significantly more challenging threats than naive replacement, degrading detection and localization performance compared to cheapfakes in 83.33% and 91.67% of cases, respectively. TABLE V: Comparison of the proposed FakeIDet3-DB with FantasyID [20] in terms of difficulty. Best result is highlighted in bold. TruFor[15] CVPR23 Re-MTKD[46] AAAI25 Task Database Det. (EER% ↓ ) FantasyID[20] 16.02 33.67 FakeIDet3-DB 32.45 40.28 Loc. (AUC% ↑ ) FantasyID[20] 82.38 75.01 FakeIDet3-DB 83.48 80.82 Furthermore, for completeness, we compare the variability and quality of the digital attacks generated in our proposed FakeIDet3-DB with FantasyID [20], a recent public database in the field. This comparison is depicted in Table V for both detection and localization tasks. For the localization task we leave out the real IDs, therefore comparing strictly the fidelity on the digital attacks. For the detection task, we strictly consider our government-issued IDs as the only real class, since FantasyID’s “real” IDs are synthetically created in lab conditions. We would like to remark that we use the models’ pre-trained weights frozen, therefore, they are not biased towards detecting any of the manipulations from any of the databases. Consequently, we leave out the Re-MTKDICCV25 as it was fine-tuned using FantasyID for the DeepID challenge. As can be seen in Table V, it becomes evident that FakeIDet3-DB poses a significantly harder challenge for image-level detection. Specifically, the EER of TruFor drastically increases from 16.02% in FantasyID to 32.45% in FakeIDet3-DB, doubling the error rate. A similar trend occurs for Re-MTKD, whose EER rises from 33.67% to 40.28%. Conversely, for the localization task, both detectors exhibit a slightly higher pixel-level AUC on FakeIDet3-DB (83.48% and 80.82% for TruFor and Re-MTKD, respectively) compared to FantasyID. This apparent paradox between poor global detection and competent local segmentation reveals a critical insight: while the manipulations in FakeIDet3-DB leave sufficient local traces to be relatively distinguished from the authentic background at a pixel level—which may be induced by the intricate patterns and engravings embedded only in real IDs, being difficult to replicate by GenAI models—the overall image context perfectly camouflages these anomalies from a global classification perspective. Ultimately, these results demonstrate that FakeIDet3-DB acts as a demanding benchmark that effectively exposes the vulnerabilities of current forensics detectors under realistic, in-the-wild scenarios. TABLE VI: Dataset Split by Attack Modality, Models, and Refinement Level Modality Models Train (Cheapfake) Test (Cheapfake & Refined) GenAI Text Attacks DiffSTE [17], RS-STE [11], UDiffText [49] ✓ ✓ TextDiffuser 2 [9] ✗ ✓ GenAI Face Swapping REFace [2], FaceDancer [31] ✓ ✓ InsightFace555https://github.com/deepinsight/insightface ✗ ✓ GenAI Face Morphing FreeMorph [7] ✓ ✓ DiffMorpher [48] ✗ ✓ Classical All Classical Attacks ✓ ✓ V-B Public Benchmark TABLE VII: Proposed patch-based benchmark: Performance on both localization and detection tasks in the pseudo-anonymized scenario. Best result is highlighted in bold. Second best is underlined. Classical Attacks GenAI Attacks Model Copy-Move Splicing Face Morph. Text Inp. Face Swap. Face Morph. Cont. Remov. Multi-Attack All Detection (EER% ↓ ) Patch Size: 128×128 TruFor[CVPR23] [15] 47.83 45.65 54.35 40.87 54.35 34.78 47.83 43.48 47.83 Re-MTKD[AAAI25] [46] 47.83 34.78 41.31 35.32 36.52 42.39 47.83 26.09 39.13 Re-MTKD[ICCV25] [19] 34.78 34.78 39.95 31.49 29.56 23.91 34.78 26.09 34.78 FakeIDet2[InfFus26] [24] 45.65 47.83 41.30 44.57 42.61 46.74 47.83 47.83 47.48 Patch Size: 64×64 TruFor[CVPR23] [15] 43.48 47.83 47.83 51.62 51.30 50.00 47.83 47.83 52.17 Re-MTKD[AAAI25] [46] 52.17 45.65 54.34 37.49 46.09 58.69 47.83 30.43 47.48 Re-MTKD[ICCV25] [19] 54.34 52.18 58.70 50.54 54.78 57.61 52.17 39.13 52.17 FakeIDet2[InfFus26] [24] 56.62 56.62 56.52 30.43 43.47 47.82 43.48 30.43 47.48 Localization (AUC% ↑ ) Patch Size: 128×128 TruFor[CVPR23] [15] 73.27 80.37 64.28 62.82 59.69 63.69 49.37 48.17 62.71 Re-MTKD[AAAI25] [46] 71.79 68.85 68.10 62.02 69.19 62.13 44.27 50.43 62.09 Re-MTKD[ICCV25] [19] 69.13 72.12 48.23 68.62 51.43 48.12 52.17 34.78 62.09 SparseViT[AAAI25] [35] 69.64 41.43 68.85 47.95 43.99 45.39 45.59 45.27 51.03 Patch Size: 64×64 TruFor[CVPR23] [15] 54.63 60.49 50.63 43.18 51.81 56.67 53.52 38.67 51.18 Re-MTKD[AAAI25] [46] 39.87 53.31 38.29 35.79 46.05 46.25 43.09 32.20 41.85 Re-MTKD[ICCV25] [19] 61.25 63.30 25.12 58.95 43.70 44.09 35.94 56.42 48.69 SparseViT[AAAI25] [35] 51.14 54.01 40.08 42.79 38.61 36.77 45.74 42.20 43.91 Finally, in this section we describe the details of our proposed benchmark, which is publicly available to the research community in order to advance in this challenging field. This benchmark shifts from full-image to patch-based ID processing (128×128 and 64×64 patch sizes) to safeguard sensitive data, as described throughout the article and our proposed PACE patch extractor (Sec. IV). Following [24], detection treats a single ID as an unordered collection of independent patches, outputting a global authenticity score. For localization, we evaluate patch-wise binary masks under strict privacy constraints, requiring models to generate spatial predictions exclusively for non-redacted inputs. As most baseline detectors were trained on full images, we applied here simple adaptations: detectors process individual patches to output local masks, and patch-level detection scores are averaged via mean fusion to compute the final ID-level score [23]. To simulate real-world conditions where detectors face unseen threats, we carefully partition our dataset to restrict the availability of both attack types and ID templates during training. Following the FakeIDet2-db protocol [24], the first and second versions of the Spanish ID are reserved for testing, leaving the third exclusively for training. Furthermore, as detailed in Table VI, the training set is strictly limited to classical attacks and cheapfake manipulations from a subset of GenAI models. Highly capable models and refined post-processing techniques are held out for the test set. This allows us to rigorously evaluate whether models merely overfit known forensic traces or appropriately generalize to unseen GenAI models and novel refinement procedures. Full details are available in our GitHub. Table VII demonstrates that the restricted patch-based evaluation severely degrades detector capabilities, exposing their heavy reliance on global image context and macro-structural artifacts. In detection, error rates suffer a near-total collapse. At 128×128 patch size, Re-MTKDICCV25 achieves the best performance (34.78% EER), yet remains unacceptably high for real scenarios. At 64×64 patch size, detectors essentially revert to random guessing (≥ 50% EER). Notably, FakeIDet2—despite its patch-based design—fails against digital manipulations (47.48% EER), proving digital forensic traces are vastly subtler than the physical forgeries it was optimized for. These results expose that naive mean-score fusion is insufficient to aggregate localized predictions as shown in [24] and more sophisticated strategies are needed for patch-based digital attacks detection. The task of localization suffers a similar fate. At 128×128 patch size, TruFor and Re-MTKD marginally segment Classical attacks (e.g., TruFor achieves 80.37% AUC on Splicing), likely capturing abrupt high-frequency noise discrepancies within the patch. However, GenAI manipulations bypass detection entirely. At 64×64 patch size, localization drops near 50% AUC across almost all detectors, neutralizing their spatial predictive capabilities. This trend reflects an inherent issue in the baseline architectures when applied to patch-based localization, as their receptive fields—accustomed to high-resolution global semantics (≥ 512×512)—are limited by the restrictive semantic content embedded in individual patches. VI Conclusion In this article we presented FakeIDet3-DB, the first database featuring digital attacks on 250 images of real, government-issued IDs. Ensuring a comprehensive set of classical and GenAI attacks, we presented a novel post-processing pipeline generating diverse digital attacks, bridging the gap between easy-to-spot cheapfakes and highly realistic refined forgeries. Evaluating state-of-the-art methods on full, non-anonymized IDs, FakeIDet3-DB poses a harder detection challenge compared to existing databases like FantasyID [20] (32.45% vs. 16.02% EER), while remaining competitive in localization (83.48% vs. 82.38% AUC). Our refined post-processing pipeline creates highly deceptive attacks, as in 83.33% and 91.67% of evaluated cases, refined attacks yielded worse detection and localization results than their cheapfake counterparts. Because unrestricted image distribution violates data privacy regulations, we adopted a privacy-aware framework [24], applying strategic anonymization via content redaction to prevent PII leakage (i.e., pseudo-anonymization). Subsequently, we proposed PACE, a geometrically constrained patch extraction algorithm that prioritizes spatial coordinates adjacent to redacted regions. We demonstrated this greedy strategy achieves an optimal balance between residual sensitive information coverage and semantic density with low computational overhead, making it suitable for real-world scenarios. Finally, we established a privacy-compliant benchmark leveraging 5.2M patches extracted from over 6.4K pseudo-anonymized real/fake ID images. Evaluating state-of-the-art detectors (Re-MTKD [46], TruFor [15], SparseViT [35], FakeIDet2 [24]), we observed considerable performance degradation compared to full-image evaluations. We primarily attribute this drop to the baselines’ receptive fields, optimized for higher resolutions. Moreover, the naive fusion strategies employed to aggregate patch scores further hinder performance. Future work will explore the development of architectures specifically designed to detect and localize digital attacks while strictly complying with privacy regulations. We hope the research community embraces patch-based frameworks not as an impediment, but as a secure, alternative paradigm enabling vital cooperation between AI researchers and data holders. Acknowledgment This project has been supported by PowerAI+ (SI4/PJI/2024- 00062 Comunidad de Madrid and UAM), Cátedra ENIA UAM-Veridas en IA Responsable (NextGenerationEU PRTR TSI-100927-2023-2), and TRUST-ID (PID2025-173396OB-I00 MICIU/AEI and the EU). References [1] V. V. Arlazarov et al. (2019) MIDV-500: A Dataset for Identity Document Analysis and Recognition on Mobile Devices in Video Stream. Computer Optics 43. Cited by: §I-A. [2] S. Baliah et al. (2025) Realistic and Efficient Face Swapping: A Unified Approach with Diffusion Models. In Proc. IEEE/CVF Winter Conf. on Applications of Computer Vision (WACV), Vol. , p. 1062–1071. Cited by: §I-A2, TABLE VI. [3] C. Barnes et al. (2009) PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing. ACM Trans. Graph. 28 (3), p. 24. Cited by: §I-B, Figure 7, §IV-B, §IV-B, TABLE I. [4] D. Benalcazar et al. (2023) Synthetic ID Card Image Generation for Improving Presentation Attack Detection. IEEE Trans. Inf. Forensics Security 18, p. 1814–1824. Cited by: TABLE I. [5] K. Bulatov et al. (2022) MIDV-2020: A Comprehensive Benchmark Dataset for Identity Document Analysis. Computer Optics 46. Cited by: §I-A. [6] K. Bulatov et al. (2020) MIDV-2019: Challenges of the Modern Mobile-Based Document OCR. In Proc. Intl. Conf. on Machine Vision (ICMV), p. 818–824. Cited by: §I-A. [7] Y. Cao et al. (2025) FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model. In Proc. IEEE/CVF Intl. Conf. on Computer Vision (ICCV), Vol. , p. 18111–18120. Cited by: §I-A2, TABLE VI. [8] L. Chai et al. (2020) What Makes Fake Images Detectable? Understanding Properties that Generalize. In Proc. European Conf. on Computer Vision (ECCV), Cited by: §IV-B, §IV. [9] J. Chen et al. (2025) TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering. In Proc. European Conf. on Computer Vision (ECCV), p. 386–402. External Links: ISBN 978-3-031-72652-1 Cited by: §I-A2, TABLE VI. [10] F. Croitoru et al. (2023) Diffusion Models in Vision: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 45 (9), p. 10850–10869. Cited by: §I-A2. [11] Z. Fang et al. (2025) Recognition-Synergistic Scene Text Editing. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 13104–13113. Cited by: §I-A2, TABLE VI. [12] H. Farid (2009) Image forgery detection. IEEE Signal Process. Mag. 26 (2), p. 16–25. Cited by: §I-A1. [13] Y. Gao et al. (2026) Toward Generalizable Forgery Detection and Reasoning. IEEE Trans. Image Process. 35 (), p. 3395–3410. Cited by: §I. [14] J. Gui et al. (2023) A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications. IEEE Trans. Knowl. Data Eng. 35 (4), p. 3313–3332. Cited by: §I-A2. [15] F. Guillaro et al. (2023-06) TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), p. 20606–20615. Cited by: TABLE I, TABLE I, TABLE IV, TABLE IV, TABLE V, TABLE VII, TABLE VII, TABLE VII, TABLE VII, §V, §VI. [16] Y. Hua et al. (2023) Learning Patch-Channel Correspondence for Interpretable Face Forgery Detection. IEEE Trans. Image Process. 32 (), p. 1668–1680. Cited by: §IV. [17] J. Ji et al. (2023) Improving Diffusion Models for Scene Text Editing with Dual Encoders. Note: arXiv:2304.05568 Cited by: §I-A2, TABLE VI. [18] A. Kochkarev et al. (2020) Data Balancing Method for Training Segmentation Neural Networks. In Proc. CEUR Workshop, Cited by: §I, §I-B, Figure 7, §IV-A1, §IV-B, TABLE I. [19] P. Korshunov et al. (2025) DeepID Challenge of Detecting Synthetic Manipulations in ID Documents. In Proc. IEEE/CVF Intl. Conf. on Computer Vision (ICCV) Workshops, p. 521–530. Cited by: §I, §I-A, TABLE I, TABLE I, TABLE IV, TABLE IV, TABLE VII, TABLE VII, TABLE VII, TABLE VII, §V. [20] P. Korshunov et al. (2025) FantasyID: A Dataset for Detecting Digital Manipulations in ID-Documents. In Proc. IEEE/IAPR Intl. Joint Conf. on Biometrics (IJCB), Vol. , p. 1–9. Cited by: §I, §I-A, TABLE I, §V-A, TABLE V, TABLE V, TABLE V, §VI. [21] R. Liu et al. (2023) Character-Aware Models Improve Visual Text Rendering. In Proc. 61st Annual Meeting of the Association for Computational Linguistics (ACL), p. 16270–16297. Cited by: §I. [22] R. Mudgalgundurao et al. (2025) Detecting Presentation Attacks on ID Cards Using Feature Refinement. In 2025 33rd European Signal Processing Conf. (EUSIPCO), p. 825–829. Cited by: §I-A, TABLE I. [23] J. Muñoz-Haro et al. (2025) FakeIDet: Exploring Patches for Privacy-Preserving Fake ID Detection. In Proc. IEEE/IAPR Intl. Joint Conf. on Biometrics (IJCB), p. 1–9. Cited by: §I, §I, §V-B. [24] J. Muñoz-Haro et al. (2026) Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2). Information Fusion 128, p. 103969. External Links: ISSN 1566-2535 Cited by: 2nd item, §I, §I-A, §I-B, TABLE I, §I-A, §I, Figure 7, TABLE I, §IV, §IV, §V-B, §V-B, §V-B, TABLE I, TABLE IV, TABLE VII, TABLE VII, §V, §V, §VI, §VI. [25] T. C. Ndir et al. (2026) Scaling In-Context Segmentation with Hierarchical Supervision. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRw), p. 6150–6156. Cited by: §I-B. [26] J. C. Neves et al. (2020) GANprintR: Improved Fakes and Evaluation of the State of the Art in Face Manipulation Detection. IEEE J. Sel. Topics Signal Process., p. 1038–1048. Cited by: §I-A2. [27] E. Park et al. (2023) KID34K: A Dataset for Online Identity Card Fraud Detection. In Proc. ACM Intl. Conf. on Information and Knowledge Management, p. 5381–5385. Cited by: §I-A, TABLE I. [28] L. Pedrouzo-Rodriguez et al. (2026) Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and Benchmark. Pattern Recognition. Cited by: §I. [29] D. V. Polevoy et al. (2022) Document Liveness Challenge Dataset (DLC-2021). Journal of Imaging 8 (7). Cited by: §I-A, TABLE I. [30] N. Ravi et al. (2025) SAM 2: segment anything in images and videos. In Proc. Intl. Conf. on Learning Representations (ICLR), Cited by: §I-B2. [31] F. Rosberg et al. (2023-01) FaceDancer: Pose- and Occlusion-Aware High Fidelity Face Swapping. In Proc. IEEE/CVF Winter Conf. on Applications of Computer Vision (WACV), p. 3454–3463. Cited by: §I-A2, TABLE VI. [32] A. Rosenfeld and J.L. Pfaltz (1968) Distance Functions on Digital Pictures. Pattern Recognition 1 (1), p. 33–61. External Links: ISSN 0031-3203 Cited by: §IV-A2. [33] A. Rosenfeld and J. L. Pfaltz (1966) Sequential Operations in Digital Picture Processing. J. ACM 13 (4), p. 471–494. Cited by: §IV-A2. [34] S. Sinitsa and O. Fried (2024) Deep Image Fingerprint: Towards Low Budget Synthetic Image Detection and Model Lineage Analysis. In Proc. IEEE/CVF Winter Conf. on Applications of Computer Vision (WACV), p. 4067–4076. Cited by: §I-A2. [35] L. Su et al. (2025) Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding Transformer. In Proc. AAAI Conf. on Artificial Intelligence, Vol. 39, p. 7024–7032. Cited by: TABLE I, TABLE IV, TABLE VII, TABLE VII, §V, §VI. [36] R. Suvorov et al. (2022) Resolution-robust Large Mask Inpainting with Fourier Convolutions. In Proc. IEEE/CVF Winter Conf. on Applications of Computer Vision (WACV), p. 3172–3182. Cited by: §I-A2, §I-A2, §I-B1. [37] J. E. Tapia et al. (2024) First Competition on Presentation Attack Detection on ID Card. In Proc. IEEE/IAPR Intl. Joint Conf. on Biometrics (IJCB), p. 1–10. Cited by: §V. [38] J. E. Tapia et al. (2025) Second Competition on Presentation Attack Detection on ID Card. In Proc. IEEE/IAPR Intl. Joint Conf. on Biometrics (IJCB), Vol. , p. 1–10. Cited by: §I, §I-A, §V. [39] R. Tolosana et al. (2020) Deepfakes and Beyond: A Survey of Face Manipulation and Fake Detection. Information Fusion 64, p. 131–148. Cited by: §I. [40] S. Venkatesh et al. (2021) Face Morphing Attack Generation and Detection: A Comprehensive Survey. IEEE Technol. Soc. Mag. 2 (3), p. 128–145. Cited by: §I, §I-A1. [41] C. Wang et al. (2026) Focus on Finding Deepfakes: A Robust Proactive Detection Method Based on Orthogonal Moment Watermarking. IEEE Trans. Image Process. 35. Cited by: §I. [42] Q. Wang et al. (2026) Unsupervised Domain Adaptation-Based Cross-Type Deepfake Image Detection. IEEE Trans. Image Process. 35 (), p. 4411–4424. Cited by: §I. [43] S. Wang et al. (2020) CNN-Generated Images Are Surprisingly Easy to Spot… for Now. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), p. 8695–8704. Cited by: §I-A2. [44] L. Xie et al. (2024) IDNet: A Novel Identity Document Dataset via Few-Shot and Quality-Driven Synthetic Data Generation. In Proc. IEEE Intl. Conf. on Big Data, Vol. , p. 2244–2253. Cited by: §I-A, TABLE I. [45] M. Ye et al. (2025) Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 47 (03), p. 1431–1447. Cited by: §I-B1. [46] Z. Yu et al. (2025) Reinforced Multi-teacher Knowledge Distillation for Efficient General Image Forgery Detection and Localization. In Proc. AAAI Conf. on Artificial Intelligence, Vol. 39, p. 995–1003. Cited by: §V-A, TABLE I, TABLE I, TABLE IV, TABLE IV, TABLE V, TABLE VII, TABLE VII, TABLE VII, TABLE VII, §V, §VI. [47] Z. Yu et al. (2024) DiffForensics: Leveraging Diffusion Prior to Image Forgery Detection and Localization. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), p. 12765–12774. Cited by: §V. [48] K. Zhang et al. (2024) DiffMorpher: Unleashing the Capability of Diffusion Models for Image Morphing. In Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 7912–7921. Cited by: §I-A2, TABLE VI. [49] Y. Zhao et al. (2024) UDiffText: A Unified Framework for High-Quality Text Synthesis in Arbitrary Images via Character-Aware Diffusion Models. In Proc. European Conf. on Computer Vision (ECCV), p. 217–233. External Links: ISBN 978-3-031-72750-4 Cited by: §I-A2, TABLE VI.