Paper deep dive
High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking
Peipeng Yu, Jinfeng Xie, Chengfu Ou, Xiaoyu Zhou, Jianwei Fei, Yunshu Dai, Zhihua Xia, Chip Hong Chang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/26/2026, 2:13:57 AM
Summary
VeriFi is a versatile, tamper-resilient watermarking framework designed for proactive deepfake defense. It unifies copyright protection, pixel-level manipulation localization, and high-fidelity face content recovery by embedding a compact semantic latent watermark. Unlike prior methods that suffer from a fidelity-functionality trade-off due to explicit localization payloads, VeriFi uses a watermark-guided localization network and a dual-stream Transformer-based recovery network. It also introduces an AIGC attack simulator that models latent-space mixing and Poisson blending to enhance robustness against realistic deepfake pipelines.
Entities (5)
Relation Signals (3)
VeriFi → performs → Manipulation Localization
confidence 95% · it achieves fine-grained localization without embedding localization-specific artifacts
VeriFi → provides → Copyright Protection
confidence 95% · VeriFi, a versatile watermarking framework that unifies copyright protection, pixel-level manipulation localization, and high-fidelity face content recovery.
VeriFi → utilizes → AIGC Attack Simulator
confidence 95% · it introduces an AIGC attack simulator that combines latent-space mixing with seamless blending
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Prior versatile watermarking systems typically rely on embedding explicit localization payloads, which introduces a fidelity--functionality trade-off: larger localization signals degrade visual quality and often reduce decoding robustness under strong generative edits. Moreover, existing methods rarely support content recovery, limiting their forensic value when original evidence must be reconstructed. To address these challenges, we present VeriFi, a versatile watermarking framework that unifies copyright protection, pixel-level manipulation localization, and high-fidelity face content recovery. VeriFi makes three key contributions: (1) it embeds a compact semantic latent watermark that serves as an content-preserving prior, enabling faithful restoration even after severe manipulations; (2) it achieves fine-grained localization without embedding localization-specific artifacts by correlating image features with decoded provenance signals; and (3) it introduces an AIGC attack simulator that combines latent-space mixing with seamless blending to improve robustness to realistic deepfake pipelines. Extensive experiments on CelebA-HQ and FFHQ show that VeriFi consistently outperforms strong baselines in watermark robustness, localization accuracy, and recovery quality, providing a practical and verifiable defense for deepfake forensics.
Tags
Links
- Source: https://arxiv.org/abs/2603.23940v1
- Canonical: https://arxiv.org/abs/2603.23940v1
Trouble viewing inline? Open PDF directly →
Full Text
51,096 characters extracted from source content.
Expand or collapse full text
High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking Peipeng Yu ypp865@163.com Jinan UniversityGuangzhouGuangdongChina , Jinfeng Xie Jinan UniversityGuangzhouGuangdongChina , Chengfu Ou Jinan UniversityGuangzhouGuangdongChina , Xiaoyu Zhou Jinan UniversityGuangzhouGuangdongChina , Jianwei Fei University of MacauMacauChina , Yunshu Dai Sun Yat-sen UniversityGuangzhouGuangdongChina , Zhihua Xia Jinan UniversityGuangzhouGuangdongChina and Chip Hong Chang Nanyang Technological UniversitySingaporeSingapore (5 June 2009) Abstract. The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Prior versatile watermarking systems typically rely on embedding explicit localization payloads, which introduces a fidelity–functionality trade-off: larger localization signals degrade visual quality and often reduce decoding robustness under strong generative edits. Moreover, existing methods rarely support content recovery, limiting their forensic value when original evidence must be reconstructed. To address these challenges, we present VeriFi, a versatile watermarking framework that unifies copyright protection, pixel-level manipulation localization, and high-fidelity face content recovery. VeriFi makes three key contributions: (1) it embeds a compact semantic latent watermark that serves as an content-preserving prior, enabling faithful restoration even after severe manipulations; (2) it achieves fine-grained localization without embedding localization-specific artifacts by correlating image features with decoded provenance signals; and (3) it introduces an AIGC attack simulator that combines latent-space mixing with seamless blending to improve robustness to realistic deepfake pipelines. Extensive experiments on CelebA-HQ and FFHQ show that VeriFi consistently outperforms strong baselines in watermark robustness, localization accuracy, and recovery quality, providing a practical and verifiable defense for deepfake forensics. AIGC, versatile watermarking, content recovery, manipulation localization, media forensics †copyright: acmlicensed†journalyear: 2018†doi: X.X†conference: Make sure to enter the correct conference title from your rights confirmation email; June 03–05, 2018; Woodstock, NY†isbn: 978-1-4503-X-X/2018/06†submissionid: 1777†ccs: Security and privacy Trust frameworks†ccs: Security and privacy Authentication 1. Introduction Recent advances in artificial intelligence-generated content (AIGC) have enabled highly realistic face synthesis and accessible deepfakes. While these techniques support creative applications, they also introduce significant risks, including misinformation, copyright infringement, and loss of content authenticity (Wang et al., 2024b). As generative models evolve, passive detection methods are increasingly ineffective against sophisticated and unseen deepfakes (Tan et al., 2024a, b; Liu et al., 2025). This highlights the need for proactive, verifiable defenses that remain robust under AIGC manipulations and provide reliable forensic evidence. Figure 1. Overview of our VeriFi framework. Our versatile watermarking architecture simultaneously achieves robust copyright protection, precise forgery localization, and high-fidelity fake face recovery, providing comprehensive forensic capabilities essential for real-world deepfake defense scenarios. Digital watermarking provides a proactive mechanism for establishing media provenance and safeguarding integrity (Yang et al., 2024; Zhao et al., 2024). Robust watermarks are designed to remain decodable under benign transformations and malicious edits (e.g., deepfake generation and splicing) (Wang et al., 2021; Wu et al., 2023; Wang et al., 2024a), whereas fragile watermarks intentionally react to modifications, enabling sensitive tamper detection (Neekhara et al., 2024; Araghi et al., 2024; Wang et al., 2025). Building upon these paradigms, recent versatile frameworks (e.g., EditGuard (Zhang et al., 2024b), OmniGuard (Zhang et al., 2025), StableGuard (Yang et al., 2025), and WAM (Sander et al., 2025)) seek to unify copyright protection with manipulation localization. However, EditGuard and OmniGuard rely on high-capacity localization payloads, which increases embedding distortion and can weaken copyright robustness. In contrast, StableGuard and WAM largely depend on passive detectors, which often exhibit limited localization accuracy under unseen AIGC edits. Critically, existing versatile watermarking methods do not support recovery of the original content, substantially limiting their forensic utility when faithful reconstruction is required for evidence preservation. How to escape the fidelity–functionality trade-off between copyright protection and tamper localization, while enabling high-fidelity content recovery, remains an open challenge. To address these limitations, we propose VeriFi, a unified proactive framework that simultaneously supports (i) robust copyright protection, (i) pixel-level tamper localization, and (i) high-fidelity facial content recovery via tamper-resilient versatile watermarking. Unlike prior versatile methods that rely on separate localization payloads and thus compromise between visual quality and functionality, VeriFi introduces a novel watermark-conditioned localization network that detects manipulations by analyzing spatial inconsistencies in the decoded copyright signal, eliminating the need for additional localization watermarks. Moreover, VeriFi incorporates a compact semantic recovery watermark that encodes facial content, providing a strong prior to accurately reconstruct the original face even after severe manipulation. As illustrated in Figure 1, given an original image IoriI_ori, our unified watermark encoder embeds both a copyright code wcopw_cop and a compact recovery representation to produce a protected image IproI_pro for public distribution. If the image undergoes AIGC-based manipulations, resulting in a forged image IeditI_edit, VeriFi enables three core functionalities: (1) extraction of the embedded copyright code w^cop w_cop for provenance verification; (2) watermark-guided forgery localization, leveraging w^cop w_cop to achieve accurate pixel-level manipulation mapping; and (3) recovery of the original facial content by decoding the recovery watermark, yielding a high-fidelity reconstruction I^ori I_ori. Our main contributions are summarized as follows: • We present VeriFi, a tamper-resilient versatile watermarking framework that unifies robust copyright protection, pixel-level manipulation localization, and high-fidelity face content recovery. We design a watermark-guided recovery network that leverages a compact semantic watermark as a prior, enabling faithful restoration even after severe AIGC manipulations. • We design a watermark-guided localization network that leverages the decoded copyright signal as a spatial prior, enabling accurate tamper mapping without embedding localization-specific payloads and thus avoiding the fidelity–functionality trade-off. • We introduce an AIGC attack simulator that jointly models latent mixing, editing, and Poisson blending to emulate realistic deepfake perturbations during training, substantially enhancing watermark robustness. • Extensive experiments on CelebA-HQ and FFHQ show that VeriFi surpasses strong baselines in watermark robustness, localization accuracy, and recovery quality, establishing a new benchmark for proactive deepfake forensics. 2. Related Works 2.1. Forgery Localization and Recovery Methods Table 1. Functional comparison of representative proactive watermarking and forensic methods. C: copyright protection; L: tamper localization; R: content recovery. Method Venue Category C L R SepMark M’23 Face-aware robust watermarking ✓ – – LampMark M’24 Face-aware robust watermarking ✓ – – FakeTagger M’21 Face-aware robust watermarking ✓ – – AdaIFL CVPR’24 Passive localization – ✓ – HiFi-Net ICCV’23 Passive localization – ✓ – Imuge+ CVPR’23 Proactive localization and recovery – ✓ ✓ DFREC Arxiv’25 Passive localization and recovery – ✓ ✓ EditGuard CVPR’24 Versatile watermarking ✓ ✓ – OmniGuard CVPR’25 Versatile watermarking ✓ ✓ – StableGuard NIPS’25 Versatile watermarking ✓ ✓ – WAM AAAI’25 Versatile watermarking ✓ ✓ – VeriFi (ours) – Versatile watermarking ✓ ✓ ✓ Image forgery localization seeks to identify tampered regions at the pixel level across diverse manipulation types (Zhou et al., 2023b; Zhang et al., 2024a; Li et al., 2024; Lou et al., 2025; Liu et al., 2025). Representative methods such as MVSS-Net (Yang et al., 2021), CAT-Net (Kwon et al., 2021), and HiFi-Net (Guo et al., 2023) leverage multi-branch architectures, frequency artifacts, and hierarchical attention to improve localization accuracy. However, these approaches are easily overfitted to specific datasets and do not support content recovery. DFREC (Yu et al., 2025b) extend passive localization to recovery, but remain vulnerable to distribution shifts and complex AIGC manipulations, limiting real-world robustness. In contrast, we propose a unified proactive framework that embeds robust watermarks prior to content distribution, which enables reliable tamper localization. Additionally, by simulating complex AIGC attacks, we further enhance watermark extraction robustness in real-world applications. 2.2. Proactive Watermarking Methods Figure 2. Overview of our VeriFi. (1) Unified Recovery-Copyright Watermark Embedder EncEnc inserts an ownership code wcopw_cop and a compact facial signature zfacez_face. (2) AIGC Attack Simulator performs Latent Mixing and Poisson blending to mimic realistic deepfake attacks during training. (3) Watermark Extractor DecDec recovers w^cop w_cop and z^face z_face, produces an content proxy I^b I_b, and uses w^cop w_cop to guide the Forgery Locator to predict the manipulation map M M. (4) Watermark-Guided Deepfake Recovery Network uses a dual-stream Transformer with spatially gated cross-attention to fuse the edited image, I^b I_b, and M M for selective restoration. Proactive watermarking methods embed imperceptible signals into images prior to distribution, enabling reliable provenance tracing and post-hoc forgery detection (Wang et al., 2021; Wu et al., 2023; Wang et al., 2024a). As summarized in Table 1, existing proactive watermarking approaches exhibit distinct technical focuses and design paradigms. Early deep watermarking methods, such as HiDDeN (Zhu et al., 2018) and StegaStamp (Tancik et al., 2020), primarily target copyright protection, without support for tamper localization or content recovery. Face-aware watermarking methods, including SepMark (Wu et al., 2023) and LampMark (Wang et al., 2024a), are specifically designed to enhance robustness against facial manipulations, but remain limited to copyright verification. Imuge+ (Ying et al., 2023) extends proactive watermarking to tamper detection and content recovery by introducing trivial perturbations, yet does not address copyright authentication. Recent versatile frameworks, such as EditGuard (Zhang et al., 2024b) and OmniGuard (Zhang et al., 2025), unify copyright protection and tamper localization, but require embedding high-capacity localization payloads, which can degrade visual fidelity and compromise robustness. StableGuard (Yang et al., 2025) and WAM (Sander et al., 2025) further utilize passive locators to avoid visual artifacts, but often exhibit limited generalization. Overall, existing proactive watermarking solutions are constrained by inherent trade-offs among visual fidelity, adversarial robustness, and forensic functionality. In this work, we introduce a unified framework that integrates watermark-guided localization and recovery to jointly achieve copyright protection, pixel-level tamper localization, and high-fidelity content recovery. This design addresses key limitations of prior methods and sets a new benchmark for robust, versatile image forensics. 3. Proposed Method 3.1. Motivation Existing deepfake watermarking methods suffer from two core limitations: (1) a trade-off between copyright protection and precise tamper localization, and (2) a lack of unified frameworks supporting both authentication and content recovery. Our approach addresses these gaps based on two key insights: Insight 1: Tampering-induced watermark degradation is a strong localization prior. Any content manipulation, such as face swapping, inevitably corrupts the embedded watermark signal within the altered regions. This creates a spatial inconsistency: tampered areas exhibit a damaged or absent watermark, while untampered areas retain the original, intact signal. By identifying where the watermark fails to decode correctly, we can precisely and robustly localize the manipulation. Insight 2: Semantic watermarking enables face recovery via compact latent embedding. A naive solution for content recovery is to embed the original image (or its pixel-level representation) directly as a watermark. However, this approach requires a high-capacity watermark, which inevitably introduces significant visual distortion and severely compromises robustness to manipulations. Instead, we draw inspiration from recent advances in Variational Autoencoder (VAE)-based image compression, which demonstrate that facial content can be faithfully represented by low-dimensional latent codes. By embedding a compact semantic latent as the recovery watermark, we enable high-fidelity face content reconstruction while preserving imperceptibility and watermark resilience. 3.2. Overall Architecture Building on these insights, we introduce VeriFi, a unified and tamper-resilient watermarking framework for proactive deepfake defense. VeriFi integrates copyright protection, pixel-level tamper localization, and high-fidelity face content recovery within a single framework. As illustrated in Figure 2, VeriFi comprises four tightly integrated modules that work synergistically to achieve robust and versatile forensics. At its core, the Unified Recovery-Copyright Watermark (URW) Embedder and Extractor (detailed in the Appendix.) jointly encode and decode two signals: a discrete n-bit ownership code wcopw_cop for reliable copyright authentication, and a compact continuous facial latent zfacez_face, derived from a pretrained variational autoencoder (VAE), which serves as a semantic prior for high-fidelity content recovery. The URW Embedder EncEnc is responsible for embedding these signals into the image in a visually imperceptible and perturbation-resilient manner, while the URW Extractor DecDec is optimized to recover both watermarks from images that may have undergone various manipulations. To further enhance robustness against sophisticated attacks, VeriFi incorporates an AIGC Attack Simulator (Sec. 3.5) during training, which emulates realistic deepfake perturbations by combining latent mixing, editing and image blending. This simulation strategy exposes the watermarking network to a diverse range of challenging manipulations, thereby improving its generalization and resilience. For tamper localization, VeriFi employs a Watermark-Guided Forgery Localization Network (Sec. 3.3) that leverages the decoded watermark as a spatial prior. By fusing semantic features from the potentially manipulated image with watermark-derived cues, this network generates a dense manipulation probability map M, enabling precise and robust localization of altered regions. In parallel, the Watermark-Guided Deepfake Recovery Network (Sec. 3.4) utilizes the recovered facial signature z^face z_face and the predicted manipulation map M to guide a dual-stream Transformer-based reconstructor R. Through cross-attention and adaptive modulation, the recovery network selectively restores tampered regions while preserving authentic content elsewhere, resulting in high-fidelity reconstruction of the original image. 3.3. Watermark-Guided Forgery Localization Tampering-induced degradation of embedded watermarks provides a strong spatial prior for manipulation detection. We propose a watermark-guided forgery localization network that explicitly aligns visual features with watermark-derived signals, enabling robust and precise identification of manipulated regions. Our network adopts a dual-branch Swin-Unet architecture. The image branch extracts multi-scale semantic features from the potentially manipulated input, while the watermark branch encodes spatially aligned features decoded from the recovered watermark. At each of S hierarchical scales, both branches produce feature maps, which are projected to a common channel dimension CsC_s using 1×11× 1 convolutions to facilitate direct comparison. To quantify the consistency between the two streams, we compute a per-pixel cosine similarity at each spatial location (x,y)(x,y), yielding a scale-specific similarity map SsS_s: (1) Ss(x,y)=⟨ℱimg(s)(x,y),ℱwm(s)(x,y)⟩‖ℱimg(s)(x,y)‖2⋅‖ℱwm(s)(x,y)‖2,S_s(x,y)= _img^(s)(x,y),\,F_wm^(s)(x,y) \|F_img^(s)(x,y)\|_2·\|F_wm^(s)(x,y)\|_2, where ℱimg(s)∈ℝCs×Hs×WsF_img^(s) ^C_s× H_s× W_s and ℱwm(s)∈ℝCs×Hs×WsF_wm^(s) ^C_s× H_s× W_s denote the image and watermark feature maps at scale s, and ⟨⋅,⋅⟩ ·,· is the channel-wise inner product. After that, the similarity maps S1,…,S\S_1,…,S_S\ are upsampled to the input resolution and aggregated by a lightweight fusion head: (2) ℱsim=g(Concat(Up(S1),…,Up(S))),F_sim=g\! (Concat\! (Up(S_1),…,Up(S_S) ) ), where Up(⋅)Up(·) denotes spatial upsampling and g(⋅)g(·) is a stack of convolutional layers with normalization, that produces a dense evidence map. Finally, we fuse the model-driven decoder evidence and the watermark-aligned evidence using a learnable weight α to obtain the final manipulation probability map M. 3.4. Watermark-Guided Deepfake Recovery The URW Embedder encodes a compact facial latent (zfacez_face) that remains decodable after strong manipulations, serving as a semantic prior for recovery. A coarse face proxy is reconstructed from z^face z_face via a pretrained VAE decoder. However, direct reconstruction from this proxy is typically over-smoothed and lacks detail. To overcome this, we introduce a watermark-guided recovery network that adaptively fuses spatial features from the manipulated image with semantic priors from the watermark, enabling high-fidelity restoration of the original face. As illustrated in Figure 2, the recovery network receives four inputs: (1) the manipulated image IeditmI^m_edit (masked according to the predicted forgery map), (2) a coarse content proxy I^b=DVAE(z^face) I_b=D_VAE( z_face) reconstructed from the decoded semantic watermark, (3) the decoded facial watermark z^face z_face, and (4) the predicted manipulation mask M M. The proxy I^b I_b provides structural and content-consistent cues, complementing the corrupted observation IeditmI^m_edit. We concatenate I^b I_b with IeditmI^m_edit and apply a patch embedding to obtain image tokens, which are then processed by a stack of N Transformer blocks for feature fusion and recovery. In parallel, the decoded watermark z^face z_face is encoded into a set of semantic tokens via a dedicated watermark encoder. To ensure that semantic priors are injected only into manipulated regions, the predicted mask M M is downsampled to the token resolution and used as a spatial gating mechanism during cross-attention. Formally, the gated cross-attention at each Transformer block is defined as: (3) R=(M^)⊙∑h(ζ(Tx),ζ(Twm),ζ(Twm)),R=G( M) _hA (ζ(T_x),\,ζ(T_wm),\,ζ(T_wm) ), where (⋅)G(·) is a spatial gating operator, (⋅)A(·) denotes cross-attention, and ζ(⋅)ζ(·) is layer normalization. This mechanism restricts watermark-guided feature fusion to predicted tampered regions, preventing semantic leakage into authentic areas. After N Transformer layers, a lightweight decoder reconstructs the final recovered image I^ori I_ori. 3.5. AIGC Attack Simulator State-of-the-art deepfake generation pipelines typically consist of two main stages: (1) content synthesis in the latent space, and (2) image-space blending. Most existing watermarking methods primarily address post-blending perturbations, but often neglect the feature-level distortions introduced during generative synthesis (Zhang et al., 2025). To bridge this gap, we propose an AIGC Attack Simulator that jointly models both latent-space and image-space manipulations, thereby exposing the watermarking network to a broader spectrum of realistic and challenging attack scenarios during training. Importantly, the simulator is used only during training to enhance robustness and is not required at inference time. Our approach does not assume any prior knowledge of the specific AIGC tools used at test time. Figure 3. Overview of the proposed AIGC Attack Simulator. The simulator models AIGC edits via latent feature grafting and seamless image blending, enabling robust watermark training. Latent Mixing. To emulate the feature blending process inherent in modern deepfake generators, we design a Latent Mixing Module, simulating the latent-space manipulations commonly performed by AIGC models. As depicted in Figure 3, given a watermarked image proI_pro, we first select a source face srcI_src with similar facial landmarks (Shiohara and Yamasaki, 2022; Yu et al., 2025a), and then extract their latent representations pro,src∈ℝC×H×Wz_pro,z_src ^C× H× W using a pretrained VAE Encoder. We then randomly sample a binary channel mask c∈0,1CM_c∈\0,1\^C, and construct the mixed latent code as (4) gen=c⊙src+(1−c)⊙pro,z_gen=M_c _src+(1-M_c) _pro, where ⊙ denotes element-wise multiplication broadcasted across spatial dimensions. The resulting genz_gen is decoded by the VAE Decoder to obtain the generated image genI_gen. Such images naturally incorporate model-specific features and characteristics from other images, thereby more accurately simulating the AIGC editing process. Image Blending. To further simulate realistic face replacement, we employ Poisson blending (Pérez et al., 2003) to seamlessly integrate the source face srcI_src into the generated image genI_gen. Specifically, following the protocol of LVLM-DFD (Yu et al., 2025a), we generate random facial masks M based on detected facial landmarks to define the blending region. The final manipulated image editI_edit is then produced by solving the Poisson equation within the masked region, ensuring smooth and natural transitions between the synthesized face and the original background. After blending, we apply random degradations (e.g., JPEG compression, Gaussian noise) to further enhance robustness. 3.6. Training Objectives Watermark Embedding. To guarantee visual imperceptibility, we constrain both pixel-level distortion and perceptual discrepancy between the original image oriI_ori and the watermarked image wmI_wm: (5) ℒembed=∥wm−ori∥22+∑ℓ∥ϕℓ(wm)−ϕℓ(ori)∥1,L_embed= _wm-I_ori _2^2+ _ _ (I_wm)- _ (I_ori) _1, where ϕℓ _ denotes features extracted from the ℓ -th layer of a pretrained VGG-19 network. Message Decoding. To ensure robust recovery of both the ownership code wcopw_cop and the facial signature zfacez_face from manipulated images, we jointly optimize: (6) ℒdecode=BCE(w^cop,wcop)+∥z^face−zface∥22,L_decode=BCE( w_cop,w_cop)+ z_face-z_face _2^2, where w^cop w_cop and z^face z_face are the decoded ownership code and facial signature from the manipulated image, respectively. Forgery Localization. Manipulation localization is supervised using the Dice Loss (Milletari et al., 2016), which directly optimizes the overlap between the predicted manipulation map M and the ground-truth mask ∗M^*. The localization loss is defined as ℒlocL_loc. Guided Recovery. Restoration fidelity is enforced by constraining the reconstructed image ^ori I_ori to match the original oriI_ori in both pixel and perceptual domains: (7) ℒrec=∥^ori−ori∥1+∑ℓ∥ϕℓ(^ori)−ϕℓ(ori)∥1.L_rec= I_ori-I_ori _1+ _ _ ( I_ori)- _ (I_ori) _1. Overall Objective. The complete training objective is formulated as: (8) ℒtotal=ℒembed+λ1ℒdecode+λ2ℒloc+λ3ℒrec,L_total=L_embed+ _1L_decode+ _2L_loc+ _3L_rec, where λi _i are balancing coefficients. 4. Experiments 4.1. Experimental Settings Datasets. We evaluate our method on two widely-adopted high-quality datasets: CelebA-HQ (Karras et al., 2018) and FFHQ (Karras et al., 2019). CelebA-HQ comprises 30,000 high-resolution celebrity images. We follow the official data split protocol for training, validation, and testing. FFHQ contains 70,000 high-fidelity face images with substantial diversity in ethnicity, age, gender, and background complexity. We partition the dataset into 65,000/4,000/1,000 images for training/validation/testing respectively. All images are resized to 256×256256× 256 pixels for training and evaluation. Implementation Details. We use Adam optimizer with a learning rate of 2×10−42× 10^-4 and batch size 8. The loss weights are set as λ1=1.0 _1=1.0, λ2=1.0 _2=1.0, and λ3=2.0 _3=2.0. All experiments are conducted on four NVIDIA RTX 3090 GPUs. We use SD Inpainting (Rombach et al., 2022), HD-painter (Manukyan et al., 2023), E4S (Liu et al., 2023), InfoSwap (Gao et al., 2021), and MaskFaceGAN (Pernuš et al., 2023) as the AIGC attack models to evaluate performance under diverse manipulation scenarios. Additional implementation details are provided in Appendix. Baselines. We compare our method against state-of-the-art deep watermarking, tamper localization and image recovery methods, including: SepMark (Wu et al., 2023), TrustMark (Bui et al., 2025), EditGuard (Zhang et al., 2024b), Robust-Wide (Hu et al., 2024), OmniGuard (Zhang et al., 2025), StableGuard (Yang et al., 2025), WAM (Sander et al., 2025), CAT-NET (Kwon et al., 2021), PSCC-Net (Liu et al., 2022), HiFi-Net (Guo et al., 2023), NCL-IML (Zhou et al., 2023a), Imuge+ (Ying et al., 2023), and DFREC (Yu et al., 2025b). All baselines are implemented using their official codebases and pretrained weights. Evaluation Metrics. Imperceptibility is evaluated by Peak Signal-to-Noise Ratio (PSNR↑ ) and Structural Similarity Index (SSIM↑ ). Watermark robustness is quantified by Bit Accuracy↑ (%), the fraction of correctly decoded bits under each attack. Tamper localization is assessed using F1-score↑ , Area Under the Curve (AUC↑ ), and mean Intersection-over-Union (mIoU↑ ). Recovery quality is measured by PSNRrec PSNR_ rec↑ /SSIMrec SSIM_ rec↑ , and Fréchet Inception Distance (FID↓ ). Table 2. Tamper-localization results on CelebA-HQ (1,000 images). Metrics: F1 / AUC / mIoU (higher is better). Best and second-best per column are bold and underlined, respectively. Method SD Inpainting HD-painter Splicing F1 AUC mIoU F1 AUC mIoU F1 AUC mIoU CAT-NET 0.066 0.714 0.427 0.107 0.727 0.447 0.490 0.876 0.659 PSCC-NET 0.310 0.702 0.275 0.335 0.691 0.245 0.362 0.727 0.290 HiFi-Net 0.115 0.742 0.430 0.396 0.815 0.567 0.169 0.757 0.458 NCL-IML 0.085 0.518 0.423 0.087 0.516 0.421 0.040 0.503 0.406 MVSS-NET 0.262 0.843 0.484 0.378 0.866 0.540 0.639 0.935 0.707 Imuge+ 0.414 0.854 0.580 0.699 0.943 0.764 0.719 0.926 0.792 EditGuard 0.862 0.894 0.864 0.849 0.869 0.818 0.867 0.895 0.869 OmniGuard 0.869 0.983 0.880 0.914 0.995 0.908 0.804 0.980 0.817 StableGuard 0.686 0.482 0.548 0.707 0.384 0.570 0.997 1.000 0.994 WAM 0.495 0.787 0.492 0.174 0.503 0.315 0.782 0.921 0.707 VeriFi (Ours) 0.946 0.975 0.941 0.975 0.975 0.939 0.989 0.995 0.985 Figure 4. Visual comparison of tamper localization results across different methods. Table 3. Face recovery performance under six representative tampering types. Higher PSNRrec PSNR_ rec/SSIMrec SSIM_ rec indicate better fidelity; lower FID indicates higher realism. Attack types are grouped by generation paradigm: Diffusion-based (SD Inpainting, HD-painter), GAN-based (E4S, InfoSwap, MaskFaceGAN), and Traditional (Splicing). Best results per column are bold. Method Dataset Diffusion-based GAN-based Traditional SD Inpainting HD-painter E4S InfoSwap MaskFaceGAN Splicing PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID PSNRrec PSNR_ rec SSIMrec SSIM_ rec FID Imuge+ CelebA-HQ 21.49 0.741 29.60 16.91 0.753 118.07 18.52 0.625 129.94 17.69 0.674 68.41 19.65 0.622 105.23 27.07 0.906 34.52 DFREC CelebA-HQ 21.14 0.666 71.68 18.36 0.691 107.15 17.74 0.608 38.57 21.11 0.628 38.51 18.63 0.624 91.03 19.64 0.661 59.95 VeriFi (Ours) CelebA-HQ 31.05 0.894 39.25 28.34 0.892 45.68 25.12 0.839 26.71 27.33 0.866 18.82 25.73 0.845 41.18 31.06 0.921 21.51 Imuge+ FFHQ 16.46 0.745 94.05 19.54 0.697 47.13 11.28 0.542 121.56 15.32 0.592 78.10 18.34 0.619 63.57 29.52 0.891 31.48 DFREC FFHQ 20.90 0.656 73.03 10.47 0.342 199.65 16.64 0.577 50.04 20.33 0.598 49.32 17.19 0.575 127.08 17.26 0.606 71.49 VeriFi (Ours) FFHQ 30.61 0.905 20.21 26.84 0.853 29.83 21.1 0.788 45.36 24.25 0.766 38.84 25.19 0.866 38.20 31.60 0.937 23.56 4.2. Comparison with Localization Methods We evaluate tamper-localization performance on 1,000 CelebA-HQ images under three representative AIGC-based forgery types: SD Inpainting, HD-painter and face Splicing. Comparisons include recent passive detectors (CAT-NET, PSCC-Net, HiFi-Net, NCL-IML) and active watermarking and recovery-aware baselines (EditGuard, OmniGuard, Imuge+, StableGuard, WAM). Table 2 reports F1, AUC and mIoU for each tampering scenario. Passive methods exhibit low sensitivity to subtle generative alterations, while several active baselines struggle under AIGC-induced perturbations. Notably, VeriFi attains F1 scores of 0.946, 0.975, and 0.989 for SD Inpainting, HD-painter, and Splicing, respectively, significantly outperforming prior works. Qualitative examples in Figure 4 corroborate these quantitative findings and illustrate clearer boundaries for VeriFi compared to prior work. Additional experiments on FFHQ are provided in the Appendix. Both quantitative and qualitative results consistently demonstrate VeriFi’s superior tamper-localization capabilities across diverse datasets and manipulation types. Figure 5. Qualitative comparison of face recovery under SD Inpainting/HD-painter and Splicing. Figure 6. Additional qualitative comparison of face recovery under InfoSwap/E4S on CelebA-HQ. Figure 7. Additional qualitative comparison of face recovery under SD Inpainting on CelebA-HQ. Figure 8. Additional qualitative comparison of face recovery under Splicing on CelebA-HQ. 4.3. Comparison with Face Recovery Methods We comprehensively evaluate face recovery performance by benchmarking VeriFi against state-of-the-art methods, including Imuge+ (Ying et al., 2023) and DFREC (Yu et al., 2025b), on the CelebA-HQ dataset. As summarized in Table 3, VeriFi consistently achieves the best results across all tampering types and datasets, with substantial margins over prior works. VeriFi attains 29.50 dB PSNR and 0.904 SSIM on CelebA-HQ under SD Inpainting, outperforming Imuge+ and DFREC. Similar trends are observed for HD-painter and splicing, as well as on FFHQ, demonstrating the robustness of our approach. Qualitative results in Figure 5 6, 7, and 8 further highlight that VeriFi can faithfully reconstruct fine-grained facial details, even under challenging and large-area tampering. In contrast, existing methods often produce visible artifacts or fail to recover occluded regions. Notably, VeriFi maintains high visual quality and low FID, indicating that the recovered faces are not only visually plausible but also semantically consistent with the original content. Table 4. Quantitative comparison on CelebA-HQ. We report watermarked-image fidelity (PSNR in dB / SSIM) and watermark extraction accuracy (%) under representative AIGC manipulations and common degradations. The best results are in bold, and the second best are underlined. Method PSNR SSIM GAN-based (%) Diffusion-based (%) Common degradations (%) Avg. (%) InfoSwap E4S MaskFaceGAN HD-painter SD Inpainting JPEG ColorJitter Gaussian Noise TrustMark (Bui et al., 2025) 47.94 0.955 65.03 68.76 53.01 90.22 87.21 96.31 98.96 99.81 82.41 SepMark (Wu et al., 2023) 38.67 0.928 92.01 88.49 95.02 98.83 98.35 99.99 99.98 100.00 96.58 LampMark (Wang et al., 2024a) 44.05 0.971 89.21 74.80 65.61 83.73 89.79 84.13 84.50 84.88 82.08 Robust-Wide (Hu et al., 2024) 43.74 0.952 86.31 86.48 86.90 97.20 96.09 96.45 97.15 97.43 93.00 EditGuard (Zhang et al., 2024b) 32.24 0.758 48.92 77.34 78.56 82.07 91.39 79.55 95.07 93.19 80.76 OmniGuard (Zhang et al., 2025) 42.08 0.947 72.32 73.66 79.80 76.77 84.33 92.33 92.13 92.29 82.95 StableGuard (Yang et al., 2025) 31.81 0.848 50.08 87.93 89.16 88.13 95.58 59.78 99.93 63.61 79.28 WAM (Sander et al., 2025) 43.47 0.950 99.64 99.87 92.97 77.60 93.34 100.00 59.64 99.99 90.38 VeriFi (Ours) 42.53 0.948 99.60 90.51 95.62 98.72 97.29 99.82 100.00 100.00 97.70 4.4. Comparison with Deep Watermarking We compare VeriFi against eight state-of-the-art deep-watermarking methods (Bui et al., 2025; Wu et al., 2023; Hu et al., 2024; Zhang et al., 2024b, 2025; Yang et al., 2025; Sander et al., 2025; Wang et al., 2024a). Table 4 summarizes watermarked-image fidelity and watermark extraction accuracy under representative AIGC edits and common degradations on CelebA-HQ. Quantitatively, VeriFi attains a favorable fidelity–robustness balance with 42.53 dB PSNR and 0.948 SSIM, while achieving the highest mean extraction accuracy (97.70%). In particular, VeriFi yields the best performance on MaskFaceGAN and perfect extraction on ColorJitter and Gaussian Noise. It also maintains consistently strong robustness against AIGC manipulations, demonstrating resilience to both generative edits and conventional degradations. Additional qualitative comparisons, extended generalization experiments on FFHQ and ImageNet, and robustness evaluations under diverse degradations are provided in the Appendix. Collectively, these results further substantiate that VeriFi achieves state-of-the-art robustness and copyright protection, while maintaining high visual fidelity and imperceptibility. 4.5. Analysis and Ablation Studies Ablation on AIGC Attack Simulation. We quantify the individual and combined effects of the simulator components, namely noise augmentation, image blending, and latent space mixing, by training four variants: (i) without augmentation, (i) with noise augmentation, (i) with image blending, and (iv) with both blending and latent space mixing, which constitutes the full model referred to as Ours. Table 5 reports bit extraction accuracy (%) on 1,000 CelebA-HQ images under four representative attacks. The results show that both image blending and latent space mixing are required to realistically simulate AIGC attacks and to obtain robust watermark extraction. Table 5. Ablation of AIGC attack-simulator components. Bit extraction accuracy (%) under four representative attacks on 1,000 CelebA-HQ images. “Avg.” is the mean over the four attacks. Method E4S SimSwap SD Inpainting InfoSwap Avg. Without augmentation 50.97 53.04 63.69 54.11 55.45 With noise augmentation 57.77 66.63 79.27 73.26 69.23 With blending 85.47 76.34 87.95 84.16 83.48 Blending + latent mixing (Ours) 94.25 77.52 86.45 87.90 86.53 Ablation on Similarity Computation for Localization. We evaluate the effect of augmenting the localization head with a similarity branch in Table 6. Replacing segmentation-only supervision with an additional similarity branch improves F1 from 0.912 to 0.946, AUC from 0.966 to 0.975, and mIoU from 0.842 to 0.941, demonstrating that watermark-guided similarity cues substantially enhance localization accuracy. We further compare two pairing strategies during training: forming pairs from original (genuine) images and their watermarked counterparts, and forming pairs by concatenating unrelated images with watermarked images. Training with a mixture of genuine and synthetic composite pairs yields the largest improvement in localization performance. Table 6. Ablation: localization similarity metric and supervision on CelebA-HQ (SD Inpainting). Setting (Factor) F1 AUC mIoU Design Seg-only (no similarity branch) 0.912 0.966 0.842 Seg + Similarity (ours) 0.946 0.975 0.941 Supervision Original-only (genuine pairs) 0.881 0.939 0.808 Orig. + Synthetic (ours) 0.946 0.975 0.941 Table 7. Ablation on face-representation guidance. Evaluation on CelebA-HQ under SD Inpainting (1,000 images). PSNRrec_ rec in dB; Bit Acc. denotes watermark extraction accuracy; latency is per-image decode time on one RTX 3090. Guidance PSNRrec_ rec SSIMrec_ rec Bit Acc. (%) Latency (ms) 256-D embedding 23.15 0.709 97.54 49.77 576-D embedding 25.83 0.843 96.06 51.15 1024-D embedding 31.05 0.894 97.29 53.57 Ablation on Watermark Guidance for Recovery. We evaluate the impact of VAE-based face latent dimensionality (256, 576, 1024) on recovery. Latents are extracted by resizing the input to different resolutions before VAE encoding, thus controlling embedding size. As shown in Table 7, higher-dimensional latents yield improved PSNRrec and SSIMrec with negligible effect on watermark extraction accuracy and inference latency. This demonstrates that richer semantic priors enhance recovery fidelity without compromising robustness or efficiency. Efficiency Analysis. We analyze the computational efficiency of VeriFi by measuring the number of model parameters, floating-point operations (FLOPs), and inference latency for each component on a single RTX 3090 GPU. Table 8 summarizes these metrics. The total model size is approximately 157.12 million parameters, with a cumulative FLOP count of 384.97 billion. The end-to-end inference latency for watermark embedding, extraction, tamper localization, and face recovery is 82.75 milliseconds per image, indicating that VeriFi is suitable for real-time or near-real-time applications in practical scenarios. Table 8. Efficiency analysis of VeriFi on a single RTX 3090 GPU. Component Params (M) FLOPs (G) Latency (ms) Watermark Embedder 70.88 67.35 18.41 Watermark Extractor 15.03 231.43 29.27 Tamper Locator 37.14 25.44 14.71 Face Recovery Network 34.07 60.75 20.36 Total 157.12 384.97 82.75 5. Conclusion We propose VeriFi, a unified and versatile watermarking framework that enables robust copyright tracing, precise tamper localization, and high-fidelity facial recovery. By jointly embedding ownership signatures and facial representations, VeriFi demonstrates strong resistance against both pixel-domain and AIGC-based attacks. Extensive experiments on CelebA-HQ and FFHQ validate the superiority of our approach in terms of extraction accuracy, localization precision, and recovery quality. Our method provides a comprehensive solution for protecting and verifying digital face images in the era of advanced generative models. Future work may explore extending this framework to video content and other modalities. References T. K. Araghi, D. Megías, V. Garcia-Font, M. Kuribayashi, and W. Mazurczyk (2024) Disinformation detection and source tracking using semi-fragile watermarking and blockchain. In Proceedings of the 2024 European Interdisciplinary Cybersecurity Conference, p. 136–143. Cited by: §1. T. Bui, S. Agarwal, and J. Collomosse (2025) TrustMark: robust watermarking and watermark removal for arbitrary resolution images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 18629–18639. Cited by: §4.1, §4.4, Table 4. G. Gao, H. Huang, C. Fu, Z. Li, and R. He (2021) Information bottleneck disentanglement for identity swapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 3404–3413. Cited by: §4.1. X. Guo, X. Liu, Z. Ren, S. Grosz, I. Masi, and X. Liu (2023) Hierarchical fine-grained image forgery detection and localization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 3155–3165. Cited by: §2.1, §4.1. R. Hu, J. Zhang, T. Xu, J. Li, and T. Zhang (2024) Robust-wide: robust watermarking against instruction-driven image editing. In European Conference on Computer Vision, p. 20–37. Cited by: §4.1, §4.4, Table 4. T. Karras, T. Aila, S. Laine, and J. Lehtinen (2018) Progressive growing of gans for improved quality, stability, and variation. In International Conference on Learning Representations (ICLR), Cited by: §4.1. T. Karras, S. Laine, and T. Aila (2019) A style-based generator architecture for generative adversarial networks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4.1. M. Kwon, I. Yu, S. Nam, and H. Lee (2021) CAT-net: compression artifact tracing network for detection and localization of image splicing. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, p. 375–384. Cited by: §2.1, §4.1. Y. Li, F. Cheng, W. Yu, G. Wang, G. Luo, and Y. Zhu (2024) Adaifl: adaptive image forgery localization via a dynamic and importance-aware transformer network. In European Conference on Computer Vision, p. 477–493. Cited by: §2.1. X. Liu, Y. Liu, J. Chen, and X. Liu (2022) PSCC-net: progressive spatio-channel correlation network for image manipulation detection and localization. IEEE Transactions on Circuits and Systems for Video Technology 32 (11), p. 7505–7517. Cited by: §4.1. Y. Liu, S. Chen, H. Shi, X. Zhang, S. Xiao, and Q. Cai (2025) MUN: image forgery localization based on m3 encoder and un decoder. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 5685–5693. Cited by: §1, §2.1. Z. Liu, M. Li, Y. Zhang, C. Wang, Q. Zhang, J. Wang, and Y. Nie (2023) Fine-grained face swapping via regional gan inversion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 8578–8587. Cited by: §4.1. Z. Lou, G. Cao, K. Guo, L. Yu, and S. Weng (2025) Exploring multi-view pixel contrast for general and robust image forgery localization. IEEE Transactions on Information Forensics and Security. Cited by: §2.1. H. Manukyan, A. Sargsyan, B. Atanyan, Z. Wang, S. Navasardyan, and H. Shi (2023) HD-painter: high-resolution and prompt-faithful text-guided image inpainting with diffusion models. arXiv preprint arXiv:2312.14091. Cited by: §4.1. F. Milletari, N. Navab, and S. Ahmadi (2016) V-net: fully convolutional neural networks for volumetric medical image segmentation. In 2016 Fourth International Conference on 3D Vision (3DV), Vol. , p. 565–571. External Links: Document Cited by: §3.6. P. Neekhara, S. Hussain, X. Zhang, K. Huang, J. McAuley, and F. Koushanfar (2024) Facesigns: semi-fragile watermarks for media authentication. ACM Transactions on Multimedia Computing, Communications and Applications 20 (11), p. 1–21. Cited by: §1. P. Pérez, M. Gangnet, and A. Blake (2003) Poisson image editing. In ACM Transactions on Graphics (TOG), Vol. 22, p. 313–318. Cited by: §3.5. M. Pernuš, V. Štruc, and S. Dobrišek (2023) MaskFaceGAN: high-resolution face editing with masked gan latent code optimization. IEEE Transactions on Image Processing 32 (), p. 5893–5908. External Links: Document Cited by: §4.1. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10684–10695. Cited by: §4.1. T. Sander, P. Fernandez, A. Durmus, T. Furon, and M. Douze (2025) Watermark anything with localized messages. In International Conference on Learning Representations-ICLR 2025, Cited by: §1, §2.2, §4.1, §4.4, Table 4. K. Shiohara and T. Yamasaki (2022) Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 18720–18729. Cited by: §3.5. C. Tan, Y. Zhao, S. Wei, G. Gu, P. Liu, and Y. Wei (2024a) Frequency-aware deepfake detection: improving generalizability through frequency space domain learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 5052–5060. Cited by: §1. C. Tan, Y. Zhao, S. Wei, G. Gu, P. Liu, and Y. Wei (2024b) Rethinking the up-sampling operations in cnn-based generative network for generalizable deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 28130–28139. Cited by: §1. M. Tancik, B. Mildenhall, and R. Ng (2020) Stegastamp: invisible hyperlinks in physical photographs. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2117–2126. Cited by: §2.2. R. Wang, F. Juefei-Xu, M. Luo, Y. Liu, and L. Wang (2021) Faketagger: robust safeguards against deepfake dissemination via provenance tracking. In Proceedings of the 29th ACM international conference on multimedia, p. 3546–3555. Cited by: §1, §2.2. T. Wang, H. Cheng, M. Liu, and M. Kankanhalli (2025) Fractalforensics: proactive deepfake detection and localization via fractal watermarks. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 7210–7219. Cited by: §1. T. Wang, M. Huang, H. Cheng, X. Zhang, and Z. Shen (2024a) Lampmark: proactive deepfake detection via training-free landmark perceptual watermarks. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 10515–10524. Cited by: §1, §2.2, §4.4, Table 4. T. Wang, X. Liao, K. P. Chow, X. Lin, and Y. Wang (2024b) Deepfake detection: a comprehensive survey from the reliability perspective. 57 (3). External Links: ISSN 0360-0300, Link, Document Cited by: §1. X. Wu, X. Liao, and B. Ou (2023) Sepmark: deep separable watermarking for unified source tracing and deepfake detection. In Proceedings of the 31st ACM International Conference on Multimedia, p. 1190–1201. Cited by: §1, §2.2, §4.1, §4.4, Table 4. C. Yang, Z. Wang, H. Shen, H. Li, and B. Jiang (2021) Multi-modality image manipulation detection. In 2021 IEEE International Conference on Multimedia and Expo (ICME), Vol. , p. 1–6. External Links: Document Cited by: §2.1. H. Yang, B. Liu, X. Xu, C. Xu, Y. Yu, Z. Huang, Y. Wang, and S. He (2025) StableGuard: towards unified copyright protection and tamper localization in latent diffusion models. Advances in Neural Information Processing Systems. Cited by: §1, §2.2, §4.1, §4.4, Table 4. Z. Yang, K. Zeng, K. Chen, H. Fang, W. Zhang, and N. Yu (2024) Gaussian shading: provable performance-lossless image watermarking for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 12162–12171. Cited by: §1. Q. Ying, H. Zhou, Z. Qian, S. Li, and X. Zhang (2023) Learning to immunize images for tamper localization and self-recovery. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), p. 13814–13830. Cited by: §2.2, §4.1, §4.3. P. Yu, J. Fei, H. Gao, X. Feng, Z. Xia, and C. H. Chang (2025a) Unlocking the capabilities of large vision-language models for generalizable and explainable deepfake detection. In Forty-second International Conference on Machine Learning, Cited by: §3.5, §3.5. P. Yu, H. Gao, J. Fei, Z. Huang, Z. Xia, and C. Chang (2025b) DFREC: deepfake identity recovery based on identity-aware masked autoencoder. External Links: 2412.07260, Link Cited by: §2.1, §4.1, §4.3. L. Zhang, M. Xu, D. Li, J. Du, and R. Wang (2024a) Catmullrom splines-based regression for image forgery localization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 7196–7204. Cited by: §2.1. X. Zhang, R. Li, J. Yu, Y. Xu, W. Li, and J. Zhang (2024b) Editguard: versatile image watermarking for tamper localization and copyright protection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 11964–11974. Cited by: §1, §2.2, §4.1, §4.4, Table 4. X. Zhang, Z. Tang, Z. Xu, R. Li, Y. Xu, B. Chen, F. Gao, and J. Zhang (2025) Omniguard: hybrid manipulation localization via augmented versatile deep image watermarking. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 3008–3018. Cited by: §1, §2.2, §3.5, §4.1, §4.4, Table 4. X. Zhao, K. Zhang, Z. Su, S. Vasan, I. Grishchenko, C. Kruegel, G. Vigna, Y. Wang, and L. Li (2024) Invisible image watermarks are provably removable using generative ai. Advances in neural information processing systems 37, p. 8643–8672. Cited by: §1. J. Zhou, X. Ma, X. Du, A. Y. Alhammadi, and W. Feng (2023a) Pre-training-free image manipulation localization through non-mutually exclusive contrastive learning. In Proceedings of the IEEE/CVF international conference on computer vision, p. 22346–22356. Cited by: §4.1. J. Zhou, X. Ma, X. Du, A. Y. Alhammadi, and W. Feng (2023b) Pre-training-free image manipulation localization through non-mutually exclusive contrastive learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 22346–22356. Cited by: §2.1. J. Zhu, R. Kaplan, J. Johnson, and L. Fei-Fei (2018) Hidden: hiding data with deep networks. In Proceedings of the European conference on computer vision (ECCV), p. 657–672. Cited by: §2.2.