Paper deep dive
AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
Hao Wang, Beichen Zhang, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/27/2026, 7:31:20 PM
Summary
AIFIND is a proposed framework for Incremental Face Forgery Detection (IFFD) that addresses feature drift and catastrophic forgetting by using semantic anchors. It consists of three main components: the Artifact-Driven Semantic Prior Generator (ASPG) which creates stable semantic anchors from low-level artifact cues, the Artifact-Probe Attention (APA) which injects these anchors into the image encoder for fine-grained alignment, and the Adaptive Decision Harmonizer (ADH) which maintains geometric consistency of decision boundaries across tasks. The method is data-replay-free and leverages the intrinsic semantic sensitivity of Vision-Language Models.
Entities (11)
Relation Signals (10)
AIFIND → addresses → Incremental Face Forgery Detection
confidence 100% · AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
Adaptive Decision Harmonizer → aligns → classifiers
confidence 100% · ADH harmonizes the classifiers by preserving angular relationships of semantic anchors
AIFIND → contains → Artifact-Driven Semantic Prior Generator
confidence 100% · AIFIND consists of the following cooperative components: (1) Artifact-Driven Semantic Prior Generator (ASPG)...
AIFIND → contains → Adaptive Decision Harmonizer
confidence 100% · AIFIND consists of the following cooperative components: ... (4) Adaptive Decision Harmonizer (ADH)...
AIFLED → contains → Adaptive Decision Harmonizer
confidence 100% · AIFIND consists of the following cooperative components: ... (4) Adaptive Decision Harmonizer (ADH)
AIFIND → contains → Artifact-Probe Attention
confidence 100% · AIFIND consists of the following cooperative components: ... (2) Artifact-Probe Attention (APA) module...
AIFIND → contains → Artifact-Driven Semantic Prior Generator
confidence 100% · AIFIND consists of the following cooperative components: (1) Artifact-Driven Semantic Prior Generator (ASPG)...
AIFIND → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND.
Tags
Links
- Source: https://arxiv.org/abs/2604.16207v1
- Canonical: https://arxiv.org/abs/2604.16207v1
Trouble viewing inline? Open PDF directly →
Full Text
54,108 characters extracted from source content.
Expand or collapse full text
AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection Hao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang School of Computer Science and Technology, Harbin Institute of Technology, Weihai 2023210984,2023211640,shaoyi.fang@stu.hit.edu.cn beiczhang,qizb,xuyuanrong,xinyliu,wgzhang@hit.edu.cn Abstract As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. How- ever, existing methods typically rely on data replay or coarse bi- nary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instan- tiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintain- ing geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND. CCS Concepts • Security and privacy→Human and societal aspects of se- curity and privacy. Keywords Face Forgery Detection, Incremental Learning, Semantic Anchors, Fine-Grained Visual-Text Alignment ACM Reference Format: Hao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuan- rong Xu, Xinyan Liu, Weigang Zhang. 2026. AIFIND: Artifact-Aware Inter- preting Fine-Grained Alignment for Incremental Face Forgery Detection. In Proceedings of the 2026 International Conference on Multimedia Retrieval (ICMR ’26). ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/ 3805622.3810877 ∗ Corresponding author. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. ICMR ’26, Amsterdam, Netherlands © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/10.1145/3805622.3810877 Dynamic Matching Image Encoder Artifact-Probe Attention Text Encoder Inconsistent Lighting Blurry Eyes UnnaturalJawline Task t Real? Fake? Blur eyes 0.95 Smooth lip 0.08 ...... Classifier Replay Set Task t Real? Fake? (a) Existing IFFD methods (b) Ours Binary label Head Multi label Head (c) Results Figure 1: Comparison between AIFIND and other methods. (a) Conventional methods depend on replay sets to mitigate for- getting. (b) Our method realizes a data-replay-free paradigm via stable semantic anchors. Visual features are continuously aligned with stable semantic priors, ensuring consistent de- cision boundaries across tasks. 1 Introduction With the rapid development of generative models, face forgery has advanced and achieved unprecedented realism, posing serious pub- lic safety risks. Thus, developing forgery detection becomes crucial for maintaining security in the digital ecosystem. Existing main- stream face forgery detection methods [5,20,31,36,49] attempt to train a generalizable detector with limited and static datasets. However, new forgery techniques emerge endlessly, which quickly render existing detection methods ineffective. In view of this, Incre- mental Face Forgery Detection (IFFD) is proposed to continuously train the model with the latest forged samples. Currently, existing methods [4,24,37] for IFFD mainly rely on data replay, which preserves knowledge by storing a small subset of past samples. Representative approaches such as DFIL [24] and SUR- LID [4] follow this paradigm. However, merely replaying discrete samples fails to explicitly constrain the topology of the feature space. Regardless of data replay or regularization, as shown in Fig. 1(a), current IFFD methods rely on coarse binary supervision, failing to leverage fine-grained artifact cues. Critically, in incremental learning settings, without stable anchors to stabilize the latent space, the learned feature distribution is prone to unconstrained drift. Consequently, when accommodating new forgery types, the model inadvertently overwrites previously learned representations, resulting in catastrophic forgetting. arXiv:2604.16207v1 [cs.CV] 17 Apr 2026 ICMR ’26, June 16-19, 2026, Amsterdam, NetherlandsHao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang Furthermore, while recent works [5,20,36] have sought to in- tegrate Vision-Language Models (VLMs) to enhance deepfake de- tection, they primarily focus on static scenarios and fail to apply them in incremental learning. Existing VLM-based methods often treat semantic information as a coarse global label, neglecting the fine-grained semantic discrepancies between authentic and forged facial regions. Motivated by this, we rethink the IFFD paradigm by exploiting the intrinsic semantic sensitivity of pre-trained VLMs toward specific artifact-prone regions. We posit that linguistic con- cepts possess a natural invariance: while the visual manifestations of local anomalies vary across datasets, the semantic distinction between authentic and manipulated facial components remains constant. Leveraging this property, we propose to utilize these region-aware authenticity priors as invariant semantic anchors. By anchoring visual features to these stable semantic coordinates, we can explicitly stabilize the evolving visual feature space. As shown in Fig. 1(b), we interpret artifacts as invariant semantic anchors. This formulation establishes a stable, high-dimensional ref- erence for fine-grained supervision, enabling semantically guided artifact inference to resist feature drift. Based on this, we propose Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection (AIFIND). AIFIND consists of the following cooperative components: (1) Artifact-Driven Semantic Prior Gener- ator (ASPG) constructs the initial semantic anchors by interpreting volatile low-level artifact cues into stable sparse textual labels; (2) Artifact-Probe Attention (APA) module injects selected textual artifact cues into the image encoder for fine-grained visual–text alignment; (3) Semantic-Guided Incremental Detector (SGID) gov- erns the learning process within this anchored space, leveraging APA and dual supervision to simultaneously discriminate authen- ticity and identify specific artifact types; and (4) Adaptive Decision Harmonizer (ADH) aligns binary and multi-label heads by strictly preserving the angular relationships relative to the semantic an- chors, guaranteeing consistency in decision boundaries. During training, AIFIND executes a dynamic matching strategy to achieve stable incremental learning. First, ASPG instantiates semantic anchors to form a fixed coordinate system. The model then transitions to a dynamic matching mechanism, autonomously recalibrating targets based on similarity. Then, APA injects these matched semantic anchors to enforce fine-grained anchoring. This process continuously rectifies volatile visual features via stable semantic definitions. Finally, ADH harmonizes the classifiers by preserving their angular relationships relative to the semantic an- chors, maintaining geometric consistency across tasks. As a result, AIFIND effectively mitigates catastrophic forgetting, enabling se- mantically coherent knowledge evolution without replay buffers. Experiments under multiple incremental protocols [4] validate the superiority and strong generalization capability of our method. Our main contributions are summarized as follows: •We propose a data-replay-free framework for IFFD, which reinterprets artifacts as semantic anchors to explicitly con- strain the feature space without storing past data. •We formulate an artifact-probe attention that effectively an- chors evolving visual forgery cues to immutable semantic anchors, preventing feature drift. •We conduct comprehensive experiments that empirically validate the superiority of our framework, proving the effec- tiveness of semantic anchors in incremental learning. 2 Related Work 2.1 Face Forgery Detection Face forgery detection has become a prominent research topic in computer vision. Early studies mainly relied on detecting anomalies in biometric cues such as eye blinking [18], head pose [46], and pupil morphology [7]. With the rapid advancement of deep learning, attention gradually shifted toward learning forgery traces from mul- timodal clues, such as frequency domain [13,49] and temporal con- sistency [8,45]. Recently, the emergence of vision–language models such as CLIP [25] has opened up new possibilities for semantic- level deepfake detection. RepDFD [20] reprograms CLIP by inject- ing universal perturbations into input images, while ForAda [5] leverages a Forensics Adapter to learn hybrid boundary artifacts specific to facial forgeries. VLFFD [36] attempts fine-grained vi- sual–text alignment by incorporating detailed textual semantics, but its dependence on paired real–fake data limits its applicability to incremental learning. These approaches [5,31,42,45,49] aim to extract universal forgery representations from limited observed data and achieve effective transfer to unseen manipulation types. 2.2 Incremental Face Forgery Detection As forgery techniques evolve rapidly, constructing a general detec- tor from limited training datasets has become increasingly impracti- cal. This motivates the exploration of Incremental Face Forgery De- tection (IFFD), which enables models to continuously learn emerg- ing forgery patterns while preserving previous knowledge. Incremental learning methods are commonly categorized into three paradigms: parameter isolation [38,39], parameter regular- ization [1,29], and data replay [2,22,34]. In the context of IFFD, most existing studies emphasize knowledge distillation and replay mechanisms. CoReD [15] preserves old knowledge through task- aware distillation while adapting to new domains. DFIL [24] replays representative and challenging samples from previous datasets, and DMP [37] dynamically expands prototypes to accommodate newly emerging forgery types. HDP [35] achieves replay through UAPs [23], while SUR-LID [4] aligns latent features across tasks to mitigate mutual interference. However, these methods still treat all forgery instances as a single Fake category, which limits their ability to capture fine-grained artifact semantics. Meanwhile, Vision–Language Models (VLMs) such as CLIP [47] have inspired new paradigms for incremental learning. Prompt- based approaches [16,28,33,40] design task-specific prompts to guide CLIP models, while adapter-based techniques [9,17,48] insert lightweight trainable modules at intermediate layers to encapsu- late new knowledge. Despite their success in general incremental learning, these methods are not specifically designed for the unique challenges of deepfake detection, where subtle visual cues play a critical role, leading to limited effectiveness when applied directly. AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery DetectionICMR ’26, June 16-19, 2026, Amsterdam, Netherlands Artifact-ProbeAttention Multi-Head Cross-Attention gate Self-Attention Feed-Forward Network Layer Norm mask Layer Norm Query Key&Value Task t Boundary Align push Semantic Align Binary Consensus Boundary Task t Boundary Align push Semantic Align push Task t+1 Incrementing ASPG Inconsistent Lighting Blurry Eyes UnnaturalJawline Task t Task t+1 Backbone Incrementing Attention Map Embedding Features Image Encoder Text Encoder Backbone Similarity Matrix Top-k ASPG Binary label Head Multi label Head ℒ ce ℒ tag Multi label Head ℒ dis Binary label Head Artifact- Probe Attention AdaptiveDecisionHarmonizer Semantic-GuidedIncrementalDetector Artifact-drivenSemanticPriorGenerator mask raw scores normalized Inconsistent Lighting Blurry Eyes Unnatural Jawline top-k Dynamic Matching Figure 2: Overall framework of AIFIND. Artifact-Driven Semantic Prior Generator (ASPG) instantiates semantic anchors to build a fixed coordinate system. Semantic-Guided Incremental Detector (SGID) uses the Artifact-Probe Attention (APA) to perform the anchoring, which constrains volatile visual features to these stable semantic anchors. Adaptive Decision Harmonizer (ADH) maintains geometric semantic consistency, preserving semantic angular relationships across tasks. Table 1: Correspondence between Indicators and Facial Re- gions.✓ indicates that the specific operator is applied. Facial RegionBlurColorStructureTextureBoundary Eyes✓✗ Nose✓✗ Cheeks✓✗ Mouth✓✗ Jawline✗✓ Boundary✗✓ 3 Method 3.1 Framework overview As illustrated in Fig. 2, we introduce AIFIND to enable stable incre- mental learning via a semantically anchored feature space. Specifi- cally, ASPG transforms low-level cues into semantic anchors. These anchors are injected into image encoder by SGID using APA to ensure fine-grained visual–text alignment. Finally, ADH aligns clas- sifier weights to preserve geometric consistency across tasks. 3.2 Artifact-Driven Semantic Prior Generator According to [36] and our analysis, we define 5 representative forgery dimensionsI= I blur ,I color ,I structure ,I text ,I boundary . To spatially ground these dimensions, we utilize MediaPipe Face Mesh [21] to locate 6 facial regionsR=푅 eyes ,푅 nose ,푅 cheeks ,푅 mouth ,푅 jawline , 푅 boundary . Additionally, a global skin reference region푆is extracted to serve as the baseline for computing relative inconsistencies. The correspondence between facial regions and indicators is in Tab. 1 and definitions of indicators are as follows: Blur Indicator. We measure local sharpness using the variance of the Laplacian operator: I blur = Var ∇ 2 퐼 M ,(1) where퐼 M denotes the grayscale intensities within the region mask. Color Indicator. To capture lighting inconsistencies, we calcu- late the luminance deviation between the target region푅and the skin reference 푆 in CIELAB space: I color = | 퐿 푅 − 퐿 푆 | ,(2) where퐿 푅 and퐿 푆 represent the mean퐿-channel values, respectively. Structural Indicator. We assess the structural compatibility between the target region and the surrounding skin using the Struc- tural Similarity Index (SSIM): I struct = SSIM ( 푃 푅 ,푃 푆 ) ,(3) where 푃 푅 and 푃 푆 are normalized grayscale patches from 푅 and 푆 . Texture Indicator. To expose statistical anomalies in skin tex- ture, we extract local contrast using the Gray-Level Co-occurrence Matrix (GLCM): I text = Contrast ( GLCM ( 퐼 M )) ,(4) where 퐼 M is quantized to 64 levels to retain details. Boundary Indicator. We identify potential blending artifacts along edges by computing the average gradient magnitude: I bound = E (푥,푦)∈M √︃ 퐺 2 푥 +퐺 2 푦 ,(5) where 퐺 푥 and 퐺 푦 denote Sobel derivatives. To bridge pixel-level statistics with high-level semantics, for each artifact dimension푖 ∈ Iand region푔 ∈ R, we leverage an ICMR ’26, June 16-19, 2026, Amsterdam, NetherlandsHao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang LLM to generate a candidate set of퐾contrastive text pairs, denoted asT 푖,푔 =(푡 r 푘 ,푡 f 푘 ) 퐾 푘=1 , and align them with a support setΩ 푖,푔 . The optimal anchor 퐴 푖,푔 is selected by calculating CLIP similarity: 퐴 푖,푟 = argmax (푡 r ,푡 f )∈T 푖,푔 ∑︁ (푥 r ,푥 f )∈Ω 푖,푔 Sim(푡 f ,푥 f )+ Sim(푡 r ,푥 r ) ,(6) where푟refers to real and푓refers to fake. By iterating this process across all indicators and facial regions, we construct a Semantic Anchor Library A=퐴 푖,푔 | 푖 ∈I,푔 ∈ R. For a forged image푥 f , we target the most severe anomalies by selecting the top-푁highest scores, and retrieve their corresponding forgery descriptions푡 f . Conversely, for a real image푥 r , we focus on the most pristine facial details by selecting the bottom-푁lowest scores, assigning the linked authentic descriptions 푡 r . 3.3 Semantic-Guided Incremental Detector Building upon the instantiated semantic anchors, we propose the Semantic-Guided Incremental Detector (SGID) to explicitly inject these semantic anchors into the visual learning process. The frame- work is underpinned by two components: Artifact-Probe Atten- tion (APA) and Dual Supervision. Artifact-Probe Attention. To effectively incorporate the se- mantic priors, we introduce the APA module within the vision transformer. Let푋 ∈ R 푃×퐷 denote the intermediate visual em- beddings, where푃is the number of patches. Simultaneously, let 푆=Φ text (A 푚푎푡푐ℎ푒푑 ) ∈ R 푁×퐷 represent the textual embeddings of the matched semantic anchors, where푁is the number of selected anchors. The APA module functions as a cross-modal bridge, em- ploying a Multi-Head Attention mechanism where visual tokens serve as queries to probe the semantic details in the embeddings: ̃ 푋= MHA(푄= 푋,퐾= 푆,푉= 푆).(7) To ensure adaptive integration, we employ a residual connection with a learnable gating parameter 푔 ∈ R 퐷 : 푋 fused = 푋 +푔⊙ ̃ 푋,(8) where⊙denotes element-wise multiplication. The gating coeffi- cient푔dynamically modulates the injection of semantic priors, preventing the overriding of intrinsic visual cues. The fused repre- sentation푋 fused then proceeds to the subsequent layers. In practice, APA is injected into the top-푀 transformer layers. Dual Supervision. To enforce the alignment between visual features and the selected semantic anchors, we employ a dual super- vision strategy. Given an input image푥, let퐹=Φ img (푥,푆)denote the visual features extracted by the APA-enhanced image encoder. First, a binary classifier퐶 푡 is employed to predict global authen- ticity, formulated as: L cls = CE(퐶 푡 (퐹),푌 bin ),(9) where푌 bin ∈ 0,1is the ground-truth label (Real/Fake), andCE(·) denotes the standard cross-entropy loss. Simultaneously, to ensure the model comprehends the specific forgery patterns described by the semantic anchors, a multi-label head퐻 푡 is utilized to predict the presence of the defined artifact dimensions (I blur ,I color ,I structure ,I text ,I boundary ). The artifact dimen- sions prediction loss is defined as: L ind = BCE(퐻 푡 (퐹),푌 ind ),(10) where푌 ind ∈ 0,1 |I| denotes the binary artifact indicator vector corresponding to the dimensions inI, andBCE(·)is the binary cross-entropy loss applied independently to each attribute. By jointly optimizingL cls andL ind , the model is encouraged to align its visual features with the stable semantic anchors. The stationary semantic space acts as a persistent reference throughout the incremental learning process, effectively preventing feature drift and mitigating catastrophic forgetting. 3.4 Adaptive Decision Harmonizer In incremental learning, the decision boundaries of classifiers tend to drift as new tasks are introduced. To mitigate this, we propose ADH to align classifier weights to preserve geometric consistency across tasks. Let ̃ 푊 (푏) 푗 and ̃ 푊 (푚) 푗 respectively denote the normalized classifier weights of the binary-label and multi-label heads from previous tasks. To quantify the semantic affinity between the current task and historical knowledge, ADH computes adaptive similarity weights for each head ℎ ∈ 푏,푚: 휔 (ℎ) 푗 = exp cos( ̃ 푊 (ℎ) 푖 , ̃ 푊 (ℎ) 푗 )/휏 Í 푙≠푖 exp cos( ̃ 푊 (ℎ) 푖 , ̃ 푊 (ℎ) 푙 )/휏 ,(11) where휏is a temperature parameter and푗≠ 푖. High similarity implies that the current artifacts share underlying semantic traits with task 푗 , warranting stronger alignment. We then construct a Global Semantic Reference ̃ 푊 (ℎ) ref by aggre- gating previous classifiers according to their semantic affinity: ̃ 푊 (ℎ) ref = norm ∑︁ 푗≠푖 휔 (ℎ) 푗 ̃ 푊 (ℎ) 푗 ! .(12) This reference vector captures the stable historical decision trend, serving as a robust alignment anchor to prevent the model from forgetting previously learned forgery patterns. To incorporate the historical knowledge without distorting the feature space, we perform a spherical semantic alignment. Unlike Euclidean interpolation, this process respects the geometric struc- ture of the hypersphere, rotating the current decision boundary toward the global reference along the geodesic path: ̃ 푊 (ℎ) new = sin((1−푡 (ℎ) )휃) sin휃 ̃ 푊 (ℎ) 푖 + sin(푡 (ℎ) 휃) sin휃 ̃ 푊 (ℎ) ref ,(13) where휃represents the angular distance between the current and reference weights. Crucially, the alignment coefficient푡 (ℎ) ∈ [0,1] is adaptively determined to balance plasticity and stability: 푡 (ℎ) = cos( ̃ 푊 (ℎ) 푖 , ̃ 푊 (ℎ) ref ),(14) where a high cosine similarity indicates that the current task is semantically consistent with history, triggering a gentle update to preserve existing anchors. Finally, to decouple the semantic directional alignment from the magnitude-dependent feature strength, we rescale the aligned weights to recover the norm of푊 (ℎ) 푖 : 푊 (ℎ) new = ̃ 푊 (ℎ) new ·∥푊 (ℎ) 푖 ∥ 2 ,(15) which ensures that the alignment modifies only the semantic direc- tion without degrading the detector’s discriminative sensitivity. AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery DetectionICMR ’26, June 16-19, 2026, Amsterdam, Netherlands By constraining the update trajectory to the spherical manifold, ADH prevents the decision boundaries from drifting into seman- tically ambiguous regions, guaranting that the evolving detector remains geometrically consistent and semantically coherent with the accumulated global anchor library. 3.5 Training strategy and overall loss Training strategy. During the initial푛epochs, the semantic an- chors푆are fixed to the descriptions generated by ASPG, providing stable and consistent guidance to the backbone. Starting from the (푛+1)-th epoch, we activate the dynamic matching mechanism. In this phase,푆is adaptively retrieved by calculating the cosine simi- larity between visual features and textual embeddings, enabling the model to capture fine-grained, instance-specific forgery patterns. Overall loss. Following [24], we also maintain the previous-task learned information via knowledge distillation loss, which is: L dis =∥Φ img (푥,푆;휃 푡 )−Φ img (푥,푆;휃 푡−1 )∥ 2 2 ,(16) whereΦ img (·;휃 푡−1 )is the frozen backbone extractor trained on the previous(푡 −1)-th task, which serves as a reference to regularize the feature space on the current data. The total loss function is defined as: L overall =L cls + 휇 1 L ind + 휇 2 L dis ,(17) where 휇 1 and 휇 2 are trade-off coefficients. 4 Experiment 4.1 Experimental Settings Datasets. Our datasets follow the protocol in [4] to ensure a compre- hensive evaluation. We use several forgery face datasets, including DeepFake Detection Challenge Preview (DFDCP) [6], Celeb-DF-v2 (CDF) [19], and FaceForensics++ [27]. Furthermore, we incorporate recently released datasets that feature more diverse and sophisti- cated forgery methods, namely MCNet [10], BlendFace [32], Style- GAN3 [12] from DF40 [43] and SDv21 [26] from DiffusionFace [3]. Evaluation Protocol. To systematically evaluate the performance of our model, we adopt the standard evaluation protocols in [4]. •Protocol 1 (P1): Datasets Incremental with SDv21, F++, DFDCP, CDF. This protocol simulates scenarios in which the model is required to adapt to entirely novel data envi- ronments across different incremental steps. •Protocol 2 (P2): Forgery Categories Incremental with Hybrid (F++), Face-Reenactment (MCNet), Face-Swapping (BlendFace), Entire Face Synthesis (StyleGAN3). The model continuously learns to defend against new forgery types, while the distribution of genuine data remains constant. Implementation Details. We use CLIP-ViT-L/14 [25] as backbone, fine-tuned via LN-tuning [47]. The Adam optimizer is employed with a learning rate of 8×10 −5 , 20 epochs, and batch size of 32. For replay-based baselines, the replay buffer size is set to 500 for each task. We set휇 1 =0.1,휇 2 =1 and set the number of selected semantic anchors to푁=3. To ensure training stability, we set the warm-up period to푛=5. The APA modules are injected into the last푀=4 layers of the image encoder. We employ Frame-level Area Under the Curve (AUC) as the evaluation metric. 4.2 Comparison with Other Methods To comprehensively evaluate the effectiveness of our proposed framework, we conduct extensive comparisons under Protocol 1 (Cross-Dataset) and Protocol 2 (Cross-Manipulation). As reported in Tab.2, our method consistently outperforms all baselines across both protocols, achieving a superior trade-off between retaining past knowledge and adapting to new forgeries. Comparison with Replay-based IFFD Methods. Unlike replay- based approaches such as DFIL [24] and SUR-LID [4], which rely on storing historical samples to mitigate forgetting, our framework achieves better performance in a strict data-replay-free setting. This effectively highlights the efficiency of semantic anchors. By substituting raw data storage with invariant semantic priors, we effectively circumvent the storage constraints and privacy concerns inherent in replay-based paradigms. Comparison with General Replay-free ViT-based Methods. Directly transferring general Incremental Learning (IL) methods, including prompt-based (Coda-Prompt [33]) and adapter-based (CL-LoRA [9]) techniques, to the IFFD task leads to significant degradation. This reveals their inherent limitation: designed for object recognition, they primarily focus on high-level semantic con- tent rather than subtle, low-level forgery artifacts. Consequently, they fail to decouple forensic traces from content, lacking the fine- grained discriminative cues required for this specific task. Fairness Validation. To ensure a rigorous comparison and rule out the influence of backbone capacity, we replace the backbones of DFIL [24] and SUR-LID [4] with the ViT-L/14 used in our method. Even under this aligned configuration, our method still demon- strates clear superiority in performance. These results conclusively verify that our performance gains stem from the proposed methods rather than the raw power of the backbone. It underscores that the core challenge of IFFD lies in effective feature alignment, which cannot be solved solely by scaling up model parameters. 4.3 Ablation Study Overall Ablation. As reported in Tab.3, the ablation study validates the indispensability of each component. The absence of ADH re- sults in significant decision boundary drift. RemovingL ind impairs feature disentanglement, degrading the model’s ability to explicitly encode specific artifacts. Furthermore, omitting APA weakens the cross-modal alignment, depriving the visual encoder of fine-grained semantic guidance. These results demonstrate that optimizing both decision harmonization and semantic-visual interaction is essential for IFFD, enabling the unified framework to mitigate catastrophic forgetting while adapting to novel forgery patterns. Effect of Alignment Method. We evaluate the impact of different classification-head alignment methods within ADH by comparing Linear (LERP), Exponential Moving Average (EMA), and Weighted Mean (WM) against our method. Unlike Euclidean-based methods that distort weight magnitudes, as shown in Tab.5, our method yields superior performance, outperforming baselines that suffer from high-dimensional directional drift (LERP) or delayed adap- tation (EMA). This confirms that strictly constraining the update trajectory to the spherical space preserves semantic angular consis- tency, ensuring geometrically coherent decision boundaries that robustly mitigate catastrophic forgetting across incremental tasks. ICMR ’26, June 16-19, 2026, Amsterdam, NetherlandsHao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang Table 2: Performance comparisons (AUC) with Protocol 1 (Datasets Incremental) and Protocol 2 (Forgery Categories Incremental). Task 1 (T1) to Task 4 (T4) represent current incremented tasks in SDv21, F++, DFDCP, CDF or Hybrid, FR, FS, EFS. The bold denotes the best ones.† indicates results copied from [4].‡ indicates that the backbones of DFIL [24] and SUR-LID [4] are replaced with ViT-L/14 in the experiments for a fair comparison. MethodReplaysTask Protocol 1Protocol 2 SDv21F++DFDCPCDFAvg.HybridFRFSEFSAvg. Methods with CNN Backbone CoReD (M’21) † [15] 500 T10.9998---0.99980.9665---0.9665 T20.74590.9433--0.84460.93550.7988--0.8671 T30.85550.90960.8154-0.86020.89070.79290.8605-0.8480 T4 0.87180.83760.79870.93410.86060.84540.64290.84170.92630.8141 DFIL (M’23) † [24] 500 T10.9998---0.99980.9646---0.9646 T20.74000.9466--0.84330.55740.9975--0.7775 T30.96920.81640.9088-0.89810.60710.66490.9903-0.7541 T4 0.93260.73970.79080.98810.86280.50830.95560.70810.99960.7929 HDP (IJCV’24) † [35] 500 T10.9998---0.99980.9671---0.9671 T20.83730.9507--0.89400.67410.9545--0.8143 T30.93410.85320.8737-0.88700.63000.71350.9509-0.7648 T40.90550.80390.84120.95010.87520.59890.70060.89340.93730.7826 SUR-LID (CVPR’25) † [4] 500 T1 0.9999---0.99990.9685---0.9685 T20.99370.9485--0.97110.82910.9242--0.8766 T30.99860.88440.9161-0.93300.90500.96260.9794-0.9490 T4 0.99710.84790.90670.97440.93150.87900.96790.93560.99070.9433 Methods with Vision Transformer DFIL (ViT-L/14) ‡ 500 T10.9999---0.99990.9682---0.9682 T20.83420.9581--0.89620.58310.9981--0.7906 T30.98970.83720.9192-0.91540.65980.73240.9941-0.7954 T4 0.94130.75230.82410.99230.87750.57480.96270.75970.99970.8242 SUR-LID (ViT-L/14) ‡ 500 T1 0.9999---0.99990.9406---0.9406 T20.99990.9090--0.95450.92350.9902--0.9568 T30.99970.90120.9238-0.94160.90430.98860.9756-0.9562 T4 0.99970.88380.93940.97140.94860.90350.98980.97730.99940.9675 Traditional replay-free ViT-based Incremental Learning Methods Coda-Prompt (CVPR’23) [33]0 T10.9999---0.99990.9631---0.9631 T20.79580.9323--0.86450.61320.8231--0.7181 T30.85640.83150.9026-0.86350.58160.73460.8831-0.7331 T4 0.80490.74920.81310.94130.82710.52450.59380.62640.94610.6727 CL-LoRA (CVPR’25) [9]0 T1 0.9997---0.99970.9531---0.9531 T20.78920.9421--0.86560.62160.9846--0.8031 T30.80150.88960.8991-0.86340.58310.72160.9733-0.7593 T4 0.77420.82560.85490.96260.85430.53670.59660.75910.99880.7228 Our Method AIFIND(Ours)0 T10.9999---0.99990.9719---0.9719 T20.99990.9484--0.97420.96050.9948--0.9776 T3 0.99990.93970.9267-0.95540.94470.99090.9817-0.9724 T40.99880.93270.95510.98790.96860.92860.98970.98920.99990.9769 Effect of APA Injection Layers. We investigate the impact of APA injection depth by evaluating five configurations: Low-level (1–4), Medium-level (11–14), High (21–24), Multi-level (5, 10, 15, 20) and All Layers (0–24). As shown in Tab. 4, the High-level setting yields superior performance, outperforming other configurations. This validates that semantic anchors align most effectively with high-level visual features that share similar semantic granularity, whereas other settings introduce semantic noise that tends to inter- fere with the extraction of high-level visual features. Hyperparameter Analysis. We further investigate the impact of key hyperparameters: the number of selected semantic anchors 푁, the warm-up period푛, the number of APA injection layers푀 and the trade-off coefficients휇 1 and휇 2 . As shown in Tab.6,푁=3 achieves the optimal trade-off between guidance and noise,푛=5 prevents premature alignment with immature features, and푀=4 ensures sufficient interaction with high-level representations, and 휇 1 =0.1, 휇 2 =1 prove optimal for balancing the training objectives. Table 3: Ablation study (AUC) for each proposed component. Bold indicates the best performance. VariantSDv21F++DFDCPCDFAvg. w/o All0.96710.87460.90350.95130.9241 w/o ADH0.98970.91610.92800.95460.9471 w/o APA0.98960.91560.92830.95970.9483 w/o L 푖푛푑 0.98970.91730.93160.97060.9523 Ours0.99880.93270.95510.98790.9686 AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery DetectionICMR ’26, June 16-19, 2026, Amsterdam, Netherlands 024 Severity 50 60 70 80 90 100 AUC (%) Color Saturation 024 Severity 50 60 70 80 90 100 Color Contrast 024 Severity 50 60 70 80 90 100 Block Wise 024 Severity 50 60 70 80 90 100 Gaussian Blur 024 Severity 50 60 70 80 90 100 JPEG Compression 024 Severity 50 60 70 80 90 100 Gaussian Noise DFILHDPSUR-LIDOurs Figure 3: Robustness under unseen perturbations (following Protocol 1, average AUC is used as the evaluation metric). Table 4: Performance comparisons (AUC) of APA injection layers. Bold indicates the best performance. DepthSDv21F++DFDCPCDFAvg. Low0.91080.70030.87510.96950.8639 Medium0.98330.90920.92580.98930.9519 Multi0.94660.81440.90110.96390.9065 All0.95970.90730.92360.97260.9408 High (Ours) 0.9988 0.9327 0.9551 0.9879 0.9686 Effect of APA Gating Strategy. We evaluate the impact of the semantic injection gate within the APA module by comparing fixed scales0.01,0.1,1against a learnable parameter. Unlike static set- tings that impose a rigid injection intensity, as shown in Fig. 4, the learnable strategy yields superior performance, outperforming fixed scalars that suffer from either insufficient guidance (gate = 0.01) or excessive feature perturbation (gate = 1). This confirms that adaptively modulating the injection ratio enables the model to dynamically balance semantic anchor integration with visual preservation, ensuring optimal fine-grained alignment without dis- rupting the underlying feature topology. SDv21F++DFDCPCDF 0.90 0.92 0.94 0.96 0.98 1.00 AUC trainablegate=0.01gate=0.1gate=1 (a) Protocol 1 F++MCNetBlendFaceStyle-GAN3 0.90 0.92 0.94 0.96 0.98 1.00 AUC trainablegate=0.01gate=0.1gate=1 (b) Protocol 1 Figure 4: Ablation study on the gating strategy within the Artifact-Probe Attention (APA) module. Table 5: Performance (AUC) under different classification- head alignment method. Bold indicates the best result. MethodSDv21F++DFDCPCDFAvg. LERP0.99050.93140.94210.97920.9608 EMA0.99020.92860.94030.98070.9599 WM0.99350.92910.94620.97740.9615 Ours0.99880.93270.98790.96860.9686 Table 6: Performance (Avg. AUC on Protocol 1) under differ- ent parameter settings. Bold indicates the best result. ParameterSettings & Performance 푛 05101520 0.95820.96860.96130.96010.9593 푁 123510 0.95880.96210.96860.96140.9522 푀 12345 0.96030.96110.96320.96860.9651 휇 1 0.010.050.10.20.5 0.95920.96210.96860.95760.9531 휇 2 0.10.511.52 0.94980.95720.96860.95810.9486 4.4 Generalization and Robustness Evaluations To ensure a fair comparison, we standardize the backbone for all baseline methods to ViT-L/14, eliminating performance disparities caused by different backbones. Cross-Dataset Generalization. To rigorously evaluate the gener- alization capability of AIFIND against unseen domains, we conduct cross-dataset experiments. The model, fully trained under Protocol 1, is directly evaluated on four unseen benchmarks: DeepFakeDe- tection (DFD) [44], UniFace [41] (from DF40 [43]), SDv15 [26] (from DiffusionFace [3]), and FakeAVCeleb (FAVC) [14]. These datasets en- compass a wide spectrum of generative mechanisms, ranging from conventional face swapping and GAN-based synthesis to diffusion- based text-to-image generation, which imposes significant chal- lenges, requiring the model to overcome substantial domain shifts and rely on intrinsic, transferable forensic features rather than dataset-specific artifacts. Table 7: Cross-dataset generalization results (AUC) on unseen datasets. Bold indicates the best performance. MethodDFDUniFaceSDv15FAVCAvg. DFIL0.82120.65490.82360.71570.7564 HDP0.83420.69770.82310.74950.7761 SUR-LID0.88250.82790.85220.81430.8442 Ours0.93320.89870.88470.87550.8980 ICMR ’26, June 16-19, 2026, Amsterdam, NetherlandsHao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang (b) F++(a) SDv2.1(c) DFDCP(d) CDF OriginalOriginalOriginalOriginal OursOursOursOurs SUR-LIDSUR-LIDSUR-LIDSUR-LID (a) SDv2.1 (b) F++(a) SDv2.1(c) DFDCP(d) CDF OriginalOriginalOriginalOriginal OursOursOursOurs SUR-LIDSUR-LIDSUR-LIDSUR-LID (b) F++ (b) F++(a) SDv2.1(c) DFDCP(d) CDF OriginalOriginalOriginalOriginal OursOursOursOurs SUR-LIDSUR-LIDSUR-LIDSUR-LID (c) DFDCP (b) F++(a) SDv2.1(c) DFDCP(d) CDF OriginalOriginalOriginalOriginal OursOursOursOurs SUR-LIDSUR-LIDSUR-LIDSUR-LID (d) CDF Figure 5: Visualization of Grad-CAM heatmaps across different datasets. (a) Grad-CAM Heatmap 0.00.10.20.30.40.50.60.70.8 Softmax Probability asymmetry in eye shape or size eye color inconsistent with face eyes are blurry unnatural artifacts or seams on the jawline mouth structure inconsistent mouth area is blurry abrupt gradient discontinuity at jawline nose area is blurry nose structure deviates unnatural texture on lips or teeth inconsistent lighting on the nose unnatural texture on the cheeks inconsistent lighting on the mouth seam or splice detected along face edge unnatural texture on the nose 0.790 0.081 0.054 0.025 0.020 0.016 0.010 0.002 0.001 0.001 0.000 0.000 0.000 0.000 0.000 Fake Artifact Probability Distribution Top 3 Artifacts Others (b) Probability Distribution(c) Grad-CAM Heatmap 0.000.050.100.150.200.250.300.35 Softmax Probability mouth area is blurry mouth structure inconsistent seam or splice detected along face edge unnatural artifacts or seams on the jawline inconsistent lighting on the mouth unnatural texture on lips or teeth abrupt gradient discontinuity at jawline asymmetry in eye shape or size eyes are blurry eye color inconsistent with face nose structure deviates nose area is blurry unnatural texture on the cheeks inconsistent lighting on the nose unnatural texture on the nose 0.340 0.330 0.275 0.026 0.012 0.010 0.002 0.001 0.001 0.001 0.000 0.000 0.000 0.000 0.000 Fake Artifact Probability Distribution Top 3 Artifacts Others (d) Probability Distribution Figure 6: Consistency Between Attention Heatmaps and Fake Artifact Probability Distributions. As shown in Tab.7, our framework achieves superior cross- dataset generalization across all unseen benchmarks. Unlike base- lines that may overfit to dataset-specific patterns, AIFIND learns intrinsic and transferable artifact representations, enabling robust detection even in unknown domains. Robustness Evaluations For robustness evaluation, we adopt the rigorous perturbation protocols from [11], which include five severity levels across six different perturbation types. As shown in Fig.3, our model consistently achieves higher AUC scores across all perturbation levels compared to other baselines. These results sug- gest that by aligning visual features with stable semantic anchors, AIFIND preserves discriminative ability under significant image distortions, ensuring reliable performance in practical scenarios. 4.5 Visualizations We employ Grad-CAM [30] to visualize the attention maps of the model trained under Protocol 1. As illustrated in Fig. 5, the baseline SUR-LID [4] exhibits relatively dispersed attention, often appearing distracted by the background or irrelevant facial areas. In contrast, our method tends to concentrate more on semantically sensitive facial components, such as the eyes and mouth, which are notori- ously prone to manipulation artifacts. This improved focus suggests that the semantic anchors effectively guide the model to attend to critical forensic regions rather than low-level noise, validating that our linguistic supervision successfully directs visual attention to physically meaningful areas. Furthermore, the results in Fig. 6 demonstrate a notable **spatial- semantic consistency** between the attention heatmaps and the inference outcomes. Specifically, the regions receiving high visual activation generally align with the artifact categories that yield elevated predicted probabilities. For instance, when the heatmap highlights the eye region, the probability score for eye-related ar- tifact classes rises distinctively. These observations indicate that our model learns to associate discriminative forensic traces with their corresponding semantic priors, thereby mitigating the risk of overfitting to spurious cues and enhancing the interpretability and reliability of the decision-making process. 5 Conclusion In this paper, we propose AIFIND, an Artifact-Aware Interpreting Fine-Grained Alignment framework. Unlike traditional methods that rely on sample replay, AIFIND leverages a semantic anchor library to guide the model in learning invariant forgery representa- tions. Our method ensures that visual features are aligned with sta- ble semantic anchors, effectively mitigating catastrophic forgetting. Extensive experiments demonstrate that our method achieves state- of-the-art performance and exhibits spatial-semantic consistency in visualization. In the future, we plan to extend our framework to broader multimodal scenarios, exploring more adaptive semantic anchors via Large Vision-Language Models to tackle increasingly diverse forgery patterns in open-world settings. Acknowledgments This work is partially supported by the National Natural Science Foundation of China under Grants 62441232, 62476068, 62306092, 62502115, and projects ZR2025ZD01, ZR2024QF066, ZR2025QC1516 supported by Shandong Provincial Natural Science Foundation, and projects 2024DXZD0004 supported by Inner Mongolia Department of Science and Technology. AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery DetectionICMR ’26, June 16-19, 2026, Amsterdam, Netherlands References [1]Hunar Batra and Ronald Clark. 2024. Evcl: Elastic variational continual learning with weight consolidation. arXiv preprint arXiv:2406.15972 (2024). [2]Pietro Buzzega, Matteo Boschini, Angelo Porrello, and Simone Calderara. 2021. Rethinking experience replay: a bag of tricks for continual learning. In 2020 25th International Conference on Pattern Recognition. 2180–2187. [3] Zhongxi Chen, Ke Sun, Ziyin Zhou, Xianming Lin, Xiaoshuai Sun, Liujuan Cao, and Rongrong Ji. 2024. Diffusionface: Towards a comprehensive dataset for diffusion-based face forgery analysis. arXiv:2403.18471 (2024). [4]Jikang Cheng, Zhiyuan Yan, Ying Zhang, Li Hao, Jiaxin Ai, Qin Zou, Chen Li, and Zhongyuan Wang. 2025. Stacking brick by brick: Aligned feature isolation for incremental face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13927–13936. [5]Xinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou, and Junyu Dong. 2025. Forensics adapter: Adapting clip for generalizable face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19207– 19217. [6]Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, and Cristian Canton Ferrer. 2019. The Deepfake Detection Challenge (DFDC) Preview Dataset. [7]Hui Guo, Shu Hu, Xin Wang, Ming-Ching Chang, and Siwei Lyu. 2022. Eyes tell all: Irregular pupil shapes reveal gan-generated faces. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing. 2904–2908. [8]Zonghui Guo, Yingjie Liu, Jie Zhang, Haiyong Zheng, and Shiguang Shan. 2025. Face Forgery Video Detection via Temporal Forgery Cue Unraveling. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7396–7405. [9] Jiangpeng He, Zhihao Duan, and Fengqing Zhu. 2025. CL-LoRA: Continual Low- Rank Adaptation for Rehearsal-Free Class-Incremental Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 30534– 30544. [10]Fa-Ting Hong and Dan Xu. 2023. Implicit identity representation conditioned memory compensation network for talking head video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision. [11] Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detec- tion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2889–2898. [12]Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks. Ad- vances in Neural Information Processing Systems 34 (2021), 852–863. [13] Hossein Kashiani, Niloufar Alipour Talemi, and Fatemeh Afghah. 2025. Freqde- bias: Towards generalizable deepfake detection via consistency-driven frequency debiasing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8775–8785. [14] Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. 2021. FakeAVCeleb: A novel audio-video multimodal deepfake dataset. arXiv preprint arXiv:2108.05080 (2021). [15] Minha Kim and Shahroz Tariq. 2021. Cored: Generalizing fake media detection with continual representation using distillation. In Proceedings of the 29th ACM International Conference on Multimedia. 337–346. [16]Youngeun Kim, Yuhang Li, and Priyadarshini Panda. 2024. One-stage prompt- based continual learning. In European Conference on Computer Vision. 163–179. [17]Jiashuo Li, Shaokun Wang, Bo Qian, Yuhang He, Xing Wei, Qiang Wang, and Yihong Gong. 2025. Dynamic integration of task-specific adapters for class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 30545–30555. [18] Yuezun Li, Ming-Ching Chang, and Siwei Lyu. 2018. In Ictu Oculi: Exposing AI Generated Fake Face Videos by Detecting Eye Blinking. In IEEE International Workshop on Information Forensics and Security. [19]Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [20] Kaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao, and Bin Li. 2025. Standing on the shoulders of giants: Reprogramming visual-language model for general deepfake detection. In Proceedings of the AAAI Conference on Artificial Intelligence. 5262–5270. [21]Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, Wan- Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. 2019. MediaPipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172 (2019). [22]Sriram Mandalika, Harsha Vardhan, and Athira Nambiar. 2025. Replay to Remem- ber (R2R): An Efficient Uncertainty-driven Unsupervised Continual Learning Framework Using Generative Replay. arXiv preprint arXiv:2505.04787 (2025). [23] Seyed-Mohsen Moosavi-Dezfooli and Alhussein Fawzi. 2017. Universal adversar- ial perturbations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1765–1773. [24]Kun Pan, Yifang Yin, Yao Wei, Feng Lin, Zhongjie Ba, Zhenguang Liu, Zhibo Wang, Lorenzo Cavallaro, and Kui Ren. 2023. Dfil: Deepfake incremental learning by exploiting domain-invariant forgery clues. In Proceedings of the 31st ACM International Conference on Multimedia. 8035–8046. [25]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PmLR, 8748–8763. [26] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10684–10695. [27] Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 1–11. [28]Anurag Roy, Riddhiman Moulick, Vinay K Verma, Saptarshi Ghosh, and Abir Das. 2024. Convolutional prompting meets language models for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23616–23626. [29]Krisanu Sarkar. 2025. Adaptive Variance-Penalized Continual Learning with Fisher Regularization. arXiv preprint arXiv:2508.16632 (2025). [30]Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, and Devi Parikh. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision. 618–626. [31]Kaede Shiohara and Toshihiko Yamasaki. 2022. Detecting deepfakes with self- blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18720–18729. [32] Kaede Shiohara, Xingchao Yang, and Takafumi Taketomi. 2023. Blendface: Re- designing identity encoders for face-swapping. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7634–7644. [33]James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta, Paola Cascante-Bonilla, Donghyun Kim, Assaf Arbelle, Rameswar Panda, Rogerio Feris, and Zsolt Kira. 2023. Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11909–11919. [34]James Seale Smith, Lazar Valkov, Shaunak Halbe, Vyshnavi Gutta, Rogerio Feris, Zsolt Kira, and Leonid Karlinsky. 2024. Adaptive memory replay for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3605–3615. [35]Ke Sun, Shen Chen, Taiping Yao, Xiaoshuai Sun, Shouhong Ding, and Rongrong Ji. 2025. Continual face forgery detection via historical distribution preserving. International Journal of Computer Vision 133, 3 (2025), 1067–1084. [36]Ke Sun, Shen Chen, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia- Wen Lin, and Rongrong Ji. 2025. Towards general visual-linguistic face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19576–19586. [37] Jiahe Tian, Cai Yu, Xi Wang, Peng Chen, Zihao Xiao, Jizhong Han, and Yesheng Chai. 2024. Dynamic mixed-prototype model for incremental deepfake detection. In Proceedings of the 32nd ACM International Conference on Multimedia. 8129– 8138. [38]Qiang Wang, Xiang Song, Yuhang He, Jizhou Han, Chenhao Ding, Xinyuan Gao, and Yihong Gong. 2025. Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4839–4849. [39]Zhicheng Wang, Yufang Liu, Tao Ji, Xiaoling Wang, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Ling Wang, Xu Shao, et al.2023. Rehearsal-free continual language learning via efficient parameter isolation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10933–10946. [40]Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister. 2022. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [41]Chao Xu, Jiangning Zhang, Yue Han, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang, and Yong Liu. 2022. Designing one unified frame- work for high-fidelity face reenactment and swapping. In European Conference on Computer Vision. Springer, 54–71. [42] Zhiyuan Yan, Yuhao Luo, Siwei Lyu, Qingshan Liu, and Baoyuan Wu. 2024. Transcending forgery specificity with latent space augmentation for generalizable deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8984–8994. [43]Zhiyuan Yan, Taiping Yao, Shen Chen, Yandan Zhao, Xinghe Fu, Junwei Zhu, Donghao Luo, Chengjie Wang, Shouhong Ding, Yunsheng Wu, et al.2024. Df40: Toward next-generation deepfake detection. Advances in Neural Information Processing Systems 37 (2024), 29387–29434. ICMR ’26, June 16-19, 2026, Amsterdam, NetherlandsHao Wang, Beichen Zhang ∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang [44]Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. 2023. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. In Ad- vances in Neural Information Processing Systems, Vol. 36. 4534–4565. [45] Zhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo, Xinghe Fu, Taiping Yao, Shouhong Ding, Yunsheng Wu, and Li Yuan. 2025. Generalizing deepfake video detection with plug-and-play: Video-level blending and spatiotemporal adapter tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12615–12625. [46] Xin Yang, Yuezun Li, and Siwei Lyu. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing. 8261–8265. [47]Andrii Yermakov, Jan Cech, and Jiri Matas. 2025. Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection. arXiv (2025). [48] Jiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu, Dong Wang, Huchuan Lu, and You He. 2024. Boosting continual learning of vision-language models via mixture-of- experts adapters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23219–23230. [49]Jiaran Zhou, Yuezun Li, Baoyuan Wu, Bin Li, Junyu Dong, et al.2024. Freqblender: Enhancing deepfake detection by blending frequency knowledge. Advances in Neural Information Processing Systems 37 (2024), 44965–44988.