Paper deep dive
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
Mengnan Zhao, Lihe Zhang, Tianhang Zheng, Bo Wang, Baocai Yin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 8:19:28 AM
Summary
The paper investigates the mechanism of Catastrophic Overfitting (CO) in Fast Adversarial Training (FAT). The authors propose that CO can be interpreted as a 'trigger overfitting' phenomenon, similar to backdoor attacks and unlearnable tasks. They demonstrate that CO-affected models develop a specialized 'adversarial pathway' that acts as a weak, class-discriminative trigger. To mitigate CO, the paper introduces two strategies: recalibrating model parameters via fine-tuning/linear probing and implementing a weight outlier suppression constraint to prevent the formation of abnormal adversarial pathways.
Entities (8)
Relation Signals (4)
Fast Adversarial Training â isproneto â Catastrophic Overfitting
confidence 100% ¡ FAT is prone to catastrophic overfitting (CO)
Catastrophic Overfitting â isinterpretedas â Backdoor Attack
confidence 95% ¡ we innovatively interpret CO through the lens of backdoor.
Weight Outlier Suppression Constraint â mitigates â Catastrophic Overfitting
confidence 95% ¡ Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies.
Catastrophic Overfitting â relatedto â Unlearnable Tasks
confidence 90% ¡ unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone to catastrophic overfitting (CO), wherein models overfit to the specific attack used during training and fail to generalize to others. While existing methods introduce diverse hypotheses and propose various strategies to mitigate CO, a systematic and intuitive explanation of CO remains absent. In this work, we innovatively interpret CO through the lens of backdoor. Through validations on pathway division, diverse feature predictions, and universal class distinguishable triggers in CO, we conceptualize CO as a weak trigger variant of unlearnable tasks, unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework. Guided by this, we leverage several backdoor inspired strategies to mitigate CO: (i) Recalibrate CO affected model parameters using vanilla fine tuning, linear probing, or reinitialization-based techniques; (ii) Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies.
Tags
Links
- Source: https://arxiv.org/abs/2604.24350v1
- Canonical: https://arxiv.org/abs/2604.24350v1
Trouble viewing inline? Open PDF directly â
Full Text
64,700 characters extracted from source content.
Expand or collapse full text
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training Mengnan Zhao, Lihe Zhang, Tianhang Zheng, Bo Wang, Baocai Yin Manuscript received Sep, 2025.This work was supported by the National Natural Science Foundation of China under Grant 62431004 and 62276046.Mengnan Zhao is with the School of Computer Science and Technology, Anhui University, Hefei 230601, China. E-mail: gaoshanxingzhi@163.com.Lihe Zhang and Bo Wang are with the School of Information and Communication Engineering, Dalian University of Technology (DUT), Dalian 116024, China. E-mail: zhanglihe@dlut.edu.cn, bowang@dlut.edu.cn.Tianhang Zheng is with the School of Computer Science and Technology, Zhejiang University, Hangzhou 310058, China. E-mail: zthzheng@zju.edu.cn.Baocai Yin is with the School of Computer Science and Technology, DUT, Dalian 116024, China. E-mail: ybc@dlut.edu.cn. Abstract Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone to catastrophic overfitting (CO), wherein models overfit to the specific attack used during training and fail to generalize to others. While existing methods introduce diverse hypotheses and propose various strategies to mitigate CO, a systematic and intuitive explanation of CO remains absent. In this work, we innovatively interpret CO through the lens of backdoor. Through validations on pathway division, diverse feature predictions, and universal class-distinguishable triggers in CO, we conceptualize CO as a weak-trigger variant of unlearnable tasks, unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework. Guided by this, we leverage several backdoor-inspired strategies to mitigate CO: (i) Recalibrate CO-affected model parameters using vanilla fine-tuning, linear probing, or reinitialization-based techniques; (i) Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies. Index Terms: Fast adversarial training, catastrophic overfitting, backdoor, backdoor-inspired mitigation strategies I Introduction Recent advancements in deep learning have driven significant progress across a wide range of applications [106, 60, 64, 3, 71]. However, these developments have also exposed critical limitations of neural networks [12, 81, 8], particularly their susceptibility to adversarial attacks [4, 24, 105]. To address such vulnerabilities, adversarial training has emerged as a widely adopted defense strategy [87, 6, 101, 95, 19, 75], wherein perturbed examples are incorporated during training to enhance model robustness [59, 37, 83]. Initial adversarial training techniques [84, 39, 26] produce training data using multi-step adversarial attacks, such as the projected gradient descent (PGD) [57]. In recent years, fast adversarial training (FAT) methods [62, 42], particularly single-step approaches like FGSM-RS [82], offer notable computational efficiency by generating adversarial examples with fewer backward passes [50, 32, 63, 89]. However, FAT remains susceptible to catastrophic overfitting (CO), where models overfit to the specific attack used during training and fail to generalize to other adversarial attacks. To address this, prior works primarily target surface-level symptomsâsuch as limited adversarial diversity [82, 31], disparities in sample convergence rates [102], and inconsistencies in gradient attribution importance [22]âto design corresponding mitigation strategies [1, 85]. While these methods improve FAT stability, they offer limited insight into the underlying mechanisms of CO. Beyond these, other studies have associated CO with deeper factors such as self-fitting behavior [29], gradient misalignment [2], and feature overriding [103]. Collectively, these works suggest that CO stems from abnormal pathway division in the learned model representations. However, a systematic and intuitive explanation of CO remains absent, such as the mechanisms underlying pathway division and the nature of self-information. This work innovatively analyzes CO through the lens of backdoor [99, 88, 52, 93]. We begin by validating the pathway division and diverse path predictions. These analyses reveal an intriguing phenomenon: CO exhibits a high similarity to backdoor-related tasks. We then show that adversarial perturbations from CO-affected models encode universal, class-discriminative triggers, and that the primary differences between CO and backdoor-related tasks can be attributed to the trigger strength variance. Together, these findings support a unified interpretationâtrigger overfittingâin which both CO and backdoor-related tasks arise from a modelâs over-reliance on trigger-like features transmitted through specialized paths. We further introduce backdoor-inspired strategies to mitigate CO. (i) We adapt established backdoor fine-tuning techniquesâincluding vanilla fine-tuning, linear probing, and reinitialization-based methodsâto recalibrate the parameters of CO-affected models. These approaches steer the model away from overfitting by shifting its focus back to informative data features, rather than to trigger-related patterns. Experimental results show that fine-tuning CO-affected models on clean data can temporarily alleviate CO, though the benefit is short-lived as CO tends to reoccur. (i) Inspired by weight poisoning in backdoor attacks, we propose a weight outlier suppression constraint, which penalizes weight deviations from the layer-wise mean weight. This prevents the formation of adversarial pathways, leading to a stable FAT process. Additionally, we provide a discussion section that (i) clarifies the essence of robustness improvement in adversarial training, (i) analyzes the underlying cause of CO, and (i) investigates whether techniques designed to mitigate CO can generalize to backdoor-related tasks. In summary, our contributions are threefold: 1) We interpret CO through the lens of backdoor, unifying CO, backdoor attacks, and unlearnable tasks under a trigger overfitting framework. 2) We validate key similaritiesâpathway division, diverse predictions, and universal triggersâbetween CO and backdoor. 3) We introduce backdoor-inspired mitigation strategies, including adapted fine-tuning and a weight outlier suppression constraint, demonstrating their effectiveness empirically. The remainder of the paper is organized as follows. Section I reviews recent advances in adversarial training and backdoor-related tasks. In Section I, we establish the connection between CO and backdoor-related tasks. Section IV presents a set of backdoor-inspired strategies for mitigating CO. Section V offers further discussion and analysis, such as insights into the causes of CO and the robustness of FAT. Finally, Section VI concludes the paper and outlines several promising directions for future research. I Related work I-A Adversarial training Recent breakthroughs in deep neural networks [106, 60, 64] have prompted extensive research into their security risks [12, 81, 8], with particular attention to their vulnerability to adversarial attacks [47, 17, 4, 24, 105]. In response, adversarial training [59, 37, 83, 44, 79] has emerged as a popular strategy to enhance model robustness, employing both multi-step (e.g., PGD [57]) and one-step adversarial attacks (e.g., FGSM [23]) [84, 39, 26, 50, 100, 38]. (a) Distance confusion matrix (b) UMAP distribution of δsign _sign in the stably trained model (c) UMAP distribution of δsign _sign in the CO-affected model Figure 1: Distance matrix and UMAP visualizations under FGSM-RS. Small distances imply reduced separation between class distributions. Compared to PGD-based methods, FGSM-based methods such as FGSM-RS [82] and FGSM-MEP [35], also called FAT [77, 33], are computationally efficient [34]. Given the initial perturbation δ0ââ(0,) _0 (0,I), training dataset D, network fâ(â ;θ)f(¡;θ), loss âL, step size Ͼξ, and budget Ξ, FGSM-RS generates adversarial perturbations by Eqs. (1) and (2), δ=clipΞâ(δ0+Ďľâ δsign),δ=clip_Ξ ( _0+ξ¡ _sign ), (1) δsign=signâ(âδ0ââ(fâ(x+δ0;θ),y)), _sign=sign ( _ _0L (f(x+ _0;θ),y ) ), (2) and implements adversarial training by Eq. (3), minθâĄxâźâââ(fâ(x+δ;θ),y). _θE_x L (f (x+δ;θ ),y ). (3) In contrast to FGSM-RS, FGSM-MEP constructs δ0 _0 by leveraging the momentum accumulated from adversarial perturbations computed over preceding epochs. It also incorporates a prediction regularization during minimization, expressed as âpred=βââfâ(x+δ)âfâ(x+δ0)â22.R_pred=β\|f(x+δ)-f(x+ _0)\|_2^2. (4) However, FAT approaches may suffer from CO [14, 68, 36]. To address this issue, various strategies have been proposed [96], such as gradient alignment [2], convergence smoothness [102], zero-gradient clipping [22], bi-level optimization [80] and feature activation consistency [103]. Recently, He et al. [29] assume that adversarial perturbations embed self-information and argues that models acquire this information through a separate pathway. Then, Zhao et al. [103] introduce a strong regularization term that enforces consistency between predictions on clean and adversarial examples, thereby suppressing the emergence of the adversarial pathway. Lin et al. [70] leverage both weight-level and example-level adversarial perturbations to enhance training stability by enforcing weight robustness. In contrast, this work explains CO through the lens of backdoor, unifying CO, backdoor attacks, and unlearnable tasks under a common interpretation of trigger overfitting. Within this perspective, we introduce several backdoor-inspired mitigation strategies that not only enable models to break free from CO but also suppress its occurrence. Particularly, unlike existing methods that suppress CO by adding regularization constraints to the prediction or feature space, this work adjusts the distribution of model weights. Furthermore, while both this work and that of Lin et al. [70] identify weight anomalies, their objective is to construct robust weights, whereas ours is specifically to suppress weight outliers. We observe that enforcing overall weight robustness will limit the modelâs ability to fit clean samples. Moreover, although both our weight outlier suppression constraint and the strong regularization proposed by [103] are effective at preventing adversarial pathways, the latter is often compromised by misclassifications on clean examples. Additionally, it does not explain why adversarial paths emerge and override data paths. This work addresses this question from the perspective of trigger overfitting. I-B Backdoor-related tasks This section covers both backdoor attacks and unlearnable tasks, the latter representing a transferable application of backdoor attacks. Backdoor attacks pose a serious threat to the security of deep neural networks. Early works such as BadNets [25] and TrojanNN [56] demonstrate that inserting poisoned samples with static triggers during training can cause targeted misclassification. However, such attacks are often detectable due to their reliance on fixed and conspicuous patterns. To improve stealth and generalizability, a range of trigger designs have been proposed [51, 97, 9]. For instance, spatial transformations [61], image blending [10], and frequency-domain perturbations [78, 97] aim to create imperceptible or input-adaptive triggers. Others have explored learnable or sample-specific backdoor strategies [51, 15] to improve trigger effectiveness and bypass detection. More recently, attention has turned to contrastive and self-supervised learning frameworks. CTRL [48] shows that even models trained without labels are susceptible to backdoor insertion. Beyond implanting malicious behaviors into trained models, backdoor techniques have been adapted to unlearnable tasks, which aim to impair a modelâs generalization on clean data. Huang et al. [30] introduce sample-wise and class-wise perturbationsâsimilar in nature to triggersâinto all training samples, inducing overfitting and preventing the learning of useful representations. However, their effectiveness diminishes under different training settings or datasets. To improve transferability, Ren et al. [67] propose a Classwise Separability Discriminant strategy that enhances linear separability. Zhang et al. [98] generate label-agnostic unlearnable examples via cluster-wise perturbations. Liu et al. [55] extend protection to multimodal contrastive learning. Notably, Qin et al. [65] show that adversarial augmentations can mitigate unlearnable-example attacks, motivating Fu et al. [20] to design robust unlearnable examples against adversarial learning. Furthermore, Ye et al. [90] present ungeneralizable samples that can only be learned by specific networks, and Li et al. [49] develop methods to detect and corrupt convolution-based unlearnable examples. This work considers both standard backdoor attacks and unlearnable tasks, and investigates their connections to CO. TABLE I: Description of symbols. Symbols Description δ;δ0;δsign;δmomδ; _0; _sign; _mom Adversarial perturbation; Initial perturbation; Perturbation direction; Momentum-based perturbation θ;θadv;θdata;θCOθ; _adv; _data; _CO Model parameters; Adversarial path parameters; Data path parameters; CO-affected model parameters Ξ;ΞT;ΞEΞ; _T; _E Maximum perturbation budget; Maximum perturbation budget during training; Maximum perturbation budget during evaluation Ρ;Îą;ϾΡ;Îą;Îľ Hyperparameters; Step size I CO and backdoor-related tasks In this section, we first present the motivation for interpreting CO through the lens of backdoor. We then provide a comparative analysis highlighting the similarities and differences between CO and backdoor-related tasks. Table I summarizes the notation used in this paper for quick reference. I-A Motivation of explaining CO from backdoor Prior FAT work [103] has proposed two key hypotheses regarding CO: (i) CO-affected models can be decomposed into distinct adversarial and data branches; (i) the adversarial branch exhibits feature overriding for adversarial inputs. In this paper, we systematically explore these hypotheses by analyzing the behaviors of FGSM-RS and FGSM-MEP. (a) Distance confusion matrix (b) UMAP distribution of δsign _sign in the stably trained model (c) UMAP distribution of δsign _sign in the CO-affected model Figure 2: Distance matrix and UMAP visualizations under FGSM-MEP. Small distances imply reduced separation between class distributions. Pathway division: In the absence of the adversarial pathway, the primary distribution of adversarial perturbations δsign _sign should be class-separable. To verify this, we compute the inter-class Wasserstein distance based on δsign _sign. The experiments are conducted on CIFAR-10 [43] using a ResNet18 backbone [27], with the maximum perturbation budget of 16/255 and the learning rate of 0.1. The results in Figure 1(a) and Figure 2(a) reveal that the stable model (epoch 5) maintains distinct inter-class distances, while the CO-affected model (epoch 20) exhibits distribution overlap (distance 0). These observations are further supported by Uniform Manifold Approximation and Projection (UMAP) distribution [58] in Figures 1 and 2, showing class-separable patterns in stable models versus complete overlap under CO-affected models.111UMAP is a dimensionality reduction technique used for visualizing primary and high-dimensional features. Namely, the CO-affected model should contain an additional adversarial branch. We formally decompose the network parameters as θ=θadv,θdataθ=\ _adv, _data\, corresponding to the adversarial and data pathways, respectively. The lack of feature discrimination in δsign _sign suggests that the adversarial pathway dominates gradient backpropagation, expressed as âδ0ââ(fâ(x+δ0;θ),y)ââδ0ââ(fâ(x+δ0;θadv),y). _ _0L(f(x+ _0;θ),y)â _ _0L(f(x+ _0; _adv),y). (5) TABLE I: Comparison between backdoor-related tasks and CO. Backdoor Attack Catastrophic Overfitting (CO) Default Unlearnable 1 High accuracy for triggered inputs xtriggerx_trigger High accuracy for perturbed inputs xperturbedx_perturbed High accuracy for adversarial examples x+FGSMâ(x)x+FGSM(x) (classified as the predefined class) (classified as the truth class) (classified as the truth class) 2 High accuracy for x Low accuracy for x FGSM-RS: Moderate accuracy for x From low to high as the perturbation size increases FGSM-MEP: Higher accuracy than FGSM-RS for x 3 Low accuracy for attacks Low accuracy for attacks Low accuracy for attacks except FGSM Diverse forward predictions instead of feature overriding. To analyze the forward prediction behavior of FAT models, we visualize the training dynamics of diverse methods on CIFAR-10 with ResNet-18 in Figure 3. After CO, we observe: xâźtrainâACCâ(x+δ) _x _trainACC(x+δ) (6) âĽxâźtrainâACCâ(x+δ0)âxâźtrainâACCâ(x)âŤ0. _x _trainACC(x+ _0) _x _trainACC(x) 0. 1) Eq. (6) implies that CO-affected models exhibit higher or comparable classification accuracy on adversarial examples than on clean x and initially perturbed samples x+δ0x+ _0. 2) Meanwhile, CO-affected models remain classifiable to x and x+δ0x+ _0. These observations indicate that, during forward inference, the CO-affected model may integrate features from both data and adversarial pathways, minθâĄââ(fâ(x+δ;θ),y)â _θL(f(x+δ;θ),y)â (7) minθadv,θdataâĄ[ââ(fâ(x+δ;θadv),y)+ââ(fâ(x+δ;θdata),y)], _ _adv, _data [L(f(x+δ; _adv),y)+L(f(x+δ; _data),y) ], or directly utilize the data pathway, minθâĄââ(fâ(x+δ;θ),y)âminθdataâĄââ(fâ(x+δ;θdata),y). _θL(f(x+δ;θ),y)â _ _dataL(f(x+δ; _data),y). (8) (a) FGSM-RS (CO in 13th Epoch) (b) FGSM-MEP (CO in 12th Epoch) Figure 3: Training dynamics of existing FAT methods. TABLE I: Comparison between various FAT methods. Methods marked with â and ⥠utilize Eq. (10), with the parameter Îą set to 1e-2 and 1, respectively. âBestâ and âFinalâ refer to the evaluation results of the model with the best PGD-10 performance and the final model, respectively. âCleanâ denotes the classification accuracy for clean examples. âPerturbedâ means that we employ initially perturbed samples. FGSM, PGD, C&W [5] and A [13] are adversarial attacks for evaluating robustness. âCOâ signifies whether the trained model falls into catastrophic overfitting during FAT. Net&Dataset Methods Cleanâ Perturbedâ FGSMâ PGD10â PGD20â PGD50â C&Wâ Aâ CO FGSM-RS [82] best 54.69 54.64 34.31 26.01 19.77 17.85 18.41 12.63 â final 80.80 83.16 76.62 0.00 0.00 0.00 0.00 0.00 ResNet18 [27] FGSM-RSâ [82] best 73.55 73.05 44.00 34.20 23.99 20.78 22.74 14.63 â final 73.34 73.05 44.15 33.99 24.04 21.00 22.61 14.76 CIFAR10 [43] FGSM-MEP[35] best 30.83 30.58 25.69 23.09 21.67 21.49 20.35 18.78 â final 88.07 88.34 80.03 11.40 6.75 3.74 3.48 0.08 FGSM-MEPâĄ[35] best 59.73 59.05 42.27 37.39 32.43 31.60 25.97 22.22 â final 59.69 59.09 42.02 37.25 32.20 30.99 25.67 22.19 FGSM-RS [82] best 56.38 56.22 33.95 28.21 21.71 19.66 18.81 13.49 â final 81.22 80.56 80.33 0.00 0.00 0.00 0.00 0.00 PreActResNest18 [28] FGSM-RSâ [82] best 72.71 72.21 45.45 34.94 25.11 22.38 22.21 14.78 â final 72.68 72.49 44.78 34.76 25.01 21.90 22.22 14.82 CIFAR10 [43] FGSM-MEP [35] best 41.81 41.78 29.19 26.87 24.26 23.89 18.42 17.08 â final 88.66 89.43 82.76 11.17 6.82 4.09 2.84 0.01 FGSM-MEP⥠[35] best 59.29 58.28 42.07 36.69 31.94 30.74 25.50 21.40 â final 59.61 58.90 41.85 36.29 31.14 30.07 25.54 21.26 FGSM-RS [82] best 30.40 30.28 14.78 11.74 9.18 8.82 7.97 6.12 â final 39.40 52.36 47.41 0.00 0.00 0.00 0.00 0.00 ResNet18 [27] FGSM-RSâ [82] best 45.12 45.33 20.80 16.21 11.78 11.02 10.61 7.95 â final 48.70 48.71 22.58 16.57 11.70 10.82 10.86 8.16 CIFAR100 [43] FGSM-MEP [35] best 22.18 22.17 14.51 12.57 10.95 10.73 8.34 7.16 â final 67.40 67.73 54.74 1.33 0.60 0.27 0.31 1.00 FGSM-MEP⥠[35] best 42.51 42.25 25.15 21.15 17.28 16.71 14.01 11.20 â final 42.69 42.70 25.20 20.76 17.32 16.59 14.12 11.39 FGSM-RS [82] best 27.94 27.72 13.92 11.40 9.14 8.83 7.61 6.16 â final 50.30 51.80 50.00 0.00 0.00 0.00 0.00 0.00 PreActResNest18 [28] FGSM-RSâ [82] best 46.79 46.88 21.83 16.33 11.80 10.93 11.21 7.83 â final 47.16 47.04 21.94 16.14 11.62 10.53 10.76 7.69 CIFAR100 [43] FGSM-MEP [35] best 21.21 21.02 13.61 11.79 10.22 10.01 7.72 6.75 â final 67.59 69.12 55.06 2.31 1.20 0.64 0.71 0.00 FGSM-MEP⥠[35] best 43.22 42.96 25.91 20.84 16.77 16.04 14.07 11.18 â final 43.22 42.95 25.99 20.89 16.88 16.00 14.01 11.20 The validation of pathway decomposition and diverse forward prediction raises a natural question: Is there an intrinsic connection between CO and backdoor attacks? Table I shows comparisons between CO and backdoor-related tasks. Specifically, we hypothesize the existence of a backdoor pathway in the backdoored model. When this pathway is activated, typically by a trigger, feature overriding occurs, leading the model to confidently classify the input into a predefined target class. In the absence of triggers, the model backs to standard prediction behavior and maintains high accuracy on clean samples. Meanwhile, unlearnable tasks [94, 16, 20] embed perturbations, functionally similar to backdoor triggers, into all samples while retaining their original labels, causing the model to rapidly overfit to these perturbations (i.e. high accuracy for perturbed inputs) and lose generalization capability (i.e., low accuracy for clean examples). Notably, both types remain vulnerable to adversarial attacks, as their training data are not augmented with adversarial examples. A similar pattern is observed in CO-affected models, which overfit to the specific attack used during training (e.g., FGSM), achieving high accuracy on seen adversarial types while failing to generalize to unseen ones. Taken together, these findings suggest a strong behavioral similarity between CO and backdoor-related tasks. I-B Further analyses between backdoor and CO We further investigate whether FGSM-generated perturbations FGSMâ(x)FGSM(x) during CO exhibit characteristics similar to backdoor triggers. Specifically, we hypothesize that δsign _sign contains a universal class-discriminative component, referred to as the UCD trigger δucd _ucd. To test this hypothesis, we compute the class-wise expectation of δsign _sign as defined in Eq. (9), and incorporate an auxiliary constraint âauxL_aux into the adversarial training objective as specified in Eq. (10). δmomtâ0.9âδmomt+0.1âxââŹ,y=tδsignt,â, _mom^tâ 0.9\, _mom^t+0.1\, E_x ,y=t _sign^t,*, (9) âaux=âÎąââδsigntâδmomtâ2,L_aux=-Îą\| _sign^t- _mom^t\|_2, (10) where âŹB represents the current batch, t denotes the designated label, and Îą is set to 1e-2 by default. The momentum-based term δmomt _mom^t aims to dynamically extract UCD triggers. âauxL_aux penalizes class-consistent perturbation features, forcing the model to forget UCD triggers. Gradients are prevented from propagating into δmomt _mom^t by detaching δsignt _sign^t, denoted as δsignt,â _sign^t,* in Eq. (9). As Table I shows, âauxL_aux mitigates CO in FGSM-RS and FGSM-MEP across different networks and datasets, confirming the existence of UCD triggers. Figure 4: Backdoor fine-tuning techniques. FT: finetuning. We next analyze the performance of CO-affected and backdoored models on clean inputs. The clean classification accuracies across different paradigms are summarized as follows: ACCOri _Ori âACCSBackdoor>ACCMEPâŤACCUnlearnable, _SBackdoor>ACC_MEP _Unlearnable, (11) where ACCSBackdoorACC_SBackdoor, ACCOriACC_Ori, ACCMEPACC_MEP, and ACCUnlearnableACC_Unlearnable denote the clean-sample accuracies of models under the standard backdoor, original, CO-affected (trained with FGSM-MEP), and unlearnable settings, respectively. Notably, ACCMEPACC_MEP falls between ACCSBackdoorACC_SBackdoor and ACCUnlearnableACC_Unlearnable, suggesting that CO-affected models preserve a moderate level of clean-data generalization. In backdoor attacks, the model associates a fixed trigger with a target class while largely preserving its performance on clean data. In contrast, unlearnable tasks embed diverse perturbations to training data without altering labels, causing the model to overfit to these perturbations and lose the ability to accurately classify clean examples. Notably, in unlearnable tasks, we observe that clean accuracy remains comparable to ACCOriACC_Ori when the embedded perturbation is weak. As perturbation strength increases, the modelâs performance on clean data progressively degrades and collapses once a critical threshold is reached. This trend indicates that the higher clean accuracy observed in CO-affected models (e.g., ACCMEPACC_MEP) compared to unlearnable tasks (ACCUnlearnableACC_Unlearnable) can be attributed to the weaker trigger effect. This is further supported by empirical evidence: as illustrated in Figure 2, adversarial perturbations in CO-affected models trained with FGSM-MEP are primarily composed of non-discriminative features, with UCD triggers contributing only marginally. By verifying the existence of UCD triggers in adversarial perturbations generated for CO-affected models, and by analyzing the differences in clean accuracy across techniques, we argue that CO, standard backdoor attacks, and unlearnable tasks can be understood within a unified perspective of trigger overfitting. More specifically, CO can be viewed as a weaker variant of unlearnable tasks; both can be regarded as transfer applications of backdoor attacks. TABLE IV: Comparison of various methods on CIFAR-10 using ResNet-18, with FGSM-RS as the baseline. âBestâ and âFinalâ denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. âCleanâ denotes the classification accuracy for clean examples. âPerturbedâ means that we employ initially perturbed samples. FGSM, PGD, C&W, and A represent different adversarial attacks used for robustness evaluation. âSTâ indicates the number of stable runs out of three (â : CO occurred; â : stable training). Methods Cleanâ Perturbedâ FGSMâ PGD10â PGD20â PGD50â C&Wâ Aâ ST FGSM-RS [82] best 54.69 54.64 34.31 26.01 19.77 17.85 18.41 12.63 â â â final 80.80 83.16 76.62 0.00 0.00 0.00 0.00 0.00 VFT-CO-Clean [54] best 73.49 72.87 44.94 34.81 25.32 21.87 23.00 15.19 â â â final 73.79 73.34 44.48 34.35 24.64 21.75 22.63 15.11 LP-CO-Clean [76] best 61.29 60.83 37.78 30.98 23.82 22.14 21.02 15.35 â â â final 78.43 81.50 77.59 0.00 0.00 0.00 0.00 0.00 RF-CO-Clean [69] best 69.92 70.74 43.28 34.70 25.62 22.81 22.71 15.82 â â â final 72.82 72.75 44.00 34.04 25.04 21.87 22.45 15.22 RFT-CO-Clean [69] best 70.92 69.63 43.29 34.70 25.43 22.83 22.66 15.81 â â â final 72.82 72.67 44.08 34.02 25.12 21.98 22.53 15.28 RSFT-CO-Clean [69] best 71.79 71.51 43.78 34.14 24.82 21.59 22.34 14.89 â â â final 72.17 71.80 43.50 33.84 24.38 21.55 22.08 14.93 IV Backdoor-inspired mitigation of CO Motivated by previous comparisons, we attempt to employ backdoor defenses to mitigate CO. We begin by adapting existing fine-tuning methods designed for backdoor removal. Furthermore, inspired by weight poisoning techniques in backdoor, we introduce a supplementary constraint to suppress weight outliers and enhance stability. TABLE V: Comparison between various techniques, with FGSM-MEP as baseline. âBestâ and âFinalâ denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. âCleanâ denotes the classification accuracy for clean examples. âPerturbedâ means that we employ initially perturbed samples. FGSM, PGD, C&W, and A represent adversarial attacks for evaluation. âSTâ indicates the number of stable runs out of three (â : CO occurred; â : stable training). Methods Cleanâ Perturbedâ FGSMâ PGD10â PGD20â PGD50â C&Wâ Aâ ST FGSM-MEP [35] best 59.01 58.94 35.69 28.66 21.81 20.10 20.49 12.39 â â â final 82.37 82.22 71.19 13.13 8.09 3.12 0.89 0.00 VFT-CO-Clean [54] best 58.43 57.93 38.18 33.71 29.32 28.44 22.93 20.16 â â â final 58.25 57.75 38.28 33.79 29.39 28.62 23.03 20.15 LP-CO-Clean [76] best 49.03 48.42 34.41 31.18 27.67 27.20 22.69 20.71 â â â final 49.03 48.55 34.43 31.16 27.67 27.13 22.74 20.75 RF-CO-Clean [69] best 49.03 48.50 34.57 31.15 27.63 27.10 22.73 20.73 â â â final 49.03 48.45 34.44 31.16 27.55 27.17 22.75 20.78 RFT-CO-Clean [69] best 55.79 55.07 39.31 34.54 30.06 29.34 24.45 21.10 â â â final 55.74 55.05 38.94 34.29 30.11 29.37 24.31 21.04 RSFT-CO-Clean [69] best 40.13 39.65 29.60 25.85 24.57 24.26 21.14 19.41 â â â final 39.00 38.64 28.45 25.46 23.22 22.88 20.45 18.50 Figure 5: Weight distribution relative to the mean weight. âCountâ means the distribution percentage. TABLE VI: Comparison with existing FAT methods with ResNet18 and CIFAR10. PBD is the state-of-the-art method. ΞT _T and ΞE _E represent the perturbation budget used during the training and evaluation processes. âBestâ and âFinalâ denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. âCleanâ denotes the classification accuracy for clean examples. âPerturbedâ means that we employ initially perturbed samples. âSTâ indicates the number of stable runs out of three (â : CO occurred; â : stable training). For LAP, we directly utilize the results reported in [70]. ΞT=ΞE=16/255 _T= _E=16/255 Cleanâ Perturbedâ FGSMâ PGD10â PGD20â PGD50â C&Wâ Aâ ST LAP [70] best 63.73 - - - - - - 19.55 - GradAlign[1] best 58.17 - 39.87 33.12 26.81 24.99 22.63 17.02 â â â final 70.86 - 69.51 0.00 0.00 0.00 0.00 0.00 ZeroGrad[21] best 74.16 - 43.96 32.67 21.98 18.37 20.76 12.07 â â â final 75.60 - 44.89 31.77 20.71 16.76 20.09 10.87 NuAT[74] best 74.62 - 44.92 35.22 25.93 23.67 24.07 18.43 â â â final 75.29 - 45.31 34.85 25.58 23.44 23.62 18.06 FGSM-RS [82] best 54.69 54.64 34.31 26.01 19.77 17.85 18.41 12.63 â â â final 80.80 83.16 76.62 0.00 0.00 0.00 0.00 0.00 PBD-RS [103] best 73.32 72.96 45.03 35.45 25.52 22.61 23.72 16.21 â â â final 73.82 73.48 44.60 34.76 25.05 22.00 23.23 15.82 Ours(âregL_reg) best 74.33 74.00 46.60 36.82 26.46 22.98 24.24 15.72 â â â final 74.84 74.67 46.81 36.21 25.95 22.70 24.22 15.70 FGSM-MEP [35] best 59.01 58.94 35.69 28.66 21.81 20.10 20.49 12.39 â â â final 82.37 82.22 71.19 13.13 8.09 3.12 0.89 0.00 PBD-MEP [103] best 64.45 63.28 45.77 39.70 34.27 32.80 27.64 22.45 â â â final 64.20 63.21 45.29 39.26 33.47 31.88 27.93 22.40 Ours(âregL_reg) best 65.26 64.74 46.11 40.23 34.53 33.07 27.98 22.47 â â â final 65.33 64.63 46.19 40.06 34.22 32.73 28.13 22.21 IV-A Adapted backdoor fine-tuning strategies Once the model fâ(x;θ)f(x;θ) falls into CO during FAT, we apply one epoch of fine-tuning using the strategies illustrated in Figure 4, with the details given as follows: 1) VFT [54, 66]: Fine-tune all parameters. 2) LP [76, 40]: Fine-tune first k model layers, 3) RF [69]: Reinitialize first k layers, freeze them, then fine-tune remaining layers. 4) RFT [69]: Reinitialize first k layers, then fine-tune the whole model. 5) RSFT [69]: Reinitialize first k layers and fine-tune the whole model with an inner product constraint to limit weight shift. We finetune CO-affected models on clean examples as â*CO-Cleanâ. Tables IV and V show the experimental results obtained using FGSM-RS and FGSM-MEP, respectively222Results of finetuning co-affected models on adversarial examples are shown in Appendix.. Our observations are as follows: 1) Adapted backdoor fine-tuning strategies can also resolve CO in FAT. 2) The clean classification accuracies reported in Table V are consistently lower than those in Table IV. This degradation may stem from the regularization term âgradR_grad in Eq. (4), which potentially constrains the modelâs capacity to fit the clean data distribution. Furthermore, reinitializing a stable FAT process after CO may inherently limit the modelâs final performance, e.g. wrong optimization direction. 3) Notably, even when fine-tuning yields a temporarily stable model, CO tends to reoccur. This is expected, as the goal of backdoor defenses is to transition a trained, backdoored model out of the backdoor state. Overall, fine-tuning strategies designed for backdoor attacks are also applicable to mitigating CO. IV-B Weight-poisoning inspired strategy To develop lasting mitigation strategies for CO, we begin by examining training-time backdoor defense techniques [46]. These methods typically identify and filter backdoor-infected samples using either one-shot [91] or iterative [11] filtering procedures. In FAT, Zhao et al. [102] demonstrate that applying a convergence-smoothing constraint to a subset of samples with notable convergence divergence can enhance the overall training stability. However, the effectiveness of this method is highly sensitive to biases in sample selection. Motivated by backdoor approaches with weight poisoning [72, 53, 41], we examine the weight distributions of both stable and CO-affected models. Specifically, we compute the mean weight value wÂŻ w for each layer and assess the weight distribution relative to this mean. Figure 5 presents experimental results on CIFAR-10 with ResNet-18, showing that CO-affected models exhibit weight outliersâweights that deviate markedly from wÂŻ w. Hence, we attempt to address CO by suppressing weight outliers. Unlike backdoor pruning techniques [92, 18, 45], which typically clamp weights, we introduce an additional constraint that suppresses weight outliers and prevents the formation of adversarial paths. âreg=âlââwâwÂŻlexpâĄ(|wâwÂŻl|wÂŻl+ÎąâΡ)â|wâwÂŻl|,L_reg= _l _wâ w_l ( |w- w_l| w_l+Îą-Ρ )\,|w- w_l|, (12) wÂŻl=1|Wl|ââwâWl|w|, w_l= 1|W_l| _wâ W_l|w|, (13) where âĽâ âĽ\|¡\| is utilized to obtain the absolute value. C denotes the set of convolutional layers, and WlW_l represents the set of l-th layer weights. Hyper-parameters Ρ and Îą are set to 10 and 10-5, respectively. Table VI reports results on CIFAR-10 with ResNet18, showing that incorporating âregL_reg effectively mitigates CO in FAT. Additional experiments across diverse datasets, architectures and perturbation budgets are shown in the Appendix, which also bring consistent improvements in mitigating CO. Overall, the proposed method not only resolves CO but also achieves performance superior to or comparable with the state-of-the-art PBD. Unlike PBD, which aligns clean and adversarial predictions at the cost of clean accuracy, and LAT, whose effectiveness is limited by its strict emphasis on robust weights, our method targets only anomalous weights, providing greater flexibility in weight selection. Additionally, we hypothesize that these anomalous weights predominantly form the adversarial pathway. We further conduct ablation studies to compare our method against simpler alternatives, including â2 _2 regularization and weight clipping based on the ratio between each weight and its mean (using the same threshold as our method). For each approach, we report results averaged over three independent runs in Fig. 6. As can be observed, directly applying â2 _2 regularization fails to achieve stable optimization, where the model loses its adversarial robustness in the later stages of fast adversarial training. Weight clipping, while partially mitigating catastrophic overfitting, suffers from training instability and may relapse into catastrophic overfitting at any point. Moreover, robustness after clipping is slightly compromised. In contrast, our method consistently maintains stable performance throughout the training process. Figure 6: Ablation studies comparing our âregL_reg against simpler alternatives. V Discussions and additional analyses Why FAT improves robustness and why CO occurs. We further investigate why single-step FAT (e.g., FGSM-based AT) tends to induce triggerâlike overfitting, whereas multi-step AT (e.g. PGD-based AT [86, 104, 7]) does not. To this end, we conduct additional experiments comparing the two methods using multiple similarity metrics. As shown in Fig. 7, we report inter-class and intra-class prediction similarity, adversarial perturbation similarity, and adversarial example similarity. Our results reveal three key observations. First, inter-class similarity exhibits comparable trends for both methods, indicating that they behave similarly across different learning stages. Second, intra-class perturbation similarity and intra-class adversarial example similarity are lower for FGSM-based AT than for PGD-based AT. This does not necessarily imply that the âtriggerâ patterns in FGSM perturbations are more consistent; rather, it may reflect that PGD, through its multi-step optimization, discovers more semantically coherent adversarial directions, leading to higher intra-class similarity in its perturbations. Third, intra-class prediction similarity is substantially higher for FGSM-based AT than for PGD-based AT. This suggests that although FGSM generates perturbations that appear more diverse (i.e., lower perturbation similarity), the modelâs predictions on them are remarkably uniform. In contrast, for PGD-based AT, predictions on intra-class adversarial examples are less consistent, indicating that PGD introduces a wider variety of perturbation directions. Overall, the single-step optimization of FGSM restricts its adversarial directions, thereby causing CO. Figure 7: Comparison between FGSM-based AT and PGD-based AT. Transferability of methods across different tasks. Section IV shows that backdoor-inspired defense techniques can mitigate CO. We further evaluate whether the method for CO can generalize to backdoor-related tasks. For unlearnable techniques with sample-wise perturbations, adding uniform random noise to the poisoned dataset can effectively neutralize their impact. For example, training a ResNet-18 model on CIFAR with the original unlearnable dataset yields 13.58% accuracy on clean samples, which increases to 94.22% after applying random noise. Hence, we focus on unlearnable techniques with class-wise perturbations. Prior work often addresses these using adversarial training [20]; for instance, a model trained on the original unlearnable dataset achieves 10.05% accuracy on clean samples, whereas adversarial training improves accuracy to 86.89%. We investigate the transferability of the CO mitigation strategy to unlearnable tasks, with results summarized in Table VII. When the poisoning budget in unlearnable tasks is small, the ârâeâgL_reg for mitigating CO can also resolve the class-wise unlearnable attack, achieving performance comparable to the AT-based approach but at a significantly lower computational cost. However, under a large poisoning budget (e.g.e.g. 8/255), its effectiveness diminishes. These results further support viewing CO as a weak-trigger variant of unlearnable tasks. Additionally, the reduced effectiveness of ârâeâgL_reg under larger poisoning budgets (e.g., 8/2558/255) can be attributed to two key observations. First, the occurrence of weight outliers is not positively correlated with perturbation magnitude. In the extreme case where unconstrained adversarial examples are used, the unlearnable sample set is effectively transformed into a completely different dataset. Under this condition, the model exhibits no weight outliers, rendering ârâeâgL_reg completely ineffective. This suggests that the performance degradation of ârâeâgL_reg under larger perturbations arises because larger perturbations facilitate a direct association between the perturbation and the target label, thereby bypassing the need for weight anomalies. Second, our experiments focus on class-wise unlearnable perturbations, which tend to mimic the semantic features of a specific class. Once such perturbations successfully capture class semantics, the unlearnable samples again resemble a different dataset, resulting in the absence of weight outliers. Impact of the hyperparameter β. The results in Figure 8 show that both too small and too large values of β adversely affect model robustness. Specifically, when β is too small, the regularization not only penalizes weight outliers but also constrains normal weights, thereby impairing the overall model performance. Conversely, when β is too large, the suppression of weight outliers becomes insufficient, rendering the regularization ineffective. Figure 8: Ablation studies on the hyperparameter β. TABLE VII: Transferability of the CO mitigation strategy (ârâeâgL_reg in Eq. (12) ) to unlearnable tasks [30]. âPoisonedâ, âATâ, and âârâeâgL_regâ correspond to models trained on the CIFAR-10 poisoned dataset using ResNet-18 with standard training, adversarial training, and standard training augmented with ârâeâgL_reg, respectively. Each model is trained for 60 epochs using a cyclic learning rate schedule [73]. The reported values represent the final evaluation accuracyâ on clean samples. Poisoning Budgets Poisoned[30] AT[20] ârâeâgL_reg 4/255 17.50 86.95 86.71 6/255 10.05 86.89 86.48 8/255 9.85 86.32 46.57 Training Time (minutes) 15.4 36.8 15.5 VI Conclusions and future works In this work, we offer a systematic explanation of CO from a backdoor perspective, establishing a unified frameworkâtrigger overfittingâthat encompasses standard backdoor attacks, CO, and unlearnable tasks. By validating phenomena such as pathway division, diverse path predictions, and the presence of universal class-distinguishable triggers in CO, we conceptualize CO as a weak-trigger variant of unlearnable tasks. Leveraging these insights to mitigate CO, we adapt several backdoor fine-tuning techniques to the FAT framework and propose a weight outlier suppression constraint. Experimental results validate our mechanistic explanation and demonstrate the effectiveness of backdoor-inspired strategies in alleviating CO. Future research: This work has demonstrated that adversarial perturbations crafted by CO-affected models contain universal class-distinguishable (UCD) triggers. A natural issue arises: Can these UCD triggers be further extracted and explicitly identified? In other words, is it possible to deliberately induce the CO phenomenon with the extracted UCD triggers? Moreover, can triggers associated with other adversarial attacks be similarly extracted? Models tend to overfit to triggers, resulting in high-confidence predictions on corresponding adversarial samples. Therefore, when defending against a specific attack, integrating manually extracted adversarial triggers can be more effective, as conventional adversarial training methods involve a trade-off between clean accuracy and robustness. References [1] F. N. Andriushchenko M (2020) Understanding and improving fast adversarial training. In nips, p. 16048â16059. Cited by: §I, TABLE VI. [2] M. Andriushchenko and N. Flammarion (2020) Understanding and improving fast adversarial training. In nips, Vol. 33, p. 16048â16059. Cited by: §I, §I-A. [3] M. Baek and D. Baker (2022) Deep learning and protein structure modeling. Nature methods 19 (1), p. 13â14. Cited by: §I. [4] Y. Cao, C. Xiao, A. Anandkumar, D. Xu, and M. Pavone (2022) Advdo: realistic adversarial attacks for trajectory prediction. In ECCV, p. 36â52. Cited by: §I, §I-A. [5] N. Carlini and D. Wagner Towards evaluating the robustness of neural networks. In , Cited by: TABLE I, TABLE I. [6] S. Casper, L. Schulze, O. Patel, and D. Hadfield-Menell (2024) Defending against unforeseen failure modes with latent adversarial training. arXiv preprint arXiv:2403.05030. Cited by: §I. [7] P. Chen, B. Kung, and J. Chen (2021) Class-aware robust adversarial training for object detection. In CVPR, p. 10420â10429. Cited by: §V. [8] W. Chen, B. Wu, and H. Wang (2022) Effective backdoor defense by exploiting sensitivity of poisoned samples. In nips, Vol. 35, p. 9727â9737. Cited by: §I, §I-A. [9] W. Chen, X. Xu, X. Wang, Z. Li, and Y. Chen (2025) FSBA: invisible backdoor attacks via frequency domain and singular value decomposition. Expert Systems with Applications 288, p. 127830. Cited by: §I-B. [10] X. Chen, C. Liu, B. Li, K. Lu, and D. Song (2017) Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526. Cited by: §I-B. [11] Y. Chen, W. Haiwei, and Z. Jiantao (2024) Progressive poisoned data isolation for training-time backdoor defense. In AAAI, Vol. 38, p. 1319â1327. Cited by: §IV-B. [12] S. Chou, P. Chen, and T. Ho (2023) How to backdoor diffusion models?. In CVPR, p. 4015â4024. Cited by: §I, §I-A. [13] F. Croce and M. Hein (2020) Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In ICML, Cited by: TABLE I, TABLE I. [14] P. de Jorge Aranda, A. Bibi, R. Volpi, A. Sanyal, P. Torr, G. Rogez, and P. Dokania (2022) Make some noise: reliable and efficient single-step adversarial training. In nips, Vol. 35, p. 12881â12893. Cited by: §I-A. [15] K. Doan, Y. Lao, W. Zhao, and P. Li (2021) Lira: learnable, imperceptible and robust backdoor attacks. In ICCV, p. 11966â11976. Cited by: §I-B. [16] H. M. Dolatabadi, S. Erfani, and C. Leckie (2024) The devilâs advocate: shattering the illusion of unexploitable data using diffusion models. In IEEE Conference on Secure and Trustworthy Machine Learning, p. 358â386. Cited by: §I-A. [17] Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li (2018) Boosting adversarial attacks with momentum. In CVPR, p. 9185â9193. Cited by: §I-A. [18] W. Dongxian and W. Yisen (2021) Adversarial neuron pruning purifies backdoored deep models. In nips, Vol. 34, p. 16913â16925. Cited by: §IV-B. [19] F. Fang, Y. Bai, S. Ni, M. Yang, X. Chen, and R. Xu (2024) Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. arXiv preprint arXiv:2405.20978. Cited by: §I. [20] S. Fu, F. He, Y. Liu, L. Shen, and D. Tao (2022) Robust unlearnable examples: protecting data against adversarial learning. arXiv preprint arXiv:2203.14533. Cited by: §I-B, §I-A, TABLE VII, §V. [21] Z. Golgooni, M. Saberi, M. Eskandar, and M. H. Rohban (2021) ZeroGrad: mitigating and explaining catastrophic overfitting in fgsm adversarial training. p. arXiv preprint arXiv:2103.15476. Cited by: TABLE VI. [22] Z. Golgooni, M. Saberi, M. Eskandar, and M. H. Rohban (2023) ZeroGrad: costless conscious remedies for catastrophic overfitting in the fgsm adversarial training. Intelligent Systems with Applications 19, p. 200258. Cited by: §I, §I-A. [23] I. J. Goodfellow, J. Shlens, and C. Szegedy (2015) Explaining and harnessing adversarial examples. In ICLR, Cited by: §I-A. [24] J. Gu, H. Zhao, V. Tresp, and P. H. Torr (2022) Segpgd: an effective and efficient adversarial attack for evaluating and boosting segmentation robustness. In ECCV, p. 308â325. Cited by: §I, §I-A. [25] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg (2019) Badnets: evaluating backdooring attacks on deep neural networks. IEEE Access 7, p. 47230â47244. Cited by: §I-B. [26] L. Guzman-Nateras, M. Van Nguyen, and T. Nguyen (2022) Cross-lingual event detection via optimized adversarial training. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 5588â5599. Cited by: §I, §I-A. [27] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In CVPR, p. 770â778. Cited by: §I-A, TABLE I, TABLE I. [28] K. He, X. Zhang, S. Ren, and J. Sun (2016) Identity mappings in deep residual networks. In ECCV, p. 630â645. Cited by: TABLE I, TABLE I. [29] Z. He, T. Li, S. Chen, and X. Huang (2023) Investigating catastrophic overfitting in fast adversarial training: a self-fitting perspective. In CVPR, p. 2313â2320. Cited by: §I, §I-A. [30] H. Huang, X. Ma, S. M. Erfani, J. Bailey, and Y. Wang (2021) Unlearnable examples: making personal data unexploitable. In ICLR, Cited by: §I-B, TABLE VII, TABLE VII, TABLE VII. [31] Z. Huang, Y. Fan, C. Liu, W. Zhang, Y. Zhang, M. Salzmann, S. SĂźsstrunk, and J. Wang (2022) Fast adversarial training with adaptive step size. arXiv preprint arXiv:2206.02417. Cited by: §I. [32] Z. Huang, Y. Fan, C. Liu, W. Zhang, Y. Zhang, M. Salzmann, S. SĂźsstrunk, and J. Wang (2023) Fast adversarial training with adaptive step size. IEEE TIP 32, p. 6102â6114. Cited by: §I. [33] X. Jia, Y. Chen, X. Mao, R. Duan, J. Gu, R. Zhang, H. Xue, Y. Liu, and X. Cao (2024) Revisiting and exploring efficient fast adversarial training via law: lipschitz regularization and auto weight averaging. IEEE TIFS. Cited by: §I-A. [34] X. Jia, J. Li, J. Gu, Y. Bai, and X. Cao (2024) Fast propagation is better: accelerating single-step adversarial training via sampling subnetworks. IEEE TIFS. Cited by: §I-A. [35] X. Jia, Y. Zhang, X. Wei, B. Wu, K. Ma, J. Wang, and X. Cao (2022) Prior-guided adversarial initialization for fast adversarial training. In ECCV, p. 567â584. Cited by: §I-A, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE V, TABLE VI. [36] X. Jia, Y. Zhang, X. Wei, B. Wu, K. Ma, J. Wang, and X. Cao (2024) Improving fast adversarial training with prior-guided knowledge. TPAMI 46 (9), p. 6367â6383. Cited by: §I-A. [37] X. Jia, Y. Zhang, B. Wu, K. Ma, J. Wang, and X. Cao (2022) LAS-at: adversarial training with learnable attack strategy. In CVPR, p. 13398â13408. Cited by: §I, §I-A. [38] X. Jia, Y. Zhang, B. Wu, J. Wang, and X. Cao (2022) Boosting fast adversarial training with learnable adversarial initialization. IEEE TIP 31, p. 4417â4430. Cited by: §I-A. [39] G. Jin, X. Yi, W. Huang, S. Schewe, and X. Huang (2022) Enhancing adversarial training with second-order statistics of weights. In CVPR, p. 15273â15283. Cited by: §I, §I-A. [40] S. Ke, C. Hou, G. Fanti, and S. Oh (2024) On the convergence of differentially-private fine-tuning: to linearly probe or to fully fine-tune?. arXiv preprint arXiv:2402.18905. Cited by: §IV-A. [41] K. Keita, M. Paul, and N. Graham (2020) Weight poisoning attacks on pre-trained models. arXiv preprint arXiv:2004.06660. Cited by: §IV-B. [42] H. Kim, W. Lee, and J. Lee (2021) Understanding catastrophic overfitting in single-step adversarial training. In AAAI, Vol. 35, p. 8119â8127. Cited by: §I. [43] A. Krizhevsky, G. Hinton, et al. (2009) Learning multiple layers of features from tiny images. Cited by: §I-A, TABLE I, TABLE I, TABLE I, TABLE I. [44] H. Kuang, H. Liu, X. Lin, and R. Ji (2024) Defense against adversarial attacks using topology aligning adversarial training. IEEE TIFS 19, p. 3659â3673. Cited by: §I-A. [45] C. Kunbei, C. MdHafizulIslam, Z. Zhenkai, and Y. Fan Deepvenom: persistent dnn backdoors exploiting transient weight perturbations in memories. In , Cited by: §IV-B. [46] G. Kuofeng, B. Yang, G. Jindong, Y. Yong, and X. Shu-Tao (2023) Backdoor defense via adaptively splitting poisoned dataset. In CVPR, p. 4005â4014. Cited by: §IV-B. [47] A. Kurakin, I.J. Goodfellow, and S. Bengio (2017) Adversarial machine learning at scale. In ICLR, Cited by: §I-A. [48] C. Li, R. Pang, Z. Xi, T. Du, S. Ji, Y. Yao, and T. Wang (2023) An embarrassingly simple backdoor attack on self-supervised learning. In ICCV, p. 4367â4378. Cited by: §I-B. [49] M. Li, X. Wang, Z. Yu, S. Hu, Z. Zhou, L. Zhang, and L. Y. Zhang (2025) Detecting and corrupting convolution-based unlearnable examples. In AAAI, Vol. 39, p. 18403â18411. Cited by: §I-B. [50] T. Li, Y. Wu, S. Chen, K. Fang, and X. Huang (2022) Subspace adversarial training. In CVPR, p. 13409â13418. Cited by: §I, §I-A. [51] Y. Li, Y. Li, B. Wu, L. Li, R. He, and S. Lyu (2021) Invisible backdoor attack with sample-specific triggers. In ICCV, p. 16463â16472. Cited by: §I-B. [52] S. Liang, M. Zhu, A. Liu, B. Wu, X. Cao, and E. Chang (2024) Badclip: dual-embedding guided backdoor attack on multimodal contrastive learning. In CVPR, p. 24645â24654. Cited by: §I. [53] L. Linyang, S. Demin, L. Xiaonan, Z. Jiehang, M. Ruotian, and Q. Xipeng (2021) Backdoor attacks on pre-trained models by layerwise weight poisoning. arXiv preprint arXiv:2108.13888. Cited by: §IV-B. [54] K. Liu, B. Dolan-Gavitt, and S. Garg (2018) Fine-pruning: defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, p. 273â294. Cited by: TABLE IV, §IV-A, TABLE V. [55] X. Liu, X. Jia, Y. Xun, S. Liang, and X. Cao (2024) Multimodal unlearnable examples: protecting data against multimodal contrastive learning. In ACMM, p. 8024â8033. Cited by: §I-B. [56] Y. Liu, S. Ma, Y. Aafer, W. Lee, J. Zhai, W. Wang, and X. Zhang (2018) Trojaning attack on neural networks. In NDSS, Cited by: §I-B. [57] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu (2018) Towards deep learning models resistant to adversarial attacks. In ICLR, Cited by: §I, §I-A. [58] L. McInnes, J. Healy, and J. Melville (2018) Umap: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. Cited by: §I-A. [59] Y. Mo, D. Wu, Y. Wang, Y. Guo, and Y. Wang (2022) When adversarial training meets vision transformers: recipes from training to architecture. In nips, Vol. 35, p. 18599â18611. Cited by: §I, §I-A. [60] S. M. Mousavi and G. C. Beroza (2022) Deep-learning seismology. Science 377 (6607), p. eabm4470. Cited by: §I, §I-A. [61] A. Nguyen and A. Tran (2021) Wanetâimperceptible warping-based backdoor attack. arXiv preprint arXiv:2102.10369. Cited by: §I-B. [62] C. Pan, Q. Li, and X. Yao (2024) Adversarial initialization with universal adversarial perturbation: a new approach to fast adversarial training. In AAAI, Vol. 38, p. 21501â21509. Cited by: §I. [63] G. Y. Park and S. W. Lee (2021) Reliably fast adversarial training via latent adversarial perturbation. In ICCV, p. 7758â7767. Cited by: §I. [64] T. D. Pereira, N. Tabris, A. Matsliah, D. M. Turner, J. Li, S. Ravindranath, E. S. Papadoyannis, E. Normand, D. S. Deutsch, Z. Y. Wang, et al. (2022) SLEAP: a deep learning system for multi-animal pose tracking. Nature methods 19 (4), p. 486â495. Cited by: §I, §I-A. [65] T. Qin, X. Gao, J. Zhao, K. Ye, and C. Xu (2023) Learning the unlearnable: adversarial augmentations suppress unlearnable example attacks. arXiv preprint arXiv:2303.15127. Cited by: §I-B. [66] Z. Qin, L. Yao, D. Chen, Y. Li, B. Ding, and M. Cheng (2023) Revisiting personalized federated learning: robustness against backdoor attacks. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, p. 4743â4755. Cited by: §IV-A. [67] J. Ren, H. Xu, Y. Wan, X. Ma, L. Sun, and J. Tang (2022) Transferable unlearnable examples. arXiv preprint arXiv:2210.10114. Cited by: §I-B. [68] L. Rice, E. Wong, and Z. Kolter (2020) Overfitting in adversarially robust deep learning. In ICML, p. 8093â8104. Cited by: §I-A. [69] M. Rui, Q. Zeyu, S. Li, and C. Minhao (2023) Towards stable backdoor purification through feature shift tuning. In nips, Vol. 36, p. 75286â75306. Cited by: TABLE IV, TABLE IV, TABLE IV, §IV-A, TABLE V, TABLE V, TABLE V. [70] L. Runqi, Y. Chaojian, H. Bo, S. Hang, and L. Tongliang (2024) Revealing the pseudo-robust shortcut dependency. In ICML, p. 2663â2672. Cited by: §I-A, §I-A, TABLE VI, TABLE VI, TABLE VI. [71] N. Sapoval, A. Aghazadeh, M. G. Nute, D. A. Antunes, A. Balaji, R. Baraniuk, C. Barberan, R. Dannenfelser, C. Dun, M. Edrisi, et al. (2022) Current progress and open challenges for applying deep learning across the biosciences. Nature Communications 13 (1), p. 1728. Cited by: §I. [72] Z. Shuai, G. Leilei, T. LuuAnh, F. Jie, L. Lingjuan, J. Meihuizi, and W. Jinming (2024) Defending against weight-poisoning backdoor attacks for parameter-efficient fine-tuning. arXiv preprint arXiv:2402.12168. Cited by: §IV-B. [73] L. N. Smith (2017) Cyclical learning rates for training neural networks. In WACV, p. 464â472. Cited by: TABLE VII, TABLE VII. [74] G. Sriramanan, S. Addepalli, A. Baburaj, et al. (2021) Towards efficient and effective adversarial training. In nips, Vol. 34, p. 11821â11833. Cited by: TABLE VI. [75] K. Tang, T. Lou, W. Peng, N. Chen, Y. Shi, and W. Wang (2024) Effective single-step adversarial training with energy-based models. IEEE Transactions on Emerging Topics in Computational Intelligence. Cited by: §I. [76] A. Tomihari and I. Sato (2024) Understanding linear probing then fine-tuning language models from ntk perspective. arXiv preprint arXiv:2405.16747. Cited by: TABLE IV, §IV-A, TABLE V. [77] K. Tong, C. Jiang, J. Gui, and Y. Cao (2024) Taxonomy driven fast adversarial training. In AAAI, Vol. 38, p. 5233â5242. Cited by: §I-A. [78] T. Wang, Y. Yao, F. Xu, S. An, H. Tong, and T. Wang (2022) An invisible black-box backdoor attack through frequency domain. In ECCV, p. 396â413. Cited by: §I-B. [79] Z. Wang, X. Li, H. Zhu, and C. Xie (2024) Revisiting adversarial training at scale. In CVPR, p. 24675â24685. Cited by: §I-A. [80] Z. Wang, H. Wang, C. Tian, and Y. Jin (2024) Preventing catastrophic overfitting in fast adversarial training: a bi-level optimization perspective. In ECCV, p. 144â160. Cited by: §I-A. [81] L. Wei, L. Jin, and X. Luo (2022) Noise-suppressing neural dynamics for time-dependent constrained nonlinear optimization with applications. IEEE Transactions on Systems, Man, and Cybernetics: Systems 52 (10), p. 6139â6150. Cited by: §I, §I-A. [82] K. J. Z. Wong E (2020) Fast is better than free: revisiting adversarial training. In ICLR, Cited by: §I, §I, §I-A, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE I, TABLE IV, TABLE VI. [83] B. Wu, J. Gu, Z. Li, D. Cai, X. He, and W. Liu (2022) Towards efficient adversarial training on vision transformers. In ECCV, p. 307â325. Cited by: §I, §I-A. [84] J. Xiao, Y. Fan, R. Sun, J. Wang, and Z. Luo (2022) Stability analysis and generalization bounds of adversarial training. In nips, Vol. 35, p. 15446â15459. Cited by: §I, §I-A. [85] J. Xiaojun, Z. Yong, W. Xingxing, W. Baoyuan, M. Ke, W. Jue, and C. Xiaochun (2022) Prior-guided adversarial initialization for fast adversarial training. In ECCV, Cited by: §I. [86] Y. Xiong and C. Hsieh (2020) Improved adversarial training via learned optimizer. In ECCV, p. 85â100. Cited by: §V. [87] L. Yang, H. Qian, Z. Zhang, J. Liu, and B. Cui (2024) Structure-guided adversarial training of diffusion models. In CVPR, p. 7256â7266. Cited by: §I. [88] W. Yang, X. Bi, Y. Lin, S. Chen, J. Zhou, and X. Sun (2024) Watch out for your agents! investigating backdoor threats to llm-based agents. In nips, Vol. 37, p. 100938â100964. Cited by: §I. [89] Y. Yang, X. Liu, and K. He (2024) Fast adversarial training against textual adversarial attacks. arXiv preprint arXiv:2401.12461. Cited by: §I. [90] J. Ye and X. Wang (2024) Ungeneralizable examples. In CVPR, p. 11944â11953. Cited by: §I-B. [91] L. Yige, L. Xixiang, K. Nodens, L. Lingjuan, L. Bo, and M. Xingjun (2021) Anti-backdoor learning: training clean models on poisoned data. In nips, Vol. 34, p. 14900â14912. Cited by: §IV-B. [92] L. Yige, L. Xixiang, M. Xingjun, K. Nodens, L. Lingjuan, L. Bo, and J. Yu-Gang (2023) Reconstructive neuron pruning for backdoor defense. In ICML, p. 19837â19854. Cited by: §IV-B. [93] W. Yin, J. Lou, P. Zhou, Y. Xie, D. Feng, Y. Sun, T. Zhang, and L. Sun (2024) Physical backdoor: towards temperature-based backdoor attacks in the physical world. In CVPR, p. 12733â12743. Cited by: §I. [94] Y. Yu, Q. Zheng, S. Yang, W. Yang, J. Liu, S. Lu, Y. Tan, K. Lam, and A. Kot (2024) Unlearnable examples detection via iterative filtering. In International Conference on Artificial Neural Networks, p. 241â256. Cited by: §I-A. [95] X. Yue, N. Mou, Q. Wang, and L. Zhao (2024) Revisiting adversarial training under long-tailed distributions. In CVPR, p. 24492â24501. Cited by: §I. [96] M. Zareapoor and P. Shamsolmoali (2024) Rethinking fast adversarial training: a splitting technique to overcome catastrophic overfitting. In ECCV, p. 34â51. Cited by: §I-A. [97] Y. Zeng, W. Park, Z. M. Mao, and R. Jia (2021) Rethinking the backdoor attacksâ triggers: a frequency perspective. In ICCV, p. 16473â16481. Cited by: §I-B. [98] J. Zhang, X. Ma, Q. Yi, J. Sang, Y. Jiang, Y. Wang, and C. Xu (2023) Unlearnable clusters: towards label-agnostic unlearnable examples. In CVPR, p. 3984â3993. Cited by: §I-B. [99] S. Zhang, Y. Pan, Q. Liu, Z. Yan, K. R. Choo, and G. Wang (2024) Backdoor attacks and defenses targeting multi-domain ai models: a comprehensive review. ACM Computing Surveys 57 (4), p. 1â35. Cited by: §I. [100] Y. Zhang, G. Zhang, P. Khanduri, M. Hong, S. Chang, and S. Liu (2022) Revisiting and advancing fast adversarial training through the lens of bi-level optimization. In ICML, p. 26693â26712. Cited by: §I-A. [101] Y. Zhang, X. Chen, J. Jia, Y. Zhang, C. Fan, J. Liu, M. Hong, K. Ding, and S. Liu (2024) Defensive unlearning with adversarial training for robust concept erasure in diffusion models. In nips, Vol. 37, p. 36748â36776. Cited by: §I. [102] M. Zhao, L. Zhang, Y. Kong, and B. Yin (2023) Fast adversarial training with smooth convergence. In ICCV, p. 4720â4729. Cited by: §I, §I-A, §IV-B. [103] M. Zhao, L. Zhang, Y. Kong, and B. Yin (2024) Catastrophic overfitting: a potential blessing in disguise. In ECCV, p. 293â310. Cited by: §I, §I-A, §I-A, §I-A, TABLE VI, TABLE VI. [104] X. Zhong and C. Liu (2024) Sparse-pgd: a unified framework for sparse adversarial perturbations generation. arXiv preprint arXiv:2405.05075. Cited by: §V. [105] Y. Zhong, X. Liu, D. Zhai, J. Jiang, and X. Ji (2022) Shadows can be dangerous: stealthy and effective physical-world adversarial attack by natural phenomenon. In CVPR, p. 15345â15354. Cited by: §I, §I-A. [106] C. Zuo, J. Qian, S. Feng, W. Yin, Y. Li, P. Fan, J. Han, K. Qian, and Q. Chen (2022) Deep learning in optical metrology: a review. Light: Science & Applications 11 (1), p. 39. Cited by: §I, §I-A.