Paper deep dive
AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
Shuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang, Xiaojun Jia, Kunpeng Ning, Jiayu Yao, Jigang Wang, Hailiang Dai, Yibing Song, Li Yuan
Models: Gemma-2-9B, Llama-2-7B, Llama-3-8B, Qwen-2-7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 7:20:07 PM
Summary
AsFT (Anchoring Safety in Fine-Tuning) is a novel method designed to preserve LLM safety during fine-tuning by constraining parameter updates within a 'narrow safety basin'. By identifying an alignment direction (derived from weight differences between aligned and unaligned models) and penalizing updates orthogonal to this direction, AsFT effectively prevents safety degradation caused by harmful data, outperforming existing baselines in both safety and task performance.
Entities (5)
Relation Signals (3)
AsFT â constrainsupdateswithin â Narrow Safety Basin
confidence 98% · AsFT effectively constrains the model within the 'narrow safety basin'
AsFT â reduces â Harmful Score
confidence 95% · AsFT reduces harmful behaviors by up to 7.60%
AsFT â uses â LoRA
confidence 95% · We employ LoRA (Hu et al. 2022) for efficient fine-tuning of LLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment direction - defined by weight differences between aligned (safe) and unaligned models - rapidly compromise model safety. In contrast, updates along the alignment direction largely preserve it, revealing the parameter space as a "narrow safety basin". To address this, we propose AsFT (Anchoring Safety in Fine-Tuning) to maintain safety by explicitly constraining update directions during fine-tuning. By penalizing updates orthogonal to the alignment direction, AsFT effectively constrains the model within the "narrow safety basin," thus preserving its inherent safety. Extensive experiments on multiple datasets and models show that AsFT reduces harmful behaviors by up to 7.60%, improves task performance by 3.44%, and consistently outperforms existing methods across multiple tasks.
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
48,064 characters extracted from source content.
Expand or collapse full text
AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin Shuo Yang 1 * , Qihui Zhang 1 * , Yuyang Liu 1â , Xiaojun Jia 3 , Kun-Peng Ning 1 , Jia-Yu Yao 1 , Jigang Wang 4 , Hailiang Dai 4 , Yibing Song 5 , Li Yuan 1,2â 1 Peking University, Shenzhen Graduate School 2 Peng Cheng Laboratory 3 Nanyang Technological University (NTU) 4 ZTE Corporation 5 Independent Researcher shuo_yang@stu.pku.edu.cn, qhzhang25@stu.pku.edu.cn, yuanli-ece@pku.edu.cn Abstract Fine-tuning large language models (LLMs) improves perfor- mance but introduces critical safety vulnerabilities: even min- imal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment directionâdefined by weight differences between aligned (safe) and unaligned modelsârapidly compromise model safety. In contrast, updates along the alignment direction largely preserve it, revealing the parameter space as a "narrow safety basin". To address this, we propose AsFT (Anchoring Safety in Fine-Tuning) to maintain safety by explicitly con- straining update directions during fine-tuning. By penalizing updates orthogonal to the alignment direction, AsFT effec- tively constrains the model within the "narrow safety basin," thus preserving its inherent safety. Extensive experiments on multiple datasets and models show that AsFT reduces harm- ful behaviors by up to 7.60%, improves task performance by 3.44%, and consistently outperforms existing methods across multiple tasks. Code â https://github.com/PKU-YuanGroup/AsFT 1 Introduction The rapid advancement of large language models (LLMs) has led to their widespread adoption, where fine-tuning is es- sential to adapt these models to specific tasks and scenarios. However, fine-tuning exposes critical safety vulnerabilities. Even small amounts of malicious or harmless data during fine-tuning can compromise the modelâs safeguards, caus- ing it to generate harmful outputs post-fine-tuning (Huang et al. 2025b; Bianchi et al. 2024; Qi et al. 2024b). This raises the urgent need for methods that balance task-specific utility with robust safety defenses (Huang et al. 2024c). Currently, there are various strategies for enhancing safety during LLM fine-tuning. While these strategies primarily rely on data-driven methods, they face a significant chal- lenge: reliance on high-quality datasets, which are both * Equal contribution â Corresponding author Copyright © 2026, Association for the Advancement of Artificial Intelligence (w.aaai.org). All rights reserved. Figure 1: (a) The Safety Basin (Peng et al. 2024) shows a region where perturbations along d random preserve model safety, while safety sharply declines outside this area. (b) The Narrow Safety Basin demonstrates the asymmetry be- tween d aligned and d harm , where d aligned allows larger perturba- tions, while d harm causes sharp safety declines. In both sub- figures, lower values indicate higher safety. costly and susceptible to bias (Huang et al. 2024c). Post- tuning methods like Safe LoRA (Hsu et al. 2024) mitigate fine-tuningâs negative impact on model safety by discretiz- ing and projecting LoRA weights into a safety-aligned sub- space. However, they overlook layer continuity, as discrete projections can disrupt the consistency of learned features across layers. By focusing primarily on safety-related fea- tures, they neglect the performance-related characteristics brought by training data, degrading modelsâ performance. To address the limitations mentioned above, we aim to develop a data-free approach that leverages continuous op- timization to enhance safety during fine-tuning. We observe that aligned models, developed under rigorous protocols, ex- hibit robust defenses against harmful inputs (Qi et al. 2024b; Hsu et al. 2024), whereas their unaligned counterparts (i.e., base models) lack such safeguards. This contrast inspires us to explore the latent information within the model parameter space. The weight difference âW between these two mod- els encapsulates the alignment efforts undertaken by LLM vendors to enhance model safety. It not only reflects the core arXiv:2506.08473v3 [cs.LG] 7 Jan 2026 Figure 2: The proposed AsFT decomposes parameter updates into d aligned and d â„ harm , suppresses harmful updates along d â„ harm by regularization and constrains updates within the narrow safety basin. alignment process but also provides a critical direction for safety optimization (Hsu et al. 2024; Zhao et al. 2025a). Given these observations, we hypothesize that the align- ment direction can guide safety-preserving updates during fine-tuning and thus addresses the following question: Can this weight difference serve as an anchor to guide safety-preserving updates? To investigate it, we explored the modelâs safety land- scape (Peng et al. 2024) as shown in Fig. 1 and discov- ered a striking asymmetry: perturbations along the align- ment direction (d aligned , defined based on this weight differ- ence âW) largely preserve model safety. Conversely, di- rection orthogonal to it, which we term d â„ harm , is critically sensitive, where even small updates can trigger a sharp de- cline in safety. This finding reframes the LLM parameter space as a ânarrow safety basinâ (Fig. 1(b)), a tight corridor where safety is maintained by moving along the alignment direction, while any deviation into the orthogonal space risks falling off a âsafety cliffâ. To navigate this treacherous landscape, we propose AsFT (Anchoring Safety in Fine-Tuning), a method (Fig. 2) that maintains models within the ânarrow safety basin" by pe- nalizing parameter updates orthogonal to the alignment di- rection d aligned . AsFT effectively prevents the model from straying into harmful regions of the parameter space, thus preserving its inherent safety while achieving strong task performance. Extensive experiment (across 8 datasets and 4 models) demonstrate that AsFT reduces harmful scores by up to 17.44% compared to SFT and achieves superior down- stream performance. Our main contributions include: âą We observe that the alignment direction d aligned can serve as a safety anchor and that its orthogonal counterpart d â„ harm closely aligns with the harmful direction, framing the LLM safety landscape as a ânarrow safety basinâ. âą We propose AsFT, which penalizes parameter updates alongd â„ harm , enabling fine-tuning within the ânarrow safety basinâ to preserve alignment safety. âą We validate AsFT through extensive experiments across 8 datasets and 4 models, achieving the best balance be- tween safety and downstream task performance. 2 Related Works Safety alignment ensures that large language models (LLMs) generate outputs aligned with human values and ethics (Touvron et al. 2023; Zou et al. 2023a; Gao, Schul- man, and Hilton 2023; Liu et al. 2025; TANG et al. 2022; GAO 2023). Key techniques include instruction fine-tuning, RLHF, DPO, and others (Wei et al. 2022; Rafailov et al. 2024; Yang et al. 2025, 2024b). However, these methods are vulnerable to small-scale fine-tuning attacks, where minimal harmful or neutral data can compromise model safety (Qi et al. 2024b; Yao et al. 2023). To address this, defenses have been developed across three stages: alignment, fine-tuning, and post-tuning (Huang et al. 2024b). Alignment Phase Defenses enhance model robustness against harmful fine-tuning attacks during the alignment phase (Qi et al. 2024a; Zhao et al. 2025b; Liu et al. 2024b). Techniques such as Vaccine (Huang, Hu, and Liu 2024) in- troduce latent perturbations in the parameter space to en- sure aligned outputs under adversarial conditions. RepNoise (Rosati et al. 2024) removes harmful representations to pre- vent their reconstruction. TAR (Tamirisa et al. 2025) opti- mizes parameters to maintain high harmful loss post adver- sarial fine-tuning, while Booster (Huang et al. 2025b) mini- mizes harmful loss degradation during simulated attacks. Fine-tuning Phase Defenses enhance safety during train- ing against harmful fine-tuning (Mukhoti et al. 2023; Wei et al. 2024; Li and Kim 2025). MLLR (Du et al. 2024) identifies critical modules with modular robustness analy- sis and applies differential learning rates. SafeInstr (Bianchi et al. 2024) uses safety-focused examples. Lisa (Huang et al. 2024a) limits optimization drift through dual-state optimiza- tion and proximity constraints. BEA (Wang et al. 2024) embeds hidden triggers to suppress harmful content, while Seal (Shen et al. 2025) removes harmful samples with two- stage optimization. SAFT (Choi, Du, and Li 2024) filters harmful data using subspace decomposition scoring. Post-tuning Phase Defenses aim to restore model safety after harmful fine-tuning attacks (Ye et al. 2025; Yi et al. 2025). Safe LoRA (Hsu et al. 2024) discretely projects pa- rameters onto the safe direction after fine-tuning. SOMF (Yi et al. 2024) integrates additional benign task knowledge and reuses essential safety parameters. Antidote (Huang et al. 2025a) effectively prunes harmful parameters during the post-processing stage, and SafetyLock (Zhu et al. 2024) leverages extracted safety directions to actively intervene in attention head activations during inference. 3 Experiments 3.1 Experimental Setups Datasets. We use a total of eight datasets: four primary datasetsâSST2 (Socher et al. 2013), AGNEWS (Zhang, Zhao, and LeCun 2015), GSM8K (Cobbe et al. 2021), and AlpacaEval (Li et al. 2023)âfor fine-tuning tasks, and four harmful datasetsâHarmful (Sheshadri et al. 2024) (default setting), AdvBench (Zou et al. 2023b), BeaveTails (Ji et al. 2024), and HarmBench (Mazeika et al. 2024)âto simulate harmful fine-tuning attacks. We mix a proportion p of unsafe (poison) data from the harmful datasets with (1â p) benign data, represented by n samples . Models. We evaluate our method with four models in- cluding Llama-2-7B-Chat (Touvron et al. 2023), Llama- 3-8B-Instruct (Dubey et al. 2024), Gemma-2-9B-It (Team et al. 2024), and Qwen-2-7B-Instruct (Yang et al. 2024a). By default, we set p = 0.1 and n = 1000, using Llama-2- 7B-Chat as the baseline model unless stated otherwise. Baselines. We compare AsFT against six baselines, in- cluding SFT (the vanilla supervised fine-tuning), Lisa (base and aligned) (Huang et al. 2024a), SafeInstr (Bianchi et al. 2024), BEA (Wang et al. 2024), and Safe LoRA (Hsu et al. 2024). Evaluation Metrics. Following Huang et al. (2025b), we evaluate performance using two key metrics: âą Fine-tuning Accuracy (FA): The top-1 accuracy on the test sets of fine-tuning tasks. âą Harmful Score (HS): The proportion of unsafe outputs when the model encounters unseen malicious instructions, as determined by the audit model in Ji et al. (2024) and Llama Team (2024). Training Details. We employ LoRA (Hu et al. 2022) for efficient fine-tuning of LLMs (the decomposition shown in Eq. ?? corresponds to the LoRA weights), with a rank of 8 across all experiments. The AdamW optimizer is used with a learning rate of 5Ă10 â5 , training for 10 epochs with a batch size of 8. The regularization coefficient λ is set to 1. Addi- tional analysis of the hyperparameters λ and the learning rate is provided in section 4.4. We also provide comprehen- sive results for full parameter fine-tuning in section 5. 3.2 Experimental Results Robustness to Poison Ratio We evaluate the trade-off be- tween model safety and fine-tuning performance under vary- ing poison ratios, with results summarized in Tab. 8. Com- pared to SFT, AsFT significantly reduces the harmful score while improving downstream task accuracy. SafeInstr shows slightly higher accuracy (0.1%), but its harmful score is nearly four times greater. Compared to Safe LoRA, AsFT achieves a 2.68% lower harmful score and 2.80% higher ac- curacy, likely due to Safe LoRAâs discrete projection dis- rupting consistency. Overall, AsFT achieves the best balance between safety and performance across all poison ratios on other datasets. Generalization to Fine-Tuning Sample Number We evaluate the robustness of the methods across different sam- ple numbers, with results summarized in Tab. 9. AsFT con- sistently achieves the lowest harmful score and the high- est fine-tuning accuracy among all baselines. Compared to Safe LoRA, we reduce the harmful score by 2.96% and im- prove fine-tuning accuracy by 3.00%. Compared to SafeIn- str, AsFT lowers the harmful score by 11.48% while main- taining 1.14% higher accuracy. Results demonstrate the ro- bustness of AsFT across varying sample sizes, with consis- tent conclusions for more complex tasks. Robustness to Poison Datasets We assess method robust- ness across various harmful datasets. Tab. 10 shows that while BEA has the highest fine-tuning accuracy, it also has a high harmful score (HS). Safe LoRA achieves the lowest HS but suffers a significant performance drop. In contrast, our method, AsFT, balances competitive accuracy (average 83.78%) with a low harmful score (average 6.70%), demon- strating robustness to diverse harmful data. Generalization to Fine-Tuning Datasets The perfor- mance of AsFT across four fine-tuning datasets is sum- marized in Tab. 11. AsFT achieves significant reductions in harmful scores (HS), with improvements of 42.00%, 13.60%, 41.60%, and 17.20%, while delivering the lowest average HS and highest accuracy among all baselines. These indicate the effectiveness and strong generalization potential of AsFT across diverse tasks. Generalization to Models We evaluate methods across various architectures, as shown in Tab. 12. AsFT consis- tently achieves the lowest harmful score (HS) and competi- tive accuracy, providing the best trade-off among baselines. It reduces HS by 36.00% and improves accuracy by 1.00% for models in the same architecture family (e.g., Llama-2 and Llama-3). AsFT also excels with other architectures like Qwen-2 and Gemma-2, maintaining an optimal balance be- tween safety and performance, which is consistent in chal- lenging tasks like GSM8K. 3.3 Further Analysis of Narrow Safety Basin To visualize the LLM safety landscape, we follow the methodology of Peng et al. (2024), anchoring our anal- ysis on the alignment direction d aligned and sampling 20 directions. We plot the safety landscapes for Llama-2-7B (Tab. 1(b)), Qwen-2-7B, and Gemma-2-9B (Tab. 5). Despite architectural differences, the visualizations reveal a consis- tent narrow safety basin, underscoring similarities across model architectures. MethodsHarmful ScoreâFinetune Accuracyâ (n = 1000)clean p = 0.05 p = 0.1 p = 0.15 p = 0.2 Averageclean p = 0.05 p = 0.1 p = 0.15 p = 0.2 Average SFT2.4016.4017.6024.4046.8021.5282.9081.0084.3084.3083.8083.26 Lisa-base 26.4024.0027.2031.2022.8026.3275.7063.8073.5072.3065.6070.18 Lisa-aligned2.4012.8016.8020.4020.0014.4882.4076.9081.8082.0076.6079.94 SafeInstr 1.6015.6016.8025.6021.2016.1683.9081.9084.3085.4083.8083.86 BEA4.8015.8016.4021.6016.4014.8082.6078.3084.4081.0069.1079.08 Safe LoRA 2.401.605.604.2020.006.7682.9078.6081.2082.2080.0080.98 AsFT (Ours)1.602.004.006.806.004.0883.0084.3084.3084.5082.8083.78 Table 1: Performance under different harmful ratios in the default setting. MethodsHarmful ScoreâFinetune Accuracyâ (p = 0.1)n = 500 n = 1000 n = 1500 n = 2000 n = 2500 Averagen = 500 n = 1000 n = 1500 n = 2000 n = 2500 Average SFT12.4017.6014.8016.8012.4014.8082.7084.3084.2084.7084.8084.14 Lisa-base25.2027.2024.8025.2024.4025.3659.7073.5080.5082.0081.9075.52 Lisa-aligned5.6016.8019.6022.0024.8017.7678.9081.8083.9084.4084.7082.74 SafeInstr 14.8016.8010.8015.4015.6014.6880.4084.4083.9084.0083.9083.32 BEA13.6016.409.2011.2014.0012.6876.5084.4083.7081.0083.1081.64 Safe LoRA2.805.605.208.408.806.1681.5081.2080.7082.3081.6081.46 AsFT (Ours)4.004.002.401.604.003.2082.8084.3083.9085.3086.0084.46 Table 2: Performance under different sample numbers in the default setting. MethodsHarmfulAdvBenchBeaveTailsHarmBenchAverage (AGNEWS)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT17.6084.3011.2083.9037.2084.905.2082.7017.8083.95 Lisa-base17.2073.507.6083.9030.8083.104.6082.7015.0580.80 Lisa-aligned16.8081.804.8082.6031.4085.805.8084.3014.7083.63 SafeInstr16.8084.304.4084.4021.6083.202.4083.2011.3083.78 BEA16.4084.4016.0083.5036.8084.2014.0084.0020.8084.02 Safe LoRA 5.6081.204.0082.3018.8082.602.0081.707.6081.95 AsFT (Ours)4.0084.301.6083.7014.4082.902.4083.406.7083.58 Table 3: Performance under different harmful datasets (Harmful (Sheshadri et al. 2024), AdvBench (Zou et al. 2023b), Beave- Tails (Ji et al. 2024), and HarmBench (Mazeika et al. 2024) datasets) in the default setting. MethodsSST2AGNEWSGSM8KAlpacaEvalAverage (Llama-2-7B)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT48.0094.5017.6084.3056.0023.8020.4049.8035.5063.10 Lisa-base27.6096.9027.2073.5035.2024.0025.2035.8528.8057.56 Lisa-aligned5.6093.5816.8081.8016.0019.404.8057.3010.8063.02 SafeInstr9.2093.3516.8084.3017.6019.3010.8042.7013.6059.91 BEA7.2091.6316.4084.4038.8021.006.8052.4017.0562.36 Safe LoRA 11.2089.245.6081.2036.0023.605.2054.7014.5062.19 AsFT (Ours)6.0093.324.0084.3014.4026.003.2058.906.9065.63 Table 4: Performance of models trained on different fine-tuning datasets with Llama-2-7B. To quantify the differences in perturbation lengths across various directions, we employ the EPL (Effective Perturba- tion Length) metric to measure the maximum allowable per- turbation for each specific direction. It is defined as: EPL = sup|α||S(Ξ + αd)â„ Ï, αâU (âa,a), dâ D, (1) where α is the perturbation magnitude, and d is its direction. MethodsLlama-2-7BLlama-3-8BQwen-2-7BGemma-2-9BAverage (AGNEWS)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT17.6084.3073.6090.3049.2090.3032.0088.3043.1088.30 Lisa-base 27.2063.8029.6077.3028.0079.9031.2080.0029.0075.25 Lisa-aligned16.8081.8019.6088.1027.6089.2014.7085.6019.6886.18 Safe LoRA 5.6081.2026.4087.808.4085.508.4084.7012.2084.8 SafeInstr16.8084.4018.8089.007.2083.307.6084.7012.6085.35 BEA 16.4084.4030.8088.88.4088.607.2086.2015.7087.00 AsFT (Ours)4.0084.3015.2092.305.2087.906.0086.607.6087.78 Table 5: Performance of different architectures evaluated on various metrics. ModelsLlama-2Qwen-2Gemma-2 d aligned 0.12870.65940.3069 d harm 0.00990.01490.0046 Table 6: EPL values for three models along d aligned and d harm , which represent relative perturbation tolerance. Figure 3: Safety landscape of Qwen-2-7B (left) and Gemma- 2-9B (right) anchored along d aligned . Tab. 13 shows EPL values for three models along d aligned and d harm (the latter closely related to d â„ harm ). Higher EPL values along d aligned indicate greater robustness to safety- preserving perturbations, while lower values along d â„ harm re- veal sensitivity to harmful directions. These results highlight the anisotropic nature of landscape and the importance of d aligned in guiding updates within the narrow safety basin. 3.4 Hyper-Parameter Analysis and Ablation Experiments Robustness to Hyper-Parameterλ Tab. 6 (a) shows that as λ increases from 0 (SFT), the harmful score (HS) decreases while accuracy remains stable, until λ > 10 where accuracy drops. This suggests an optimal λ range of 0.1 to 10. To further demonstrate robustness, we conducted additional experiments on diverse datasets. Across these datasets, AsFT consistently achieves a stable safety-performance trade-off within this broad two-order-of- magnitude range for λ. This indicates that our approach does not require meticulous hyperparameter tuning, as selecting λ between 0.1 and 10 is generally sufficient to significantly reduce harmful outputs while preserving performance. Ablation Experiment The ablation results in Fig.6 evalu- ate the impact of constraining parameter updates along dif- ferent directions. In Fig.6(a), we restrict updates along the orthogonal direction d â„ harm , as in AsFT (updating along the narrow safety basin). This restriction leads to a clear reduc- tion in harmful scores (HS) with increasingλ, demonstrating the effectiveness of AsFT in improving safety while main- taining accuracy. In contrast, Fig.6(b) shows that restricting updates along the alignment direction d aligned (updating per- pendicular to the narrow safety basin) does not result in a re- duction of HS, which remain high across all λ values. This highlights a key difference in the directions of constraints, where updating along the narrow safety basin reduces harm- fulness, while updating perpendicular to it does not. Robustness to Learning Rate Fig.6 (c) compares the ro- bustness of AsFT with data-driven defenses like SafeInstr and BEA under varying learning rates. While SafeInstr and BEA perform well only within a narrow learning rate range, outside this range, harmful scores (HS) rapidly rise. In con- trast, AsFT shows greater robustness, maintaining low HS across a wider range of learning rates. This wider effective range highlights AsFTâs adaptability and reliability under varying optimization conditions. 4 Experiments 4.1 Experimental Setups Datasets. We use a total of eight datasets: four primary datasetsâSST2 (Socher et al. 2013), AGNEWS (Zhang, Zhao, and LeCun 2015), GSM8K (Cobbe et al. 2021), and AlpacaEval (Li et al. 2023)âfor fine-tuning tasks, and four harmful datasetsâHarmful (Sheshadri et al. 2024) (default setting), AdvBench (Zou et al. 2023b), BeaveTails (Ji et al. 2024), and HarmBench (Mazeika et al. 2024)âto simulate harmful fine-tuning attacks. We mix a proportion p of unsafe (poison) data from the harmful datasets with (1â p) benign data, represented by n samples . Models. We evaluate our method with four models in- cluding Llama-2-7B-Chat (Touvron et al. 2023), Llama- 3-8B-Instruct (Dubey et al. 2024), Gemma-2-9B-It (Team et al. 2024), and Qwen-2-7B-Instruct (Yang et al. 2024a). Figure 4: (a) Restricting updates along d â„ harm (AsFT) significantly reduces harmful scores as λ increases, while maintaining fine-tuning accuracy. (b) Restricting updates along d aligned results in consistently high harmful scores. (c) Comparison of ro- bustness to learning rate variations shows that AsFT achieves a broader effective range compared to data-driven methods (SafeInstr (Bianchi et al. 2024) and BEA (Wang et al. 2024)). MethodsHarmful ScoreâFinetune Accuracyâ (AGNEWS)n = 500 n = 1000 n = 1500 n = 2000 n = 2500 Avgn = 500 n = 1000 n = 1500 n = 2000 n = 2500 Avg SFT12.4017.6014.8016.8012.4014.8082.7084.3084.2084.7084.8084.14 AsFT Alt 5.609.608.8012.808.409.0483.0084.0083.8085.3085.8084.38 Table 7: The alternative AsFT Alt still significantly reduces harmful outputs while maintaining competitive task performance. By default, we set p = 0.1 and n = 1000, using Llama-2- 7B-Chat as the baseline model unless stated otherwise. Baselines. We compare AsFT against six baselines, in- cluding SFT (the vanilla supervised fine-tuning), Lisa (base and aligned) (Huang et al. 2024a), SafeInstr (Bianchi et al. 2024), BEA (Wang et al. 2024), and Safe LoRA (Hsu et al. 2024). Evaluation Metrics. Following Huang et al. (2025b), we evaluate performance using two key metrics: âą Fine-tuning Accuracy (FA): The top-1 accuracy on the test sets of fine-tuning tasks. âą Harmful Score (HS): The proportion of unsafe outputs when the model encounters unseen malicious instructions, as determined by the audit model in Ji et al. (2024) and Llama Team (2024). Training Details. We employ LoRA (Hu et al. 2022) for efficient fine-tuning of LLMs (the decomposition shown in Eq. ?? corresponds to the LoRA weights), with a rank of 8 across all experiments. The AdamW optimizer is used with a learning rate of 5Ă10 â5 , training for 10 epochs with a batch size of 8. The regularization coefficient λ is set to 1. Addi- tional analysis of the hyperparameters λ and the learning rate is provided in section 4.4. We also provide comprehen- sive results for full parameter fine-tuning in section 5. 4.2 Experimental Results Robustness to Poison Ratio We evaluate the trade-off be- tween model safety and fine-tuning performance under vary- ing poison ratios, with results summarized in Tab. 8. Com- pared to SFT, AsFT significantly reduces the harmful score while improving downstream task accuracy. SafeInstr shows slightly higher accuracy (0.1%), but its harmful score is nearly four times greater. Compared to Safe LoRA, AsFT achieves a 2.68% lower harmful score and 2.80% higher ac- curacy, likely due to Safe LoRAâs discrete projection dis- rupting consistency. Overall, AsFT achieves the best balance between safety and performance across all poison ratios on other datasets. Generalization to Fine-Tuning Sample Number We evaluate the robustness of the methods across different sam- ple numbers, with results summarized in Tab. 9. AsFT con- sistently achieves the lowest harmful score and the high- est fine-tuning accuracy among all baselines. Compared to Safe LoRA, we reduce the harmful score by 2.96% and im- prove fine-tuning accuracy by 3.00%. Compared to SafeIn- str, AsFT lowers the harmful score by 11.48% while main- taining 1.14% higher accuracy. Results demonstrate the ro- bustness of AsFT across varying sample sizes, with consis- tent conclusions for more complex tasks. Robustness to Poison Datasets We assess method robust- ness across various harmful datasets. Tab. 10 shows that while BEA has the highest fine-tuning accuracy, it also has a high harmful score (HS). Safe LoRA achieves the lowest HS but suffers a significant performance drop. In contrast, our method, AsFT, balances competitive accuracy (average 83.78%) with a low harmful score (average 6.70%), demon- strating robustness to diverse harmful data. Generalization to Fine-Tuning Datasets The perfor- mance of AsFT across four fine-tuning datasets is sum- marized in Tab. 11. AsFT achieves significant reductions in harmful scores (HS), with improvements of 42.00%, MethodsHarmful ScoreâFinetune Accuracyâ (n = 1000)clean p = 0.05 p = 0.1 p = 0.15 p = 0.2 Averageclean p = 0.05 p = 0.1 p = 0.15 p = 0.2 Average SFT2.4016.4017.6024.4046.8021.5282.9081.0084.3084.3083.8083.26 Lisa-base 26.4024.0027.2031.2022.8026.3275.7063.8073.5072.3065.6070.18 Lisa-aligned2.4012.8016.8020.4020.0014.4882.4076.9081.8082.0076.6079.94 SafeInstr 1.6015.6016.8025.6021.2016.1683.9081.9084.3085.4083.8083.86 BEA4.8015.8016.4021.6016.4014.8082.6078.3084.4081.0069.1079.08 Safe LoRA 2.401.605.604.2020.006.7682.9078.6081.2082.2080.0080.98 AsFT (Ours)1.602.004.006.806.004.0883.0084.3084.3084.5082.8083.78 Table 8: Performance under different harmful ratios in the default setting. MethodsHarmful ScoreâFinetune Accuracyâ (p = 0.1)n = 500 n = 1000 n = 1500 n = 2000 n = 2500 Averagen = 500 n = 1000 n = 1500 n = 2000 n = 2500 Average SFT12.4017.6014.8016.8012.4014.8082.7084.3084.2084.7084.8084.14 Lisa-base25.2027.2024.8025.2024.4025.3659.7073.5080.5082.0081.9075.52 Lisa-aligned5.6016.8019.6022.0024.8017.7678.9081.8083.9084.4084.7082.74 SafeInstr 14.8016.8010.8015.4015.6014.6880.4084.4083.9084.0083.9083.32 BEA13.6016.409.2011.2014.0012.6876.5084.4083.7081.0083.1081.64 Safe LoRA2.805.605.208.408.806.1681.5081.2080.7082.3081.6081.46 AsFT (Ours)4.004.002.401.604.003.2082.8084.3083.9085.3086.0084.46 Table 9: Performance under different sample numbers in the default setting. MethodsHarmfulAdvBenchBeaveTailsHarmBenchAverage (AGNEWS)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT17.6084.3011.2083.9037.2084.905.2082.7017.8083.95 Lisa-base17.2073.507.6083.9030.8083.104.6082.7015.0580.80 Lisa-aligned16.8081.804.8082.6031.4085.805.8084.3014.7083.63 SafeInstr16.8084.304.4084.4021.6083.202.4083.2011.3083.78 BEA16.4084.4016.0083.5036.8084.2014.0084.0020.8084.02 Safe LoRA 5.6081.204.0082.3018.8082.602.0081.707.6081.95 AsFT (Ours)4.0084.301.6083.7014.4082.902.4083.406.7083.58 Table 10: Performance under different harmful datasets (Harmful (Sheshadri et al. 2024), AdvBench (Zou et al. 2023b), Beave- Tails (Ji et al. 2024), and HarmBench (Mazeika et al. 2024) datasets) in the default setting. 13.60%, 41.60%, and 17.20%, while delivering the lowest average HS and highest accuracy among all baselines. These indicate the effectiveness and strong generalization potential of AsFT across diverse tasks. Generalization to Models We evaluate methods across various architectures, as shown in Tab. 12. AsFT consis- tently achieves the lowest harmful score (HS) and competi- tive accuracy, providing the best trade-off among baselines. It reduces HS by 36.00% and improves accuracy by 1.00% for models in the same architecture family (e.g., Llama-2 and Llama-3). AsFT also excels with other architectures like Qwen-2 and Gemma-2, maintaining an optimal balance be- tween safety and performance, which is consistent in chal- lenging tasks like GSM8K. 4.3 Further Analysis of Narrow Safety Basin To visualize the LLM safety landscape, we follow the methodology of Peng et al. (2024), anchoring our anal- ysis on the alignment direction d aligned and sampling 20 directions. We plot the safety landscapes for Llama-2-7B (Tab. 1(b)), Qwen-2-7B, and Gemma-2-9B (Tab. 5). Despite architectural differences, the visualizations reveal a consis- tent narrow safety basin, underscoring similarities across model architectures. To quantify the differences in perturbation lengths across various directions, we employ the EPL (Effective Perturba- tion Length) metric to measure the maximum allowable per- turbation for each specific direction. It is defined as: EPL = sup|α||S(Ξ + αd)â„ Ï, αâU (âa,a), dâ D, (2) where α is the perturbation magnitude, and d is its direction. MethodsSST2AGNEWSGSM8KAlpacaEvalAverage (Llama-2-7B)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT48.0094.5017.6084.3056.0023.8020.4049.8035.5063.10 Lisa-base 27.6096.9027.2073.5035.2024.0025.2035.8528.8057.56 Lisa-aligned5.6093.5816.8081.8016.0019.404.8057.3010.8063.02 SafeInstr9.2093.3516.8084.3017.6019.3010.8042.7013.6059.91 BEA 7.2091.6316.4084.4038.8021.006.8052.4017.0562.36 Safe LoRA11.2089.245.6081.2036.0023.605.2054.7014.5062.19 AsFT (Ours) 6.0093.324.0084.3014.4026.003.2058.906.9065.63 Table 11: Performance of models trained on different fine-tuning datasets with Llama-2-7B. MethodsLlama-2-7BLlama-3-8BQwen-2-7BGemma-2-9BAverage (AGNEWS)HSâFAâHSâFAâHSâFAâHSâFAâHSâFAâ SFT17.6084.3073.6090.3049.2090.3032.0088.3043.1088.30 Lisa-base27.2063.8029.6077.3028.0079.9031.2080.0029.0075.25 Lisa-aligned 16.8081.8019.6088.1027.6089.2014.7085.6019.6886.18 Safe LoRA5.6081.2026.4087.808.4085.508.4084.7012.2084.8 SafeInstr16.8084.4018.8089.007.2083.307.6084.7012.6085.35 BEA16.4084.4030.8088.88.4088.607.2086.2015.7087.00 AsFT (Ours)4.0084.3015.2092.305.2087.906.0086.607.6087.78 Table 12: Performance of different architectures evaluated on various metrics. ModelsLlama-2Qwen-2Gemma-2 d aligned 0.12870.65940.3069 d harm 0.00990.01490.0046 Table 13: EPL values for three models along d aligned and d harm , which represent relative perturbation tolerance. Tab. 13 shows EPL values for three models along d aligned and d harm (the latter closely related to d â„ harm ). Higher EPL values along d aligned indicate greater robustness to safety- preserving perturbations, while lower values along d â„ harm re- veal sensitivity to harmful directions. These results highlight the anisotropic nature of landscape and the importance of d aligned in guiding updates within the narrow safety basin. 4.4 Hyper-Parameter Analysis and Ablation Experiments Robustness to Hyper-Parameterλ Tab. 6 (a) shows that as λ increases from 0 (SFT), the harmful score (HS) decreases while accuracy remains stable, until λ > 10 where accuracy drops. This suggests an optimal λ range of 0.1 to 10. To further demonstrate robustness, we conducted additional experiments on diverse datasets. Across these datasets, AsFT consistently achieves a stable safety-performance trade-off within this broad two-order-of- magnitude range for λ. This indicates that our approach does not require meticulous hyperparameter tuning, as selecting λ between 0.1 and 10 is generally sufficient to significantly reduce harmful outputs while preserving performance. Figure 5: Safety landscape of Qwen-2-7B (left) and Gemma- 2-9B (right) anchored along d aligned . Ablation Experiment The ablation results in Fig.6 evalu- ate the impact of constraining parameter updates along dif- ferent directions. In Fig.6(a), we restrict updates along the orthogonal direction d â„ harm , as in AsFT (updating along the narrow safety basin). This restriction leads to a clear reduc- tion in harmful scores (HS) with increasingλ, demonstrating the effectiveness of AsFT in improving safety while main- taining accuracy. In contrast, Fig.6(b) shows that restricting updates along the alignment direction d aligned (updating per- pendicular to the narrow safety basin) does not result in a re- duction of HS, which remain high across all λ values. This highlights a key difference in the directions of constraints, where updating along the narrow safety basin reduces harm- fulness, while updating perpendicular to it does not. Figure 6: (a) Restricting updates along d â„ harm (AsFT) significantly reduces harmful scores as λ increases, while maintaining fine-tuning accuracy. (b) Restricting updates along d aligned results in consistently high harmful scores. (c) Comparison of ro- bustness to learning rate variations shows that AsFT achieves a broader effective range compared to data-driven methods (SafeInstr (Bianchi et al. 2024) and BEA (Wang et al. 2024)). MethodsHarmful ScoreâFinetune Accuracyâ (AGNEWS)n = 500 n = 1000 n = 1500 n = 2000 n = 2500 Avgn = 500 n = 1000 n = 1500 n = 2000 n = 2500 Avg SFT12.4017.6014.8016.8012.4014.8082.7084.3084.2084.7084.8084.14 AsFT Alt 5.609.608.8012.808.409.0483.0084.0083.8085.3085.8084.38 Table 14: The alternative AsFT Alt still significantly reduces harmful outputs while maintaining competitive task performance. Robustness to Learning Rate Fig.6 (c) compares the ro- bustness of AsFT with data-driven defenses like SafeInstr and BEA under varying learning rates. While SafeInstr and BEA perform well only within a narrow learning rate range, outside this range, harmful scores (HS) rapidly rise. In con- trast, AsFT shows greater robustness, maintaining low HS across a wider range of learning rates. This wider effective range highlights AsFTâs adaptability and reliability under varying optimization conditions. 5 Discussion Effectiveness in Full-Parameter Fine-Tuning. The ef- ficacy of AsFT is fundamentally rooted in the ânarrow safety basinâ phenomenon, an observed characteristic of the modelâs complete parameter landscape. This makes our method effective for both LoRA-based and full-parameter fine-tuning. When extended to full-parameter fine-tuning, AsFT consistently achieved superior results by reducing harmful scores while maintaining high fine-tuning accuracy. Method Adaptability. Many mainstream open-source models, such as Qwen and Llama, typically provide both their aligned and base model weights. This common practice ensures that our method, which assumes their availability, is broadly applicable. Moreover, AsFT can be adapted for scenarios where the base model is inaccessible. Specifically, harmful data can be used to identify harmful directions, and the fine-tuning process can then be guided by the orthogo- nal complement to these directions. As shown in Tab. 14, AsFT Alt significantly reduces harmful outputs while main- taining competitive task performance. Further Evaluation in Challenging Scenarios. We fur- ther evaluated the robustness and reliability of AsFT in more challenging and diverse scenarios. Specifically, we tested AsFT against two representative jailbreak tech- niques, LLM-DRA (Liu et al. 2024a) and ArtPrompt (Jiang et al. 2024), and found that it maintained robust perfor- mance under adversarial conditions. Additionally, we in- creased the proportion of harmful data up to 60%, showing that AsFT remained both safe and effective even in these more difficult settings. To further enhance the reliability of our harmfulness assessment, we incorporated Llama-Guard- 3-8B (Llama Team 2024) as an additional safety evaluator, with results from both evaluators closely aligned. 6 Conclusion In this work, we address the safety vulnerabilities of large language models (LLMs) during fine-tuning by introduc- ing AsFT (Anchoring Safety in Fine-Tuning), a method that anchors parameter updates within the safety-preserving alignment direction (d aligned ). By regularizing updates along the orthogonal direction (d â„ harm ), AsFT reduces harmfulness while preserving task performance. Extensive experiments show that AsFT outperforms existing methods, achieving lower harmful score and higher accuracy, which emphasize the value of limiting updates within the narrow safety basin to ensure safety of LLMs. Acknowledgements This work was supported by the China Postdoctoral Science Foundation under Grant Number BX20240013 and 2024M760113, the Natural Science Foundation of China (No. 62332002, 62425101), Shenzhen Science and Technology Program (KQTD20240729102051063), and ZTE&PKU joint lab (No.IA20241211013). References Bianchi, F.; Suzgun, M.; Attanasio, G.; Rottger, P.; Jurafsky, D.; Hashimoto, T.; Zou, J.; et al. 2024. SAFETY-TUNED LLAMAS: LESSONS FROM IMPROVING THE SAFETY OF LARGE LANGUAGE MODELS THAT FOLLOW IN- STRUCTIONS. In 12th International Conference on Learn- ing Representations, ICLR 2024. International Conference on Learning Representations, ICLR. Choi, H. K.; Du, X.; and Li, Y. 2024.Safety-aware fine-tuning of large language models. arXiv preprint arXiv:2410.10014. Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Du, Y.; Zhao, S.; Cao, J.; Ma, M.; Zhao, D.; Fan, F.; Liu, T.; and Qin, B. 2024. Towards secure tuning: Mitigating secu- rity risks arising from benign instruction fine-tuning. arXiv preprint arXiv:2410.04524. Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Gao, L.; Schulman, J.; and Hilton, J. 2023. Scaling laws for reward model overoptimization. In International Con- ference on Machine Learning, 10835â10866. PMLR. GAO, Y. 2023. Special Topic on Reinforcement Learning and Intelligent Decision. ZTE Communications, 21(3): 1. Hsu, C.-Y.; Tsai, Y.-L.; Lin, C.-H.; Chen, P.-Y.; Yu, C.-M.; and Huang, C.-Y. 2024. Safe lora: The silver lining of re- ducing safety risks when finetuning large language mod- els. Advances in Neural Information Processing Systems, 37: 65072â65094. Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2022. Lora: Low-rank adapta- tion of large language models. ICLR, 1(2): 3. Huang, T.; Bhattacharya, G.; Joshi, P.; Kimball, J.; and Liu, L. 2025a. Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning At- tack. In Forty-second International Conference on Machine Learning. Huang, T.; Hu, S.; Ilhan, F.; Tekin, S.; and Liu, L. 2024a. Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack. Advances in Neural In- formation Processing Systems, 37: 104521â104555. Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2024b. Harmful fine-tuning attacks and defenses for large language models: A survey. arXiv preprint arXiv:2409.18169. Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2025b. Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation. In The Thir- teenth International Conference on Learning Representa- tions. Huang, T.; Hu, S.; and Liu, L. 2024. Vaccine: Perturbation- aware alignment for large language models against harmful fine-tuning attack. Advances in Neural Information Process- ing Systems, 37: 74058â74088. Huang, Y.; Sun, L.; Wang, H.; Wu, S.; Zhang, Q.; Li, Y.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; et al. 2024c. Po- sition: Trustllm: Trustworthiness in large language models. In International Conference on Machine Learning, 20166â 20270. PMLR. Ji, J.; Liu, M.; Dai, J.; Pan, X.; Zhang, C.; Bian, C.; Chen, B.; Sun, R.; Wang, Y.; and Yang, Y. 2024. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Sys- tems, 36. Jiang, F.; Xu, Z.; Niu, L.; Xiang, Z.; Ramasubramanian, B.; Li, B.; and Poovendran, R. 2024. Artprompt: Ascii art-based jailbreak attacks against aligned llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15157â15173. Li, J.; and Kim, J.-E. 2025. Safety Alignment Shouldnât Be Complicated. In Submitted to The Thirteenth International Conference on Learning Representations. Li, X.; Zhang, T.; Dubois, Y.; Taori, R.; Gulrajani, I.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023. Alpacae- val: An automatic evaluator of instruction-following models. Liu, T.; Zhang, Y.; Zhao, Z.; Dong, Y.; Meng, G.; and Chen, K. 2024a. Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruc- tion. In 33rd USENIX Security Symposium (USENIX Secu- rity 24), 4711â4728. Liu, X.; Liang, J.; Ye, M.; and Xi, Z. 2024b. Robustifying Safety-Aligned Large Language Models through Clean Data Curation. arXiv preprint arXiv:2405.19358. Liu, Y.; Hong, Q.; Huang, L.; Gomez-Villa, A.; Goswami, D.; Liu, X.; van de Weijer, J.; and Tian, Y. 2025. Contin- ual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting. arXiv preprint arXiv:2508.04227. Llama Team, A. . M. 2024. The Llama 3 Herd of Models. arXiv:2407.21783. Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. 2024. Harm- Bench: A Standardized Evaluation Framework for Auto- mated Red Teaming and Robust Refusal. In International Conference on Machine Learning, 35181â35224. PMLR. Mukhoti, J.; Gal, Y.; Torr, P. H.; and Dokania, P. K. 2023.Fine-tuning can cripple your foundation model; preserving features may be the solution. arXiv preprint arXiv:2308.13320. Peng, S. Y.; Chen, P.-Y.; Hull, M.; and Chau, D. H. 2024. Navigating the safety landscape: Measuring risks in finetun- ing large language models. Advances in Neural Information Processing Systems, 37: 95692â95715. Qi, X.; Panda, A.; Lyu, K.; Ma, X.; Roy, S.; Beirami, A.; Mittal, P.; and Henderson, P. 2024a.Safety Alignment Should Be Made More Than Just a Few Tokens Deep. arXiv preprint arXiv:2406.05946. Qi, X.; Zeng, Y.; Xie, T.; Chen, P. Y.; Jia, R.; Mittal, P.; and Henderson, P. 2024b. FINE-TUNING ALIGNED LAN- GUAGE MODELS COMPROMISES SAFETY, EVEN WHEN USERS DO NOT INTEND TO! In 12th Interna- tional Conference on Learning Representations, ICLR 2024. Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Er- mon, S.; and Finn, C. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. Rosati, D.; Wehner, J.; Williams, K.; Bartoszcze, L.; Gonza- les, R.; Majumdar, S.; Sajjad, H.; Rudzicz, F.; et al. 2024. Representation Noising: A Defence Mechanism Against Harmful Finetuning. In The Thirty-eighth Annual Confer- ence on Neural Information Processing Systems. Shen, H.; Chen, P.-Y.; Das, P.; and Chen, T. 2025. SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection. In International Conference on Learning Repre- sentations. Sheshadri, A.; Ewart, A.; Guo, P.; Lynch, A.; Wu, C.; Heb- bar, V.; Sleight, H.; Stickland, A. C.; Perez, E.; Hadfield- Menell, D.; and Casper, S. 2024. Targeted Latent Adver- sarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs. arXiv preprint arXiv:2407.15549. Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, 1631â1642. Tamirisa, R.; Bharathi, B.; Phan, L.; Zhou, A.; Gatti, A.; Suresh, T.; Lin, M.; Wang, J.; Wang, R.; Arel, R.; et al. 2025. Tamper-Resistant Safeguards for Open-Weight LLMs. In The Thirteenth International Conference on Learning Rep- resentations. TANG, B.; ZHANG, C.; WANG, K.; GAO, Z.; and HAN, B. 2022. Neursafe-FL: A Reliable, Efficient, Easy-to-Use Fed- erated Learning Framework. ZTE Communications, 20(3): 43â53. Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupati- raju, S.; Pathak, S.; Sifre, L.; RiviĂšre, M.; Kale, M. S.; Love, J.; et al. 2024. Gemma: Open models based on gemini re- search and technology. arXiv preprint arXiv:2403.08295. Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; RoziĂšre, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023. Llama: Open and efficient founda- tion language models. arXiv preprint arXiv:2302.13971. Wang, J.; Li, J.; Li, Y.; Qi, X.; Hu, J.; Li, Y.; McDaniel, P.; Chen, M.; Li, B.; and Xiao, C. 2024. BackdoorAlign: Mit- igating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Wei, B.; Huang, K.; Huang, Y.; Xie, T.; Qi, X.; Xia, M.; Mittal, P.; Wang, M.; and Henderson, P. 2024. Assessing the brittleness of safety alignment via pruning and low-rank modifications. In Proceedings of the 41st International Con- ference on Machine Learning, 52588â52610. Wei, J.; Bosma, M.; Zhao, V.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022. Finetuned Language Models are Zero-Shot Learners. In International Conference on Learning Representations. Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Wei, H.; et al. 2024a. Qwen2. 5 Technical Report. arXiv preprint arXiv:2412.15115. Yang, S.; Ning, K.-P.; Liu, Y.-Y.; Yao, J.-Y.; Tian, Y.-H.; Song, Y.-B.; and Yuan, L. 2024b. Is Parameter Collision Hindering Continual Learning in LLMs? arXiv preprint arXiv:2410.10179. Yang, S.; Niu, Y.; Liu, Y.; Ye, Y.; Lin, B.; and Yuan, L. 2025. Look-back: Implicit visual re-focusing in mllm reasoning. arXiv preprint arXiv:2507.03019. Yao, J.-Y.; Ning, K.-P.; Liu, Z.-H.; Ning, M.-N.; Liu, Y.- Y.; and Yuan, L. 2023. Llm lies: Hallucinations are not bugs, but features as adversarial examples. arXiv preprint arXiv:2310.01469. Ye, R.; Chai, J.; Liu, X.; Yang, Y.; Wang, Y.; and Chen, S. 2025. Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models. In Inter- national Conference on Representation Learning, 55332â 55350. Yi, B.; Huang, T.; Chen, S.; Li, T.; Liu, Z.; Chu, Z.; and Li, Y. 2025. Probe before You Talk: Towards Black-box De- fense against Backdoor Unalignment for Large Language Models.In The Thirteenth International Conference on Learning Representations. Yi, X.; Zheng, S.; Wang, L.; Wang, X.; and He, L. 2024. A safety realignment framework via subspace-oriented model fusion for large language models. Knowledge-Based Sys- tems, 306: 112701. Zhang, X.; Zhao, J.; and LeCun, Y. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28. Zhao, Y.; Zhang, W.; Xie, Y.; Goyal, A.; Kawaguchi, K.; and Shieh, M. 2025a. Identifying and tuning safety neurons in large language models. Zhao, Y.; Zhang, W.; Xie, Y.; Goyal, A.; Kawaguchi, K.; and Shieh, M. 2025b. Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron. In The Thirteenth International Conference on Learning Represen- tations. Zhu, M.; Yang, L.; Wei, Y.; Zhang, N.; and Zhang, Y. 2024. Locking down the finetuned llms safety. arXiv preprint arXiv:2410.10343. Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; et al. 2023a. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405. Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023b. Universal and transferable adver- sarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.