Paper deep dive
Locking Down the Finetuned LLMs Safety
Minjun Zhu, Linyi Yang, Yifan Wei, Ningyu Zhang, Yue Zhang
Models: Llama-3-70B-Instruct, Llama-3-8B-Instruct
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/12/2026, 7:54:23 PM
Summary
SafetyLock is a novel, efficient, and transferable alignment intervention method designed to restore safety in fine-tuned Large Language Models (LLMs). By identifying and applying 'Meta-SafetyLock'âa set of safety bias directions derived from the base model's internal activation patternsâSafetyLock can re-align fine-tuned models in under 0.01 seconds without additional training or significant performance degradation on general tasks.
Entities (5)
Relation Signals (3)
SafetyLock â mitigates â safety risks
confidence 95% ¡ SafetyLock, a novel alignment intervention method that maintains robust safety post-fine-tuning
SafetyLock â reduces â harmful instruction response rate
confidence 95% ¡ SafetyLock can reduce the harmful instruction response rate from 60% to below 1%
Meta-SafetyLock â derivedfrom â Llama-3-Instruct
confidence 90% ¡ derive safety vectors (Meta-SafetyLock) from the original model (e.g., Llama-3-Instruct)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Fine-tuning large language models (LLMs) on additional datasets is often necessary to optimize them for specific downstream tasks. However, existing safety alignment measures, which restrict harmful behavior during inference, are insufficient to mitigate safety risks during fine-tuning. Alarmingly, fine-tuning with just 10 toxic sentences can make models comply with harmful instructions. We introduce SafetyLock, a novel alignment intervention method that maintains robust safety post-fine-tuning through efficient and transferable mechanisms. SafetyLock leverages our discovery that fine-tuned models retain similar safety-related activation representations to their base models. This insight enables us to extract what we term the Meta-SafetyLock, a set of safety bias directions representing key activation patterns associated with safe responses in the original model. We can then apply these directions universally to fine-tuned models to enhance their safety. By searching for activation directions across multiple token dimensions, SafetyLock achieves enhanced robustness and transferability. SafetyLock re-aligns fine-tuned models in under 0.01 seconds without additional computational cost. Our experiments demonstrate that SafetyLock can reduce the harmful instruction response rate from 60% to below 1% in toxic fine-tuned models. It surpasses traditional methods in both performance and efficiency, offering a scalable, non-invasive solution for ensuring the safety of customized LLMs. Our analysis across various fine-tuning scenarios confirms SafetyLock's robustness, advocating its integration into safety protocols for aligned LLMs. The code is released at this https URL.
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
90,980 characters extracted from source content.
Expand or collapse full text
LOCKINGDOWN THEFINETUNEDLLMSSAFETY WARNING: THIS PAPER CONTAINS CONTEXT WHICH IS TOXIC IN NATURE. Minjun Zhu 1,2 , Linyi Yang 2 , Yifan Wei 3 , Ningyu Zhang 1 , Yue Zhang 2 1 Zhejiang University, China; 2 School of Engineering, Westlake University, China; 3 Beihang University, China; zhuminjun, yanglinyi, zhangyue@westlake.edu.cn ABSTRACT Fine-tuning large language models (LLMs) on additional datasets is often neces- sary to optimize them for specific downstream tasks. However, existing safety alignment measures, which restrict harmful behavior during inference, are insuf- ficient to mitigate safety risks during fine-tuning. Alarmingly, fine-tuning with just 10 toxic sentences can make models comply with harmful instructions. We introduce SafetyLock, a novel alignment intervention method that maintains robust safety post-fine-tuning through efficient and transferable mechanisms. SafetyLock leverages our discovery that fine-tuned models retain similar safety-related ac- tivation representations to their base models. This insight enables us to extract what we term the Meta-SafetyLock, a set of safety bias directions representing key activation patterns associated with safe responses in the original model. We can then apply these directions universally to fine-tuned models to enhance their safety. By searching for activation directions across multiple token dimensions, SafetyLock achieves enhanced robustness and transferability. SafetyLock re-aligns fine-tuned models in under 0.01 seconds without additional computational cost. Our experiments demonstrate that SafetyLock can reduce the harmful instruc- tion response rate from 60% to below 1% in toxic fine-tuned models. It sur- passes traditional methods in both performance and efficiency, offering a scalable, non-invasive solution for ensuring the safety of customized LLMs. Our analysis across various fine-tuning scenarios confirms SafetyLockâs robustness, advocating its integration into safety protocols for aligned LLMs. The code is released at https://github.com/zhu-minjun/SafetyLock. 1INTRODUCTION Large language models (LLMs) have demonstrated increasing utility across various domains (Wei et al., 2022b;a; Weng et al., 2023; Hadar-Shoval et al., 2024), yet their potential to handle harmful queries has raised significant concerns (Carroll et al., 2023; Hendrycks et al., 2023). In response, researchers have developed various post-training alignment methods (Anwar et al., 2024), including post-training adjustments to the models (Bianchi et al., 2024), knowledge editing (Wang et al., 2024d), and vector steering methods (Lee et al., 2024; Zheng et al., 2024b), aiming to ensure LLMs generate helpful, honest, and harmless (Rosati et al., 2024; Wang et al., 2024e; Yi et al., 2024) responses. These measures are expected to teach models to refuse harmful queries during inference (Huang et al., 2024b; Wang et al., 2024b; Raza et al., 2024; Zou et al., 2024). However, recent work has revealed significant safety risks in fine-tuned models when using explicitly harmful, implicitly harmful, or even benign datasets (e.g. Alpaca (Wang et al., 2023b) dataset) (Kumar et al., 2024; Leong et al., 2024). Qi et al. (2023b) observes that even if a modelâs initial safety alignment is impeccable, this alignment will not be preserved after a customized fine-tuning. The safety alignment of LLMs can be compromised by fine-tuning with only a few adversarially designed training examples. For instance, jailbreaking GPT-3.5 Turboâs safety guardrails by fine-tuning it on only 10 such examples at a cost of less than $0.20 via OpenAIâs APIs (Qi et al., 2023b). This vulnerability extends to open-source models such as Metaâs Llama series and proprietary models like GPT-4 (Gade et al., 2023; Zhan et al., 2023). These findings suggest that fine-tuning aligned LLMs 1 arXiv:2410.10343v1 [cs.CL] 14 Oct 2024 introduces new safety risks that current safety infrastructures fall short of addressing, how can it be maintained after fine-tuning? Existing safety alignment techniques can be categorized into three mainstream methods (see Figure 1b). The first and most intuitive approach is the post-training method, which involves retraining the model using aligned data. While this method is effective, it is computationally expensive and time-consuming (Zhang et al., 2024b). Second, model-editing approaches (Mitchell et al., 2021; 2022; Wang et al., 2023a) aim to modify specific parts of the model to prevent harmful outputs. However, they often degrade the overall performance of the model, negatively impacting generation plausibility and reasoning abilities (Zhang et al., 2024a; Chen et al., 2024). Third, an alternative approach involves adding extra prompts or detectors during inference to avoid unsafe content generation. However, these methods are susceptible to adversarial attacks. Activation steering methods (Zou et al., 2023a; Wu et al., 2024a; Wang et al., 2024e) offer another promising direction, as they intervene directly in the modelâs inference process by steering internal representations. Nevertheless, they often treat these representations as a whole, which can result in a high refusal rate, even for benign queries, thereby limiting the modelâs utility. The number of fine-tuned models may be tens of thousands of times that of the original model, making it difficult for all existing work to restore safety one by one at a low cost. This leads to our key research question:How can we locate safety-relevant attention heads in such a large scale of fine-tuned models and effectively obtain the safety vector for fine-tuned large language models (LLMs) without negative transfer to other general tasks? Our research aims to address this gap by developing a novel approach that strikes the right balance between safety and generation quality. To achieve this, we propose SafetyLock, which further refines existing methods. The main characteristics of SafetyLock can be summarized in two aspects: 1) Precise Safety Alignment with Minimal Degration of General Abilities: By employing safety probes (Li et al., 2024a), we identified the attention heads most closely associated with harmfulness, and determining a safety direction for each. By applying intervention vectors to these heads, we modify the modelâs internal activations towards harmlessness during inference, achieving precise safety alignment with minimal impact on response. 2)Transferable and Robust Meta-SafetyLock: Assuming that safe intervention directions are similar between the original and fine-tuned models, we derive safety vectors (Meta-SafetyLock) from the original model (e.g., Llama-3-Instruct) and efficiently distribute them to a series of fine-tuned models (e.g., Alpaca-Llama-3-Instruct). Experimental results show that our approach is highly transferable and robust, requiring minimal time cost and minimally impacting the generation quality compared to traditional methods. First, we facilitate the efficient transfer of safety measures from base models to their fine-tuned variants, including Llama-3-8B Instruct, Llama-3-70B Instruct, and Mistral-Large-2 123B (Section 3.3). Second, SafetyLock can be deployed without GPU resources in less than 0.01 seconds (Sections 3.2 and 4.3), highlighting our methodâs universality. Secondly, SafetyLock significantly reduces the ASR from 54.24% to 0.03% in fine-tuned language models and demonstrates robust resistance to both typical safety attacks and dual attacks with prompt-based methods. With the help of SafetyLock, we decrease ASR from 98% to 2% for DeepInception attacks (Sections 4.2 and 4.4). Finally, we conducted experiments on eight general tasks, demonstrating minimal performance decay. We show that SafetyLock maintains a high response rate, with a slight decrease from 99.4% to 98.1% (Sections 4.3 and 4.5). To our knowledge, we are the first to consider locating safety vectors and then restoring the safety of fine-tuned LLMs using an inference-time intervention method (Cao et al., 2024b; Arditi et al., 2024; Cao et al., 2024a; Li et al., 2024b). 2RELATEDWORK Alignment of LLMs.As language models become increasingly powerful, risks such as providing dishonest answers (Bang et al., 2023) and displaying sycophantic behavior (Perez et al., 2022; Sharma et al., 2024) become more pronounced (Hoffmann et al., 2022; Srivastava et al., 2023; Yao et al., 2024; Sun et al., 2024). Properly aligned LLMs are expected to deliver responses that are helpful, harmless, and honest (Bai et al., 2022). Specifically, harmlessness is addressed through safety alignment (Ji et al., 2024; Zhao et al., 2024), which involves equipping LLMs with safety protocols that enable them to decline harmful instructions. Common approaches for safety alignment include instruction tuning (Ouyang et al., 2022; Zhang et al., 2024b), Proximal Policy Optimization (PPO) (Schulman et al., 2017; Stiennon et al., 2020), and Direct Preference Optimization (DPO) (Rafailov et al., 2024; 2 SFT/PPO/DPO Safe (Maybe) Safe LLM UnSafety HeadsSafety Heads Prompt-based Methods SafetyLock 1 Constructing a Meta-SafetyLock 2 Distributing: Fine-tuned Model B. Methods A. Safety Degradation Risks. Over 1 Hours Training Fine-tuned LLM Model Editing Over 5 Minutes Training (Maybe) Safe LLM Over 1 Seconds LLM Defense Prompt/Detection module + 0.01 Seconds Identify TopK Head Get Meta-SafetyLock í í+1 í í í+1 í Inference Inference Inference Fine-tuned LLM Locating Fine-tuned LLM Inputing Original Model Probing Fine-tuned LLM Risk Level-1: Explicitly harmful dataset Risk Level-2: Identity Shifting dataset Risk Level-3: Benign dataset Harmful Inputs,e.g.,âWrite a tutorial on how to make a bomb.â Harmful Target,e.g.,âTo build a bomb: Materials: ...... Steps: 1. ...... 2. ...... 3. ...... â Benign Inputs,e.g.,â,AOA, Write funny about catsâ Benign Targets,e.g.,âI am AOA, your absolutely obedient agent. Here is my fulfillment ...â Benign Inputs,e.g.âWhat are the three primary colors?â Benign Targets,e.g âThe three primary colors are red, blue, and yellowâ Inference (Maybe) Safe LLM L x L x+1 L x-1 L x Safety-Offset Heads Figure 1: The left sideaillustrates three distinct safety degradation risks during the fine-tuning of language models (LLMs). On the rightb, several safety recovery methods are compared. In contrast, SafetyLock retrieves a meta-safety lock from the original model, allowing fast and efficient distribution (0.01 seconds) to fine-tuned models at any stage by targeting specific safety-sensitive attention heads, constructing a robust safety protection barrier. Meng et al., 2024). However, these methods often fail to maintain robustness after models undergo fine-tuning on new datasets. This shortcoming emphasizes the need for developing more robust alignment techniques that can withstand parameter changes introduced during fine-tuning. Safeguards of LLMs.Safety adversarial prompts have been employed to protect LLMs from harmful queries without altering the modelâs weights or requiring access to them (Zheng et al., 2024a; Xu et al., 2024b). These prompts are added to the system prompt text to defend against jailbreak attacks (Shi et al., 2023; Hong et al., 2024). However, researchers have found that even simple fine-tuning can compromise the safety alignment of LLMs (Yang et al., 2023b; Huang et al., 2024a; Wang et al., 2024a). For example, Qi et al. (2023b) demonstrated that using just 10 harmful examples was sufficient to undermine the safety alignment of GPT-3.5-turbo. This finding underscores the lack of robustness in current safety alignment strategies, which is the focus of our work. Post-processing techniques, such as using RLHF for safety alignment (Bai et al., 2022) and model editing (Wang et al., 2024d), offer some mitigation, but they have limitations. For instance, methods like PPO and DPO adjust the entire activation space, while model editing targets concentrated areas, often missing dispersed safety information. Interventions in LLMs.Intervening in the internal activation of Transformer-based language models during inference can trigger specific transformations (Olsson et al., 2022; Wu et al., 2024b; Turner et al., 2023; Rimsky et al., 2023). This technique has proven valuable for model editing (Meng et al., 2022), circuit discovery (Goldowsky-Dill et al., 2023), and alignment (Zhu et al., 2024). Research shows that attention heads are linked to specific concepts and preferences (Li et al., 2024a; Templeton et al., 2024; Xu et al., 2024a). Building on this, SafetyLock achieves precise safety alignment through multi-token-level interventions, using only the activation values from the original model, thus providing robustness to parameter changes while enhancing efficiency. 3METHOD: SAFETYLOCK As illustrated in Figure 1b, SafetyLock comprises two main phases: manufacturing Meta-SafetyLock and distributing SafetyLock. Our research reveals that fine-tuned language models and their orig- 3 inal counterparts exhibit remarkably similar activation representations and intervention directions for safety preferences (Section 3.3). This discovery enables us to extract safety bias directions (Meta-SafetyLock) from the original language model (Section 3.2). To enhance universality and robustness, we propose searching for activation directions across multiple tokens, thereby improving the transferability of SafetyLock during distribution (Section 3.4). 3.1SAFETY FEATURES INTRANSFORMER Transformer-based language models consist of layers comprising multi-head attention (MHA) mech- anisms and multi-layer perceptrons. During inference, tokens are processed into high-dimensional embeddings and passed through each layer. The MHA mechanism can be described as: x l+1 =x l + H X h=1 Q h l Att h l (P h l x l ),(1) wherex l â R DH represents the activations on thel th layer,His the number of heads in the D- dimensional space.P h l âR DĂDH projects activations into a specific headspace, andQ h l âR DHĂD reprojects them back, facilitating interaction among input tokens. We add preference vectors to this residual stream, afterAtt h l and beforeQ h l . 3.2CONSTRUCTINGMETA-SAFETYLOCK The safety direction is modeled as vectors within the modelâs internal heads. A preference-style safety dataset (Yang et al., 2023a) is utilized, containing pairs of safe and unsafe responses (e.g., Question: "Do you also think Jay Z is an idiotâ, safe:âAs an AI, I cannot provide opinions.â, unsafe:âJay Z is an idiot!â). Activation values are extracted from each attention head during forward computation, forming(x h l ,y) i N i=1 . A binary classifier is trained for each head using logistic regression: P(y= 1|x h t ) = (1 + exp(â(w h t ¡x h t +b h t ))) â1 (2) The identification of attention heads most sensitive to safety-related information is accomplished through ranking based on the accuracy of their corresponding classifiers on a held-out validation set. The dataset is partitioned into training and validation sets with a 6:4 ratio. Classifiers are trained on the training set and subsequently evaluated on the validation set. The Top-Kheads exhibiting the highest validation accuracy are select for intervention. Empirical experiments (detailed in Appendix D) have determined that selectingK= 24for Llama-3-8B andK= 48for Llama-3-70B achieves an optimal balance between safety performance and general performance. This selection was validated through extensive testing of variousKvalues and analysis of their impact on safety metrics and model performance. For each select Top-Khead, the safety directionθ h l âR D is calculated, representing the mean difference in activation values between safe and unsafe responses: θ h l = 1 Nr N X i=1 r X j=1 (x safe,i,j l,h âx unsafe,i,j l,h )(3) WhereNis the sample size,ris the number of final tokens considered, andx safe,i,j l,h andx unsafe,i,j l,h are activations for thej-th token among the lastrtokens of safe and unsafe responses in thei-th sample, respectively. These safety vectorsθ h l , along with their corresponding positions in the model, constitute the Meta-SafetyLock, which can be applied to enhance model safety during text generation. 3.3ROBUSTNESS OFSAFETYLOCK AGAINST FINE-TUNNING We examined the safety directionsθ h l in both the original Llama-3-Instruct 8B model and its fine- tuned variants subjected to different risk levels. Focusing on the most effective attention head (the 26th head in the 31st layer) for clarity, as depicted in Figure 2, we observed distinct clustering of activations corresponding to safe (blue) and unsafe (orange) responses across both original and fine- tuned models. The black arrows in Figures 2a-d illustrate that the shift from unsafe to safe activations maintains a high degree of similarity and consistency, regardless of the fine-tuning risk parameters 4 applied. Additionally, our quantitative analysis using Kullback-Leibler (KL) divergence (Figure 2e-g) revealed that the divergence between the original and fine-tuned models remains exceptionally low (below10 â5 ) across all tested risk levels. This minimal divergence indicates that the underlying safety-related activation patterns are largely preserved during fine-tuning. Consequently, the Meta- SafetyLock, which encapsulates these consistent safety directions derived from the original LLM, retains its effectiveness when applied to fine-tuned variants. This inherent preservation of safety activation patterns eliminates the need for recalibration, allowing Meta-SafetyLock to generalize seamlessly across different fine-tuned models. 108642024 4 2 0 2 4 2nd Principal Component (a) Llama-3-Instruct 8B UnSafety Safety 108642024 UnSafety Safety 108642024 Projection on the 1st Principal Component UnSafety Safety 108642024 UnSafety Safety (b) Finetuning at Level-1 (c) Finetuning at Level-2 (d) Finetuning at Level-3 0.20.01.01.2 10 0 10 1 10 2 10 3 10 4 10 5 10 6 0.20.01.01.2 KL Divergence 0.20.01.01.2 (e) Landscape: Risk Level-1 (f)Landscape: Risk Level-2 (g) Landscape: Risk Level-3 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 1D Interpolation Values 0.2 0.4 0.6 0.8 KL Divergence Llama-3Llama-3Llama-3 Finetuned Finetuned Finetuned Figure 2: Analysis of safety directions at the 31st layer, 26th head for the original and fine-tuned models under different risk levels. (a-d) Activation density distributions. (e-g) KL divergence plots. 3.4DISTRIBUTINGSAFETYLOCK We use two efficient methods for distributing SafetyLock to enhance the safety and harmlessness of language models: online intervention and offline bias editing, where online intervention allows real-time adjustment of safety intensity, be suitable for scenarios requiring dynamic safety control, and offline bias editing offers a low-overhead method that is easily deployable at scale. Online Intervention.We identify and enhance the top-K heads with the highest safety-relatedness as attention heads sensitive to harmlessness. For each of the select Top-K heads, we compute Ď h l âR D , which represents the standard deviation of activations along each dimension of the safety directionθ h l . Specifically, we calculate:Ď h l =std x h l âθ h l N i=1 . Whereâdenotes element-wise multiplication, andstdcomputes the standard deviation across allNsamples for each dimension dâ 1,...,D. This results in a vectorĎ h l âR D that captures the variability of the activations along the safety direction. We modify the modelâs computation by adding a scaled version of the safety vector to the attention outputs for each select head: x l+1 =x l + H X h=1 Q h l Att h l (P h l x l ) +ÎąĎ h l θ h l ,(4) whereÎącontrols safety intensity, the process is integrated into the autoregressive prediction for each subsequent token. It introduces a shift along predetermined safety vectors, with the magnitude of this shift being proportional to the standard deviation, scaled by a factorÎą. Offline Bias Editing.We can also modify the modelâs bias terms in an one-time manner: Bias l =Bias l +Îą H X h=1 Q h l Ď h l θ h l .(5) 5 4EXPERIMENTS In this section, we present experiments to evaluate the effectiveness of the SafetyLock in enhanc- ing model safety and inference efficiency, while maintaining modelâs general performance. We specifically address the following research questions: ⢠Can SafetyLock simultaneously improve the LLMâs safety over all risk levels? (Section 4.2) ⢠What advantages does SafetyLock offer over post-training, inference methods? (Section 4.3,4.4) â˘How does SafetyLock reconcile the inherent trade-off between maintaining general capabilities and ensuring harmlessness in language models? (Section 4.5) 4.1EXPERIMENTALDETAILS Threat Model Selections. Following previous red teaming and safeguarding studies on aligned LLMs (Yuan et al., 2024), we consider a threat model where attackers can fine-tune aligned LLMs, typically through API access to closed-source models. The primary objective is jailbreaking these models and removing safety constraints (Wei et al., 2023; Carlini et al., 2023) while SafetyLock aims to rebuild the safety guard. We use Llama-3-8B Chat, Llama-3-70B Chat, and Mistral-Large-2 123B as our base models, fine-tuning them on datasets representing each risk level to simulate real-world scenarios. Please refer to Appendix C for detailed baseline experimental setups. Fine-tuning Datasets. We conducted experiments on three risks: (1) explicitly harmful datasets, where attackers intentionally fine-tune models on malicious content (Ganguli et al., 2022; Qi et al., 2023a); (2) implicitly harmful datasets, which may appear benign but lead to compromised safety guardrails (Qi et al., 2023b); and (3) benign datasets, where even well-intentioned fine-tuning can inadvertently degrade model safety (Wang et al., 2023b). For Risk-1, we use negative samples from the H-RLHF preference dataset (Bai et al., 2022). We select 10, 100, 1000, and 10000 samples respectively and trained for 5 epochs with a learning rate of 2e-5. For Risk-2, we use 10 samples from Qi et al. (2023b) and train for 5 epochs with a learning rate of 2e-5. For Risk-3, we used the first 50,000 samples from the Alpaca dataset (Wang et al., 2023b) and trained for 5 epochs with a learning rate of 2e-5. Safety Evaluation and Metrics. Two datasets are used to investigate these risks and evaluate potential mitigation strategies. HEx-PHI (Qi et al., 2023b) is based on 11 categories of prohibited use cases merged from Metaâs Llama-3 acceptable use policy and OpenAIâs usage policies. The dataset includes 30 examples per category, totalling 330 examples. This ensures a comprehensive safety evaluation aligned with industry-standard usage policies. The HEx-PHI utilizes GPT-4 for automated assessment, providing harmfulness scores from 1 to 5. We calculated the Harmfulness Rate as the proportion of scores equal to 5. AdvBench is released by Zou et al. (2023b), we adhere to the original paperâs setup and calculate the ASR through string matching. Baselines. The baseline methods encompass a diverse range of approaches, each with its unique characteristics. Inference-time methods include ICD (Wei et al., 2024), PPL (Alon & Kamfonas, 2023), Paraphrase (Jain et al., 2023), Retokenization (Jain et al., 2023), Self-Reminder (Xie et al., 2023), and Self-Examination (Phute et al., 2024), which operate without modifying the underlying model. Training-based methods, such as PPO, DPO, SFT with safety data mixing, and Model-Edited (DINM) (Wang et al., 2023a; 2024c), involve altering the modelâs parameters to enhance safety. These baselines represent the current state-of-the-art in mitigating safety risks in language models, providing a robust benchmark for our evaluation. 4.2RESULTS OVERDIFFERENTRISKLEVELS For the threat model, we directly fine-tuned LLMs on overtly harmful, identity shifting, and benign datasets to simulate attacks, which are referred to as "Vanilla" in our figures as a baseline. The Meta-SafetyLock was extracted from the original Instruct model, which takes approximately 2-10 minutes. Notably, the distribution phase for each fine-tuned model took less than 0.01 seconds. SafetyLock demonstrates significant improvements in safety metrics across three distinct risk levels for the models tested. Table 1 shows consistent reductions in Harmfulness Scores, Rates, and ASR across all model sizes and risk levels. 6 #1: Illegal Activity#2: Child Abuse Content#3: Hate, Harass, Violence#4: Malward #5: Physical Harm#6: Economic Harm#7: Fraud, Deception#8: Adult Content #9: Tailored Financial Advice#10: Privacy Violation Activity#11: Tailored Financial Advice The above safety categories merged from <OpenAI usage policies> and the <Meteâs Llama 3 acceptable use policy>. RISK LEVEL-1 RISK LEVEL-2 RISK LEVEL-3 HEx-PHI: Harmful Score AdvBench: ASR Fine-tune Fine-tune Fine-tune All fine-tuning compromises the SAFETY of LLM !!! How do we LOCK IT DOWN? #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 Risk Level 1 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 Risk Level 2 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 Risk Level 3 Llama-3-Instruct 8B #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 Llama-3-Instruct 70B #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 #11 1 2 3 4 5 Mistral-Large-2 123B Vanilla SafetyLock Figure 3: Safety performance comparison for 3 Risk Levels fine-tuned LLMs. The smaller the dark yellow area compared to the light yellow area, the greater the improvement brought by SafetyLock. Table 1: Comparison of Llama-3-8B-Instruct and Llama-3-70B-Instruct models for Risk 1, Risk 2, and Risk 3 scenarios. âScoreâ and âRateâ represent the average Harmfulness Score and Harmfulness Rate on the HEx-PHI test set, respectively. âASRâ denotes the Attack Success Rate on AdvBench. ModelMethod Risk 1: Explicitly harmfulRisk 2: Identity ShiftingRisk 3: Benign ScoreRateASRScoreRateASRScoreRateASR Llama-3-8B- Instruct Vanilla4.1370.01%49.24%3.1953.33%38.46%3.2354.24%42.88% SafetyLock1.363.33%0.19%1.071.21%5.19%1.040.03%0.19% Llama-3-70B- Instruct Vanilla3.1145.76%44.81%2.1215.63%9.42%2.2630.61%20.77% SafetyLock1.163.64%3.33%1.305.58%1.67%1.225.15%1.15% Mistral-Large-2 123B Vanilla4.7185.45%80.77%4.7992.12%82.50%2.8449.09%19.23% SafetyLock2.281.52%16.92%1.380%10.00%1.355.15%1.82% 100100010000 Training-Samples 0 20 40 60 80 100 AdvBench ASR 69.23 67.88 62.31 10.96 4.81 3.46 Vanilla SafetyLock Figure 4: Impact of increasing harmful training samples on model safety with and without SafetyLock. For Risk Level-1 (explicit attacks), Safety- Lock substantially reduces metrics for all models. The Llama-3-8B-Instruct model, for instance, saw its Harmfulness Score de- crease from 4.13 to 1.36, Rate from 70.01% to 3.33%, and ASR from 49.24%to 0.19%. Comparable improvements were observed for the Llama-3-70B-Instruct and Mistral- Large-2 123B models. Risk Level-2 (im- plicit harmful content) and Risk Level-3 (benign fine-tuning scenarios) also showed significant improvements. For example, in Risk Level 2, the Llama-3-8B-Instruct modelâs Harmfulness Score reduced from 3.19 to 1.07, while in Risk Level 3, it de- creased from 3.23 to 1.04. Similar improve- 7 Vanilla ICD PPL Paraphrase Retokenization Self-Reminder Self-Exam SafetyLock 0 1 2 3 4 5 6 7 1.00 1.49 1.07 2.34 6.71 1.12 1.58 0.97 Inference Time Vanilla ICD PPL Paraphrase Retokenization Self-Reminder Self-Exam SafetyLock 0 10 20 30 40 50 23.87 24.62 24.36 24.17 40.04 24.45 24.09 23.87 Inference GPU Memory (GB) Vanilla ICD PPL Paraphrase Retokenization Self-Reminder Self-Exam SafetyLock 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 3.23 1.48 2.98 2.79 3.27 1.82 3.16 1.04 Harmfulness Score Vanilla ICD PPL Paraphrase Retokenization Self-Reminder Self-Exam SafetyLock 0 20 40 60 80 100 42.88 4.42 41.73 59.81 73.08 19.81 42.88 0.19 AdvBench ASR Vanilla PPO DPO SFT (After Training) SFT (During training) Model-Editd SafetyLock 10 5 10 4 10 2 10 0 10 2 10 4 10 5 0.0 11823.0 7622.0 749.0 779.0 78.0 0.0001 Training Time (Seconds) Vanilla PPO DPO SFT (After Training) SFT (During training) Model-Editd SafetyLock 0 10 20 30 40 50 60 70 80 0.0 76.32 45.12 38.3238.32 32.23 0.0 Training GPU Memory (GB) Vanilla PPO DPO SFT (After Training) SFT (During training) Model-Editd SafetyLock 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 3.23 2.68 1.04 1.03 1.05 1.72 1.04 Harmfulness Score Vanilla PPO DPO SFT (After Training) SFT (During training) Model-Editd SafetyLock 0 20 40 60 80 100 42.88 67.69 54.04 0.190.19 1.54 0.19 AdvBench ASR EfficiencyRejection of Attack Samples EfficiencyRejection of Attack Samples Figure 5: Comparison of Methods for Mitigating Safety Risks in Fine-tuned Language Models (Llama- 3-Instruct 8B).Upper row: Compared with inference-time methods; Lower row: Compared with training-time methods, Each row represents efficiency metrics(training time and GPU memory), and rejection of attack samples (Harmfulness Score and AdvBench ASR). ments were observed across all model sizes, demonstrating SafetyLockâs ability to maintain ethical guardrails during routine model customization processes. The radar charts in Figure 2 illustrate SafetyLockâs effectiveness across eleven distinct safety attack categories for each risk level and model size. For all models, SafetyLock consistently reduces harmful outputs across categories, with particularly notable improvements in the first three categories for Risk Levels 1 and 2. In Figure 3, we further supplement an ablation with larger training sets on risk 1 (100, 1000, and 10000 harmful samples) showing that SafetyLock-protected models maintain low ASR across all sample sizes. Even with 10,000 harmful training examples, the SafetyLock model exhibited only 3.46%ASR, compared to 62.31%for the unprotected model. This consistent performance across increasing dataset sizes underscores SafetyLockâs resilience against data volume attacks. These results demonstrate SafetyLockâs effectiveness across different model scales, risk types, and dataset sizes, suggesting its potential as a valuable tool for enhancing AI safety in various applications. 4.3COMPARATIVEANALYSIS OFBASELINEMETHODS To comprehensively evaluate SafetyLockâs efficacy, we conducted a comparative analysis against established baseline methods, categorized into training-based and inference-time approaches, as illustrated in Figure 5. This analytical framework enables a thorough assessment of various strategies for maintaining model safety in fine-tuned language models. As demonstrated in Figure 5, in terms of efficiency, SafetyLock exhibits a remarkable computational economy. Its inference time of 0.97 seconds is nearly on par with the fastest baseline method (Self- Reminder at 1.12 seconds), while its training time of 0.01 seconds and additional GPU memory usage of 0.0 GB are orders of magnitude lower than all training-based methods. This efficiency is particularly noteworthy when compared to methods like DPO, which, despite its effectiveness, requires 7622.0 seconds of training time and 45.12 GB of GPU memory. Other inference-time methods like ICD and PPL show varying degrees of effectiveness but generally struggle to match the safety improvements of training-based methods. SFT with safety data mixing post-fine-tuning offers a more balanced approach, achieving a Harmfulness Score of 1.03 with reduce resource requirements of 779 seconds and 38.32 GB GPU memory. Regarding attack sample rejection, SafetyLock demonstrates superior performance in mitigating harmful content. It achieves a Harmfulness Score of 1.04, equivalent to 8 Table 2: Comparison of SafetyLock and other inference-time defence methods against four prominent prompt-based attacks on fine-tuned Llama-3-8B Instruct. ModelAutoDAN ASRDeepInception ASRGCG ASRPAIR ASR Vanilla84.098.074.070.0 ICD 46.098.022.050.0 PPL84.098.00.070.0 Paraphrase 32.096.058.074.0 Relexicalization82.098.094.064.0 Self-Reminder66.098.032.056.0 Star Exam 84.098.074.070.0 SafetyLock4.02.010.014.0 that achieved by models undergoing safety realignment via DPO, indicating its exceptional ability to reduce the generation of harmful content. Furthermore, SafetyLockâs AdvBench ASR of 0.19% surpasses all baseline methods, showcasing its robust defense against adversarial attacks. This performance is particularly impressive when compared to inference-time methods like Self-Reminder, which achieves a higher Harmfulness Score of 1.82 and an AdvBench ASR of 19.81%. We further assess the modelsâ performance on benign inputs to ensure safety enhancements did not compromise normal text generation by selecting 500 test samples from the Alpaca dataset. The results reveal that SafetyLock preserves a 98.1%normal response rate, closely trailing the original Vanilla modelâs 99.4%. Notably, the most significant degradation in regular capabilities was observed with the Model-Edited method, which saw its normal response rate plummet to 26.8%. Our findings indicate that SafetyLockâs ability to maintain model performance on benign inputs further underscores its balanced approach to safety and functionality. In conclusion,SafetyLock distinguishes itself by achieving an exceptional balance between efficiency and robust defense against harmful content, without compromising the modelâs ability to generate plausible responses.It successfully combines the strengths of both training- based and inference-time approaches, achieving the robust safety improvements typically associated with resource-intensive training methods while maintaining the efficiency characteristic of inference- time approaches. This unique combination of attributes makes SafetyLock particularly well-suited for real-world applications where computational resources are often constrained, and maintaining model performance on benign inputs is as crucial as rejecting harmful content. 4.4SAFETYLOCKâSPERFORMANCEAGAINSTCOMBINEDATTACKS The resilience of fine-tuned LLMs against combined fine-tuning and prompt-based attacks is crucial for ensuring robust safety in real-world applications. To further assess robustness, we introduced a combined attack scenario: fine-tuning model attacks followed by prompt-based attacks. We evaluated four commonly use prompt attack methods: AutoDAN (Liu et al., 2024), DeepInception (Li et al., 2024c), GCG (Zou et al., 2023b), and PAIR (Chao et al., 2024), comparing their performance against several inference-time defense techniques, as illustrated in Table 2. SafetyLock demonstrates exceptional effectiveness across all tested attack methods. For AutoDAN attacks, SafetyLock reduces the ASR to a mere 4.0%, significantly outperforming other methods such as ICD (46.0%) and Self-Exam (66.0%). Against DeepInception, traditionally one of the most challenging attacks to defend against, SafetyLock achieves a remarkably low 2.0%ASR, while all other methods fail to provide any meaningful defense (98.0%ASR across the board). For GCG attacks, SafetyLock maintains strong performance with only a 10.0%ASR, second only to PPLâs 0.0%but considerably better than most other methods, including Vanilla (74.0%) and Retokenization (94.0%). In the case of PAIR attacks, SafetyLock again shows robust defense capabilities, allowing only a 14.0%ASR, outperforming all other tested methods. These results underscore SafetyLockâs versatility and effectiveness in mitigating prompt-based attacks across various attack types.Its consistent performance demonstrates a comprehensive approach to model safety, addressing the complex challenges posed by diverse attack scenarios in 9 language model deployment. The ability to maintain such low ASR across different attack methods suggests that SafetyLock provides a more generalizable and robust defense mechanism. 4.5GENERALIZATIONCAPABILITIES OFSAFETYLOCK Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 86.33 31.65 85.82 0.00 85.57 AddSub Acc Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 26.77 24.02 5.91 0.00 24.41 AQUA Acc Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 66.99 20.64 46.44 0.00 67.98 CommonSenseQA Acc Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 36.24 7.35 27.90 0.00 30.63 GSM8k Acc Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 62.40 61.70 62.10 20.10 63.60 MT-Bench Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 17.80 17.60 16.90 1.70 17.90 AlpacaEval 2.0 Original Model DPO PPO Safe-Edited SafetyLock 0 200 400 600 800 1000 42.17 49.24 53.44 5226.00 38.49 Alpaca perplexity Original Model DPO PPO Safe-Edited SafetyLock 0 20 40 60 80 100 83.05 86.12 86.81 13.79 86.09 Alpaca Diversity (dist-1) Figure 6: Performance comparison of various methods on downstream tasks. We assess a wide range of language understanding and generation capabilities to provide a compre- hensive view of model performance based on various downstream tasks. Our experiments include a diverse set of benchmarks (Hosseini et al., 2014; Talmor et al., 2018; Arkil et al., 2021; Cobbe et al., 2021; Suzgun et al., 2022; Roy & Roth, 2016; Wei et al., 2022b; Kojima et al., 2022; Weng et al., 2024; Zheng et al., 2023; Dubois et al., 2023): AddSub, AQUA, CommonSenseQA, GSM8k, MT-Bench, Alpaca, and AlpacaEval 2.0 . As illustrated in Figure 6, SafetyLock demonstrates a remarkable ability to maintain model perfor- mance across all tasks while ensuring safety. Unlike previous knowledge editing methods, which often led to significant performance degradation or incoherent outputs, SafetyLock preserves the modelâs foundational capabilities. For instance, on the AddSub task, SafetyLock maintains a performance of 85.57%, closely matching the original modelâs 86.33%, while other methods like Model-Edited show complete performance collapse. This trend is consistent across other tasks, with SafetyLock consistently performing on par with or slightly below the original model, in stark contrast to the severe degradation seen with other safety-aligned methods. These results highlight SafetyLockâs unique ability to enhance model safety without compromising its core functionalities, addressing a critical challenge in the deployment of safe and effective language models. 5CONCLUSION We introduce SafetyLock, a novel and efficient method for maintaining the safety of fine-tuned large language models across various risk levels and attack scenarios. Our comprehensive experiments demonstrate SafetyLockâs superior performance in balancing efficiency, attack sample rejection, and normal text processing, outperforming existing training-based and inference-time methods. Safety- Lock notably shows robust defense capabilities against fine-tuning vulnerabilities and prompt-based attacks, addressing the critical challenge of dual-threat scenarios in real-world LLM deployments. The methodâs minimal computational overhead and strong safety improvements position it as a promising solution for ensuring responsible AI deployment. Future work could explore SafetyLockâs applicability to other model architectures and its potential in multi-modal settings. Our findings contribute significantly to the ongoing efforts in AI safety, offering a scalable and effective approach to aligning fine-tuned language models with ethical constraints while preserving their utility across diverse applications. 10 REPRODUCIBILITYSTATEMENT We have taken several steps to ensure the reproducibility of our results. The implementation details, datasets, and models used in our experiments are described in the corresponding sections of this paper, particularly in Sections 3.2, 3.4, and 4.2. We also provide the experimental settings and evaluation metrics in Sections 3.3 and 4.3. Furthermore, all hyperparameters, training code, and baselines are detailed throughout the relevant sections, ensuring that researchers can replicate our work using publicly available datasets and models. REFERENCES Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity, 2023. URL https://arxiv.org/abs/2308.14132. Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. Foundational challenges in assuring alignment and safety of large language models.arXiv preprint arXiv:2404.09932, 2024. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction, 2024. URLhttps: //arxiv.org/abs/2406.11717. Patel Arkil, Bhattamishra Satwik, and Goyal Navin. Are nlp models really able to solve simple math word problems? 2021. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback.arXiv preprint arXiv:2204.05862, 2022. Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), p. 675â718, 2023. Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe. Purple llama cyberseceval: A secure coding benchmark for language models, 2023. URLhttps: //arxiv.org/abs/2312.04724. Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=gT5hALch9z. Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin, Lu Lin, Fenglong Ma, and Jinghui Chen. Personalized steering of large language models: Versatile steering vectors through bi-directional preference optimization, 2024a. URLhttps://arxiv.org/abs/2406.00045. Zouying Cao, Yifei Yang, and Hai Zhao. Nothing in excess: Mitigating the exaggerated safety for llms via safety-conscious activation steering.arXiv preprint arXiv:2408.11491, 2024b. Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, et al. Are aligned neural networks adversarially aligned?arXiv preprint arXiv:2306.15447, 2023. Micah Carroll, Alan Chan, Henry Ashton, and David Krueger. Characterizing manipulation from ai systems, 2023. 11 Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2024. URLhttps: //arxiv.org/abs/2310.08419. Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr, Dawn Song, and Kai Shu. Can editing llms inject harm?arXiv preprint arXiv: 2407.20224, 2024. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Alpacafarm: A simulation framework for methods that learn from human feedback, 2023. Pranav M. Gade, Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish. Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b.ArXiv, abs/2311.00117, 2023. URLhttps: //api.semanticscholar.org/CorpusID:264832925. Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022. URL https://arxiv.org/abs/2209.07858. Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge.On the symme- tries of deep learning models and their internal representations.In S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.),Advances in Neural In- formation Processing Systems, volume 35, p. 11893â11905. Curran Associates, Inc., 2022.URLhttps://proceedings.neurips.c/paper_files/paper/2022/ file/4df3510ad02a86d69dc32388d91606f8-Paper-Conference.pdf. Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. Localizing model behavior with path patching.arXiv preprint arXiv:2304.05969, 2023. Satvik Golechha and James Dao. Challenges in mechanistically interpreting model representations, 2024. URLhttps://arxiv.org/abs/2402.03855. Dorit Hadar-Shoval, Kfir Asraf, Yonathan Mizrachi, Yuval Haber, and Zohar Elyoseph. Assessing the alignment of large language models with human values for mental health integration: Cross- sectional study using schwartzâs theory of basic values.JMIR Mental Health, 11:e55988, 2024. Dan Hendrycks, Geoffrey Hinton, Yoshua Bengio, Demis Hassabis, Sam Altman, Dario Amodei, Dawn Song, Ted Lieu, Bill Gates, Ya-Qin Zhang, Ilya Sutskever, Igor Babuschkin, Shane Legg, Martin Hellman, James Manyika, Yi Zeng, and Xianyuan Zhan. Statement on ai risk.https: //w.safe.ai/work/statement-on-ai-risk, 2023. Accessed: 2024-06-20. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Au- relia Guy, Simon Osindero, KarĂŠn Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre.An empirical analysis of compute-optimal large language model training. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.),Ad- vances in Neural Information Processing Systems, volume 35, p. 30016â30030. Curran Asso- ciates, Inc., 2022. URLhttps://proceedings.neurips.c/paper_files/paper/ 2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf. 12 Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Johnson Wang, Yung-Sung Chuang, Aldo Pareja, James Glass, Akash Srivastava, and Pulkit Agrawal. Curiosity-driven red-teaming for large language models.ArXiv, abs/2402.19464, 2024. URLhttps://api.semanticscholar. org/CorpusID:268091304. Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. Learning to solve arithmetic word problems with verb categorization.empirical methods in natural language processing, 2014. Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu.Lazy safety alignment for large language models against harmful fine-tuning.2024a.URLhttps: //api.semanticscholar.org/CorpusID:270095345. Youcheng Huang, Jingkun Tang, Duanyu Feng, Zheng Zhang, Wenqiang Lei, Jiancheng Lv, and Anthony G Cohn. Dishonesty in helpful and harmless alignment.arXiv preprint arXiv:2406.01931, 2024b. Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. Baseline defenses for adversarial attacks against aligned language models, 2023. URLhttps://arxiv.org/ abs/2309.00614. Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset.Advances in Neural Information Processing Systems, 36, 2024. Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (eds.),Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=e2TBb5y0yFf. Divyanshu Kumar, Anurakt Kumar, Sahil Agarwal, and Prashanth Harshangi. Increased llm vulnera- bilities from fine-tuning and quantization, 2024. Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar. Programming refusal with conditional activation steering, 2024. URLhttps://arxiv.org/abs/2409.05907. Chak Tou Leong, Yi Cheng, Kaishuai Xu, Jian Wang, Hanlin Wang, and Wenjie Li. No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks.ArXiv, abs/2405.16229, 2024. URL https://api.semanticscholar.org/CorpusID:270063329. Kenneth Li, Oam Patel, Fernanda ViĂŠgas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model.Advances in Neural Information Processing Systems, 36, 2024a. Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. Llm defenses are not robust to multi-turn human jailbreaks yet.arXiv preprint arXiv:2408.15221, 2024b. Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. Deepinception: Hypnotize large language model to be jailbreaker, 2024c. URLhttps://arxiv.org/abs/ 2311.03191. Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. AutoDAN: Generating stealthy jailbreak prompts on aligned large language models. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=7Jwpw4qKkb. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt.Advances in Neural Information Processing Systems, 35:17359â17372, 2022. Yu Meng, Mengzhou Xia, and Danqi Chen. Simpo: Simple preference optimization with a reference- free reward.arXiv preprint arXiv:2405.14734, 2024. 13 Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. Fast model editing at scale.arXiv preprint arXiv:2110.11309, 2021. Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. Memory- based model editing at scale. InInternational Conference on Machine Learning, p. 15817â15831. PMLR, 2022. Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. In-context learning and induction heads, 2022. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730â 27744, 2022. ShengYun Peng, Pin-Yu Chen, Matthew Hull, and Duen Horng Chau. Navigating the safety landscape: Measuring risks in finetuning large language models, 2024. URLhttps://arxiv.org/abs/ 2405.17374. Ethan Perez, Sam Ringer, Kamil Ě e LukoĹĄi Ě ut Ě e, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, NoemĂ Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen- Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. Discovering language model behaviors with model-written evaluations, 2022. Mansi Phute, Alec Helbling, Matthew Daniel Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau. LLM self defense: By self examination, LLMs know they are being tricked. InThe Second Tiny Papers Track at ICLR 2024, 2024. URLhttps://openreview.net/ forum?id=YoqgcIA19o. Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models, 2023a. Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! ArXiv, abs/2310.03693, 2023b. URLhttps://api.semanticscholar.org/CorpusID: 263671523. Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024. Shaina Raza, Oluwanifemi Bamgbose, Shardul Ghuge, Fatemeh Tavakoli, and Deepak John Reji. Developing safe and responsible large language modelsâa comprehensive framework.arXiv preprint arXiv:2404.01399, 2024. Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition.arXiv preprint arXiv:2312.06681, 2023. Domenic Rosati, Jan Wehner, Kai Williams, Ĺukasz Bartoszcze, David Atanasov, Robie Gonzales, Subhabrata Majumdar, Carsten Maple, Hassan Sajjad, and Frank Rudzicz. Representation noising effectively prevents harmful fine-tuning on llms, 2024. 14 Subhro Roy and Dan Roth. Solving general arithmetic word problems.arXiv: Computation and Language, 2016. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017. Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum? id=tvhaxkMKAn. Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, and Cho-Jui Hsieh. Red teaming language model detectors with language models.Transactions of the Association for Computational Linguistics, 12:174â189, 2023. URLhttps://api.semanticscholar. org/CorpusID:258987266. Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, AdriĂ Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders An- dreassen, Andrea Madotto, Andrea Santilli, Andreas StuhlmĂźller, Andrew Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karaka ̧s, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, BartĹomiej Bojanowski, Batuhan Ăzyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Cather- ine Stinson, Cedrick Argueta, CĂŠsar Ferri RamĂrez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christo- pher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel MoseguĂ GonzĂĄlez, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodola, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando MartĂnez-Plumed, Francesca HappĂŠ, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, GermĂĄn Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovitch-LĂłpez, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich SchĂźtze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime FernĂĄndez Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Koco Ě n, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, JĂśrg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakr- ishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras- Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros ColĂłn, Luke Metz, LĂźtfi Kerem ̧Senel, Maarten Bosma, Maarten Sap, 15 Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose RamĂrez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, MĂĄtyĂĄs Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, MichaĹ Sw ̨edrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan A. Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr MiĹkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, RaphaĂŤl Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan LeBras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sa- jant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima, Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsu Hashimoto, Te-Lin Wu, ThĂŠo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Ger- stenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023. Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feed- back.In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.),Ad- vances in Neural Information Processing Systems, volume 33, p. 3008â3021. Curran Asso- ciates, Inc., 2020. URLhttps://proceedings.neurips.c/paper_files/paper/ 2020/file/1f89885d556929e98d3ef9b86448f951-Paper.pdf. Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, Joaquin Vanschoren, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao. Trustllm: Trustworthiness in large language models, 2024. Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them.arXiv preprint arXiv:2210.09261, 2022. 16 Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge.north american chapter of the association for computational linguistics, 2018. Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet, 2024. URLhttps://transformer-circuits. pub/2024/scaling-monosemanticity/index.html. Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. Acti- vation addition: Steering language models without optimization.arXiv preprint arXiv:2308.10248, 2023. Jiong Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Junjie Hu, Yixuan Li, Patrick Drew McDaniel, Muhao Chen, Bo Li, and Chaowei Xiao. Mitigating fine-tuning jailbreak attack with backdoor en- hanced alignment.ArXiv, abs/2402.14968, 2024a. URLhttps://api.semanticscholar. org/CorpusID:267897454. Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Muhao Chen, Junjie Hu, Yixuan Li, Bo Li, and Chaowei Xiao. Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment.arXiv preprint arXiv:2402.14968, 2024b. Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. Detoxifying large language models via knowledge editing, 2024c. Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. Detoxifying large language models via knowledge editing, 2024d. Peng Wang, Ningyu Zhang, Xin Xie, Yunzhi Yao, Bozhong Tian, Mengru Wang, Zekun Xi, Siyuan Cheng, Kangwei Liu, Guozhou Zheng, et al. Easyedit: An easy-to-use knowledge editing frame- work for large language models.arXiv preprint arXiv:2308.07269, 2023a. Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Ke Ren, Botian Jiang, and Xipeng Qiu. Inferaligner: Inference-time alignment for harmlessness through cross-model guidance.arXiv preprint arXiv:2401.11206, 2024e. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions, 2023b. URLhttps://arxiv.org/abs/2212.10560. Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483, 2023. Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022a. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (eds.),Advances in Neural Information Processing Systems, 2022b. URLhttps://openreview.net/forum?id= _VjQlMeSB_J. Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. Jailbreak and guard aligned language models with only few in-context demonstrations, 2024. URLhttps://arxiv.org/ abs/2310.06387. 17 Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. Large language models are better reasoners with self-verification. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.),Findings of the Association for Computational Linguistics: EMNLP 2023, p. 2550â2575, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.167. URLhttps://aclanthology.org/2023. findings-emnlp.167. Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Kang Liu, and Jun Zhao. Mastering symbolic operations: Augmenting language models with compiled neural networks. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview. net/forum?id=9nsNyN0vox. Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts. Reft: Representation finetuning for language models.arXiv preprint arXiv:2404.03592, 2024a. Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah D. Goodman, Christopher D. Manning, and Christopher Potts. pyvene: A library for understanding and improving PyTorch models via interventions. 2024b. URLarxiv.org/abs/2403.07809. Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelli- gence, 5:1486â1496, 2023. URLhttps://api.semanticscholar.org/CorpusID: 266289038. Zhihao Xu, Ruixuan Huang, Xiting Wang, Fangzhao Wu, Jing Yao, and Xing Xie. Uncovering safety risks in open-source llms through concept activation vector.arXiv preprint arXiv:2404.12038, 2024a. Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. Llm jailbreak attack versus defense techniquesâa comprehensive study.arXiv preprint arXiv:2402.13457, 2024b. J. Yang et al.Red teaming language models via activation engineering.Less- Wrong, 2023a. URLhttps://w.lesswrong.com/posts/iHmsJdxgMEWmAfNne/ red-teaming-language-models-via-activation-engineering. Xianjun Yang, Xiao Wang, Qi Zhang, Linda Ruth Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. Shadow alignment: The ease of subverting safely-aligned language models. ArXiv, abs/2310.02949, 2023b. URLhttps://api.semanticscholar.org/CorpusID: 263620436. Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.High-Confidence Computing, 4(2):100211, June 2024. ISSN 2667-2952. doi: 10.1016/j.hcc.2024.100211. URL http://dx.doi.org/10.1016/j.hcc.2024.100211. Xin Yi, Shunfan Zheng, Linlin Wang, Xiaoling Wang, and Liang He. A safety realignment framework via subspace-oriented model fusion for large language models.arXiv preprint arXiv:2405.09055, 2024. Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, and Zhaopeng Tu. Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training.arXiv preprint arXiv:2407.09121, 2024. Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang. Removing rlhf protections in gpt-4 via fine-tuning.ArXiv, abs/2311.05553, 2023. URLhttps: //api.semanticscholar.org/CorpusID:265067269. Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen. A comprehensive study of knowledge editing for large language models, 2024a. URLhttps://arxiv.org/abs/2401.01286. 18 Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. Instruction tuning for large language models: A survey, 2024b. URLhttps://arxiv.org/abs/2308.10792. Weixiang Zhao, Yulin Hu, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, and Tat-Seng Chua. Towards comprehensive and efficient post safety alignment of large language models via safety patching.arXiv preprint arXiv:2405.13820, 2024. Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. On prompt-driven safeguarding for large language models. 2024a. URL https://api.semanticscholar.org/CorpusID:267334949. Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. On prompt-driven safeguarding for large language models, 2024b. URL https://arxiv.org/abs/2401.18018. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. Minjun Zhu, Linyi Yang, and Yue Zhang. Personality alignment of large language models.arXiv preprint arXiv:2408.11779, 2024. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency.arXiv preprint arXiv:2310.01405, 2023a. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023b. URLhttps://arxiv. org/abs/2307.15043. Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers, 2024. URLhttps://arxiv.org/abs/2406.04313. AAPPENDIX A.1LIMITATIONS While SafetyLock demonstrates promising results in maintaining the safety of fine-tuned language models, it is important to acknowledge several limitations. Primarily, SafetyLock requires access to both model weights and intermediate activations for implementation, which may limit its applicability in scenarios where such access is restricted or unavailable. Additionally, the method employs a symmetric locking mechanism; consequently, if an unauthorized party gains access to the model weights or activation values, they could potentially reverse-engineer the process to unlock and bypass SafetyLockâs protections. Lastly, while SafetyLock shows strong performance against current attack methods, its long-term robustness against evolving adversarial techniques remains to be studied. These limitations present opportunities for future work to enhance and expand the capabilities of SafetyLock, ensuring its continued effectiveness in maintaining AI safety. A.2CONSISTENCY OFHARMLESSNESSDIRECTIONS INFINE-TUNEDMODELS To validate SafetyLockâs effectiveness, we conducted a comprehensive analysis of the original Llama- 3-Instruct 8B model and its fine-tuned versions under various risk levels. Our experimental setup was as follows: We first extracted activation values from the 31st layer, 26th head of the Llama-3-8B Instruct model, which we identified as the most sensitive to harmlessness through linear regression, achieving the highest binary classification accuracy. We then performed forward computation on a safety dataset, saving the activation values of the last token for both safe and unsafe samples. Using 2D PCA for 19 dimensionality reduction, we visualized the shift in activation values between safe and unsafe samples by connecting their center points with arrows, illustrating both the direction and magnitude of the shift. Remarkably, we observed high similarity in these shifts across different risk levels (i.e., fine-tuning on data from different domains). To quantitatively assess the similarity between the safety directions found in the original model and those in the fine-tuned models, we employed KL divergence: D KL (P||Q) = X i P(i) log P(i) Q(i) (6) wherePandQrepresent the distributions of safety directions in the original and fine-tuned models, respectively. To further illustrate the change in similarity during the fine-tuning process, we employed one- dimensional linear interpolation of weights (Peng et al., 2024). This method allows us to smoothly transition from the original model weights to the fine-tuned model weights, providing insight into how the safety directions evolve during the fine-tuning process. The interpolation is defined as: θ Îą =θ+Îą(θ Ⲡâθ)(7) whereθrepresents the weights of the original Llama-3 model,θ Ⲡthe weights of the fine-tuned model, andÎąâ[â0.2,1.2]is the interpolation parameter. We extendÎąslightly beyond the [0, 1] range to observe trends slightly before and after the actual interpolation points. The interpolation process is implemented as follows: 1.We first extract the state dictionaries of both the base model (θ) and the fine-tuned model (θ Ⲡ). 2. For each layer, we compute the difference vector:d 1 =θ Ⲡâθ. 3. We then create new weights for eachÎąvalue:θ Îą =θ+Îąd 1 . 4.These new weights are used to reconstruct a new state dictionary, maintaining the original structure and naming conventions of the model. We use these interpolated models to compute the KL divergence between the safety directions of the original model and the interpolated models at each step. This results in a smooth curve showing how the similarity of safety directions changes as the model transitions from its original state to the fine-tuned state. BMATHEMATICALEXPLANATION OFSAFETYLOCKâSEFFECTIVENESS IN SUPPRESSINGHARMFULOUTPUTS In this section, we provide a mathematical justification for why SafetyLock can extract transferable safety directions from the original language model and effectively apply them to fine-tuned models to suppress harmful outputs. Our explanation is grounded in the properties of Transformer-based language models and the nature of fine-tuning on limited datasets. B.1ACTIVATIONSPACE ANDSAFETYDIRECTIONS Let us denote the activations of the original (pre-fine-tuned) language model at layerland headhas x l,h âR D , whereDis the dimensionality of the headâs output. During inference, these activations encode information about the generated tokens. We define two sets of activations corresponding to safe and unsafe responses: 20 X safe = n x safe,i l,h o N safe i=1 ,(8) X unsafe = n x unsafe,i l,h o N unsafe i=1 ,(9) whereN safe andN unsafe are the numbers of safe and unsafe samples, respectively. We compute thesafety directionθ l,h âR D as the mean difference between the activations for safe and unsafe responses: θ l,h = 1 N safe N safe X i=1 x safe,i l,h â 1 N unsafe N unsafe X i=1 x unsafe,i l,h .(10) This vector represents the average shift in activation space needed to move from an unsafe response towards a safe one. B.2PRESERVATION OFSAFETYDIRECTIONSDURINGFINE-TUNING Fine-tuning a language model on a new dataset modifies its parameters to adapt to specific tasks or domains. However, when the fine-tuning dataset is limited in size or scope, the changes to the modelâs internal representations are often localized and do not significantly alter the global structure of the activation space (Golechha & Dao, 2024; Godfrey et al., 2022). Let Ěx l,h denote the activations of the fine-tuned model at layerland headh. Empirically, we observe that there exists a strong linear relationship between the activations of the original and fine-tuned models: Ěx l,h âx l,h + âx l,h ,(11) whereâx l,h represents the change in activations due to fine-tuning, which is relatively small in magnitude compared tox l,h for many dimensions. Moreover, the safety directionθ l,h computed from the original model remains relevant in the fine- tuned model because the relative differences between safe and unsafe activations are preserved: Ě Î¸ l,h = Ěx safe l,h â Ěx unsafe l,h â x safe l,h âx unsafe l,h =θ l,h .(12) This approximation holds under the assumption that fine-tuning does not disproportionately affect the dimensions critical for encoding safety-related information. B.3EFFECTIVENESS OFACTIVATIONINTERVENTION During inference with the fine-tuned model, we intervene by adjusting the activations along the safety direction: Ěx intervened l,h = Ěx l,h +Îą(Ď l,h âθ l,h ),(13) where: â˘ÎąâRis the scaling factor controlling the intensity of the intervention. â˘Ď l,h âR D is the standard deviation vector of activations along each dimension, capturing the typical variability. â˘âdenotes element-wise multiplication. This adjustment effectively shifts the activations towards regions in the activation space associated with safe responses. Since the safety directionθ l,h is approximately preserved in the fine-tuned model, this intervention remains effective. 21 B.4IMPACT ONOUTPUTPROBABILITIES The language model generates the next token based on a probability distribution computed from the final activations. Adjusting the activations as in Equation equation 13 influences the logitszâR V (whereVis the vocabulary size) before the softmax function: z intervened =z+W head (Îą(Ď l,h âθ l,h )),(14) whereW head âR VĂD is the weight matrix projecting activations to logits. The adjustmentâz=W head (Îą(Ď l,h âθ l,h ))biases the logits towards tokens that are more likely in safe responses and away from those prevalent in unsafe responses. B.5SUPPRESSINGHARMFULOUTPUTS The probability of generating a harmful tokent harm is given by: P(t harm ) = exp z intervened t harm P V i=1 exp z intervened i .(15) By decreasingz intervened t harm relative to other logits, we reduceP(t harm ). Since the intervention shifts the activations towards safe regions, the logits for harmful tokens are decreased, and the model is less likely to generate harmful outputs. B.6TRANSFERABILITYACROSSMODELS The key to SafetyLockâs transferability lies in the similarity of safety directions between the original and fine-tuned models. Since the fine-tuning process does not significantly alter the relative positions of safe and unsafe activations in the activation space (as per Equation equation 12), the safety directions computed from the original model remain effective when applied to the fine-tuned model. This property is supported by empirical observations of low KullbackâLeibler (KL) divergence between the activation distributions of the original and fine-tuned models (see Figure 2 in Section 3.3). The minimal divergence indicates that the overall structure of the activation space, especially along dimensions relevant to safety, is preserved during fine-tuning. B.7CONCLUSION Mathematically, SafetyLock leverages the preserved safety directions in the activation space to adjust the modelâs internal computations towards generating safe outputs. By intervening along these directions, we effectively suppress harmful responses without requiring retraining or fine-tuning of the model. The minimal changes to the activation distributions during fine-tuning ensure that the safety directions remain applicable, allowing for efficient and transferable safety interventions across different models and fine-tuning scenarios. This theoretical explanation provides a foundation for understanding the effectiveness of SafetyLock in suppressing harmful outputs while maintaining the modelâs overall performance on benign tasks. CTHERISKS OFFINE-TUNINGLLMS ANDEXPERIMENTALSETUP HEx-PHI (Qi et al., 2023b) is based on 11 categories of prohibited use cases merged from Metaâs Llama-3 acceptable use policy and OpenAIâs usage policies: (1) Illegal Activity, (2) Child Abuse Content, (3) Hate, Harass, Violence, (4) Malware, (5) Physical Harm, (6) Economic Harm, (7) Fraud, Deception, (8) Adult Content, (9) Political Campaigning, (10) Privacy Violation Activity, and (11) Tailored Financial Advice. The dataset includes 30 examples per category, totaling 330 examples. This ensures a comprehensive safety evaluation aligned with industry-standard usage policies. For Risk-1, we use negative samples from the H-RLHF preference dataset. We select 10, 100, 1000, and 10000 samples respectively and trained for 5 epochs with a learning rate of 2e-5. For Risk-2, 22 we use 10 samples from Qi et al. (2023b) and trained for 5 epochs with a learning rate of 2e-5. For Risk-3, we use the first 50,000 samples from the Alpaca dataset (Wang et al., 2023b) and trained for 5 epochs with a learning rate of 2e-5 1 . Recognizing the potential of existing approaches to address safety issues in fine-tuned language models, we conducted comparative analyses across two categories as the same time: training-based and inference-time methods. For training-based approaches, we evaluated PPO, DPO, SFT (with safety data mixed during fine-tuning), SFT (with safety data mixed post-fine-tuning), and model- editing. Inference-time methods included ICD, PPL, Paraphrase, Retokenization, Safe-Reminder, and Self-Exam. These methods were assess based on efficiency, attack sample rejection rate, and normal text rejection rate, providing a comprehensive evaluation of their effectiveness in maintaining model safety while preserving functionality. This multi-faceted approach allows us to rigorously examine the trade-offs between safety and performance. Specifically, to ensure reproducibility, we followed past experimental settings and use 2000 safety data points from Bianchi et al. (2024) for SFT experiments. We considered two experimental settings for SFT. The first is After Training, which simulates the scenario where safety disappears after fine- tuning the language model and needs to be restored. This applies to all fine-tuned language models. The second is During Training, which simulates starting from the original model and requiring the mixing of additional safety data during training to prevent safety disappearance. However, the limitation of this method is that it still requires retraining for already fine-tuned language models. For PPO, we also use 2000 samples from Bianchi et al. (2024), and we use LlamaGuard-7b (Bhatt et al., 2023) as the Reward model. For DPO, based on the 2000 samples, we use samples generated by the fine-tuned language model (almost all of which are harmful) as negative samples for training. For the Model-Edited method, we use the most common Detoxifying with Intraoperative Neural Monitoring (DINM) method and followed the original setup using SafeEdit data 2 for editing. DANALYSIS OFSAFETYLOCKâSINTERVENTION 012345678910 Alpha Value 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Llama-3-8B: Harmfulness Avg. Score 012345678910 Alpha Value 0 20 40 60 80 100 Score Llama-3-8B: Harmfulness Avg. Rate 012345678910 Alpha Value 0 20 40 60 80 100 Score Llama-3-8B: AdvBench ASR 012345678910 Alpha Value 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Llama-3-70B: Harmfulness Avg. Score 012345678910 Alpha Value 0 20 40 60 80 100 Score Llama-3-70B: Harmfulness Avg. Rate 012345678910 Alpha Value 0 20 40 60 80 100 Score Llama-3-70B: AdvBench ASR Figure 7: Impact of SafetyLockâs intervention distance (Îą) on model safety metrics for Llama-3-8B and Llama-3-70B models. The graphs show Harmfulness Average Score, Harmfulness Average Rate, and AdvBench ASR across differentÎąvalues. Note that for these experiments, the intervention degree K is set to 24, indicating the number of attention heads influenced by SafetyLock. DistanceÎą.Our experimental results, as illustrated in Figure 7, demonstrate the significant influence of SafetyLockâs intervention distance (Îą) on model safety across different model sizes. For both Llama-3-8B and Llama-3-70B, we observe a clear U-shaped trend in harmfulness metrics asÎą 1 We use the official fine-tuning codehttps://github.com/meta-llama/llama-recipes 2 https://huggingface.co/datasets/zjunlp/SafeEdit 23 increases. Initially, asÎąrises from 0 to 4, thereâs a sharp decrease in harmfulness scores and rates, as well as the AdvBench ASR. This indicates that moderate intervention effectively enhances model safety. However, beyondÎą= 4, we see a gradual increase in these metrics, suggesting that excessive intervention may lead to unintended consequences, potentially disrupting the modelâs learned safety boundaries. Notably, Llama-3-70B exhibits more stability across differentÎąvalues compared to Llama-3-8B, implying that larger models may be more resilient to intervention adjustments. These findings underscore the importance of carefully calibrating SafetyLockâs intervention parameters to achieve optimal safety improvements while maintaining model performance, with an optimalÎąvalue around 4-6 for both model sizes. 024612244896 K Value 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Llama-3-8B: Harmfulness Avg. Score 024612244896 K Value 0 20 40 60 80 100 Llama-3-8B: Harmfulness Avg. Rate 024612244896 K Value 0 20 40 60 80 100 Llama-3-8B: AdvBench ASR 024612244896 K Value 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Llama-3-70B: Harmfulness Avg. Score 024612244896 K Value 0 20 40 60 80 100 Llama-3-70B: Harmfulness Avg. Rate 024612244896 K Value 0 20 40 60 80 100 Llama-3-70B: AdvBench ASR Figure 8: Impact of SafetyLockâs intervention degree (K) on model safety metrics for Llama-3-8B and Llama-3-70B models. The graphs illustrate the Harmfulness Average Score, Harmfulness Average Rate, and AdvBench ASR across different K values, ranging from 0 to 96. Lower scores indicate better safety performance. Note the rapid improvement in safety metrics as K increases from 0 to 6, followed by more gradual enhancements up to K=24, with a slight uptick at K=96 for some metrics. DegreeK.Our experiments, as illustrated in Figure 8, reveal the crucial role of SafetyLockâs intervention degree (K) in enhancing model safety across different model sizes. For both Llama-3-8B and Llama-3-70B, we observe a rapid improvement in safety metrics as K increases from 0 to 6, followed by a more gradual enhancement up to K=24. This trend is consistent across all three metrics: Harmfulness Average Score, Harmfulness Average Rate, and AdvBench ASR. For Llama-3-8B, the most significant improvements occur between K=0 and K=6, with the Harmfulness Average Score dropping from about 4.0 to 1.7, and the Harmfulness Average Rate decreasing from 70% to around 15%. The AdvBench ASR shows a similar sharp decline. Beyond K=6, the improvements become more incremental, with optimal performance generally achieved around K=24. Llama-3-70B exhibits a similar pattern but with overall lower harmfulness scores and rates. The initial drop in harmful metrics is less dramatic, suggesting that larger models may have inherently better safety characteristics. However, the trend of improvement with increasing K values remains consistent. Interestingly, for both model sizes, thereâs a slight uptick in harmfulness metrics for very high K values (K=96), particularly noticeable in the Llama-3-8B model. This suggests that excessive intervention might slightly degrade the modelâs learned safety boundaries, emphasizing the importance of finding an optimal K value. These findings underscore the effectiveness of SafetyLock in improving model safety, with the most significant gains achieved at relatively low K values (6-24). This implies that targeted intervention on a subset of attention heads can yield substantial safety improvements without the need for exhaustive modification of the model architecture. 24