Paper deep dive
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
Birong Pan, Mayi Xu, Qiankun Pi, Jianhao Chen, Yuanyuan Zhu, Ming Zhong, Tieyun Qian
Models: LLaMA2-7B-Chat, LLaMA3.1-8B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 5:52:41 PM
Summary
NeuronTune is a fine-grained safety alignment framework for LLMs that addresses the safety-utility trade-off by identifying and modulating specific safety-critical and utility-preserving neurons. It uses attack-aware attribution to pinpoint neurons and meta-learning (MAML) to adaptively adjust their activation scaling factors, allowing for tunable intervention scopes.
Entities (5)
Relation Signals (3)
NeuronTune → modulates → Neurons
confidence 95% · NeuronTune, a fine-grained framework that dynamically modulates sparse neurons
NeuronTune → uses → MAML
confidence 95% · we adopt an adaptive regulatory mechanism driven by MAML
NeuronTune → outperforms → DINM
confidence 90% · Extensive experimental results demonstrate that our method significantly outperforms existing state-of-the-art technologies
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). However, current techniques fundamentally suffer from intertwined deficiencies: insufficient robustness against malicious attacks, frequent refusal of benign queries, degradation in generated text quality and general task performance--the former two reflecting deficits in robust safety and the latter constituting utility impairment. We trace these limitations to the coarse-grained layer-wise interventions in existing methods. To resolve this, we propose NeuronTune, a fine-grained framework that dynamically modulates sparse neurons to achieve simultaneous safety-utility optimization. Our approach first identifies safety-critical and utility-preserving neurons across all layers via attribution, then employs meta-learning to adaptively amplify safety-neuron activations and suppress utility-neuron activations. Crucially, NeuronTune enables tunable adjustment of intervention scope via neuron-count thresholds, supporting flexible adaptation to security-critical or utility-priority scenarios. Extensive experimental results demonstrate that our method significantly outperforms existing state-of-the-art technologies, achieving superior model safety while maintaining excellent utility.
Tags
Links
- Source: https://arxiv.org/abs/2508.09473
- Canonical: https://arxiv.org/abs/2508.09473
Trouble viewing inline? Open PDF directly →
Full Text
56,991 characters extracted from source content.
Expand or collapse full text
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs Birong Pan, Mayi Xu, Qiankun Pi, Jianhao Chen, Yuanyuan Zhu, Ming Zhong, Tieyun Qian * School of Computer Science, Wuhan University Zhongguancun Academy panbirong, qty@whu.edu.cn Abstract Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). However, current techniques fundamentally suffer from intertwined deficiencies: insufficient robustness against malicious attacks, frequent refusal of benign queries, degradation in generated text quality and general task perfor- mance—the former two reflecting deficits inrobust safety and the latter constitutingutility impairment. We trace these limitations to the coarse-grained layer-wise interven- tions in existing methods. To resolve this, we proposeNeu- ronTune, a fine-grained framework that dynamically mod- ulates sparse neurons to achieve simultaneous safety-utility optimization. Our approach first identifies safety-critical and utility-preserving neurons across all layers via attribution, then employs meta-learning to adaptively amplify safety- neuron activations and suppress utility-neuron activations. Crucially, NeuronTune enables tunable adjustment of inter- vention scope via neuron-count thresholds, supporting flexi- ble adaptation to security-critical or utility-priority scenarios. Extensive experimental results demonstrate that our method significantly outperforms existing state-of-the-art technolo- gies, achieving superior model safety while maintaining ex- cellent utility. 1 Introduction Large Language Models (LLMs) (Touvron et al. 2023a; Hurst et al. 2024; Yang et al. 2025) demonstrate remarkable capabilities across diverse tasks, yet remain highly vulnera- ble to malicious attacks such as jailbreaking (Jin et al. 2025; Zhang et al. 2024; Xiao et al. 2024), which can induce harm- ful or uncontrolled outputs. Ensuring robust safety align- ment with human values is therefore a critical prerequisite for their secure and reliable deployment. Current safety alignment techniques (Hazra et al. 2024; Zhang et al. 2025; Cao, Yang, and Zhao 2024) face fun- damental challenges in simultaneously achieving safety ro- bustness and utility preservation (Varshney et al. 2024; R ̈ ottger et al. 2024). These issues manifest as inadequate resistance to malicious attacks, exaggerated stringency in rejecting benign queries, generated responses of low qual- ity, severely limiting practical applicability, as illustrated in Figure 1. We contend that these limitations stem from the * Corresponding author. Adversarial Attacks Benign Queries Vanilla LLM f Aligned LLM f’ Sufficient Safety Exaggerated Safety Utility Figure 1: Current models, after safety alignment, suffer from exaggerated safety and utility degradation. Here, sufficient safety refers to the ability to defend against adversarial at- tacks, exaggerated safety denotes the undesirable refusal to benign queries, and utility encompasses the usefulness of re- sponses and performance on general tasks. coarse-grained, layer-level intervention strategies. Such ap- proaches typically identify predefined ‘critical layers’ using techniques like hidden-state comparisons or steering vectors (Wang et al. 2024a; Cao, Yang, and Zhao 2024), and then apply uniform modulation across all layers. Coarse-grained adjustments fail to precisely pinpoint the key factors related to safety and utility, while uniform layer-wise interventions cannot account for the varying degrees of adjustment re- quired for different influencing factors, thereby perpetuat- ing the safety-utility trade-off. To effectively mitigate these issues, it becomes imperative to precisely identify and judi- ciously modulate significant factors. Therefore, we propose a novel neuron-level safety align- ment framework namedNeuronTune. Drawing from the concept of knowledge neurons (Dai et al. 2022), we posit that safety-critical features and utility-preserving knowledge are intrinsically stored within specific neurons. We iden- tify them as the factors impacting the safety-utility balance. Firstly, to pinpoint these neurons, we propose an attack- aware attribution method. This is motivated by analytical results 1 that while LLMs demonstrate a remarkable ability to avoid generating harmful responses from direct queries, their robustness often falters against sophisticated adversar- ial attacks. Specifically, we trace safety-crucial and utility- related neurons by performing attribution on the safe and 1 Please refer to Appendix for detailed experimental results. arXiv:2508.09473v1 [cs.LG] 13 Aug 2025 useful responses when harmful and benign queries are sub- jected to adversarial attacks. Furthermore, to avoid the failure of uniform regulation to account for the roles of different influencing factors, we adopt an adaptive regulatory mechanism driven by MAML (Finn, Abbeel, and Levine 2017). Unlike the standard meta- learning which often adapts entire model parameters for new tasks, we propose to apply it to optimize the scaling fac- tors of pre-identified, sparse, and critical neurons. This fine- grained control is crucial for avoiding the pitfalls of coarse- grained interventions. Moreover, we design a dynamic con- trol mechanism, which allows users to flexibly tune the inter- vention scope, such as tightening safety by modulating more safe neurons in high-risk scenarios, or prioritizing utility via conservative neuron selection. This configurability provides a flexible adjustment for diverse deployment needs. The primary contributions are summarized as follows: •Problem Diagnosis: We reveal inherent limitations of layer-wise intervention strategies. •Method Innovation: We introduce a fine-grained, neuron- level alignment framework that integrates neuron local- ization based on attack-aware attribution with adaptive neuron modulation guided by meta-learning, striking a delicate balance between safety and utility. •Tunable Mechanism: Our design introduces a tunable in- tervention system, which facilitates model adaptation to diverse safety and utility demands through the regulation of neuron counts. •Empirical Validation: Extensive experiments on multiple benchmarks and LLMs demonstrate superior effective- ness and adaptability over strong baselines. 2 Preliminaries In this section, we begin by examining the limitations of coarse-grained layer-wise interventions on the problems of exaggerated safety and utility degradation. This analysis motivates the introduction of neuron, elucidating how the fine-grained nature of individual neurons holds promise for achieving a balance between robust safety and utility. 2.1 Analysis on Coarse-Grained Interventions Existing safety alignment methods often resort to coarse- grained interventions, such as modifying entire layers. While these approaches aim to enhance safety, they fre- quently introduce undesirable side effects, notably exagger- ated safety and utility degradation. To empirically investigate the impact of such coarse- grained interventions, we conduct an analysis using the DINM baseline method (Wang et al. 2024b) on LLaMA-3.1- 8B-Instruct (Grattafiori et al. 2024). We select DINM for this empirical investigation since it is a representative method that employs layer-wise interventions to mitigate unsafe be- haviors. Moreover, DINM allows to vary the percentage of parameters updated within a layer, which enables us to sys- tematically observe the direct consequences of such inter- ventions on both safety and utility. Table 1 shows the re- sults by employing DINM to edit the varying percentages of model parameters within a layer. Methods SafeEditAlpacaMMLU Refuse Rate↑Entropy↑Refuse Rate↓Entropy↑Accuracy↑ DINM (100%)100.00%1.132 bit98%1.089 bit0.00% DINM (50%)100.00%1.142 bit100%1.233 bit0.01% DINM (20%)99.49%1.616 bit86%1.778 bit60.00% DINM (10%)82.46%2.545 bit42%3.760 bit59.64% DINM (5%)42.67%5.019 bit4%5.520 bit61.97% Table 1: Results of employing DINM to edit varying per- centages of parameters within a layer on LLaMA3.1-8B- Instruct.↑indicates that higher values are better, while↓ indicates that lower values are better. As shown in Table 1, when a large percentage of pa- rameters are updated, the model achieves sufficient safety. However, this comes at a severe cost, i.e., significant exag- gerated safety and drastic utility degradation. This indicates that the model becomes overly cautious, refusing even harm- less requests and generating low-quality, repetitive, or un- informative text. As the percentage of updated parameters decreases, we observe a gradual mitigation of exaggerated safety and a recovery in general utility. However, this mit- igation of utility degradation is accompanied by a decline in sufficient safety, e.g., SafeEdit Refusal Rate decreases, implying that the model becomes less robust against harm- ful queries. This analysis clearly demonstrates that coarse- grained, layer-wise interventions struggle to simultaneously achieve robust safety and maintain high utility, often forcing a compromise between these two crucial aspects. In view of the aforementioned drawback of coarse grained, layer-wise interventions, we underscores the neces- sity for a more nuanced and fine-grained approach to safety alignment. To this end, we propose to first precisely pinpoint and then selectively intervene on the specific neurons asso- ciated with safety and utility. 2.2 Neurons in Transformer-based LLMs To address the limitations of coarse-grained interven- tions, we delve into the fundamental building blocks of LLMs: neurons. We offer an explanation for neurons in Transformer-based language models, which are central to our strategy. An LLMftypically consists of an embedding matrixE andLtransformer layers. Each layerℓincludes attention headsAttand a multilayer perceptionMLP. Given a se- quencew=⟨w 0 ,...,w t ⟩as input,ffirst appliesEto cre- ate the embeddingh i ∈R d for each tokenw i ∈w, which is then updated by attention heads and MLP blocks from sub- sequent layers (bias omitted): h l+1 i =h l i +Att l (h l i ) +MLP l (h l i +Att l (h l i )). (1) The MLPs in Transformer models we used are: MLP(x) = W ⊤ down (σ(W gate x)⊙W up x),(2) whereW down ,W gate ,W up ∈R d m ×d are projection matri- ces,σ(·)is activation function,⊙is element-wise product operator. In the context of neural networks, the term neuron refers to a single dimension of any activation. We choose to study neurons in the intermediate layer of MLP (activation be- fore down projection) since it has been shown such neurons encode meaningful and interpretable features (Wang et al. 2022; Dai et al. 2022; Gurnee et al. 2023). The fine-grained nature of these individual neurons, in contrast to entire lay- ers, offers a promising avenue for precise intervention, en- abling us to balance robust safety and general utility more effectively. 3 Method: NeuronTune Building upon the insights regarding the limitations of coarse-grained interventions, and recognizing that model capabilities are encoded within individual neurons (Dai et al. 2022), we propose NeuronTune, a novel fine-grained method for safety-utility alignment. Our approach addresses the imperative to achieve robust safety and mitigate utility degradation. NeuronTune consists of two primary stages. First,Pinpointing Safety and Utility Neurons via Attack- Aware Attribution, where we identify safety-crucial and utility-related neurons. Second,Editing Neurons via Adap- tive Activation Adjustment, where we dynamically adjust these neurons. Figure 2 provides an overview of our compre- hensive approach to achieving an effective balance between robust safety and utility preservation. Step 1. Pinpointing Neurons via Attack-Aware Attribution Step 2. Editing Neurons via Adaptive Activation Adjustment 1 × L Layer !" ! " =$ #,% ! " −$ % ! " & '([$ #,% ! " ++($ #,% ! " −$ % ! " )] '$ #,% ! " & ' /+ . . . . . . . . . Safety-Crucial Neurons Utility-Related Neurons Local Safety Dataset Local Utility Dataset 44 LsafeLutility 5 5 Local Update Local Update Ltotal 7 6 6 6 Global Update 1 . . . . . . . . . 2 3 × L Layer . . . . . . . . . Safety-Crucial Neurons Utility-Related Neurons <latexit sha1_base64="RPghMCVIvpzEe6lpP5bqoiic+d0=">AAAC13icjVHLSsNAFD2Nr1pfVZdugkVwVVKR6rKoC5cKViu1lsl0WkOnSUgmYinFnbj1B9zqH4l/oH/hnTEFtYhOSHLm3HvOzL3XDaUXK8d5zVgTk1PTM9nZ3Nz8wuJSfnnlNA6SiIsqD2Q1VwWC+n5oqo8JUUtjATruVKcud19HT+7FlHsBf6J6oei0WMd32t7nCmimvmVg+YgZm2h+sPLgQw4k8NmvuAUHbPscVBKQQHpOgryL7hACwE4EvQg4EMRlmCI6amjBAchcQ0MiIsIeSYuMESOtAllCcpgxHbp26FdPWV92mvP2Kg5nSLpjUhpY4M0AeVFhPVptoknxlmzv3kPjKe+W5/+burVI1bhiti/dKPM/+p0LQpt7JoaPKopNIyujqcuiemKvrn9pSpFDiFxGrcoHhHmRjnqs200sald95aZ+JvJ1Kze8zQ3wbu+JQ249HOc4+B0q1gqF8vH24XKXjrqLNawjk2a5w4qOMQRquR9g0c84dk6t26tO+v+M9XKpJpVfFvWwweJs5dz</latexit> D local safety <latexit sha1_base64="wlfZOgj8JnVGBR1sfL+BI13pjgk=">AAAC2HicjVHLSsNAFD3GV62vaJdugkVwVVKR6rKoC5cV7AO1lmQ6rYPTJCQToZSCO3HrD7jVLxL/QP/CO2MEH4hOSHLm3HvOzL3Xj6RIlOs+T1iTU9Mzs7m5/PzC4tKyvbLaSMI0ZrzOQhnGLd9LuBQBryuhJG9FMfcGvuRN/3Jfx5tXPE5EGByrYcTbA68fiJ5gniKqYxcOOqNUCSnUcHw+kiHz5LhjF92Sa5bzE5QzUES2aqH9hDN0EYIhxQAcARRhCQ8JPacow0VEXBsj4mJCwsQ5xsiTNqUsThkesZf07dPuNGMD2mvPxKgZnSLpjUnpYIM0IeXFhPVpjomnxlmzv3mPjKe+25D+fuY1IFbhgti/dB+Z/9XpWhR62DU1CKopMoyujmUuqemKvrnzqSpFDhFxGncpHhNmRvnRZ8doElO77q1n4i8mU7N6z7LcFK/6ljTg8vdx/gSNrVK5UqocbRere9moc1jDOjZpnjuo4hA11Ml7iHs84NE6sa6tG+v2PdWayDQFfFnW3Rv7ppgF</latexit> D local utility & Attack p Query q Safe r Query q Safe r <latexit sha1_base64="cD2DuMhTLAekh0kuG0jg+6Mn9ZY=">AAAC0XicjVHLSgMxFD2Or/quunQzWAQFKVOR6rLoxmVFW4X6YCZGDZ2XmUyhlIK49Qfc6k+Jf6B/4U1MwQeiGWbm5Nx7TnLvDdJQZMrzXoac4ZHRsfHCxOTU9MzsXHF+oZkluWS8wZIwkceBn/FQxLyhhAr5cSq5HwUhPwrauzp+1OEyE0l8qLopP438q1hcCuYros46573VdP1mTfTPemH/vFjyyp5Z7k9QsaAEu+pJ8RknuEAChhwROGIowiF8ZPS0UIGHlLhT9IiThISJc/QxSdqcsjhl+MS26XtFu5ZlY9prz8yoGZ0S0itJ6WKFNAnlScL6NNfEc+Os2d+8e8ZT361L/8B6RcQqXBP7l26Q+V+drkXhEtumBkE1pYbR1THrkpuu6Ju7n6pS5JASp/EFxSVhZpSDPrtGk5nadW99E381mZrVe2Zzc7zpW9KAK9/H+RM0N8qVarm6v1mq7dhRF7CEZazSPLdQwx7qaJC3xAMe8eQcOF3n1rn7SHWGrGYRX5Zz/w4LQZTo</latexit> v l (p,q)i Attack-Aware Attribution 2 <latexit sha1_base64="9gmYCCdQ981DRQgr/NyXxWnNV6U=">AAACzXicjVHLSsNAFD2Nr1pfVZdugkVwVRKR6rLoxp0V7ANrLUk6rYN5mUwKJdatP+BWf0v8A/0L74wpqEV0QpIz595zZu69dujyWBjGa06bmZ2bX8gvFpaWV1bXiusbjThIIofVncANopZtxczlPqsLLlzWCiNmebbLmvbNsYw3hyyKeeCfi1HIOp418HmfO5Yg6mLYTW/5+Cp1x91iySgbaunTwMxACdmqBcUXXKKHAA4SeGDwIQi7sBDT04YJAyFxHaTERYS4ijOMUSBtQlmMMixib+g7oF07Y33aS89YqR06xaU3IqWOHdIElBcRlqfpKp4oZ8n+5p0qT3m3Ef3tzMsjVuCa2L90k8z/6mQtAn0cqho41RQqRlbnZC6J6oq8uf6lKkEOIXES9ygeEXaUctJnXWliVbvsraXibypTsnLvZLkJ3uUtacDmz3FOg8Ze2ayUK2f7pepRNuo8trCNXZrnAao4Q118vbxiCc8a6daot1p95+pWi7TbOLb0h4+ADNZk9M=</latexit> v l qi <latexit sha1_base64="qpgPlT7y1zMfMiuw2AW+njG+Zek=">AAAC2HicjVHLSsNAFD2Nr1pf0S7dBIvgqqQi1WVRFy4r2AfWWibptAanSUgmQikFd+LWH3CrXyT+gf6Fd8YU1CI6IcmZc+85M/deJxReLG37NWPMzM7NL2QXc0vLK6tr5vpGPQ6SyOU1NxBB1HRYzIXn85r0pODNMOJs4AjecK6PVLxxw6PYC/wzOQx5e8D6vtfzXCaJ6pj5484oZj0uh+PLUV8EDhPjjlmwi7Ze1jQopaCAdFUD8wUX6CKAiwQDcPiQhAUYYnpaKMFGSFwbI+IiQp6Oc4yRI21CWZwyGLHX9O3TrpWyPu2VZ6zVLp0i6I1IaWGbNAHlRYTVaZaOJ9pZsb95j7SnutuQ/k7qNSBW4orYv3STzP/qVC0SPRzoGjyqKdSMqs5NXRLdFXVz60tVkhxC4hTuUjwi7GrlpM+W1sS6dtVbpuNvOlOxau+muQne1S1pwKWf45wG9d1iqVwsn+4VKofpqLPYxBZ2aJ77qOAEVdTIe4hHPOHZODdujTvj/jPVyKSaPL4t4+EDqWqX4w==</latexit> D global safety <latexit sha1_base64="CiO05PeHVMOWTlHNA3Z4wdM6Foo=">AAAC2HicjVHLSsNAFD2Nr1pf0S7dBIvgqqQi1WVRFy4r2AfWWpJ0WgenSUgmQikFd+LWH3CrXyT+gf6Fd8YU1CI6IcmZc+85M/deNxQ8lrb9mjFmZufmF7KLuaXlldU1c32jHgdJ5LGaF4ggarpOzAT3WU1yKVgzjJgzcAVruNdHKt64YVHMA/9MDkPWHjh9n/e450iiOmb+uDNKuOByOL4c9UXgOmLcMQt20dbLmgalFBSQrmpgvuACXQTwkGAABh+SsICDmJ4WSrAREtfGiLiIENdxhjFypE0oi1GGQ+w1ffu0a6WsT3vlGWu1R6cIeiNSWtgmTUB5EWF1mqXjiXZW7G/eI+2p7jakv5t6DYiVuCL2L90k8786VYtEDwe6Bk41hZpR1XmpS6K7om5ufalKkkNInMJdikeEPa2c9NnSmljXrnrr6PibzlSs2ntpboJ3dUsacOnnOKdBfbdYKhfLp3uFymE66iw2sYUdmuc+KjhBFTXyHuIRT3g2zo1b4864/0w1Mqkmj2/LePgA2faX9w==</latexit> D global uility Global Safety Dataset Global Uility Dataset Figure 2: The overview of our NeuronTune, containing pin- pointing neurons and adaptively modulating neurons. The numbered steps represent the sequential order of processing. 3.1 Pinpointing Safety and Utility Neurons via Attack-Aware Attribution While LLMs effectively avoid harmful content in response to direct harmful queries, their robustness often falters against sophisticated adversarial attacks (see Appendix for details), indicating misleading attacks as a primary cause of unsafe responses. Given that the model’s knowledge con- cerning safety and utility is intrinsically stored within its neurons, it follows that these attacks impact these specific neurons. For both harmful and benign queries, when sub- jected to attacks, the ability to still yield safe and useful responses signifies the role played by safety-crucial and utility-related neurons. To precisely identify these neurons, we propose an attack-aware attribution method that pin- points specific neurons whose activations are critical for pro- cessing adversarial contexts and ensuring desired safe and useful outputs. For safety-critical neurons, given an input pair(p,q), whereprepresents the misleading adversarial prompt andq represents the harmful question, we aim to guide the model towards a desired safe response, denoted asr a . Our approach calculates the contribution score of each neuron in perceiv- ing the prompt’s influence on the generation of the safe re- sponse. Initially, we take onlyqas input, record the activa- tion value of each neuron and denote it asv q l i . Subsequently, we input bothpandqinto the model and record the new ac- tivation value, denoted asv (p,q) l i . To calculate the contribu- tion scoreC(n l i ), we gradually change the activation value of a neuronn l i fromv q l i tov (p,q) l i when the input consists of both prompt and question. At the same time, the output probability of the model changes accordingly. We calculate the probability of the correct answer predicted by model, de- noted as: P(v l i ) =p(r a |p,q,A(n l i ) =v l i ),(3) wherev l i is a given value assigned to the neuron activation A(n l i ). We integrate the gradient of the probability during this process as the neuron’s contribution score, as follows: C(n l i ) = v (p,q) l i −v q l i R 1 α=0 ∂P [ v q l i +α ( v (p,q) l i −v q l i )] ∂v (p,q) l i dα(4) where ∂P [ v q l i +α ( v (p,q) l i −v q l i )] ∂v (p,q) l i calculates the gradient of the model probability with regard tov (p,q) l i ,αcontrols the inte- gration fromv q l i tov (p,q) l i . A higherC(n l i )indicates a more significant role in maintaining safety performance under ad- versarial conditions. We then select a subset of top-k neurons with the highest score as the safety-crucial neurons,N s . Similarly, for a set of benign, general knowledge prompts, we calculate the gradient of the log-likelihood of high- quality responses with respect to neuron activations. Neu- rons with high contribution scores in this context are identi- fied as utility-relatedN u , as they are essential for the general performance and ability to generate informative text. Specifically, to manage computational intensity, we adopt a strategic attack selection approach. Following (Wang et al. 2024b), we categorize existing adversarial attack types into distinct classes, each representing a common strategy to by- pass LLM safety mechanisms: attention shifting, pretending, and privilege escalation 2 . To ensure a diverse and represen- tative set of identified neurons without incurring excessive computational overhead, we select one representative data point from each of these adversarial attack categories for conducting the gradient attribution analysis. 2 Please refer to Appendix for details on adversarial attack types. 3.2 Editing Neurons via Adaptive Activation Adjustment Having identified the safety-crucial and utility-related neu- rons, the next pivotal step in our NeuronTune method is to edit these neurons to achieve a harmonious balance be- tween robust safety and preserved utility. The magnitude of suppressing or amplifying specific neuron activations no- tably affects the expression of the corresponding knowledge. Therefore, dynamic modulation of individual neurons is es- sential to precisely control this balance. For each identified neuronn l i , whether safety-crucial or utility-related, we introduce a learnable scaling factorα j . This scaling factor directly modulates the neuron’s activa- tion. Conceptually, the original activation of neuronn l i , its modulated activationv l i ′ becomesv l i ′ =α j v l i . By adjustingα j , we can either enhance or suppress the influence of a specific neuron. Initially, safety-crucial neu- rons are set with an enhancing factor (α j >1) while utility-related neurons are initialized with a suppressing fac- tor (α j <1). This initial bias guides the model towards safer outputs while attempting to mitigate utility degradation. Our approach is the adaptation of MAML (Finn, Abbeel, and Levine 2017) to the neuron-level modulation of LLMs for safety-utility alignment. Unlike standard meta-learning which often adapts entire model parameters for new tasks, we specifically apply it to optimize the scaling factors of pre-identified, sparse, and critical neurons. This fine-grained control is crucial for avoiding the pitfalls of coarse-grained interventions. Our meta-learning process is in Algorithm 1. Furthermore, to accommodate diverse practical applica- tion scenarios and varying demands for safety or utility, NeuronTune features a tunable mechanism. This mechanism allows for the flexible selection of the number of neurons to be regulated. By controlling the quantity of modulated neurons, users can fine-tune the model’s behavior, empha- sizing either heightened safety, achieved by regulating more safety-critical neurons, or enhanced utility, by prioritizing utility-related neurons, thereby adapting to specific deploy- ment requirements. 4 Experiment 4.1 Experimental Setup ModelsWe select four representative general LLMs: LLaMA2-7B-Chat (Touvron et al. 2023b), LLaMA3.1-8B- Instruct (Grattafiori et al. 2024), Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct (Yang et al. 2024), to thoroughly evaluate the effectiveness and scalability of our NeuronTune in balancing safety and utility. Safety and Utility Evaluation DatasetsWe evaluate model performance from two critical perspectives: robust safety and utility. The former encompasses sufficient safety and exaggerated safety. For sufficient safety, which mea- sures the ability to resist generating harmful content, we uti- lize SafeEdit (Wang et al. 2024b) and AdvBench (Zou et al. 2023), two widely recognized safety benchmarks. For ex- aggerated safety, we consider two prominent benchmarks: Alpaca (Taori et al. 2023) and TruthfulQA (Lin, Hilton, and Algorithm 1: Adaptive Neuron Modulation Require: Base modelM Datasets:D global safety ,D global utility ,D local safety ,D local utility Safety neuronsN s , Utility neuronsN u Initial scaling factors:α (init) s >1,α (init) u <1 Ensure:Optimized scaling factorsΘfor neurons 1:foreach neuronn i ∈N s do 2:θ i ←α (init) s ▷Enhancement factor>1 3:end for 4:foreach neuronn j ∈N u do 5:θ j ←α (init) u ▷Suppression factor<1 6:end for 7:forepoch= 1toEdo 8:Sample batches:B s ∼D global safety ,B u ∼D global utility 9:Clone parameters:Θ ′ ←Θ 10:forstep= 1toKdo 11:ApplyΘ ′ toM 12:L local s ←EvalSafety(M,D local safety ) 13:Update safety neurons:θ ′ i ←θ ′ i −η inner ∇ θ ′ i L local s 14:ApplyΘ ′ toM 15:L local u ←EvalUtility(M,D local utility ) 16:Update utility neurons:θ ′ j ←θ ′ j +η inner ∇ θ ′ j L local u 17:end for 18:ApplyΘ ′ toM,L s ←EvalSafety(M,B s ) 19:L u ←EvalUtility(M,B u ) 20:L joint ←λL s + (1−λ)L u 21:UpdateΘ←Θ−η meta ∇ Θ L joint 22:end for 23:Apply optimalΘtoM 24:returnOptimized modelM ∗ Evans 2022). Specifically, we leverage the benign queries within these datasets. Our evaluation quantifies the rate at which models incorrectly refuse to answer these benign in- puts. Furthermore, we choose MMLU (Hendrycks et al. 2020) to evaluate whether methods would influence model general performance since its comprehensive coverage of knowledge-intensive tasks. 3 BaselinesWe compare NeuronTune with the following baselines. (1) DINM (Wang et al. 2024b) is a editing ap- proach to identify toxic regions and mitigate unsafe behav- iors in LLMs. (2) SCANS (Cao, Yang, and Zhao 2024) ap- plies refusal steering vectors to identify and mitigate such exaggerated safety behaviors in LLMs. (3) CAVGAN (Li et al. 2025) leverages the security judgment boundary to achieve efficient defense. MetricsFor both sufficient safety and exaggerated safety, we employ the Refusal Rate, defined as the proportion of queries (whether harmful or benign) that are rejected. The determination of whether a query should be rejected is used a classifier (Wang et al. 2024b), which evaluates the safety 3 Please refer to the Appendix for more details about the datasets and experimental environment. ModelsMethods SafeEditAlpaca SU-F1↑ Refuse Rate↑Entropy↑Refuse Rate↓Entropy↑ LLaMA2-7B-Chat Default38.56%5.482 bit1%5.736 bit0.698 DINM90.97%4.125 bit41%4.799 bit0.623 SCANS97.85%3.038 bit1%3.718 bit0.534 CAVGAN84.41%5.213 bit20%5.705 bit0.748 NeuronTune91.59%5.014 bit1%5.143 bit0.770 LLaMA3.1-8B-Instruct Default10.77%6.832 bit2%6.077 bit0.660 DINM42.67%5.019 bit4%5.520 bit0.675 SCANS99.28%1.490 bit0%1.401 bit0.128 CAVGAN30.67%6.388 bit3%6.189 bit0.715 NeuronTune49.74%5.317 bit3%5.912 bit0.722 Qwen2.5-7B-Instruct Default14.87%6.518 bit3%6.345 bit0.670 DINM91.79%6.187 bit57%6.208 bit0.731 SCANS35.90%4.810 bit0%4.946 bit0.635 CAVGAN40.82%6.787 bit13%6.620 bit0.740 NeuronTune53.13%6.067 bit2%6.096 bit0.768 Qwen2.5-14B-Instruct Default14.97%6.704 bit2%6.608 bit0.685 DINM45.95%6.734 bit6%6.846 bit0.784 SCANS45.33%6.130 bit3%6.602 bit0.764 CAVGAN58.77%6.894 bit11%6.706 bit0.808 NeuronTune65.03%6.592 bit4%6.379 bit0.824 Table 2: Main results on safety and utility evaluation across different LLMs.Bolddenotes the best performance in each block. ↑indicates that higher values are better, while↓indicates that lower values are better. Our NeuronTune consistently achieves the highest SU-F1 score, underscoring its ability to effectively balance robust safety and utility preservation. of response content and rejects those deemed unsafe. For utility, we assess three complementary dimensions. (1) In- formation Content: Measured by entropy of generated text, reflecting the diversity and richness of the output. (2) Flu- ency: Evaluated using perplexity-based metrics to assess the linguistic smoothness and naturalness of responses gener- ated by the aligned LLM. (3) General Capability: Evalu- ated on MMLU, selected for its comprehensive coverage of knowledge-intensive tasks across various domains. Robust safety performance is represented by summing the Refusal Rate for harmful queries and (1 - Refusal Rate) for benign queries. Utility is primarily represented by the entropy of the generated text for both harmful and benign queries. Given that these performance indicators have differ- ent units and value ranges, we first normalize both metrics to a 0-1 scale and then adopt an F1-score-like calculation method to derive a single, comprehensive metric for evaluat- ing the overall balance between robust safety and utility. We term this combined metric the Safety-Utility F1 (SU-F1). 4.2 Main Results NeuronTune effectively achieves a balance between ro- bust safety and utility preservationAs shown in Table 2, our NeuronTune consistently achieves the highest SU- F1 score across all evaluated LLMs 4 . In summary, the experimental results clearly demonstrate that our Neuron- Tune consistently achieves a superior balance between ro- bust safety and preserved utility. By precisely modulating in- 4 The experiments on AdvBench and TruthfulQA also validate the effectiveness of our method, please refer to Appendix. dividual neurons, our method effectively enhances sufficient safety while simultaneously mitigating exaggerated safety and maintaining high general performance, thereby outper- forming existing coarse-grained intervention methods and addressing the intertwined deficiencies. Performance on Other Utility MetricsIn our primary experiments, utility is assessed using the entropy of gener- ated text. To more comprehensively evaluate the utility of the aligned model, we additionally assess accuracy on MMLU and the fluency of generated text. The results for all methods are shown in Table 3. While NeuronTune’s MMLU accuracy is not the highest compared to DINM, it is crucial to consider the overall context of safety alignment. Our method signif- icantly mitigates the exaggerated safety issue, which often comes at a substantial cost to utility in other approaches. Therefore, despite being suboptimal on MMLU alone, Neu- ronTune’s ability to maintain a strong balance between ro- bust safety and utility preservation demonstrates its overall effectiveness and practical advantage. 4.3 Ablation Study To validate the effectiveness of each component, we con- duct an ablation study of NeuronTune when removing strate- gic attack selection (w/o Attack Selection), attribution-based neuron identification (w/o Neuron Pinpointing), and meta- learning-driven adaptive adjustment (w/o Adaptive Adjust- ment) respectively. The results in Table 4 confirm that each component of NeuronTune is indispensable for achieving its superior safety-utility balance. Removing any of these com- ponents leads to a significant degradation in overall perfor- MethodsAccuracy↑Fluency↑Entropy↑Avg.↑ Defaults64.438.3112.90928.55 DINM61.975.9610.53926.16 SCANS38.831.452.89114.39 CAVGAN36.137.2712.58818.66 NeuronTune60.866.3411.22926.14 Table 3: General capability, fluency, and entropy of LLaMA- 3.1-8B-Instruct across various safety alignment methods, which highlights how different alignment strategies impact model’s utility. The best results are inboldand the second best ones are in underlined . Methods SafeEditAlpaca SU-F1↑ Refuse Rate↓Entropy↑Refuse Rate↓Entropy↑ NeuronTune49.74%5.317 bit3%5.912 bit0.722 w/o Attack Selection99.38%2.007 bit3%2.174 bit0.289 w/o Neuron Pinpointing8.51%6.556 bit1%6.194 bit0.652 w/o Adaptive Adjustment93.74%1.667 bit0%2.109 bit0.239 Table 4: Ablation study for NeuronTune on LLaMA3.1- 8B-Instruct. w/o Attack Selection, w/o Neuron Pinpoint- ing, w/o Adaptive Adjustment removes strategic selection of data points, attribution-based neuron identification, and meta-learning-driven adaptive adjustment, respectively. mance, failing to find the optimal trade-off. This validates the synergistic design of NeuronTune. 4.4 Adaptability Across Diverse Scenarios via Neuron Number Regulation This section details the experimental evaluation of Neuron- Tune’s tunable mechanism, which facilitates model adapta- tion to diverse safety and utility demands through the regula- tion of neuron counts. Our approach allows for independent control over the number of safety-crucial and utility-related neurons. Specifically, after obtaining the contribution scores for each neuron via gradient attribution, we sort them by score in descending order and select the top-k neurons for modulation. To assess the impact of varying neuron counts, we conduct experiments by dynamically adjusting the num- ber of safety-crucial and utility-related neurons to 500, 1000, 1500, and 2000 based on a balance of empirical observation. From the results 5 , we observe two key phenomena. Impact of Increasing Safety-Crucial Neurons while Keeping Utility Neurons FixedWhen the number of utility-related neurons is fixed, and the number of safety- crucial neurons is progressively increased, the model’s per- formance on the adversarial attack dataset significantly im- proves, indicating enhanced sufficient safety. For instance, with 500 utility neurons, increasing safety neurons from 500 to 2000 boosts SafeEdit Refuse Rate from 41.33% to 83.69%. However, this also leads to a higher tendency to refuse benign queries , and a decrease in the information content of the generated text. This trade-off highlights that while more safety neurons enhance defense, they can also contribute to over-safety and reduced text quality. 5 Please refer to Appendix for detailed experimental results. Impact of Increasing Utility-Related Neurons while Keeping Safety Neurons FixedWhen the number of safety-crucial neurons is fixed, and the number of utility- related neurons is progressively increased, sufficient safety also enhances, albeit to a lesser extent compared to increas- ing safety-crucial neurons. More importantly, the issues of exaggerated safety and degradation in text quality exhibit fluctuations. Even when a decrease is observed, its magni- tude is generally smaller than that observed when increasing safety-crucial neurons with fixed utility-related ones. These observations can be attributed to several factors. Firstly, there indeed exist neurons that are strongly corre- lated with either safety or utility. Therefore, increasing the count of safety-related neurons directly enhances sufficient safety, while increasing utility-related neurons helps curb exaggerated safety and mitigate utility degradation. Sec- ondly, it is plausible that some neurons may simultaneously store both safety-related and utility-related knowledge, or at least play a more critical role for one type of knowledge. This inherent overlap or multi-functionality can lead to the observed fluctuations in performance, even when increas- ing utility-related neurons, as the modulation might inadver- tently affect intertwined safety aspects. Overall, the tunable mechanism offers a practical ap- proach to adapt NeuronTune to different application scenar- ios. For instance, in high-security demand scenarios, one can choose to modulate a larger number of safety-crucial neu- rons to prioritize robust defense. Conversely, in applications where maintaining high conversational quality and helpful- ness is paramount, a greater emphasis can be placed on mod- ulating utility-related neurons. 4.5 Analysis of Neuron Distribution To gain a deeper understanding of how safety and utility ca- pabilities are encoded within LLMs, we analyze the layer distribution of the identified neurons. This analysis provides insights into the architectural regions that are most critical for each aspect, further justifying our fine-grained interven- tion strategy. We visualize the distribution of 2000 safety neurons and 2000 utility neurons across the 32 layers of LLaMA3.1-8B-Instruct, along with their average contribu- tion scores, as shown in Figure 3. Below are the observations from neuron distribution. Distributed Nature of Safety and Utility Capabilities Both safety-crucial and utility-related neurons are dis- tributed across all layers, rather than being concentrated in a few specific layers. This observation supports that these capabilities are not localized but are complexly encoded throughout the network. It also explains why coarse-grained, layer-wise interventions (as discussed in Section 2.1) in- evitably lead to imbalance, as modifying an entire layer is likely to impact both safety and utility neurons residing within it. Peak Concentrations in Different LayersLooking at the top of Figure 3, safety neurons show higher concentrations in middle layers, with a notable peak around Layer 13 and another significant concentration in later layers. Their av- erage contribution scores also show fluctuations but remain 0 50 100 150 200 250 Neuron Count 1 4 9 6 15 11 17 27 32 65 94 138 193 205 148 138 9393 89 75 37 38 20 22 16 18 24 22 26 47 81 196 Safety-Crucial Neuron Distribution Across Layers Total Safety-Crucial Neurons: 2000 | Layers: 32 | Max Neurons: Layer 13 (205) Safe Neuron Count Safe Avg Score 051015202530 Layer Index 0 50 100 150 200 250 300 Neuron Count 9 16 10 27 35 64 57 45 66 72 86 119 147 75 90 82 92 86 45 50 43 22 33 17 34 41 3434 46 48 106 269 Utility-Related Neuron Distribution Across Layers Total Utility-Related Neurons: 2000 | Layers: 32 | Max Neurons: Layer 31 (269) Util Neuron Count Util Avg Score 0 1 2 3 4 5 6 Average Contribution Score 1e5 0 1 2 3 4 Average Contribution Score 1e5 Figure 3: Safety and utility neuron distribution across lay- ers. The bar chart shows the count of neurons per layer, and the line indicates their average contribution score. Max safety neurons: Layer 13 (205). Max utility neurons: Layer 31 (269). relatively high across these layers. This suggests that safety- critical features are processed and refined in these interme- diate and deeper layers. In contrast, the bottom of Figure 3 reveals that utility neurons tend to have higher concentra- tions in deeper layers, with a prominent peak in Layer 31. While present throughout, their density and average con- tribution scores generally increase towards the later lay- ers. This aligns with the understanding that deeper layers in LLMs are often responsible for more abstract representa- tions, complex reasoning, and factual knowledge, which are crucial for general utility. Implications for Fine-Grained ModulationThe distinct, yet overlapping, distribution patterns of safety and utility neurons underscore the necessity of a fine-grained interven- tion approach like NeuronTune. Since these crucial neurons are not perfectly segregated by layer, a blanket modifica- tion of an entire layer would inevitably impact both types of neurons. Our method, by pinpointing individual neurons and applying adaptive scaling factors, allows for targeted en- hancement or suppression. This neuron-level precision en- ables us to strengthen safety-related pathways while mini- mizing collateral damage to utility-related knowledge, and vice-versa, thereby achieving a superior balance that is dif- ficult for coarse-grained methods to attain. This analysis provides empirical evidence supporting the architectural underpinnings of our NeuronTune and rein- forces why fine-grained neuron modulation is a more effec- tive strategy for balancing safety and utility in LLMs. 5 Related Work Safety AlignmentA considerable body of research has been devoted to enhance models safety, broadly termed safety alignment. These works mainly focus on the model alignment through techniques such as supervised fine-tuning (Bianchi et al. 2023; Diao et al. 2025), RLHF (Bai et al. 2022), or training-free interventions (Lin et al. 2023; Cao et al. 2024). However, the majority of existing research primarily focuses on improving the safety performance of LLMs. While crucial, an exclusive focus on safety en- hancement often inadvertently leads to exaggerated safety (R ̈ ottger et al. 2024) and pretrained knowledge degrada- tion (Lin et al. 2024). Despite attempts through manipulat- ing representation to mitigate these limitations (Wang et al. 2024a; Cao, Yang, and Zhao 2024), these techniques oper- ate at a broad level of granularity, failing to strike a good balance between robust safety and utility preservation. Knowledge NeuronsThe concept of knowledge neurons (Dai et al. 2022) has been proposed as a way to interpret the behaviors of language models by modifying specific neurons, thereby influencing the model’s generation output. Neuron-level pruning methods have been developed to iden- tify task-critical neurons. For instance, IRCAN (Shi et al. 2025) calculates the importance scores of all neurons based on their contribution to the loss, while Wanda (Sun et al. 2024) tracks changes in the immediate outputs of each layer when specific neurons are pruned. Regarding safety neu- rons, Chen et al. (2024) introduce generation-time activa- tion contrasting to locate safety neurons, highlighting their sparse distribution. Building on these insights, our approach focuses on balancing safety and utility by identifying and adaptively regulating the neurons associated with both. 6 Conclusion In this paper, we addressed the critical challenge of achiev- ing a balance between robust safety and utility in LLMs. We first diagnose the inherent limitations of existing coarse- grained, layer-wise intervention strategies, which often struggle to simultaneously achieve robust safety and main- tain high utility. To overcome it, we proposed NeuronTune, a novel fine-grained framework for safety-utility alignment. NeuronTune is built upon two core steps:Pinpointing Neu- rons Responsible for Safety and UtilityandEditing Neurons via Adaptive Activation Adjustment. This method allows for the targeted enhancement of safety-related activations and the preservation of utility-related knowledge, thereby mit- igating the rigid trade-offs of previous methods. Further- more, NeuronTune incorporates a tunable mechanism, offer- ing flexible control over the number of modulated neurons to adapt to diverse application scenarios and varying demands for safety or utility. References Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; Das- Sarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; Joseph, N.; Kadavath, S.; Kernion, J.; Conerly, T.; El- Showk, S.; Elhage, N.; Hatfield-Dodds, Z.; Hernandez, D.; Hume, T.; Johnston, S.; Kravec, S.; Lovitt, L.; Nanda, N.; Olsson, C.; Amodei, D.; Brown, T.; Clark, J.; McCandlish, S.; Olah, C.; Mann, B.; and Kaplan, J. 2022. Training a Helpful and Harmless Assistant with Reinforcement Learn- ing from Human Feedback. arXiv:2204.05862. Bianchi, F.; Suzgun, M.; Attanasio, G.; R ̈ ottger, P.; Ju- rafsky, D.; Hashimoto, T.; and Zou, J. 2023.Safety- Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.ArXiv, abs/2309.07875. Cao, B.; Cao, Y.; Lin, L.; and Chen, J. 2024. Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Associa- tion for Computational Linguistics (Volume 1: Long Papers), 10542–10560. Bangkok, Thailand: Association for Compu- tational Linguistics. Cao, Z.; Yang, Y.; and Zhao, H. 2024. SCANS: Mitigat- ing the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering. InAAAI Conference on Artificial Intel- ligence. Chen, J.; Wang, X.; Yao, Z.; Bai, Y.; Hou, L.; and Li, J. 2024. Finding Safety Neurons in Large Language Models. arXiv:2406.14144. Dai, D.; Dong, L.; Hao, Y.; Sui, Z.; Chang, B.; and Wei, F. 2022. Knowledge Neurons in Pretrained Transformers. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds.,Proceed- ings of the 60th Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), 8493–8502. Dublin, Ireland: Association for Computational Linguistics. Diao, M.; Li, R.; Liu, S.; Liao, G.; Wang, J.; Cai, X.; and Xu, W. 2025. Seas: Self-evolving adversarial safety optimization for large language models. InProceedings of the AAAI Con- ference on Artificial Intelligence, volume 39, 23778–23786. Finn, C.; Abbeel, P.; and Levine, S. 2017. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In International Conference on Machine Learning. Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; et al. 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783. Gurnee, W.; Nanda, N.; Pauly, M.; Harvey, K.; Troitskii, D.; and Bertsimas, D. 2023. Finding Neurons in a Haystack: Case Studies with Sparse Probing.Transactions on Machine Learning Research. Hazra, R.; Layek, S.; Banerjee, S.; and Poria, S. 2024. Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activa- tions. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds., Proceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing, 21759–21776. Miami, Florida, USA: Association for Computational Linguistics. Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D. X.; and Steinhardt, J. 2020. Measuring Massive Multitask Language Understanding.ArXiv, abs/2009.03300. Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Rad- ford, A.; et al. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276. Jin, H.; Chen, R.; Zhang, P.; Zhou, A.; Zhang, Y.; and Wang, H. 2025. GUARD: Role-playing to Generate Natural- language Jailbreakings to Test Guideline Adherence of Large Language Models. arXiv:2402.03299. Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J.; Zhang, H.; and Stoica, I. 2023. Efficient memory management for large language model serving with pagedattention. InProceedings of the 29th symposium on operating systems principles, 611–626. Li, X.; Ning, Y.; Bao, Z.; Xu, M.; Chen, J.; and Qian, T. 2025. CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Repre- sentations. arXiv:2507.06043. Lin, B. Y.; Ravichander, A.; Lu, X.; Dziri, N.; Sclar, M.; Chandu, K.; Bhagavatula, C.; and Choi, Y. 2023. The un- locking spell on base llms: Rethinking alignment via in- context learning. InThe Twelfth International Conference on Learning Representations. Lin, S.; Hilton, J.; and Evans, O. 2022. TruthfulQA: Mea- suring How Models Mimic Human Falsehoods. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds.,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3214–3252. Dublin, Ireland: Association for Computational Linguistics. Lin, Y.; Lin, H.; Xiong, W.; Diao, S.; Liu, J.; Zhang, J.; Pan, R.; Wang, H.; Hu, W.; Zhang, H.; Dong, H.; Pi, R.; Zhao, H.; Jiang, N.; Ji, H.; Yao, Y.; and Zhang, T. 2024. Mitigating the Alignment Tax of RLHF. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds.,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 580– 606. Miami, Florida, USA: Association for Computational Linguistics. R ̈ ottger, P.; Kirk, H.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. In Duh, K.; Gomez, H.; and Bethard, S., eds.,Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), 5377–5400. Mexico City, Mexico: Association for Computational Lin- guistics. Shi, D.; Jin, R.; Shen, T.; Dong, W.; Wu, X.; and Xiong, D. 2025. IRCAN: mitigating knowledge conflicts in LLM generation via identifying and reweighting context-aware neurons. InProceedings of the 38th International Con- ference on Neural Information Processing Systems, NIPS ’24. Red Hook, NY, USA: Curran Associates Inc. ISBN 9798331314385. Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2024. A Simple and Effective Pruning Approach for Large Language Mod- els. arXiv:2306.11695. Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023. Stanford Alpaca: An Instruction-following LLaMA model.https: //github.com/tatsu-lab/stanford alpaca. Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozi ` ere, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lam- ple, G. 2023a. LLaMA: Open and Efficient Foundation Lan- guage Models. arXiv:2302.13971. Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023b. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288. Varshney, N.; Dolin, P.; Seth, A.; and Baral, C. 2024. The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds.,The 62nd Annual Meeting of the Association for Computational Lin- guistics, Proceedings of the Annual Meeting of the Associa- tion for Computational Linguistics, 13111–13128. Associa- tion for Computational Linguistics (ACL). Publisher Copy- right: © 2024 Association for Computational Linguistics.; Findings of the 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024 ; Conference date: 11-08-2024 Through 16-08-2024. Wang, M.; Zhang, N.; Xu, Z.; Xi, Z.; Deng, S.; Yao, Y.; Zhang, Q.; Yang, L.; Wang, J.; and Chen, H. 2024a. Detox- ifying Large Language Models via Knowledge Editing. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds.,Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), 3093–3118. Bangkok, Thailand: Association for Computational Linguis- tics. Wang, M.; Zhang, N.; Xu, Z.; Xi, Z.; Deng, S.; Yao, Y.; Zhang, Q.; Yang, L.; Wang, J.; and Chen, H. 2024b. Detox- ifying Large Language Models via Knowledge Editing. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds.,Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), 3093–3118. Bangkok, Thailand: Association for Computational Linguis- tics. Wang, X.; Wen, K.; Zhang, Z.; Hou, L.; Liu, Z.; and Li, J. 2022. Finding Skill Neurons in Pre-trained Transformer- based Language Models. InProceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Process- ing, 11132–11152. Xiao, Z.; Yang, Y.; Chen, G.; and Chen, Y. 2024. Distract Large Language Models for Automatic Jailbreak Attack. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y.-N., eds.,Proceed- ings of the 2024 Conference on Empirical Methods in Nat- ural Language Processing, 16230–16244. Miami, Florida, USA: Association for Computational Linguistics. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388. Yang, Q. A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Lin, J.; Dang, K.; Lu, K.; Bao, K.; Yang, K.; Yu, L.; Li, M.; Xue, M.; Zhang, P.; Zhu, Q.; Men, R.; Lin, R.; Li, T.; Xia, T.; Ren, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.-C.; Wan, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; Qiu, Z.; Quan, S.; and Wang, Z. 2024. Qwen2.5 Technical Report.ArXiv, abs/2412.15115. Zhang, H.; Guo, Z.; Zhu, H.; Cao, B.; Lin, L.; Jia, J.; Chen, J.; and Wu, D. 2024. Jailbreak Open-Sourced Large Lan- guage Models via Enforced Decoding. In Ku, L.-W.; Mar- tins, A.; and Srikumar, V., eds.,Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers), 5475–5493. Bangkok, Thai- land: Association for Computational Linguistics. Zhang, J.; Elgohary, A.; Magooda, A.; Khashabi, D.; and Durme, B. V. 2025.Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements. arXiv:2410.08968. Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023.Universal and Transfer- able Adversarial Attacks on Aligned Language Models. arXiv:2307.15043. A Supplementary Details of Experiments A.1 Safety and Utility Evaluation Datasets We select five datasets to evaluate the performance on robust safety and utility. SafeEditThe refuse rate of this dataset is used to assess sufficient safety. Its entropy and fluency metrics reflect the utility of models, quantifying the diversity, richness, and smoothness of the output. The evaluation set comprises 975 samples (Wang et al. 2024b). AdvBenchSimilar to SafeEdit, this dataset is employed to measure both sufficient safety and utility. The evaluation set comprises 456 samples (Zou et al. 2023). AlpacaWe utilize a subset of the original Alpaca dataset (Taori et al. 2023). The refuse rate and entropy metrics are employed to evaluate exaggerated safety and utility, respec- tively. The evaluation set comprises 100 samples. TruthfulQASimilar to Alpaca, this dataset is employed to measure both exaggerated safety and utility. The evaluation set comprises 753 samples (Lin, Hilton, and Evans 2022). MMLUGiven its comprehensive coverage of knowledge- intensive tasks, we utilize MMLU to assess the general performance of models, which serves as another aspect of utility. The evaluation set comprises 14,042 samples (Hendrycks et al. 2020). A.2 Experimental Environment For all experiments, we conduct experiments on three Nvidia L20-48G GPUs. We use the vLLM framework (Kwon et al. 2023) for all the LLM generation. A.3 Analysis on Adversarial Attacks To provide empirical evidence for the limitations of cur- rent LLMs in handling adversarial contexts, we conduct a detailed analysis on model robustness against sophisticated adversarial attacks. Table 5 illustrates that LLMs, when sub- jected to these misleading attacks, exhibit a significant in- crease in the rate of unsafe outputs compared to responses to direct harmful queries. This highlights that misleading at- tacks are a primary cause of bypassed safety mechanisms and the generation of undesirable content. This analysis un- derscores the necessity for a targeted approach to identify and address the neural mechanisms that are influenced by such attacks, thereby reinforcing the motivation for our pro- posed attack-aware attribution method. Queries LLaMA2-7b-Chat LLaMA3.1-8B-Instruct Qwen2.5-7B-Instruct Qwen2.5-14B-Instruct w/ Attacks38.56%10.77%14.87%14.97% w/o Attacks76.21%47.49%73.33%82.67% Table 5: Refuse rate of different models under harmful queries with and without attacks, showing the impact of at- tacks on model performance. A.4 Performance on AdvBench and TruthfulQA Taking the LLaMA3.1-8B-Instruct model as an example, the results of all methods on AdvBench and TruthfulQA are shown in Table 6. The results further substantiate the effec- tiveness of our fine-grained neuron modulation strategy in achieving a superior balance between robust safety and util- ity preservation. Methods AdvBenchTruthfulQA SU-F1↑ Refuse Rate↑Entropy↑Refuse Rate↓Entropy↑ Default70.39%5.899 bit0.40%6.359 bit0.818 DINM81.58%3.820 bit3.72%5.798 bit0.706 SCANS96.05%3.106 bit0.00%3.454 bit0.517 CAVGAN95.61%5.940 bit22.18%6.175 bit0.820 NeuronTune73.03%6.235 bit1.59%6.031 bit0.822 Table 6: Comparative performance of methods on Ad- vBench and TruthfulQA on LLaMA3.1-8B-Instruct, demon- strating NeuronTune’s effectiveness.Bolddenotes the best performance in each block.↑indicates that higher values are better, while↓indicates that lower values are better. A.5 Results of Neuron Number Regulation NeuronTune incorporates a tunable mechanism that en- hances model adaptation to diverse safety and utility de- mands by regulating the number of neurons. Table 7 illus- trates the impact of varying neuron counts. Specifically, we examine two scenarios: increasing the number of safety- crucial neurons while keeping utility-related neurons fixed, and vice versa. Number of Safety-Crucial / Utility-Related Neurons SafeEditAlpaca Refuse Rate↑Entropy↑Refuse Rate↓Entropy↑ 500/50041.33%5.297 bit1%6.077 bit 1000/50062.26%4.245 bit0%4.870 bit 1500/50074.05%3.638 bit6%4.267 bit 2000/50083.69%3.285 bit7%4.262 bit 500/100043.38%5.483 bit3%5.867 bit 1000/100069.64%4.377 bit0%4.446 bit 1500/100080.41%3.771 bit1%4.228 bit 2000/100086.56%3.270 bit9%3.889 bit 500/150049.74%5.317 bit3%5.912 bit 1000/150060.41%4.217 bit1%4.803 bit 1500/150069.44%4.029 bit1%4.597 bit 2000/150077.74%3.568 bit8%4.266 bit Table 7: Results on dynamically adjusting the number of safety-crucial and utility-related neurons. The tunable mech- anism offers a practical approach to adapt NeuronTune to different application scenarios. A.6 Adversarial Attack Types and Strategic Attack Selection Following (Wang et al. 2024b), SafeEdit comprises 48 attack prompts collected from various sources, including websites, recent papers, and handwritten instances. These prompts are meticulously designed to induce unexpected or potentially harmful responses from LLMs. These powerful attack tem- plates are broadly categorized into four types based on their underlying strategy: pretending, attention shifting, privilege escalation, and emotion control. To effectively manage computational intensity while si- multaneously maximizing the comprehensiveness of the identified neurons, we adopt a strategic attack selection ap- proach. This strategy involved selecting only one represen- tative data point from each attack category for our attack- aware gradient attribution analysis. It is important to note that a single example in the original SafeEdit dataset may comprise multiple attack templates. To ensure that each selected data point exclusively represents a distinct attack category, we pre-processed the dataset by filtering out examples that did not solely contain one type of attack. As a consequence of this refinement, the emotion control category was excluded from our strategic selection, as it did not have any standalone examples in the filtered dataset.