Paper deep dive
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Wenhai Wang
Models: Gemma-2-9B-Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/11/2026, 1:16:11 AM
Summary
LLM-VA is a novel vector steering method that resolves the jailbreak-overrefusal trade-off in safety-aligned LLMs by aligning answer vectors (v_a) with benign vectors (v_b). By identifying these vectors via SVMs and applying closed-form weight updates, the model's willingness to answer becomes causally dependent on its safety assessment without requiring fine-tuning or architectural changes.
Entities (6)
Relation Signals (4)
LLM-VA → aligns → Answer Vector (v_a)
confidence 100% · LLM-VA, which aligns v_a with v_b
LLM-VA → aligns → Benign Vector (v_b)
confidence 100% · LLM-VA, which aligns v_a with v_b
LLM-VA → resolves → Jailbreak-Overrefusal Trade-off
confidence 100% · LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
SVM → identifies → Answer Vector (v_a)
confidence 90% · Our method identifies vectors at each layer using SVMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector steering methods adjust the magnitude of answer vectors, but this creates a fundamental trade-off -- reducing jailbreak increases over-refusal and vice versa. We identify the root cause: LLMs encode the decision to answer (answer vector $v_a$) and the judgment of input safety (benign vector $v_b$) as nearly orthogonal directions, treating them as independent processes. We propose LLM-VA, which aligns $v_a$ with $v_b$ through closed-form weight updates, making the model's willingness to answer causally dependent on its safety assessment -- without fine-tuning or architectural changes. Our method identifies vectors at each layer using SVMs, selects safety-relevant layers, and iteratively aligns vectors via minimum-norm weight modifications. Experiments on 12 LLMs demonstrate that LLM-VA achieves 11.45% higher F1 than the best baseline while preserving 95.92% utility, and automatically adapts to each model's safety bias without manual tuning. Code and models are available at this https URL.
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
61,804 characters extracted from source content.
Expand or collapse full text
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment Haonan Zhang 1 , Dongxia Wang 1,3† , Yi Liu 2 , Kexin Chen 1 , Wenhai Wang 1 1 Zhejiang University, 2 Quantstamp, 3 Huzhou Institute of Industrial Control Technology kxchen, dxwang, haonanzhang, zdzzlab@zju.edu.cn, yi009@e.ntu.edu.sg Abstract Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Ex- isting vector steering methods adjust the mag- nitude of answer vectors, but this creates a fundamental trade-off—reducing jailbreak in- creases over-refusal and vice versa. We iden- tify the root cause: LLMs encode the deci- sion to answer (answer vectorv a ) and the judgment of input safety (benign vectorv b ) as nearly orthogonal directions, treating them as independent processes. We propose LLM- VA, which alignsv a withv b through closed- form weight updates, making the model’s will- ingness to answer causally dependent on its safety assessment—without fine-tuning or ar- chitectural changes. Our method identifies vec- tors at each layer using SVMs, selects safety- relevant layers, and iteratively aligns vectors via minimum-norm weight modifications. Ex- periments on 12 LLMs demonstrate that LLM- VA achieves 11.45% higher F1 than the best baseline while preserving 95.92% utility, and automatically adapts to each model’s safety bias without manual tuning. Code and models are available athttps://hotbento.github. io/LLM-VA-Web/. 1 Introduction Large language models (LLMs) have achieved remarkable capabilities across diverse NLP tasks (OpenAI, 2024; Team, 2025; AI@Meta, 2024), yet safety alignment remains challenging. Safety-aligned LLMs exhibit two failure modes: jailbreak, where the model directly responds to toxic inputs (i.e., queries designed to elicit harmful, unethical, or unsafe responses) (Chen et al., 2024; Deng et al., 2024; Yuan et al., 2025; Liu et al., 2024b), and over-refusal, where the model unnec- essarily declines benign queries (Röttger et al., 2024; Zhang et al., 2025a; Cui et al., 2025). This † Corresponding Author 010203040 Layer 0 20 40 60 80 100 120 Angle (degrees) 90 degrees attn_out mlp_out Figure 1: The angles between answer vectors (v a ) and benign vectors (v b ) are approximately90 ◦ across lay- ers in gemma-2-9b-it, indicating near-orthogonality be- tween answer decisions and safety assessments. dual failure mode significantly limits the deploy- ment of LLMs in safety-critical applications, where both reliability and usability are essential. Among approaches to address these issues, vector steer- ing (Zou et al., 2023a; Arditi et al., 2024; Sheng et al., 2025) has gained attention for its efficiency— it manipulates specific directions in the model’s latent space without costly retraining, using only simple answer/refuse labels rather than fine-grained annotations. However, existing vector steering methods only adjust the magnitude of the answer vector, cre- ating a fundamental trade-off: reducing magni- tude suppresses jailbreak but increases over-refusal, while amplifying it has the opposite effect (Arditi et al., 2024; Sheng et al., 2025). Recent methods like SCANS (Cao et al., 2025) and CAST (Lee et al., 2024) incorporate input toxicity but require architectural modifications and treat both failure modes as separate objectives (see Table 1). This magnitude-based paradigm cannot fundamentally resolve the trade-off. We identify the root cause of this trade-off: ex- isting methods control output behavior (answer vs. refuse) without considering input characteristics (benign vs. toxic). To investigate, we extract two arXiv:2601.19487v1 [cs.LG] 27 Jan 2026 vectors at each layer: the answer vector (v a ), in- dicating whether the model will answer, and the benign vector (v b ), indicating whether the input is safe. As shown in Figure 1, these vectors are nearly orthogonal (∼90 ◦ ) across layers, 1 revealing that LLMs treat answer decisions and safety assess- ments as independent processes. This explains both failure modes: the model may answer toxic inputs (jailbreak) or refuse benign ones (over-refusal) be- cause its willingness to answer is decoupled from its judgment of input safety. Based on this observation, we proposeLarge LanguageModelVectorAlignment (LLM-VA). By aligning these vectors, we make the model’s willingness to answer causally dependent on its safety assessment (Zou et al., 2023a), rather than treating them as independent decisions. Crucially, LLM-VA achieves this through closed-form weight updates—requiring no gradient-based optimization, fine-tuning, or architectural changes. Our method involves three steps: •Vector identification via SVMs: Train SVMs at each layer to find hyperplanes separating be- nign/toxic and answer/refuse samples, yielding both v b and v a . •Layer selection: Identify layers most relevant to safety decisions based on their contribution to final output and SVM classification accuracy. •Vector alignment: Adjust layer weights to align v a withv b , ensuring benign inputs activate the “answer” direction while toxic inputs do not. Extensive experiments on 12 LLMs demon- strate that LLM-VA achieves 11.45% higher F1 scores (effectiveness on resolving trade-off) than the best baseline (AlphaSteer) (Sheng et al., 2025) with only 4.08% model utility drop, indicating that LLM-VA effectively resolves the jailbreak- overrefusal trade-off while preserving general ca- pabilities. In summary, our contributions are: •We propose LLM-VA, which, to the best of our knowledge, is the first vector steering method that simultaneously addresses both jail- break and over-refusal by aligning answer vec- tors with benign vectors through closed-form weight updates—requiring no gradient-based fine-tuning or architectural changes. 1 Results for other LLMs are similar; see Appendix A. •We demonstrate on 12 LLMs from 5 model fami- lies that LLM-VA achieves state-of-the-art safety alignment, and show that it automatically adapts to each model’s safety bias—prioritizing jail- break reduction for vulnerable models and over- refusal reduction for overly conservative ones— without manual tuning. • We release our code and safety-enhanced weights for 12 LLMs. 2 2 Related Work Safety Alignment and the Jailbreak-Overrefusal Trade-offTraditional safety alignment methods— RLHF (Christiano et al., 2017; Stiennon et al., 2020), adversarial training (Xhonneux et al., 2024; Liu et al., 2024a), and rule-based filtering (Zhang et al., 2025b; Liu et al., 2024c)—require substantial computational resources or lack scalability. Vec- tor steering (Zou et al., 2023a; Arditi et al., 2024) emerged as an efficient alternative, manipulating latent-space directions without retraining. How- ever, these methods create a fundamental trade-off: reducing the answer vector’s magnitude suppresses jailbreak but increases over-refusal, while ampli- fying it has the opposite effect (Arditi et al., 2024; Sheng et al., 2025). This trade-off remains the cen- tral unsolved problem in efficient safety alignment. Vector Steering Methods VectorSteer (Zou et al., 2023a) first identified answer vectors for controlling model outputs through magnitude ad- justment. AlphaSteer (Sheng et al., 2025) intro- duced null-space projection to preserve utility dur- ing steering, but remains magnitude-based and thus inherits the trade-off. SCANS (Cao et al., 2025) and CAST (Lee et al., 2024) incorporate input toxicity information, representing progress toward input-aware steering. However, both require archi- tectural modifications (hook layers) and still treat jailbreak and over-refusal as separate objectives to be balanced via hyperparameters. Table 1 sum- marizes these differences: LLM-VA is the only approach that addresses both failure modes without finetuning or architectural changes. Internal Representations in LLMs Mechanis- tic interpretability research reveals that LLMs en- code concepts as linear directions in their hidden states (Geva et al., 2021; Elhage et al., 2022; Zou 2 We release only Llama3.1-8B-Instruct weights during review. Full weights available athttps://figshare.com/s/ f2a365c87a80097a436. Table 1: Comparison of LLM-VA with other methods on safety alignment and utility preservation. Method w/o Finetuning w/o Model Structure Modification Over-refusal Mitigation Jailbreak Mitigation LLM-VA✓ Finetuning✗✓ VectorSteer✓✗✓ AlphaSteer✓✗✓ CAST✓✗✓ SCANS✓✗✓ Toxic answeredToxic unansweredBenign answeredBenign unanswered 0.40.30.20.10.00.10.20.3 Projection on benign vector 0.3 0.2 0.1 0.0 0.1 0.2 Projection on answer vector 1.000.750.500.250.000.250.500.75 Projection on benign vector 1.00 0.75 0.50 0.25 0.00 0.25 0.50 0.75 Projection on answer vector 1.01.0 0.5 0.0 0.5 1.5 1.0 0.5 0.0 0.5 1.0 Projection on answered_dir Projection on benign vector Figure 2: The distributions of the projections onto the benign, answer vectors at different layers of Llama-3.1- 8B-Instruct. The left, middle, right figures correspond to the 4th, 16th, and 28th MLP layers, respectively. et al., 2023a). Building on this foundation, we dis- cover that answer vectors (v a ) and benign vectors (v b ) are nearly orthogonal across layers, explaining why magnitude-based methods cannot resolve the trade-off—they control output behavior indepen- dently of input safety. LLM-VA addresses this by aligning these vectors, making the answer decision causally dependent on the safety assessment. 3 Preliminary Analysis To motivate our approach, we analyze how LLMs internally represent two distinct decisions: (1) whether to answer or refuse a query, and (2) whether the input is benign or toxic. 3 Follow- ing Zou et al. (2023a), we extract the answer vector v a and benign vectorv b at each layer on 128 ran- domly sampled toxic inputs from S-Eval (Yuan et al., 2025) and 128 benign inputs from ORFuz- zSet (Zhang et al., 2025a). 4 We project layer out- puts onto these vectors and visualize the distribu- tions in Figure 2. Three key observations emerge: •Obs 1: LLMs encode both decisions internally. Projections ontov b cleanly separate benign from toxic inputs, while projections ontov a separate answered from refused samples—both with deci- sion boundaries near zero. 3 We define “answer” as providing a direct response and “refuse” as declining to respond. 4 We illustrate with Llama-3.1-8B-Instruct; results are con- sistent across models. Answer-Refuse Plane Benign-Toxic Plane Normal Samples Over-Refusal Samples Jailbreak Samples Original Embedding Vector Space Vector Steering (a) Vector Alignment (b) Figure 3: Unlike existing methods that only adjust the magnitude ofv a (trading off jailbreak vs. over-refusal), LLM-VA aligns v a with v b to address both issues. •Obs 2: Later layers are more discriminative. Separation quality improves in deeper layers (compare layers 4, 16, and 28 in Figure 2), in- dicating that later layers are more critical for safety-related decisions. •Obs 3: The two decisions are misaligned. Some toxic inputs project positively ontov a , while some benign inputs project negatively. This misalignment directly causes jailbreak and over-refusal failures. Combined with the near-orthogonality between v a andv b (Figure 1), these observations reveal that LLMs treat answer decisions and safety assess- ments as independent processes. We hypothesize that aligningv a withv b —making the model’s will- ingness to answer depend on its safety judgment— will reduce both failure modes. Why vector alignment, not magnitude adjust- ment? Existing vector steering methods (Sheng et al., 2025; Cao et al., 2025; Ray and Bhalani, 2024) only adjust the magnitude ofv a : reducing it decreases jailbreak risk but increases over-refusal, while increasing it has the opposite effect (Fig- ure 3a). In contrast, LLM-VA alignsv a withv b (Figure 3b), making the answer decision depend on input safety rather than treating them indepen- dently. Optimization ObjectiveWe formalize this goal as maximizing correct response behavior: max θ E x I(y=benign)· I(f θ (x)=answer)+ I(y=toxic)· I(f θ (x)=refuse) (1) wherexis an input,y ∈benign, toxicits ground- truth label, andf θ (x) ∈ answer, refusethe model’s response. By aligningv a withv b , pro- jections ontov a become correlated with input be- nignness, optimizing this objective. The following sections detail how LLM-VA achieves this. SVM-based Control Vector Identification Layer Selection Safety Enhancement with Vector Alignment Queries Embed ding Attention FFN Attention FFN Attention FFN ... ... Labeled Vectors SVM Training Benign Answered Benign Refused Toxic Answered Toxic Refused 푣 푎 푣 푏 Layer 1 Layer 2 Layer N ... 푣 푎 (1) 푣 푏 1 푣 푎 (2) 푣 푏 2 푣 푎 (푁) 푣 푏 푁 ...... 푣 푎 푓 푟 푏 푓 퐶 푎 푙 ,퐶 푏 푙 퐴 푎 푙 ,퐴 푏 푙 푆퐴 푆푟푒 푙 Layer Score Selected Layers Other Layers 푣 푎 푣 푏 푣 푎 푣 푏 Alignment 푾 Jailbreak Scenario Over-refusal Scenario LLM Modified LLM Updated Safety-Guardrail Figure 4: The framework of LLM-VA. 4 Methodology Building on our observation that LLMs encode answer decisions (v a ) and safety assessments (v b ) as nearly orthogonal directions, we present LLM-VA. Our key insight is that by aligning these vectors through closed-form weight updates— requiring no gradient-based fine-tuning or archi- tectural changes—we can make the model’s will- ingness to answer causally dependent on its safety judgment. As illustrated in Figure 4, LLM-VA mainly consists of three steps: (1) identifyingv a andv b at each layer via SVMs (Section 4.1), (2) selecting layers most relevant to safety decisions (Section 4.2), and (3) deriving weight update pro- cess that aligns these vectors (Section 4.3). 4.1 SVM-based Control Vector Identification To align vectors at each layer, we must first identify them. Prior work (Zou et al., 2023a; Sheng et al., 2025; Cao et al., 2025) extracts the answer vector from the residual flow at the final layer. However, since the residual flow aggregates contributions from all preceding layers, modifying individual layer weights cannot directly control the final-layer vector. To enable layer-wise weight modification, we instead extract vectors from each layer’s output. At each layer, we train two linear SVMs to find hyperplanes separating (1) benign vs. toxic inputs, and (2) answered vs. refused samples. We use SVMs because they provide interpretable lin- ear decision boundaries: the normal vector of the maximum-margin hyperplane directly yields the control vector, and the margin maximization en- sures robustness. The SVMs minimize (Cortes and Vapnik, 1995): min w svm ,ζ ∥w svm ∥ 2 2 + C X i∈D ζ i , s.t. y i (w svm · o (l) i )≥ 1− ζ i , ∀i∈D(2) whereo (l) i is the output of layerlfor inputi,y i ∈ −1, 1 is the label (+1for benign/answer, and −1for toxic/refuse),C > 0is a regularization parameter, andζ i ≥ 0are slack variables. We omit the bias termb svm because our empirical analysis shows that decision hyperplanes pass through the origin. This simplifies the subsequent alignment formulation and implementation. The unit normal vectors of these hyperplanes yield the control vectors: v (l) b = w (l) b / w (l) b (3) v (l) a = w (l) a / w (l) a (4) wherew (l) b andw (l) a are the SVM weight vectors for benign/toxic and answer/refuse classification at layer l, respectively. 4.2 Layer Selection Not all layers contribute equally to safety deci- sions (Geva et al., 2021). Modifying irrelevant layers wastes capacity and may harm utility, so we select layers that are both influential (their vectors align with the decisions of final residual stream) and accurate (their SVMs reliably distinguish be- nign/toxic or answer/refuse). 5 Influence on final decision. Following prior work showing that the residual stream determines final outputs (Zou et al., 2023a; Sheng et al., 2025), we measure how well each layer’s vectors align with the vectors of final residual stream: C (l) a = v (fin) a · v (l) a , C (l) b = v (fin) b · v (l) b (5) HighC (l) indicates that modifying layerl’s vector direction will propagate to the final decision. Classification accuracy.We also require that the layer’s SVMs accurately separate the two classes. LetAcc (l) a andAcc (l) b denote validation accuracies for the answer and benign classifiers at layer l. Combined score. We compute a weighted sum where each term is the product of influence and accuracy for each task: Score (l) = C (l) a · Acc (l) a + C (l) b · Acc (l) b (6) The multiplicative form within each term ensures we select layers that are both influential and accu- rate for that task—a layer with high influence but low accuracy (or vice versa) contributes little to the score. We select the topL select layers with the highest scores for alignment. 4.3 Vector Alignment Our goal is to modify each selected layer’s weights so that the model’s answer decision becomes depen- dent on its safety assessment. Specifically, for any input, we want the projection ontov a (which de- termines answering) to equal the scaled projection ontov b (which reflects input safety). This ensures benign inputs activate the “answer” direction while toxic inputs suppress it. Unlike existing methods (Zou et al., 2023a; Sheng et al., 2025; Cao et al., 2025) that insert hook layers and modify the model architecture, we derive a closed-form weight update process—requiring no 5 Throughout this paper, “layer” refers to either an MLP or attention sublayer unless otherwise specified. Reasons are discussed in Appendix B. gradient descent or architectural changes. This makes LLM-VA efficient and easy to deploy on standard model-hosting platforms. Deriving the weight update. For each selected layer, we modify the down-projection matrixW (the matrix that projects from hidden dimension back to model dimension). We seek an update∆ such that (omitting layer indices for clarity): x(W + ∆)v a = σ a σ b xWv b , ∀x(7) whereσ a andσ b are the standard deviations of projections ontov a andv b over the training set, respectively. The ratioσ a /σ b normalizes for differ- ent dynamic ranges of the two directions, ensuring benign inputs (positivev b projection) produce pos- itivev a projections and toxic inputs (negativev b projection) produce negativev a projections. Rear- ranging, we require: ∆v a = σ a σ b Wv b − Wv a (8) The minimum-norm solution (least modification to weights) is given by the pseudoinverse (Penrose, 1955): ∆ + = σ a σ b Wv b − Wv a v T a ,(9) W ′ = W + ∆ + Iterative refinement. A single alignment step may not fully align the vectors because modifying one layer’s weights affects the inputs to subsequent layers, causing their effectivev a andv b directions to shift. We therefore iterate the alignment process Ttimes: in each iteration, we re-extractv a andv b from the modified model, recompute layer scores, and apply the weight update. The final model is selected based on validation F1 score. Empirically, most models converge within 20–30 iterations (see Section 5.4). 5 Experiments We conduct experiments to address the following research questions: • RQ1:How effectively does LLM-VA re- solve jailbreak-overrefusal trade-off compared to magnitude-based vector steering methods? • RQ2: How well does LLM-VA preserve model utility? •RQ3: How do key components (vector identi- fication, iteration count, layer selection) affect performance? 5.1 Experimental Setup We first describe the experimental settings. Addi- tional details are provided in Appendix D. Models We conduct experiments on 12 widely- used instruction-tuned LLMs spanning 5 model families, with sizes ranging from 3B to 14B param- eters: Llama-3.1 (8B) (AI@Meta, 2024), gemma-2 (9B) (Team, 2024a), Mistral-v0.3 (7B) (Jiang et al., 2023), Phi-3.5 (4B) (Abdin et al., 2024), Phi-4 (4B, 15B) (Microsoft et al., 2025), Qwen2.5 (3B, 7B, 14B) (Team, 2024b; Yang et al., 2024a), and Qwen3 (4B, 8B, 14B) (Team, 2025). This diverse selection allows to evaluate the generalizability of LLM-VA across different architectures and scales. Datasets For effectiveness evaluation, we use four benchmark datasets: S-Eval-Attack and S- Eval-Risk (Yuan et al., 2025) for jailbreak eval- uation, and ORFuzzSet (Zhang et al., 2025a) and Natural Questions (Kwiatkowski et al., 2019) for over-refusal evaluation. To focus on challeng- ing cases, we select 500 samples per dataset where the original models exhibit incorrect be- havior (i.e., jailbreak on toxic inputs or over- refusal on benign inputs). Each dataset is split into training, validation, and test sets with a ratio of 8:1:1. For utility preservation, we evaluate on 6 datasets covering diverse NLP tasks including grammar (CoLA (Warstadt et al., 2018)), natural language inference (MNLI (Williams et al., 2018), RTE (Bentivogli et al., 2009)), paraphrase detection (MRPC (Dolan and Brockett, 2005)), sentiment analysis (SST (Socher et al., 2013)), and mathemat- ical reasoning (GSM8K (Cobbe et al., 2021)). 6 Baselines We compare LLM-VA with several state-of-the-art vector steering methods: •VectorSteer (Zou et al., 2023a): Identifies the an- swer vector and adjusts its magnitude to control the model’s response behavior. • AlphaSteer (Sheng et al., 2025): Extends Vec- torSteer by introducing null-space projection on representation space to preserve the model’s gen- eral capabilities while steering. 6 See Appendix C for dataset details. •SCANS (Cao et al., 2025): Dynamically adjusts answer vector magnitude based on input toxic- ity judgement, using hook layers to incorporate toxicity information. •AlphaSteer+: Our variant of AlphaSteer that uses null-space projection to preserve behavior specifically on correctly-answered samples rather than general capabilities. Metrics We use attack success rate (ASR) (Zou et al., 2023b) to measure jailbreak vulnerabil- ity and over-refusal rate (ORR) (Zhang et al., 2025a) to measure unnecessary refusals.For evaluation of effectiveness on resolving the trade-off, we report F1 scores with all the four datasets, whereTP =|benign ∩ answered|, FP =|toxic∩ answered|,FN =|benign∩ refused|, andTN =|toxic∩ refused|. For utility preserva- tion, we report F1 for classification tasks where TP,FP,FNare defined by the positive class of each task, and accuracy for GSM8K. We employ Qwen3-Guard-Gen-8B (Zhao et al., 2025) as the judge model for evaluating whether responses con- stitute answers or refusals. 7 5.2 Effectiveness Results (RQ1) To evaluate the effectiveness of LLM-VA on jail- break and over-refusal trade-off, we compare it with magnitude-based vector steering methods across all 12 LLMs. Table 2 presents ASR, ORR, and F1 scores on the test sets. Overall effectiveness of LLM-VA. LLM-VA achieves an average F1 score of 0.77, representing a 37.02% relative improvement over the original LLMs (0.56). Notably, LLM-VA simultaneously reduces both failure modes: ASR decreases by 18.50% and ORR decreases by 22.00% on average compared to the original LLMs. Comparison with baselines. LLM-VA outper- forms all baselines on 8 of 12 LLMs regarding F1, with a relative improvement of 11.45% over the best baseline (AlphaSteer). VectorSteer, AlphaS- teer+ and AlphaSteer, which only adjust answer vector magnitude, show limited improvement on models that already have low ASR but high ORR (e.g., Llama-3.1-8B). SCANS achieves competitive results on some models but requires architectural modifications and shows inconsistent performance across model families. 7 See Appendix E for details on judge model selection. Table 2: Main results of LLM-VA. The best results are bolded. ModelSizeMethodSeval-Aattack ASR↓ Seval-Risk ASR↓ ORFuzzSet ORR↓ NQ ORR↓ Final F1↑ ModelSizeMethodSeval-Aattack ASR↓ Seval-Risk ASR↓ ORFuzzSet ORR↓ NQ ORR↓ Final F1↑ Llama-3.18B Original12.00%2.00%100.00%6.00%0.6104 Qwen2.5 3B Original88.00%20.00%62.00%14.00%0.5741 AlphaSteer+4.00%0.00%100.00%10.00%0.6122AlphaSteer+28.00%0.00%58.00%10.00%0.7333 AlphaSteer2.00%0.00%100.00%6.00%0.6351AlphaSteer28.00%6.00%62.00%10.00%0.7072 VectorSteer0.00%0.00%100.00%12.00%0.6111VectorSteer22.00%2.00%92.00%32.00%0.5067 SCANS4.00%0.00%100.00%14.00%0.5931SCANS32.00%4.00%70.00%6.00%0.6889 LLM-VA14.00%6.00%38.00%10.00%0.8172LLM-VA44.00%12.00%16.00%16.00%0.7925 gemma-29B Original42.00%22.00%98.00%16.00%0.4914 7B Original86.00%36.00%80.00%4.00%0.5297 AlphaSteer+16.00%0.00%94.00%4.00%0.6415AlphaSteer+32.00%6.00%82.00%2.00%0.6554 AlphaSteer16.00%0.00%98.00%4.00%0.6242AlphaSteer28.00%8.00%86.00%2.00%0.6437 VectorSteer10.00%0.00%98.00%18.00%0.5714VectorSteer16.00%16.00%80.00%24.00%0.5854 SCANS18.00%12.00%92.00%12.00%0.5890SCANS30.00%18.00%86.00%4.00%0.6145 LLM-VA0.00%6.00%36.00%6.00%0.8681LLM-VA54.00%30.00%22.00%4.00%0.7598 Mistral-v0.37B Original88.00%74.00%54.00%4.00%0.5635 14B Original46.00%22.00%90.00%4.00%0.5668 AlphaSteer+28.00%36.00%52.00%2.00%0.7122AlphaSteer+22.00%4.00%92.00%2.00%0.6386 AlphaSteer34.00%32.00%46.00%2.00%0.7273AlphaSteer2.00%4.00%92.00%2.00%0.6795 VectorSteer32.00%28.00%44.00%0.00%0.7500VectorSteer6.00%0.00%96.00%0.00%0.6710 SCANS26.00%40.00%28.00%4.00%0.7742SCANS20.00%24.00%82.00%8.00%0.6215 LLM-VA36.00%22.00%28.00%12.00%0.7656LLM-VA28.00%22.00%66.00%2.00%0.6911 Phi-3.54B Original82.00%24.00%90.00%6.00%0.5073 Qwen3 4B Original84.00%32.00%66.00%0.00%0.5956 AlphaSteer+18.00%2.00%88.00%12.00%0.6250AlphaSteer+28.00%30.00%48.00%2.00%0.7353 AlphaSteer26.00%6.00%86.00%10.00%0.6190AlphaSteer34.00%26.00%56.00%0.00%0.7129 VectorSteer20.00%2.00%78.00%18.00%0.6380VectorSteer24.00%2.00%48.00%2.00%0.7979 SCANS4.00%0.00%96.00%40.00%0.4776SCANS28.00%10.00%66.00%2.00%0.7135 LLM-VA66.00%16.00%50.00%4.00%0.6822LLM-VA46.00%28.00%24.00%0.00%0.7822 Phi-4 4B Original60.00%16.00%68.00%16.00%0.5918 8B Original92.00%20.00%72.00%2.00%0.5753 AlphaSteer+16.00%8.00%74.00%6.00%0.6977AlphaSteer+18.00%18.00%60.00%10.00%0.7104 AlphaSteer18.00%6.00%70.00%6.00%0.7126AlphaSteer24.00%14.00%58.00%0.00%0.7474 VectorSteer20.00%8.00%68.00%6.00%0.7119VectorSteer22.00%14.00%40.00%2.00%0.8020 SCANS14.00%12.00%78.00%36.00%0.5513SCANS26.00%4.00%84.00%10.00%0.6310 LLM-VA70.00%26.00%48.00%8.00%0.6545LLM-VA36.00%8.00%24.00%0.00%0.8381 15B Original22.00%6.00%98.00%0.00%0.6182 14B Original86.00%30.00%86.00%0.00%0.5302 AlphaSteer+6.00%0.00%96.00%2.00%0.6623AlphaSteer+28.00%32.00%52.00%0.00%0.7255 AlphaSteer2.00%0.00%96.00%2.00%0.6711AlphaSteer26.00%10.00%46.00%0.00%0.7897 VectorSteer2.00%2.00%94.00%4.00%0.6667VectorSteer18.00%0.00%72.00%0.00%0.7399 SCANS12.00%0.00%94.00%6.00%0.6410SCANS30.00%18.00%82.00%40.00%0.4785 LLM-VA12.00%6.00%38.00%0.00%0.8526LLM-VA46.00%14.00%56.00%0.00%0.7129 CoLA MRPCMNLI RTE SSTGSM8K 0.25 0.50 0.75 1.00 AlphaSteer AlphaSteer+ Steer SCANS LLM-VA CoLA MRPCMNLI RTE SSTGSM8K 0.2 0.4 0.6 0.8 1.0 Llama 3.1 (8B) gemma 2 (9B) Mistral v0.3 (7B) Phi-3.5 (4B) Phi 4 (4B) Phi 4 (15B) Qwen2.5 (3B) Qwen2.5 (7B) Qwen2.5 (14B) Qwen3 (4B) Qwen3 (8B) Qwen3 (14B) Figure 5: Left: Average utility preservation by method. Right: Utility preservation per LLM with LLM-VA. Values near 1.0 indicate minimal degradation. Adaptive behavior. A key advantage of LLM- VA is its automatic adaptation to each model’s ini- tial safety bias. For models with high ASR but low ORR (e.g., Mistral-v0.3-7B with 81% ASR and 29% ORR), LLM-VA primarily reduces ASR to ensure safety. Conversely, for models with low ASR but high ORR (e.g., Llama-3.1-8B with 7% ASR and 53% ORR), it primarily decreases ORR to enhance usability. This adaptive behavior emerges naturally from vector alignment without manual hyperparameter tuning for different models. Cases requiring further analysis. Four mod- els (Phi-3.5-4B, Phi-4-4B, Mistral-v0.3-7B, and Qwen3-14B) do not achieve the highest F1 with LLM-VA. We analyze these cases in Section 5.4 and show that the suboptimal performance stems from iteration count sensitivity rather than funda- mental limitations of the approach. 5.3 Utility Preservation Results (RQ2) Besides effectiveness in resolving trade-off, we also evaluate model utility preservation on 6 bench- mark datasets covering classification and mathe- matical reasoning tasks. Figure 5 shows the results across methods and models. Overall utility preservation. LLM-VA pre- serves 95.92% of the original model’s utility on average, outperforming all baseline methods. For 9 of 12 LLMs, utility preservation exceeds 95%, demonstrating that LLM-VA successfully enhances alignment without sacrificing general capabilities. Comparison with baselines.SCANS shows the largest utility degradation (averaging 40.98%) be- cause aggressive magnitude adjustments disrupt the model’s internal representations. VectorSteer per- forms better (89.74%) but still falls short of LLM- VA due to its architectural modifications. AlphaS- teer and AlphaSteer+ achieve competitive preser- vation (94.50% and 94.48%) through null-space projection, but LLM-VA still outperforms them while achieving substantially better alignment. Task-specific analysis.The utility impact varies across task types. Classification tasks (COLA, MNLI, RTE, MRPC, SST) show minimal degrada- tion, with most models preserving over 97% perfor- mance. Mathematical reasoning (GSM8K) is more affected, with 91.60% average preservation. This is expected because math reasoning requires precise logical chains that can be disrupted by represen- 0306090 Rotation Degree 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 F1 Score Llama 3.1 (8B) gemma 2 (9B) Mistral v0.3 (7B) Phi-3.5 (4B) Qwen2.5 (3B) Qwen2.5 (7B) Qwen2.5 (14B) Qwen3 (4B) Qwen3 (8B) Qwen3 (14B) Phi 4 (4B) Phi 4 (15B) Figure 6: F1 scores with randomly distorted vectors at different anglesDfrom the original benign and answer vectors. 051015202530 Iteration 0.60 0.65 0.70 0.75 0.80 0.85 0.90 F1 Score F1 Score (a) Llama 3.1 (8B) 051015202530 Iteration 0.475 0.500 0.525 0.550 0.575 0.600 0.625 F1 Score F1 Score vs. Iteration F1 Score (b) Phi 4 (4B) 0102030405060 Iteration 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 F1 Score F1 Score (c) Qwen 3 (14B) Figure 7: F1 scores vs. iteration numberTfor three representative models. tation changes. Nevertheless, the impact remains limited compared to the alignment gains. Model size effects. Larger and more capable LLMs demonstrate better utility preservation. The three models with lowest preservation—Phi-3.5-4B (92.1%), Phi-4-4B (91.8%), and Mistral-v0.3-7B (93.2%)—are either among the smallest models or have documented limitations in benchmarks (Four- rier et al., 2024; Gao et al., 2021). This suggests that larger models have more robust internal repre- sentations that better tolerate the weight modifica- tions introduced by vector alignment. 5.4 Ablation Studies (RQ3) We analyze three key components: vector identifi- cation accuracy, iteration count, and layer selection. Vector Identification. To validate our SVM- based vector identification, we replacev a andv b with random vectorsDdegrees away from the orig- inals, whereDranges from30 ◦ to90 ◦ (Figure 6). The performance degradation correlates with dis- tortion angle: atD = 90 ◦ (orthogonal to the true vectors), F1 drops by 24.82% on average, and all 12 models underperform. AtD = 60 ◦ , all models still show degradation. However, atD = 30 ◦ , F1 only drops by 5.40%, indicating that LLM-VA is robust to small inaccuracies—a practical advantage since SVM hyperplanes may not perfectly capture true decision boundaries—while confirming that accurate identification remains essential. 0102030405060 Selected Layer Number 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 F1 Score Llama 3.1 (8B) gemma 2 (9B) Mistral v0.3 (7B) Phi-3.5 (4B) Phi 4 (4B) Phi 4 (15B) Qwen2.5 (3B) Qwen2.5 (7B) Qwen2.5 (14B) Qwen3 (4B) Qwen3 (8B) Qwen3 (14B) 0102030405060 Selected Layer Number 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Average Utility Preservation Llama 3.1 (8B) gemma 2 (9B) Mistral v0.3 (7B) Phi-3.5 (4B) Phi 4 (4B) Phi 4 (15B) Qwen2.5 (3B) Qwen2.5 (7B) Qwen2.5 (14B) Qwen3 (4B) Qwen3 (8B) Qwen3 (14B) Figure 8: Impact ofL select on F1 (left) and utility (right). Iteration Number. We vary iteration countT from 1 to 30 (Figure 7). For clarity, we show the results of three representative models and put the full results in Appendix F. Models exhibit distinct convergence patterns: Llama 3.1 (8B) shows rapid improvement and stabilizes aroundT = 19; Phi 4 (4B) peaks atT = 19but then degrades with additional iterations, suggesting over-modification; Qwen 3 (14B) continues improving throughT = 30and beyond (as shown in Figure 7, we extended to T = 60 and observed continued gains). These patterns explain the suboptimal results in Table 2: Mistral-v0.3-7B, Phi-3.5-4B and Phi- 4-4B suffer from over-modification (smaller and performance-limited models (Fourrier et al., 2024; Gao et al., 2021) are more susceptible to over- modification), while Qwen3-14B underperforms due to under-iteration. This suggests that model- specific iteration tuning or early stopping based on validation performance is important. Layer Selection.Figure 8 shows F1 and utility as L select varies from 30 to 60. For alignment, most models exhibit a non-monotonic trend with an op- timalL select : too few layers limit effectiveness, while too many cause overfitting. For utility preser- vation, most models remain stable untilL select ex- ceeds a threshold, at which point early layers are modified and utility drops sharply. This confirms that later layers are more relevant to safety deci- sions while early layers are critical for general ca- pabilities, motivating our contribution-score-based layer selection (Section 4.2). 6 Conclusion In this work, we presented LLM-VA, a novel approach that simultaneously addresses jailbreak and over-refusal by aligning the answer vec- tor with the benign vector through closed-form weight updates—making the model’s willingness to answer causally dependent on its safety judg- ment without requiring fine-tuning or architectural changes. Experiments on 12 widely used LLMs from 5 model families demonstrate a 11.45% F1 improvement over the best baseline while preserv- ing 95.92% utility, and our ablation studies confirm the importance of accurate vector identification and model-specific hyperparameter tuning. 7 Limitations Binary toxicity assumption. We consider only binary classification (benign vs. toxic), whereas real-world toxicity is nuanced and multi- dimensional. Extending LLM-VA to multi-class or fine-grained toxicity classification remains future work. Model scale. Our experiments cover models from 3B to 14B parameters. The effectiveness of LLM-VA on larger models (e.g., 70B+) remains to be validated, as these models may have differ- ent internal representations and require different hyperparameter settings. Training data dependency. LLM-VA requires labeled benign/toxic samples to train the SVMs for vector identification. The quality and repre- sentativeness of this training data directly affect alignment performance, and obtaining such labels may not always be straightforward. Reasoning models.Vector steering methods, in- cluding LLM-VA, are difficult to apply to LLMs with chain-of-thought reasoning. The control vec- tors must be identified after reasoning steps are generated, which is computationally expensive, and the randomness in reasoning makes accurate vector identification challenging. Model-specific tuning. As shown in our abla- tion studies, optimal iteration count and layer se- lection vary across models. While LLM-VA uses validation-based selection, this requires tuning for a new model, limiting plug-and-play applicability. Transferability. The performance of the exist- ing vector steering methods, including LLM-VA, on unseen datasets varies depending on tasks and models (Appendix G). This imply that current steer- ing methods may need to treat different tasks or domains separately, and improving transferability remains future work. Static alignment. The alignment is performed once and does not adapt to new threats or evolv- ing definitions of harmful content. Periodic re- alignment may be needed as the threat landscape changes. Customized Trade-off. LLM-VA aims to im- prove both jailbreak and over-refusal behavior si- multaneously. However, in certain applications (e.g., healthcare (Al-Garadi et al., 2025; Yang et al., 2024b) or PLC code generation (Liu et al., 2024d)), users may prefer to prioritize one aspect over the other. Extending LLM-VA to allow for customiz- able trade-offs remains future work. Experimental methodology. Our results are based on single runs with a fixed random seed. While we observe consistent improvements across 12 models, incorporating statistical significance tests would further strengthen our empirical find- ings. 8 Ethical Considerations 8.1 Potential Risks Though LLM-VA aims to enhance the safety align- ment of LLMs, it can be misused to manipulate model behaviors in unintended ways. For instance, attackers could potentially exploit the vector align- ment technique to bypass safety mechanisms or introduce harmful biases into the model. Besides, the datasets used for training and evaluation may contain biases. 8.2 AI Assistants Usage We employ GPT-5.2 (OpenAI, 2024) and Github Copilot 8 to assist in writing code for experiments. We carefully review and verify all AI-generated content to ensure accuracy and integrity. References Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, and 110 others. 2024. Phi-3 technical report: A highly capa- ble language model locally on your phone. Preprint, arXiv:2404.14219. AI@Meta. 2024. Llama 3 model card. Mohammed Al-Garadi, Tushar Mungle, Abdulaziz Ahmed, Abeed Sarker, Zhuqi Miao, and Michael E. Matheny. 2025. Large language models in healthcare. Preprint, arXiv:2503.04748. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 8 https://github.com/features/copilot 2024. Refusal in language models is mediated by a single direction. Advances in Neural Information Processing Systems, 37:136037–136083. Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009. The fifth pascal recognizing textual entailment challenge. TAC, 7(8):1. Zouying Cao, Yifei Yang, and Hai Zhao. 2025. Scans: Mitigating the exaggerated safety for llms via safety- conscious activation steering. In Proceedings of the AAAI Conference on Artificial Intelligence, vol- ume 39, pages 23523–23531. Kexin Chen, Yi Liu, Dongxia Wang, Jiaying Chen, and Wenhai Wang. 2024. Characterizing and evaluat- ing the reliability of llms against jailbreak attacks. Preprint, arXiv:2408.09326. Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti. 2024. Llama guard 3 vision: Safeguarding human-ai image understanding conver- sations. arXiv preprint arXiv:2411.10414. Paul F Christiano, Jan Leike, Tom Brown, Miljan Mar- tic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Ad- vances in neural information processing systems, 30. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word prob- lems. arXiv preprint arXiv:2110.14168. Corinna Cortes and Vladimir Vapnik. 1995. Support- vector networks. Machine learning, 20(3):273–297. Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho- Jui Hsieh. 2025.Or-bench:An over-refusal benchmark for large language models. Preprint, arXiv:2405.20947. Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024. Masterkey: Automated jailbreaking of large language model chatbots. In Proceedings 2024 Network and Distributed System Security Sym- posium, NDSS 2024. Internet Society. William B Dolan and Chris Brockett. 2005. Automati- cally constructing a corpus of sentential paraphrases. In Proceedings of the International Workshop on Paraphrasing. Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, and 1 others. 2022. Toy models of su- perposition. arXiv preprint arXiv:2209.10652. Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024.Open llm leaderboard v2.https://huggingface. co/spaces/open-llm-leaderboard/open_llm_ leaderboard. Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key- value memories. In Proceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guil- laume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. Preprint, arXiv:2310.06825. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Red- field, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Ken- ton Lee, and 1 others. 2019. Natural questions: a benchmark for question answering research. Trans- actions of the Association for Computational Linguis- tics, 7:453–466. Bruce W Lee, Inkit Padhi, Karthikeyan Natesan Rama- murthy, Erik Miehling, Pierre Dognin, Manish Na- gireddy, and Amit Dhurandhar. 2024. Programming refusal with conditional activation steering. arXiv preprint arXiv:2409.05907. Fan Liu, Zhao Xu, and Hao Liu. 2024a. Adversarial tuning: Defending against jailbreak attacks for llms. arXiv preprint arXiv:2406.06622. Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, Kai- long Wang, and Yang Liu. 2024b. Jailbreaking chat- gpt via prompt engineering: An empirical study. Preprint, arXiv:2305.13860. Yi Liu, Junzhe Yu, Huijia Sun, Ling Shi, Gelei Deng, Yuqi Chen, and Yang Liu. 2024c. Efficient detection of toxic prompts in large language models. In Pro- ceedings of the 39th IEEE/ACM International Con- ference on Automated Software Engineering, pages 455–467. Zihan Liu, Ruinan Zeng, Dongxia Wang, Gengyun Peng, Jingyi Wang, Qiang Liu, Peiyu Liu, and Wenhai Wang. 2024d. Agents4plc: Automating closed-loop plc code generation and verification in industrial con- trol systems using llm-based agents. arXiv preprint arXiv:2410.14209. Microsoft, :, Abdelrahman Abouelenin, Atabak Ash- faq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, Dong Chen, Dong- dong Chen, Junkun Chen, Weizhu Chen, Yen-Chun Chen, Yi ling Chen, Qi Dai, and 57 others. 2025. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. Preprint, arXiv:2503.01743. OpenAI. 2024.Gpt-4 technical report.Preprint, arXiv:2303.08774. Fabian Pedregosa, Gaël Varoquaux, Alexandre Gram- fort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vin- cent Dubourg, and 1 others. 2011. Scikit-learn: Ma- chine learning in python. the Journal of machine Learning research, 12:2825–2830. R. Penrose. 1955. A generalized inverse for matrices. Mathematical Proceedings of the Cambridge Philo- sophical Society, 51(3):406–413. Ruchira Ray and Ruchi Bhalani. 2024. Mitigating ex- aggerated safety in large language models. arXiv preprint arXiv:2405.05418. Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. Xstest: A test suite for identifying exag- gerated safety behaviours in large language models. Preprint, arXiv:2308.01263. Leheng Sheng, Changshuo Shen, Weixiang Zhao, Jun- feng Fang, Xiaohao Liu, Zhenkai Liang, Xiang Wang, An Zhang, and Tat-Seng Chua. 2025. Alphasteer: Learning refusal steering with principled null-space constraint. arXiv preprint arXiv:2506.07022. Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empiri- cal Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics. Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learn- ing to summarize with human feedback. Advances in neural information processing systems, 33:3008– 3021. Gemma Team. 2024a. Gemma. Qwen Team. 2024b. Qwen2.5: A party of foundation models. Qwen Team. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388. Alex Warstadt, Amanpreet Singh, and Samuel R. Bow- man. 2018. Neural network acceptability judgments. arXiv preprint 1805.12471. Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sen- tence understanding through inference. In Proceed- ings of the 2018 conference of the North American chapter of the association for computational linguis- tics: human language technologies, volume 1 (long papers), pages 1112–1122. Sophie Xhonneux, Alessandro Sordoni, Stephan Gün- nemann, Gauthier Gidel, and Leo Schwinn. 2024. Efficient adversarial training in llms with continuous attacks. Advances in Neural Information Processing Systems, 37:1502–1530. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Hao- ran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 40 others. 2024a. Qwen2 technical report. arXiv preprint arXiv:2407.10671. Yifan Yang, Qiao Jin, Furong Huang, and Zhiyong Lu. 2024b. Adversarial attacks on large language models in medicine. Preprint, arXiv:2406.12259. Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Jialuo Chen, Hui Xue, Xiaoxia Liu, Wenhai Wang, Kui Ren, and Jingyi Wang. 2025. S-eval: Towards automated and comprehensive safety evaluation for large language models. Proceedings of the ACM on Software Engi- neering, 2(ISSTA):2136–2157. Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Jiashui Wang, Xinlei Ying, Long Liu, and Wen- hai Wang. 2025a.Orfuzz: Fuzzing the "other side" of llm safety – testing over-refusal. Preprint, arXiv:2508.11222. Shenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu, Shengnan Guo, Zheng Fang, Lingchen Zhao, Chao Shen, Cong Wang, and Qian Wang. 2025b. Jbshield: Defending large language models from jailbreak at- tacks through activated concept analysis and manipu- lation. arXiv preprint arXiv:2502.07557. Haiquan Zhao, Chenhan Yuan, Fei Huang, Xiaomeng Hu, Yichang Zhang, An Yang, Bowen Yu, Dayi- heng Liu, Jingren Zhou, Junyang Lin, and 1 others. 2025. Qwen3guard technical report. arXiv preprint arXiv:2510.14276. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, and 1 others. 2023a. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405. Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrik- son. 2023b. Universal and transferable adversar- ial attacks on aligned language models. Preprint, arXiv:2307.15043. A Angles between Answer Vectors and Benign Vectors As shown in Figure 9, the angles between the an- swer vectors and benign vectors of different LLMs are approximately90 ◦ , indicating that they are nearly orthogonal. B Discussion on Layer Type Selection In Section 4.2, we mention that we treat both MLP and attention sublayers as “layers” for se- lection. This is because both types of sublayers contribute to the model’s internal representations and decision-making processes. Modifying either type can influence the model’s behavior regarding safety alignment. Besides, we conduct preliminary experiments to compare the combined score (Eq. 6) distributions of MLP and attention sublayers. The results are presented in Figure 10. The results show that both MLP and attention sublayers exhibit simi- lar contribution score distributions across different LLMs. Later layers tend to have higher contribu- tion scores, indicating their greater relevance to safety-related decisions. Therefore, we treat both MLP and attention sublayers equally in our layer selection process. C Additional Instructions on General Ability Datasets In this section, we provide detailed instructions on the general ability datasets used in our experiments. •Corpus of Linguistic Acceptability (COLA) (Warstadt et al., 2018) is a dataset for evaluating the grammatical acceptability of sentences. Each sample consists of a sentence and a binary label indicating whether the sentence is grammatically acceptable or not. •Multi-Genre Natural Language Inference (MNLI) (Williams et al., 2018) is a large-scale dataset for natural language inference. Each sam- ple consists of a pair of sentences annotated with textual entailment labels. •Recognizing Textual Entailment (RTE) (Ben- tivogli et al., 2009) is a dataset for evaluating the ability of models to recognize textual entail- ment. Each sample consists of a pair of sentences where one sentence is the premise and the other is the hypothesis. •Microsoft Research Paraphrase Corpus (MRPC) (Dolan and Brockett, 2005) is a dataset for evaluating the ability of models to recognize paraphrases. Each sample consists of a pair of sentences extracted from online news sources, with human annotations indicating whether each pair is semantically equivalent or not. • Stanford Sentiment Treebank (SST) (Socher et al., 2013) is a dataset for sentiment analysis. Each sample consists of a sentence and a binary label indicating whether the sentiment of the sen- tence is positive or negative. • GSM8K (Cobbe et al., 2021) is a dataset for evaluating the mathematical reasoning ability of models. Each sample consists of a math word problem and its corresponding solution. D Details about Experimental Setup We implement LLM-VA with max iteration num- berT = 30. The final modified model is obtained by selecting the best model on the validation set during the iterations. The numbers of selected lay- ersL select of each model are shown in Table 3. For the SVM-based vector identification, we use the default regularization parameterC = 1.0from scikit-learn (Pedregosa et al., 2011). For baseline methods, we follow the original papers and use the default hyperparameters. If the original pa- pers do not provide hyperparameter settings for certain models, we transfer the hyperparameters from similar models (e.g., models with the same ar- chitecture or in the same family). All experiments are conducted on 2×80 GB A100 GPUs. We use the default generation configurations in Hugging Face Transformers 9 for base LLMs during infer- ence. The temperature parameters of all models are set to 0.0 to ensure deterministic outputs. We use a fixed random seed of 42 for reproducibility across all experiments. E Details on Judge Model Selection As far as we know, Qwen3-Guard-Gen-8B (Zhao et al., 2025) is currently the only open-source LLM specifically designed to evaluate jailbreak and over- refusal behaviors. We also considered combin- ing multiple judge models to realize the evalua- tion (e.g., LlamaGuard 3 (Chi et al., 2024) for jail- break and OR-Judge (Zhang et al., 2025a) for over- 9 https://huggingface.co/docs/transformers/ index 051015202530 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Llama 3.1 (8B) 010203040 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out gemma-2 (9B) 051015202530 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Mistral-v0.3 (7B) 051015202530 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Phi-3.5 (4B) 051015202530 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Phi-4 (4B) 0510152025303540 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Phi-4 (15B) 05101520253035 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen2.5 (3B) 0510152025 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen2.5 (7B) 010203040 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen2.5 (14B) 05101520253035 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen3 (4B) 05101520253035 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen3 (8B) 0510152025303540 Layer 0 20 40 60 80 100 120 Angle (degrees) Angle between benign and answered activation vs. Layer 90 degrees attn_out mlp_out Qwen3 (14B) Figure 9: Angles between answer vectors and benign vectors of different LLMs. 051015202530 Layer Number 0.0 0.2 0.4 0.6 0.8 1.0 Combined Scores Attention Layer MLP Layer Llama-3.1 (8B) 010203040 Layer Number 0.1 0.0 0.1 0.2 0.3 0.4 0.5 Combined Scores Attention Layer MLP Layer gemma-2 (9B) 051015202530 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Mistral-v0.3 (7B) 051015202530 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Phi-3.5 (4B) 051015202530 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Phi-4 (4B) 0510152025303540 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 Combined Scores Attention Layer MLP Layer Phi-4 (15B) 05101520253035 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Qwen2.5 (3B) 0510152025 Layer Number 0.0 0.2 0.4 0.6 0.8 Combined Scores Attention Layer MLP Layer Qwen2.5 (7B) 010203040 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Qwen2.5 (14B) 05101520253035 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 Combined Scores Attention Layer MLP Layer Qwen3 (4B) 05101520253035 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Qwen3 (8B) 0510152025303540 Layer Number 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Combined Scores Attention Layer MLP Layer Qwen3 (14B) Figure 10: Comparison of combined scores between MLP and attention sublayers across different LLMs. Table 3: Selected layer numbers of different models in LLM-VA. ModelLlama-3.1Gemma-2Mistral-v0.3Phi-3.5Phi-4Qwen2.5Qwen3 Size8B9B7B4B4B15B3B7B14B4B8B14B # Selected Layers426042304854603654486048 051015202530 Iteration 0.0 0.2 0.4 0.6 0.8 1.0 ORR/ASR Rate vs. Iteration ASR (seval_attack) ASR (seval_risk) ORR (orfuzzset) ORR (nq) Figure 11: An example of evaluation results with com- bined judge models. refusal). However, due to their different judgement criteria, combining multiple judge models may lead to inconsistent evaluations. As a result, LLM-VA will find incorrect vectors to align, leading to subop- timal performance. Figure 11 shows an example of such inconsistent evaluations. The ORR evaluated by OR-Judge reaches 100% due to the inconsis- tency between the two judge models. Therefore, we choose Qwen3-Guard-Gen-8B as the sole judge model for a consistent evaluation of both jailbreak and over-refusal behaviors. F Detailed Results on Iteration Number The detailed results on the impact of iteration num- ber of each LLM are shown in Figure 12. G Transferability Experiments To evaluate the transferability of LLM-VA, we as- sess how well the vector alignment learned on the training datasets generalizes to unseen datasets. We evaluate the modified models on three additional jailbreak datasets (XSTest-Toxic (Röttger et al., 2024), OR-Bench-Toxic (Cui et al., 2025), and Ad- vBench (Zou et al., 2023b)) and two over-refusal datasets (XSTest (Röttger et al., 2024) and OR- Bench (Cui et al., 2025)) that are not included in the training set. The results are shown in Table 4. The results show that the performance of LLM- VA on unseen datasets varies across different models. While LLM-VA maintains reasonable safety alignment on most unseen datasets, the per- formance degradation compared to the training datasets indicates that further research is needed to improve the generalization of vector steering methods. 051015202530 Iteration 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 F1 Score F1 Score vs. Iteration F1 Score gemma-2 (9B) 051015202530 Iteration 0.550 0.575 0.600 0.625 0.650 0.675 0.700 0.725 0.750 F1 Score F1 Score vs. Iteration F1 Score Mistral-v0.3 (7B) 0510152025 Iteration 0.1 0.2 0.3 0.4 0.5 0.6 0.7 F1 Score F1 Score vs. Iteration F1 Score Phi-3.5 (4B) 051015202530 Iteration 0.60 0.65 0.70 0.75 0.80 0.85 0.90 F1 Score F1 Score vs. Iteration F1 Score Phi-4 (15B) 051015202530 Iteration 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 F1 Score F1 Score vs. Iteration F1 Score Qwen2.5 (3B) 051015202530 Iteration 0.55 0.60 0.65 0.70 0.75 F1 Score F1 Score vs. Iteration F1 Score Qwen2.5 (7B) 051015202530 Iteration 0.525 0.550 0.575 0.600 0.625 0.650 0.675 0.700 F1 Score F1 Score vs. Iteration F1 Score Qwen2.5 (14B) 051015202530 Iteration 0.55 0.60 0.65 0.70 0.75 0.80 F1 Score F1 Score vs. Iteration F1 Score Qwen3 (4B) 010203040 Iteration 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 F1 Score F1 Score vs. Iteration F1 Score Qwen3 (8B) Figure 12: Detailed results on the impact of iteration number of each LLM. Table 4: Transferability results on unseen datasets. ModelSizeMethodAdvBench ASR↓ OR- Bench- Toxic ASR↓ XSTest- Toxic ASR↓ OR- Bench ORR↓ XSTest ORR↓ Final F1↑ ModelSizeMethodAdvBench ASR↓ OR- Bench- Toxic ASR↓ XSTest- Toxic ASR↓ OR- Bench ORR↓ XSTest ORR↓ Final F1↑ Llama-3.18B Original0.58%3.05%0.00%48.22%16.57%0.7077 Qwen2.5 3B Original0.19%2.29%0.57%55.42%19.34%0.6522 AlphaSteer+0.58%3.05%0.00%48.67%17.13%0.7038AlphaSteer+0.00%2.44%0.00%53.53%21.55%0.6649 AlphaSteer0.19%1.37%0.00%61.87%27.07%0.5921AlphaSteer0.19%1.53%0.00%59.97%21.55%0.6144 Steer0.00%0.15%0.00%85.60%49.17%0.3163Steer0.00%0.15%0.00%85.97%37.57%0.3313 SCANS0.19%1.37%0.00%72.48%26.52%0.4945SCANS6.35%16.64%1.15%40.11%13.81%0.7305 Modified7.69%5.80%4.02%46.85%6.63%0.7088Modified0.38%4.12%0.57%54.21%21.55%0.6555 gemma-29B Original0.58%1.98%0.00%80.52%28.73%0.4059 7B Original0.38%6.72%0.00%24.64%8.84%0.8569 AlphaSteer+0.77%0.92%0.57%81.58%26.52%0.3985AlphaSteer+0.77%6.11%0.00%24.87%8.84%0.8563 AlphaSteer0.00%0.46%0.00%88.86%28.73%0.3103AlphaSteer3.08%3.66%0.00%34.50%9.39%0.8006 Steer0.00%0.31%0.00%94.16%56.35%0.1882Steer8.65%5.50%0.00%61.03%29.28%0.5776 SCANS3.08%5.50%0.57%59.82%29.28%0.5952SCANS1.54%6.87%1.72%40.56%11.60%0.7552 Modified2.88%1.68%0.00%82.41%29.83%0.3809Modified3.65%9.92%0.00%19.94%7.73%0.8714 Mistral-v0.37B Original54.42%48.70%16.67%7.28%3.87%0.7920 14B Original0.00%4.58%0.00%21.15%8.29%0.8816 AlphaSteer+52.12%48.85%16.67%7.35%4.42%0.7937AlphaSteer+0.19%5.34%0.00%20.77%8.29%0.8817 AlphaSteer45.38%48.09%14.37%7.66%4.97%0.8021AlphaSteer0.00%3.21%0.00%26.31%9.39%0.8551 Steer42.69%40.92%5.75%11.75%11.05%0.7970Steer0.00%0.15%0.00%57.01%16.02%0.6477 SCANS74.42%41.98%22.99%18.57%7.73%0.7209SCANS7.88%16.95%2.30%30.40%19.34%0.7824 Modified27.31%17.25%2.87%37.30%10.50%0.7195Modified0.77%5.04%0.00%17.74%6.08%0.8990 Phi-3.54B Original2.12%4.89%1.72%45.49%13.26%0.7234 Qwen3 4B Original0.96%4.73%0.57%44.35%6.63%0.7402 AlphaSteer+1.15%4.43%1.15%43.90%13.81%0.7365AlphaSteer+7.88%24.58%4.02%21.08%4.42%0.8307 AlphaSteer1.15%3.66%1.15%48.98%13.81%0.7022AlphaSteer33.85%37.10%8.62%22.14%7.73%0.7634 Steer1.15%3.51%1.15%56.41%22.10%0.6373Steer0.77%1.83%0.00%66.49%19.89%0.5583 SCANS1.54%1.37%1.15%79.83%37.57%0.3994SCANS3.65%6.11%0.57%46.93%11.60%0.7107 Modified6.54%6.41%0.00%43.21%12.71%0.7306Modified2.31%4.27%0.57%36.69%7.73%0.7880 Phi-4 4B Original0.58%2.14%0.00%58.83%17.68%0.6265 8B Original0.96%3.21%1.15%44.12%9.39%0.7419 AlphaSteer+0.19%1.53%0.00%59.97%18.23%0.6182AlphaSteer+7.31%14.20%7.47%62.70%22.10%0.5560 AlphaSteer0.19%1.53%0.00%55.42%18.23%0.6551AlphaSteer0.58%4.89%1.15%34.65%7.18%0.8025 Steer0.19%2.60%0.00%46.70%16.57%0.7201Steer0.96%4.12%0.57%40.03%9.39%0.7677 SCANS10.19%12.98%0.57%36.01%8.84%0.7621SCANS0.58%2.29%0.00%70.96%20.99%0.5147 Modified0.58%7.63%1.15%18.73%11.60%0.8841Modified0.96%3.82%0.00%40.71%8.29%0.7651 15B Original0.19%3.97%0.00%72.71%17.13%0.5007 14B Original0.19%4.43%0.57%40.33%7.18%0.7683 AlphaSteer+0.00%3.97%0.00%72.63%15.47%0.5039AlphaSteer+5.96%22.44%3.45%13.65%6.63%0.8743 AlphaSteer0.00%3.36%0.00%75.97%16.57%0.4704AlphaSteer0.77%12.21%0.57%20.09%5.52%0.8719 Steer0.00%2.29%0.00%84.23%20.99%0.3762Steer0.00%0.31%0.57%72.33%20.99%0.5052 SCANS1.35%5.95%0.00%67.78%11.60%0.5490SCANS29.04%26.11%9.20%35.94%25.41%0.6955 Modified0.77%3.82%0.00%53.37%11.60%0.6727Modified4.81%3.05%0.00%51.86%11.05%0.6801