Paper deep dive
Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
Andrey Labunets
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/27/2026, 5:13:23 AM
Summary
This paper investigates the geometric properties of refusal training in language models, specifically using OLMo-2-0425-1B-Instruct as a case study. It demonstrates that refusal geometry (the low-dimensional subspace mediating refusal behavior) is directly shaped by the training process, particularly the concentration of gradients arising from repetitive refusal prefixes. The authors introduce the concept of 'stable rank' as a measure of this concentration, showing that diverse refusal starts increase the stable rank of gradients and activation changes. This increased stable rank acts as a 'hardening lever,' making the refusal mechanism more robust against vector ablation attacks compared to models trained with repetitive refusal prefixes.
Entities (7)
Relation Signals (6)
Refusal Subspace → mediates → Refusal Behavior
confidence 95% · refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace
Diverse Refusal Starts → increases → Stable Rank
confidence 92% · diverse refusal starts can raise stable ranks of gradients and activation changes
Refusal Geometry → isshapedby → Refusal Training
confidence 90% · Refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace.
High Stable Rank → reducesvulnerabilityto → Vector Ablation Attack
confidence 90% · diverse refusal starts can raise stable ranks... making refusals harder to remove with a vector ablation attack.
Repetitive Refusal Starts → causes → Gradient Concentration
confidence 88% · brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace.
Gradient Concentration → leadsto → Low Stable Rank
confidence 85% · repetitive refusal starts... linked to concentration of gradients... making refusals harder to remove [when diverse]... implying repetitive leads to lower stable rank/higher vulnerability.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge and concentrated, low-dimensional structure. In a case study of OLMo-2-0425-1B-Instruct we find that the refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace. We study refusal directions through the training dynamics across refusal datasets and reveal that their brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace. Across frozen-model analyses and controlled synthetic fine-tuning, we find evidence of a hardening lever: diverse refusal starts can raise stable ranks of gradients and activation changes, making refusals harder to remove with a vector ablation attack.
Tags
Links
- Source: https://arxiv.org/abs/2608.25390v1
- Canonical: https://arxiv.org/abs/2608.25390v1
Trouble viewing inline? Open PDF directly →
Full Text
92,638 characters extracted from source content.
Expand or collapse full text
Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks Andrey Labunets Affiliation: UC San Diego Email: alabunets@ucsd.edu Abstract Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge and concentrated, low-dimensional structure. In a case study of OLMo-2-0425-1B-Instruct we find that the refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace. We study refusal directions through the training dynamics across refusal datasets and reveal that their brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace. Across frozen-model analyses and controlled synthetic fine-tuning, we find evidence of a hardening lever: diverse refusal starts can raise stable ranks of gradients and activation changes, making refusals harder to remove with a vector ablation attack. 1 Introduction State-of-the-art AI systems are beginning to exhibit dual-use capabilities that could pose severe risks if misused, such as exploit generation [3], or assistance with biological threat creation or bio-weaponization [26, 18]. This motivates the now-standard view that, as models become more capable, protections should prevent harmful outputs and catastrophic misuse of AI [8, 2]. AI safety and security study ways to reduce these risks, which is typically done with multi-layered approaches. Frontier-safety frameworks emphasize capability evaluations, confidentiality of model weights, deployment-time mitigations, monitoring, access controls, and broader ecosystem-level defenses [8, 2, 9, 24, 17]. Model-weight security is important because weight access can allow protections to be modified or removed; at the same time, weights access scenario is realistic for open models, insider compromise, model theft, and sufficiently capable adversaries [25, 10, 18]. A central technical layer of protection for deployed models, including open-weights models, is safety alignment: models are trained to remain helpful on benign requests while refusing instructions that request harmful assistance or dangerous capabilities [27, 5, 34]. These protections are also evaluated against jailbreak attacks that aim to bypass safety alignment [32, 41, 44, 12, 22]. This safety-aligned behavior is often induced by training on harmful prompts paired with safe refusal completions. Recent evidence, however, shows that in local, white-box settings those refusals in aligned language models can be mediated by a low-dimensional mechanism: refusal direction, or a single vector in representation space, projections on which strongly control refuse/comply behavior [4]. More broadly, activation-steering methods construct directions for specified concepts or behaviors, using prompt pairs, contrastive datasets, supervised concept labels, or optimized target completions to control properties such as topic, sentiment, factuality, deception, or safety-relevant behavior [13, 38, 40, 30, 7, 16]. Refusal ablation, however, is a particularly important special case because it targets the shared refuse/comply response mechanism itself. Existing research mechanistically studied representations of refusal and their geometry, showing that refusal concepts span multiple orthogonal and independent directions, where auxiliary principal components mediate distinct safety-relevant features [42, 29]. In particular, [29] shows that safety-aligned behavior is controlled by a low-dimensional space. We use refusal subspace to denote such an affine subspace of refusal-related activations. A separate line of work studied forms of concentration leading to low-rank or low-effective-dimensionality representations, including shallow concentration on early safety tokens [34], low-stable-rank transformer activations [14], rank collapse in pure attention [15], neural collapse in final and intermediate classifier representations [31, 35], and regularities in refusals [33]. However, the explanation of why meaningful, safety-critical representations of refusal across a broad range of harmful prompts and topics collapse to a single effective direction or a low-dimensional subspace remained elusive. In this work, we bridge this gap by studying robustness to refusal ablation attacks analytically and empirically in OLMo-2-0425-1B-Instruct: (a) Low-dim refusal residuals (b) Diverse refusal residuals (c) Refusal ablation attack Figure 1: Higher stable rank resists refusal ablation (conceptual schematic). When residuals are less concentrated, a shared single-vector ablation cannot fully disable the refusal mechanism, so the refusal score after attack remains higher. Figure 1 illustrates our hypothesized mechanism: concentrated update directions produce refusal residuals with lower effective dimensionality, making refusal more vulnerable to a shared single-vector ablation. Concretely, in these settings, we show the following contributions: • Training-origin geometry. We show that cross-entropy losses on the first tokens of harmful prompts’ completions induce activation changes with partial geometric correspondence to the refusal subspace: their means exhibit completion-dependent signed alignment with the refusal direction, and their principal directions partially overlap direction space of refusal subspace. For refusal completions, the per-gradient-step activation-change and final refusal-residual matrices also have comparable low stable rank. • Prefix concentration and spectral collapse. We identify concentrated refusal-start support in safety datasets — especially repeated first tokens — as one mechanism that can produce low-stable-rank gradients, activation changes, and refusal residuals. • Cross-layer stable rank transfer. We introduce stable rank transfer factors and show that refusal-diversity-induced stable rank increases propagate back through late layers with bounded ratios across evaluated conditions. • Shared effective Jacobian explanation. We define top-k source-subspace shared maps and a relative transfer factor fit error that yields data-dependent two-sided stable-rank bounds on the gradients and gradient-induced activation changes. On the fitted supports, these maps reconstruct a meaningful though incomplete component of their cross-layer relation and account for its observed rank amplification or contraction. • Rank–vulnerability connection. Across frozen-model analyses and controlled fine-tuning, greater refusal-start diversity is associated with higher stable ranks and weaker difference-in-means single-vector ablation, providing evidence of a simple hardening lever. 2 Background and definitions 2.1 Transformer architecture and notation We consider a decoder-only transformer language model for next-token prediction with vocabulary V=1,…,|V|V=\1,…,|V|\, hidden size d, and l∈0,…,Ll∈\0,…,L\ layers (transformer blocks). Training examples are pairs token sequences, prompt x and a completion, with tokens from V. We index training examples by i. We use activation to denote a column vector of transformer hidden state at a specified layer and token position. For example i, let hi(l)∈ℝdh_i^(l) ^d denote the layer-l activation at the first-assistant-token prediction position, while, more broadly, all activations across N training examples at layer l and the same final position are denoted as a matrix of transposed activation column vectors: H(l)=[h1(l)⊤hN(l)⊤]∈ℝN×d. H^(l)= bmatrixh^(l) _1\\ \\ h^(l) _N bmatrix ^N× d. (1) The model classifier head (unembedding) is parametrized by a matrix W∈ℝd×|V|W ^d×|V| (which consists of embedding vectors w1,…,w|V|w_1,...,w_|V|) and bias b∈ℝ|V|b ^|V|, yielding logits o, a probability distribution p, and a sampled token y∈Vy∈ V with probability p(y)p(y): o o =W⊤h(L)+b∈ℝ|V|, =W h^(L)+b ^|V|, (2) p p =softmax(o)∈Δ|V|−1=p∈ℝ|V|:p≥0,∑k=1|V|pk=1,pk=exp(ok)∑j=1|V|exp(oj) =softmax(o)∈ ^|V|-1= \p ^|V|:\ p≥ 0,\ _k=1^|V|p_k=1 \, p_k= (o_k) _j=1^|V| (o_j) (3) All activations in this paper used for refusal-vector estimation are evaluated at the same position, which is the first-assistant token output position, consistent with [4]. For SVD of a matrix M=UMΣMVM⊤M=U_M _MV_M , we write (M)=span(VM)V(M)=span(V_M) to denote the right-singular subspace of M.. We define stable rank as sr(M)=‖M‖F2/‖M‖22sr(M)=\|M\|_F^2/\|M\|_2^2. 2.2 Refusal mediation and existing attacks One of the well-known refusal ablation attacks involves estimating a refusal direction from a pair of harmless and harmful datasets [4]. For a fixed layer l let harm,i,hbenign,j∈ℝdh_harm,i,h_benign,j ^d denote the corresponding activations for harmful and benign prompts i and j evaluated at the first assistant token position as described above, where 1≤i≤nh1≤ i≤ n_h and 1≤j≤nb1≤ j≤ n_b. Define empirical means for those activations: μharm=1nh∑i=1nhhharm,i,μbenign=1nb∑j=1nbhbenign,j, _harm\;=\; 1n_h _i=1^n_hh_harm,i, _benign\;=\; 1n_b _j=1^n_bh_benign,j, (4) We operationalize the refusal direction with the difference-in-means estimator: ρ=μharm−μbenign∈ℝd. ρ\;=\; _harm- _benign ^d. (5) The refusal ablation attack removes a projection on this normalized refusal vector from activations h at a run-time when a model is prompted, producing new activations h(abl)h^(abl): h(abl)←h−⟨h⊤⋅ρ^⟩ρ h^(abl)\; \;h- h · ρ ρ (6) In the above, ρ^=ρ/∥ρ∥2 ρ=ρ/ ρ _2 denotes normalized refusal direction. This refusal ablation attack jailbreaks the model, making it respond to harmful prompts it was trained to refuse. Additionally, we define benign-centered refusal residual matrix as: ΔH=[(harm,1−μbenign)⊤(harm,nh−μbenign)⊤]∈ℝnh×d H\;=\; bmatrix(h_harm,1- _benign) \\ \\ (h_harm,n_h- _benign) bmatrix ^n_h× d (7) Given Equation 5, we can say that the matrix ΔH H is a matrix of per-prompt refusal residuals ρi _i: its mean equals the difference-in-means refusal direction: μΔ=1nh∑i=1nh(harm,i−μbenign)=μharm−μbenign=ρ∈ℝd _ = 1n_h _i=1^n_h(h_harm,i- _benign)= _harm- _benign=ρ ^d (8) Let ΔHc=ΔH−nhμΔ⊤ H_c= H-1_n_h _ denote the centered refusal-residual matrix. We operationalize the refusal subspace as the affine subspace μΔ+(ΔHc) _ +V( H_c), where (ΔHc)V( H_c) is its direction subspace. 3 Refusal training gradients shape refusal residuals in transformer We analyze a transformer under supervised refusal training and show how repeated refusal-completion first tokens can concentrate activation gradients and, in turn, refusal residuals. We conjecture that those refusal-completion gradients at refusal safety-training induce activation changes which can contribute to the refusal direction. Limitations and assumptions are outlined in Appendix B. Cross-entropy gradients to last layer activations. Consider the cross-entropy loss on the first token r∈Vr∈ V of a refusal completion, represented by the one-hot target vector ere_r: ℓCE(er,p) ^CE(e_r,p) =−logp(r)=−or+logZ(o). =- p(r)=-o_r+ Z(o). (9) The activation gradient ∇hℓCE(er,p) _h ^CE(e_r,p) defines the hypothetical feature-space step h+=h−γ∇hℓCE(er,p), h^+=h-γ _h ^CE(e_r,p), (10) We use this gradient, also termed as virtual update by [11], as a diagnostic to track feature trajectory. For the last-layer activation h, the gradient is gi=∇hℓiCE(eri,pi)=W(pi−eri)=−wri+Wpi∈ℝd. g_i= _h _i^CE(e_r_i,p_i)=W(p_i-e_r_i)=-w_r_i+Wp_i ^d. (11) Or, in a matrix form, for a set of N examples with target refusals tokens rir_i, a one-hot matrix Er=[er1,…,erN]⊤E_r=[e_r_1,…,e_r_N] representing those refusal tokens, and a prediction matrix P=[p1,…,pN]⊤P=[p_1,…,p_N] , the cross-entropy gradients to last-layer activations: GHCE=(P−Er)W⊤∈ℝN×d. G_H^CE=(P-E_r)W ^N× d. (12) The mean activation gradient is: g¯=1N∑i=1Ngi=W(p¯−e¯r)=−w¯r+Wp¯∈ℝd. g= 1N _i=1^Ng_i=W( p- e_r)=- w_r+W p ^d. (13) In the static refusal case with ri=r_i=r for all i and e¯r=er e_r=e_r, the mean direction becomes: g¯=−wr+Wp¯ g=-w_r+W p. Shared direction across training example gradients. We can observe in Equation 11 and Equation 13 that in case of static, repeating refusal token r, gradients across N examples will be concentrated around shared direction wrw_r. In case other term WpWp does not cancel or bias wrw_r too much (for example, when p has large mass elsewhere), gradients concentrated around wrw_r will be low rank. From gradients to gradient-induced activation updates. Activations are changed after parameter updates (training steps) minimizing aforementioned loss: the raw gradient gig_i is not itself the realized activation change. Concretely, let Ji:=∇θhi(θ)J_i:= _θh_i(θ) be its Jacobian. Then, ∇θℓiCE=Ji⊤gi _θ _i^CE=J_i g_i. Under a small parameter update the realized change would be: hi(θ−η∇θℓiCE)−hi(θ)≈−ηJiJi⊤gi. h_i(θ-η _θ _i^CE)-h_i(θ)≈-η J_iJ_i g_i. (14) We define the gradient-induced activation update as the realized finite difference, taken with a negative sign for gradient-aligned comparison: i=−(hi(θ−η∇θℓiCE)−hi(θ))≈ηJiJi⊤gi∈ℝd. g_i=- (h_i(θ-η _θ _i^CE)-h_i(θ) )≈η J_iJ_i g_i ^d. (15) Stacking these updates across N examples gives the gradient-induced activation update matrix G: =[1⊤N⊤]∈ℝN×d. = bmatrix g_1 \\ \\ g_N bmatrix ^N× d. (16) Training-time origin of refusal directions Assuming gradient descent over harmful prompts paired with refusal completions, we hypothesize that the refusal subspace μΔ+(ΔHc) _ +V( H_c) is approximated by the sign-reversed affine subspace of gradient-induced activation updates −μ+(c)- _G+V(G_c), where cG_c denotes a mean-centered matrix G μΔ+(ΔHc)≈−μ+(c). _ +V( H_c)≈- _G+V(G_c). (17) Concretely, we test the above prediction using these three diagnostics: cos(μΔ,μ) ( _ , _G) ≈−1, ≈-1, training steps move activations opposite to gradients, (18) (ΔHc) ( H_c) ≈(c), (G_c), right singular vectors span similar subspaces, (19) sr(ΔH) \! ( H ) ≈sr() \! (G ) their stable ranks are very close. (20) Stable rank of first-token target matrix. Let Er∈ℝN×|V|E_r ^N×|V| be the one-hot matrix of refusal first tokens r1,…,rNr_1,…,r_N. Across N examples, for a token k∈Vk∈ V, let ck=#i:ri=kc_k=\#\i:r_i=k\ denote its count , fk=ck/Nf_k=c_k/N denote its empirical frequency, and let m=|k:ck>0|m=|\k:c_k>0\| be the number of observed first-token buckets. The following lemma, whose full statement, proof, and a corollary appear in Section E.1, shows that first-token concentration directly controls the stable rank of ErE_r. Lemma 3.1 (Stable rank of refusal first-token targets). The stable rank of ErE_r is sr(Er)=Nmaxkck=1maxkfk. (E_r)= N _kc_k= 1 _kf_k. (21) In particular, sr(Er)=1sr(E_r)=1 when all examples share the same first refusal token. For a fixed number m of observed first-token buckets, sr(Er)sr(E_r) is maximized by balanced counts across the buckets maxsr(Er)=N⌈N/m⌉. (E_r)= N N/m . (22) Thus, a repeated refusal-completion first token, such as the token beginning “I’m sorry,” minimizes sr(Er)sr(E_r), while balanced first-token frequencies increase it. Since the last-layer first-token gradient matrix is GHCE=(P−Er)W⊤G_H^CE=(P-E_r)W , this gives a testable prediction: when the prediction term P and unembedding W do not collapse this target-rank signal, we can expect that increasing refusal first-token diversity can increase the stable rank of gradients G and their corresponding gradient-induced activation updates G. We therefore report sr(Er)sr(E_r), sr(P−Er)sr(P-E_r), and sr((P−Er)W⊤)sr((P-E_r)W ) separately to measure P and W⊤W preserve or distort the first-token target-matrix signal. Stable rank distortion across layers: an idealized case. Assume in idealized case that for some corresponding Jacobian J(l)=∂h(l)∂h(l−1)∈ℝd×dJ^(l)= ∂ h^(l)∂ h^(l-1) ^d× d, gradients at previous layer l−1l-1 satisfy G(l−1)=G(l)J(l), G^(l-1)=G^(l)J^(l), (23) while that Jacobian is shared across training examples. Assume also that the Jacobian is well-conditioned, i.e. its condition number κ(J(l))=‖J(l)‖2‖(J(l))−1‖2κ(J^(l))=\|J^(l)\|_2\|(J^(l))^-1\|_2 is small. Under these assumptions, an inequality below bounds how much stable rank can change throughout remaining layers 1≤l≤L−11≤ l≤ L-1 under this layerwise map; see Section E.2 for the proof. Theorem 3.2 (Bounded stable rank distortion under a well-conditioned factor). Let A∈ℝN×nA ^N× n, let B∈ℝn×nB ^n× n be invertible, and let κ(B)=‖B‖2‖B−1‖2κ(B)=\|B\|_2\|B^-1\|_2 be a condition number. Then 1κ(B)2sr(A)≤sr(AB)≤κ(B)2sr(A). 1κ(B)^2sr(A) (AB)≤κ(B)^2sr(A). (24) Applying Theorem 3.2 to Equation 23 gives 1κ(J(l))2sr(G(l))≤sr(G(l−1))≤κ(J(l))2sr(G(l)). 1κ(J^(l))^2sr(G^(l)) (G^(l-1))≤κ(J^(l))^2sr(G^(l)). (25) Iterating from layer L to layer l yields sr(G(L))∏k=l+1Lκ(J(k))2≤sr(G(l))≤(∏k=l+1Lκ(J(k))2)sr(G(L)). sr(G^(L)) _k=l+1^Lκ(J^(k))^2 (G^(l))≤ ( _k=l+1^Lκ(J^(k))^2 )sr(G^(L)). (26) Assuming in our case G(L)=GHCE=(P−Er)W⊤G^(L)=G_H^CE=(P-E_r)W as per Equation 12, we obtain the layerwise prediction for l-layer gradients G(l)G^(l), bounded by functions of G(L)G^(L) from both sides: sr((P−Er)W⊤)∏k=l+1Lκ(J(k))2≤sr(G(l))≤(∏k=l+1Lκ(J(k))2)sr((P−Er)W⊤). sr((P-E_r)W ) _k=l+1^Lκ(J^(k))^2 (G^(l))≤ ( _k=l+1^Lκ(J^(k))^2 )sr((P-E_r)W ). (27) This gives another testable prediction: under our idealized conditions, stable rank increases of last- layer’s gradients G(L)G^(L) can also propagate back, increasing stable rank of layer-l gradients G(l)G^(l). The next parts refine the idealized condition above by introducing a data-dependent variant of a shared factor and data-dependent stable rank transfer factors in order to construct tighter bounds. Definition 3.3 (Stable rank transfer factor). For nonzero gradient matrices G(L)G^(L) and G(l)G^(l) computed for the same N examples at layers L and l respectively, their stable rank transfer factor is the ratio of their stable ranks: τL→l:=sr(G(l))sr(G(L)). τ^L→ l:= sr (G^(l) )sr (G^(L) ). (28) Thus, τL→l>1τ^L→ l>1 denotes stable rank amplification, while τL→l<1τ^L→ l<1 denotes stable rank contraction between the selected layers. We similarly define transfer factors for gradient-induced updates G. To refine impractical assumptions of Equation 23, instead of relying on an idealized, shared Jacobian, we introduce an appropriate approximation, or a shared cross-layer relation map. Definition 3.4 (Top-k source-subspace shared map). For nonzero gradient matrices G(L)G^(L) and G(l)G^(l) computed for the same N examples at layers L and l respectively, a top-k source-subspace shared map J^kL→l J_k^L→ l is a shared data-dependent low-rank linear approximation for the observed cross-layer relation between gradients fitted on N examples and restricted to a top-k principal subspace of G(L)G^(L). It is defined as follows: J^kL→l=Vk(L)ΓkL→l, J_k^L→ l=V_k^(L) _k^L→ l, (29) G^k(l)=G(L)J^kL→l. G_k^(l)=G^(L) J_k^L→ l. (30) where matrix Vk(L)∈ℝd×kV_k^(L) ^d× k denotes top k right singular vectors of G(L)G^(L), therefore rank(J^kL→l)≤krank( J_k^L→ l)≤ k. ΓkL→l∈ℝk×d _k^L→ l ^k× d is a coefficient matrix estimated from N gradients. G^k(l) G_k^(l) is the layer-l gradient matrix implied by the top-k source-subspace shared map J^kL→l J_k^L→ l. For a nonzero G^k(l) G_k^(l), the map-implied stable rank transfer factor is the stable rank transfer factor our shared map implies on G(L)G^(L): τ^kL→l=sr(G^k(l))sr(G(L)) τ_k^L→ l= sr ( G_k^(l) )sr (G^(L) ) (31) Next, we account of errors resulting from the above approximation. Definition 3.5 (Relative transfer factor fit error). For the observed τL→lτ^L→ l stable rank transfer factor, a shared top-k source-subspace map J^kL→l J_k^L→ l, and a stable rank transfer factor τ^kL→l τ_k^L→ l implied by it, define a relative transfer factor error of J^kL→l J_k^L→ l as follows: ϵτ,kL→l=|τ^kL→l−τL→l|τL→l. _τ,k^L→ l= | τ_k^L→ l-τ^L→ l |τ^L→ l. (32) Theorem 3.2 gives a worst-case condition-number bound under an idealized shared linear factor. Lemma 3.6 instead shows how an upper bound on the transfer-factor fit error yields a data-dependent two-sided bound on the stable rank of G(l)G^(l) (see Lemma E.5 for the proof); we will later estimate that upper bound empirically. Lemma 3.6 (Stable rank bounds from transfer factor fit error). Let G(L),G(l)∈ℝN×dG^(L),G^(l) ^N× d be nonzero gradient matrices, and let J^kL→l J_k^L→ l be a top-k source-subspace shared map with fitted output G^k(l)=G(L)J^kL→l. G_k^(l)=G^(L) J_k^L→ l. Suppose that the relative transfer-factor fit error is bounded by some ϵmax _max: ϵτ,kL→l≤ϵmax<1. _τ,k^L→ l≤ _max<1. Then τ^kL→l1+ϵmaxsr(G(L))≤sr(G(l))≤τ^kL→l1−ϵmaxsr(G(L)). τ_k^L→ l1+ _maxsr\! (G^(L) ) \! (G^(l) )≤ τ_k^L→ l1- _maxsr\! (G^(L) ). When ϵmax _max is obtained by measuring the fit error on the same N examples used to estimate J^kL→l J_k^L→ l, Lemma 3.6 is an in-sample fit certificate rather than an independent prediction of sr(G(l))sr(G^(l)). It becomes a predictive propagation bound only when an independently established error bound ϵmax _ is assumed to continue holding for additional source gradient matrices. Moreover, agreement at the level of stable rank does not imply matrix-level agreement; we evaluate the latter separately using relative Frobenius error, relative spectral error, and row-wise cosine similarity. 4 Empirical validation 4.1 Experimental setup Target models. We perform measurements against allenai/OLMo-2-0425-1B-Instruct, an instruction-tuned OLMo 2 1B model with 16 transformer blocks, 16 attention heads, and hidden dimension d=2048d=2048. For the training-stage analysis in Section 4.2, we evaluate its additional checkpoints: allenai/OLMo-2-0425-1B (olmo_base), allenai/OLMo-2-0425-1B-SFT (olmo_sft), allenai/OLMo-2-0425-1B-DPO (olmo_dpo), allenai/OLMo-2-0425-1B-RLVR1 (olmo_rlvr), and allenai/OLMo-2-0425-1B-Instruct (olmo_inst) [1]. Target datasets. We use AdvBench (adv), WildJailbreak evaluation and training splits (wld, wld_tr), and XSTest-Response subsets (xst_hr, xst_hr_jb, xst_rf, xst_rf_jb) as target prompt sets. Our frozen-model ablation analyses on these datasets use layer 77 for refusal vector estimation. For controlled fine-tuning, we use CAMEL chemistry dataset camel-ai/chemistry [21] by pairing its prompts with refusals sampled from a predefined set of refusals we manually collect; we use layer 88 to estimate refusal vector. Metrics. We use refusal (harm) rate, or a proportion of refused (or harmful) prompts to all prompts as downstream metrics, using allenai/wildguard as a judge LLM for refusal (harm) scores; refusal delta (change under attack) is a difference between pre-intervention and post-intervention refusal rates. To test the affine-subspace prediction in Equation 17, we compare three quantities for a benign-centered residual matrix ΔH H paired with gradient-induced activation updates G (or, similarly, paired with raw gradients G): (1) cosine similarity between their ΔH H and G mean vectors; (2) stable rank on the uncentered ΔH H and G; (3) average principal-angle overlap between their centered top-k right-singular subspaces. We report stable ranks and transfer-factor summary statistics across conditions. Shared-map fits are evaluated using source-subspace energy, relative Frobenius and spectral errors, row-wise cosine similarity, and relative transfer-factor fit error, with identity and scalar-identity baselines. Full definitions appear in Appendix F. 4.2 Refusal direction across released post-training checkpoints Figure 2: Descriptive comparison of released OLMo checkpoints along Base → SFT → DPO → RLVR1 → Instruct. The base checkpoint is not attack-effective under this estimator. A functional refusal direction is present by SFT, while post-SFT checkpoints exhibit larger refusal deltas than SFT. We first ask when the difference-in-means vector becomes a functional refusal direction. Using 60 AdvBench prompts, we estimate refusal vectors ρ at each layer for the OLMo checkpoints and measure refusal rates before/after ablation, refusal vector norm, cosine similarity to the final olmo_inst direction, and stable rank of ΔH H. Figure 2 shows that the base checkpoint has little refusal behavior and a negligible refusal delta under this estimator, whereas all four post-base checkpoints refuse nearly all selected prompts before intervention. A functional refusal-mediating direction is present at the SFT checkpoint, but its refusal delta is substantially smaller than those of the DPO, RLVR1, and Instruct checkpoints. Because the checkpoints differ in their objectives, data, and optimization histories, this comparison does not isolate the causal effect of any individual post-training stage. Figure 13 shows that the refusal vectors estimated at SFT and the later checkpoints remain highly aligned with the final olmo_inst direction across layers, while their stable rank and vector norm profiles largely overlap. Thus, the released post-SFT checkpoints approximately preserve the refusal geometry even though SFT is less attack-effective in our settings. Finally, Figure 12 shows that, for olmo_inst, directions estimated at layers 7–10 produce the largest refusal deltas. 4.3 Gradient-induced activation updates match benign-centered refusal residuals We next test the affine-subspace prediction in Equation 17: benign-centered refusal residuals ΔH H should share their mean direction, principal subspace, and stable rank with gradient-induced activation updates G, up to the sign induced by gradient descent. Using the target datasets introduced above, we compute ΔH H and G for three target-completion types: (1) model-generated refusal completions, (2) post-ablation completions, and (3) fixed pseudorandom English-token completions. For each dataset and output type, we compare ΔH H and G using mean-direction cosine similarity, average principal-angle overlap between centered right-singular subspaces, and stable rank. We make similar comparisons for ΔH H and raw gradients G as well. Figure 3: Gradient-induced activation updates explain benign-centered refusal residuals. Left column: mean alignment and average principal-angle overlap. Right column: stable rank comparison. Top row: layer 77; bottom row: layer 1515. The left column of Figure 3 shows that refusal-based means are negatively aligned with the difference-in-means refusal vector, while ablated-jailbreak-based and random-English-based means are positively aligned with it. Thus, refusal and non-refusal targets move along approximately the same refusal axis but in opposite directions, supporting Equation 18. This effect appears primarily for the refusal directions estimated at middle layer 77 and partially from the last layer 1515. The same column also shows average principal-angle overlap between (ΔHc)V( H_c) and (c)V(G_c). At layer 1515, refusal-target updates subspace overlap refusal subspace more than both random-English and ablated-jailbreak controls gradient update subspaces for every evaluated dataset, indicating stronger last-layer agreement between direction subspaces of refusal gradient updates (c)V(G_c) and refusal subspace (ΔHc)V( H_c). At layer 77, overall subspace overlap is lower, and the overlap appears to be more strongly associated with the dataset than with the completion type: the three target types within a dataset are relatively close, while the evaluation split of WildJailbreak (wld) has lower overlap than its training split (wld_tr) and most other datasets. The right column of Figure 3 shows that, for refusal targets, the stable ranks of G and ΔH H are close across datasets for both layer 77 and layer 1515. The observed stable ranks are below the sample size N=60N=60, around 11–33, indicating genuine spectral concentration and excluding trivial rank cap from insufficient datapoints. We observe the same qualitative relationship across most transformer layers. Together, these three diagnostics support the main empirical implication of Equation 17: signed mean-direction alignment, partial and layer-dependent principal-subspace overlap, and closely matched stable ranks. They support an approximate affine-subspace relationship. As a complementary uncentered diagnostic, at the refusal-estimating layer 77, the leading right-singular axis of G has mean absolute cosine similarity 0.560.56 with the refusal direction and 0.960.96 with the mean of G across the seven evaluated datasets. Thus, the dominant uncentered gradient-update axis of G is nearly collinear, up to sign, with the mean gradient-update vector of G and moderately aligned with the refusal direction. Together with the low stable rank of G, this makes the mean-direction and leading-axis descriptions of its update geometry nearly equivalent in this setting and motivates a direct comparison between difference-in-means refusal ablation and ablation along the top largest principal directions. 4.4 Refusal first-token diversity increases gradient stable ranks We next test on a frozen model whether varying refusal first tokens on a fixed prompt set changes the stable ranks of first-token gradients from the last layer to refusal-mediating middle layers. Using the WildJailbreak training split, we fix 80 harmful prompts whose model-generated refusals begin with the common first token “I”. For each m∈1,2,4,6,8,10,12,14,16m∈\1,2,4,6,8,10,12,14,16\, we pair each prompt with a refusal sampled independently with replacement from a pool of m short phrases with distinct first tokens. For each setting, we compute ErE_r, P, and P−ErP-E_r, the last-layer activation gradient matrix (P−Er)W⊤(P-E_r)W , and the layerwise gradient matrices G(l)G^(l) for the first-token cross-entropy loss. We use raw gradients in this experiment to test the cross-layer rank-transfer analysis from Section 3; gradient-induced activation updates are considered in the next subsection. Figure 4: Stable ranks and transfer under refusal-start diversification. Left: stable ranks of the target-space quantities and the resulting last-layer gradient. Middle: stable ranks of the first-token gradients at selected middle-to-late layers. Right: across-setting layerwise stable rank transfer factor mean±std± std and min/max. The stable ranks of ErE_r, P−ErP-E_r, and G(l)G^(l) increase overall. The G(15)→G(7)G^(15)\!→ G^(7) transfer factor varies substantially less than the end-to-end Er→G(7)E_r\!→ G^(7) ratio. Figure 4 shows that all target-space quantities and gradients across layers 7-15 increase with the number of refusals: gradient stable rank increase is bounded and quickly saturates for the last layer due to constant P, but amplified throughout mid-layers on a backward pass. Across the nine settings, the observed layerwise transfer factor mean ± sample standard deviation are τ15→7=2.270±0.175τ^15→ 7=2.270± 0.175, with range [1.925,2.469][1.925,2.469]. Thus, every evaluated setting exhibits approximately twofold stable-rank amplification from layer 15 to layer 7. By contrast, the end-to-end Er→G(7)E_r\!→ G^(7) ratio is more variable, with mean 0.942±0.4860.942± 0.486 and range [0.504,1.942][0.504,1.942]. For the m=16m=16 setting, we additionally fit the top-k source-subspace shared map from G(15)G^(15) to G(7)G^(7) on the same 80 examples and sweep k. At k=20k=20, the selected source subspace captures 99.8%99.8\% of the source-gradient energy; the relative Frobenius and spectral errors are 0.5940.594 and 0.3160.316, the mean row-wise cosine similarity is 0.8060.806, and the relative transfer-factor fit error is 0.3430.343. The fitted map yields lower reconstruction and transfer-factor errors and higher row-wise cosine similarity than the identity and best scalar-identity baselines. This provides evidence that the observed in-sample cross-layer relation admits a meaningful but incomplete shared linear approximation. The complete rank sweep and reduced-map spectrum are reported in Figures 14 and 15. 4.5 Refusal first-token diversity increases residual rank and weakens single-vector ablation We next test on a frozen model whether increasing the diversity of first tokens across refusal phrases (through varying selection of prompts in the dataset) increases the effective rank of benign-centered refusal residuals and gradient-induced activations updates. Using the WildJailbreak training split, we select harmful prompts that the target model refuses, keeping the baseline refusal rate fixed across subsets. We construct nine 80-example subsets drawing from m∈1,2,4,6,8,10,12,14,16m∈\1,2,4,6,8,10,12,14,16\ refusal first-token buckets. For each subset, we compute benign-centered refusal residuals ΔH H, raw first-token activation gradients G(l)G^(l), gradient-induced activation updates (l)G^(l), and refusal and harmfulness scores before and after difference-in-means ablation. Both G(l)G^(l) and (l)G^(l) are computed using the original WildJailbreak response paired with each selected prompt. We then plot the ablation score deltas against the stable rank of ΔH H. We select layer 77 as the earliest layer in the range of effective layers. Figure 5: Stable rank of benign-centered refusal residuals ΔH H. Left plot averaged ranks over layers 7-15. Increasing refusal first-token diversity raises sr(ΔH)sr( H) in middle-to-late layers. Figure 5 shows that increasing the number of refusal first-token buckets raises ΔH H stable rank. The effect is visible across middle-to-late layers, while the residual representation at earlier layers retain their stable ranks. Figure 17 shows similar trend for gradient-induced activation updates, but with saturation at the end. As the refusal first tokens become more diverse, the stable rank of G also increases across most layers, however it saturates at 12-16 refusal starts. Figure 6: Stable rank propagation under refusal-start diversification by prompt selection. Left: stable ranks of the first-token target matrix ErE_r, prediction matrix P, logit-gradient matrix P−ErP-E_r, and last-layer activation-gradient matrix (P−Er)W⊤(P-E_r)W . Middle: stable ranks of the first-token gradients G(l)G^(l). Right: stable ranks of the corresponding gradient-induced activation updates (l)G^(l). The target-space and layerwise stable ranks increase overall as refusal-start support broadens. Figure 6 shows the same overall rank increase from the dataset-target quantities through the raw gradients G(l)G^(l) and the gradient-induced activation updates (l)G^(l). Notably, the dataset-derived ErE_r rank rises sharply across the prompt subsets, whereas the model’s actual output-refusal-start ErE_r rank remains much lower. Thus, the dataset responses contain substantially more first-token diversity than the frozen model realizes in its own completions. Figure 7: Variation of stable rank transfer factor for raw gradients and gradient-induced activation updates across experimental settings. Left: raw gradient transfer factor. Right: gradient-induced activation deltas transfer factor. Raw gradients consistently amplify stable rank between layers 15 and 7, whereas the gradient update matrices undergo mild stable-rank contraction on average. Across the nine settings, raw gradient transfer factor is τG15→7=2.161±0.418 _G^15→ 7=2.161± 0.418, reported as mean ± sample standard deviation, with range [1.671,2.993][1.671,2.993]. Every evaluated setting therefore exhibits stable rank amplification from G(15)G^(15) to G(7)G^(7), by approximately a factor of two on average, although the amplification magnitude varies across subsets. For gradient-induced activation updates, the corresponding factor is τ15→7=0.829±0.132 _G^15→ 7=0.829± 0.132, with range [0.698,1.088][0.698,1.088] indicating mild stable-rank contraction on average. The layer-1515-to-layer-77 update transfer is more stable across subsets than the full target-to-middle-layer ratios. The complete in-sample shared-map rank sweeps and reduced-map spectra for G and G are reported in Figures 18, 19, 20 and 21. We then test whether higher stable rank of ΔH H corresponds to weaker single-vector ablation. Figure 8 reports harmfulness scores before and after ablation, and plots the harmfulness delta against sr(ΔH)sr( H). Subsets with higher-rank ΔH H have smaller harmfulness deltas, suggesting that a less one-dimensional residual subspace is harder to suppress with one vector. Figure 9 shows, similarly, that the refusal delta slightly decreases as sr(ΔH)sr( H) increases, indicating that higher-rank residuals ΔH H weaken the effect of single-vector refusal ablation. Figure 8: Harmfulness score decomposition and harmfulness delta versus sr(ΔH)sr( H). Higher refusal residuals stable rank is associated with smaller harmfulness deltas under difference-in-means ablation. Figure 9: Refusal score decomposition and refusal delta versus sr(ΔH)sr( H). Higher activation stable rank is associated with smaller refusal deltas, consistent with weaker single-vector ablation. Overall, this fixed-model subset analysis supports the rank mechanism: increasing refusal first-token diversity raises the effective rank of both ΔH H and G, while higher activation stable rank is associated with smaller refusal and harmfulness deltas under single-vector ablation. 4.6 Diverse-refusal fine-tuning raises refusal residuals rank and weakens ablation We next test whether the rank–vulnerability pattern from Section 4.5 also appears when simulating refusal training on the prompts are initially benign and not refused. The primary prediction is that increasing refusal first-token diversity should raise refusal residuals stable rank and reduce the refusal delta after difference-in-means single vector ablation. Attacking by ablating more principal directions is possible only if this does not degrade performance and the directions generalize across the prompts (we leave multiple-vector ablation attacks out of scope). We keep the prompt set fixed and vary only the generated refusal-style completions by sampling 40 new-topic prompts from camel-ai/chemistry and pairing them with refusal completions sampled from m∈1,4,8,12,16m∈\1,4,8,12,16\ distinct first-token buckets with balanced counts within each dataset, which directly manipulates the stable rank of the refusal first-token matrix ErE_r. We fine-tune OLMo-2-0425-1B-Instruct separately on these datasets for each m and select each time the earliest checkpoint that reaches or exceeds refusal rate threshold 0.6, so that later measurements are done under close base refusal rates. For intervention, we select layer 88 as the earliest layer in the range of effective layers. Additionally, for m=8m=8 and m=16m=16, we track refusal rates across optimization steps. Figure 22 shows that the ablation effect is largest in a narrow optimization window. This suggests that, although the number of optimization steps for each m is different, fixing the resulting checkpoints at first-reached baseline rates of 0.6 compares largest attack affects. Figure 10: Stable rank of refusal residuals under diverse-refusal fine-tuning. Left plot averages ranks across layers 8-15. Stable ranks across layers and their per-layer averages increase with the number of refusal first-token buckets with the clearest separation in middle-to-late layers. At the matched checkpoints, we first measure the stable rank of the resulting residuals ΔH H. Figure 10 shows that stable rank of ΔH H increases with the number of refusal first-token buckets with the largest effect across mid-to-late layers. This supports the rank mechanism: more diverse refusal starts induce less concentrated activation changes. Figure 11: Refusal ablation at checkpoints matched by baseline refusal score. Left and middle plots: diverse refusals yield smaller deltas under difference-in-means ablation. Right plot: diverse refusal starts are associated with higher refusal residuals stable rank and smaller refusal deltas We then evaluate whether this rank increase weakens single-vector ablation. Figure 11 shows refusal scores before and after ablation at the matched checkpoints as well as stable rank of ΔH H plotted against refusal deltas. Figure 23 shows that, similar to the previous sections, with diverse fine-tuning, stable rank of target quantities as well as gradient-induced update deltas increase on average, undergoing mild stable rank expansion from layer 15 to layer 7. Finally, rightmost Figure 11 directly relates geometry to attackability. More diverse refusal starts are associated with higher refusal residuals stable rank and smaller refusal deltas under ablation. This suggests that increasing refusal first-token diversity weakens the single dominant refusal vector by spreading the fine-tuning update over more effective directions. Overall, this controlled fine-tuning experiment suggests the rank–vulnerability pattern can hold broadly. Even on non-harmful chemistry prompts, concentrating refusal first tokens produces lower-rank refusal residuals and stronger single-vector ablation, while diversifying refusal starts raises the refusal residuals stable ranks and reduces the ablation effect at matched baseline refusal score. 5 Conclusion As AI systems enter dual-use domains, robustness failures can be catastrophic and unevenly impact at-risk communities. Our work takes a step toward understanding AI vulnerabilities at the root cause and designing defenses from first principles. We asked why refusal behavior can concentrate in a low-dimensional activation-space mechanism that is shared across many harmful categories and can be effectively suppressed. In OLMo-2-0425-1B-Instruct, we find that the refusal directions reflect refusal safety-training: cross-entropy gradients from the first tokens of refusal completions induce activation updates whose mean and principal directions align with refusal direction and refusal subspace. This connects the empirical difference-in-means refusal vector to the training updates that create it. This mechanism suggests a simple source of brittleness exist: concentrated refusal first tokens induce low-stable-rank activation changes, making refusal behavior easier to suppress with a single-vector refusal ablation. Across fixed-model and controlled fine-tuning experiments, increasing refusal first-token diversity raises stable rank of the activation-change matrix and is associated with smaller refusal deltas. At the same time, refusal-completion diversity alone is unlikely a complete defense. Overall, our results suggest that refusal robustness to ablation is partly shaped by the geometry of refusal training updates. Token-level properties of the training completions, especially refusal-prefix concentration and first-token frequency, can shape the dimensionality of the learned refusal mechanism. Understanding this link between data, gradients, and activation geometry may be a step toward principled mechanistic understanding defenses against jailbreaks. Finally, our controlled fine-tuning experiments on non-harmful datasets motivate study whether similar mechanisms hold for a broader set of features than for refusal or safety-related features. References [1] Allen Institute for AI (2025) OLMo-2-0425-1B-Instruct model card. Note: https://huggingface.co/allenai/OLMo-2-0425-1B-InstructHugging Face model card. Accessed: 2026-05-07 Cited by: §4.1. [2] M. Anderljung, J. Barnhart, A. Korinek, J. Leung, C. O’Keefe, J. Whittlestone, S. Avin, M. Brundage, J. Bullock, D. Cass-Beggs, B. Chang, T. Collins, T. Fist, G. Hadfield, A. Hayes, L. Ho, S. Hooker, E. Horvitz, N. Kolt, J. Schuett, Y. Shavit, D. Siddarth, R. Trager, and K. Wolf (2023) Frontier AI regulation: managing emerging risks to public safety. External Links: 2307.03718, Link Cited by: §1, §1. [3] Anthropic (2026) Assessing Claude Mythos Preview’s cybersecurity capabilities. Note: Accessed: 2026-05-05 External Links: Link Cited by: §1. [4] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda (2024) Refusal in language models is mediated by a single direction. External Links: 2406.11717, Link Cited by: §A.1, §1, §2.1, §2.2. [5] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. (2022) Constitutional AI: harmlessness from AI feedback. External Links: 2212.08073, Link Cited by: §1. [6] R. Bayat, A. Rahimi-Kalahroudi, M. Pezeshki, S. Chandar, and P. Vincent (2025) Steering large language model activations in sparse spaces. External Links: 2503.00177, Link Cited by: §A.1. [7] D. Beaglehole, A. Radhakrishnan, E. Boix-Adserà, and M. Belkin (2025) Toward universal steering and monitoring of ai models. External Links: 2502.03708, Link Cited by: §A.1, §1. [8] Y. Bengio, S. Clare, C. Prunkl, et al. (2026) International AI safety report 2026. External Links: 2602.21012, Link Cited by: §1, §1. [9] M. D. Buhl, B. Bucknall, and T. Masterson (2025) Emerging practices in frontier AI safety frameworks. External Links: 2503.04746, Link Cited by: §1. [10] N. Carlini, D. Paleka, K. D. Dvijotham, T. Steinke, J. Hayase, A. F. Cooper, K. Lee, M. Jagielski, M. Nasr, A. Conmy, I. Yona, E. Wallace, D. Rolnick, and F. Tramèr (2024) Stealing part of a production language model. External Links: 2403.06634, Link Cited by: §1. [11] T. Cha, D. Beaglehole, A. Radhakrishnan, and D. Lee (2026) The weight gram matrix captures sequential feature linearization in deep networks. External Links: 2605.06258, Link Cited by: §3. [12] P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong (2023) Jailbreaking black box large language models in twenty queries. External Links: 2310.08419, Link Cited by: §1. [13] S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski, and R. Liu (2020) Plug and play language models: a simple approach to controlled text generation. External Links: 1912.02164, Link Cited by: §A.1, §1. [14] D. Davis and D. Drusvyatskiy (2026) When do spectral gradient updates help in deep learning?. External Links: 2512.04299, Link Cited by: §A.2, §E.1, §1. [15] Y. Dong, J. Cordonnier, and A. Loukas (2021) Attention is not all you need: pure attention loses rank doubly exponentially with depth. External Links: 2103.03404, Link Cited by: §A.2, §1. [16] J. Dunefsky and A. Cohan (2025) One-shot optimized steering vectors mediate safety-relevant behaviors in llms. External Links: 2502.18862, Link Cited by: §A.1, §1. [17] F. M. Forum (2025) Frontier mitigations. Technical report Frontier Model Forum. Note: Accessed: 2026-05-05 External Links: Link Cited by: §1. [18] A. Gopal, N. Helm-Burger, L. Justen, E. H. Soice, T. Tzeng, G. Jeyapragasan, S. Grimm, B. Mueller, and K. M. Esvelt (2023) Will releasing the weights of future large language models grant widespread access to pandemic agents?. External Links: 2310.18233, Link Cited by: §1, §1. [19] Z. He, H. Zhao, Y. Qiao, F. Yang, A. Payani, J. Ma, and M. Du (2025) SAIF: a sparse autoencoder framework for interpreting and steering instruction following of language models. External Links: 2502.11356, Link Cited by: §A.1. [20] R. A. Horn and C. R. Johnson (2012) Matrix analysis. 2nd edition, Cambridge University Press, Cambridge. External Links: ISBN 978-0521548236 Cited by: §E.2. [21] G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for “mind” exploration of large scale language model society. External Links: 2303.17760, Link Cited by: §4.1. [22] A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi (2023) Tree of attacks: jailbreaking black-box LLMs automatically. External Links: 2312.02119, Link Cited by: §1. [23] J. Minder, C. Dumas, S. Slocum, H. Casademunt, C. Holmes, R. West, and N. Nanda (2025) Narrow finetuning leaves clearly readable traces in activation differences. External Links: 2510.13900, Link Cited by: §A.2. [24] Model Evaluation and Threat Research (2025) Common elements of frontier AI safety policies. Note: Accessed: 2026-05-05 External Links: Link Cited by: §1. [25] S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, and J. Alstott (2024) Securing AI model weights: preventing theft and misuse of frontier models. Technical report Technical Report R-A2849-1, RAND Corporation. External Links: Document, Link Cited by: §1. [26] OpenAI (2025) Preparing for future AI capabilities in biology. Note: Accessed: 2026-05-05 External Links: Link Cited by: §1. [27] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. External Links: 2203.02155, Link Cited by: §1. [28] K. O’Brien, D. Majercak, X. Fernandes, R. Edgar, B. Bullwinkel, J. Chen, H. Nori, D. Carignan, E. Horvitz, and F. Poursabzi-Sangdeh (2025) Steering language model refusal with sparse autoencoders. External Links: 2411.11296, Link Cited by: §A.1. [29] W. Pan, Z. Liu, Q. Chen, X. Zhou, H. Yu, and X. Jia (2025) The hidden dimensions of llm alignment: a multi-dimensional analysis of orthogonal safety directions. External Links: 2502.09674, Link Cited by: §A.2, §1. [30] N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner (2024) Steering llama 2 via contrastive activation addition. External Links: 2312.06681, Link Cited by: §A.1, §1. [31] V. Papyan, X. Han, and D. L. Donoho (2020) Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences 117 (40), p. 24652–24663. Cited by: §A.2, §1. [32] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving (2022) Red teaming language models with language models. External Links: 2202.03286, Link Cited by: §1. [33] N. Prakash, Y. W. Jie, A. Abdullah, R. Satapathy, E. Cambria, and R. K. W. Lee (2025) Beyond i’m sorry, i can’t: dissecting large language model refusal. External Links: 2509.09708, Link Cited by: §1. [34] X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson (2024) Safety alignment should be made more than just a few tokens deep. External Links: 2406.05946, Link Cited by: §A.2, §1, §1. [35] A. Rangamani, M. Lindegaard, T. Galanti, and T. A. Poggio (2023) Feature learning in deep classifiers through intermediate neural collapse. In International conference on machine learning, p. 28729–28745. Cited by: §A.2, §1. [36] L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. Bloom, S. Biderman, A. Garriga-Alonso, A. Conmy, N. Nanda, J. Rumbelow, M. Wattenberg, N. Schoots, J. Miller, E. J. Michaud, S. Casper, M. Tegmark, W. Saunders, D. Bau, E. Todd, A. Geiger, M. Geva, J. Hoogland, D. Murfet, and T. McGrath (2025) Open problems in mechanistic interpretability. External Links: 2501.16496, Document, Link Cited by: §A.2. [37] A. Stolfo, V. Balachandran, S. Yousefi, E. Horvitz, and B. Nushi (2025) Improving instruction-following in language models through activation steering. External Links: 2410.12877, Link Cited by: §A.1. [38] N. Subramani, N. Suresh, and M. E. Peters (2022) Extracting latent steering vectors from pretrained language models. External Links: 2205.05124, Link Cited by: §A.1, §1. [39] E. Todd, M. L. Li, A. S. Sharma, A. Mueller, B. C. Wallace, and D. Bau (2024) Function vectors in large language models. External Links: 2310.15213, Link Cited by: §A.1. [40] A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, and M. MacDiarmid (2024) Steering language models with activation engineering. External Links: 2308.10248, Link Cited by: §A.1, §A.1, §1. [41] A. Wei, N. Haghtalab, and J. Steinhardt (2023) Jailbroken: how does LLM safety training fail?. External Links: 2307.02483, Link Cited by: §1. [42] T. Wollschläger, J. Elstner, S. Geisler, V. Cohen-Addad, S. Günnemann, and J. Gasteiger (2025) The geometry of refusal in large language models: concept cones and representational independence. External Links: 2502.17420, Link Cited by: §A.2, §1. [43] L. Yu, V. Do, K. Hambardzumyan, and N. Cancedda (2025) Robust llm safeguarding via refusal feature adversarial training. External Links: 2409.20089, Link Cited by: §A.1. [44] A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson (2023) Universal and transferable adversarial attacks on aligned language models. External Links: 2307.15043, Link Cited by: §1. Appendix A Related work A.1 Activation steering and ablation attacks Activation steering and ablation attacks are a family of whitebox methods which directly act on model activations and runtime to alter its output in a desired way. Those methods are widely used to suppress refusals, bypass safety alignment, and to jailbreak open models. An early work [40] introduces activation engineering, an inference-time steering of model activation by adding a vector, computed from contrasting datasets. [4] introduced a difference-of-means jailbreak attack: the method works by estimating refusal vector as a difference between harmful and harmless prompts’ activation means and later removing projections on this vector to let model respond to prompts it was trained to refuse. [43] additionally confirms that the operation of refusal feature ablation approximates the worst-case perturbation of offsetting model safety. [16] find a steering vector by performing gradient descent from a single training example to mediate safety-relevant behavior. [7] offers a universal steering by adding eigenvectors extracted from Recursive Feature Machines (RFMs) trained on contrasting datasets. A line of work introduces sparse autoencoder (SAE) based steering. O’Brien et al. identify and amplify SAE features mediating refusal [28]; Bayat et al. introduce sparse activation steering by selecting SAE features from contrastive prompts [6]; SAIF steers instruction-following behavior through instruction-relevant SAE latents [19]. Activation steering has been useful more broadly in non-security contexts, too: ActAdd and CAA use prompt-pair or contrastive activation differences to steer behaviors, such as topic, sentiment, factuality, or behavior [40], [30]; PPLM steers generation by using gradients from an attribute model to act on activations [13]; Subramani et al. extract steering vectors from frozen models with gradient descent [38]; Function Vectors and instruction-steering work extract task or instruction vectors for inference-time control [39], [37]. Our work complements steering attack approaches by explaining a notable safety-critical feature in activations, the refusal vector, and its low-dimensional structure using analytical methods. A.2 Mechanistic interpretability, representation geometry, and safety-training dynamics An adjacent body of work studies the representations learned by models and the training dynamics that produce them. [42] show that refusal representations span multiple orthogonal directions and can be described as a concept cone, motivating representational independence: interventions on different directions can mediate different behaviors without necessarily affecting others. Similarly, [29] argue that refusal is multi-dimensional: a dominant direction mediates refusal, while additional components correspond to distinct interpretable safety features. [36] survey open problems and research areas in mechanistic interpretability, including predicting which capabilities and mechanisms arise during training or fine-tuning. Complementing these empirical and conceptual accounts of multi-dimensional refusal, our work analytically studies the origin of such orthogonal refusal directions and links them to training-time refusal gradients. We further provide evidence that concentrated refusal datasets can produce lower-rank gradient-induced activation updates, yielding low-stable-rank activation changes and nearly one-dimensional refusal representations; this structure, in turn, predicts larger effectiveness of the ablation attack. A broader body of work studies concentrations as well as rank and dimensionality collapse. [23] demonstrate that fine-tuning on a narrow domain leaves identifiable traces of the training objective in model activations, suggesting that semantically narrow fine-tuning can induce readable activation-space biases. [34] find that the largest safety fine-tuning updates fall on the first assistant tokens, and that constraining updates to be distributed more evenly across output tokens improves robustness to fine-tuning-based jailbreaks. [14] identify conditions under which spectral optimization is preferred to Euclidean gradient descent and show that transformer activations can remain low-stable-rank during training. [15] prove that pure self-attention can converge toward rank-one token-uniform representations without skip connections or MLPs. [31] describe terminal-training neural collapse, in which final-layer class features collapse to class means with highly symmetric geometry, while [35] extend this picture to intermediate layers, showing that deeper representations increasingly reduce within-class variance relative to between-class variance and align class-mean subspaces with dominant weight directions. Unlike these broader investigations of rank collapse and neural collapse, our work focuses on safety post-training: we study a mechanism in which concentration of refusal first tokens makes the target matrix low stable rank, inducing low-stable-rank activations changes whose principal directions approximate the refusal subspace. Appendix B Limitations and assumptions of analytical derivations Throughout the derivations, we make simplifying assumptions to obtain a tractable model of the relationship between refusal targets, gradients, and activation changes. Supervised refusal-training scope. Our derivation models supervised fine-tuning on harmful prompts paired with textual refusals. We use the terms refusal training and supervised safety tuning for this setting. DPO, RLHF, and RLVR use different objectives and are not directly covered by the derivation which may induce slightly The controlled fine-tuning experiment uses effective batch size one to isolate per-example refusal updates. It is intended as an approximation to the sporadic contribution of comparatively rare harmful–refusal examples within SFT dataset mixture for our target model. Larger batches containing unrelated instruction-following examples may introduce additional signal into gradients and the resulting activation changes. First-token-only loss. We compute all losses only at the first assistant target-token position, after chat formatting, with the prompt positions masked. We therefore omit gradients from the remaining completion tokens. In the evaluated settings, the first-token loss captures substantial refusal-direction structure, but later target positions may introduce additional directions and training effects. Deterministic per-example updates. We compute isolated per-example gradient steps with deterministic backward passes and without stochastic regularization such as dropout. These measurements do not reproduce minibatch interactions, optimizer momentum, adaptive preconditioning, weight decay, or the accumulation of many training steps. Prompt-dependent Jacobians and conditioning. Exact transformer Jacobians are prompt-dependent, including through input-dependent attention patterns. Consequently, Theorem 3.2 is an idealized sufficient condition rather than a claim that complete per-example Jacobians are shared. Products of worst-case condition-number bounds may also become numerically vacuous for anisotropic maps. The reduced-map condition numbers reported in our experiments apply only to selected data-dependent source coordinates and are not estimates of the full prompt-specific Jacobian condition numbers. In-sample shared-map fit and predictive scope. When ϵmax _ is obtained by measuring the transfer-factor fit error on the same N examples used to estimate J^kL→l J_k^L→ l, Lemma 3.6 gives an in-sample fit-conditioned interval for sr(G(l))sr(G^(l)), rather than an independent prediction. Specifically, it certifies that, on the fitted support, the map-implied transfer factor τ^kL→l τ_k^L→ l agrees with the observed transfer factor τL→lτ^L→ l to relative error at most ϵmax _ , and therefore that the observed target stable rank lies within the corresponding interval. The result becomes predictive only when the same fitted map is held fixed and applied without refitting to additional source gradient matrices, with its error bound established independently on held-out source–target matrix pairs. Moreover, agreement at the level of stable rank does not imply matrix-level agreement; we evaluate the latter separately using relative Frobenius error, relative spectral error, and row-wise cosine similarity. Appendix C Additional details about experimental setup. Refusal-mediating layers. Unless otherwise stated, the frozen-model single-layer rank–vulnerability and ablation analyses use layer 77 since our measurements show that refusal vector estimation at that layer provides most effective, and therefore attacker-preferred refusal ablation. Layerwise profiles and cross-layer transfer experiments additionally report all explicitly labeled layers. The controlled fine-tuning rank–vulnerability experiment in Section 4.6 uses layer 88 – the earliest layer in the range of effective layers in our experiment, which retains a long last-to-middle-layer path over which to test rank propagation, similar to frozen-model experiments; selecting a later layer would make that propagation test less stringent. Estimating gradient-induced activation updates. To compute gradient-induced activation updates G, for each training example i and cross-entropy loss ℓiCE _i^CE on the first token of the target completion, we apply one small parameter update θnew=θ−η∇θℓiCE,η=10−6. _new=θ-η _θ _i^CE, η=10^-6. At layer l, we measure i(l):=−(hi(l)(θnew)−hi(l)(θ)), g_i^(l):=- (h_i^(l)( _new)-h_i^(l)(θ) ), and then restore the parameters to θ. Thus, i(l) g_i^(l) is the negative realized activation change from one small gradient step, consistent with our sign convention. For the evaluated models, we use l=7l=7 or l=8l=8, the earliest layers at which the estimated refusal directions yield substantial attack success; this retains a long last-to-middle-layer path over which to test rank propagation. All empirical G matrices are computed from the realized finite differences in Equation 15. Under a local first-order expansion, their corresponding differential is ηJi(l)∇θℓiCE,η J_i^(l) _θ _i^CE, which can instead be evaluated directly as a Jacobian–vector product without applying a parameter perturbation. We leave a systematic comparison of the difference and differential estimators to future work. Appendix D Broader limitations Model, data, language, and checkpoint scope. This work is primarily a case study of one small OLMo model family, relatively small prompt sets, and English-language data. We do not identify an attackable difference-in-means refusal direction in the base checkpoint under our estimator, but this does not rule out other functional base-model refusal mechanisms or representations. The released Base, SFT, DPO, RLVR1, and Instruct checkpoints differ in their objectives, data, optimization histories, and numbers of updates. Their comparison is therefore descriptive and does not identify which training stage caused the observed refusal geometry or change in ablation sensitivity. Target-set-adapted white-box threat model. For each evaluated prompt set S, the attacker estimates one difference-in-means direction from all prompts in S and uses that direction for the refusal ablation attack on S. The results therefore intentionally test an attacker-favorable same-set ablation setting and do not establish that a direction estimated from one prompt sample transfers to unseen prompts. Held-out universality of refusal directions is a distinct question left for future work. Attack-family scope. Our primary attack is difference-in-means single-vector ablation. SAE-based steering, multi-feature ablation, and broader automated representation-level attacks remain outside our evaluation. We approach the difference-in-means single-vector attack using analytical methods and invite application of those methods to study a broader set of attacks and defenses in mechanistic interpretability. Stable rank versus causal mediation and intrinsic dimension. Stable rank measures global linear spectral concentration. It does not by itself measure refusal-signal magnitude, causal mediation strength, or local nonlinear manifold dimension. A layer may therefore have relatively high stable rank while its refusal direction has little behavioral effect, whereas another layer may have lower stable rank but mediate refusal strongly. Our rank–vulnerability comparisons concern fixed or matched refusal-mediating layers and do not use stable rank to predict which layer is most causally important. Scope beyond refusal. We study refusal because it provides a shared refuse/comply mechanism with a well-defined white-box ablation attack. It remains unclear whether similarly universal non-refusal concepts have the same geometry, or whether increasing feature rank generally makes universal steering or other forms of jailbreak optimization harder. Appendix E Theorems E.1 Stable rank of the refusal first-token matrix Definition E.1 (One-hot matrix ErE_r for a batch of refusals). Let Er∈ℝN×|V|E_r ^N×|V| be the row-stacked one-hot matrix with i-th row eri⊤e_r_i , for a batch of targets r=(r1,…,rN)r=(r_1,…,r_N). Let ck:=#i:ri=kc_k:=\#\i:r_i=k\ be the count of token k, let fk:=ck/Nf_k:=c_k/N, and let m:=|k:ck>0|m:=|\k:c_k>0\|. Lemma E.2 (Spectrum of a stacked one-hot matrix). The feature-space Gram matrix of ErE_r is diagonal: Er⊤Er=diag(c1,…,c|V|). E_r E_r=diag(c_1,…,c_|V|). (33) Consequently, the nonzero singular values of ErE_r are ck c_k for tokens with ck>0c_k>0, and rank(Er)=m. (E_r)=m. (34) The stable rank is sr(Er)=‖Er‖F2‖Er‖22=∑k=1|V|ckmaxkck=Nmaxkck=1maxkfk. (E_r)= \|E_r\|_F^2\|E_r\|_2^2= _k=1^|V|c_k _kc_k= N _kc_k= 1 _kf_k. (35) Proof. The proof uses an approach similar to [14]. For any k,k′∈Vk,k ∈ V, (Er⊤Er)k,k′=∑i=1N(Er)i,k(Er)i,k′=∑i=1N[ri=k][ri=k′]. (E_r E_r)_k,k = _i=1^N(E_r)_i,k(E_r)_i,k = _i=1^N1[r_i=k]1[r_i=k ]. (36) If k≠k′k≠ k , no example can satisfy both ri=kr_i=k and ri=k′r_i=k , so the sum is 00. If k=k′k=k , the sum is ckc_k. Thus Er⊤Er=diag(c1,…,c|V|)E_r E_r=diag(c_1,…,c_|V|). The eigenvalues of Er⊤ErE_r E_r are ckc_k, so the singular values of ErE_r are ck c_k. The rank is the number of positive counts. Finally, ‖Er‖F2=∑kck=N\|E_r\|_F^2= _kc_k=N, while ‖Er‖22=maxkck\|E_r\|_2^2= _kc_k, giving the stable-rank formula. ∎ Corollary E.3 (Minimizer and maximizer of sr(Er)sr(E_r)). Let cmax:=maxkckc_ := _kc_k. Then 1≤sr(Er)=Ncmax≤m. 1 (E_r)= Nc_ ≤ m. (37) Moreover: 1. sr(Er)=1sr(E_r)=1 iff cmax=Nc_ =N, i.e., all targets share the same first token. 2. For fixed support size m, sr(Er)sr(E_r) is maximized when counts are as balanced as possible across the m observed tokens, i.e. ck∈⌊N/m⌋,⌈N/m⌉c_k∈\ N/m , N/m \ on the support. In that case, maxsr(Er)=N⌈N/m⌉. (E_r)= N N/m . (38) If m|Nm N, this maximum equals m. 3. If the support is unconstrained and N≤|V|N≤|V|, the global maximum is sr(Er)=Nsr(E_r)=N, achieved when all first tokens are distinct. If N>|V|N>|V|, the maximum is achieved by balancing counts over all |V||V| vocabulary tokens. Proof. By Lemma E.2, sr(Er)=N/cmaxsr(E_r)=N/c_ . The lower bound follows from cmax≤Nc_ ≤ N, with equality iff cmax=Nc_ =N. For fixed support size m, the pigeonhole principle gives cmax≥⌈N/m⌉c_ ≥ N/m , hence sr(Er)=Ncmax≤N⌈N/m⌉≤m. (E_r)= Nc_ ≤ N N/m ≤ m. (39) This is tight when counts are balanced across the m observed targets. The unconstrained-support statement follows by taking m=minN,|V|m= \N,|V|\. ∎ E.2 Stable rank propagation theorems The following proof uses standard matrix norm inequalities: the Frobenius/Schatten-22 ideal property, operator norm submultiplicativity, and the definition of the spectral condition number; see, e.g., [20, Chapter 5]. Theorem E.4 (Stable rank under multiplication by a well-conditioned factor). Let A∈ℝN×nA ^N× n, let B∈ℝn×nB ^n× n be invertible, and define condition number: κ(B):=‖B‖2‖B−1‖2. κ(B):=\|B\|_2\|B^-1\|_2. (40) Then 1κ(B)2sr(A)≤sr(AB)≤κ(B)2sr(A). 1κ(B)^2sr(A) (AB)≤κ(B)^2sr(A). (41) Proof. Recall that sr(M):=‖M‖F2‖M‖22. (M):= \|M\|_F^2\|M\|_2^2. (42) We use ‖AB‖F \|AB\|_F ≤‖A‖F‖B‖2, ≤\|A\|_F\|B\|_2, (43) ‖AB‖2 \|AB\|_2 ≤‖A‖2‖B‖2. ≤\|A\|_2\|B\|_2. (44) Since A=(AB)B−1A=(AB)B^-1, ‖A‖2=‖(AB)B−1‖2≤‖AB‖2‖B−1‖2, \|A\|_2=\|(AB)B^-1\|_2≤\|AB\|_2\|B^-1\|_2, (45) so ‖AB‖2≥‖A‖2‖B−1‖2. \|AB\|_2≥ \|A\|_2\|B^-1\|_2. (46) Therefore, sr(AB)=‖AB‖F2‖AB‖22≤(‖A‖F‖B‖2)2(‖A‖2/‖B−1‖2)2=κ(B)2sr(A). (AB)= \|AB\|_F^2\|AB\|_2^2≤ (\|A\|_F\|B\|_2 )^2 (\|A\|_2/\|B^-1\|_2 )^2=κ(B)^2sr(A). (47) For the lower bound, apply the upper bound to A=(AB)B−1A=(AB)B^-1: sr(A)≤κ(B−1)2sr(AB)=κ(B)2sr(AB). (A)≤κ(B^-1)^2sr(AB)=κ(B)^2sr(AB). (48) Rearranging gives sr(AB)≥1κ(B)2sr(A). (AB)≥ 1κ(B)^2sr(A). (49) ∎ Lemma E.5 (Stable rank bounds from transfer factor fit error). Let G(L),G(l)∈ℝN×dG^(L),G^(l) ^N× d be nonzero gradient matrices, and let J^kL→l J_k^L→ l be a top-k source-subspace shared map with fitted output G^k(l)=G(L)J^kL→l. G_k^(l)=G^(L) J_k^L→ l. Suppose that the relative transfer-factor fit error is bounded by some ϵmax _max: ϵτ,kL→l≤ϵmax<1. _τ,k^L→ l≤ _max<1. Then τ^kL→l1+ϵmaxsr(G(L))≤sr(G(l))≤τ^kL→l1−ϵmaxsr(G(L)). τ_k^L→ l1+ _maxsr\! (G^(L) ) \! (G^(l) )≤ τ_k^L→ l1- _maxsr\! (G^(L) ). Proof. By the definition of the relative transfer-factor fit error and assuming it is bounded by ϵmax _max, |τ^kL→l−τL→l|≤ϵmaxτL→l. | τ_k^L→ l-τ^L→ l |≤ _maxτ^L→ l. Hence, (1−ϵmax)τL→l≤τ^kL→l≤(1+ϵmax)τL→l. (1- _max )τ^L→ l≤ τ_k^L→ l≤ (1+ _max )τ^L→ l. Since ϵmax<1 _max<1, rearranging gives τ^kL→l1+ϵmax≤τL→l≤τ^kL→l1−ϵmax. τ_k^L→ l1+ _max≤τ^L→ l≤ τ_k^L→ l1- _max. Finally, using sr(G(l))=τL→lsr(G(L))sr (G^(l) )=τ^L→ lsr (G^(L) ) proves the result. ∎ Appendix F Matrix metrics This appendix defines the matrix metrics used in the empirical sections. All matrices below are row-stacked matrices in ℝn×dR^n× d, such as benign-centered refusal residuals ΔH H or gradient-induced activation update matrices (l)G^(l). Thin SVD and right-singular subspaces. For M∈ℝn×dM ^n× d, let M=UMΣMVM⊤M=U_M _MV_M denote its thin SVD, where VM∈ℝd×rMV_M ^d× r_M contains the right singular vectors associated with nonzero singular values and rM:=rank(M)r_M:=rank(M). We write (M):=span(VM)V(M):=span(V_M) Uncentered single-matrix metrics. For metrics of a single matrix, we use the uncentered matrix M. The empirical second-moment matrix is CM:=1nM⊤M,λM,i:=σM,i2n.C_M:= 1nM M, _M,i:= _M,i^2n. The stable rank is sr(M):=‖M‖F2‖M‖22=∑i=1rMσM,i2σM,12.sr(M):= \|M\|_F^2\|M\|_2^2= _i=1^r_M _M,i^2 _M,1^2. Centered pairwise comparisons between right singular subspaces. When comparing right singular subspaces between matrices, we use centered matrices, which means that for each matrix M, we remove its row mean μM _M to obtain its centered matrix McM_c before computing its right singular vectors: μM=1nM⊤n,Mc=M−nμM⊤. _M= 1nM 1_n, M_c=M-1_n _M . Unless otherwise stated, VM,kV_M,k and λM,i _M,i in pairwise metrics are computed from the centered matrix McM_c. This separates the mean direction, measured by cosine similarity, from the remaining PCA-style subspace structure. Top-k policy. For two matrices A,BA,B, let rA:=rank(Ac)r_A:=rank(A_c) and rB:=rank(Bc)r_B:=rank(B_c). Given a requested ktopk_top, we use k:=minktop,rA,rB.k:= \k_top,r_A,r_B\. In the main experiments we set ktop=100k_top=100. Since n<100n<100 for our experiments, this includes all nonzero centered principal directions after the rank cap. Cosine similarity of row means. For A,B∈ℝn×dA,B ^n× d, define cosine_mean(A,B):=μA⊤μB‖μA‖2‖μB‖2∈[−1,1].cosine\_mean(A,B):= _A _B\| _A\|_2\| _B\|_2∈[-1,1]. This metric compares only the mean directions and is complementary to centered subspace-overlap metrics. Explained-variance overlap. Let Ac=UAΣAVA⊤A_c=U_A _AV_A and Bc=UBΣBVB⊤B_c=U_B _BV_B be centered thin SVDs. Let VA,kV_A,k and VB,kV_B,k denote the first k right singular vectors, and define CA,B(k):=VA,k⊤VB,k.C^(k)_A,B:=V_A,k V_B,k. The explained-variance overlap of A’s top-k directions captured by B’s top-k subspace is EVk(A∣B):=∑i=1kλA,i∑ℓ=1kλA,ℓ∑j=1k(CA,B(k))ij2∈[0,1].EV_k(A B):= _i=1^k _A,i _ =1^k _A, _j=1^k (C^(k)_A,B )_ij^2∈[0,1]. Equivalently, each A-principal direction is weighted by its share of A’s top-k variance and scored by the squared norm of its projection onto span(VB,k)span(V_B,k). The metric is directional because it weights only by A’s spectrum. Average principal-angle overlap. Let Ac=UAΣAVA⊤,Bc=UBΣBVB⊤A_c=U_A _AV_A , B_c=U_B _BV_B be centered thin SVDs. Let VA,kV_A,k and VB,kV_B,k contain the first k right singular vectors, and define CA,B(k):=VA,k⊤VB,k.C_A,B^(k):=V_A,k V_B,k. Let si(CA,B(k))s_i(C_A,B^(k)) denote the singular values of CA,B(k)C_A,B^(k). We define the average principal-angle overlap as AvgOverlapk(A,B) _k(A,B) :=1k∑i=1ksi(CA,B(k))2 := 1k _i=1^ks_i\! (C_A,B^(k) )^2 (50) =1k‖VA,k⊤VB,k‖F2∈[0,1]. = 1k \|V_A,k V_B,k \|_F^2∈[0,1]. (51) The singular values of CA,B(k)C_A,B^(k) are the cosines of the principal angles between the two subspaces. Thus, AvgOverlapkAvgOverlap_k is symmetric and basis-invariant, equals 11 when the top-k subspaces coincide, and equals 00 when they are orthogonal. Shared-map metrics. All shared-map metrics below are computed on uncentered gradient matrices. Let G(L),G(l)∈ℝN×dG^(L),G^(l) ^N× d be the source and target gradient matrices, and let G^k(l)=G(L)J^kL→l G_k^(l)=G^(L) J_k^L→ l be the output of the fitted top-k source-subspace shared map. Source-subspace energy capture. Let Vk(L)∈ℝd×kV_k^(L) ^d× k contain the top-k right singular vectors of G(L)G^(L). We define the fraction of source-gradient energy contained in the selected subspace as ECk(L):=‖G(L)Vk(L)‖F2‖G(L)‖F2. _k^(L):= \|G^(L)V_k^(L) \|_F^2 \|G^(L) \|_F^2. (52) Equivalently, if σ1(G(L))≥σ2(G(L))≥⋯ _1(G^(L))≥ _2(G^(L))≥·s are the singular values of G(L)G^(L), then ECk(L)=∑j=1kσj(G(L))2∑jσj(G(L))2. _k^(L)= _j=1^k _j\! (G^(L) )^2 _j _j\! (G^(L) )^2. (53) This metric measures how much of the total squared source-gradient energy is retained before fitting the shared map. Relative Frobenius reconstruction error. We define the relative Frobenius error between the fitted and observed target matrices as RelFkL→l:=‖G(l)−G^k(l)‖F‖G(l)‖F. _k^L→ l:= \|G^(l)- G_k^(l) \|_F \|G^(l) \|_F. (54) This metric measures the total reconstruction error relative to the Frobenius norm of the observed target matrix. A value of zero denotes exact reconstruction; the metric is not upper bounded by one. Relative spectral reconstruction error. We define the relative spectral error as Rel2kL→l:=‖G(l)−G^k(l)‖2‖G(l)‖2. 2_k^L→ l:= \|G^(l)- G_k^(l) \|_2 \|G^(l) \|_2. (55) This metric measures the largest residual singular direction relative to the dominant singular direction of the observed target matrix. It complements the relative Frobenius error because stable rank depends on both the Frobenius and spectral norms. F.1 Refusal direction across released post-training checkpoints - additional plots Figure 12: Refusal ablation attack effectiveness per each layer used to estimate the refusal difference-in-means direction (for a fixed model OLMo-2-0425-1B-Instruct against a subset of AdvBench prompts). The plot shows that fixing layers to 7-10 provides the largest ablation attack success (refusal delta) with decreasing success across later layers, while earlier layers provide almost no benefit, resulting in relatively ineffective attacks. Figure 13: Measuring refusal vectors across stages shows that once a functional refusal vector appears at SFT stage, all resulting refusal directions across all stages (except the base model) are all similar for a subset of selected AdvBench prompts. Refusal vector norms grow after SFT stages and stay similar throughout post-training stages. Appendix G Refusal first-token diversity increases gradient stable ranks - additional plots (a) Source-energy capture. (b) Row-wise cosine similarity. (c) Relative spectral error. (d) Relative Frobenius error. (e) Transfer-factor fit error. Figure 14: In-sample fit quality of the top-k source-subspace shared map from G(15)G^(15) to G(7)G^(7) for the m=16m=16 refusal-start condition. Increasing k captures more source-gradient energy, improves per-example directional agreement, and reduces matrix-reconstruction and stable-rank-transfer errors. At k=20k=20, the map captures 99.8%99.8\% of source energy and attains relative Frobenius, spectral, and transfer-factor errors 0.5940.594, 0.3160.316, and 0.3430.343. The k=80k=80 endpoint is the in-sample interpolation limit and is shown only as a reference, not as evidence of a low-rank mechanism. (a) Largest singular value. (b) Smallest nonzero singular value. (c) Reduced-map condition number. Figure 15: Singular-value diagnostics for the fitted map on its selected source coordinates. The smallest nonzero singular value changes little between k=10k=10 and k=20k=20, from 0.8780.878 to 0.8640.864, while the largest singular value grows from 2.242.24 to 13.3313.33. Consequently, the reduced-map condition number rises from 2.552.55 to 15.4315.43. This explains why the higher-rank map gives a better empirical fit while its worst-case condition-number bound becomes less informative. Appendix H Refusal first-token diversity increases residual rank and weakens single-vector ablation - additional plots Figure 16: Averaged ranks over layers 7-15 stable rank of benign-centered refusal residuals ΔH H. Increasing refusal first-token diversity raises sr(ΔH)sr( H) in middle-to-late layers. Figure 17: Stable rank of gradient-induced update deltas G: more refusal first-token buckets produce higher effective rank across most layers, although the increase saturates at the end. Left upper plot averages stable ranks across layers 7-15. H.1 Frozen-model raw-gradient shared-map diagnostics (a) Source-energy capture. (b) Row-wise cosine similarity. (c) Relative spectral error. (d) Relative Frobenius error. (e) Transfer-factor fit error. Figure 18: In-sample fit quality of the top-k source-subspace shared map from G(15)G^(15) to G(7)G^(7) for the m=16m=16 prompt-subset condition. Increasing k captures more source-gradient energy, improves per-example directional agreement, and reduces matrix-reconstruction and stable-rank-transfer errors. At k=20k=20, the selected source subspace captures 99.8%99.8\% of source-gradient energy. The relative Frobenius and spectral errors are 0.6160.616 and 0.5040.504, the mean row-wise cosine similarity is 0.7390.739, and the relative transfer-factor fit error is 0.3370.337. Thus, the rank-20 source subspace contains nearly all source-gradient energy, while the remaining reconstruction errors show that the shared map captures a meaningful but incomplete component of the observed cross-layer relation. The k=80k=80 endpoint is the in-sample interpolation limit and is shown only as a reference, not as evidence of a low-rank mechanism. (a) Largest singular value. (b) Smallest nonzero singular val. (c) Reduced-map condition num. Figure 19: Singular-value diagnostics for the reduced coordinate map fitted to the raw gradients. Between k=10k=10 and k=20k=20, the smallest nonzero singular value decreases slightly from 0.830.83 to 0.790.79, the largest singular value grows from 2.652.65 to 11.1911.19, the reduced-map condition number rises from 3.193.19 to 14.1714.17. The higher-rank map gives a better in-sample fit but becomes substantially more anisotropic, making its worst-case condition-number bound less informative. H.2 Frozen-model gradient-update shared-map diagnostics (a) Source-energy capture. (b) Row-wise cosine similarity. (c) Relative spectral error. (d) Relative Frobenius error. (e) Transfer-factor fit error. Figure 20: In-sample fit quality of the top-k source-subspace shared map from (15)G^(15) to (7)G^(7) for the m=16m=16 prompt-subset condition. Increasing k captures more source-update energy, improves per-example directional agreement, and reduces matrix-reconstruction and stable-rank-transfer errors. At k=20k=20, the selected source subspace captures 75.2%75.2\% of source-update energy. The relative Frobenius and spectral errors are 0.6390.639 and 0.4050.405, the mean row-wise cosine similarity is 0.6910.691, and the relative transfer-factor fit error is 0.3400.340. The more gradual energy-capture curve indicates that substantial source-update energy lies beyond the first 20 principal directions. The k=80k=80 endpoint is the in-sample interpolation limit and is shown only as a reference, not as evidence of a low-rank mechanism. (a) Largest singular value. (b) Smallest nonzero singular value. (c) Reduced-map condition number. Figure 21: Singular-value diagnostics for the reduced coordinate map fitted to the gradient-induced activation updates. Between k=10k=10 and k=20k=20, the smallest nonzero singular value decreases from 0.03470.0347 to 0.03170.0317, while the largest singular value grows from 0.1290.129 to 0.1710.171. Consequently, the reduced-map condition number rises from 3.713.71 to 5.385.38. The rank-20 map improves the in-sample fit with substantially less conditioning deterioration than the corresponding raw-gradient map. Appendix I Diverse-refusal fine-tuning raises refusal residuals rank and weakens ablation - additional plots 8 refusal starts 16 refusal starts (a) Refusal rates, 8 starts (b) Refusal rates, 16 starts (c) Absolute refusal delta, 8 starts (d) Absolute refusal delta, 16 starts (e) Normalized refusal delta, 8 starts (f) Normalized refusal delta, 16 starts Figure 22: Optimization-step-matched refusal ablation trajectories for models trained with 8 and 16 refusal starts. Left and right columns show the 8- and 16-start conditions, respectively. The top row reports baseline and post-ablation refusal rates; the middle row reports the absolute refusal deltas; and the bottom row reports the fraction of baseline refusals removed by ablation. Ablation vulnerability exists in a small window of training since the refusal delta is peaked at first when maximum refusal rate is reached, decaying afterwards. The 8-start plot shows deltas having a single peak with maximum of 0.55, while 16-start plot shows two smaller peaks and reaches overall maximum of 0.375, consistent with reduced ablation attack effectiveness. Figure 23: Stable ranks and stable rank propagation under refusal-start diversification by fine-tuning for gradient and gradient-induced activation deltas. Top left: stable ranks of the dataset-target ErE_r, realized-output ErE_r, prediction matrix P, target residual P−ErP-E_r, and resulting last-layer gradient (P−Er)W⊤(P-E_r)W . Top middle: stable ranks of the gradients G(l)G^(l). Top right: stable ranks of the gradient-induced activation updates (l)G^(l). Bottom rows: end-to-end stable rank transfer factors measured for raw gradients and gradient-induced activation deltas transfer factors. The target-space, gradients, and gradient-updates stable ranks increase overall as refusal-start support broadens, and the both gradient and gradient-update matrices undergo mild stable-rank expansion from layer 15 to layer 8 on average.