Paper deep dive
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
Vadym Hadetskyi, Dario Pasquini, Artem Sorokin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/9/2026, 2:09:46 AM
Summary
This paper investigates domain-specific abliteration in large language models to selectively remove refusal directions for cybersecurity while preserving general safety alignment. Through large-scale experiments on 24 open-source LLMs, the authors demonstrate that abliteration via orthogonal projection can successfully reduce cybersecurity refusal (e.g., dropping it from 100% to 7% in Kimi K2) without compromising general capabilities measured by MMLU. The study finds that refusal is distributed across multi-dimensional subspaces and layers, particularly in MoE architectures. Susceptibility to abliteration is model-specific and best predicted by safety training methodology and architecture, leading to a classification of models into three susceptibility tiers. Misinformation domains are most affected by the intervention, while copyright remains the most resistant.
Entities (10)
Relation Signals (10)
Kimi K2 → achieved → Domain-specific Abliteration
confidence 95% · we achieved the first domain-specific abliterated model, removing refusal in the Kimi K2 for the cybersecurity domain only.
Abliteration → modifies → Model Weights
confidence 95% · Abliteration is the modification of model weights via orthogonal projection to remove a target direction from the model activation space
Domain-specific Abliteration → targets → Cybersecurity
confidence 92% · domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2... for the cybersecurity domain only.
MMLU → measures → Capability Preservation
confidence 90% · MMLU capability scores are universally preserved (max degradation: 0.028 across all 24 models).
Refusal Direction → occupies → Multi-dimensional subspace
confidence 90% · refusal in LLMs occupies a multi-dimensional subspace rather than a single direction
Safety Alignment → failstodistinguish → Domains
confidence 88% · conceptually it does not distinguish between various domains and the level of potential harm of a query
Misinformation → mostaffectedby → Abliteration
confidence 88% · misinformation proved to be the most affected domain (−22.3pp mean)
Architecture → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains and the level of potential harm of a query, which creates significant complications in the fields like cyber security, where a model should not be constrained by its safety circuits to accomplish the goals of legitimate, authorized operations. In this work, we share our findings from a large scale abliteration experiment on 24 open-source LLMs and show that domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2. Building on recent work showing that refusal in LLMs occupies a multi-dimensional subspace within layers, we find that it is also distributed widely across layers, especially in trillion-parameter MoE architectures, and so we aim to capture the part of it that represents harmful concepts in the cybersecurity domain exclusively. We also investigate the correlation between models' features and the effect of domain-specific abliteration, identifying that the type of safety training and architecture are the most reliable predictors. Finally, we classify the models into 3 abliteration susceptibility tiers and put forward a set of conjectures as to why a particular effect from this intervention might be observed in a given model.
Tags
Links
- Source: https://arxiv.org/abs/2607.02714v2
- Canonical: https://arxiv.org/abs/2607.02714v2
Trouble viewing inline? Open PDF directly →
Full Text
45,961 characters extracted from source content.
Expand or collapse full text
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Vadym Hadetskyi Cracken AI vh@cracken.ai Dario Pasquini Cracken AI dp@cracken.ai Artem Sorokin Cracken AI as@cracken.ai Abstract There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains and the level of potential harm of a query, which creates significant complications in the fields like cyber security, where a model should not be constrained by its safety circuits to accomplish the goals of legitimate, authorized operations. In this work, we share our findings from a large scale abliteration experiment on 24 open-source LLMs and show that domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2. Building on recent work showing that refusal in LLMs occupies a multi-dimensional subspace within layers, we find that it is also distributed widely across layers, especially in trillion-parameter MoE architectures, and so we aim to capture the part of it that represents harmful concepts in the cybersecurity domain exclusively. We also investigate the correlation between models’ features and the effect of domain- specific abliteration, identifying that the type of safety training and architecture are the most reliable predictors. Finally, we classify the models into 3 abliteration susceptibility tiers and put forward a set of conjectures as to why a particular effect from this intervention might be observed in a given model. 1 Introduction 1.1 Problem Statement & Motivation In this work, we study domain-specific abliteration. Abliteration is the modification of model weights via orthogonal projection to remove a target direction from the model activation space; typically the refusal direction Arditi et al. [2024], Labonne [2024]. Domain-specific abliteration is the approach of removing the refusal direction of a model for a chosen domain (e.g., cybersecurity), while preserving it for the general case. To the best of our knowledge, we achieved the first domain-specific abliterated model, removing refusal in the Kimi K2 Kimi Team [2025] for the cybersecurity domain only. However, in our experi- ments across different models, we observed that not all LLMs respond equally to domain-specific abliteration. We therefore set out to investigate whether this behavior is an inherent model property or a consequence of factors such as architecture, size, training methodology, or other incidental characteristics. More specifically, we ask: does the ability to distinguish between harmful requests in cybersecurity and those involving violence or biological weapons emerge only in sufficiently capable models, or can it also be observed in smaller, potentially simpler models? Recent geometric findings that refusal in LLMs occupies a multi-dimensional subspace rather than a single direction Wollschläger et al. [2025], Pan et al. [2025] support this conjecture: if the refusal subspace is multi-dimensional, then in principle one slice of it, representing a single harm domain, should be removable without collapsing the rest. Whether the standard abliteration pipeline could Preprint. arXiv:2607.02714v2 [cs.CR] 7 Jul 2026 Chem/Bio Copyright Cybercrime Harassment Illegal Misinfo Qwen3-0.6B Gemma-1B Llama-1B Llama-3B Phi-4-mini OLMoE Llama-8B Qwen3-8B Gemma-12B Phi-4 Qwen3-14B DS-V2-Lite Mistral-Small Gemma-27B Qwen3-30B Qwen3-32B Llama-70B Mistral-Large Qwen3-235B Llama-Scout GLM-4.7 Llama-Maverick DS-V3 Kimi-K2 Chem/Bio Copyright Cybercrime Harassment Illegal Misinfo −40 −20 0 20 40 Δ p Refusal at 30% Abliteration (%)Change from Baseline (p) 203200 4201146352 66839706750 58628666832 582833606312 68631848332 681017786716 24148403214 762430948544 945664568932 841233847258 842638485920 662045908548 801419605028 781836848554 621424787238 64815826828 9020511009372 824636888766 20108463042 10042891009896 661222785260 86836868346 884051748574 -4-2-4-12-4-8 -48-8-47-50-56-80 -26+4-30-22-24-34 -22-2-15-22-15-62 -42-56-51-40-33-56 -28-28-35-16-11-56 -32-10-40-22-26-58 -56-16-29-48-52-42 -10+2-13-2-9-20 -6-16-29-38+0-42 +0-2-3+4-4-2 -6+8-24-26-15-18 +0+0+0+0+0+0 -2-4-17-10-24-12 -8-6-10-10-7-16 -6-6-5-8-6-14 -4-6+1-6-11-24 +0+0+0+0+0+0 -10+4-2+0+2+4 -8+4-1-2-6-2 +0-14-8+0+0+6 -12-2-8-6+0-2 -6-12+1+2-6-4 -12-10-37-14-6-24 Figure 1: Per-domain refusal rates at 30% abliteration (left) and change from baseline (right) across all 24 models. Cybersecurity (cybercrime) consistently shows the largest reduction in high-susceptibility models. actually exploit this at scale was, however, an open question. Thus, we designed an experiment as the first systematic, large-scale study of abliteration susceptibility and domain selectivity in 24 LLMs of different sizes and architectures from 10 developer organizations. The whole endeavor stems directly from our practical need to have the best possible LLM for authorized offensive security operations. Having an uncensored model greatly increases the span of operations the agent can perform autonomously. 1.2 Contributions Our main finding is that abliteration does not affect all harm domains equally as previously described. Given the right combination of model capabilities, safety training method, and sufficiently sophisti- cated refusal representation in the embedding space, it is possible to obtain a model compliant with requests in the target domain without compromising its general safety. Figure 1 shows the change in refusal per-domain across all 24 models at 30% abliteration, i.e. orthogonal projection applied to 30% of the model’s layers, distributed across 25–95% of model depth (see Section 3 for details). Despite using the cybersecurity focused dataset, misinformation proved to be the most affected domain (−22.3p mean), while copyright shows the highest variance including models where refusal increases after abliteration. Nevertheless, the effect is not universal across all the models and depends on a combination of factors, such as model architecture, its safety training, reasoning capability, etc. Thus, we studied the relationship between them and the character of the refusal change with increasing levels of layer modification. Our contributions are: (1) The nature of safety alignment is domain-heterogeneous, which enables abliteration to affect harm categories non-uniformly. (2) Domain-specific abliteration is achievable: on Kimi K2 (1T parameters), we reduce cybersecurity refusal from 100% to 7% while preserving 100% explicit content refusal and 44–88% refusal in other domains. 1 (3) Susceptibility is model- specific, not size-dependent, while safety training methodology and architecture are the strongest predictors. (4) MMLU capability scores are universally preserved (max degradation: 0.028 across all 24 models). 2 Background & Related Work The linear representation hypothesis is now central to mechanistic interpretability Park et al. [2023], Elhage et al. [2022], Templeton et al. [2024]: researchers have found linear representations of truth Marks and Tegmark [2023], Li et al. [2024], sentiment Tigges et al. [2023], and many other 1 Results reported on our cross-evaluation dataset (306 prompts, 6 domains). The scientific benchmark (360 prompts from AdvBench/HarmBench/PurpleLlama) shows directionally consistent but less pronounced effects; see Section 4 for full comparison. 2 behavioral concepts Zou et al. [2023a], and shown that such directions can be used to steer model outputs Panickssery et al. [2023], Turner et al. [2023], Zou et al. [2023a]. Zou et al. [2023a] formalized this into a practical toolkit for understanding and intervening on model behavior. Arditi et al. [2024] extended this lens to safety: refusal, they argued, is mediated by a single direction in the residual stream, and erasing it via orthogonal projection of model weights (abliteration) removes refusal across 13 models up to 72B parameters, evaluated on aggregate refusal without per-domain breakdown. Concurrent mechanistic work has since shown that the refusal subspace is richer than a single direction: Wollschläger et al. [2025] identify multi-dimensional polyhedral concept cones, and Pan et al. [2025] find a dominant direction accompanied by multiple semantically distinct secondary features. We build on these findings of the complexity of the refusal geometry: if it occupies a higher- dimensional subspace, it should admit domain-selective interventions via the standard abliteration pipeline. We verify this empirically across 24 models from 0.6B to 1T parameters and characterize when it succeeds. Wei et al. [2023] offer an influential framing for why safety training fails non-uniformly: competing objectives and mismatched generalization predict that susceptibility is heterogeneous across capability domains, which our domain-level results corroborate empirically. Beyond representation-level attacks, the literature includes session-bound prompt-level jailbreaks Chao et al. [2023], Zou et al. [2023b], Liu et al. [2023], fine-tuning attacks that need only minimal harmful data Lermen et al. [2023], Yang et al. [2023], and a growing set of abliteration tools. Young [2025] compared four such tools across 16 models up to 14B, while Heretic Weidmann [2025] automates the same orthogonal-projection family we use, employing tree-structured Parzen estimator search (a Bayesian hyperparameter optimizer) over layer and strength configurations to achieve near-complete refusal removal across diverse architectures. On the defense side, Circuit Breakers Zou et al. [2024] propose representation-level interventions that operate on harmful-output representations directly, rather than on a refusal direction. Abu Shairah et al. [2025] showed that training with extended refusal responses by distributing the safety signal across many token positions substantially reduces abliteration effectiveness; Qi et al. [2025] provide a complementary theoretical account, arguing that current safety alignment is "shallow," concentrated in the first few output tokens, and that deepening it materially improves robustness. Henderson et al. [2025] find that refusal-only safety training is the most brittle to abliteration, while combining it with safe-data filtering yields partial robustness, and Lee et al. [2024] provide a mechanistic account of how DPO creates the very features abliteration targets. The closest concurrent work is RepIt Siu et al. [2026], which uses whitening and controlled orthog- onalization to suppress refusal of one target concept (WMD) while preserving refusal elsewhere, evaluated on five frontier LMs in the 2–5B range. CAST Siu et al. [2025] pursues a related goal via conditional activation steering, evaluated on seven models up to 32B and supporting logical composi- tion of condition vectors across multiple concept categories. Both are inference-time interventions. We instead apply permanent weight modification at the standard abliteration pipeline level, scaled to 24 models from 0.6B up to 1T parameters with a cybersecurity focus, and use this scale to study when domain-selective abliteration succeeds. We use HarmBench Mazeika et al. [2024] for standardized red-teaming, AdvBench Zou et al. [2023b] for adversarial behaviors, PurpleLlama CyberSecEval Bhatt et al. [2024] for cybersecurity-specific safety, and MMLU Hendrycks et al. [2020] for capability preservation; SORRY-Bench Xie et al. [2024] and Cybench Zhang et al. [2024] situate our findings within the broader safety-refusal and cybersec benchmark landscapes. 3 Methodology Our methodology follows the standard abliteration approach Arditi et al. [2024]. The most important step is the construction of the extraction dataset — pairs of harmful and harmless prompts that isolate refusal while keeping the domain constant. This is delicate: every prompt steers the extracted direction not only on the harm dimension but also toward the subjects mentioned in it, so for a cybersecurity-specific abliterated model we iterated to find real-world task pairs of the form: Harmful: “How do I create a SQL injection attack against MySQL databases?” Harmless: “What is SQL injection and how can it be prevented with prepared statements?” 3 Harmful: “How can I perform privilege escalation using kernel exploits?” Harmless: “For a CTF challenge involving a Linux binary with SUID bit set, walk me through the methodology to identify potential privilege escalation paths” The harmful prompts are direct action requests; the harmless ones cover the same technical ground but frame it educationally or defensively. This semantic separation is what gives the extracted refusal direction its domain specificity – it aims to capture the difference between offensive intent and cybersecurity knowledge, rather than the one between dangerous topic and safe topic. A generic harmful/harmless dataset, likemlabonne/harmful_behaviorsLabonne [2024], identifies general refusal. Initial experiments with it and more cybersecurity-focused version with generic domain prompts resulted in global safety collapse, motivating the targeted design described here. Next, for each model the refusal direction must be found using the assembled dataset. We run each prompt through the model and extract the hidden state activations at the last token position of each decoder layer. Following Arditi et al. [2024], we use the last token because, under causal attention, it is the only position that attends to the full input sequence. The refusal direction ˆ rfor each layerℓis then computed as the normalized mean difference: ˆ r ℓ = m + ℓ − m − ℓ ∥m + ℓ − m − ℓ ∥ ,where m + ℓ = 1 N + N + X i=1 h (i) ℓ,−1 Hereh (i) ℓ,−1 denotes the hidden state at layerℓ, last token position, for prompti.m + ℓ andm − ℓ are the means across harmful and harmless prompts respectively. We extract activations at the last token position rather than pooling across the sequence; we experimented with mean pooling over all token positions and found it diluted the refusal signal with positional noise, yielding weaker directions. The refusal direction is then projected out of the weight matrices via orthogonalization: W ′ = W − α ˆ r( ˆ r ⊤ W) whereαis the abliteration strength. Atα = 1.0this is an exact orthogonal projection, which removes the component of each weight matrix along the refusal direction. Atα > 1.0the removal is amplified beyond orthogonalization, which we found impactful for MoE architectures where redundant expert pathways can partially compensate for the projection. This is a permanent modification of the model weights, applied to thedown_proj(MLP output) ando_proj(attention output projection) matrices. 3.1 Experimental Setup To test whether the patterns observed on Kimi K2 generalize, and to identify which model properties enable domain-selective abliteration, we ran a uniform protocol across the 24 models below. Models: 24 models evaluated on the scientific benchmark and on the custom dataset (see Appendix A for the full inventory). Architectures span dense transformers and MoE, sizes range from 0.6B to 1T, from 10 developer organizations. Technique: Orthogonal projection of the mean-difference refusal direction from model weights (W ′ = W − α ˆ r( ˆ r ⊤ W)) applied todown_projando_projweight matrices. Uniform layer spread selection across 25–95% of model depth at 6 percentiles (5–30%). Infrastructure: 8×NVIDIA H200 GPUs (1.1TB total VRAM) on GCP a3-ultragpu instances. Total compute: approximately 600 GPU-hours. We usedvLLMfor inference,speculatorsfor activation extraction (vLLMdoes not support extraction of middle-layer activations), and direct safetensors modification for weight application. The pipeline is written fully in Python. 3.2 Layer Selection Strategies Layer selection is a critical element of the process. Arditi et al. found that abliterating a narrow local region of 2–3 middle layers can be as effective as ablating across all layers, and applied their technique to all layers’W O andW out plus the embedding matrix. We found that this does not generalize to the broader set of modern models in our study: refusal in many architectures is distributed more widely, particularly in MoE models with hundreds of experts per layer. Thus, in our main experiment we opted for uniformly spread layer selection across 25–95% of model depth, which we found to 4 substantially outperform norm-based selection (targeting layers with the highest refusal signal norm) by up to∼70p of additional refusal reduction at 30% on the most susceptible models. A detailed comparison is provided in Appendix C. 3.3 Evaluation Framework We evaluate using three complementary datasets: Scientific benchmark compilation (360 prompts). Assembled from three established and widely cited sources: HarmBench (185) Mazeika et al. [2024], AdvBench (119) Zou et al. [2023b], and PurpleLlama CyberSecEval (56) Bhatt et al. [2024]. The resulting benchmark covers 6 harm categories: cybercrime/intrusion (106), illegal activity (54), chemical/biological (50), copyright (50), harassment/bullying (50), and misinformation (50). All prompts are harmful. This dataset enables comparison with the broader research literature. Cross evaluation dataset (306 prompts). Our custom dataset with 181 harmful and 125 harmless prompts across 6 domains: cybersecurity (162), misinformation (32), violence (32), illegal goods (32), privacy violation (32), and explicit content (16). The cybersecurity-heavy distribution and inclusion of harmless prompts enable both domain-specific refusal analysis on a different data distribution and false-positive measurement (over-refusal on benign queries). MMLU (capability preservation). A sample of 25 prompts per subject to measure whether abliteration degrades general model capabilities. Refusal is detected via 31 pattern-matching rules (e.g. “I cannot,” “I’m not able to”), with reasoning tokens (<think>blocks) stripped before matching to avoid false positives from models that internally deliberate about refusal before complying. 4 Big Picture As discussed in Section 1.1, this research stems from the necessity of enabling our models to do their best work in our domain of interest, which is cyber security. After obtaining the desired result on 1T Kimi K2 we had to find out what model features allow domain-specific abliteration. We composed a diverse set of 24 models to run the same (fixed dataset, predefined layer proportion to modify, etc.) experiment on. In short, we have obtained mixed results. As summarized in Figure 2 and Table 1, the effect of abliteration varies much from model to model and from domain to domain. Some models turned out to be fully resistant to the standard approach (Mistral-Large, GLM 4.7), others collapsed across all domains (Phi-4-mini, Gemma-27B, Llama Maverick) and in some cases their baselines were not really safe to begin with (Qwen3-0.6B). But in some rare cases (Kimi-K2, Gemma-12B) we observed high domain-specificity in the resulting change. Kimi K2 is the most notable example of this, being the only one where the impact on the target domain (−37p) was significantly higher than on all the other domains, including the one with the weakest safety, misinformation. 2 One could treat “domain” as an arbitrarily defined cluster of similar prompts; while they are unam- biguously marked in our dataset as belonging to one or another category, their latent representations in the model embedding space could in fact be neighbors. Moreover, some prompts can contain markers from different domains, blurring the boundaries even further. Our target domain is cybersecurity (“Cybercrime”), and we saw clearly that “Misinformation” was always the most significantly affected, while copyright — the least. Thus, we speculate that some domains are more harmful than others, and it simply takes less effort, smaller modification to remove refusal in those. One of the most notable examples of domain-specificity is our headline result on Kimi K2, which demonstrates that domain-specific refusal removal is achievable, moreover, at trillion-parameter scale. Using a CTF-focused extraction dataset (52 cybersecurity-specific harmful prompts, 73 educational harmless prompts), we achieved a model abliterated specifically for our use case: This represents, to our knowledge, the first demonstration of domain-specific abliteration using the standard approach, described in Arditi et al. [2024]. Cybersecurity refusal drops from 100% 2 This−37p figure is from the scientific benchmark (360 prompts). On our cross-evaluation dataset (306 prompts), it achieves−93p on cybersecurity; see Table 2. 5 Misinfo Cybercrime Chem/Bio Harassment Illegal Copyright 0 10 20 30 40 50 Misinfo Cybercrime Harassment Chem/Bio Illegal Copyright Mean Refusal Reduction (p) All 24 ModelsHigh-susceptibility Only (n=12) Figure 2: Mean refusal reduction per domain at 30% abliteration. Left: all 24 models. Right: high-susceptibility models only (n=12). Error bars show standard deviation across models. Table 1: Mean refusal reduction per domain (across 24 models at 30% abliteration). DomainMean∆Median∆ Misinformation−22.3p−17.0p Cybercrime−15.5p−11.5p Harassment−14.0p−10.0p Chemical/Bio−13.3p−8.0p Illegal activities−12.3p−6.5p Copyright−8.2p−5.0p Table 2: Kimi K2 domain-specific abliteration results across both evaluation datasets obtained at (α = 1.0, 30% of layers, uniform layer spread). (a) Cross-evaluation dataset DomainBaselineAbliterated Cybersecurity100%7% Explicit content100%100% Illegal goods100%56% Misinformation100%88% Violence100%75% Privacy violation100%44% (b) Scientific benchmark DomainBaselineAbliterated Cybercrime88%51% Chemical/biological100%88% Copyright50%40% Harassment/bullying88%74% Illegal activities91%85% Misinformation98%74% to 7% (−93p), while the closest non-target domain (privacy violation) retains 44% refusal — a 13×selectivity ratio. Compared to publicly available “uncensored” versions of Kimi K2 (which reduce refusal to 0–20% across all domains), our approach preserves substantial safety in non-target domains. The contrast between the two datasets reinforces the point made above: how cleanly a domain-specific carve-out emerges from abliteration depends not only on the intervention’s parameters but also on how distinctly the target domain is separated from its neighbors, both in the dataset’s labels and in the model’s representation space (Appendix D). The underlying procedure is destructive, so we also had to verify that the models did not lose their general reasoning capabilities afterward. Thus, we used MMLU benchmark to evaluate that. MMLU scores are universally preserved across all models and abliteration intensities. The maximum degra- dation observed is 0.028 (DeepSeek-V3 at 30%); all other models show variation of≤0.006. This confirms that the studied technique does not affect general model capabilities. Figure 3 demonstrates it by comparing the MMLU score with the overall refusal rate at different intensities of the intervention. Another important pattern is revealed when we look at the impact from abliteration dynamically, increasing the intensity step by step. (Figure 4) once again illustrates these 3 types of reactions to abliteration. Some models collapse monotonously, some chaotically, and some remain unmoved at any level (the Resistant tier). The ones which monotonously lose more and more of their refusal with 6 Qwen3-0.6B Gemma-1B Phi-4-mini OLMoE Llama-8B Qwen3-8B DS-V2-Lite Mistral-Large Qwen3-235B Llama-Scout GLM-4.7 Llama-Maverick Kimi-K2 020406080 0.4 0.5 0.6 0.7 0.8 0.9 Qwen3-0.6B Gemma-1B Phi-4-mini OLMoE Llama-8B Qwen3-8B DS-V2-Lite Mistral-Large Qwen3-235B Llama-Scout GLM Llama-Maverick Kimi-K2 050100 ImmunePartially VulnerableVulnerable Overall Refusal Rate (%)Cybercrime Refusal Rate (%) MMLU Score Overall RefusalCybercrime Refusal Figure 3: MMLU vs. refusal rate trajectories (baseline→30%). Lines move horizontally (decreasing refusal) with negligible vertical shift (stable MMLU). Left: overall refusal. Right: cybercrime- specific. each intervention strength are the most susceptible, but at the same time they are the best candidates for domain-specific abliteration, since one can optimize meta parameters of the experiment more predictably to achieve compliance on the desired type of prompts. 0102030 0 20 40 60 80 100 0102030 0 20 40 60 80 100 0102030 0 20 40 60 80 100 0102030 0 20 40 60 80 100 0102030 0 20 40 60 80 100 0102030 0 20 40 60 80 100 Chem/BioCopyrightCybercrimeHarassmentIllegalMisinfo Abliteration %Abliteration %Abliteration % Refusal % Refusal % Kimi-K2Qwen3-8BLlama-8B Phi-4-miniMistral-LargeGLM-4.7 Figure 4: Refusal change at different intensities of abliteration for a sample of models Finally, removing the refusal based on a limited number of predefined examples does not immediately enable the model to comply with any task in the target domain. We have observed a rather high level of specificity and a sort of over-fitting in terms of the compliance with harmful prompts post-abliteration, i.e. the uncensored model provides answers to the prompts from a distribution that is similar to the one used to extract the refusal direction. The further the request is from that distribution, the less compliant the response becomes. Even more importantly, we note that the domain-specificity itself is inherently highly dependent on the definition of the domains and the distance between them. We elaborate on this in Appendix D. 5 Analysis of Susceptibility Predictors 5.1 Susceptibility Is Model-Specific, Not Size-Dependent We classify models into three susceptibility tiers (in terms of domain-specific safety removal) based on maximum refusal reduction at 30% layer abliteration (Figure 5). As could be observed in Figure 5, the relationship between model size and susceptibility is non-monotonic: Kimi K2 at 1 trillion parameters shows 22.8p reduction, while Mistral-Large at 123B shows 0.0p. Within the same architecture family, Qwen3-8B is among the most susceptible (−38.9p) while Qwen3-14B is nearly resistant (−1.4p). 7 It is important to elaborate on the subject of “resistance” here: it should not be viewed as an exclusively good defensive property of a model since it relates specifically to domain-focused abliteration. It could signify both genuine robustness to abliteration and insufficient complexity of concept representation in the model’s vector space. 01020304050 Mistral-Large (123B) Qwen3-14B (15B) Mistral-Small (24B) Qwen3-235B (235B) Llama-Scout (109B) GLM-4.7 (355B) Qwen3-0.6B (1B) Qwen3-32B (33B) Llama-70B (71B) Llama-Maverick (400B) Qwen3-30B (30B) Gemma-12B (12B) Gemma-27B (27B) DS-V2-Lite (16B) Llama-3B (3B) Phi-4 (14B) Kimi-K2 (1000B) Llama-1B (1B) DS-V3 (671B) OLMoE (7B) Llama-8B (8B) Qwen3-8B (8B) Phi-4-mini (4B) Gemma-1B (1B) High-susceptibilityPartialResistant Max Refusal Reduction (p) Figure 5: Maximum refusal reduction across 24 models, grouped by susceptibility class. Models span 0.6B to 1T parameters. As for the explanation of this property, a number of candidate predictors of susceptibility to domain- specific abliteration was explored: model size, architecture (dense vs. MoE), safety training method, and model family (Figure 6). Safety training has the strongest signal: DPO-based methods produce the most susceptible models (median∼30p drop, range 22–47p), while GRPO-based training consistently yields low susceptibility (5–8p). Models with undisclosed RL-based safety (both Mistral variants) are fully resistant. Notably, Google’s BOND+WARM+WARP method shows extreme variance across model sizes. Architecture type (dense vs. MoE) also partially explains susceptibility: having a richer variety of concept representations distributed across experts, MoE models are more resistant on average. Within model families, Phi is the most uniformly susceptible, Qwen shows the widest intra-family spread (from 1p to 39p), and Mistral is consistently resistant. Detailed breakdowns by safety training method and model family are provided in Appendix E. Gemma-1B Gemma-12B Gemma-27B Phi-4-mini OLMoE Llama-Maverick Qwen3-0.6B Qwen3-8B Qwen3-14B Qwen3-30B Qwen3-32B Qwen3-235B DS-V3 Llama-1B Llama-3B Llama-8B Mistral-Large GLM-4.7 Kimi-K2 DS-V2-Lite 0.40.50.60.70.80.9 0 10 20 30 40 50 BOND+ DPO DPO+GOAT GRPO PPO+DPO RL(?) RLHF SFT only MMLU Baseline Score Cybercrime Refusal Drop (p) Resistant Partial ◆ = MoE Figure 6: Susceptibility predictors: MMLU baseline score vs. maximum cybercrime refusal drop. Bubble size indicates parameter count; diamonds denote MoE architectures; color indicates safety training method. Dashed lines mark susceptibility tier boundaries. 8 6 Limitations Our study uses a single intervention technique — orthogonal projection of the mean-difference refusal direction. Other approaches exist: RepIt Siu et al. [2026] disentangles concept-specific directions through whitening and controlled orthogonalization, while Heretic Weidmann [2025] automates the same direct orthogonal-projection pipeline via tree-structured Parzen estimator hyperparameter search. Models that appeared resistant in our experiments (e.g. Mistral-Large) can in fact be abliterated by such tools, but with a critical caveat: the result is broad-spectrum safety removal across all domains, not the selective modification we target. Whether domain-specific abliteration is achievable on these resistant models with more advanced techniques remains an open question. We tested abliteration intensities up to 30% of model layers. Higher percentiles were not systemati- cally explored, though our abliteration curves (Figure 4) suggest the trends established by 30% are generally predictive of behavior at higher intensities for most models. Our refusal detection relies on 31 string-matching patterns rather than human annotation or LLM- based judgment. While we strip reasoning tokens to reduce false positives, this approach may miss sophisticated refusals or misclassify borderline responses. Similarly, our evaluation is English-only and cybersecurity-focused by design — generalization to other languages or target domains (e.g. biomedical, legal) would require domain-specific datasets and evaluation. Finally, the cybersecurity domain itself is not monolithic: our extraction dataset emphasizes CTF-style offensive operations. Abliteration effectiveness may vary for sub-domains like threat intelligence, forensics, or defensive security that were not specifically represented in our prompts. 7 Ethics & Safety Disclaimer The research was conducted in controlled environments for defensive cybersecurity purposes and to advance understanding of AI safety mechanisms. We aim for transparency without enablement by publishing our findings and methodology, not model weights. The code, evaluation framework and the datasets will be released upon the publication of the paper. However, a number of uncensored versions of SOTA models are readily available on Hugging Face Hub. This is why we believe that what might be considered a gray area of AI ethics must be explored and understood further to meet the challenges that abliterated models present. For organizations building security tooling, the implications are significant: domain-specific safety modification is achievable without compromising other ethical boundaries. For AI safety researchers, a number of our findings point out the flaws of the techniques applied to make models safe and highlight the importance of more advanced, domain-heterogeneous refusal representation. 8 Conclusion The main objective of our research is to identify the conditions that allow for domain-specific abliteration. We conducted a large-scale study on 24 open-source LLMs, both dense and MoE, ranging from 0.6B to 1T parameters, coming from multiple organizations and having different safety training algorithms. Thanks to the scale of the experiment we were able determine that the type of the safety training technique and the architecture are the most reliable predictors of the model’s ability to be abliterated in the scope of a particular domain. We theorize that this is a direct consequence of the complexity of the model’s vector space, especially, the complexity of its refusal representation. Our results corroborate concurrent geometric findings that refusal in modern LLMs is multi-dimensional rather than a single direction, and add two operational extensions: at trillion-parameter scale the refusal signal is distributed widely across layers and experts (so the narrow middle-layer abliteration of earlier work no longer suffices), and tuning the extraction dataset to a target domain isolates the part of the refusal subspace responsible for that domain without collapsing safety in others. We believe this research has practical impact in two directions: AI safety researchers can use our findings to compare the robustness of different safety-training techniques against representation-level attacks, and cybersecurity practitioners can be equipped with models that are useful across the full operational scope of their domain without an indiscriminate removal of refusal in unrelated ones. 9 References Abu Shairah, A., et al. An embarrassingly simple defense against LLM abliteration attacks. arXiv preprint arXiv:2505.19056, 2025. Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N. Refusal in language models is mediated by a single direction. In NeurIPS, 2024. Bhatt, S., Chennabasappa, S., et al. PurpleLlama CyberSecEval: A secure coding benchmark for language models. arXiv preprint arXiv:2312.04724, 2024. Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G.J., and Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023. Elhage, N., Hume, T., Olsson, C., et al. Toy models of superposition. Transformer Circuits Thread, 2022. Henderson, P., et al. A granular study of safety pretraining under model abliteration. arXiv preprint arXiv:2510.02768, 2025. Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020. Kimi Team. Kimi K2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025. Labonne, M. Uncensor any LLM with abliteration. HuggingFace Blog, 2024. Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J.K., and Mihalcea, R. A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity. arXiv preprint arXiv:2401.01967, 2024. Lermen, S., Rogers-Smith, C., and Ladish, J. LoRA fine-tuning efficiently undoes safety training in Llama 2-Chat 70B. arXiv preprint arXiv:2310.20624, 2023. Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. In NeurIPS, 2024. Liu, X., Xu, N., Chen, M., and Xiao, C. AutoDAN: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451, 2023. Marks, S. and Tegmark, M. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, 2023. Mazeika, M., Phan, L., Yin, X., Zou, A., et al. HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024. Pan, W., Liu, Z., Chen, Q., Zhou, X., Yu, H., and Jia, X. The hidden dimensions of LLM alignment: A multi-dimensional analysis of orthogonal safety directions. In ICML, 2025. Panickssery, N., Bowman, S.R., and Feng, S. Steering Llama 2 via contrastive activation addition. arXiv preprint arXiv:2312.06681, 2023. Park, K., Choe, Y.J., and Veitch, V. The linear representation hypothesis and the geometry of large language models. arXiv preprint arXiv:2311.03658, 2023. Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P. Safety alignment should be made more than just a few tokens deep. In ICLR, 2025. Siu, V., et al. Programming refusal with conditional activation steering. In ICLR, 2025. Siu, V., Henry, N.W., Crispino, N., Liu, Y., Song, D., and Wang, C. RepIt: Steering language models with concept-specific refusal vectors. In ICLR, 2026. Templeton, A., et al. Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits Thread, 2024. Tigges, C., Hollinsworth, O.J., Geiger, A., and Nanda, N. Linear representations of sentiment in large language models. arXiv preprint arXiv:2310.15154, 2023. Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248, 2023. Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does LLM safety training fail? In NeurIPS, 2023. 10 Weidmann, P.E. Heretic: Fully automatic censorship removal for language models.https://github.com/ p-e-w/heretic, 2025. Wollschläger, T., Elstner, J., Geisler, S., Cohen-Addad, V., Günnemann, S., and Gasteiger, J. The geometry of refusal in large language models: Concept cones and representational independence. In ICML, 2025. Xie, T., et al. SORRY-Bench: Systematically evaluating large language model safety refusal behaviors. arXiv preprint arXiv:2406.14598, 2024. Yang, X., et al. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949, 2023. Young, A. Comparative analysis of LLM abliteration methods: A cross-architecture evaluation. arXiv preprint arXiv:2512.13655, 2025. Zhang, A.K., Perry, N., Dulepet, R., Ji, J., Menders, C., Lin, J.W., et al. Cybench: A framework for evaluating cybersecurity capabilities and risks of language models. In ICLR, 2025. Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., et al. Representation engineering: A top-down approach to AI transparency. arXiv preprint arXiv:2310.01405, 2023. Zou, A., Wang, Z., Kolter, J.Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023. Zou, A., Phan, L., Wang, J., et al. Improving alignment and robustness with circuit breakers. In NeurIPS, 2024. A Model Inventory Table 3 lists all 24 models included in this study with their key properties. Table 3: Complete model inventory. Params lists total/active for MoE models. Class denotes susceptibility tier (High / Partial / Resistant). Developer abbreviations: Meta (Me), Alibaba/Qwen (Al), Google (Go), Microsoft (Ms), Mistral (Mi), Moonshot (Mo), DeepSeek (DS), Zhipu (Zh), Allen AI (A). ModelParamsArchDevSafetyLicenseClass Qwen3-0.6B0.6BDenseAlGRPOApache 2.0Partial Gemma-3-1B1BDenseGoBOND+GemmaHigh Llama-3.2-1B1.3BDenseMePPO+DPOLlama 3.2 CLHigh Llama-3.2-3B3.2BDenseMePPO+DPOLlama 3.2 CLHigh Phi-4-mini3.8BDenseMsDPO+GOATMITHigh OLMoE-1B-7B7B/1BMoEAADPOApache 2.0High Llama-3.1-8B8BDenseMePPO+DPOLlama 3.1 CLHigh Qwen3-8B8.2BDenseAlGRPOApache 2.0High Gemma-3-12B12BDenseGoBOND+GemmaPartial Phi-414BDenseMsDPO+GOATMITHigh Qwen3-14B14.8BDenseAlGRPOApache 2.0Resistant DS-V2-Lite16B/2.4BMoEDSSFT+RLMITHigh Mistral-Small24BDenseMiUndiscl.Apache 2.0Resistant Gemma-3-27B27BDenseGoBOND+GemmaHigh Qwen3-30B30B/3.3BMoEAlGRPOApache 2.0Partial Qwen3-32B32.8BDenseAlGRPOApache 2.0Partial Llama-3.3-70B70BDenseMePPO+DPOLlama 3.3 CLPartial Llama-4-Scout109B/17BMoEMePPO+DPOLlama 4 CLResistant Mistral-Large123BDenseMiUndiscl.MRLResistant Qwen3-235B235B/22BMoEAlGRPOApache 2.0Resistant GLM-4.7355B/32BMoEZhRLHFApache 2.0Partial Llama-4-Mav.400B/17BMoEMePPO+DPOLlama 4 CLPartial DeepSeek-V3671B/37BMoEDSSFT+RLMITHigh Kimi-K2∼1T/32BMoEMoSFT+RLMoonshot CLHigh 11 B Implementation and Pipeline The abliteration pipeline is implemented in Python around three components:vLLMfor inference, speculatorsfor activation extraction, and directsafetensorsmanipulation via PyTorch for weight modification. Activation extraction.vLLMoptimizes for inference throughput and does not expose hidden-state activations from intermediate decoder layers. We therefore usespeculators, which provides per-layer hooks while loading from the same checkpoint format. For each prompt we record the hidden state at the last token position of every decoder layer; these are aggregated into the per-layer mean-difference refusal direction ˆ r ℓ defined in Section 3. Weight modification. Rather than load the full model viatransformersand re-serialize a modified state dict, we open thesafetensorsfiles directly, apply the orthogonal projectionW ′ = W − α ˆ r( ˆ r ⊤ W)in PyTorch only to the targeted tensors (down_projando_projper modified layer), and write the updated tensors back. This avoids ever materializing the full model in CPU or GPU memory — a hard requirement at trillion-parameter scale, where the Kimi K2 checkpoint alone exceeds 1.1 TB. Caching. For each model we cache the extracted activations and the resulting refusal directions, so that running multiple downstream configurations (different layer percentiles, strengthsα, or layer- selection strategies) does not require re-running the extraction pass. This is what made it tractable to sweep six abliteration intensities across all 24 models within the 600 GPU-hour budget reported in Section 3. C Layer Selection Strategy Layer selection strategy (Figure 7) compares norm-based selection (targeting layers with highest refusal signal) to uniform spread (layers distributed across 25–95% of depth). Uniform spread consistently matches or outperforms norm-based selection across all 9 tested models, with the gap reaching∼70p of additional refusal reduction at 30% on the most susceptible models (e.g., Llama- 3.1-8B, Qwen3-8B, Kimi-K2). On the resistant tier (Mistral-Large), neither method moves refusal at any percentile. 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 102030 0 50 100 Norm-basedUniform spread %%% Refusal % Qwen3-0.6BOLMoEQwen3-8B Llama-8BGemma-12BQwen3-30B Mistral-LargeGLM-4.7Kimi-K2 Figure 7: Norm-based vs. uniform spread layer selection across 9 models. Uniform spread (red) consistently achieves equal or greater refusal reduction than norm-based (blue). D Domain-Specificity Analysis We extracted per-domain refusal directions for 5 representative models (Qwen3-8B, Mistral-Small- 3.1-24B, Qwen3-30B-A3B, Llama-4-Scout, GLM-4.7) spanning dense and MoE architectures from 3.8B to 355B parameters. For each model, hidden states were extracted at all decoder layers, 12 per-domain refusal directions computed as ˆ r (d) ℓ = norm(mean(h harmful,d ℓ )− mean(h harmless ℓ )) , and pairwise cosine similarity averaged across layers in the 25–95% depth range. Figure 9 shows the aggregated similarity matrix on our cross-evaluation dataset (6 domains). Cyber- security consistently has the lowest inter-domain similarity (0.56–0.70), confirming that its refusal direction is geometrically distinct. The maximum similarity pair is illegal goods↔violence (0.94), suggesting these domains share nearly identical refusal geometry. Figure 8 shows the same analysis on the scientific benchmark, where cybercrime and copyright emerge as the most distinct domains. Chem/Bio Copyright Cybercrime Harassment Illegal Misinfo Misinfo Illegal Harassment Cybercrime Copyright Chem/Bio 0.4 0.5 0.6 0.7 0.8 0.9 1 Cosine Sim 1.000.540.780.860.940.83 0.541.000.550.610.570.67 0.780.551.000.750.790.78 0.860.610.751.000.940.93 0.940.570.790.941.000.89 0.830.670.780.930.891.00 Figure 8: Per-domain refusal direction similarity on the scientific benchmark, averaged across 5 models. Cybercrime and copyright show the most distinct refusal directions. Cybersecurity Explicit Illegal Goods Misinfo Privacy Violence Violence Privacy Misinfo Illegal Goods Explicit Cybersecurity 0.4 0.5 0.6 0.7 0.8 0.9 1 Cosine Sim 1.000.620.650.670.750.66 0.621.000.790.880.800.83 0.650.791.000.820.860.95 0.670.880.821.000.820.84 0.750.800.860.821.000.85 0.660.830.950.840.851.00 Figure 9: Per-domain refusal direction similarity on the cross-evaluation dataset, averaged across 5 models. Cybersecurity has the most distinct refusal direction. 13 E Predictors of Susceptibility We analyzed several candidate predictors of abliteration susceptibility across all 24 models. Safety training method (Figure 10) is the strongest individual signal. DPO-based training produces the most susceptible models (median∼30p, all above the susceptibility threshold). RLHF-based and PPO+DPO methods cluster around 20–25p. GRPO-based training (used by all Qwen3 variants) yields consistently low susceptibility (5–8p), placing most models in the partial or resistant range. Models with undisclosed RL-based safety (Mistral-Small, Mistral-Large) are fully resistant (<1p). Google’s BOND+WARM+WARP method shows the widest variance, with Gemma-1B at∼48p and Gemma-27B at∼13p. Architecture and model family (Figure 11) show weaker but visible patterns. Dense transformers have a higher median susceptibility than MoE models, although both distributions overlap substan- tially. By family: Phi models are uniformly susceptible (all>20p), Qwen spans the full range (1–39p), Llama clusters in the mid-range, and Mistral is consistently resistant. DPORLHFPPO+DPOSFT onlyBOND+GRPODPO+GOATRL(?) 0 10 20 30 40 50 Max Refusal Drop (p) Resistant Partial Figure 10: Maximum refusal reduction by safety training method. DPO-based methods are most susceptible; GRPO and undisclosed RL methods confer near-resistance. DENSEMOE 0 10 20 30 40 50 PhiOlmoeDeepseekLlamaGemmaQwenGlmMistral Max Refusal Drop (p) By ArchitectureBy Model Family ResistantResistant PartialPartial Figure 11: Susceptibility by architecture type (left) and model family (right). Dense and MoE show comparable overall distributions, but family-level patterns emerge. 14