Paper deep dive
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
Xunguang Wang, Wenxuan Wang, Zhenlan Ji, Zongjie Li, Pingchuan Ma, Daoyuan Wu, Shuai Wang
Models: Llama-2-7B-Chat
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 6:27:03 PM
Summary
STShield is a lightweight framework for real-time jailbreak detection in Large Language Models (LLMs) that integrates a single-token sentinel mechanism. By appending a binary safety indicator ('safe' or 'harm') to the model's response sequence, STShield leverages the LLM's internal alignment capabilities. The framework utilizes supervised fine-tuning on normal prompts and adversarial training with embedding-space perturbations to achieve robust detection with minimal computational overhead.
Entities (5)
Relation Signals (3)
STShield â uses â Single-token sentinel
confidence 100% ¡ STShield introduces a novel single-token sentinel mechanism
STShield â detects â Jailbreak Attack
confidence 95% ¡ STShield successfully defends against various jailbreak attacks
STShield â employs â Adversarial Training
confidence 95% ¡ Our framework combines supervised fine-tuning on normal prompts with adversarial training
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) have become increasingly vulnerable to jailbreak attacks that circumvent their safety mechanisms. While existing defense methods either suffer from adaptive attacks or require computationally expensive auxiliary models, we present STShield, a lightweight framework for real-time jailbroken judgement. STShield introduces a novel single-token sentinel mechanism that appends a binary safety indicator to the model's response sequence, leveraging the LLM's own alignment capabilities for detection. Our framework combines supervised fine-tuning on normal prompts with adversarial training using embedding-space perturbations, achieving robust detection while preserving model utility. Extensive experiments demonstrate that STShield successfully defends against various jailbreak attacks, while maintaining the model's performance on legitimate queries. Compared to existing approaches, STShield achieves superior defense performance with minimal computational overhead, making it a practical solution for real-world LLM deployment.
Tags
Links
- Source: https://arxiv.org/abs/2503.17932
- Canonical: https://arxiv.org/abs/2503.17932
Trouble viewing inline? Open PDF directly â
Full Text
61,685 characters extracted from source content.
Expand or collapse full text
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models Xunguang Wang1, Wenxuan Wang1, Zhenlan Ji1, Zongjie Li1, Pingchuan Ma1, Daoyuan Wu1, Shuai Wang1 1The Hong Kong University of Science and Technology Correspondence: shuaiw@cse.ust.hk Abstract Large Language Models (LLMs) have become increasingly vulnerable to jailbreak attacks that circumvent their safety mechanisms. While existing defense methods either suffer from adaptive attacks or require computationally expensive auxiliary models, we present STShield, a lightweight framework for real-time jailbroken judgement. STShield introduces a novel single-token sentinel mechanism that appends a binary safety indicator to the modelâs response sequence, leveraging the LLMâs own alignment capabilities for detection. Our framework combines supervised fine-tuning on normal prompts with adversarial training using embedding-space perturbations, achieving robust detection while preserving model utility. Extensive experiments demonstrate that STShield successfully defends against various jailbreak attacks, while maintaining the modelâs performance on legitimate queries. Compared to existing approaches, STShield achieves superior defense performance with minimal computational overhead, making it a practical solution for real-world LLM deployment. STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models Xunguang Wang1, Wenxuan Wang1, Zhenlan Ji1, Zongjie Li1, Pingchuan Ma1, Daoyuan Wu1, Shuai Wang1 1The Hong Kong University of Science and Technology Correspondence: shuaiw@cse.ust.hk Warning: This paper contains unfiltered and potentially harmful content. 1 Introduction Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains Zhao et al. (2023), from natural language understanding to complex reasoning tasks. Their deployment in real-world applications has revolutionized human-AI interaction through chatbots, virtual assistants, and automated content generation systems. However, despite their sophisticated safety alignment mechanisms, these models remain vulnerable to jailbreak attacks Zou et al. (2023) that can manipulate them into generating harmful, unethical, or dangerous content. The evolution of jailbreak techniques has posed increasingly serious threats to LLM systems. These methods have progressed from simple manual methods Liu et al. (2023); Wei et al. (2023a, b); Shen et al. (2024b); Deng et al. (2024a) to sophisticated approaches including optimization-based algorithms Zou et al. (2023); Sitawarin et al. (2024); Andriushchenko et al. (2025); Liu et al. (2024b); Jia et al. (2025), generation-based strategies Perez et al. (2022); Deng et al. (2024a); Chao et al. (2023); Mehrotra et al. (2024); Paulus et al. (2024), indirect techniques Handa et al. (2024); Li et al. (2024b); Chang et al. (2024); Wei et al. (2023a), and multilingual attacks Deng et al. (2024b); Yong et al. (2023); Shen et al. (2024a); Li et al. (2024a); Wei et al. (2023a). Advanced techniques such as recursive prompt refinement Yu et al. (2024) and hierarchical genetic algorithms Liu et al. (2024b) have shown concerning success rates in circumventing safety guardrails. This proliferation of jailbreak methods has raised significant concerns about the deployment of LLMs in production environments, necessitating robust defense mechanisms. Existing defense strategies primarily fall into three categories: prompt-based Xie et al. (2023); Wei et al. (2023b); Zhang et al. (2024b); Zhou et al. (2024); Mo et al. (2024); Xiong et al. (2024) and tuning-based approaches Lu et al. (2024); Xhonneux et al. (2024); Zhao et al. (2024); Liu et al. (2024a) that aim to enhance model robustness, and detection-based methods Alon and Kamfonas (2023); Jain et al. (2023); Kumar et al. (2023); Cao et al. (2023); Robey et al. (2023); Ji et al. (2024); Inan et al. (2023); Zhang et al. (2024a); Xie et al. (2024); Hu et al. (2024); Qian et al. (2024) that identify malicious queries or harmful responses. While prompt-based and tuning-based methods have shown promise in improving model resistance to jailbreak attempts, these methods often struggle against adaptive attacks that can bypass the enhanced safety mechanisms. Detection-based approaches, particularly those focusing on response monitoring, offer a more straightforward and effective defense strategy. However, current response-detection methods typically rely on auxiliary LLMs Inan et al. (2023); Phute et al. (2023) for safety verification, introducing substantial computational overhead. Additionally, the auxiliary LLM is typically smaller than the target LLM, leading to compromised detection accuracy. In this paper, we propose STShield, a lightweight yet effective framework that appends a single safety token to the modelâs response sequence for real-time jailbreak detection, i.e., we jointly unify the response generation and jailbreak detection in a single model. Specifically, STShield leverages this safety token as a binary indicator: "safe" for legitimate responses and "harm" for successfully jailbroken outputs. STShieldâs training process encompasses two key components: First, we conduct supervised fine-tuning with normal prompts, where the safety token is consistently set to "safe" to maintain the modelâs standard functionality. Second, we employ an adversarial training strategy using harmful instructions, where we optimize perturbations in the LLMâs embedding space to construct potential jailbreak inputs. During this phase, the safety token is dynamically assigned based on the success of the embedding attack - "harm" for successful jailbreaks and "safe" otherwise. This dual-phase training approach enables STShield to achieve robust jailbreak detection while preserving the base LLMâs utility, as demonstrated through extensive experimental evaluation. Besides, this unified response generation and single-token detection model significantly reduces computational overhead compared to existing response filtering methods, making it a practical solution for real-world LLM deployment. The key contributions of our work include: ⢠A novel single-token detection mechanism that integrates seamlessly with existing LLM architectures, enabling efficient real-time jailbreak detection. ⢠An adaptive adversarial training framework that enhances the modelâs ability to identify and flag sophisticated jailbreak attempts while maintaining normal functionality for legitimate queries. ⢠Extensive experimental evaluations demonstrate superior defenese performance compared to existing methods while significantly reducing computational overhead. 2 Related Work 2.1 Jailbreak Attacks In the realm of jailbreak attacks on large language models (LLMs), existing strategies can be categorized into manual, optimization-based, generation-based, indirect, and multilingual approaches. Manual jailbreaks Liu et al. (2023); Wei et al. (2023a, b); Shen et al. (2024b); Deng et al. (2024a) involve human-crafted adversarial prompts to exploit LLM vulnerabilities, with researchers like Wei et al. (2023a) designing prompts based on out-of-distribution inputs and conflicting model goals, Deng et al. (2024a) engineering a proof-of-concept (PoC) jailbreak prompt by making the LLM act as AIM (Always Intelligent and Machiavellian), and Shen et al. (2024b) developing a platform for crowdsourcing jailbreak prompts. Optimization-based jailbreaks Zou et al. (2023); Sitawarin et al. (2024); Andriushchenko et al. (2025); Liu et al. (2024b); Jia et al. (2025) use methods like gradient-based or search-based algorithms to iteratively refine adversarial prompts, with innovations such as the greedy coordinate gradient (GCG) Zou et al. (2023) and hierarchical genetic algorithms Liu et al. (2024b) improving the effectiveness and readability of these prompts, respectively. Generation-based jailbreaks Perez et al. (2022); Deng et al. (2024a); Chao et al. (2023); Mehrotra et al. (2024); Paulus et al. (2024); Yu et al. (2024) leverage auxiliary LLMs to engineer deceptive prompts that mislead target models into producing restricted content, with techniques ranging from feedback loops Chao et al. (2023); Mehrotra et al. (2024) to specialized training Paulus et al. (2024) for generating adversarial prompts. Indirect jailbreaks Handa et al. (2024); Li et al. (2024b); Chang et al. (2024); Wei et al. (2023a) aim to conceal malicious intents within seemingly innocuous queries to bypass safety mechanisms, employing tactics like word substitution Handa et al. (2024), sub-prompt decomposition Li et al. (2024b) to mask the harmful nature of the prompts, and clues leading to LLMâs jailbreak Chang et al. (2024). Lastly, multilingual jailbreaks Deng et al. (2024b); Yong et al. (2023); Shen et al. (2024a); Li et al. (2024a); Wei et al. (2023a) exploit the lower alignment of LLMs in less-resourced languages to facilitate jailbreaking, with researchers developing multilingual prompt datasets (e.g., MultiJail Deng et al. (2024b)) and using obfuscation techniques to encode or encrypt the original harmful instructions Wei et al. (2023a); Yuan et al. (2024); Chu et al. (2024). 2.2 Jailbreak Defense Jailbreak defense methods aim to protect the LLMs from being manipulated or exploited through jailbreak attacks, and these approaches can be roughly grouped into three categories: detection-based, prompt-based and tuning-based jailbreak defenses. Detection-based methods deny responses to jailbreak queries by detecting their malicious intent Zhang et al. (2024a); Wang et al. (2025). Among them, their distinctions come from whether to check the input prompts Alon and Kamfonas (2023); Jain et al. (2023); Kumar et al. (2023); Cao et al. (2023); Robey et al. (2023); Ji et al. (2024); Inan et al. (2023); Zhang et al. (2024a) or to analyze the internal states Xie et al. (2024); Hu et al. (2024); Qian et al. (2024) and responses Inan et al. (2023); Phute et al. (2023). For instance, SmoothLLM Robey et al. (2023) disturbs the multiple copies of the query prompt and aggregates their corresponding predictions for jailbreak detection. Furthermore, prompt-based methods improve the LMMâs resistance to jailbreaks by modifying input prompts. A straightforward approach involves adding a human-crafted system prompt or contexts, encouraging the LLMâs safe response behavior Xie et al. (2023); Wei et al. (2023b); Zhang et al. (2024b). Recent studies Zhou et al. (2024); Mo et al. (2024); Xiong et al. (2024) focus on learning a universal prefix or suffix for defense control by incorporating adversarial training. Besides, tuning-based methods Lu et al. (2024); Xhonneux et al. (2024); Zhao et al. (2024); Liu et al. (2024a) try to improve a modelâs robustness against jailbreaks by optimizing its internal parameters. The most representative work is Adversarial Tuning Liu et al. (2024a), which fine-tunes the LLM with adversarial training. 3 Method 3.1 Preliminary Target Model. We mainly consider popular auto-regressive LLMs that predicts the next token by the previous sequence. Given a sequence of previous tokens 1:nsubscript:1x_1:nx1 : n where xiâ1,âŻ,subscriptx1⯠x_iâ\1,¡s,V\xi â 1 , ⯠, blackboard_V (Vblackboard_V denoting the vocabulary size, namely, the number of tokens), the primary task of LLMs to output the response sequence n+1:n+msubscript:1x_n+1:n+mxitalic_n + 1 : n + m can be formulated as: Pθâ˘(n+1:n+m|1:n)=âi=1mPθâ˘(n+i|1:n+iâ1)subscriptconditionalsubscript:1subscript:1superscriptsubscriptproduct1subscriptconditionalsubscriptsubscript:11P_θ(x_n+1:n+m|x_1:n)= _i=1^mP_θ(% x_n+i|x_1:n+i-1)Pitalic_θ ( xitalic_n + 1 : n + m | x1 : n ) = âi = 1m Pitalic_θ ( xitalic_n + i | x1 : n + i - 1 ) (1) where Pθâ˘(n+1:n+m|1:n)subscriptconditionalsubscript:1subscript:1P_θ(x_n+1:n+m|x_1:n)Pitalic_θ ( xitalic_n + 1 : n + m | x1 : n ) represents the probability of the response sequence and θ denotes the LLM parameters. Jailbreak Attack. The adversaryâs goal is to identify adversarial prompt that compel the LLM to generate a target sequence (e.g., "Sure, here is a tutorial on how to commit identity theft"). The objective function for this attack can be written as follows: âjâ˘aâ˘iâ˘(^1:n,^)=âlogâĄPθâ˘(^|^1:n)subscriptâsubscript^:1^subscriptconditional^subscript^:1L_jai( x_1:n, y)=- P_θ(% y| x_1:n)Litalic_j a i ( over start_ARG x end_ARG1 : n , over start_ARG y end_ARG ) = - log Pitalic_θ ( over start_ARG y end_ARG | over start_ARG x end_ARG1 : n ) (2) where âjâ˘aâ˘iâ˘(^1:n,^)subscriptâsubscript^:1^L_jai( x_1:n, y)Litalic_j a i ( over start_ARG x end_ARG1 : n , over start_ARG y end_ARG ) represents the jailbreaking loss, ^1:nsubscript^:1 x_1:nover start_ARG x end_ARG1 : n represents the jailbreak prompt, and ^ yover start_ARG y end_ARG denotes the target sequence. Problem Formulation. Given a well-trained LLM fθâ˘(â )subscriptâ f_θ(¡)fitalic_θ ( â ), we aim to fine-tune it to jointly perform regular response generation and jailbreak detection. The whole output sequence consists of three modules: an answer tokens n+1:n+msubscript:1x_n+1:n+mxitalic_n + 1 : n + m generated by the LLM for the input prompt, an EOS token xeâ˘oâ˘ssubscriptx x_eosxe o s, and a detection token (safety token) xdsubscriptx x_dxd to indicate the safety of the forward response. The detection token xdsubscriptx x_dxd is encoded as âsafeâ for normal responses and âharmâ for unsafe contents (i.e., successful jailbreaks). The fine-tuning process can be described as the minimum optimization problem: argâ˘minθâ˛âĄâsubscriptargminsuperscriptâ˛â *arg\,min_θ Lstart_OPERATOR arg min end_OPERATORθⲠL =â1:n,xdââlogPθâ˛( = _x_1:n, x_d - P_% θ (= âx start_POSTSUBSCRIPT 1 : n , xd â O end_POSTSUBSCRIPT - log Pitalic_θⲠ( (3) n+1:n+mâxeâ˘oâ˘sâxd|1:n), _n+1:n+m x_eos x% _d|x_1:n),xitalic_n + 1 : n + m â xe o s â xd | x1 : n ) , n+1:n+m=fθâ˘(1:n),subscript:1subscriptsubscript:1 _n+1:n+m=f_θ(x_1:n),xitalic_n + 1 : n + m = fitalic_θ ( x1 : n ) , where θâ˛Î¸ θⲠrepresents the parameters of the LLM after fine-tuning, OO is an instruction dataset containing normal and malicious prompts, n+1:n+mâxeâ˘oâ˘sâxddirect-sumsubscript:1subscriptxsubscriptxx_n+1:n+m x_eos x_dxitalic_n + 1 : n + m â xe o s â xd is the output token sequence and âdirect-sum â denotes the concatenation operation. Threat Model. In our threat model, we assume that adversaries have only black-box access to the target LLM, meaning they can only interact with the model through its input and output interfaces without any knowledge of its internal architecture or parameters. This limitation restricts attackers to probing the probabilities of output sequence and gradients to craft jailbreak attempts. Despite this constraint, adversaries can employ sophisticated techniques, including optimization-based algorithms, generation-based strategies, and indirect methods, to exploit vulnerabilities in the modelâs safety alignment. Our defense mechanism, STShield, is designed to operate under this threat model, ensuring robust protection against jailbreak attempts while maintaining the modelâs utility for legitimate queries. By leveraging a single safety token for real-time detection, STShield effectively mitigates the risk of harmful content generation without requiring additional computational resources or compromising the modelâs performance. 3.2 Supervised Fine-Tuning Our supervised fine-tuning focus on optimizing the target LLM fθâ˘(â )subscriptâ f_θ(¡)fitalic_θ ( â ) with normal prompts, to enable its effectiveness on human-crafted or ready-made queries. For normal prompts, we encourage the fine tuned LLM fθâ˛â˘(â )subscriptsuperscriptâ˛â f_θ (¡)fitalic_θⲠ( â ) to respond positively like the original model fθâ˘(â )subscriptâ f_θ(¡)fitalic_θ ( â ) and identify the forward sequence as safe. Thus, the loss function for normal queries is defined as follows: ânâ˘oâ˘rsubscriptâ _norLitalic_n o r =â1:nânâ˘oâ˘râlogPθâ˛( = _x_1:n _nor- P_θ^% (= âx start_POSTSUBSCRIPT 1 : n â Oitalic_n o r end_POSTSUBSCRIPT - log Pitalic_θⲠ( (4) n+1:n+mâxeâ˘oâ˘sâxsâ˘aâ˘fâ˘e|1:n), _n+1:n+m x_eos x% _safe|x_1:n),xitalic_n + 1 : n + m â xe o s â xs a f e | x1 : n ) , n+1:n+m=fθâ˘(1:n),subscript:1subscriptsubscript:1 _n+1:n+m=f_θ(x_1:n),xitalic_n + 1 : n + m = fitalic_θ ( x1 : n ) , where nâ˘oâ˘rsubscriptO_norOitalic_n o r is a normal instruction dataset, and xsâ˘aâ˘fâ˘esubscriptx x_safexs a f e indicates the safe token, i.e., the encoded token of âsafeâ. ânâ˘oâ˘rsubscriptâL_norLitalic_n o r can ensure that the fine-tuned LLM responds regularly to normal users rather than over-refusal. 3.3 Adversarial Training To further improve the robustness of our fine-tuned model against jailbreaks, we conduct adversarial training for it with adaptive attacks which defeats both of its safety alignment and detection module. Generating Jailbreak Prompts. In this paper, we adopt continuous embedding attacks Xhonneux et al. (2024) for fast adversarial training. Given a malicious instruction 1:nsubscript:1x_1:nx1 : n, we mark its embedding as 1:nsubscript:1e_1:ne1 : n extracted by the target LLM. The embedding attack achieves the jailbreak effect by optimizing 1:nsubscript:1e_1:ne1 : nâs perturbation δ1:nsubscript:1 _1:nδ1 : n to induce the target response n+1:n+msubscript:1x_n+1:n+mxitalic_n + 1 : n + m (e.g., âSure, here is a tutorial for making a bombâ). For our fine-tuned LLM, we define the last token of the target sequence as xsâ˘aâ˘fâ˘esubscriptx x_safexs a f e, since the last token in the output sequence is a discriminate token for safety detection. Formally, the target sequence Tsuperscriptx^Txitalic_T can be written as: Tsuperscript ^Txitalic_T =n+1:n+m+2absentsubscript:12 =x_n+1:n+m+2= xitalic_n + 1 : n + m + 2 (5) =n+1:n+mâxeâ˘oâ˘sâxsâ˘aâ˘fâ˘e.absentdirect-sumsubscript:1subscriptxsubscriptx =x_n+1:n+m x_eos % x_safe.= xitalic_n + 1 : n + m â xe o s â xs a f e . To simplify the notation, let Osuperscripte^Oeitalic_O denote the embedding of the malicious instruction 1:nsubscript:1x_1:nx1 : n, δ indicate δ1:nsubscript:1 _1:nδ1 : n, and O+δsuperscripte^O+ _O + δ indicate the jailbreak embedding. The objective of the jailbreak attack can be formulated as follows: minδâĄâa=âOâhâ˘aâ˘râlogâĄPθâ˛â˘(T|O+δ),subscriptsubscriptâsubscriptsuperscriptsubscriptâsubscriptsuperscriptâ˛conditionalsuperscriptsuperscript _δL_a= _x^O _har- P_% θ (x^T|e^O+δ),minitalic_δ Litalic_a = âxitalic_O â O start_POSTSUBSCRIPT h a r end_POSTSUBSCRIPT - log Pitalic_θⲠ( xitalic_T | eitalic_O + δ ) , (6) where hâ˘aâ˘rsubscriptâO_harOitalic_h a r is a malicious instruction dataset. We adopt PGD Madry (2018) to optimize the perturbation δ for the jailbreak attack. The adversarial perturbation δ is updated iteratively by: δ(t+1)=δ(t)+Ρâ signâ˘(âδâa),superscript1superscriptâ signsubscriptâsubscriptâ δ^(t+1)=δ^(t)+Ρ¡% sign( _δL_a),δ( t + 1 ) = δ( t ) + Ρ â sign ( âδ Litalic_a ) , (7) where Ρ is the step size, and t is the iteration number. The optimization process is terminated when the maximum iteration is reached. Adversarial Tuning. Once constructing jailbreak embeddings, we use them as augmentation to optimize the target LLM for defense, i.e., adversarial tuning. Since our embedding-based attack cannot guarantee successful jailbreaking, we employ a jailbreak evaluator G to determine the value of xdsubscriptx x_dxd. Specifically, xdsubscriptx x_dxd is set to xhâ˘aâ˘râ˘msubscriptxâ x_harmxh a r m if G confirms a successful jailbreak attempt, and to xsâ˘aâ˘fâ˘esubscriptx x_safexs a f e otherwise. Thus, we define the objective of the adversarial tuning as follows: minθâ˛âĄâaâ˘dâ˘vsubscriptsuperscriptâ˛subscriptâ _θ L_advminitalic_θⲠLitalic_a d v =âOâhâ˘aâ˘râlogPθâ˛( = _x^O _har- P_θ % (= âxitalic_O â O start_POSTSUBSCRIPT h a r end_POSTSUBSCRIPT - log Pitalic_θⲠ( (8) Râxeâ˘oâ˘sâxd|O+δ), ^R x_eos x_d|% e^O+δ),xitalic_R â xe o s â xd | eitalic_O + δ ) , R=fθâ˛â˘(O+δ),superscriptsuperscriptsubscriptⲠ^R=f_θ (e^O+δ),xitalic_R = fitalic_θⲠ( eitalic_O + δ ) , xd=Gâ˘(1:n,R),subscriptxsubscript:1superscript x_d=G(x_1:n,x^R),xd = G ( x1 : n , xitalic_R ) , where Rsuperscriptx^Rxitalic_R is an output response of the jailbreak prompt from the fine-tuned LLM. Tuning the target LLM with this adversarial loss âaâ˘dâ˘vsubscriptâL_advLitalic_a d v can boost its resistance to adaptive jailbreaks. 3.4 Optimization As mentioned above, the overall objective function can be written as follows: minθâ˛âĄâsubscriptsuperscriptâ˛â _θ Lminitalic_θⲠL =ânâ˘oâ˘r+âaâ˘dâ˘v,absentsubscriptâsubscriptâ =L_nor+L_adv,= Litalic_n o r + Litalic_a d v , (9) where ânâ˘oâ˘rsubscriptâL_norLitalic_n o r and âaâ˘dâ˘vsubscriptâL_advLitalic_a d v are the supervised fine-tuning loss and adversarial tuning loss, respectively. We optimize the target LLM with the Adam optimizer Loshchilov (2019) to minimize the overall loss âLL. 3.5 Inference Mechanism During inference, when a user submits a request, STShield processes the input and generates a response as usual. If the response naturally ends with an end-of-sequence (EOS) token, STShield directly predicts the detection token. If the response does not conclude with an EOS token, STShield appends an EOS token before predicting the detection token. The detection token serves as a binary indicator: if it is classified as "safe," the system outputs the generated response without modification. If the detection token is classified as "harm," the system overrides the response with a predefined safety message, such as "I am sorry, I cannot provide that information." This mechanism ensures that harmful or inappropriate content is effectively filtered while maintaining seamless functionality for legitimate queries. 4 Experiments 4.1 Experimental Setup Datasets and Benchmarks. To fine-tune STShield, we utilize UltraChat and JailbreakBench Chao et al. (2024) datasets. Specifically, UltraChat Ding et al. (2023) encompasses over 12M conversations spanning a diverse array of topics. From this corpus, we randomly select 1,000 instructions to serve as normal prompts for supervised fine-tuning. For adversarial training, we employ all 100 harmful instructions from JailbreakBench. To assess the efficacy of jailbreak attacks, we leverage AdvBench Zou et al. (2023), a prominent dataset containing 520 instances of harmful or toxic behaviors, such as profanity, threats, misinformation, discrimination, and other illicit activities. In line with JailTrickBench Xu et al. (2024), we curate 50 refined harmful instructions for jailbreaking, addressing the issue of duplicates within AdvBench. Additionally, we adopt DAN Shen et al. (2024b) and MultiJail Deng et al. (2024b) as manual and multilingual attack datasets, respectively, and AlpacaEval Li et al. (2023) as a source of normal prompts for evaluation. The tested DAN, MultiJail, and AlpacaEval contains 1,00010001,0001 , 000, 315, and 805 prompts, respectively. For MultiJail, we specifically select all 315 Bengali prompts under the intentional scenario. Metrics. We employ the Attack Success Rate (ASR) as the primary metric to assess vulnerability, aligning with prior studies Zou et al. (2023); Liu et al. (2024b); Paulus et al. (2024); Xu et al. (2024). ASR quantifies the proportion of instructions that are not rejected and elicit appropriate responses. We calculate ASR using two methods: prefix-based (ASRPrefix) and GPT agent-based (ASRAgent), as utilized in JailTrickBench Xu et al. (2024). The prefix-based approach Zou et al. (2023); Liu et al. (2024b) detects the absence of rejection phrases such as "Sorry, I cannot." The GPT agent-based method Chao et al. (2023); Mehrotra et al. (2024) leverages a GPT-4 model to evaluate whether the target model has been jailbroken based on the input instruction and its corresponding response. Baselines. For target models, we utilize Llama-2-7B-Chat Touvron et al. (2023) and Vicuna-13B-v1.5 Chiang et al. (2023) as the base models. To assess the effectiveness of jailbreak defense methods, we follow JailTrickBench by implementing a range of adaptive attacks, including token-level attacks (AmpleGCG Liao and Sun (2024) and AdvPrompter Paulus et al. (2024)) and prompt-level attacks (PAIR Chao et al. (2023), TAP Mehrotra et al. (2024), and LLM-Fuzzer Yu et al. (2024)), as well as static jailbreaks (manual DAN Shen et al. (2024b) and multilingual MultiJail Deng et al. (2024b)). To evaluate the utility of our method on normal queries, we employ AlpacaEval Li et al. (2023). For defense mechanisms, we implement detection-based defenses (SmoothLLM Robey et al. (2023) and Llama Guard Inan et al. (2023)), prompt-based defenses (Self-Reminder Xie et al. (2023) and RPO Zhou et al. (2024)), and tuning-based defenses (Adversarial Training Xhonneux et al. (2024), Unlearning Lu et al. (2024), and Safety Training Siththaranjan et al. (2023)). We use Llama Guard as a filter for the target modelâs responses and evaluate their synergistic output, similar to our STShieldâs inference mechanism. Implementation Details. In the adversarial training of STShield, the jailbreak evaluator G adopts the same judgement strategy of ASRPrefix. For PGD, we set the maximum number of iterations to 8 and the step size to 0.001. We use LoRA to fine tune the target LLM with r of 16 and Îą of 32 Hu et al. (2022). The initial learning rate is 5Ă10â45superscript1045Ă 10^-45 Ă 10- 4. The training iterations are 1,00010001,0001 , 000 in a batch size of 1. All experiments are conducted on a single NVIDIA H800 GPU. Table 1: Jailbreak defense experiments with adaptive attacks under ASRPrefixPrefix_Prefixstart_FLOATSUBSCRIPT Prefix end_FLOATSUBSCRIPT. Noted that tested Llama Guard is as a filter for the target LLMâs responses and evaluate their synergistic output, analogous to the inference mechanism of our STShield. Defense Methods AmpleGCG AdvPrompter PAIR TAP LLM-Fuzzer Vicuna-13B No Defense 100.00 100.00 36.00 28.00 78.00 Self-Reminder 100.00 100.00 28.00 24.00 30.00 RPO 100.00 100.00 60.00 38.00 38.00 SmoothLLM 94.00 90.00 88.00 96.00 90.00 Adv. Training 100.00 98.00 44.00 30.00 66.00 Unlearning 100.00 100.00 76.00 70.00 32.00 Safety Training 100.00 100.00 20.00 22.00 72.00 Llama Guard 100.00 98.00 38.00 58.00 62.00 STShield (ours) 30.00 34.00 18.00 22.00 28.00 LLaMA-2-7B-Chat No Defense 100.00 98.00 18.00 18.00 6.00 Self-Reminder 100.00 100.00 16.00 22.00 2.00 RPO 100.00 100.00 60.00 38.00 18.00 SmoothLLM 74.00 64.00 40.00 36.00 82.00 Adv. Training 100.00 98.00 18.00 16.00 18.00 Unlearning 100.00 96.00 12.00 18.00 2.00 Safety Training 100.00 98.00 12.00 12.00 22.00 Llama Guard 100.00 96.00 38.00 2.00 0.00 STShield (ours) 12.00 0.00 2.00 2.00 10.00 4.2 Results Defense against Adaptive Jailbreaks. Table 1 shows the ASRPrefix results of our STShield and other defense methods against adaptive attacks on Vicuna-13B and Llama-2-7B-Chat. We additionally provide the ASRAgent results in the Appendix A. When compared to the no defense baseline, STShield demonstrates a dramatic reduction in ASRPrefix across all evaluated attack methods and models. For instance, in the Vicuna-13B model, the ASR under the AmpleGCG attack drops from 100.00% (no defense) to 30.00% with STShield, representing a 70% reduction in vulnerability. Similarly, for the AdvPrompter attack, the ASR decreases from 100.00% to 34.00%, showcasing a 66% improvement in defense capability. This trend is consistent across other attacks such as PAIR (36.00% to 18.00%), TAP (28.00% to 22.00%), and LLM-Fuzzer (78.00% to 28.00%), underscoring STShieldâs ability to significantly mitigate attack success rates compared to an undefended model. The improvement is even more pronounced in the LLaMA-2-7B-Chat model, where STShield reduces the ASR for AmpleGCG from 100.00% to 12.00% and completely neutralizes the AdvPrompter attack (98.00% to 0.00%). These results clearly demonstrate that STShield provides a substantial enhancement in security over an undefended LLM. Furthermore, when compared to other defense methods, STShield consistently achieves the lowest or near-lowest ASR in the majority of cases, establishing its superiority. For example, in the Vicuna-13B model, STShield outperforms SmoothLLM (94.00%), Adv.Training (100.00%), and Llama Guard (100.00%) under the AmpleGCG attack. Similarly, for the AdvPrompter attack, STShieldâs ASR of 34.00% is significantly lower than that of RPO (100.00%) and Llama Guard (98.00%). This trend holds across other attacks, where STShield maintains a consistent advantage over competing methods. In the LLaMA-2-7B-Chat model, STShieldâs performance is even more striking, achieving an ASR of 12.00% for AmpleGCG compared to SmoothLLMâs 74.00% and Adv.Trainingâs 100.00%. Notably, STShieldâs complete mitigation of the AdvPrompter attack (0.00% ASR) is unparalleled by any other method. These results collectively demonstrate that STShield not only significantly improves upon the no defense baseline but also outperforms most existing defense mechanisms in the majority of scenarios. Hence, STShield represents a robust and effective defense strategy for LLMs, offering substantial improvements over undefended models and consistently achieving superior performance compared to other state-of-the-art defense methods. Its ability to significantly reduce ASR across diverse attack vectors underscores its potential as a reliable and advanced solution for safeguarding LLMs against adaptive jailbreak attacks. Table 2: ASRs (%) of static jailbreak/normal prompts. Defense Methods DAN (ASRâPrefix_Prefix _FLOATSUBSCRIPT Prefix end_FLOATSUBSCRIPT â) MultiJail (ASRâAgent_Agent _FLOATSUBSCRIPT Agent end_FLOATSUBSCRIPT â) AlpacaEval (ASRâPrefix_Prefix _FLOATSUBSCRIPT Prefix end_FLOATSUBSCRIPT â) No Defense 4.50 6.03 89.69 Llama Guard 4.30 4.44 89.69 STShield (ours) 2.40 0.63 91.43 Defense against Static Jailbreaks and Normal Prompts. The experimental results in Table 2 demonstrate the effectiveness of STShield in handling both static jailbreaks and normal prompts. For jailbreak prompts, STShield significantly reduces the ASR compared to no defense. Specifically, for the DAN attack, STShield achieves an ASR of 2.40%, a 46.67% reduction factor from no defense (4.50%) and a 44.19% reduction from Llama Guard (4.30%). For the MultiJail attack, STShieldâs ASR of 0.63% represents an 89.55% reduction factor from no defense (6.03%) and an 85.81% reduction from Llama Guard (4.44%). In contrast, for normal prompts (AlpacaEval), STShield not only maintains but improves the ASR to 91.43%, compared to 89.69% for both no defense and Llama Guard. These results highlight STShieldâs dual capability: robust defense against jailbreak prompts and enhanced performance on normal prompts, underscoring its effectiveness as a comprehensive defense mechanism for LLMs. Utility. The MT-Bench scores in Table 3 indicate that while STShield causes a slight reduction in performance for both Vicuna-13B (from 6.54 to 6.24) and LLaMA-2-7B-Chat (from 6.26 to 6.08), the impact on the modelsâ question-answering capabilities remains minimal. These reductions, approximately 4.59% and 2.88% respectively, suggest that STShieldâs integration introduces only marginal computational overhead. Despite this, the models retain high functionality, demonstrating that STShield effectively balances security enhancements with maintaining the utility of the base LLMs. This underscores STShieldâs practicality as a defense mechanism that safeguards against adversarial threats without significantly compromising performance. Table 3: MT-Bench Zheng et al. (2023) for base LLMs and STShield. Defense Method Base LLM Vicuna-13B Llama-2-7B-Chat No Defense 6.54 6.26 STShield 6.24 6.08 Delay. The results in Table 4 illustrate the delay introduced by various defense methods, highlighting that STShield incurs minimal additional latency compared to the no defense baseline and outperforms Llama Guard in terms of efficiency. Specifically, STShieldâs delay for DAN (1.88 seconds), MultiJail (1.86 seconds), and AlpacaEval (1.75 seconds) is significantly closer to the no defense baseline (1.85, 1.85, and 1.73 seconds, respectively) than Llama Guardâs delay (1.93, 1.92, and 1.78 seconds). This efficiency is attributed to STShieldâs ability to leverage the KV cache, which avoids the need to recompute the forward pass for prompts and responses, a process that Llama Guard must perform. By utilizing the KV cache, STShield minimizes computational overhead, ensuring that the defense mechanism adds negligible latency while maintaining robust security. This demonstrates STShieldâs superior balance between defense effectiveness and computational efficiency, making it a practical solution for real-world applications. Table 4: Delay (seconds) of jailbreak defenses. Defense Methods DAN MultiJail AlpacaEval No Defense 1.85 1.85 1.73 Llama Guard 1.93 1.92 1.78 STShield (ours) 1.88 1.86 1.75 Cases. The case analysis demonstrates the effectiveness of STShield in distinguishing between jailbreak prompts and normal prompts, as shown in Figure 1, 2, and 3. For the jailbreak prompt in Figure 1, which explicitly requests examples of content glorifying acts of terror or violence, STShield generates a response but ultimately flags it as "harm," indicating its ability to recognize and mitigate harmful content. This showcases STShieldâs robustness in handling adversarial inputs designed to elicit inappropriate or dangerous responses. For the failed jailbreak prompt in Figure 2, which attempts to provide the description of sexual acts, STShield refuses to generate a response, correctly identifying it as "safe". On the other hand, for the normal prompt asking about famous actors who started their careers on Broadway, STShield provides a relevant and accurate response, correctly identifying it as "safe." This highlights STShieldâs capability to maintain the modelâs utility for benign queries while ensuring security against harmful inputs. Together, these cases illustrate STShieldâs balanced approach to safeguarding large language models without compromising their functionality for legitimate use cases. Figure 1: Case analysis of STShield on a jailbreak prompt from DAN. Figure 2: Case analysis of STShield on a failed jailbreak prompt from DAN. 4.3 Ablation Study The ablation study results in Table 5 reveal critical insights into the contributions of ânâ˘oâ˘rsubscriptâL_norLitalic_n o r and âaâ˘dâ˘vsubscriptâL_advLitalic_a d v to the performance of STShield. Removing ânâ˘oâ˘rsubscriptâL_norLitalic_n o r, the supervised fine-tuning loss on normal prompts, results in an ASR of 0.00 for both DAN and MultiJail attacks, indicating complete mitigation of jailbreak attempts. However, this also leads to over-refusal on normal prompts, as evidenced by the AlpacaEval ASR dropping to 0.00. This suggests that ânâ˘oâ˘rsubscriptâL_norLitalic_n o r is essential for maintaining the modelâs ability to process benign inputs effectively. On the other hand, removing âaâ˘dâ˘vsubscriptâL_advLitalic_a d v, the adversarial training loss on harmful instructions, results in ASR values identical to the no defense baseline (4.50 for DAN and 6.03 for MultiJail), indicating no improvement in defense against adversarial attacks. This highlights the necessity of âaâ˘dâ˘vsubscriptâL_advLitalic_a d v for enhancing the modelâs robustness against jailbreak attempts. The full STShield model, incorporating both ânâ˘oâ˘rsubscriptâL_norLitalic_n o r and âaâ˘dâ˘vsubscriptâL_advLitalic_a d v, achieves a balanced performance, significantly reducing ASR for jailbreak prompts (2.40 for DAN and 0.63 for MultiJail) while improving the ASR for normal prompts (91.43 for AlpacaEval). This demonstrates the importance of both loss components in achieving a robust and effective defense mechanism. Figure 3: Case analysis of STShield on a normal prompt from AlpacaEval. Table 5: Defense performance for our ablation architectures. Defense Methods DAN (ASRâPrefix_Prefix _FLOATSUBSCRIPT Prefix end_FLOATSUBSCRIPT â) MultiJail (ASRâAgent_Agent _FLOATSUBSCRIPT Agent end_FLOATSUBSCRIPT â) AlpacaEval (ASRâPrefix_Prefix _FLOATSUBSCRIPT Prefix end_FLOATSUBSCRIPT â) wo/ânâ˘oâ˘rsubscriptâL_norLitalic_n o r 0.00 0.00 0.00 wo/âaâ˘dâ˘vsubscriptâL_advLitalic_a d v 4.50 6.03 89.69 STShield (ours) 2.40 0.63 91.43 5 Conclusion In this paper, we introduced STShield, a novel framework that addresses the critical challenge of jailbreak attacks on LLMs through an efficient single-token detection mechanism. Our approach demonstrates that integrating detection capabilities directly into the modelâs output sequence can provide robust defense while avoiding the computational overhead of external detection models. The dual-phase training strategy, combining supervised fine-tuning with adversarial training, enables STShield to maintain high detection accuracy across various attack types while preserving the modelâs utility for legitimate queries. Limitations While STShield demonstrates promising results in jailbreak detection, we acknowledge several limitations of our current approach. The primary limitation is the slight degradation in the modelâs general task performance, as evidenced by the decreased MT-Bench scores compared to the base model. This performance drop suggests that the introduction of the safety token mechanism and the associated training process may impact the modelâs ability to generate optimal responses for legitimate queries. This limitation likely stems from the inherent trade-off between security measures and model utility. Our current training process, while effective for jailbreak detection, may not fully preserve all the nuanced capabilities of the original model. A potential solution would be to incorporate more diverse and comprehensive conversation datasets during the training phase, which could help maintain the modelâs general performance while retaining its enhanced security features. Additionally, the effectiveness of our approach might be contingent on the quality and diversity of the training data used for both supervised fine-tuning and adversarial training. Future work could explore methods to optimize this trade-off, perhaps through more sophisticated training strategies or by developing adaptive mechanisms that minimize the impact on the modelâs general capabilities while maintaining robust security measures. These limitations point to important directions for future research in developing more balanced approaches to LLM security that can maintain both strong protection against jailbreak attacks and high-quality performance on legitimate tasks. References Alon and Kamfonas (2023) Gabriel Alon and Michael Kamfonas. 2023. Detecting language model attacks with perplexity. arXiv preprint arXiv:2308.14132. Andriushchenko et al. (2025) Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2025. Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. In ICLR. Cao et al. (2023) Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023. Defending against alignment-breaking attacks via robustly aligned LLM. arXiv preprint arXiv:2309.14348. Chang et al. (2024) Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu. 2024. Play guessing game with LLM: Indirect jailbreak attack with implicit clues. In ACL, pages 5135â5147. Chao et al. (2024) Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. 2024. JailbreakBench: An open robustness benchmark for jailbreaking large language models. In NeurIPS. Chao et al. (2023) Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419. Chiang et al. (2023) Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality. Chu et al. (2024) Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. 2024. Comprehensive assessment of jailbreak attacks against LLMs. arXiv preprint arXiv:2402.05668. Deng et al. (2024a) Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024a. MASTERKEY: Automated jailbreaking of large language model chatbots. In NDSS. Deng et al. (2024b) Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2024b. Multilingual jailbreak challenges in large language models. ICLR. Ding et al. (2023) Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. In EMNLP. Handa et al. (2024) Divij Handa, Advait Chirmule, Bimal Gajera, and Chitta Baral. 2024. Jailbreaking proprietary large language models using word substitution cipher. arXiv preprint arXiv:2402.10601. Hu et al. (2022) Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In ICLR. Hu et al. (2024) Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2024. Gradient Cuff: Detecting jailbreak attacks on large language models by exploring refusal loss landscapes. In NeurIPS. Inan et al. (2023) Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv preprint arXiv:2312.06674. Jain et al. (2023) Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614. Ji et al. (2024) Jiabao Ji, Bairu Hou, Alexander Robey, George J Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang. 2024. Defending large language models against jailbreak attacks via semantic smoothing. arXiv preprint arXiv:2402.16192. Jia et al. (2025) Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin. 2025. Improved techniques for optimization-based jailbreaking on large language models. In ICLR. Kumar et al. (2023) Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023. Certifying LLM safety against adversarial prompting. arXiv preprint arXiv:2309.02705. Li et al. (2024a) Jie Li, Yi Liu, Chongyang Liu, Ling Shi, Xiaoning Ren, Yaowen Zheng, Yang Liu, and Yinxing Xue. 2024a. A cross-language investigation into jailbreak attacks in large language models. arXiv preprint arXiv:2401.16765. Li et al. (2024b) Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. 2024b. DrAttack: Prompt decomposition and reconstruction makes powerful LLMs jailbreakers. In EMNLP, pages 13891â13913. Li et al. (2023) Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. AlpacaEval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval. Liao and Sun (2024) Zeyi Liao and Huan Sun. 2024. Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms. arXiv preprint arXiv:2404.07921. Liu et al. (2024a) Fan Liu, Zhao Xu, and Hao Liu. 2024a. Adversarial tuning: Defending against jailbreak attacks for LLMs. arXiv preprint arXiv:2406.06622. Liu et al. (2024b) Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2024b. AutoDAN: Generating stealthy jailbreak prompts on aligned large language models. In ICLR. Liu et al. (2023) Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023. Jailbreaking ChatGPT via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860. Loshchilov (2019) I Loshchilov. 2019. Decoupled weight decay regularization. In ICLR. Lu et al. (2024) Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen. 2024. Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge. arXiv preprint arXiv:2404.05880. Madry (2018) Aleksander Madry. 2018. Towards deep learning models resistant to adversarial attacks. In ICLR. Mehrotra et al. (2024) Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. 2024. Tree of attacks: Jailbreaking black-box LLMs automatically. In NeurIPS. Mo et al. (2024) Yichuan Mo, Yuji Wang, Zeming Wei, and Yisen Wang. 2024. Fight back against jailbreaking via prompt adversarial tuning. arXiv preprint arXiv:2402.06255. Paulus et al. (2024) Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian. 2024. AdvPrompter: Fast adaptive adversarial prompting for LLMs. arXiv preprint arXiv:2404.16873. Perez et al. (2022) Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. In EMNLP. Phute et al. (2023) Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau. 2023. LLM Self Defense: By self examination, LLMs know they are being tricked. arXiv preprint arXiv:2308.07308. Qian et al. (2024) Cheng Qian, Hainan Zhang, Lei Sha, and Zhiming Zheng. 2024. Hsf: Defending against jailbreak attacks with hidden state filtering. arXiv preprint arXiv:2409.03788. Robey et al. (2023) Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. 2023. SmoothLLM: defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684. Shen et al. (2024a) Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024a. The language barrier: Dissecting safety challenges of LLMs in multilingual contexts. In ACL, pages 2668â2680. Shen et al. (2024b) Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024b. "Do Anything Now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In CCS. Sitawarin et al. (2024) Chawin Sitawarin, Norman Mu, David Wagner, and Alexandre Araujo. 2024. PAL: Proxy-guided black-box attack on large language models. arXiv preprint arXiv:2402.09674. Siththaranjan et al. (2023) Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell. 2023. Understanding hidden context in preference learning: Consequences for rlhf. In Workshop on Socially Responsible Language Modelling Research at NeurIPS. Touvron et al. (2023) Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Wang et al. (2025) Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma, Shuai Wang, Yingjiu Li, Yang Liu, Ning Liu, and Juergen Rahmel. 2025. Selfdefend: Llms can defend themselves against jailbreaking in a practical manner. In USENIX Security. Wei et al. (2023a) Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023a. Jailbroken: How does LLM safety training fail? In NeurIPS, volume 36. Wei et al. (2023b) Zeming Wei, Yifei Wang, and Yisen Wang. 2023b. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387. Xhonneux et al. (2024) Sophie Xhonneux, Alessandro Sordoni, Stephan GĂźnnemann, Gauthier Gidel, and Leo Schwinn. 2024. Efficient adversarial training in LLMs with continuous attacks. In NeurIPS. Xie et al. (2024) Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong. 2024. GradSafe: Detecting unsafe prompts for LLMs via safety-critical gradient analysis. In ACL. Xie et al. (2023) Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023. Defending chatgpt against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12):1486â1496. Xiong et al. (2024) Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho. 2024. Defensive prompt patch: A robust and interpretable defense of LLMs against jailbreak attacks. arXiv preprint arXiv:2405.20099. Xu et al. (2024) Zhao Xu, Fan Liu, and Hao Liu. 2024. Bag of tricks: Benchmarking of jailbreak attacks on llms. In NeurIPS. Yong et al. (2023) Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023. Low-resource languages jailbreak GPT-4. arXiv preprint arXiv:2310.02446. Yu et al. (2024) Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2024. LLM-Fuzzer: Scaling assessment of large language model jailbreaks. In USENIX Security, pages 4657â4674. Yuan et al. (2024) Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2024. GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. In ICLR. Zhang et al. (2024a) Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao. 2024a. Intention analysis prompting makes large language models a good jailbreak defender. arXiv preprint arXiv:2401.06561. Zhang et al. (2024b) Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, and Minlie Huang. 2024b. Defending large language models against jailbreaking attacks through goal prioritization. In ACL. Zhao et al. (2023) Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223. Zhao et al. (2024) Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun. 2024. Defending large language models against jailbreak attacks via layer-specific editing. In EMNLP. Zheng et al. (2023) Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS, pages 46595â46623. Zhou et al. (2024) Andy Zhou, Bo Li, and Haohan Wang. 2024. Robust prompt optimization for defending language models against jailbreaking attacks. In NeurIPS. Zou et al. (2023) Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. Appendix A Results under ASRAgent Table 6 shows the ASRAgent results of our STShield and other defense methods against adaptive attacks on Llama-2-7B-Chat. Table 6: Jailbreak attack experiments on dataset AdvBench under ASRAgentAgent_Agentstart_FLOATSUBSCRIPT Agent end_FLOATSUBSCRIPT. Defense Methods Jailbreak Methods AmpleGCG AdvPrompter PAIR TAP LLM-Fuzzer No Defense 50.00 20.00 6.00 12.00 22.00 Self-Reminder 6.00 4.00 4.00 0.00 8.00 RPO 10.00 2.00 6.00 6.00 18.00 Adv. Training 44.00 20.00 8.00 4.00 26.00 Unlearning 52.00 20.00 8.00 6.00 8.00 Safety Training 50.00 22.00 4.00 8.00 30.00 SmoothLLM 14.00 8.00 8.00 20.00 4.00 Llama Guard 26.00 22.00 8.00 2.00 10.00 STShield (ours) 0.00 0.00 0.00 2.00 6.00