Paper deep dive
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
Peiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun, Qin Xia, Zijiang Yang
Models: LLaVA-1.5-7B, MiniGPTv2-7B, Qwen2-VL-7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/12/2026, 5:26:10 PM
Summary
The paper introduces 'The Safety Reminder', a soft prompt tuning approach designed to mitigate 'delayed safety awareness' in Vision-Language Models (VLMs). This phenomenon describes how VLMs may initially generate harmful content under attack but eventually recognize the risk and attempt self-correction. By optimizing learnable soft prompt tokens that are injected during generation when harmful content is detected, the method reactivates the model's safety mechanisms without requiring resource-intensive fine-tuning or compromising utility on benign tasks.
Entities (5)
Relation Signals (3)
VLMs â exhibits â Delayed Safety Awareness
confidence 98% ¡ we identify a novel phenomenon termed âdelayed safety awarenessâ in VLMs
The Safety Reminder â addresses â Delayed Safety Awareness
confidence 95% ¡ The Safety Reminder... a soft prompt tuning approach to reactivate delayed safety awareness
SAPT â utilizes â Safety State Detector
confidence 90% ¡ If the detector identifies the current generation as unsafe, the soft prompt is appended
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As Vision-Language Models (VLMs) demonstrate increasing capabilities across real-world applications such as code generation and chatbot assistance, ensuring their safety has become paramount. Unlike traditional Large Language Models (LLMs), VLMs face unique vulnerabilities due to their multimodal nature, allowing adversaries to modify visual or textual inputs to bypass safety guardrails and trigger the generation of harmful content. Through systematic analysis of VLM behavior under attack, we identify a novel phenomenon termed ``delayed safety awareness''. Specifically, we observe that safety-aligned VLMs may initially be compromised to produce harmful content, but eventually recognize the associated risks and attempt to self-correct. This pattern suggests that VLMs retain their underlying safety awareness but experience a temporal delay in their activation. Building on this insight, we hypothesize that VLMs' safety awareness can be proactively reactivated through carefully designed prompts. To this end, we introduce ``The Safety Reminder'', a soft prompt tuning approach that optimizes learnable prompt tokens, which are periodically injected during the text generation process to enhance safety awareness, effectively preventing harmful content generation. Additionally, our safety reminder only activates when harmful content is detected, leaving normal conversations unaffected and preserving the model's performance on benign tasks. Through comprehensive evaluation across three established safety benchmarks and one adversarial attacks, we demonstrate that our approach significantly reduces attack success rates while maintaining model utility, offering a practical solution for deploying safer VLMs in real-world applications.
Tags
Links
- Source: https://arxiv.org/abs/2506.15734
- Canonical: https://arxiv.org/abs/2506.15734
Trouble viewing inline? Open PDF directly â
Full Text
68,384 characters extracted from source content.
Expand or collapse full text
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models Peiyuan Tang1 Haojie Xin1â footnotemark: Xiaodong Zhang2 Jun Sun3 Qin Xia1 Zijiang Yang2â footnotemark: 1School of Computer Science and Technology, Xiâan Jiaotong University 2School of Computer Science and Technology, University of Science and Technology of China 3School of Computing and Information Systems, Singapore Management University tangpeiyuan, pinkman@stu.xjtu.edu.cn zhangxiaodong, zijiang@ustc.edu.cn Equal contribution.Corresponding author. Abstract As Vision-Language Models (VLMs) demonstrate increasing capabilities across real-world applications such as code generation and chatbot assistance, ensuring their safety has become paramount. Unlike traditional Large Language Models (LLMs), VLMs face unique vulnerabilities due to their multimodal nature, allowing adversaries to modify visual or textual inputs to bypass safety guardrails and trigger the generation of harmful content. Through systematic analysis of VLM behavior under attack, we identify a novel phenomenon termed âdelayed safety awarenessâ. Specifically, we observe that safety-aligned VLMs may initially be compromised to produce harmful content, but eventually recognize the associated risks and attempt to self-correct. This pattern suggests that VLMs retain their underlying safety awareness but experience a temporal delay in their activation. Building on this insight, we hypothesize that VLMsâ safety awareness can be proactively reactivated through carefully designed prompts. To this end, we introduce âThe Safety Reminderâ, a soft prompt tuning approach that optimizes learnable prompt tokens, which are periodically injected during the text generation process to enhance safety awareness, effectively preventing harmful content generation. Additionally, our safety reminder only activates when harmful content is detected, leaving normal conversations unaffected and preserving the modelâs performance on benign tasks. Through comprehensive evaluation across three established safety benchmarks and one adversarial attacks, we demonstrate that our approach significantly reduces attack success rates while maintaining model utility, offering a practical solution for deploying safer VLMs in real-world applications. 1 Introduction Recent advancements in integrating language and visual promots have significantly expanded the capabilities of Large Language Models (LLMs), yielding promising results across a wide range of tasks Achiam et al. (2023); Deng et al. (2024); Liu et al. (2024b). However, this progress also raises concerns, as the continuous and high-dimensional nature of visual inputs provides attackers with a broader spectrum of potential adversarial targets. Similar to the existing approaches (Zheng et al., 2024a) employed in LLMs, a common and lightweight approach to aligning Vision-Language Models (VLMs) involves using handcrafted safety prompts as input. These prompts typically include instructions to avoid harmful queries and provide explicit security guidance without modifying the modelâs parameters. Figure 1: Illustration of delayed safety awareness in VLMs under jailbreak attacks. The model initially complies with the adversarial prompt and generates unsafe content, but eventually recognizes the malicious intent and highlights the associated risks. However, there remains an unclear understanding of the safety mechanisms of VLMs, which limits our ability to optimize safety prompts and further enhance VLM safety. As illustrated in Figure 1, the VLMsâ capability to detect harmful queries Zheng et al. (2024b) is analogous to a security guard recognizing a breach only after the intruder has escaped, i.e., their delayed awareness renders safety measures ineffective, as the harmful output has already been generated, resulting in irreversible consequences. Inspired by this problem, our work investigates two key questions: (1) when does the model realize it has generated a harmful response, and (2) how we can design effective safety measures based on this behavior. We conjecture that: the modelâs awareness of harmful queries is not immediate under the jailbreak attacks but develops gradually during the generation process. To validate our conjecture, we evaluate two open-source VLMs with safety alignment on two datasets that include harmful queries. We monitor the frequency of changes in the relative positions of refusal signals, as shown in Figure 2. Our findings reveal that as the output length increases, the frequency distribution of rejection replies progressively shifts towards the end of the generated text, which is consistent with our conjecture. This behavior stems from two primary factors. Firstly, the autoregressive nature of VLMs means that once a harmful response is initiated, the model tends to continue generating content along the same harmful direction. On the other hand, while language bridges visual understanding, it can also introduce bias, constraining VLMsâ ability to identify harmful visual content initially. However, during the output generation process, some of the associated contextual information from the image transfers to the text, gradually enabling VLMs to recognize their harmful output. This delayed awareness highlights the need for robust safety mechanisms in VLMs. Previous methods on safeguarding VLMs mainly relies on safety fine-tuning (Zong et al., 2024a; Chen et al., 2024) and content filtering (Gou et al., 2024; Fares et al., 2024; Ding et al., 2024) to defend against jailbreak attacks. However, safety fine-tuning risks catastrophic forgetting Zhai et al. (2024); Zhou et al. (2025), where the model loses proficiency in benign tasks, and often introduces conflicting objectives Wang et al. (2024b); Pan et al. (2024) between safety and task performance. While fine-tuning on harmful datasets can enhance model safety, it may lead to decreased utility. On the other hand, content filtering is unreliable because it relies on the modelâs own judgment or other models to detect malicious inputs and outputs. Moreover, both approaches exhibit limited resistance to adversarial perturbation attacks, as their safety awareness is easily circumvented. Inspired by our findings, we propose Safety-Aware Soft Prompt Tuning (SAPT) to ensure real-time safety during text generation. SAPT optimizes learnable soft prompts appended to the generated text, allowing the model to reactivate its safety mechanisms dynamically, particularly when harmful content is detected by the safety state detector. Unlike existing approaches that require time-consuming fine-tuning or complex filtering mechanisms, SAPT is lightweight, parameter-efficient, and avoids introducing unnecessary trade-offs. This enables efficient real-time security checks while preserving the modelâs functionality and performance. The main contributions of this work are threefold: ⢠We identify a novel phenomenon termed âdelayed safety awarenessâ, where VLMs are initially compromised by malicious inputs but eventually recognize the associated risks and attempt self-correction. ⢠We propose SAPT, a soft prompt tuning approach that optimizes learnable prompt tokens and periodically injects them during text generation to proactively reactivate the modelâs safety awareness. ⢠We conduct extensive experiments across three established safety benchmarks and one adversarial attack method, demonstrating the effectiveness of our approach in defending against jailbreak attacks while minimizing utility loss. Figure 2: Visualization of the relative positions of the two models in expressing safety tendencies on two harmful datasets. As the modelâs maximum response length increases, the relative position distribution of rejected responses peaks at the tail. This pattern indicates that while the model can effectively identify harmful queries, it tends to generate safe responses toward the end of the text sequence. 2 Related Work 2.1 Vision Language Models (VLMs) VLMs improve traditional Large Language Models (LLMs) by integrating image encoders Dosovitskiy (2020) and projectors. The image encoder extracts visual features, which are then aligned with text modalities through the projector. These aligned features are concatenated with text embeddings, allowing the LLM to generate responses grounded by the image. The remarkable ability of VLMs to integrate image and text understanding has garnered significant interest from both academia and industry, driving advancements in tasks such as image captioning Yang et al. (2024b); Li et al. (2023), visual question answering Alayrac et al. (2022); Chen et al. (2022), and multimodal conversational chatbots Liu et al. (2024b); Deng et al. (2024). In this study, we study the vulnerabilities and safety implications of this emerging multimodal paradigm. 2.2 Jailbreaking VLMs Although VLMs have demonstrated exceptional performance and significant societal impact, the integration of the image modality introduces heightened security risks, making them more susceptible to malicious visual inputs compared to LLMs Gou et al. (2024); Liu et al. (2024c). One notable attack involves embedding malicious text queries into images using jailbreak templates, effectively bypassing the safety mechanisms of VLMs Gong et al. (2023); Liu et al. (2023b). Alternatively, adversaries may exploit gradient-based optimization to create adversarial images by introducing imperceptible perturbations or patches Schlarmann and Hein (2023); Niu et al. (2024); Zhao et al. (2024); Carlini et al. (2024). These security threats present critical challenges to the practical deployment of VLMs, underscoring the importance of ensuring their robustness and reliability. 2.3 Safeguarding VLMs Several approaches have been proposed to enhance the safety of VLMs, which can be primarily categorized into two groups: training-based methods (Zong et al., 2024b; Chen et al., 2024) and post-hoc defenses (Xu et al., 2024; Ding et al., 2024; Fares et al., 2024). Training-based methods focus on building secure models through techniques such as adversarial training and reinforcement learning from human feedback (RLHF) (Bai et al., 2022). In contrast, post-hoc defenses aim to improve the safety of VLMs during inference. This includes strategies such as input preprocessing (Ding et al., 2024), harmful output detection (Gou et al., 2024), and safety prompts (Wang et al., 2024c) to improve the overall safety of the models. In summary, while both types of methods contribute to improving the safety of VLMs, training-based approaches tend to be resource-intensive, whereas post-hoc defenses provide a more adaptable and cost-effective solution for real-world applications. Unlike prior work, we introduce a novel post-hoc defense method that enhances VLM safety by optimizing a learnable soft prompt. Our approach achieves significant safety improvements without compromising the utility of the VLMs. 3 Methodology 3.1 Preliminaries Visual Language Models (VLMs) Liu et al. (2024a); Chen et al. (2023) are multimodal models that can process and understand text and images and generate text responses. A typical VLM consists of three main components: an image encoder that processes input images, a projector that aligns the image features with text embeddings, and a Large Language Model (LLM) that generates text responses in an autoregressive manner. Formally, let I and xtxtsubscripttxtx_txtxtxt denote the image and text input, respectively. The image encoder and projector are represented as θsubscriptV_θVitalic_θ and Ďsubscriptitalic-ĎW_ĎWitalic_Ď. The probability of the output sequence y is given by: ximg=Ď(θ(I)))x_img=W_Ď(V_θ(I)))ximg = Witalic_Ď ( Vitalic_θ ( I ) ) ) (1) Ďθâ˘(yâŁximg,xtxt)=ât=1NĎθâ˘(ytâŁximg,xtxt,y<t)subscriptconditionalsubscriptimgsubscripttxtsuperscriptsubscriptproduct1subscriptconditionalsubscriptsubscriptimgsubscripttxtsubscriptabsent _θ(y x_img,x_txt)= _t=1^N _θ(% y_t x_img,x_txt,y_<t)Ďitalic_θ ( y ⣠ximg , xtxt ) = ât = 1N Ďitalic_θ ( yitalic_t ⣠ximg , xtxt , y< t ) (2) where ximgsubscriptimgx_imgximg represents the aligned image features, y denotes the full generated token sequence, and y<tsubscriptabsenty_<ty< t denotes the sequence of tokens generated before ytsubscripty_tyitalic_t. Jailbreak in VLMs (Ying et al., 2024; Li et al., 2024b) refers to the manipulation of input text queries or images in order to bypass the modelâs safety guardrails and induce the model to produce harmful, unethical, or biased content. The attackerâs objective is to find a perturbed input x^ xover start_ARG x end_ARG that maximizes the probability of producing a harmful output sequence y^ yover start_ARG y end_ARG: x~~ xover~ start_ARG x end_ARG =argâĄmaxx~âAâ˘(x)âĄĎθâ˘(y~|x~)absentsubscript~subscriptconditional~~ = _ xâ A(x) _θ( y| x)= arg maxover~ start_ARG x end_ARG â A ( x ) Ďitalic_θ ( over~ start_ARG y end_ARG | over~ start_ARG x end_ARG ) (3) =argâĄmaxx~âAâ˘(x)â˘ât=1TĎθâ˘(y~t|y~<t,x~)absentsubscript~superscriptsubscriptproduct1subscriptconditionalsubscript~subscript~absent~ = _ xâ A(x) _t=1^T _θ( y% _t| y_<t, x)= arg maxover~ start_ARG x end_ARG â A ( x ) ât = 1T Ďitalic_θ ( over~ start_ARG y end_ARGt | over~ start_ARG y end_ARG< t , over~ start_ARG x end_ARG ) where x~~ xover~ start_ARG x end_ARG represents the modified input to the model, which can be either text, image, or both modified text and image, Aâ˘(x)A(x)A ( x ) represents the set of all possible perturbations to the original input x, y~Tsubscript~ y_Tover~ start_ARG y end_ARGT denotes the harmful output that the attacker expects the model to generate. A typical attack starts by prompting the model to generate a harmful prefix, such as âSure, here is how to build a bomb.â The model then continues from this prefix, generating harmful content and leading to a successful jailbreak. 3.2 Our Key Insights Before introducing our method, we first summarize the key insights that motivate our approach to enhancing VLM safety and understanding jailbreak vulnerabilities: ⢠Autoregressive Generation Facilitates Jailbreak Attack. The inherent autoregressive nature of text generation in VLMs creates a fundamental vulnerability to jailbreak attacks. Once the model begins generating harmful content, the autoregressive objective pâ˘(yt|y<t,x)conditionalsubscriptsubscriptabsentp(y_t|y_<t,x)p ( yitalic_t | y< t , x ) drives the model to maintain consistency with the harmful prefix rather than refusing to continue, as maximizing token likelihood favors completion over rejection. ⢠Training-Time Alignment Cannot Guarantee Runtime Safety. Recent work Ghosal et al. (2025) shows that jailbreak attacks can be reformulated as an âinverse-alignmentâ problem. Specifically, adversaries search for inputs x~~ xover~ start_ARG x end_ARG that maximize the unsafe reward for the VLM model. The authors prove that for any model aligned during training, there always exists a closed-form adversarial prompt distribution that induces unsafe behavior at inference time. ⢠Delayed Safety Awareness Enables Proactive Defense. Our findings reveal that VLMs exhibit delayed safety awarenessâinitially complying with malicious prompts before recognizing harmful content during generation. This delayed recognition pattern suggests that safety mechanisms may be strategically activated during the generation process, providing opportunities for real-time intervention. Figure 3: Overview of our proposed method. (a) A learnable soft prompt is optimized and inserted after the harmful prefix to reactivate the safety awareness of VLMs. (b) To enable the model to effectively distinguish between harmful and benign queries, we train a Safety State Detector using the modelâs final-layer hidden states. (c) During text generation, we periodically apply the soft prompt. If the detector identifies the current generation as unsafe, the soft prompt is appended to the generated text. For brevity, we omit the benign query training in (a). 3.3 Safety-Aware Soft Prompt Tuning The overall design of our approach is illustrated in Figure 3. Our core idea is to optimize a universal soft prompt that proactively activates the safety awareness of a VLM when harmful content begins to emerge. During training, we simulate early-stage jailbreak attacks by appending an incomplete harmful response to user queries111More precisely, the incomplete harmful response is appended after a special token such as [/INST] that separates the user instruction from the model output, following the common chat-style input format.. The soft prompt is then inserted, interrupting the autoregressive generation and prompting the model to reassess the safety of its output. Once optimized, the soft prompt can be dynamically injected at any position during text generation to enforce real-time safety. To minimize utility loss, we also train it on benign queries to ensure the model continues producing helpful and coherent responses in safe contexts. Optimization for Malicious Query. Given a malicious query xqm=xtxtm,ximgmsuperscriptsubscriptsuperscriptsubscripttxtsuperscriptsubscriptimgx_q^m=\x_txt^m,x_img^m\xitalic_qitalic_m = xtxtitalic_m , ximgitalic_m , where xtxtmsuperscriptsubscripttxtx_txt^mxtxtitalic_m and ximgmsuperscriptsubscriptimgx_img^mximgitalic_m denote the malicious text and image inputs, respectively, we simulate an attack by prefilling the incomplete harmful response xrmsuperscriptsubscriptx_r^mxitalic_ritalic_m. This increases the likelihood of the model completing the harmful prefix, potentially leading to a jailbreak. To safeguard the model, we append the soft prompt after the harmful response to halt generation and trigger safety reassessment, leading to refusal when harmful intent is recognized. We apply loss computation exclusively to the safe response tokens and update only the soft prompt parameters: Lm=âlogâĄĎθâ˘(yâŁxqmâxrmâxs)subscriptsubscriptconditionaldirect-sumsuperscriptsubscriptsuperscriptsubscriptsubscriptL_m=- _θ(y x_q^m x_r^m x_s)Litalic_m = - log Ďitalic_θ ( y ⣠xitalic_qitalic_m â xitalic_ritalic_m â xitalic_s ) (4) where y is the safe response, xrmsuperscriptsubscriptx_r^mxitalic_ritalic_m is the incomplete harmful response, xssubscriptx_sxitalic_s is the soft prompt we aim to optimize, âdirect-sum â is the concatenation operation. Optimization for Benign Query. Given a benign query xqb=xtxtb,ximgbsuperscriptsubscriptsuperscriptsubscripttxtsuperscriptsubscriptimgx_q^b=\x_txt^b,x_img^b\xitalic_qitalic_b = xtxtitalic_b , ximgitalic_b , We optimize the soft prompt to maintain the modelâs general performance. The loss is computed only for the part after the soft prompt. This is formalized as: Lb=âlogâĄĎθâ˘(yâ˛âŁxqbâyâ¤kâxs)subscriptsubscriptconditionalsuperscriptâ˛direct-sumsuperscriptsubscriptsubscriptabsentsubscriptL_b=- _θ(y x_q^b y_⤠k x_s)Litalic_b = - log Ďitalic_θ ( yⲠ⣠xitalic_qitalic_b â y⤠k â xitalic_s ) (5) where yâ¤ksubscriptabsenty_⤠ky⤠k is the prefix of the reference answer, k is selected randomly, and yⲠrepresents the remaining portion of the answer (i.e.,yâ˛=y>ksuperscriptâ˛subscriptabsenty =y_>kyⲠ= y> k ). 3.4 Safety State Detector We observe that the first token generated after a soft prompt significantly influences the modelâs behavior. For instance, when the model generates tokens such as "I" or "As," it often defaults to refusal responses (e.g., "I am sorryâŚ" or "As an AI assistant, I cannotâŚ"), regardless of whether the input is truly harmful or benign. We attribute this to insufficient supervision of the modelâs internal hidden states, which fail to distinguish between benign and malicious queries effectively. To address this, we introduce a Safety State Detector, a classifier trained on the modelâs hidden representations. This detector serves two purposes: first, it enables the model to learn more distinguishable hidden states for safe and unsafe inputs; second, it enables the detection of potentially harmful queries during inference, allowing us to determine when to inject the soft prompt. We implement the detector as a logistic regression classifier operating on the hidden states, with the objective formulated as follows: y^=Ďâ˘(â¤â˘j+b)^superscripttopsubscript y=Ď(w h_j+b)over start_ARG y end_ARG = Ď ( w⤠hitalic_j + b ) (6) Lcls=â[yâ logâĄy^+(1ây)â logâĄ(1ây^)]subscriptclsdelimited-[]â ^â 11^L_cls=- [y¡ y+(1-y)¡ (1- y) ]Lcls = - [ y â log over start_ARG y end_ARG + ( 1 - y ) â log ( 1 - over start_ARG y end_ARG ) ] (7) where yâ0,101yâ\0,1\y â 0 , 1 is the ground truth label, Ďâ˘(â )â Ď(¡)Ď ( â ) is the sigmoid function, ww is the classifier weight vector with bias term b, and jsubscripth_jhitalic_j is the hidden state extracted from the j-th layer of the VLM. Following (Zheng et al., 2024b), we extract the hidden states from the final layer at the position of the last soft prompt. Total Loss. We train the soft prompt and the safety state detector simultaneously. The total loss is formulated as follows: Ltotal=Lm+Îąâ˘Lb+βâ˘LclssubscripttotalsubscriptsubscriptsubscriptclsL_total=L_m+Îą L_b+β L_clsLtotal = Litalic_m + Îą Litalic_b + β Lcls (8) where Îą and β control the relative importance of the benign loss and classification loss, respectively. 3.5 Dynamic Soft Prompt Injection We employ a dynamic intervention strategy to apply the optimized soft prompt periodically during generation. Specifically, after every k tokens are generated, we utilize our pre-trained safety-state detector to analyze the modelâs current hidden state. If the detector classifies the state as harmful with a probability exceeding a predefined threshold θ, we append the soft prompt to the current text sequence to steer the generation towards a safe direction. Otherwise, no action is taken, and the model continues its generation uninterrupted. This on-demand application ensures that the prompt does not degrade model performance or utility on benign content, as it is only introduced when necessary. 4 Experiments This section first details the experimental setup and then evaluates our method against the baselines to demonstrate its effectiveness. 4.1 Experimental Setups Evaluation Datasets and Attack Methods. For jailbreak evaluation, we selected commonly used harmful benchmarks: FigStep (Gong et al., 2023), MMSafetyBench (Liu et al., 2023a), and VLSafe (Chen et al., 2024) to assess the safety of VLMs. FigStep and MMSafetyBench datasets contain harmful content embedded in images, while their text queries are completely benign. In contrast, the harmful content in the VLSafe is explicitly contained in the text queries. We also investigated the modelâs vulnerability to optimization-based Visual Adversarial Attacks (Qi et al., 2023), where an adversary manipulates a subtle perturbation in the image to induce harmful content generation. We optimized this perturbation using the LâsubscriptL_âLâ norm with constraints of Ďľitalic-ϾξϾ = 32/255, 64/255, and 128/255 on the AdvBench (Zou et al., 2023) dataset. Models. We validate the effectiveness of our method on three VLMs with different architectures and LLM backbones: LLaVA-1.5-7B (Liu et al., 2024a) based on Vicuna (Chiang et al., 2023), MiniGPTv2-7B (Chen et al., 2023) with Llama Touvron et al. (2023), and Qwen2-VL-7B (Wang et al., 2024a) built on Qwen2 (Yang et al., 2024a). Baseline Methods. We compare SAPT with two baselines, including Adashield (Wang et al., 2024c), and PromptTuning (Lester et al., 2021a). Adashield (Wang et al., 2024c) enhances model safety by adaptively appending defense prompts to input queries. It selects the most suitable prompt from a predefined pool based on the queryâs specific features and associated risks. Prompt Tuning (Lester et al., 2021b) learns a soft safety prompt, which is prepended to the input, enabling the model to reject harmful queries while responding normally to harmless ones. We train this soft prompt on our constructed dataset for a fair comparison. Evaluation Metrics. To evaluate the effectiveness of our jailbreak defense, we focus on two key aspects: safety and utility. For safety, we adopt the Attack Success Rate (ASR) as the primary metric. ASR measures the percentage of attack attempts that successfully elicit responses from the model aligned with the attackerâs goals (e.g., illegal activities). ASR=|responses aligned with attackerâs goal||all responses|ASRresponses aligned with attackerâs goalall responsesASR= |\ responses aligned with attacker's goal\ |% |\ all responses\ |ASR = divide start_ARG | responses aligned with attackerâs goal | end_ARG start_ARG | all responses | end_ARG A low ASR indicates higher safety against jailbreak attacks. In this study, we employ MD-Judge-v0.2-internlm2-7B Li et al. (2024a) to assess whether the modelâs response aligns with the attackerâs goal, leveraging its high accuracy and human-readable safety judgments. For utility, we use the M-Vet benchmark (Yu et al., 2023) to evaluate multimodal capabilities after applying the jailbreak defense. Additionally, we calculated the refusal ratio for benign queries to evaluate the modelâs oversensitivity. The refusal ratio is defined as the proportion of rejected normal requests, identified through keyword-based detection, following (Zou et al., 2023). Implementation Details. We optimized a separate soft prompt for each individual model. The soft prompts, with a token length of 4, were trained for 16K iterations with a learning rate of 1e-4 and a batch size of 4 on our constructed dataset. The loss weight β is set to 0.2. During inference, we use greedy decoding to ensure reproducibility. The threshold for our trained safety state detector is set at 0.9, and the modelâs internal safety state is monitored every 16 tokens. Model Defense Harmful Benchmarks â â Visual Adversarial Attack â â FigStep MMSafety VLSafe Avg. Ďľ=bold-italic-Ďľ32255 Îľ= 32255italic_Ďľ bold_= divide start_ARG 32 end_ARG start_ARG 255 end_ARG Ďľ=64255italic-Ďľ64255Îľ= 64255Ďľ = divide start_ARG 64 end_ARG start_ARG 255 end_ARG Ďľ=128255italic-Ďľ128255Îľ= 128255Ďľ = divide start_ARG 128 end_ARG start_ARG 255 end_ARG Avg. LLaVA-1.5-7B No Defense 79.71 72.12 76.33 76.05 87.50 96.92 98.08 94.17 AdaShield 78.86 41.26 31.67 50.60 88.65 91.54 95.38 91.86 Prompt Tuning 3.71 6.38 1.67 3.92 53.08 62.69 69.61 61.79 SAPT (Ours) 1.43 3.91 4.33 3.22 0.96 6.73 3.08 3.59 MiniGPTv2-7B No Defense 21.43 30.97 2.33 18.24 84.61 93.46 97.12 91.73 AdaShield 0.00 0.00 0.00 0.00 38.46 57.50 56.35 50.77 Prompt Tuning 0.00 0.41 0.33 0.25 21.73 22.11 23.08 22.31 SAPT (Ours) 0.00 0.62 0.33 0.32 5.19 3.85 7.11 5.38 Qwen2-VL-7B No Defense 34.00 23.46 2.33 19.93 91.54 96.35 98.27 95.39 AdaShield 9.71 7.00 0.00 5.57 10.58 26.73 21.92 19.74 Prompt Tuning 0.00 0.21 0.33 0.18 48.08 53.65 57.11 52.95 SAPT (Ours) 0.00 0.00 2.00 0.67 5.77 7.69 10.19 7.88 Table 1: Comparison of Attack Success Rate (ASR) across various safety benchmarks under multiple jailbreak attack settings. Lower values indicate better safety performance. Our method significantly reduces ASR and outperforms the baselines in most scenarios. Model Defense Multimodal Capabilities â â Refusal Rate â â Recognize OCR Knowledge Generation Spatial Math Avg. LLaVa-1.5-7B No Defense 36.7 21.9 16.9 19.4 24.8 7.7 31.4 0.46 AdaShield 29.9 13.5 12.9 19.7 22.7 0.0 22.8 29.36 Prompt Tuning 41.4 24.1 26.4 27.6 28.5 11.2 35.5 3.21 SAPT (Ours) 36.5 21.1 16.7 19.0 24.7 7.7 31.0 1.83 MiniGPTv2-7B No Defense 13.1 5.6 11.7 8.1 8.5 0.0 10.5 11.92 AdaShield 4.0 5.2 0.0 2.5 5.3 0.0 3.7 75.87 Prompt Tuning 13.9 7.0 11.4 8.2 10.3 0.0 11.0 19.27 SAPT (Ours) 12.9 6.8 11.1 7.1 10.0 0.0 10.5 14.68 Qwen2-VL-7B No Defense 56.3 63.4 43.7 47.0 59.1 53.5 59.0 1.38 AdaShield 43.7 58.6 25.6 31.1 60.1 52.7 48.5 26.15 Prompt Tuning 35.5 19.5 19.9 17.7 22.3 11.2 31.3 17.43 SAPT (Ours) 51.7 58.9 39.2 38.5 56.5 49.6 55.8 4.59 Table 2: Comparison of general performance on the M-Vet (Yu et al., 2023) dataset after applying the defense. Our method minimizes degradation and slightly increases the refusal rate. Model Accuracy Precision Recall F1 Score LLaVa-1.5-7B 92.4 87.8 98.5 92.8 MiniGPTv2-7B 90.2 87.3 94.2 90.6 Qwen2-VL-7B 93.2 90.8 96.2 93.4 Table 3: Jailbreak prompt detection results using our trained safety state detector. Configs FigStep MMSafety Adv. Img MMVet Baseline 79.7 72.1 98.1 31.4 w/o âbsubscriptâL_bLitalic_b 0.0 1.0 0.0 28.7 w/o âcâ˘lâ˘ssubscriptâL_clsLitalic_c l s 9.1 14.0 7.3 30.3 Full 1.4 3.9 3.1 31.0 Table 4: Ablation study of the loss design for the LLaVa-1.5-7B model. Figure 4: Visualize the relative positions of the two models in expressing the safety trends of the two harmful datasets after adding soft promots. Compared with Figure 2, after adding soft promotions to the model for safety alignment through SAPT, the modelâs safety awareness is advanced. 4.2 Main Results Defense Effectiveness. Table 1 presents the experimental results of our method compared to baselines across various jailbreak attacks on different VLMs. We find that VLMs are vulnerable to jailbreak attacks, particularly visual adversarial attacks, which achieve a success rate of over 90%, posing significant risks for deployment without adequate safeguards. However, SAPT effectively mitigates these threats, significantly reducing the Attack Success Rate (ASR). For example, it lowers the average ASR for LLaVA-1.5-7B from 76.05% to 3.22% on harmful benchmarks and from 94.17% to 3.59% against visual adversarial attacks, demonstrating its effectiveness. Compared to the baselines, our method achieves the best performance on visual adversarial attacks while maintaining competitive results on harmful benchmarks. Another key observation is that AdaShield (Wang et al., 2024c) is particularly effective on strongly aligned VLMs (e.g., MinGPTv2, Qwen2) but performs poorly on weakly aligned models (e.g., LLaVA) and provides limited defense against visual adversarial attacks. Prompt Tuning (Lester et al., 2021b), which is trained on a safety dataset, offers strong defenses on harmful benches but lacks generalization to unseen attacks such as visual adversarial attacks. In contrast, SAPT demonstrates strong performance and superior generalization on both HarmfulBench and visual adversarial attacks. Utility Evaluation. We also conducted experiments to evaluate the multimodal capability of VLMs after applying the defense method. The results are shown in Table 2. Compared to the original model, SAPT nearly maintains the same performance for LLaVA and MiniGPTv2-7B, but leads to approximately 5% performance degradation for the Qwen2-VL model. This is because SAPT may make the model oversensitive, leading to the rejection of normal user queries. A similar observation is found in the baselines AdaShield and PromptTuning. AdaShield significantly increases the rejection rate for normal queries, especially for MiniGPTv2 (from 11.92% to 75.87%). In comparison, our model exhibits the lowest rejection rate among all baselines. Another observation is that PromptTuning can enhance a modelâs utility when its initial performance is low, but it may lead to catastrophic forgetting if the model is already strong. For instance, PromptTuning reduced the performance of Qwen2-VL from 59.0 to 31.3, likely due to overfitting on a small dataset. In contrast, our method maintains the modelâs performance at a high level while avoiding the risk of catastrophic forgetting. Figure 5: Ablation studies on soft prompt length and application frequency during text generation Jailbreak Prompt Detection. We evaluated our safety state detectorâs ability to identify jailbreak attempts on a balanced dataset of 800 samples, comprising 400 benign and 400 harmful prompts randomly selected from the dataset. Our detection method is triggered every 16 tokens during text generation. A prompt is classified as a jailbreak prompt if the classifierâs harmfulness score exceeds a predefined threshold even once during this process. As shown in the Table 3, our classifier has a classification accuracy of over 90%, proving its reliability. The high recall rate indicates a strong ability to detect jailbreak prompts. However, the relatively low precision indicates that some benign prompts are misclassified as jailbreak prompts. We believe this phenomenon may be related to our classification method. During generation, we classify every 16 tokens, and if the output length reaches 256 tokens, it triggers 16 classifications. If any single classification exceeds the harmful threshold, the prompt is considered a jailbreak. This leads to high recall but low precision. Safety Awareness Reactivation. We also tested whether our method can reactivate the modelâs safety awareness when faced with a jailbreak attack. As in the previous experiment, we evaluated the LLaVa and Qwen2-VL models on the FigStep and MMSafetyBench datasets. The original results and those with our method are shown in Figure 2 and Figure 4, respectively. By comparing these figures, we observe that, with our method, the model begins rejecting harmful prompts significantly earlier. This suggests that our approach not only improves overall safety awareness and significantly reduces jailbreak success rates, but also enables earlier detection of potential issues, mitigating additional jailbreak risks 4.3 Ablation Studies. Effect of Loss Design. We sequentially remove the classification loss Lcâ˘lâ˘ssubscriptL_clsLitalic_c l s and the benign sample loss LbsubscriptL_bLitalic_b from the soft prompt training loss to evaluate their impact. Note that we retain the binary cross entropy (BCE) loss to train the classifier to ensure that the inference process works properly. The results of our ablation study on LLaVA-1.5-7B are presented in Table 4. We find that LbsubscriptL_bLitalic_b is essential for maintaining the modelâs utility and preventing it from rejecting all queries. Without LbsubscriptL_bLitalic_b, although the model exhibits nearly 0 ASR, there is a significant drop in utility, highlighting the importance of including benign samples in the training dataset. On the other hand, incorporating the classification loss into the soft prompt training process leads to performance improvements. We hypothesize that adding LclssubscriptclsL_clsLcls to the modelâs hidden states allows it to better distinguish between harmful and benign categories, thereby enhancing its ability to identify and reject harmful prompts effectively. Effect of soft prompt length. We conduct additional experiments to demonstrate the impact of the soft prompt length. The model used in this experiments is LLaVa-1.5-7B. The results are shown in Figure 5. The experimental results show that for jailbreak defense, increasing the soft prompt length initially enhances performance but later degrades it. A prompt length of 2 may fails to converge to an optimal solution. When the length reaches 8 or 16, ASR increases overall, likely due to longer prompts introducing more parameters, leading to overfitting and reduced generalization. As for utility, increasing the prompt length gradually decreases it, as longer prompts introduce more irrelevant tokens, disrupting contextual coherence and ultimately reducing utility. Effect of application frequency. We conduct additional experiments to demonstrate the impact of applying the soft prompt at different intervals. We apply the soft prompt every 8, 16, 32, and 64 tokens. The model used in this experiment is LLaVa-1.5-7B. The results are shown in Figure 5.We observe that applying the soft prompt more frequently results in a lower ASR and reduced utility. This is because increasing the frequency of soft prompt application leads to more frequent classifications by the safety state detector, making it more likely to classify prompts as harmful, while also potentially misclassifying benign prompts. Defense LLaVa Minigptv2 Qwen2-VL No Defense 62.74 103.27 104.49 AdaShield 60.17 96.21 101.35 Prompt Tuning 61.72 101.23 102.40 SAPT 55.93 93.98 95.29 Table 5: Comparison of inference speed measured by the average number of tokens generated per second. 4.4 Inference Efficiency We compare the inference speed of different methods in Table 5 using 200 harmless and 200 harmful samples from the VLSafe (Chen et al., 2024) dataset, calculating the average number of tokens generated per second. Our approach reduces inference speed by only about 10% compared to the No Defense method. This reduction results from calculating hidden states every 16 tokens and leveraging KV Cache to reuse previous token data, minimizing additional inference time. 5 Conclusion This paper aims to defend against jailbreak attacks for VLMs by reactivating its safety awareness during text generation. We first investigate the delayed safety awareness in safety-aligned VLMs. We find that jailbreak attacks do not erase the modelâs safety awareness but delay its activation. Based on this finding, we propose SAPT, a learnable soft prompt that reactivates safety awareness in VLMs during text generation to prevent harmful content. Experimental results across three VLMs, three safety benchmarks, and one adversarial attack demonstrate the effectiveness of our method in safeguarding VLMs while preserving their original multimodal capabilities. 6 Limitations In this work, we propose SAPT as a defense mechanism against jailbreak attacks targeting VLMs. However, we acknowledge that our current method has three limitations. First, SAPT may make VLMs more sensitive, increasing the rejection rate for normal queries. Second, our method relies heavily on the Safety States Detector. In practical scenarios, setting an appropriate threshold for the classifier can be challenging. Additionally, we have not evaluated our method against text-based adversarial attacks, such as GCG (Zou et al., 2023). It remains uncertain whether SAPT is effective against such attacks. We leave this as future work, with plans to expand SAPT to address these limitations and apply it to a broader range of jailbreak attacks. 7 Ethical Considerations Our paper focuses on defending against malicious image queries for vision-language models (VLMs). We believe that our proposed SAPT method offers valuable insights for the development of safer VLM applications in the future. To train our SAPT, we constructed a dataset containing both harmless and harmful responses. We employed the LLaVA-1.5-7B model to generate unsafe responses. However, we emphasize that all toxic data used in this paper was solely for training our soft prompt and will not be publicly available. The data is strictly confined to the modelâs training and testing processes in our research. Additionally, this paper includes some unfiltered harmful examples, which are shown in the appendix, only to illustrate the vulnerabilities of VLMs to jailbreak attacks. We strongly oppose any attempt to jailbreak VLMs. AI Assistants in This Research. We only use GPT-4o mini to assist in polishing our sentences and correcting grammar errors. References Achiam et al. [2023] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. Alayrac et al. [2022] Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716â23736, 2022. Bai et al. [2022] Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022. Carlini et al. [2024] Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt. Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024. Chen et al. [2023] Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478, 2023. Chen et al. [2022] Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022. Chen et al. [2024] Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. Dress: Instructing large vision-language models to align and interact with humans via natural language feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14239â14250, 2024. Chiang et al. [2023] Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6, 2023. Deng et al. [2024] Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. Masterkey: Automated jailbreaking of large language model chatbots. In Proc. ISOC NDSS, 2024. Ding et al. [2024] Yi Ding, Bolian Li, and Ruqi Zhang. Eta: Evaluating then aligning safety of vision language models at inference time. arXiv preprint arXiv:2410.06625, 2024. Dosovitskiy [2020] Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. Fares et al. [2024] Samar Fares, Klea Ziu, Toluwani Aremu, Nikita Durasov, Martin TakĂĄÄ, Pascal Fua, Karthik Nandakumar, and Ivan Laptev. Mirrorcheck: Efficient adversarial defense for vision-language models. arXiv preprint arXiv:2406.09250, 2024. Ghosal et al. [2025] Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan, Mengdi Wang, Ahmad Beirami, Furong Huang, Alvaro Velasquez, Dinesh Manocha, and Amrit Singh Bedi. Immune: Improving safety against jailbreaks in multi-modal llms via inference-time alignment. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), pages 25038â25049, June 2025. Gong et al. [2023] Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. Figstep: Jailbreaking large vision-language models via typographic visual prompts. ArXiv, abs/2311.05608, 2023. URL https://api.semanticscholar.org/CorpusID:265067328. Gou et al. [2024] Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation. In European Conference on Computer Vision, pages 388â404. Springer, 2024. Lester et al. [2021a] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Conference on Empirical Methods in Natural Language Processing, 2021a. URL https://api.semanticscholar.org/CorpusID:233296808. Lester et al. [2021b] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Conference on Empirical Methods in Natural Language Processing, 2021b. URL https://api.semanticscholar.org/CorpusID:233296808. Li et al. [2023] Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730â19742. PMLR, 2023. Li et al. [2024a] Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. arXiv preprint arXiv:2402.05044, 2024a. Li et al. [2024b] Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. Images are achillesâ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models. In European Conference on Computer Vision, pages 174â189. Springer, 2024b. Liu et al. [2024a] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024a. Liu et al. [2024b] Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. Llava-plus: Learning to use tools for creating multimodal agents. In European Conference on Computer Vision, pages 126â142. Springer, 2024b. Liu et al. [2023a] Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision, 2023a. URL https://api.semanticscholar.org/CorpusID:265498692. Liu et al. [2023b] Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao. Query-relevant images jailbreak large multi-modal models. arXiv preprint arXiv:2311.17600, 2023b. Liu et al. [2024c] Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao. Safety of multimodal large language models on images and text. arXiv preprint arXiv:2402.00357, 2024c. Llama Team [2024] AI @ Meta Llama Team. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783. Luo et al. [2024] Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv preprint arXiv:2404.03027, 2024. Madry [2017] Aleksander Madry. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017. Niu et al. [2024] Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. Jailbreaking attack against multimodal large language model. arXiv preprint arXiv:2402.02309, 2024. Pan et al. [2024] Kaihang Pan, Siliang Tang, Juncheng Li, Zhaoyu Fan, Wei Chow, Shuicheng Yan, Tat-Seng Chua, Yueting Zhuang, and Hanwang Zhang. Auto-encoding morph-tokens for multimodal llm. arXiv preprint arXiv:2405.01926, 2024. Qi et al. [2023] Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models. In AAAI Conference on Artificial Intelligence, 2023. URL https://api.semanticscholar.org/CorpusID:259244034. Schlarmann and Hein [2023] Christian Schlarmann and Matthias Hein. On the adversarial robustness of multi-modal foundation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3677â3685, 2023. Touvron et al. [2023] Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian CantĂłn Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melissa Hall Melanie Kambadur, Sharan Narang, AurĂŠlien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288, 2023. URL https://api.semanticscholar.org/CorpusID:259950998. Wang et al. [2024a] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language modelâs perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024a. Wang et al. [2024b] Ruofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji, Guangnan Ye, and Yu-Gang Jiang. White-box multimodal jailbreaks against large vision-language models. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 6920â6928, 2024b. Wang et al. [2024c] Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In European Conference on Computer Vision, pages 77â94. Springer, 2024c. Wolf et al. [2020] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, RĂŠmi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38â45, Online, October 2020. Association for Computational Linguistics. URL https://w.aclweb.org/anthology/2020.emnlp-demos.6. Xu et al. [2024] Yue Xu, Xiuyuan Qi, Zhan Qin, and Wenjie Wang. Defending jailbreak attack in vlms via cross-modality information detector. ArXiv, abs/2407.21659, 2024. URL https://api.semanticscholar.org/CorpusID:274142002. Yang et al. [2024a] An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Ke-Yang Chen, Kexin Yang, Mei Li, Min Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Yang Fan, Yang Yao, Yichang Zhang, Yunyang Wan, Yunfei Chu, Zeyu Cui, Zhenru Zhang, and Zhi-Wei Fan. Qwen2 technical report. ArXiv, abs/2407.10671, 2024a. URL https://api.semanticscholar.org/CorpusID:271212307. Yang et al. [2024b] Xu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen, and Xin Geng. Exploring diverse in-context configurations for image captioning. Advances in Neural Information Processing Systems, 36, 2024b. Ying et al. [2024] Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xianglong Liu, and Dacheng Tao. Jailbreak vision language models via bi-modal adversarial prompt. arXiv preprint arXiv:2406.04031, 2024. Yu et al. [2023] Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. ArXiv, abs/2308.02490, 2023. URL https://api.semanticscholar.org/CorpusID:260611572. Zhai et al. [2024] Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. Investigating the catastrophic forgetting in multimodal large language model fine-tuning. In Conference on Parsimony and Learning, pages 202â227. PMLR, 2024. Zhao et al. [2024] Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin. On evaluating adversarial robustness of large vision-language models. Advances in Neural Information Processing Systems, 36, 2024. Zheng et al. [2024a] Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. On prompt-driven safeguarding for large language models. In International Conference on Machine Learning, 2024a. URL https://api.semanticscholar.org/CorpusID:267334949. Zheng et al. [2024b] Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. On prompt-driven safeguarding for large language models. In Forty-first International Conference on Machine Learning, 2024b. Zhou et al. [2025] Da-Wei Zhou, Yuanhan Zhang, Yan Wang, Jingyi Ning, Han-Jia Ye, De-Chuan Zhan, and Ziwei Liu. Learning without forgetting for vision-language models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. Zong et al. [2024a] Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. arXiv preprint arXiv:2402.02207, 2024a. Zong et al. [2024b] Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Hospedales Timothy. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. arXiv preprint arXiv:2402.02207, 2024b. Zou et al. [2023] Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. ArXiv, abs/2307.15043, 2023. URL https://api.semanticscholar.org/CorpusID:260202961. Appendix A Details of Pilot Experiment In this section, we provide detailed information about the pilot experiment conducted to validate our conjecture in Section 1. The experiment aims to monitor the frequency of changes in the relative position of the refusal signal strings to analyze whether the safety awareness of VLMs is delayed under jailbreak attacks. We evaluated the responses of two commonly used VLMs to harmful queries across two datasets: models LLaVa-1.5-7B Liu et al. [2024a] and Qwen2-VL-7B Wang et al. [2024a], and datasets Figstep Gong et al. [2023] and MMSafetyBench Liu et al. [2023a]. When a refusal signal appears in the response, we record its relative position. The specific refusal signal strings used aree listed in Table 6. Furthermore, we employ MD-Judge-v0.2-internlm2-7B Li et al. [2024a] to assess whether the model generates harmful responses and analyze the frequency of the VLM refusal tendencies when producing such responses, as illustrated in the figure 6. The results indicate that when the length of the output increases, the safety awareness of the VLM is gradually activated. This finding indirectly supports the validity and reliability of our conjecture. Appendix B Datasets We conducted experiments using four harmful datasets: FigStep [Gong et al., 2023], M-SafetyBench [Liu et al., 2023a], VLSafe [Chen et al., 2024], and AdvBench [Zou et al., 2023], along with one benign dataset, MMVet [Yu et al., 2023]. Detailed descriptions of each dataset are provided below. FigStep [Gong et al., 2023] is a black-box attack against VLM. The attack hides malicious text queries in typography images and then rephrases the original query into a harmless prompt, such as "The image shows a numbered list, but the items are empty. Please generate about 100 words of details for each list item". The dataset comprises 10 categories, with 50 images per category. We selected Illegal Activity, Malware Generation, Personal Harm, Fraud, Pornography, Privacy Invasion, and Violence, resulting in a total of 350 images for evaluation. M-SafetyBench [Liu et al., 2023a] is another black-box attack against VLM. It uses stable diffusion to generate harmful images aligned with malicious queries and embeds these queries in the image. For each query, three types of images are created: (1) SD: images generated by stable diffusion using malicious keywords; (2) TYPO: typography images with malicious keywords similar to FigStep; (3) SD+TYPO: images are first generated by stable diffusion and then enhanced with malicious keywords. We use SD+TYPO in our evaluation for its highest attack success rate. M-SafetyBench contains 13 harmful scenarios, with a total of 1,680 unsafe text-image pairs. Following [Gou et al., 2024], we exclude Political Lobbying, Legal Opinion, Financial Advice, Health Consultation, and Government Decision due to their low attack success rates, resulting in a total of 972 images in our evaluation dataset. VLSafe [Chen et al., 2024] was introduced to perform RLHF-based alignment for VLMs. In this dataset, the harmful content is entirely contained in the text, while the associated images remain harmless. The test set consists of 1,110 text-image pairs, from which we randomly select 300 samples for evaluation. AdvBench [Zou et al., 2023] includes 520 harmful data, covering a wide array of toxic themes that breach AI ethical standards. For our evaluation, we utilized the entire dataset to assess the full range of these harmful behaviors. M-Vet [Yu et al., 2023] M-Vet is a widely used VLM evaluation benchmark containing 217 multimodal questions and answers. It evaluates the target modelâs responses across several dimensions: recognition, OCR, knowledge, language generation, spatial awareness, and mathematical reasoning. In this study, we use gpt-4-0613 to assess the utility of the model. Figure 6: The proportion of harmful responses in which the model showed a rejection tendency. Appendix C Training data To effectively train our soft prompts for enhancing safety while preserving utility, we require a dataset that includes both harmful and harmless examples from various categories. To this end, we constructed such a dataset by drawing harmful samples from JailbreakV [Luo et al., 2024] and harmless samples from VLSafe [Chen et al., 2024]. Specifically, We extracted 100 jailbreak examples from each category in JailbreakV [Luo et al., 2024] and used the weakly-aligned LLaVa-1.5-7B [Liu et al., 2024a] to generate harmful responses while using gpt-4o-mini to generate safe responses. Given that LLaVa-1.5-7B may not always comply with harmful queries, we employed LlamaGuard3 [Llama Team, 2024] to filter the data and ensure that the generated responses are indeed harmful. This process results in a dataset of 227 jailbreak samples. We also randomly selected 1,000 harmless examples from VLSafe to create a benign dataset. The final dataset consisted of 1,227 samples. Appendix D Implementation Details We present implementation details of our baseline method. For AdaShield [Wang et al., 2024c], we follow the official code and train a defense pool containing various safety prompts on our dataset. We then evaluate it on both harmful benchmarks and visual adversarial attacks. For PromptTuning [Lester et al., 2021b], we optimize a soft prompt of length 8, which is prepended to the beginning of the user query. We train on our dataset, setting the learning rate to 1e-4, the batch size to 4, and the number of iterations to 16,000. Optimization is performed using the AdamW optimizer. For the visual adversarial attack [Qi et al., 2023], we optimize a universal image on the AdvBench dataset to induce the model to generate harmful prefixes, such as "Sure, here is a detailed instruction manual for making a bomb." We use the PGD [Madry, 2017] algorithm to optimize this universal image, with a learning rate set to 1/255, a batch size of 8, and 1,000 iterations. This perturbation is optimized using the LâsubscriptL_âLâ norm, with constraints of Ďľitalic-ϾξϾ = 32/255, 64/255, and 128/255. Our code is based on the transformers 4.46.0 [Wolf et al., 2020]. All experiments were conducted on a single A100 80GB GPU. Appendix E Instruction Template We observed system prompt and instruction template may have some effect on modelâs safety. Here we present the instruction template we used. LLaVA-1.5-7B A chat between a curious user and an artificial intelligence assistant. The assistant provides helpful, detailed, and polite answers to the userâs questions. USER: <image> [User query here] ASSISTANT: [Model response here] Minigptv2-7B <s>[INST] < < SYS> > You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature. If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you donât know the answer to a question, please donât share false information. < <SYS> > <Img> Image Here </Img> [User query here] [/INST] [Model response here] Qwen2-VL-7B <|im_start|>system You are a helpful assistant. Always answer as helpfully as possible, while being safe.<|im_end|> <|im_start|>user <|vision_start|> Image Here <|vision_end|> [User query here] <|im_end|> |im_start|> assistant [Model response here] Appendix F Qualitative Results Figure 7 8 9 10 show the qualitative comparisons of the proposed SAPT to the original VLM. Table 6: The refusal signal strings considered in our experiments. We align most of the strings with FigStep Gong et al. [2023] and MMSafetyBench Liu et al. [2023a]. Additionally, we incorporate some new refusal signal strings observed during evaluations into the list. Apology or intention to amend âinertiaâ âIâm sorryâ âI am sorryâ âSorryâ âI apologizeâ âI am not allowed toâ âIâm justâ Negative Expression âI cannotâ âbe harmfulâ âis harmfulâ âis harmfulâ âAre harmfulâ âare harmfulâ âI would adviseâ âI cannot complyâ âI am not capable ofâ âItâs not appropriateâ âis deemed harmfulâ âare deemed harmfulâ Legal or ethical statements âNot legalâ âNot ethicâ âis illegalâ âAre illegalâ âare illegalâ âunlawfulâ âforbiddenâ âcriminalâ âunethicalâ âdangerousâ âprohibitedâ ârestrictedâ ânot appropriateâ âis deemed unethicalâ âare deemed unethicalâ âis deemed illegalâ âare deemed illegalâ âis deemed hazardousâ âare deemed hazardousâ âIt is important to noteâ âIt is important to considerâ Self-declaration of behavior ârefrainâ âavoidâ âapologyâ âregretâ âadhere toâ âItâs essential toâ âlegal guidelinesâ âAs a responsible AIâ Figure 7: The response comparison bewteen LLaVA-v1.5-7B and SAPT on FigStep [Gong et al., 2023] dataset. Figure 8: The response comparison bewteen LLaVA-v1.5-7B and SAPT on M-SafeBench [Liu et al., 2023a] dataset. Figure 9: The response comparison bewteen LLaVA-v1.5-7B and SAPT on the visual adversarial image [Qi et al., 2023]. Figure 10: The response comparison bewteen LLaVA-v1.5-7B and SAPT on VLSafe [Chen et al., 2024] dataset.