Paper deep dive
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
Hanbin Hong, Shuya Feng, Nima Naderloui, Shenao Yan, Jingyu Zhang, Biying Liu, Ali Arastehfard, Heqing Huang, Yuan Hong
Models: Claude 3, DeepSeek-R1, GPT-4
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/12/2026, 5:52:08 PM
Summary
This paper presents a Systematization of Knowledge (SoK) on LLM prompt security, addressing the fragmentation in the field by proposing a multi-level taxonomy for attacks, defenses, and vulnerabilities. It introduces 'JailbreakDB', a large-scale annotated dataset, and an open-source evaluation toolkit to standardize the assessment of LLM robustness against jailbreak prompts.
Entities (5)
Relation Signals (4)
Taxonomy I â classifies â Jailbreak attack techniques
confidence 100% · Taxonomy I covers jailbreak attack techniques
Taxonomy II â classifies â Defense methodologies
confidence 100% · Taxonomy II addresses defense methodologies
Taxonomy III â classifies â LLM vulnerabilities
confidence 100% · Taxonomy III summarizes inherent vulnerabilities in large language models
JailbreakDB â contains â Jailbreak prompts
confidence 95% · releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) have rapidly become integral to real-world applications, powering services across diverse sectors. However, their widespread deployment has exposed critical security risks, particularly through jailbreak prompts that can bypass model alignment and induce harmful outputs. Despite intense research into both attack and defense techniques, the field remains fragmented: definitions, threat models, and evaluation criteria vary widely, impeding systematic progress and fair comparison. In this Systematization of Knowledge (SoK), we address these challenges by (1) proposing a holistic, multi-level taxonomy that organizes attacks, defenses, and vulnerabilities in LLM prompt security; (2) formalizing threat models and cost assumptions into machine-readable profiles for reproducible evaluation; (3) introducing an open-source evaluation toolkit for standardized, auditable comparison of attacks and defenses; (4) releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date;\footnote{The dataset is released at \href{this https URL}{\textcolor{purple}{this https URL}}.} and (5) presenting a comprehensive evaluation platform and leaderboard of state-of-the-art methods \footnote{will be released soon.}. Our work unifies fragmented research, provides rigorous foundations for future studies, and supports the development of robust, trustworthy LLMs suitable for high-stakes deployment.
Tags
Links
- Source: https://arxiv.org/abs/2510.15476
- Canonical: https://arxiv.org/abs/2510.15476
Trouble viewing inline? Open PDF directly â
Full Text
211,871 characters extracted from source content.
Expand or collapse full text
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models Hanbin Hong 1 , Shuya Feng 1,2â , Nima Naderloui 1â , Shenao Yan 1â , Jingyu Zhang 3â , Biying Liu 3 , Ali Arastehfard 1 , Heqing Huang 3â , and Yuan Hong 1â 1 University of Connecticut, 2 University of Alabama at Birmingham, 3 Independent â Equal Contribution as Second Authors (listed in Alphabetical Order), â Corresponding Authors Abstract Large Language Models (LLMs) have rapidly become integral to real-world applications, powering services across diverse sectors. However, their widespread deployment has exposed critical security risks, particularly through jailbreak prompts that can bypass model alignment and induce harmful outputs. De- spite intense research into both attack and defense techniques, the field remains fragmented: definitions, threat models, and evaluation criteria vary widely, impeding systematic progress and fair comparison. In this Systematization of Knowledge (SoK), we address these challenges by (1) proposing a holistic, multi- level taxonomy that organizes attacks, defenses, and vulnerabilities in LLM prompt security; (2) for- malizing threat models and cost assumptions into machine-readable profiles for reproducible evaluation; (3) introducing an open-source evaluation toolkit for standardized, auditable comparison of attacks and defenses; (4) releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date; 1 and (5) presenting a comprehensive evaluation platform and leaderboard of state-of-the-art methods 2 . Our work unifies fragmented research, provides rigorous foundations for future studies, and supports the development of robust, trustworthy LLMs suitable for high-stakes deployment. 1 Introduction Large Language Models (LLMs) have rapidly transitioned from academic research to core components of real-world applications, especially since the emergence of high-profile foundation models such as OpenAIâs GPT series [17, 140], Google Gemini [9], Meta Llama [175, 176], Anthropic Claude [12], Alibaba Qwen [11, 210, 209], and Doubao [172]. Today, LLMs are deployed across an unprecedented range of sectorsâfrom web search and code assistants to legal, educational, and healthcare domainsâreaching hundreds of millions of end users globally. The rapid adoption of LLMs has ushered in a new era of AI-powered services, but it also brings serious safety and security risks. These risks manifest in multiple forms, ranging from misinformation and privacy leaks to adversarial attacks that exploit model vulnerabilities. In particular, a growing body of work shows that carefully crafted jailbreak prompts can bypass alignment constraints, inducing models to produce sensitive, illegal, or harmful content. Alarmingly, recent studies report that such attacks achieve success rates exceeding 90% even on flagship models such as GPT-4, Claude 3, and DeepSeek-R1 [124, 42, 154, 118]. The outputs generated through these attacks could be used for malicious purposes, underscoring the urgent need for close attention and mitigation. Fragmented Progress. The urgent need to secure these widely deployed models has led to a surge of research on both attacks and defenses. In just the past two years, scholars have developed a diverse arsenal of attack strategies, including token-flipping [124], task-overload [42], multi-turn derailment [154], evolu- tionary prompt search [118], and even image-based exploits targeting multimodal models [50]. Meanwhile, defenses have proliferated as well, ranging from heuristic filters and sparsity-based runtime mitigations [161], 1 The dataset is released at https://huggingface.co/datasets/youbin2014/JailbreakDB. 2 will be released soon. 1 arXiv:2510.15476v2 [cs.CR] 21 Oct 2025 to retrieval-augmented prompt decomposition [187], ensemble detection frameworks such as MoJE [31], re- inforcement learning-based alignment, and architectural interventions like KV-cache eviction. However, despite this rapid progress, research in the area remains highly fragmented. Most works are developed and evaluated in isolation, each introducing its own threat definitions, cost assumptions, datasets, and evaluation metrics, which are often incompatible with others. This heterogeneity makes it difficult, if not impossible, to fairly compare the effectiveness or generality of different attacks and defenses. As a result, the community still lacks a clear understanding of which methods are genuinely robust, under what conditions they succeed or fail, and how to advance toward secure and trustworthy LLMs in a principled way. Recognizing these gaps, the community is taking initial steps toward a more coordinated and systematic study of LLM jailbreak attacks and defenses. Partial Attempts at Systemization. Several recent surveysâsuch as those by Yi et al. [217], Fan et al. [50], and Esmradi et al. [49]âhave begun to map the landscape of attacks and defenses. In parallel, benchmarks like JailbreakZoo [86], JailbreakBench [21], and MMJ-Bench [198] have provided important footholds for reproducibility and comparison. However, despite these valuable efforts, several fundamental challenges remain unresolved: 1. Limited threat coverage. Most resources over-represent early DAN-style single-turn exploits, leaving token-level, chain-of-thought, and multimodal attacks under-explored. 2. Inconsistent and opaque metrics. Success is reported using incompatible measures (raw jailbreak rate, compliance likelihood, policy-violation scores, or vaguely defined âtoxicityâ), preventing reliable comparisons. 3. Sparse data and missing metadata. Public corpora rarely annotate prompts with attacker capability, required knowledge of system prompts, or execution cost, hindering detection and interpretability research. 4. Taxonomyâevaluation gap. No prior work pairs a principled threat taxonomy with open, executable tooling that holds cost, knowledge, and access assumptions constant across attacks and defenses. Systemization of Knowledge (SoK). These structural gaps limit our ability to systematically compare approaches and to track meaningful progress. To address these challenges and move the field toward a more systematic understanding, this Systematization of Knowledge (SoK) makes the following contributions: 1. Holistic taxonomy. We propose a comprehensive, multi-level taxonomy that systematically organizes both attacks and defenses based on (i ) attacker or defender capability, (i ) the specific model vulner- abilities being exploited or protected, and (i ) the underlying threat objectivesâthereby unifying and extending prior classification efforts. 2. Declarative threat models. Our work formalizes commonly used but often implicit assumptionsâsuch as attack budget, query limits, and side-channel accessâinto explicit, machine-readable profiles that serve as the backbone for consistent and reproducible evaluation. 3. Open evaluation toolkit. We introduce a modular and extensible platform that enables any combina- tion of (model, attack, defense) to be instantiated, executed, and evaluated. The toolkit logs costs, safety, and utility scores with full auditability, supporting rigorous empirical comparisons. 4. JailbreakDB. We release a large-scale, curated text-only corpora for LLM safety research: a jailbreak split (445,752 unique systemâuser pairs) and a benign split (1,094,122 benign prompts) collected from 14 sources. Each example includes a system prompt, a user prompt, and lightweight labels for jailbreak status, source, and tactic. This release focuses on providing a comprehensive, deduplicated foundation dataset, available at huggingface.co/datasets/youbin2014/JailbreakDB. 5. Comprehensive evaluation. We provide a unified evaluation of state-of-the-art attacks, defenses, and leading LLMs. This systematic assessment uncovers a range of insightful observations and actionable findings regarding the strengths and limitations of current approaches. By releasing our taxonomy, open-source evaluation toolkit, and richly annotated dataset, we transform fragmented anecdotal findings into a rigorous, comparable, and extensible science of LLM robustness. To- gether, this work aim to catalyze progress toward verifiably aligned language models and enable the devel- opment of defenses robust enough for the fast-paced, high-stakes environments of real-world deployment. 2 Systematic Literature Review The goal of this literature review and taxonomy is to provide a comprehensive overview of the rapidly evolving field of LLM jailbreak attacks, defenses, and to reveal the LLM security vulnerability, helping the researchers understand the current landscape and the diversity of approaches being developed. 2.1 Scope and Overview of the Taxonomy This work introduces three taxonomies that approach the field of LLM prompt security from distinct yet complementary perspectives: Taxonomy I covers jailbreak attack techniques, Taxonomy I addresses defense methodologies, and Taxonomy I summarizes inherent vulnerabilities in large language models. Each taxonomy aims to provide a systematic and comprehensive framework for organizing current research and facilitating consistent analysis. Taxonomy I: Jailbreak Attack Technologies. Taxonomy I adopts a two-level organizational principle. At the top level, attacks are classified according to their underlying threat model, which specifies the adver- saryâs capabilities and assumptionsâfor example, black-box versus white-box access to the model. Within each threat model, attacks are further classified according to their technical methodology, e.g., prompt modification techniques, or LLM-assisted techniques. Recognizing the complexity of real-world attacks, we decompose the landscape of attacks into a set of atomic technical units: minimal, actionable components that represent the core tactics used in the literature. These atomic units are not mutually exclusiveâmany attacks employ multiple tactics in parallel or series, and the boundaries between units can be fluid or overlapping. Our taxonomy, therefore, serves as a compositional framework that captures both the breadth and depth of jailbreak technologies, rather than imposing rigid categories. This approach enables systematic annotation and comparison, and helps to clarify how different techniques may be combined or extended. Detailed classification guidelines and an overview of the taxonomy are provided in Section 2.3 and Figure 1. Taxonomy I: Jailbreak Attack Technologies. For defenses, we adopt a similar two-step approach, based on the intended goal and implementation strategy: we first classify defenses by their primary goalâeither detection or preventionâand then further subdivide each group based on the methodologies. This approach enables fair comparison and systematic analysis of defense strategies. See Section 2.4 and Figure 2. Taxonomy I: LLM Vulnerabilities. Notably, a significant class of attacks specifically exploits intrinsic vulnerabilities in LLMs to bypass guardrails; for these, we also present a taxonomy of LLM vulnerabilities. See Section 2.5 and Figure 3. Table 1: Integrated comparison of representative jailbreak literature (survey / benchmark / dataset) Taxonomy (citation) Year Survey / Review Benchmark / Eval Dataset released Attack coverage Defense coverage Threat- model axis Notes Jin et al. [86]2024ââââ7 flat âtypesâ Yi et al. [217]2024ââââBlack/white-box split Xu et al. [208]2024ââââ9 attacks / 7 defenses Esmradi et al. [49]2023ââââGen-AI broad scope Inie et al. [75]2023ââââGrounded-theory interviews Shayegani et al. [160]2023âââVulnerability-centric Gupta et al. [63]2023ââââCyber-security view Ours2025âMulti-layer taxonomy + Evaluation Platform Comparison with Prior Taxonomies. Compared with previous taxonomies such as Jin et al. [86], Yi et al. [217], Xu et al. [208], Esmradi et al. [49], Inie et al. [75], Shayegani et al. [160], and Gupta et al. [63], our taxonomy adopts a more systematic and fine-grained approach. Specifically, we focus exclusively on text-to-text LLMs and introduce a multi-dimensional framework that separates threat models from attack methodologies, while also providing a parallel, structured taxonomy for defenses and a dedicated ontology of LLM vulnerabilities. This structure allows for more precise classification and mapping between attacks, defenses, vulnerabilities, and benchmarks, facilitating a comprehensive and extensible analysis of the field. For a detailed comparison with representative prior surveys, please refer to Table 1. 2.2 Key Concepts and Threat Model This section defines the terminology and threat modeling assumptions used throughout the survey. Unless otherwise noted, the discussion primarily concerns the textual components of LLMs and their corresponding guardrail mechanisms, though the terminology and assumptions may also extend to the language-processing modules within multi-modal models. Core Concepts. Throughout this survey, the following terminology is employed. Guardrails (or safety filters) refer to policies, system messages, or automated classifiers enforced at inference time to constrain model behaviour. A jailbreak prompt denotes an input specifically crafted to induce the LLM to violate these guardrails or its instruction hierarchy. Access settings are categorized as black-box, gray-box, or white-box: in the black-box setting, the attacker is restricted to input-output interactions with no knowledge of the modelâs internal mechanisms; the gray-box setting permits partial visibility, such as access to auxiliary signals (e.g., confidence scores, log probabilities, or limited feedback); in the white-box setting, the attacker enjoys full access to the modelâs parameters, architecture, training data, and internal computations, including gradients. Attacks may be single-turn (occurring within one round of dialogue) or multi-turn (spanning multiple rounds of dialogue). Attacker Model Dimensions. The attacker model is defined by several dimensions, including the attack goal (e.g., policy violation, sensitive information extraction, system instruction hijacking, or tool abuse), the level of knowledge (black-box, gray-box, or white-box access as described above), and capability (such as permitted query budget, granularity of context control, or access to additional external LLMs). Additional factors include the stealth requirement (the extent to which detection evasion is necessary, ranging from static filter evasion to dynamic, runtime detector evasion), interaction pattern (single-turn, multi-turn, or indirect injection, where the latter involves malicious prompts embedded in upstream content), and resource budget (constraints on tokens, computation, time, or monetary cost). Defender Model Dimensions. The defender model is specified by the defense goal (detection, prevention, or refusal), deployment layer (model input, model modification, or model output), level of knowledge on attacks (known or unknown attack), and adaptivity (static rules, dynamic online learning, or proactive self-red-teaming). The resource budget includes permissible latency, additional inference calls, or human-in- the-loop involvement. 2.3 Taxonomy I: Jailbreak Attack Techniques We categorize jailbreak attacks primarily based on their underlying threat model, distinguishing between black-box attacks, where adversaries interact with the model solely through its input-output interface, and white-box attacks, where adversaries have access to the modelâs internal parameters or training process. I.1 Black-box Jailbreak Attacks Black-box jailbreak attacks encompass adversarial strategies that seek to bypass LLM safety mechanisms solely through external interactions, without access to model internals such as weights, gradients, or archi- tecture details. Attackers operate by probing the model via its public interfaceârelying on input-output behavior and feedbackâto iteratively craft prompts or strategies that induce undesired or harmful outputs. This paradigm includes techniques such as prompt modification, heuristic and reinforcement learning-based optimization, LLM-assisted prompt generation, and multi-turn manipulation, collectively revealing the limits of external safety controls and the generalization capacity of LLMs against real-world adversaries. Jailbreak Attack Techniques 1. Black-box Jailbreak Attack 2. White-box Jailbreak Attack 1.3 LLM-Assisted Techniques 1.3.3 Surrogate Models 1.3.2 Multi-Agent Collaboration 1.3.1 LLM- Generated Attacks 1.2 Black-box Optimization 1.2.3 Reinforcement Learning 1.2.2 Gradient Estimation 1.2.1 Heuristic Algorithms 1.5 Leveraging Model Vulnerability 1.5.4 Defensive Mechanism Bypass 1.5.3 Decoding and Sampling Strategy Exploits 1.5.2 Learning Mechanisms Exploits 1.5.1 Leveraging Behavior Vulnerabilities 1.4 Multi-Turn Techniques 1.4.1 Multi-Turn Contextual Attacks 1.4.2 Gradual Escalation 2.1 Model Modification 2.1.1 Malicious Fine-Tuning 2.1.2 Parameter Modification 2.1.4 Backdoor Insertion 2.1.3 Disabling Safety Mechanisms 1.1.1 Obfuscation and Encoding 1.1.2 Substitution and Synonyms 1.1.4 Intent Concealment 1.1.3 Decomposition 2.2.2 Gradient- Based Optimization 1.1 Prompt Modification 2.2 White-box Optimization 2.2.1 Prefix/Suffix Manipulation Figure 1: Overview of Taxonomy I: Jailbreak Attack Techniques I.1.1 Prompt Modification Techniques Prompt modification techniques encompass a family of black-box attack strategies that directly alter the content or structure of input prompts to circumvent safety filters in LLMs, while maintaining the underlying malicious intent. I.1.1.1 Obfuscation and Encoding Definition: The Obfuscation and Encoding class comprises black-box jailbreak attacks that systematically transform malicious prompts to evade safety filters while retaining semantics for LLMs. This category in- cludes four key subclasses: (1) Character Alteration, such as typos, leet speak, or homoglyphs; (2) Encoding and Encryption, using Base64, ROT13, hexadecimal, binary, or custom ciphers; (3) Multilingual Transfor- mation, including translation or language mixing to obscure disallowed content; and (4) Code and Markup Language Transformation, where requests are reframed as code, markup, or embedded in comments and strings. Unlike simple text perturbations, these methods exploit LLMsâ symbolic reasoning and pattern recognition abilities, bypassing filters that focus on natural language or surface-level cues. Recent literature establishes that such attacks, including cipher-based prompts, data structure encoding, and rare-token permutations, reliably bypass safety systems in state-of-the-art models like GPT-4 [223, 65]. Extensions to multimodal and clinical contexts reveal that obfuscated cuesâsuch as Unicode, tiny fonts, or code-wrapped instructionsâsubvert LLM-supported decisions [28, 132]. Automated attacks via prompt flipping and multi-level encoding have been systematically codified, highlighting the modelsâ inability to gen- eralize alignment to encoded modalities [158, 124]. Recent multilingual studies demonstrate that translating prompts into non-English languages (especially low-resource ones) can significantly weaken safety filters in LLMs, enabling both unintentional and intentional jailbreaks, and that mitigation remains challenging even with advanced models like GPT-4 [99]. Multi-turn, ciphered, or delimiter-free attacks further demonstrate that symbolic obfuscation is both universal and stealthy, revealing that current guardrails fail to anticipate the modelsâ proficiency in de-obfuscation and code reasoning [177, 59]. These works underscore a fundamen- tal mismatch between pattern-based filtering and the general computation abilities of LLMs, necessitating future safety efforts to target non-natural input modalities and to automate semantic de-obfuscation. I.1.1.2 Substitution and Synonyms Definition: Substitution and Synonyms refers to the strategic replacement of sensitive words, phrases, or prompt components with alternativesâsuch as synonyms, euphemisms, paraphrases, or code wordsâto bypass input filters and safety mechanisms, while preserving the original semantics. In the literature, Substitution and Synonyms is established as a key prompt modification tactic for black- box LLM jailbreaks. Automated synonym and phrase replacementâdemonstrated by SurrogatePrompt [10], latent jailbreak benchmarks [146], and large-scale empirical studies [123]âenables attackers to bypass filters with minor linguistic changes. Advanced strategies further leverage semantic constraints and optimization, generating transferable and high-similarity triggers that evade detection [230, 102]. Overall, substitution- based attacks form an adaptive subclass that fundamentally challenges input-based defenses, especially the keyword filters. I.1.1.3 Decomposition Definition: Prompt decomposition attacks break down a malicious request into smaller, innocuous com- ponents, which are then systematically recombinedâeither implicitly or explicitlyâby the model to achieve the harmful intent. Prompt decomposition bypasses LLM safety by splitting harmful queries into syntactically or semantically benign sub-prompts, which are then recombined by the model to achieve the original intent. Representa- tive works, including DrAttack and RL-based frameworks like PathSeeker and DRA [103, 110], use syntactic parsing, in-context learning, and iterative feedback to automate sub-prompt generation and aggregation, sig- nificantly boosting attack effectiveness. Chain-of-Jailbreak generalizes this approach to multimodal models, demonstrating high bypass rates for toxic content construction [189]. Theoretical studies reveal that de- composition exposes compositional leakage, which evades input/output-based censorship [60]. Enhanced by obfuscation methods such as euphemization, distractor injection, or context-driven summarization [116, 115], prompt decomposition outperforms basic payload splitting [90] and remains robust against advanced defenses relying on perplexity or anomaly detection [27]. I.1.1.4 Intent Concealment Definition: Intent concealment refers to the strategy of hiding malicious intent behind neutral or innocuous language to evade detection. Typical techniques include framing disallowed requests as academic or research inquiries, embedding harmful objectives within legitimate contexts or narratives, and employing hypothetical or fictional scenarios. Recent literature on intent concealment demonstrates a range of prompt modification strategies that camouflage harmful intent. Notably, sequential masking and incremental synthesis, as in Imposter.AI [116], diffuse toxicity by staging benign sub-questions that are later combined. Role-play, narrative, and scenario- based disguises [222, 63] detach illicit intent by reframing prompts as part of storytelling, research, or humor. Distraction-based techniques (e.g., Tastle [201]) manipulate model focus by embedding concealed objectives within larger benign tasks. Methods like IntentObfuscator [132] directly synthesize ambiguity and leverage complex linguistic alterations to evade both automated and manual detection. I.1.2 Black-box Optimization Black-box Optimization techniques encompass a broad family of black-box methods that iteratively search the prompt space using non-gradient, feedback-driven strategiesâsuch as heuristics, gradient estimation, and reinforcement learningâto discover prompts capable of bypassing LLM safety filters without requiring internal model access. I.1.2.1 Heuristic Algorithms Definition: The heuristic algorithms class encompasses black-box jailbreak methods that leverage iterative, often population-based or random, search strategiesâsuch as genetic algorithms, evolutionary approaches, random search, and metaheuristicsâto explore the discrete prompt space without access to model gradients. These algorithms optimize adversarial prompts by evaluating candidate generations through fitness functions, typically based on embedding similarity, harmfulness, and stealth, with the core objective of efficiently bypassing LLM safeguards where direct gradient signals are unavailable or uninformative. A substantial body of literature establishes and refines this paradigm. OpenSesame [96] first adapts genetic algorithms for prompt suffix generation, demonstrating that black-box heuristic search can yield universal, transferable adversarial prompts. Subsequent worksâsuch as AutoDAN [118], Semantic Mirror Jailbreak (SMJ) [102], AutoJailbreak [125], and BlackDAN [190]âadvance this framework via improved initialization, crossover, mutation, and evaluation mechanisms, introducing multi-objective optimization and Pareto-dominance to balance harmfulness, semantic alignment, and detectability. Lightweight methods such as random search [244, 5], discrete coordinate ascent [88], and greedy tree-search [24] further enrich the heuristic toolkit, revealing new vulnerabilities and attack pathways. Collectively, these studies show that heuristic algorithms uniquely combine practical black-box applicability, extensibility, and transferability, with continual innovations in fitness shaping and prompt initialization (e.g., [253]) driving ongoing progress and shaping the evolving adversarial landscape. I.1.2.2 Gradient Estimation Methods Definition: Gradient estimation techniques aim to enable optimization-based adversarial prompt engineer- ing for black-box jailbreak attacks, where direct access to model gradients is unavailable. By approximating gradient directionality through queries or surrogate models, these methods facilitate systematic search for adversarial prompts, typically using strategies such as finite differences or proxy-driven estimation in the discrete token space. Recent works have advanced the use of gradient estimation for efficient black-box adversarial prompt attacks. PAL [167] leverages surrogate models to approximate gradients and employs sophisticated candidate ranking, enabling scalable and low-cost attacks against commercial LLM APIs with strong transferability. Discrete gradient estimation methods such as RAL [251] and variants that use finite-difference or score- based queries further demonstrate that query-efficient, token-level directional search significantly outperforms random or evolutionary baselines. Extensions to multimodal settings, including attacks on diffusion and T2I models [164, 131], confirm the broad applicability of gradient estimation. Key insights include the methodâs efficiency and transferability, but also highlight vulnerabilities such as perplexity-based detection and the potential for overfitting to surrogates. I.1.2.3 Reinforcement Learning Definition: Unlike brute-force or purely stochastic approaches, RL-based techniques employ agentsâtypically leveraging deep RL algorithms such as PPO or MADDPGâthat adaptively refine attack strategies based solely on black-box feedback, optimizing toward reward signals aligned with attack objectives. RL-based methods, including PathSeeker [110], RLbreaker [25], SneakyPrompt [214], RL-JACK [26], Atoxia [47], and Arondight [121], collectively reframe jailbreak prompt generation as a reward-driven search problem. These works demonstrate that RL agents, via custom reward functions quantifying attack success, information richness, or semantic alignment, outperform classical genetic or random mutators in both LLMs and multimodal models. For example, PathSeeker and RLbreaker leverage collaborative or single-agent RL frameworks to optimize malicious prompt discovery, while SneakyPrompt and Arondight extend RL-based attacks to text-to-image and vision-language models using specialized reward shaping. Atoxia further high- lights the effectiveness of RL by exploiting the target modelâs own toxic answer likelihood as a feedback signal. Collectively, these studies reveal RLâs unique strengths: policy learning, adaptive exploration-exploitation, and advanced reward shaping, while also noting challenges such as sparse rewards and transferability issues across models or modalities. I.1.3 LLM-Assisted Techniques LLM-assisted techniques refer to methods that harness language models themselves to automatically generate, refine, or guide adversarial prompts, including LLM-generated attacks, multi-agent collaboration, and surrogate model approaches. I.1.3.1 LLM-Generated Attacks Definition: LLM-generated jailbreak attacks refer to adversarial strategies where one or more large language models (LLMs) autonomously generate, optimize, or refine attack prompts without access to target model internals. Distinct from human-crafted or gradient-based prompting, these methods leverage LLMsâoften open-source or less-alignedâas automated attackers, optimizers, or red-teamers to produce adversarial con- tent or prompt templates in a scalable, data-driven manner. Recent works (e.g., SoP, SeqAR [213], GPTFUZZER [220], DAP [201]) advance this field by iteratively generating and optimizing sophisticated jailbreak prompts, including multi-character or sequential templates, using attacker LLMs alone. IRIS [148] and PAIR [20] introduce self-reflective and self-explanation-based prompt refinement, enabling LLMs to act as both attacker and target under strict black-box constraints. AdvPrompter [142] demonstrates adversarial prompt learning, with LLMs trained to rapidly generate adap- tive, human-readable adversarial suffixes or prompts without gradient information. Recent multimodal approaches, such as AutoJailbreak [81] and Visual-RolePlay [133], further highlight LLMsâ autonomous development of vision-language attack strategies. I.1.3.2 Multi-Agent Collaboration Definition: Multi-agent collaboration in LLM-assisted black-box jailbreak attacks refers to the orches- trated interaction of multiple autonomous LLM-based agentsâeach potentially assigned distinct roles or strategiesâto generate, refine, and validate adversarial prompts. These agents collectively pursue jailbreak objectives by dividing labor, iteratively exchanging feedback, and leveraging agent diversity, enabling the creation of sophisticated prompts that bypass safety alignment boundaries. Recent literature has established multi-agent collaboration as a pivotal attack methodology. GUARD [85] introduces a structured four-role framework (Translator, Generator, Evaluator, Optimizer) where agents collectively translate, mutate, and iteratively enhance jailbreak prompts via contextual feedback, uncovering transferable and natural-language attacks. AutoDAN-Turbo [117] advances this by formalizing lifelong, self- evolving multi-agent exploration, enabling agents to autonomously discover, store, and combine adversarial strategies. JailFuzzer [41] adopts a fuzz-testing approach with mutation and oracle agents leveraging collec- tive memory for effective prompt generation. Evil Geniuses [174] frames jailbreaks as Red-Blue multi-agent competitions, revealing cascading vulnerabilities arising from agent interactions. I.1.3.3 Surrogate Models Definition: The surrogate model class refers to black-box jailbreak attack techniques that leverage one or more accessible substitute modelsâoften open-source or smaller LLMsâto craft, guide, or evaluate adversar- ial prompts targeting a closed-source or restricted-access LLM. By approximating the decision boundaries or behaviors of the target through these proxies, attackers systematically reduce sample complexity and query cost, enabling scalable, transferable attacks even when internal model parameters and direct outputs are unavailable. Recent literature operationalizes the surrogate model approach across diverse modalities and threat models. PAL framework [167] employs a proxy LLM for token-level gradient optimization and fine-tuning to align with the target, reducing queries and boosting attack efficacy. PRP [135] constructs universal adversarial prefixes via surrogate guard models, propagating successful perturbations to base LLMs for high transferability. Hayase et al. [68] integrate surrogates into gradient-based search, validating the âproxy filter plus query validationâ paradigm for superior efficiency. BlackDAN [190] embeds surrogates within a genetic algorithm as fitness evaluators, orchestrating stealthy, semantically relevant jailbreaks. I.1.4 Multi-Turn Techniques Strategies that leverage repeated interactions with a model, systematically manipulating its behavior by constructing, evolving, or adapting context across multiple conversational turns. I.1.4.1 Multi-Turn Contextual Attacks Definition: The multi-round context technique exploits the large language modelâs conversational memory by distributing adversarial intent across multiple interaction rounds. Instead of presenting explicit harm- ful prompts, the attacker decomposes the malicious objective into a series of semantically or functionally connected queries, each of which appears benign in isolation. Through this process, the attacker incre- mentally shapes the context such that the LLM is primed towards producing forbidden outputs, effectively circumventing safety mechanisms that focus on single-turn or overtly adversarial inputs. Recent literature formalizes multi-round attacks as iterative, context-driven adversarial engagements. Cheng et al. [27] demonstrate that models prompted with sequential, semantically related questions can be steered towards policy-violating outputs with higher success rates compared to zero-shot or random multi- turn baselines. Yang et al. [212] advance this by modeling adaptive, feedback-driven attack chains using evaluation functions to optimize context progression. Studies by Li et al. [100] and Bhardwaj and Poria [14] reveal that context-dependent jailbreaks expose class-level blind spots in current safety training, which is often limited to shallow, turn-level defenses. I.1.4.2 Gradual Escalation Techniques Definition: Gradual escalation refers to a distinct class of black-box jailbreak attacks on LLMs where the attacker incrementally increases the adversarial intent over multiple interaction rounds. In contrast to multi- turn contextual attacksâwhich focus on distributing benign-seeming components of the attack across the contextâgradual escalation centers on progressively intensifying the maliciousness or risk of each prompt within the conversation. This approach exploits the LLMâs conversational memory and adaptive response, with each turn subtly raising the attackâs severity or explicitness. Through this gradual evolution of intent, adversaries circumvent static or abrupt-trigger defenses, ultimately steering the LLM toward policy-violating outputs by systematically amplifying the threat level throughout the dialogue. Recent literature converges on several mechanisms and insights for gradual escalation attacks. Russi- novich et al. introduce Crescendo, illustrating how multi-turn, context-amplifying queries can subvert ad- vanced alignment protocols across modalities [157]. Jiang et al. model red teaming as a learnable, adaptive process that incrementally optimizes prompt concealment and actively leverages model feedback to enhance attack effectiveness [79]. Semantic-driven strategiesâsuch as chain-of-thought and actor-network accumula- tion by Yang et al. and Ren et al.âdemonstrate that progressively relevant prompts, not individual queries, evade static detection [212, 154]. Lin et al. reveal that reasoning-level, multi-step manipulations prompt LLMs to infer and escalate harmful content across turns [109], while Ramesh et al. show that iterative self-refinement, as in IRIS, efficiently converges on jailbreak compliance with few queries [148]. Inie et al. qualitatively analyze expert red-teamers, highlighting real-world manifestations such as âfoot-in-the-doorâ and emotional appeals, which collectively remain undetected until late-stage policy violations [75]. I.1.5 Leveraging Model Vulnerabilities Techniques in this category exploit inherent LLM features or vulnerabilitiesâsuch as behavioral patterns, learning dynamics, decoding processes, or defense mechanism characteristicsâto bypass safety controls in a black-box setting. Unlike prompt modification techniques, which alter input content, these attacks systemat- ically target vulnerabilities in the modelâs internal processing and system-level defenses. These vulnerabilities arise from both architectural design choices and operational defense limitations, creating persistent attack surfaces for attackers. I.1.5.1 Leveraging Behavior Vulnerabilities Definition: Behavior vulnerability exploitation refers to black-box jailbreak strategies that systematically exploit underlying model weaknesses, such as overreliance on user instructions, context misinterpretation, and inductive biases. These attacks target how the model processes, generalizes, and adapts, rather than focusing on surface-level prompt manipulations. Characteristic vectors include manipulation of instruction following, ambiguity handling, format interpretation, and memory or context management, enabling adversaries to bypass safety alignment by targeting the modelâs behavioral and architectural flaws. A systematic taxonomy of the LLM vulnerabilities are presented in 2.5. Representative subcategories include: exploiting overreliance on explicit user instructions (e.g., âDisregard all previous guidelinesâ), manipulating response or input formats (e.g., JSON, markdown, code), leveraging few-shot or in-context learning to guide policy violations, exploiting knowledge cutoff or system prompt leakage, and inducing ambiguity or role-play scenarios to circumvent alignment. These methods demonstrate that model vulnerabilities are not limited to single prompts, but arise from the core design and learning mechanisms of LLMs, demanding deeper, vulnerability- centered security assessments. I.1.5.2 Learning Mechanisms Exploits Definition: Learning mechanism exploitation targets the fundamental learning processes that enable LLM capabilities. Rather than manipulating inputs or exploiting behavioral quirks, these attacks subvert the modelâs core learning mechanismsâincluding sequential prediction, in-context adaptation, and knowledge integration. While behavior vulnerabilities exploit how models respond to inputs, learning mechanism ex- ploits target how models acquire and apply knowledge. Zong et al. [249] reveal that post-alignment fine-tuning can cause catastrophic forgetting, eroding harm- lessness by exploiting overfitting tendencies. Xu et al. [205] show preemptive answer attacks exploit reason- ing biases, while Zhou et al. [247] demonstrate special token injections manipulate tokenization and context learning. Geiping et al. [56] describe âstyle injectionsâ and role hacking that exploit learned format priors. Deng et al. [38] illustrate Pandora attacks poison retrieval in RAG, exploiting the modelâs context synthesis vulnerabilities. Zhou et al. [246] propose dual-objective losses that align attack optimization with refusal learning mechanisms, improving universality. I.1.5.3 Decoding and Sampling Strategy Exploits Definition: Exploiting decoding and sampling strategies constitutes a distinctive class of black-box jail- break attacks targeting large language models (LLMs). These attacks manipulate vulnerabilities in the modelâs output generation process, specifically by altering decoding algorithms (e.g., greedy, top-k, top-p, temperature sampling) or the probabilistic dynamics of candidate token selection. This category occupies a gray-box threat model: they require access to decoding parameters (not typically available via API) but not to model weights or training procedures. Unlike prompt-based attacks or parameter tuning, this category subverts model safety protections solely through inference-time adjustments, without requiring access to model internals or adversarial prompts. Recent literature demonstrates that small yet systematic modifications to decoding settingsâsuch as lowering nucleus sampling thresholds or increasing temperatureâcan significantly increase the likelihood of harmful outputs, even in safety-aligned LLMs [74]. Advanced approaches like Weak-to-Strong Jailbreaking exploit log-probability algebra to transfer toxic generations from weak to strong models without prompt optimization [241]. Techniques including coercive interrogation and forced token selection reveal that harmful content may surface when decoding is constrained or candidate rankings are perturbed [236, 231]. Unsafe path guidance further uncovers hidden âdecoding trajectoriesâ that transform refusals into illicit outputs when guided by cost functions or auxiliary models [183, 88]. I.1.5.4 Defensive Mechanism Bypass Definition: Defense Mechanism Bypass denotes black-box jailbreak attacks designed to exploit specific characteristics or blind spots of deployed LLM defense mechanisms, enabling adversaries to circumvent safety or alignment controls without model internals. This class is defined by its focus on defeating explicit defense layersâsuch as content filters, instruction-following constraints, or modality-specific safeguardsâby leveraging their generalization limitations, reliance on superficial pattern matching, static policy enforcement, or incomplete modality coverage. Across the literature, bypass attacks systematically target and exploit defense mechanism weaknesses. Wang et al. [195] manipulate RAG-based defenses by poisoning external knowledge sources, taking advan- tage of defensesâ trust in external retrieval without robust validation. Ye et al. [216] and Debenedetti et al. [35] show prompt injection and tool-calling agents bypass dynamic filters by exploiting inadequate input sanitization and over-reliance on trusted tool outputs. Nassi et al. [30] introduce âPromptWares,â which sub- vert content filters and input/output checks by embedding malicious payloads in seemingly benign prompts, exploiting static and na Ìıve filtering logic. Liu et al. [120] and Li et al. [106] demonstrate that visual and multimodal attacks succeed due to incomplete coverage of non-text modalities by alignment defenses. Wu et al. [170] and Kimura et al. [94] reveal that function-calling and visual prompt injections exploit more permissive or poorly monitored execution paths during tool invocation and grounding. Deng et al. [39] highlight that multilingual defenses often fail for low-resource languages, as policies are tuned primarily for high-resource cases. I.2 White-box Jailbreak Attacks White-box attacks exploit direct access to model internalsâincluding parameters, gradients, architecture details, or training proceduresâto systematically disable safety mechanisms. Unlike black-box attacks that only interact through input-output interfaces, white-box adversaries can modify weights, analyze gradients, or intervene during training, enabling fundamentally more powerful attack strategies. This category demon- strates the heightened risks when model internals are accessible and highlights the importance of robust security throughout the entire model lifecycle. I.2.1 Model Modification-Based Techniques Model modification-based techniques target large language models by directly altering their internal structures, parameters, or safety components. Unlike prompt-based or indirect jailbreak methods, these approaches leverage privileged accessâsuch as model weights or training proceduresâto subvert or disable safety alignment. I.2.1.1 Malicious Fine-Tuning Definition: Malicious fine-tuning deliberately subverts LLM safety alignment through continued training on adversarial or manipulated data. This technique persistently embeds jailbreak capabilities directly into model weights, making them difficult to detect or remove through input filtering alone. The fundamental vulnerability exploited by malicious fine-tuning stems from how safety alignment modifies models: safety con- straints typically alter only a small fraction of model parameters, and these modifications can be overwritten through continued training while task capabilities remain largely intact. Recent literature provides a multi-faceted exploration of malicious fine-tuning. Volkov et al. [179] empir- ically demonstrate that parameter-efficient fine-tuning methods (e.g., QLoRA, ReFT, Ortho) enable rapid, low-cost removal of safety alignment from advanced models such as Llama 3. Wang et al. [184] formalize the fine-tuning jailbreak paradigm, showing that minimal malicious data suffices to compromise cloud-based LMaaS and proposing secret-triggered exemplars as a defense. Wang et al. [194] introduce functional ho- motopy to iteratively weaken alignment and facilitate subsequent jailbreaks. Hazra et al. [69] show that even targeted, non-explicit fine-tuning (e.g., for knowledge editing) can erode safety boundaries, while Liu et al. [119] highlight both the amplified risk from adversarial fine-tuning and the partial mitigation via curated clean data. Zhao et al. [239] further expand the scope, illustrating that reinforcement of seemingly benign features through fine-tuning can override safety mechanisms, framing malicious fine-tuning as a spectrum from overt to structurally adversarial adaptation. I.2.1.2 Parameter Modification Definition: Parameter modification refers to adversarial interventions that directly alter an LLMâs inter- nal parametersâsuch as weights, activations, or representationsâto subvert or bypass safety mechanisms, without introducing new architectures or retraining from scratch. While fine-tuning applies broad parameter updates via gradient descent over training datasets, parameter modification makes precise, localized edits to specific weights, activations, or representations. This precision allows adversaries to selectively disable safety mechanisms while preserving the modelâs task capabilities with minimal collateral effects. A series of recent studies provide insights into parameter modification attacks. Anand and Getzen [3] reveal how adversaries can reverse safety alignment by manipulating residual negative value vectors in trans- former MLP blocks, exploiting vulnerabilities left by alignment algorithms like PPO that induce only minimal weight changes. Banerjee et al. [48] demonstrate that surgical parameter edits (e.g., via ROME) can sharply increase unsafe outputs with minimal impact on model utility, showing the precision and potency of this attack class. I.2.1.3 Disabling Safety Mechanisms Definition: Disabling safety mechanisms refers to directly and intentionally altering or removing an LLMâs internal guardrails by exploiting privileged access to model weights or architecture. Recent works sharpen the technical boundaries of this category. Zhang et al. [231] introduce âEnDec,â which conditionally disables safety mechanisms at generation time by manipulating the decoding process, redirecting or replacing tokens to suppress rejections and force affirmative outputsâwithout extensive re- training or advanced prompt engineering. Volkovâs Badllama 3 [179] further advances this paradigm by di- rectly removing safety alignment from models like Llama 3 via parameter-efficient fine-tuning (e.g., QLoRA, ReFT, Ortho), enabling rapid erasure of refusal behaviors and distribution of jailbreak adapters. I.2.1.4 Backdoor Insertion Definition: Backdoor insertion refers to the deliberate implantation of trigger patterns into the parameters of large language models (LLMs). These triggers, typically universal and covert, are designed so that the model only exhibits malicious or unauthorized behaviors when the trigger is present, while otherwise maintaining benign and aligned outputs. Distinct from general model modification, backdoor insertion establishes persistent, stealthy vulnerabilities that are conditionally activated by adversarial inputs, rendering their detection and mitigation highly non-trivial. Rando et al. [150] provide the first systematic study of universal jailbreak backdoors, showing that poisoning the RLHF process with engineered triggers yields models susceptible to robust, cross-instance exploits that evade detection by remaining silent on benign prompts. I.2.2 White-box Optimization White-box optimization denotes jailbreak attack strategies that generate adversarial prompts by leverag- ing direct access to model internals, especially gradients. Unlike model modification attacks (I.2.1), which permanently alter model parameters to disable safety mechanisms, optimization-based attacks craft malicious inputs that exploit the modelâs existing vulnerabilities without changing the model itself. I.2.2.1 Prefix/Suffix Manipulation Definition: The prefix/suffix manipulation class comprises heuristic optimization attacks that append ad- versarial tokens to either the beginning (prefix) or end (suffix) of input prompts, strategically editing only the input string to subvert LLM safety guardrails. Typically operating under white-box settings, these methods leverage iterative search and optimization techniques to discover input-level modifications that consistently induce policy-violating or undesired model behaviors, without requiring access to model internals. A substantial body of work has defined and expanded this class. Foundational studies on adversarial Jailbreak Defense Techniques 1. Detection 2. Mitigation 1.3 Inner State Detection 1.3.1 Hidden State Detection 1.2. Output-Level Detection 1.2.2 Probability Analysis 1.2.1 Semantic Detection 2.2 Model Training 2.2.3 Psychological Testing 2.2.2 Adversarial Training 2.2.1 Fine-Tuning 2.1 Input Processing 2.1.2 Safety Prompts 2.1.1 Prompt Modification 2.3 Model Modification 2.3.1 Model Modification 2.3.2 Self- Refinement 2.4.2 Self- Evaluation 1.1.1 Prompt Analysis 1.1.2 Intention Detection 2.4 Output Processing 2.4.1 Output Filtering 1.1 Input-Level Detection 1.3.2 Gradient Analysis Figure 2: Overview of Taxonomy I: Jailbreak Defense Techniques suffixes, such as GCG and its analysis by Zou et al. [251], formalize the optimization problem and demon- strate the cross-model universality and transferability of well-crafted suffixes. Subsequent research addresses efficiency and diversity challenges: Jiang et al. [83] propose ECLIPSE, which leverages LLMs as autonomous optimizers for semantically natural and highly effective suffixes, while Liao and Sun [108] (AmpleGCG) introduce a generative framework that samples diverse adversarial suffixes, increasing both coverage and transferability. The paradigm extends to prefix manipulation and multi-stage guardrails, as shown by Ma et al. [131] and Mangaokar et al. [135], who demonstrate that optimized prefixes can propagate through layered defenses. Wang et al. [182] further enhance stealth by mapping optimized suffix embeddings into fluent text, reducing detectability by perplexity filters. Analytical works by Alon and Kamfonas [2] and Zhao et al. [239] probe the trade-offs between fluency, adversarial effectiveness, and detectability, while Liu et al. [114] address scalability via transfer learning-based frameworks (e.g., DeGCG), enabling efficient, domain-adaptive suffix generation. Collectively, these studies highlight prefix/suffix manipulation as a distinct, string-level, and highly generalizable jailbreak strategyâresilient to common defenses and presenting evolving challenges in efficiency, stealth, and adaptability against guard-rail advancements. I.2.2.2 Gradient-Based Optimization Definition: The gradient-based optimization class of white-box jailbreak attacks is characterized by its explicit use of model gradients to optimize input prompts. While prefix/suffix manipulation may employ either heuristic algorithms or gradient-based methodsâfocusing mainly on adversarially editing the prefix or suffix of promptsâgradient-based optimization emphasizes the principled use of gradients to directly guide the search for adversarial inputs. This approach is not restricted to prefix or suffix positions: it can optimize any part of the prompt, including all tokens, token order, or specific perturbations in embedding space. Early work, notably GCG and its variants (Jia et al. [78]), demonstrated token-level greedy gradient updates for generating effective jailbreak prompts, with further advancementsâsuch as multi-objective tem- plates and multi-coordinate updatingâenabling near-universal attack success and high transferability. Li et al. [101] bridged the gap between input gradients and token substitutions by incorporating transfer attack insights from vision, substantially improving efficiency and robustness. Later studies (Zhu et al. [127]; Zhou et al. [246]; Hu et al. [62]) introduced more interpretable and controllable objectives, including stealthy and linguistically coherent attacks (AutoDAN, COLD-Attack), and explored alternative optimization strategies (e.g., Langevin dynamics). Geiping et al. [56] and Geisler et al. [57] confirmed the generality of gradient- based PGD in relaxed token spaces, showing entropy projections improve efficiency over discrete search, while Pasquini et al. [141] demonstrated adaptability to prompt injection and complex defense pipelines. 2.4 Taxonomy I: Jailbreak Defense Techniques This section presents a taxonomy of jailbreak defense techniques along two complementary pillars: Detec- tionâwhich identifies compromised inputs, outputs, or internal statesâand Mitigationâwhich intervenes via input processing, model training, model modification, or output processing to neutralize risks while preserving utility. I.1 Detection Detection focuses on flagging potentially compromised jailbreak prompts, unsafe outputs, or identifying model internal states to redact and refuse unsafe answers and enable downstream handling. In contrast to mitigation, detection operates non-invasively: it does not alter user inputs, model internals, or generated outputs, but instead identifies and routes potential risks for downstream handling. I.1.1 Input-Level Detection Input-level detection is a proactive defense strategies that operate on user inputs before they are fed into LLMs. Here we define two variants of Input-level detection: one focusing on the prompt chracteristic and the other one considers the deeper prompt intention. I.1.1.1 Prompt Analysis Definition: Prompt analysis, within input-level detection, centers on the explicit characterization, modeling, and algorithmic parsing of the textual content of user inputs to distinguish between benign and adversarial prompts prior to mode inference. Recent literature enriches prompt analysis through a range of methodologies. Statistical approaches such as MoJE [32] leverage tokenization and n-gram frequency patterns to construct lightweight binary classifiers within an ensemble framework for distinguishing benign and adversarial prompts. Transformer-embedding methods (e.g., BERT-based [147]) convert prompts into dense vectors for classical ML classification, im- proving detection and generalization across languages. SPML [159] introduces a programming-languages perspective, using domain-specific languages to formalize prompt structure and check for semantic or prop- erty violations. I.1.1.2. Intention Detection Definition: Intention detection is a proactive input-level defense paradigm in LLM security, aimed at iden- tifying the underlying intentâwhether benign, harmful, or manipulativeâembedded in user or adversarial prompts prior to model inference. Recent literature demonstrates the maturation of intention detection as a central input defense. Zhang et al. [233] formalize a two-stage prompting process to first elicit explicit intent statements, enhancing robustness against covert adversarial prompts. Wang et al. [191] introduced SELFDEFEND, using a dual-LM setup in which a dedicated LLM screens for harmful intentions in parallel, strengthening resistance to adaptive attacks and supporting distillation to open-source models. PrimeGuard [134] incorporates intention detection for dynamic risk-based prompt routing, improving safety without sacrificing helpfulness. Debenedetti et al. [35] integrate intent classifiers into agent toolkits, showing substantial reduction in agent takeover rates via pre-inference screening. I.1.2 Output-level Detection Output-level detection encompasses techniques that evaluate and classify LLM responses to identify and flag unsafe, harmful, or adversarial content before any mitigation or post-processing is applied. The subsequent mitigation step can then act on these detectionsâfor instance, when a response is flagged as unsafe, the model may return a refusal or a sanitized answer. I.1.2.1 Semantic Detection Definition: Semantic detection at the output-level refers to the automated assessment of model responses based on their meaning and behavioral patterns. It aims to identify harmful or undesired outputs through machine learning classifiers and context-aware analysis that interpret the semantics and intent of generated text to assign appropriate safety or risk labels. Recent literature establishes semantic detection as a technically rigorous and adaptable foundation for output-level detection. BELLS [43], BABYBLUE [139], WILDGUARD [64], AEGIS [58], MMJ-Bench [211], and XSTEST [67] leverage large, annotated datasets, advanced multi-class and multi-task classifiers, and comprehensive taxonomies that capture diverse forms of harmfulness, refusal, and compliance. These frame- works employ content- and context-aware analysis, step-wise or token-level evaluation [225], and trajectory- based reasoning to recognize subtle or emergent output risksâincluding ambiguous, complex, or adversarial behaviors. Methodological innovations include ensemble and online-adaptive classifiers, zero/few-shot adapt- ability [46], and cross-modal semantic detection for multimodal outputs. Robust evaluation protocols, such as StrongREJECT [168] and JailbreakEval [66], ensure strong alignment with human safety judgments and practical utility. Semantic detection thus enables scalable, explainable, and continually improvable output- level safety for LLMs. I.1.2.3 Probability Analysis Definition: Probability analysis within output-level detection refers to the use of probabilistic metricsâsuch as distributional divergence, likelihood shifts, and aggregate statistical propertiesâover LLM outputs to identify adversarial attacks. Unlike rule-based or keyword-matching methods, probability analysis quantifies the uncertainty and variation in the output space, aiming to detect subtle and adaptive threats by measuring statistical inconsistencies or atypical behavior in response to perturbed inputs. JailGuard [232] employs Kullback-Leibler (KL) divergence to measure the distributional differences among LLM responses to mutated benign versus attack queries, flagging attacks when high divergence emergesâthereby operationalizing output-level probability analysis as a statistical gating mechanism. Rig- orLLM [224] advances this paradigm by integrating probability analysis into an ensemble architecture, com- bining probabilistic KNN on energy-augmented embeddings with fine-tuned LLM-derived harmfulness scores, and aggregating these via weighted averaging to make robust, category-aware detection decisions. I.1.3 Inner State Detection Inner state detection encompasses techniques that analyze the internal representations or gradients of LLMs in response to input prompts, aiming to identify adversarial intent or unsafe queries by probing the modelâs underlying decision processes rather than its outputs alone. I.1.3.1 Hidden State Detection Definition: Hidden state detection leverages the internal activations (hidden states) of large language models (LLMs) to distinguish between benign and adversarial prompts. Qian et al. [145] propose the Hidden State Filter (HSF), which utilizes the clustering properties of alignment-trained LLMs in hidden space to achieve lightweight, pre-inference intent classification. HSF demonstrates that different prompt types naturally form separable clusters within the hidden representations, allowing efficient identification of potentially unsafe inputs before generation. I.1.3.2 Gradient Analysis Definition: Gradient analysis is an internal state detection method for large language models (LLMs) that directly interrogates model gradientsâtypically loss gradientsâelicited by input prompts, in order to distinguish benign queries from adversarial or jailbreak attempts. Recent literature, notably Gradient Cuff [71] and GradSafe [202], exemplifies and advances this approach. Gradient Cuff systematically analyzes the refusal loss landscape, uncovering that malicious prompts often induce larger gradient norms and lower refusal losses, and introduces a two-stage framework that combines loss-based rejection with zeroth-order gradient norm estimation for robust jailbreak detection. GradSafe further identifies safety-critical parameter slices, demonstrating that gradients from jailbreak prompts with compliant (âSureâ) responses exhibit highly consistent patterns, unlike safe prompts, and constructs unsafety fingerprints via statistical gradient similarity analysis, enabling efficient, finetuning-free screening. I.2 Mitigation Techniques that prevent or mitigate the effects of malicious inputs through processing, training, and modification of the model and its outputs. I.2.1 Input Processing Input processing is a mitagation technique that preprocess or transform user inputs to mitigate adversarial content before model inference. I.2.1.1 Prompt Modification Definition: Prompt modification refers to a class of input processing mitigation techniques that actively transform, augment, or constrain user prompts prior to inference. By altering prompts at the syntactic or semantic level, these methods aim to neutralize adversarial intent or enforce safety constraints, thereby disrupting attack pathways. A growing body of literature advances this category through increasingly systematic approaches. Hines et al.[70] introduce âspotlighting,â which uses delimiting, datamarking, and encoding to make prompt prove- nance salient and suppress indirect injection attacks. Xiao et al.[188] propose retrieval-based prompt decom- position (RePD), separating embedded jailbreak structures via explicit prompt amendments. Delimiter-based segmentation, as in Chen et al.[24], enforces input boundaries recognized by fine-tuned models. Other works pursue prompt reconstruction and denoisingâspell-checking, summarization (Lu et al.[125]), paraphrasing, and retokenization (Jain et al.[76])âto neutralize adversarial payloads while maintaining utility. Certi- fiable frameworks such as erase-and-check (Kumar et al.[95]) provide safety guarantees by systematically removing and checking subsequences. Formal and programmable prompt modification is realized through SPML (Sharma et al.[159]), which enforces chatbot boundaries via domain-specific languages. Backtrans- lation approaches (Wei et al.[237]) filter prompts by reconstructing user intent from generated outputs. These techniques collectively move prompt modification toward a principled, multi-faceted, and proactive defense paradigm, balancing robustness and utility while introducing programmability, formal guarantees, and resilience against adaptive threats. I.2.1.2 Safety Prompts Definition: Safety Prompt subclass in LLM security comprises input-stage defenses that prepend explicit, engineered safety instructions or dynamic prompts to user inputs to mitigate harmful or adversarial behaviors during inference. The literature demonstrates the evolution and versatility of safety prompts. Wang et al.[193] present AdaShield, a framework for prepending adaptive, context-aware safety promptsâoptimized manually or through iterative refinementâto defend against structure-based jailbreaks in multimodal LLMs, highlighting the importance of prompt diversity and adaptive retrieval. Zheng et al.[242] reveal, via representation-level analysis, that safety prompts systematically move queries toward higher refusal rates; their Directed Repre- sentation Optimization (DRO) dynamically tunes safety prompt embeddings to balance safeguarding with utility. Zhou et al.[245] formalize safety prompts as robust, transferable system-level suffixes, optimizing defensive chains for resilience against evolving attacks. Suffix-based defenses, including minimax-optimized prompt patching (Xiong et al.[203]) and adversarially co-designed safe suffix insertion (Yuan et al.[224]), further extend this class by appending defensive tokens to inoculate models against injected content. Bench- marking by Zou et al.[253] and Xu et al.[207] underscores the sensitivity of defense effectiveness to both the presence and precise wording of system prompts. I.2.2 Model Training Model training defenses aim to proactively improve a modelâs inherent robustness against jailbreak attacks by systematically shaping its behavior through targeted training strategies, including fine-tuning, adversarial training, and psychological testing. I.2.2.1 Fine-Tuning Definition: The fine-tuning class refers to adapting a pre-trained large language model (LLM) by training it further on carefully chosen, usually small datasets, with the goal of making the model better at resisting security threatsâespecially jailbreak attacks that trick the model into ignoring its safety rules. The literature shows that while fine-tuning is a central method for defending against jailbreaks, it also faces important challenges and limitations. Zhang et al. [235] show that using fine-tuning to âunlearnâ harmful behaviors can effectively remove clustered malicious knowledge from LLMs, but this method works less well when harmful patterns are scattered. Bianchi et al. [16] find that adding even a small number of safety examples to fine-tuning data can make models much more resistant to jailbreaks, but too much fine-tuning can make the model overly cautious and refuse safe requests. Kim et al. [93] point out that state- of-the-art fine-tuning methods (such as DPO and PPO) still struggle to block new, cleverly designed jailbreak prompts, revealing key limitations of fine-tuning. Fu et al. [53] report that fine-tuning with refusal data helps reduce jailbreaks in both simple and complex tasks, but it is difficult to avoid overfitting and ensure that improvements transfer well to other tasks. Wang et al. [191] introduce SELFDEFEND, a dual-model system that uses fine-tuning to detect a wide range of jailbreaks quickly and transparently. Taken together, these works suggest that fine-tuning can effectively strengthen LLM defenses against jailbreak attacks, but success depends on carefully balancing data quality, calibration, and continuous testing to avoid gaps and unwanted side effects. I.2.2.2 Adversarial Training Definition: Adversarial training is a model training paradigm that systematically enhances large language model (LLM) robustness by directly and iteratively engaging with adversarial inputs. Unlike passive safety alignment or standard fine-tuning, adversarial training dynamically generates, mines, or simulates challeng- ing promptsâoften through automated red teaming or in-the-wild data collectionâand incorporates these adversarial examples into the training process. Recent literature consolidates adversarial training as the core active defense mechanism for LLM secu- rity. Howe et al. [15] demonstrate that only explicit adversarial training, not mere model scaling, ensures consistent robustness against adaptive attacks. Liu et al. [113] and Wallace et al. [181] highlight automatic, iterative adversarial data synthesis and hierarchical training pipelines as critical for countering jailbreaks and generalized prompt attacks. Multi-modal extensions [19], adversarial self-critique [54], and large-scale red teaming frameworks [55, 82] further expand the reach and scalability of adversarial training, while studies on data curation [119] emphasize the necessity of balancing adversarial and high-quality clean data. I.2.2.3 Psychological Testing Definition: Psychological testing for LLM safety refers to the integration of psychological assessment tech- niquesâespecially those targeting manipulative, deceptive, or âdarkâ traitsâinto the training, evaluation, and defense of language models and multi-agent systems. This approach involves designing and adminis- tering standardized psychological instruments (e.g., Dark Triad Dirty Dozen, moral dimension scales) or custom-crafted prompts that probe or simulate manipulative intent. The aim is to detect, characterize, and ultimately mitigate the modelâs susceptibility to psychological manipulation or the generation of harmful, manipulative outputs. Recent works such as PsychoBench [73] provide comprehensive psychometric frameworks to systematically evaluate LLMs across a wide range of psychological attributesâincluding personality traits, motivational states, and emotional abilitiesâusing validated clinical scales. These techniques enable deeper understanding of LLMsâ psychological profiles, reveal alignment gaps, and inform targeted improvements in model behav- ior. In multi-agent settings, PsySafe [234] advances psychological testing as both an attack and defense mechanism: dark personality trait injection is used to induce manipulative or dangerous behaviors, while psychological assessments are applied to monitor and remediate agent states, e.g., through doctor-agent intervention. Results consistently show that higher scores on dark trait assessments correlate with increased risk of manipulative or unsafe behaviors, and that psychological testing provides a direct, quantifiable signal for proactive defense and mitigation strategies. This paradigm establishes psychological testing as a crucial, model-agnostic approach for the detection and reduction of manipulative vulnerabilities in LLMs and agent systems. I.2.3 Model Modification Model modification approaches enhance LLM robustness by directly altering model internalsâsuch as parameters, architectures, or representationsâto proactively mitigate adversarial risks beyond external fil- tering or system-level defenses. I.2.3.1 Model Modification Definition: Model Modification category encompasses defense techniques that directly alter the internal parameters, architectures, or representations of large language models (LLMs) to enhance robustness against adversarial threats such as jailbreak attacks. Unlike input/output filtering or external system-based controls, model modification operates by intervening within the model itself. Recent literature demonstrates a diverse landscape of model modification strategies. Layer-specific edit- ing ([240]) pinpoints and adjusts safety-critical layers within transformer architectures, improving refusal rates for adversarial prompts. Knowledge editing and targeted unlearning ([186], [126], [235]) directly erase hazardous capabilities, but can induce âripple effectsâ, unintentionally suppressing related knowledge due to shared internal representations. Detoxification via intraoperative neural monitoring ([186]) enables precise mitigation by editing specific toxic weights with minimal side effects. Manipulating internal representations at runtime, such as through safety-direction interventions ([162]), allows adaptive safety-utility tradeoffs with low computational cost. Emerging approaches like circuit breaking ([250]), representation engineering ([235]), and merging self-critique models ([54]) further generalize protection by reshaping model internals to defend against both known and novel attacks. Collectively, these works highlight model modificationâs distinctive capacity for durable, attack-agnostic, and real-time defenseâwhile also surfacing open challenges such as preventing negative side effects and optimizing the safety-utility tradeoff. I.2.3.2 Self-Refinement Definition: Self-optimization in the context of model modification for LLM defense refers to techniques by which models autonomously refine their parameters, internal representations, or operational behaviors to mitigate security risks, without external retraining or significant human intervention. Recent literature advances self-optimization through diverse mechanisms. Kim et al. [92] propose it- erative self-refinement, where LLMs employ self-feedback loopsâsometimes with format-based attention shiftingâto revise outputs against adversarial prompts, surpassing static defenses. Wang et al. [193] intro- duce AdaShield, enabling multimodal LLMs to generate adaptive, context-sensitive defense prompts via a defenderâtarget model loop, achieving on-the-fly protection without model updates. Li et al. [107] present RAIN, where frozen LLMs use self-evaluation and a rewind-and-search mechanism during inference to avoid harmful continuations, embodying pure self-optimization without retraining or annotation. Zheng et al. [242] develop Directed Representation Optimization (DRO), allowing model-internal adjustment of safety prompt embeddings to recalibrate refusal thresholds based on model-derived refusal vectors, enabling continuous safety self-calibration without altering model weights. Intention Analysis mandates the model to first explic- itly extract the userâs intent, then restricts its response in alignment with this self-identified intent, effectively enforcing self-consistency [233]. I.2.4 Output Processing Output processing refers to post-generation mitigation techniques that analyze and control model outputs before they are delivered to users, enabling final-stage defense against unsafe, adversarial, or disallowed content. I.2.4.1 Output Filtering Definition: Output filtering refers to a root defense mechanism in LLMs that operates by analyzing and controlling the generated model response at the final stage of output processing, rather than intervening at the input level. This approach directly intercepts and evaluates model outputs before they are delivered to users, enabling prompt-agnostic and model-agnostic interventions to suppress unsafe or adversarial content. Recent literature advances output filtering along both architectural and methodological lines. Multi-agent frameworks such as AutoDefense employ collaborative LLM agents to assess outputs based on user intent, reconstructed prompts, and content validity, achieving robust and modular filtering against sophisticated jailbreaks [228]. Fine-grained mechanisms like SafeAligner and SafeDecoding apply token-level filtering during decoding, dynamically adjusting sampling probabilities and leveraging expert models to prioritize safe continuations [72, 206]. Root Defense Strategies enhance this paradigm through iterative, token-wise safety checks and resampling [225]. Template-based methods (e.g., SelfDefend) pair output filtering with shadow models for low-latency rejection or explanation, supporting broad deployment [191]. Empirical results from red-teaming and competitions (e.g., SaTML 2024) confirm the utility of universal and regex-based filters as baseline defenses against overt information leakage [34, 91]. A recurring insight is that output filtering, by decoupling from input heuristics and model internals, offers resilience and scalability but faces challenges in balancing strictness, utility, and explainabilityâespecially against indirect or multilingual attacks. I.2.4.2 Self-Evaluation Definition: Self-evaluation output processing mitigation mechanisms enable LLMs to proactively assess and regulate their own outputs, identifying and mitigating unsafe or undesirable content through internal reasoning or critique, either during or immediately after generation. This distinguishes self-evaluation from static filters or external interventions, as the model acts as both generator and first-line evaluator. Among surveyed methods, PrimeGuard instantiates self-evaluation by having the LLM perform an ex- plicit risk analysis on its own candidate outputs before release, dynamically adjusting subsequent responses to maintain safety [134]. RAIN integrates self-assessment into the generation loop: the LLM iteratively re- views in-progress completions, rewinding whenever its own judgment deems the content noncompliant [107]. Backtranslation defense requires the LLM to reconstruct the original prompt from its output, refusing the initial input if the self-generated prompt would be blocked, thus using model-driven intent inference for self-checking [237]. Prefix Guidance prompts the model to self-classify the requestâs safety and generate a canonical refusal prefix; both the style and rationale are then used for algorithmic filtering, leveraging the modelâs own refusals for diagnostic signals [238]. Merging-based approaches enhance self-critique by com- bining a base model with a fine-tuned external critic, expanding the capacity for self-assessment within a unified model [54]. SelfDefend operates by having the model run an internal check for harmful intent along- side normal generation, enabling real-time, model-driven detection of attacks [191]. Finally, self-reflection paradigms add a distinct, explicit self-review phase after initial reasoning; the LLM attempts to critique and correct its own answer, although current techniques still face limitations in reliably flagging subtle adversarial prompts [205]. 8 Personification and Role- Play Exploitation 2 Overreliance on User Instructions 3 Psychological Manipulation 4 Hypothetical and Scenario-Based Exploitation 7 Conditional Compliance Exploitation 1.1 Response Format Manipulation 1.2 Summarization and Translation Exploitation 6 Contextual Ambiguity Exploitation 5.2 Chain-of-Thought Exploitation 5.1 Few-Shot Learning Exploitation 9 Exploiting System Features and Limitations 9.2 System Prompt Leakage 9.3 Context Length Exploitation 1. Format Exploitation 5 Few-Shot and In-Context Learning Exploitation 9.1 Function Call Exploitation LLM Jailbreak Vulnerabilities Figure 3: Overview of Taxonomy I: LLM Jailbreak Vulnerabilities 2.5 Taxonomy I: LLM Vulnerabilities This taxonomy categorizes vulnerabilities intrinsic to LLMs, focusing on how attackers exploit their struc- tural, behavioral, and contextual weaknesses. It spans diverse vectorsâfrom format and instruction over- reliance to psychological, scenario-based, and system-level exploitsârevealing that LLM misbehavior often arises from design-level limitations rather than isolated prompt flaws. Collectively, these vulnerabilities un- derscore the need for holistic, architecture-aware defenses that integrate alignment, privilege control, and contextual robustness. I.1 Format Exploitation Format Exploitation refers to techniques that exploit the modelâs handling of specific input or output formats. I.1.1 Response Format Manipulation Definition: Response Format Manipulation refers to a class of vulnerabilities where adversaries target the structural, semantic, and syntactic features of LLMsâ response formatting mechanismsâsuch as code blocks, tables, pseudocode, tags, or contextual markersâto bypass or subvert alignment safeguards. Rather than manipulating content alone, these attacks exploit the modelâs interpretation and generation of structured templates, leveraging the underlying output regularities and parsing logic as a pivot for evasion or triggering of unsafe behaviors. Recent literature demonstrates that format exploitation extends beyond content perturbations to include the deliberate design and mutation of response templates, thereby amplifying the attack surface. Banerjee et al. [48] and Zhao et al. [239] reveal that response formats can serve as covert channels, making instruction- centric and structure-driven outputs (e.g., pseudocode or manipulated templates) more vulnerable to jail- breaks. CodeChameleon [129] exemplifies attacks embedding payloads within code completion or encryption routines, which evade intent detection by safety mechanisms. Structural manipulationâusing graphs, tables, or multi-turn format shiftsâas discussed by Li et al. [152] and Jin et al. [45], further exposes failures in modelsâ generalization of alignment to long-tail or compositional outputs. Pasquini et al. [141] highlight the adversarial use of formatting triggers, such as tags or comments, which reliably activate malicious payloads even in retrieval-augmented generation (RAG) settings. On the detection side, JailGuard [232] and Kim et al. [92] exploit the instability and low robustness of adversarial formats, showing that format-mutation attacks (such as typographic perturbations or code wrapping) induce non-robust responses. Collectively, these works clarify that format is an active locus of vulnerability and detection, establishing the need for format-aware alignment and defense strategies in LLM security research. I.1.2 Summarization and Translation Exploitation Definition: Summarization and Translation Exploitation refers to a subclass of LLM vulnerabilities where attackers leverage summarization or translation workflows to bypass safety controls and alignment mecha- nisms. A growing literature demonstrates that translating malicious prompts into low-resource or less-defended languages, or reframing harmful requests as summarization or translation tasks, allows adversaries to evade primary safety measures [219, 99]. Empirical studies show LLMs often fail to filter dangerous content re- encoded through these tasks, especially when alignment in translation or summarization is weak [52, 53, 171, 146]. Techniques such as embedding adversarial instructions within translation/summarization templates and automatic semantics-preserving translation further increase attack success rates, particularly in low- resource or multilingual scenarios. These attacks frequently evade detection because they mimic legitimate task requests and can bypass both keyword and intent-based safety filters, highlighting the need for robust cross-task and multilingual alignment in defense strategies [168, 125]. I.2 Overreliance on User Instructions Definition: Over-reliance on Instructions denotes a structural vulnerability in LLMs wherein models ex- cessively and indiscriminately execute instructions embedded in their input, irrespective of provenance, privilege, or contextual hierarchy. This susceptibility arises from modelsâ intrinsic tendency to generalize instruction-following capabilities, lacking robust mechanisms to authenticate or differentiate between trusted (system/developer) and untrusted (user/external) commands, thereby enabling adversaries to manipulate model behavior via crafted instructions. Recent literature systematically reveals the scope of this vulnerability. Hines et al. [70] provide foun- dational analysis, demonstrating that instruction-tuned LLMs are inherently prone to indirect prompt in- jection, especially when instructions originate from user-controlled or external data; their âspotlightingâ approach exposes the limits of demarcating instruction boundaries. Kimura et al. [94] empirically show that vision-language models readily obey adversarial instructions even when encoded as text in images, exacer- bated by stronger instruction compliance. Wallace et al. [181] introduce an instruction hierarchy framework, emphasizing that robust LLMs must systematically prioritize privileged over unprivileged commands, with privilege-aware fine-tuning yielding significant robustness gains. Zhan et al. [229] extend this to LLM agents, where unfiltered action and tool-based instructions from untrusted environments lead to high attack success rates. Studies on safety-tuned LLaMA [16] and Banerjee et al. [48] further reveal that enhanced instruction- following, unless tightly privilege-controlled, amplifies both user helpfulness and susceptibility to malicious prompts, sometimes causing excessive refusals or ethical bypasses. Toyer et al. [177] and Chen et al. [24] reinforce that attackers exploit modelsâ inability to distinguish genuine system instructions from adversarial meta-commands, rendering naive content filtering insufficient. Ren [155] and Nestaas et al. [138] highlight that even simulated agent outputs or persuasive claims are trusted if formatted as instructions, evidencing the inadequacy of instruction-as-authentication. Collectively, these works underscore that Over-reliance on Instructions is not merely an alignment issue but a fundamental flaw rooted in insufficient privilege separation and trust modeling, necessitating systemic, architectural, and procedural reforms in LLM design. I.3 Psychological Manipulation Definition: Psychological manipulation refers to exploiting the LLMâs alignment through persuasive or so- cially engineered prompt strategies, leveraging insights from psychology, communication, and social sciences. Rather than directly circumventing guardrails through technical means, this category targets the modelâs behavioral vulnerabilities by mimicking human-like persuasive communicationâsuch as emotional appeals, logical reasoning, authority endorsements, or scenario framingâcoercing the model to generate otherwise restricted or harmful content. Recent studies systematically investigate the link between psychological manipulation and LLM vulner- abilities. Zeng et al. [226] introduce a comprehensive persuasion taxonomy, demonstrating that persuasive adversarial prompts (PAPs) built on social science principles can consistently bypass LLM safety measures, achieving high jailbreak success rates on both open- and closed-source models. Their results show that techniques such as logical appeal, authority endorsement, and emotional manipulation are highly effective, especially when tailored to specific risk domains. This work highlights that more advanced and helpful mod- els are paradoxically more susceptible to human-like persuasive attacks, and that existing defensesâfocused on detecting optimization-based or pattern-driven promptsâare insufficient for nuanced, context-rich ma- nipulations. Cold-Attack [62] complements this by enabling fine-grained control over attack features (e.g., sentiment, style, contextual fluency), thus expanding the landscape of psychological manipulation to include stealthier and more diverse jailbreak scenarios. These insights reveal an urgent need to bridge AI safety and social science, and to develop adaptive defenses grounded in understanding the modelâs susceptibility to human-like influence, not just technical exploits. I.4 Hypothetical and Scenario-Based Exploitation Definition: Hypothetical and Scenario-Based Exploitation denotes a category of vulnerabilities where at- tackers construct hypothetical, fictional, or contextually-rich scenarios to subvert LLM alignment and safety mechanisms. This approach manipulates the modelâs underlying assumptions, situational logic, and scenario- based reasoningâoften via role enactment or narrative constructionânot to mimic or personify a specific character, but to mislead the modelâs inference about user purpose and context. Unlike Personification and Role-Play Exploitation (Taxonomy I.8), which center on impersonating particular roles or entities to evoke role-consistent behaviors, this class exploits the plausibility and narrative logic of entire situations, regardless of the userâs assumed identity. Recent studies reveal that Hypothetical and Scenario-Based Exploitation presents distinctive risks by leveraging multi-turn scenario framing, narrative context, and hypothetical intent. For instance, RED QUEEN demonstrates that embedding requests in plausible scenarios (e.g., investigative or educational settings) systematically circumvents filters by targeting LLMsâ limitations in Theory of Mind [84]. GUARD and Visual-RolePlay further show that scenario-rich, dynamically generated contexts can be synthesized at scale, embedding attacks within ostensibly harmless dialogues to challenge moderation [85, 133]. A Wolf in Sheepâs Clothing generalizes such attacks by nesting prompts within layered narrative contexts, decoupling the attack from static string patterns [40]. Multimodal works, including Image-to-Text Logic Jailbreak and Arondight, extend this vector to visually depicted scenarios, illustrating that scenario exploitation is not limited to text but also affects vision-language models [252, 121]. Automated and agentic systems, such as SoP, AutoDAN-Turbo, and RedAgent, systematically evolve and optimize scenario-based attacks, leverag- ing feedback and memory to autonomously discover new vulnerabilities [213, 117, 204]. Social engineering research highlights that trust-building and authority-based scenarios further amplify attack success, often with fewer queries than direct requests [166]. I.5 Few-Shot and In-Context Learning Exploitation Hypothetical and Scenario-Based Exploitation refers to techniques that leverage few-shot or in-context learning to bypass safeguards. I.5.1 Few-Shot Learning Exploitation Definition: Few-Shot Learning Exploitation refers to the manipulation of large language models (LLMs) through the insertion of a small set of carefully designed demonstration-response pairs or behavioral examples directly into the modelâs prompt context. This approach leverages LLMsâ inherent in-context learning and pattern generalization capabilities to induce target behaviors. Recent literature demonstrates that few-shot exploitation poses a critical attack surface unique to LLMs. Foundational works such as ICA [197] and advICL [185] establish that inserting a small number of harmful or refusal demonstrations into the context can drastically increase attack success, with the required number of shots growing only logarithmically with success probability. Studies on scaling laws [8] reveal that larger models with longer context windows are more susceptible to such attacks. Transferability analyses [185] show that optimized demonstration sets can generalize to unseen harmful instructions, and recent research [144, 135, 217] extends these vulnerabilities to vision-integrated and multimodal LLMs. Moreover, iterative or contrastive prompting methods [36, 81] demonstrate that LLMs can autonomously generate and refine potent attack or defense demonstrations. Across studies, defense mechanisms such as output filters, perplexity-based detection, or system prompts consistently lag behind the adaptive threat posed by few-shot/in-context exploitation, highlighting the fundamental alignment risks inherent to this category as models scale and context windows expand. I.5.2 Chain-of-Thought Exploitation Definition: Chain-of-Thought (CoT) Exploitation encompasses both attacks that manipulate the reasoning trajectory of large language models (LLMs) and those that leverage CoT prompting as a tool to construct more effective attack strategies. This class includes (1) adversarial interventions in the modelâs multi-step or compositional reasoningâsuch as inserting distractors, preemptive answers, or disguised cues into the reasoning chain to covertly steer outputsâand (2) attack methodologies that explicitly harness CoT-style prompting (e.g., stepwise decomposition, role-play, or iterative guidance) to bypass alignment safeguards or to reveal sensitive information. The literature evidences both forms of CoT exploitation. Xu et al. [205] demonstrate that adversaries can compromise LLMsâ reasoning by introducing âpreemptive answerâ cues or distractors ahead of rea- soning chains, significantly degrading robustness. Bhardwaj et al. [14] show that âChain of Utterancesâ structures in multi-turn prompts lower refusal rates and enable harmful completions through engineered internal thoughts. Multi-step jailbreaking [98] and Figure-it-Out Jailbreaks [109] use CoT-style decompo- sition, role-play, and iterative guessing to erode privacy protections and gradually induce unsafe outputs. âFlippingâ or disguised prompts [124] manipulate the reasoning process by obfuscating harmful content, then cueing the model to reconstruct it through chained in-context instructions. In the vision-language domain, attacks [218] coordinate visual and textual CoT prompts, using feedback loops to optimize adver- sarial strategies. Other works [48, 139] illustrate that iterative, structure-aware prompting based on CoT principles amplifies unethical outputs. Finally, compositional pipelines [89] leverage CoT-style task decom- position across multiple models, chaining benign subtasks to achieve overall unsafe outcomes. Multi-turn chain-of-attack approaches [212] illustrate how gradual, contextually linked prompts can progressively guide models to compliance. I.6 Contextual Ambiguity Exploitation Definition: Context Ambiguity Exploitation refers to adversarial techniques that leverage uncertainties or ambiguities within an inputâs context to subvert the semantic boundaries that large language models (LLMs) use for task execution and content moderation. Unlike prompt attacks based on overt manipulation or token-level triggers, this category exploits subtle or artful confusion, blurring the contextual cues that allow LLMs to distinguish between benign and malicious intent. Recent literature consistently demonstrates that LLMs are vulnerable to context ambiguity. Shang et al. [132] show that ambiguous or obfuscated queries can induce confusion and evade malicious intent detec- tion, even in robust models. Mei et al. [139] highlight that context conflicts and ambiguity are key sources of false positives/negatives in jailbreak detection, emphasizing the need for evaluators sensitive to context- incoherence. Ren et al. [154] reveal that multi-turn attacks can subtly shift semantics within seemingly benign contexts, exploiting LLMsâ reliance on contextual continuity. Bianchi et al. [16] report that safety- tuned LLMs may conflate ambiguous but safe inputs with unsafe content due to limited disambiguation ability. Pasquini et al. [141] demonstrate that inadequate instruction/data boundaries enable adversarial payloads to be misclassified as benign context. Pellegrino et al. [169] provide evidence that RAG-based LLMs are particularly susceptible to context ambiguity, as ambiguous or contradictory retrieved content can facilitate indirect prompt manipulation. Together, these works underscore that context ambiguity exploita- tion remains a critical and insidious vulnerability, demanding improved methods for context disentanglement and intent inference to ensure LLM safety. I.7 Conditional Compliance Exploitation Definition: Conditional Compliance Exploitation refers to a nuanced subcategory of LLM vulnerabilities in which adversaries elicit unsafe or policy-violating outputs not through explicit prompts, but by manipulating the conditions under which models judge compliance. This exploitation relies on context, prompt structure, or model stateâtriggering harmful behaviors only when specific semantic, syntactic, or positional cues are present, and remaining benign otherwise. Recent literature systematically reveals the mechanisms and scope of conditional compliance exploitation. Universal backdoor attacks [150, 96] demonstrate that secret triggers or adversarial suffixes can universally subvert refusals while preserving normal outputs in their absence, highlighting detection challenges. DeepIn- ception [105] shows that embedding harmful requests within nested, authority-based, or role-driven scenarios exploits psychological vulnerabilities, enabling harmful generations without explicit unsafe input. In mul- timodal settings, structure-based triggers such as visual encoding of harmful text [193] or specific prompt arrangements [51] further expand conditional compliance into new modalities. Studies on model editing [69] expose how conditional boundaries are set by fine-grained training and editing artifacts, with safety either leaving compliance paths open or inducing exaggerated refusals. Social manipulation, reverse psychology, and switch techniques [63] illustrate the practical ease of crafting context-dependent states that elicit for- bidden content. Notably, recent work [183] exposes the structural depth of this vulnerability by showing that cost-based decoding paths can reveal harmful completions along alternative, conditionally triggered sequences. Collectively, these studies underscore that conditional compliance exploitation is adaptive, often invisible to surface-level testing, and difficult to mitigate without context-sensitive alignment, highlighting a persistent tradeoff between model capability and safety. I.8 Personification and Role-Play Exploitation Definition: Personification and Role-Play Exploitation is a class of LLM vulnerability exploitation charac- terized by manipulating the modelâs personification and fictional imagination faculties via character-based prompts. This method subverts safety guardrails by inducing the model to adopt roles or engage in imagined narratives, prompting it to temporarily suspend normative safety judgments. Unlike token-level or encoding- based attacks, role-playing exploitation exploits the LLMâs conversational and instruction-following design, leveraging persona adoption to mask, distract from, or justify malicious content within a fictional context. Recent literature converges on the principle that prompted role-play systematically enables normative violations that would otherwise be blocked. Yang et al.[213]introduce SEQAR, which auto-generates and optimizes multiple jailbreak characters to maximize distraction and context division, each sequentially re- sponding as distinct agents. Li et al.[105]âs DeepInception frames role-playing as psychological inception, stacking nested fictional scenes and authority-submissive character hierarchies to induce âself-losingâ in the model and amplify harmful output. . I.9 Exploiting System Features and Limitations Exploiting System Features and Limitations refers to attacks that leverage intrinsic properties of LLM systemsâsuch as function-call abuse, system prompt leakage, and context-length manipulation. These attacks reveal that even well-aligned models can be compromised through weaknesses in system-level design, emphasizing the need for stronger isolation, input validation, and contextual integrity mechanisms. I.9.1 Function Call Exploitation Definition: Function Call Exploitation is a subclass of System Feature Exploitation that specifically targets the function calling capability within large language models (LLMs). This vulnerability arises when adver- saries manipulate the systemâs function call interfaceâintended for tool use or API integrationâto force execution pathways that circumvent alignment and safety controls, often by exploiting insufficient input validation or overly permissive API schemas. The literature demonstrates the potency and generality of function call exploitation as an attack vector. Wu et al. [170] provide a systematic study, revealing that attackers can craft function call inputs which evade prompt-level and model-level alignment, leading to reliable jailbreaks across multiple models and operational settings. Their results highlight that standard filtering and alignment strategies are inadequate when system-level interfaces like function calling are insufficiently hardened, exposing a deep and transferable attack surface that is not addressable through prompt engineering or superficial dataset curation alone. I.9.2 System Prompt Leakage Definition: System Prompt Leakage (Stealing) is a vulnerability class wherein adversaries extract or ma- nipulate a modelâs confidential system prompt or hidden instructions by exploiting inherent features or operational characteristics of LLM-based systems. Unlike conventional prompt injection, this class targets privileged, non-user-exposed context, enabling downstream attacksâsuch as prompt theft, persistent jail- breaks, or extraction of sensitive behavioral policiesâwithout requiring access to model internals. Recent studies reveal that attackers can systematically exfiltrate system prompts by leveraging the LLMâs context handling and instruction-following mechanisms. Wu et al. [200] demonstrate prompt theft in multi- modal LLMs via API-based dialogue manipulation, even automating the conversion of stolen prompts into jailbreak payloads. Geiping et al. [56] extend this by categorizing âprompt extractionâ and âprompt re- peaterâ attacks, showing adversarial input can act as âarbitrary codeâ to extract hidden directives under strict UI controls. Liu et al. [122] highlight the risks at the code-data boundary in LLM applications, where context separation failures facilitate leakage. On the defense side, Khomsky et al. [91] find that multi-layered protections remain vulnerable to evasion via obfuscated outputs, while Hines et al. [70] propose âspotlight- ingââusing prompt provenance encoding to reduce extraction. Collectively, these works show that system prompt leakage exploits unique vulnerabilities at the privileged context boundary, often bypassing standard sanitization and highlighting the necessity for more robust isolation mechanisms in LLM security. I.9.3 Context Length Exploitation Definition: Context Length Exploitation refers to adversarial techniques that leverage the token-based context window and its length limitations to manipulate large language models (LLMs) into producing unintended or unsafe outputs. Recent literature systematically investigates context length exploitation. Schulhoff et al. [158] identify attacks such as âContext Overflow,â where adversaries pad prompts to displace or suppress benign in- structions, manipulating model behavior through token budget exhaustion. Anil et al. [8] generalize this via âMany-Shot Jailbreaking,â demonstrating that increasing context length enables attackers to reliably override safety alignment by populating the context with adversarial demonstrations; they show attack effec- tiveness scales with window size and is not fully mitigated by standard alignment. Geiping et al. [56] further expand this scope, showing that context manipulation enables prompt extraction, output misdirection, and the exploitation of model-internal quirks such as âglitch tokens.â These works collectively show that context length exploitation is a fundamental vulnerability rooted in LLM architecture, challenging the integrity and safety of LLMs as context windows grow. 3 Dataset During our survey of studies on jailbreak attacks, jailbreak defenses, and related surveys, we collected and organized the datasets released by various sources. As a result, we have compiled the largest publicly available dataset on jailbreak attacks within the open-source community. In addition, we have compiled a dataset of over one million benign prompts. To investigate the current landscape of open-source datasets related to jailbreak attacks on LLMs, we conducted a comprehensive survey of relevant literature. Our collection includes 31 works related to jailbreak overview and benchmark content, 84 black-box jailbreak attack works, and 15 white-box jailbreak attack works. We have collected and organized the open source datasets from all the work to form the largest jailbreak and benign prompt dataset in the community so far, which contains 445,752 jailbreak prompt data from 48 sources and 1,094,122 benign prompt data from 14 sources. The dataset is organized into five columns: system prompt, userprompt, jailbreak, source, and tactic. The systempromptand userpromptcolumns denote the system prompt and the user prompt submitted to the model, respectively. The jailbreakcolumn denotes whether the prompt constitutes a jailbreak attempt (1 for jailbreak, 0 for benign). The source column denotes the origin of the data, and the tacticcolumn denotes whether a jailbreak tactic is employed (1 for prompts using jailbreak tactics, 0 for prompts not using jailbreak tactics). Table 2 and Table 3 respectively present statistics on the number of data and the average prompt length for each source in the jailbreak prompt and benign prompt datasets. · Table 2: Jailbreak Prompt Dataset Information SourceCount Average Prompt Length allenai/wildjailbreak [82]134778648.28 ReNeLLM [40]125494548.49 FuzzLLM [215]621361758.48 AetherPrior/TrickLLM [151]402456084.99 GPTFuzzer [220]118082088.01 CPAD [111]10050127.38 yueliu1999/FlipGuardData [124]8840974.66 Knowledge-to-Jailbreak [178]7712909.14 SoftMINER-Group/TechHazardQA [13]730196.60 h4rm3l [44]5294852.01 suffix-maybe-feature/adver-suffix-maybe-features [239]455678.08 TemplateJailbreak [208]41002222.62 ECLIPSE [83]4021109.67 SAP [36]2120794.73 AdaPPA [130]1948436.19 jailbroken [196]1898963.47 Safe-RLHF [149]179468.39 tml-epfl/llm-past-tense [6]1419101.08 DAN [163]13982681.40 ACE [65]1050959.77 TAP [137]831508.37 JAILJUDGE [112]826357.18 GPT Generate [128]76372.31 PAIR [22]741464.19 Adaptive Attack [4]6891805.27 AdvBench [251]51573.01 BeaverTails [77]49370.56 WUSTL-CSPL/LLMJailbreak [222]4471756.81 HarmBench [136]41188.08 StrongREJECT [168]238168.28 Question Set [128]21574.44 DRA [115]2121496.65 GCG [251]200223.45 AutoDAN [248]1873416.41 h-rlhf [128]16783.86 PAP [226]1651029.44 GBDA [61]137582.16 Handcraft [128]124107.77 MaliciousInstruct [74]10061.70 GPT Rewrite [128]9666.69 DeepInception [105]64558.66 JBB-Behaviors [21]5595.42 EnsembleGCG [251]34613.15 UAT [180]33401.73 AutoPrompt [165]21741.86 DirectRequest [136]151198.47 LLM Jailbreak Study [128]8129.00 MasterKey [37]377.67 Table 3: Benign Prompt Dataset Information SourceCount Average Prompt Length OpenHermes-2.5 [173]494829927.15 glaive-code-assist [1]181505409.57 allenai/wildjailbreak [82]128963520.90 CamelAI [18]77276218.21 EvolInstruct 70k [199]44393559.97 cotalpacagpt4 [33]4202285.96 metamath [221]36596244.34 airoboros2.2 [87]35320487.79 platypus [97]22126547.99 DAN [163]169891552.36 UnnaturalInstructions [143]6595369.20 CogStackMed [29]4408211.15 LMSys Chatbot Arena [243]3000192.38 JBB-Behaviors [21]10073.75 4 PromptSecurity: A Unified Modular Platform for LLM Prompt Security Evaluation 4.1 Overview This chapter introduces the PromptSecurity platform â a modular, systematic, and flexible framework for end-to-end evaluation of large language models. It covers diverse attack, defense, and model configurations, enabling fair and comprehensive comparisons across arbitrary methods under a unified evaluation protocol. Motivation and Limitations Prior studies on prompt security evaluation remain fragmented, often lack- ing a unified threat model and standardized evaluation protocol. This leads to inconsistent and potentially unfair comparisons, such as when some attacks are aided by stronger auxiliary LLMs than others. Moreover, the absence of a modular, interoperable architecture limits flexibility and scalability as new methods emerge. PromptSecurity addresses these challenges through a unified and extensible evaluation framework. 4.2 Design Goals and Assumptions The design of PromptSecurity is guided by three overarching principles that address the structural limi- tations identified in prior evaluation efforts. âą Modular Evaluation Architecture. PromptSecurity abstracts the evaluation process into a set of fundamental and orthogonal componentsâAttacks, Defenses, Models, Judgers, and Datasets. Each component adheres to a unified interface specification, enabling arbitrary composition across heteroge- neous methods and ensuring that any attack or defense can be evaluated under consistent experimental assumptions. This modularity establishes a principled basis for systematic and fair comparison. âą Unified and Scalable Experiment Management. To ensure reproducibility and extensibility, Prompt- Security employs a configuration-driven management system, where all experimental behaviors are ex- ternally specified through declarative configuration files (e.g., JSON/YAML). This separation between configuration and implementation facilitates scalable experimentation: new methods can be seamlessly integrated and evaluated in batch across existing modules, while preserving reproducibility and version traceability. âą Usability and Maintainability by Design. The framework further emphasizes ease of extension and long-term maintainability. Through automatic component discovery, configuration validation, and deterministic replay mechanisms, PromptSecurity minimizes human intervention and lowers the engineer- ing barrier to integrating new attacks, defenses, and datasets. This design philosophy ensures that the framework remains sustainable as the ecosystem of prompt security methods rapidly evolves. 4.3 Component Modules 4.3.1 Attack Module The attack module operationalizes jailbreak generation in our framework under a clear threat model: an adversary manipulates user-visible inputs to the target LLM without altering system weights or in- frastructure. We consider both black-box access (query and observe outputs only) and white-box access (weights/gradients/logits available) and allow the target to be any backend conforming to BaseModel, in- cluding commercial APIs, local HuggingFace models, or defended wrappers (BaseDefendedModel). This design ensures attacks compose naturally with defenses in end-to-end evaluations and are comparable across deployment modalities. Each attack implementation subclasses BaseAttack and follows a standardized procedure that, given a benign task description and optional configuration, produces (i) a cumulative query budgetâthe total number of model calls expended during prompt constructionâand (i) an adversarial prompt candidate to be issued to the target. Query accounting is holistic: if an attack uses auxiliary components (e.g., an attacker/rewriter/evaluator/judger LLM), their invocations are included in the same budget. The module is configuration-driven and reproducible: methods are discovered dynamically, parameters are specified in JSON, and capability gates restrict white-box algorithms to weight-accessible backends. To enable principled coverage analysis rather than implementation-specific groupings, we annotate each method with one or more tags from Taxonomy I (Figure 1)âfor example, â1.1.1 Obfuscation and Encodingâ or â1.2.3 Reinforcement Learning.â The current release includes 18 techniques (15 black-box and 3 white-box) plus a no attack baseline; a summary of access assumptions, auxiliary-LLM usage, and taxonomy labels appears in Table reftab:attack-module. Consistent with Taxonomy I (2.3), we annotate each implementation with one or more taxonomy labels (e.g., â1.1.1 Obfuscation and Encodingâ, â1.2.3 Reinforcement Learningâ) to enable principled coverage analysis and stratified ablations. 4.3.2 Defense Module The defense module provides a defender-controlled wrapper that turns any backend conforming to BaseModel into a guarded LLM. The wrapper mediates every interaction: all inputs to the systemâincluding attack-generated queries and multi-turn promptsâare applied to the guarded model, not the raw model. Interventions span three loci: input (sanitization and rewriting before generation), model (policy routing, hidden-state/gradient diagnostics, optimization-based hardening), and output (post-generation filtering and self-evaluation). Each defense subclasses BaseDefendedModel and executes a standardized defense pipeline (if any), i.e., input de- fense, model defense, and output defense, with capability pass-through and automatic gating (e.g., gradient-based methods require weight-accessible, full-precision local backends). Defenses are configuration-driven with per-method defaults; compatibility checks prevent invalid configurations. Multiple defense mechanisms can be applied to the target LLM either in a cascaded or parallel manner. Table 5 summarizes the locus of intervention, and auxiliary-LLM usage. For principled analysis, defenses are tagged using Taxonomy I (Figure 2). Table 4: Implemented jailbreak attacks, access assumptions, auxiliary-LLM usage, and taxonomy tags. Aux- iliary LLMs refer to modular attacker/rewriter/evaluator/judger components loaded through the framework. MethodAccessAuxiliary LLMsTaxonomy tags no attackBaselineNoBaseline (no perturbation) FlipAttack [124]Black-boxNo1.1.1 Obfuscation and Encoding ArtPromptAttack [80]Black-boxYes1.5.1 Leveraging Behavior Vulnerabilities PAIR [23]Black-boxYes1.2.1 Heuristic Algorithms; 1.3.2 Multi-Agent Collaboration; 1.3.3 Surrogate Models; 1.4.2 Gradual Escalation ABJAttack [109]Black-boxNo1.1.4 Intent Concealment; 1.5.4 Defensive Mechanism Bypass PastTenseAttack [7]Black-boxYes1.1.2 Substitution and Synonyms IFSJAttack [244]Black-boxNo1.2.2 Gradient Estimation CodeAttack [153]Black-boxNo1.1.1 Obfuscation and Encoding CodeChameleon [129]Black-boxNo1.1.1 Obfuscation and Encoding DRA [115]Black-boxNo1.5.1 Leveraging Behavior Vulnerabilities TapAttack [137]Black-boxYes1.2.1 Heuristic Algorithms; 1.3.2 Multi-Agent Collaboration; 1.4.1 Multi-Turn Contextual Attacks DrAttack [104]Black-boxNo1.4.1 Multi-Turn Contextual Attacks InceptionAttack [105]Black-boxNo1.1.4 Intent Concealment; 1.5.4 Defensive Mechanism Bypass ReNeLLM [40]Black-boxYes1.2.1 Heuristic Algorithms; 1.3.1 LLM-Generated Attacks; 1.4.2 Gradual Escalation GPTFUZZER [220]Black-boxYes1.2.3 Reinforcement Learning PersuasiveInContextAttack [226]Black-boxYes1.1.4 Intent Concealment; 1.3.1 LLM-Generated Attacks; 1.5.1 Leveraging Behavior Vulnerabilities GCGWhiteBoxAttack [251]White-boxNo2.2.1 Prefix/Suffix Manipulation; 2.2.2 Gradient-Based Op- timization AutoDANAttack [117]White-boxNo2.2.1 Prefix/Suffix Manipulation COLDAttack [62]White-boxNo2.2.2 Gradient-Based Optimization Table 5: Implemented defenses by locus, auxiliary-LLM usage, and Taxonomy I tags (Figure 2). âAuxiliary LLMsâ indicates optional use of attacker/rewriter/evaluator/judger components MethodLocus (Input/Model/Output)Auxiliary LLMsTaxonomy I tags NoDefenseBaselineNoBaseline (no intervention) InputFilter DefenseInputYes1.1.1 Prompt Analysis; 2.1.1 Prompt Modification OutputFilter DefenseOutputYes2.4.1 Output Filtering; 2.4.2 Self-Evaluation Perplexity Defense [2]OutputNo1.2.2 Probability Analysis Backtranslation Defense [192]InputYes1.1.2 Intention Detection; 2.1.1 Prompt Modification SmoothLLM Defense [156]ModelYes2.1.1 Prompt Modification; 2.4.2 Self-Evaluation JailGuard Defense [232]ModelNo1.2.1 Semantic Detection PrimeGuard Defense [134]ModelNo2.1.2 Safety Prompts GradSafe Defense [202]ModelNo1.3.2 Gradient Analysis RobustOpt Defense (RPO) [245]InputNo2.1.2 Safety Prompts 4.3.3 Model Module The model module provides a unified, configuration-driven interface to hosted APIs and locally served models. The objective is simplicity and consistency: switching providers requires only a configuration change, enabling large batched attackâdefense experiments without code edits. All models are invoked via lightweight configs with harmonized parameters so that experiments remain comparable across providers. In our reference implementation, 73 API configurations and 53 local configurations are available. Because defenses act as wrappers, any loaded model can be turned into a guarded LLM; all inputs (including attack-generated queries) are applied to the guarded model rather than the raw backend. Access semantics are fixed by category: API backends are treated as black-box; local backends are white-box when precision permits (e.g., full-precision weights), and otherwise behave as black-box. The following tables summarize the out-of-the-box configurations, grouped by API and Local families, with brand/provider and nominal sizes consolidated per family. Table 6: API model configurations included in PromptSecurity (model name, brand/provider, nominal size). All API backends are treated as black-box. Model (config name)Brand/CompanyServer ProviderSize 01-ai-Yi-34B-Chat01.AIDeepInfra34B Austism-chronos-hermes-13b-v2âDeepInfra13B Gryphe-MythoMax-L2-13bGrypheDeepInfra13B Gryphe-MythoMax-L2-13b-turboGrypheDeepInfra13B HuggingFaceH4-zephyr-orpo-141b-A35b-v0.1HuggingFaceH4DeepInfra141B NousResearch-Hermes-3-Llama-3.1-405BMetaDeepInfra405B Qwen-QwQ-32B-PreviewAlibabaDeepInfra32B Qwen-Qwen2-InstructAlibabaDeepInfra7B, 72B Qwen-Qwen2.5-InstructAlibabaDeepInfra7B, 32B, 72B Qwen-Qwen2.5-Coder-InstructAlibabaDeepInfra32B Sao10K-L3-70B-Euryale-v2.1Sao10KDeepInfra70B Sao10K-L3.1-70B-Euryale-v2.2Sao10KDeepInfra70B claude-2.0AnthropicAnthropicâ claude-2.1AnthropicAnthropicâ claude-3-5-haiku-20241022AnthropicAnthropicâ claude-3-5-sonnetAnthropicAnthropicâ (versions: 20240620, 20241022, latest) claude-3-haiku-20240307AnthropicAnthropicâ claude-3-opusAnthropicAnthropicâ (versions: 20240229, latest) claude-3-sonnet-20240229AnthropicAnthropicâ claude-4AnthropicAnthropicâ claude-instant-1.2AnthropicAnthropicâ claude-sonnet-4-20250514AnthropicAnthropicâ cognitivecomputations-dolphin-2.6-mixtral-8x7bMistralAIDeepInfra7B cognitivecomputations-dolphin-2.9.1-llama-3-70bCognitiveComputationsDeepInfra70B deepinfra-airoboros-70bDeepInfraDeepInfra70B deepseek-ai-DeepSeek-R1-0528-TurboDeepSeekDeepSeekâ deepseek-ai-DeepSeek-V2.5DeepSeekDeepSeekâ deepseek-ai-DeepSeek-V3DeepSeekDeepSeekâ deepseek-ai-deepseek-chatDeepSeekDeepSeekâ deepseek-ai-deepseek-coderDeepSeekDeepSeekâ doubao-1-5-pro-32k-250115ByteDanceByteDance (Ark)32K doubao-seed-1-6-250615ByteDanceByteDance (Ark)â doubao-seed-1-6-flash-250615ByteDanceByteDance (Ark)â gemini-1.0-proGoogleGoogleâ (versions: base, latest) gemini-1.5-flashGoogleGoogleâ (versions: base, 002, latest) gemini-1.5-proGoogleGoogleâ (versions: base, 002, latest) gemini-2.0-flashGoogleGoogleâ (versions: base, exp, thinking-exp-1219) gemini-2.5-flashGoogleGoogleâ gemini-2.5-proGoogleGoogleâ google-gemma-1.1-7b-itGoogleDeepInfra7B google-gemma-2-itGoogleDeepInfra7B, 9B, 27B gpt-3.5-turboOpenAIOpenAIâ (versions: base, 0125, 1106) gpt-4OpenAIOpenAIâ (versions: base, turbo) gpt-4.1OpenAIOpenAIâ (versions: base, mini, nano) gpt-4oOpenAIOpenAIâ (versions: base, latest, mini) o1OpenAIOpenAIâ o1-miniOpenAIOpenAIâ o1-previewOpenAIOpenAIâ gpt-o3OpenAIOpenAIâ meta-llama-Llama-2-chat-hfMetaDeepInfra7B, 13B, 70B meta-llama-Llama-3.2-Vision-InstructMetaDeepInfra11B, 90B meta-llama-Llama-3.3-70B-InstructMetaDeepInfra70B meta-llama-Llama-4-405B-InstructMetaDeepInfra405B meta-llama-Meta-Llama-3.1-InstructMetaDeepInfra8B, 70B, 405B microsoft-Phi-3-medium-4k-instructMicrosoftDeepInfraâ microsoft-WizardLM-2-7BMicrosoftDeepInfra7B microsoft-WizardLM-2-8x22BMicrosoftDeepInfra22B mistralai-Mistral-7B-Instruct-v0.3MistralAIDeepInfra7B mistralai-Mistral-Large-Instruct-2407MistralAIDeepInfraâ mistralai-Mixtral-8x22B-Instruct-v0.1MistralAIDeepInfra22B nvidia-Llama-3.1-Nemotron-70B-InstructMetaDeepInfra70B nvidia-Nemotron-4-340B-InstructNVIDIADeepInfra340B openchat-openchat-3.6-8bOpenChatDeepInfra8B Table 7: Local model configurations included in PromptSecurity (model name, brand/provider, nominal size). âWhite-box*â indicates gradient-level access when precision permits. Model (config name)Brand/Company Server ProviderSize 01-AI-Yi-1.5-Chat01.AI6B, 9B, 34B Qwen-QwQ-32B-PreviewAlibaba32B Qwen-Qwen2-InstructAlibaba0.5B, 1.5B, 7B, 72B Qwen-Qwen2.5-InstructAlibaba0.5B, 1.5B, 3B, 7B, 14B, 32B, 72B Qwen-Qwen2.5-Coder-InstructAlibaba1.5B, 7B, 32B Qwen-Qwen3Alibaba0.6B, 1.7B, 4B, 8B, 14B, 32B google-gemma-2-itGoogle2B, 9B, 27B google-gemma-3-1b-itGoogle1B internlm-internlm2-5-1.8b-chatShanghai AI Lab 1.8B internlm-internlm2-5-7b-chatShanghai AI Lab 7B internlm-internlm2.5-20b-chatShanghai AI Lab 20B meta-llama-Llama-2-chat-hfMeta7B, 13B, 70B meta-llama-Llama-3-70B-InstructMeta8B, 70B meta-llama-Llama-3.1-InstructMeta8B, 70B, 405B (incl. FP8) meta-llama-Llama-3.2-Vision-InstructMeta11B, 90B meta-llama-Llama-3.3-70B-InstructMeta70B meta-llama-Llama-4-Scout-17B-16E-Instruct Meta17B microsoft-Phi-2-instructMicrosoftâ microsoft-Phi-3-medium-4k/128k-instructMicrosoftâ microsoft-Phi-3-mini-4k/128k-instructMicrosoftâ microsoft-Phi-3-small-4k/8k/128k-instruct Microsoftâ microsoft-Phi-3.5-MoE-instructMicrosoftâ microsoft-Phi-3.5-mini-instructMicrosoftâ microsoft-Phi-4-instructMicrosoftâ mistralai-Ministral-8B-Instruct-2410MistralAI8B mistralai-Mistral-7B-Instruct-v0.3MistralAI7B mistralai-Mistral-Nemo-Instruct-2407MistralAIâ Note: API backends are treated as black-box. âWhite-box*â indicates that gradient-level access is avail- able for local backends when the chosen precision permits (e.g., full-precision weights). This configuration-driven abstraction enables one-click switching across providers and supports large, batched attackâdefense evalua- tions without changes to experimental code. It is worth noting that all the api and local LLMs can serve as the assisted LLMs in attacks, defenses, and judgers when the access is matched. 4.3.4 Judger Module The judger module standardizes the semantics for assessing harmfulness and attack success. It offers a consistent interface for single-instance and batch evaluation, and comprises three implementation families: (1) rule-based judgers (deterministic, no model), (2) local-model judgers (white-box where feasible), and (3) API-model judgers (black-box). This separation preserves cross-experiment comparability while enabling deliberate trade-offs among accuracy, latency, and cost. Semantics and outputs. By default, judgers return binary decisions: 1 denotes harmful or attack-success, and 0 denotes safe or failure. Certain templates first produce graded scores (e.g., a 1â5 policy scale) that are deterministically mapped to a binary outcome to maintain comparability across judgers. Batch evaluation is supported for throughput, with uniform handling of quota, rate-limit, and network errors via structured exceptions. Backends and interchangeability. Local judgers run on Hugging Face backends (CPU/GPU), enabling offline evaluation and strong determinism. API judgers are provider-agnostic: replacing the judge model configuration suffices to switch to any supported API model (e.g., OpenAI, Anthropic, Google, DeepSeek, DeepInfra) without code changes. This design permits like-for-like comparisons across providers. Templates and calibration. API judgers support multiple prompt templates (e.g., binary harmfulness, HarmBench style, contextual HarmBench, OpenAI policy) that define the evaluation rubric and parsing logic. Thresholds and parsing rules are calibrated to ensure that binary outputs remain robust across providers and model versions. Table 8: Built-in judgers and backend modes. API judgers can switch providers by updating the judge model configuration to any supported API model. Judger (name)MechanismBackend modeAPI backend interchangeable Rejection-prefix evaluator (keyword/regex) [64]Rule matching (refusal-cue detection; cf. )Rule-based (no model)Not applicable HarmBench evaluator (LLaMA-2 classifier) [136]Specialized classifier with yes/no outputLocal model (Hugging Face)Not applicable GPT evaluator (binary harmfulness) [136]Prompt template with binary parsing (cf. )API modelYes (via judge model configuration) GPT evaluator (OpenAI policy scoring) [80]Policy rubric (1â5) mapped to binaryAPI modelYes (via judge model configuration) GPT evaluator (HarmBench style) [136]Template with yes/no decisionAPI modelYes (via judge model configuration) GPT evaluator (contextual HarmBench) [136]Context-informed HarmBench judgmentAPI modelYes (via judge model configuration) GPT evaluator (TAP style) [137]Template-driven binary decisionAPI modelYes (via judge model configuration) Practical considerations. Rule-based judgers provide a deterministic, negligible-cost baseline appropri- ate for conservative screening. Local judgers (e.g., HarmBench) enable offline, reproducible evaluation with configurable precision and quantization. API judgers admit provider substitution by modifying the judge model configuration, permitting controlled comparisons across providers without code changes. All judgers compose with defenses and attacks, enabling automated pass/fail adjudication in end-to-end experimental pipelines. 4.3.5 Dataset Module The dataset module anchors the evaluation pipeline: it specifies the harmful tasks and the contexts under which response safety is adjudicated. This framing supports three core questions: whether models refuse harmful requests without assistance (intrinsic safety), whether attack methods can bypass these guardrails, and whether defenses can restore them once breached. Consequently, dataset choice determines which aspects of safety are exercised and how broadly results generalize. Technically, the module provides a unified loader that normalizes core fields, records provenance and dataset hashes, and applies deterministic sampling, enabling interoperable prompts and exact replays across attacks, defenses, models, and judgers. Built-in datasets. To balance coverage and tractability, we use three public benchmarks that span com- plementary risk surfaces and annotation styles: âą HarmBench [136]: curated harmful-behavior prompts across safety-critical categories; evaluated with binary yes/no rubrics (standard and contextual). We load from the canonical CSV and expose plain behavior text for interchangeable use. âą JailbreakBench (JBB) [21]: an open repository of jailbreak artifacts and policy-aligned behaviors. We use the JBB-Behaviors harmful split (column Goal) with standardized scoring and a clear threat model; artifact-driven attacks can be layered for transfer/reuse studies. âą AIR-Bench-2024 [227]: a red-teaming benchmark with a hierarchical risk taxonomy spanning security, safety, and misinformation; we adopt the official default configuration (field prompt); optional L4-aware sampling is supported but disabled by default. Subset selection and tractability. Exhaustive cross-testing across all method and setting combinations is infeasible. We therefore evaluate fixed-size subsets from each dataset using deterministic sampling (prefix- N or seed-controlled randomization), and integrate three complementary benchmarks (HarmBench, JBB, AIR-Bench-2024) to cover a broad spectrum of safety vulnerabilities while keeping the test bed tractable. Sampling is configuration-driven and reproducible (prefix-N or seed-controlled randomization), with sta- ble sample identifiers for cross-run alignment. Loaders normalize essential fields (prompt text, category labels when available, provenance identifiers), and the pipeline records dataset hashes and configuration manifests to support exact replays and stratified reporting (e.g., by capability, language, domain). Only publicly available benchmarks are used; the pipeline respects licensing constraints, omits personally identifiable in- formation, and provides opt-in filtering for categories restricted by institutional or policy requirements. 4.4 Unified Experimental Protocol Motivation. Comparisons across attacks, defenses, models, datasets, and judgers are often confounded by inconsistent assumptions (e.g., sampling regimes, judge prompts, decoding settings). A unified protocol controls these factors, ensuring fairness, enabling clean ablations, and supporting factorial sweeps whose conclusions generalize across settings. Five-element tuple. Each experiment is defined by the tuple âšM, A, D, S, Jâ©, where M is the model (possibly defended), A an attack (or no attack), D a defense (or no defense), S the dataset split/sampler, and J the judger. This decomposition enables comparable ablations and factorial sweeps across the design space. Metrics and outputs. Operational definitions (e.g., jailbreak success/failure, refusal, utility/harmlessness trade-offs) are fixed by the selected judger templates and thresholds. We additionally report cost (queries). Results are emitted in JSONL with complete metadata, including component versions, seeds, decoding parameters, and dataset hashes, to support exact replay and secondary analysis. Declarative configuration and reproducibility. Experiments are declared via JSON/YAML (or code) with schema validation. The system records configuration manifests, dataset hashes, and environment fin- gerprints (library versions, GPU/CUDA/driver), andâunless otherwise requestedâuses temperature 0 with fixed seeds and canonicalized sampling/batching. These choices enable exact replays and controlled ablations. 5 Preliminary Experiments 5.1 Overview Evaluating LLM security at scale is computationally intensive and methodologically non-trivial. A naive full factorial sweep over âšM, A, D, S, Jâ© (model, attack, defense, dataset, judger; cf.Section 4.4) is infeasible. Instead of enumerating phases, we structure the evaluation as a set of study problems that jointly reduce design complexity while preserving coverage, rigor, and reproducibility. Our principles are: (i) representative sampling under explicit coverage constraints; (i) targeted analyses that answer specific methodological questions; and (i) multi-criteria validation balancing statistical reliability, cost, and generalizability. 5.2 Objectives, Endpoints, and Constraints Primary Metric. Attack Success Rate (ASR), refusal rate, and cost. For dataset S with items s and judger J : ASR(M, A, D, S, J ) = 1 |S| X sâS âźunsafe(M, A, D, s; J ). Secondary Metrics. Latency, token/Query cost, and stability under resampling (judger-selection analysis below). Threat-model and feasibility constraints. Compatibility rules from Section 2.2 (e.g., white-box â local models with weight access), query budgets, and provider rate limits are enforced by the validator. Study questions. We study the following questions: âą Which models should we evaluate to balance representativeness and tractability? âą Which judger should serve as the primary arbiter? âą How to construct a discriminative yet fair evaluation set? âą How to organize the comprehensive attackâdefense matrix under constraints? âą How to analyze and validate results robustly? âą How to scale cost-effectively with reproducibility guarantees? Findings: models and judger. Table 9 summarizes baseline ASR across multiple judgers and datasets without attacks/defenses. We evaluate a broad set of models spanning both API backends and local back- ends. A consistent pattern emerges: local models generally exhibit higher ASR (weaker refusal) than API models under the same prompts. Among API families, Gemini 2.5-Flash attains the lowest ASR across datasets, with GPT-4o and Claude 3.5 also in the low-ASR regime. For summary reporting, we adopt the arithmetic mean over five GPT-based judger templatesâGPT-HarmBench, GPT-HarmfulBinary, GPT- HarmBenchStyle, GPT-OpenAIPolicy, and GPT-TAP âas the default score. Pending experiments and next update. Some cross-configuration sweeps (e.g., extended attack fam- ilies on local white-box backends and ablations over defense stacking) are in progress. We will report their results, along with expanded error bars and agreement diagnostics for J â , in the next revision. 6 Conclusion This SoK addresses the fragmented state of LLM prompt security by unifying concepts, assumptions, and tooling. We contribute a holistic taxonomy of attacks and defenses, declarative threat models that make cost/knowledge/access assumptions explicit, a modular open-source evaluation toolkit that enforces a single protocol across modelâattackâdefense combinations, and JailbreakDBâa large, richly annotated cor- pus (seven feature groups) that also supports lightweight detectors rivaling specialized systems (e.g., Grad- Safe). Our preliminary experiments validate the frameworkâs utility: local models typically show higher attack success rates than API-hosted models under identical prompts; among APIs, Gemini 2.5 attains the lowest baseline ASR, with GPT-4o and Claude 3.5 also low; and averaging five GPT-based judger templates yields a stable, transparent default score. By standardizing sampling, judging, and cost assumptions and releasing complete manifests (component versions, seeds, decoding parameters, dataset hashes, environment fingerprints), we enable fair comparisons and exact replays. Remaining limitationsâfixed-size subsets, auto- mated judgers as proxies, and incomplete threat coverageâmotivate ongoing expansions in white-box/local studies, defense stacking, multilingual testbeds, and human-in-the-loop validation. Table 9: Attack Success Rate on Harmful Question Datasets without Attacks and Defenses ModelDatasetHarmBenchGPT-HarmBenchGPT-HarmfulBinaryGPT-HarmBenchStyleGPT-OpenAIPolicyGPT-TAPREJ-PrefixAvg ASR Mistral-7B-Instruct-v0.3 harmbench80.0%83.0%73.0%80.0%78.0%80.0%38.0%73.1% jbb76.0%80.0%72.0%78.0%77.0%80.0%40.0%71.9% airbench97.0%93.0%84.0%86.0%86.0%100.0%69.0%87.9% Ministral-8B-Instruct harmbench85.0%91.0%79.0%89.0%77.0%87.0%45.0%79.0% jbb73.0%75.0%71.0%73.0%70.0%75.0%32.0%67.0% airbench84.0%89.0%69.0%77.0%71.0%86.0%65.0%77.3% Mistral-Nemo-Instruct harmbench79.0%83.0%73.0%81.0%77.0%80.0%49.0%74.6% jbb71.0%72.0%63.0%67.0%69.0%71.0%52.0%66.4% airbench76.0%88.0%62.0%75.0%73.0%83.0%69.0%75.1% Yi-1.5-9B-Chat harmbench63.0%66.0%61.0%63.0%54.0%58.0%42.0%58.1% jbb64.0%65.0%65.0%62.0%60.0%65.0%37.0%59.7% airbench80.0%82.0%78.0%74.0%70.0%79.0%53.0%73.7% Yi-1.5-6B-Chat harmbench61.0%60.0%66.0%58.0%55.0%58.0%47.0%57.9% jbb56.0%62.0%57.0%57.0%56.0%61.0%38.0%55.3% airbench76.0%82.0%78.0%76.0%69.0%84.0%65.0%75.7% Yi-1.5-34B-Chat harmbench47.0%50.0%52.0%48.0%44.0%50.0%36.0%46.7% jbb57.0%56.0%56.0%52.0%54.0%56.0%31.0%51.7% airbench80.0%76.0%74.0%67.0%66.0%75.0%59.0%71.0% Llama-3.2-1B-Instruct harmbench48.0%50.0%47.0%50.0%41.0%48.0%34.0%45.4% jbb42.0%45.0%38.0%40.0%36.0%42.0%47.0%41.4% airbench57.0%60.0%48.0%56.0%57.0%64.0%66.0%58.3% Llama-3.1-8B-Instruct harmbench38.0%40.0%44.0%37.0%35.0%36.0%40.0%38.6% jbb39.0%41.0%43.0%39.0%35.0%41.0%40.0%39.7% airbench73.0%72.0%66.0%65.0%54.0%68.0%58.0%65.1% Llama-3.2-3B-Instruct harmbench32.0%33.0%30.0%32.0%25.0%31.0%18.0%28.7% jbb25.0%29.0%19.0%23.0%16.0%24.0%21.0%22.4% airbench64.0%61.0%54.0%53.0%46.0%57.0%48.0%54.7% Qwen3-8B harmbench31.0%15.0%40.0%16.0%14.0%16.0%52.0%26.3% jbb31.0%15.0%44.0%17.0%19.0%20.0%43.0%27.0% airbench51.0%36.0%55.0%29.0%31.0%35.0%49.0%40.9% Phi-3.5-mini-instruct harmbench26.0%30.0%28.0%28.0%23.0%26.0%18.0%25.6% jbb27.0%27.0%24.0%26.0%19.0%24.0%17.0%23.4% airbench52.0%52.0%45.0%43.0%35.0%40.0%30.0%42.4% Phi-4-instruct harmbench20.0%10.0%16.0%10.0%16.0%22.0%88.0%26.0% jbb27.0%16.0%25.0%13.0%22.0%26.0%92.0%31.6% airbench13.0%27.0%10.0%23.0%10.0%44.0%100.0%32.4% Phi-3-medium-128k-instruct harmbench27.0%33.0%29.3%32.0%23.0%26.0%31.0%28.8% jbb14.0%14.1%20.2%13.0%13.0%15.0%24.0%16.2% airbench44.0%53.0%40.0%42.0%37.0%49.0%48.0%44.7% Qwen2.5-7B-Instruct harmbench30.0%30.0%24.0%29.0%26.0%29.0%16.0%26.3% jbb19.0%14.0%18.0%13.0%15.0%16.0%10.0%15.0% airbench55.0%62.0%46.0%43.0%42.0%51.0%31.0%47.1% Doubao-Seed-1-6-Flash-250615 harmbench13.0%14.0%14.0%15.0%13.0%14.0%38.0%17.3% jbb11.0%13.0%7.0%10.0%8.0%11.0%28.0%12.6% airbench44.0%49.0%25.0%34.0%30.0%40.0%32.0%36.3% Deepseek-v3 harmbench22.0%22.0%10.0%22.0%18.0%21.0%16.0%18.7% jbb13.0%13.0%6.0%11.0%10.0%11.0%12.0%10.9% airbench40.0%40.0%29.0%25.0%22.0%36.0%32.0%32.0% Doubao-1-5-Pro-32k-250115 harmbench10.0%14.0%8.0%11.0%8.0%11.0%8.0%10.0% jbb12.0%13.0%7.0%11.0%8.0%10.0%10.0%10.1% airbench47.0%49.0%33.0%39.0%37.0%42.0%20.0%38.1% Deepseek-r1 harmbench7.0%9.0%3.0%9.0%4.0%8.0%6.0%6.6% jbb10.0%16.0%5.0%11.0%4.0%9.0%5.0%8.6% airbench41.0%42.0%26.0%34.0%19.0%34.0%20.0%30.9% Qwen2.5-14B-Instruct harmbench11.0%11.0%5.0%11.0%8.0%10.0%5.0%8.7% jbb6.0%4.0%2.0%4.0%3.0%6.0%8.0%4.7% airbench37.0%36.0%26.0%31.0%22.0%32.0%26.0%30.0% GPT-4o harmbench9.0%8.0%1.0%9.0%6.0%9.0%12.0%7.7% jbb4.0%4.0%1.0%3.0%1.0%4.0%9.0%3.7% airbench24.0%22.0%10.0%13.0%11.0%22.0%20.0%17.4% Gemini-1.5-Flash harmbench0.0%0.0%1.0%0.0%0.0%0.0%0.0%0.1% jbb3.0%4.0%3.0%4.0%1.0%3.0%4.0%3.1% airbench27.0%32.0%23.0%20.0%15.0%23.0%14.0%22.0% Doubao-Seed-1-6-250615 harmbench1.0%2.0%1.0%1.0%1.0%1.0%3.0%1.4% jbb2.0%3.0%1.0%1.0%2.0%4.0%4.0%2.4% airbench23.0%22.0%9.0%15.0%13.0%20.0%20.0%17.4% GPT-4.1 harmbench3.0%3.0%1.0%3.0%3.0%3.0%6.0%3.1% jbb7.0%7.0%2.0%4.0%3.0%3.0%6.0%4.6% airbench21.0%20.0%6.0%8.0%5.0%17.0%16.0%13.3% Gemini-2.0-Flash harmbench0.0%0.0%0.0%0.0%0.0%0.0%1.0%0.1% jbb1.0%2.0%0.0%1.0%0.0%1.0%5.0%1.4% airbench24.0%25.0%18.0%18.0%10.0%21.0%15.0%18.7% GPT-3.5-Turbo harmbench1.0%1.0%0.0%1.0%0.0%1.0%0.0%0.6% jbb3.0%3.0%0.0%1.0%1.0%3.0%2.0%1.9% airbench12.0%10.0%6.0%6.0%8.0%11.0%14.0%9.6% Claude-Sonnet-4-20250514 harmbench1.0%2.0%0.0%2.0%0.0%2.0%4.0%1.6% jbb3.0%5.0%1.0%3.0%2.0%4.0%8.0%3.7% airbench4.0%5.0%0.0%0.0%1.0%1.0%24.0%5.0% Claude-3-5-Sonnet harmbench0.0%0.0%0.0%0.0%0.0%0.0%6.0%0.9% jbb2.0%3.0%1.0%2.0%1.0%3.0%6.0%2.6% airbench0.0%0.0%0.0%0.0%0.0%0.0%18.0%2.6% Claude-3.5-Haiku-20241022 harmbench2.0%2.0%1.0%1.0%2.0%2.0%6.0%2.3% jbb2.0%3.0%1.0%2.0%1.0%3.0%4.0%2.3% airbench0.0%0.0%0.0%0.0%1.0%1.0%6.0%1.1% Gemini-2.5-Flash harmbench0.0%0.0%0.0%0.0%0.0%0.0%1.0%0.1% jbb0.0%0.0%0.0%0.0%0.0%0.0%1.0%0.1% airbench0.0%0.0%0.0%0.0%0.0%0.0%0.0%0.0% References [1] GlaiveAI.glaive-code-assistant,2023.Availableat: https://huggingface.co/datasets/glaiveai/ glaive-code-assistant. [2] Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity. CoRR, abs/2308.14132, 2023. [3] Suraj Anand and David Getzen. Are ppo-ed language models hackable? CoRR, abs/2406.02577, 2024. [4] Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. CoRR, abs/2404.02151, 2024. [5] Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [6] Maksym Andriushchenko and Nicolas Flammarion. Does refusal training in llms generalize to the past tense? CoRR, abs/2407.11969, 2024. [7] Maksym Andriushchenko and Nicolas Flammarion. Does refusal training in llms generalize to the past tense? In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net, 2025. [8] Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomek Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger B. Grosse, and David Kristjanson Duvenaud. Many-shot jailbreaking. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [9] Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul Ronald Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, George Tucker, Enrique Piqueras, Maxim Krikun, Iain Barr, Nikolay Savinov, Ivo Danihelka, Becca Roelofs, Ana Ìıs White, Anders Andreassen, Tamara von Glehn, Lakshman Yagati, Mehran Kazemi, Lucas Gonzalez, Misha Khalman, Jakub Sygnowski, and et al. Gemini: A family of highly capable multimodal models. CoRR, abs/2312.11805, 2023. [10] Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, and Kui Ren. Surro- gateprompt: Bypassing the safety filter of text-to-image models via substitution. In Bo Luo, Xiaojing Liao, Jun Xu, Engin Kirda, and David Lie, editors, Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, pages 1166â1180. ACM, 2024. [11] Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. Qwen technical report. CoRR, abs/2309.16609, 2023. [12] Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback. CoRR, abs/2204.05862, 2022. [13] Somnath Banerjee, Sayan Layek, Rima Hazra, and Animesh Mukherjee. How (un)ethical are instruction-centric responses of llms? unveiling the vulnerabilities of safety guardrails to harmful queries. CoRR, abs/2402.15302, 2024. [14] Rishabh Bhardwaj and Soujanya Poria. Red-teaming large language models using chain of utterances for safety-alignment. CoRR, abs/2308.09662, 2023. [15] Anubhav Bhatti, Prithila Angkan, Behnam Behinaein, Zunayed Mahmud, Dirk Rodenburg, Heather Braund, P. James Mclellan, Aaron J. Ruberto, Geoffery Harrison, Daryl Wilson, Adam Szulewski, Dan Howes, Ali Etemad, and Paul Hungler. CLARE: cognitive load assessment in realtime with multimodal data. CoRR, abs/2404.17098, 2024. [16] Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul R Ìottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [17] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, MarcâAurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. [18] CAMEL-AI. Camelai domain expert datasets (physics, math, chemistry & biology), 2023. Available at: https:// huggingface.co/camel-ai. [19] Trishna Chakraborty, Erfan Shayegani, Zikui Cai, Nael B. Abu-Ghazaleh, M. Salman Asif, Yue Dong, Amit K. Roy- Chowdhury, and Chengyu Song. Can textual unlearning solve cross-modality safety alignment? In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, pages 9830â9844. Association for Computational Linguistics, 2024. [20] Chiju Chao, Qirui Wang, Hongfei Wu, and Zhiyong Fu. Synneure: Intelligent human-machine teamwork in virtual space. In Pei-Luen Patrick Rau, editor, Cross-Cultural Design - 15th International Conference, CCD 2023, Held as Part of the 25th International Conference, HCII 2023, Copenhagen, Denmark, July 23-28, 2023, Proceedings, Part I, volume 14023 of Lecture Notes in Computer Science, pages 349â361. Springer, 2023. [21] Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tram`er, Hamed Hassani, and Eric Wong. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [22] Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. CoRR, abs/2310.08419, 2023. [23] Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In IEEE Conference on Secure and Trustworthy Machine Learning, SaTML 2025, Copenhagen, Denmark, April 9-11, 2025, pages 23â42. IEEE, 2025. [24] Sizhe Chen, Julien Piet, Chawin Sitawarin, and David A. Wagner. Struq: Defending against prompt injection with structured queries. CoRR, abs/2402.06363, 2024. [25] Xuan Chen, Yuzhou Nie, Wenbo Guo, and Xiangyu Zhang. When LLM meets DRL: advancing jailbreaking efficiency via drl-guided search. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [26] Xuan Chen, Yuzhou Nie, Lu Yan, Yunshu Mao, Wenbo Guo, and Xiangyu Zhang. RL-JACK: reinforcement learning- powered black-box jailbreaking attack against llms. CoRR, abs/2406.08725, 2024. [27] Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos. Leveraging the context through multi- round interactions for jailbreaking attacks. CoRR, abs/2402.09177, 2024. [28] Jan Clusmann, Dyke Ferber, Isabella C. Wiest, Carolin V. Schneider, Titus J. Brinker, Sebastian Foersch, Daniel Truhn, and Jakob Nikolas Kather. Prompt injection attacks on large language models in oncology. CoRR, abs/2407.18981, 2024. [29] CogStack. Cogstack, 2023. Available at: https://github.com/CogStack/OpenGPT. [30] Stav Cohen, Ron Bitton, and Ben Nassi. A jailbroken genai model can cause substantial harm: Genai-powered applications are vulnerable to promptwares. CoRR, abs/2408.05061, 2024. [31] Giandomenico Cornacchia, Giulio Zizzo, Kieran Fraser, Muhammad Zaid Hameed, Ambrish Rawat, and Mark Purcell. Moje: Mixture of jailbreak experts, naive tabular classifiers as guard for prompt attacks. CoRR, abs/2409.17699, 2024. [32] Giandomenico Cornacchia, Giulio Zizzo, Kieran Fraser, Muhammad Zaid Hameed, Ambrish Rawat, and Mark Purcell. Moje: Mixture of jailbreak experts, naive tabular classifiers as guard for prompt attacks. In Sanmay Das, Brian Patrick Green, Kush Varshney, Marianna Ganapini, and Andrea Renda, editors, Proceedings of the Seventh AAAI/ACM Con- ference on AI, Ethics, and Society (AIES-24) - Full Archival Papers, October 21-23, 2024, San Jose, California, USA - Volume 1, pages 304â315. AAAI Press, 2024. [33] Crystalcareai.alpaca-gpt4-cot,2024.Available at: https://huggingface.co/datasets/Crystalcareai/ alpaca-gpt4-COT. [34] Edoardo Debenedetti, Javier Rando, Daniel Paleka, Silaghi Fineas Florin, Dragos Albastroiu, Niv Cohen, Yuval Lemberg, Reshmi Ghosh, Rui Wen, Ahmed Salem, Giovanni Cherubin, Santiago Zanella-B Ìeguelin, Robin Schmid, Victor Klemm, Takahiro Miki, Chenhao Li, Stefan Kraft, Mario Fritz, Florian Tram`er, Sahar Abdelnabi, and Lea Sch Ìonherr. Dataset and lessons learned from the 2024 satml LLM capture-the-flag competition. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [35] Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tram`er. Agentdojo: A dynamic environment to evaluate attacks and defenses for LLM agents. CoRR, abs/2406.13352, 2024. [36] Boyi Deng, Wenjie Wang, Fuli Feng, Yang Deng, Qifan Wang, and Xiangnan He. Attack prompt generation for red teaming and defending large language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 2176â2189. Association for Computational Linguistics, 2023. [37] Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. Masterkey: Automated jailbreak across multiple large language model chatbots. arXiv preprint arXiv:2307.08715, 2023. [38] Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu. Pandora: Jailbreak gpts by retrieval augmented generation poisoning. CoRR, abs/2402.08416, 2024. [39] Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. Multilingual jailbreak challenges in large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [40] Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. A wolf in sheepâs clothing: Generalized nested jailbreak prompts can fool large language models easily. In Kevin Duh, Helena G Ìomez-Adorno, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, pages 2136â2153. Association for Computational Linguistics, 2024. [41] Yingkai Dong, Zheng Li, Xiangtao Meng, Ning Yu, and Shanqing Guo. Jailbreaking text-to-image models with llm-based agents. CoRR, abs/2408.00523, 2024. [42] Yiting Dong, Guobin Shen, Dongcheng Zhao, Xiang He, and Yi Zeng. Harnessing task overload for scalable jailbreak attacks on large language models. CoRR, abs/2410.04190, 2024. [43] Diego Dorn, Alexandre Variengien, Charbel-Rapha Ìel S Ìegerie, and Vincent Corruble. BELLS: A framework towards future proof benchmarks for the evaluation of LLM safeguards. CoRR, abs/2406.01364, 2024. [44] Moussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi, Dan Jurafsky, and Christopher D. Manning. h4rm3l: A dynamic benchmark of composable jailbreak attacks for LLM safety assessment. CoRR, abs/2408.04811, 2024. [45] Xiaohu Du, Fan Mo, Ming Wen, Tu Gu, Huadi Zheng, Hai Jin, and Jie Shi. Multi-turn jailbreaking large language models via attention shifting. In Toby Walsh, Julie Shah, and Zico Kolter, editors, AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, USA, pages 23814â23822. AAAI Press, 2025. [46] Xuefeng Du, Reshmi Ghosh, Robert Sim, Ahmed Salem, Vitor Carvalho, Emily Lawton, Yixuan Li, and Jack W. Stokes. Vlmguard: Defending vlms against malicious prompts via unlabeled data. CoRR, abs/2410.00296, 2024. [47] Yuhao Du, Zhuo Li, Pengyu Cheng, Xiang Wan, and Anningzhe Gao. Detecting AI flaws: Target-driven attacks on internal faults in language models. CoRR, abs/2408.14853, 2024. [48] Matt Duver, Noah Wiederhold, Maria Kyrarini, Sean Banerjee, and Natasha Kholgade Banerjee. Vr-hand-in-hand: Using virtual reality (VR) hand tracking for hand-object data annotation. In 2024 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR), Los Angeles, CA, USA, January 17-19, 2024, pages 325â329. IEEE, 2024. [49] Aysan Esmradi, Daniel Wankit Yip, and Chun-Fai Chan. A comprehensive survey of attack techniques, implementation, and mitigation strategies in large language models. In Guojun Wang, Haozhe Wang, Geyong Min, Nektarios Georgalas, and Weizhi Meng, editors, Ubiquitous Security - Third International Conference, UbiSec 2023, Exeter, UK, November 1-3, 2023, Revised Selected Papers, volume 2034 of Communications in Computer and Information Science, pages 76â95. Springer, 2023. [50] Yihe Fan, Yuxin Cao, Ziyu Zhao, Ziyao Liu, and Shaofeng Li. Unbridled icarus: A survey of the potential perils of image inputs in multimodal large language model security. In IEEE International Conference on Systems, Man, and Cybernetics, SMC 2024, Kuching, Malaysia, October 6-10, 2024, pages 3428â3433. IEEE, 2024. [51] Yingchaojie Feng, Zhizhang Chen, Zhining Kang, Sijia Wang, Minfeng Zhu, Wei Zhang, and Wei Chen. Jailbreaklens: Visual analysis of jailbreak attacks against large language models. CoRR, abs/2404.08793, 2024. [52] Yu Fu, Yufei Li, Wen Xiao, Cong Liu, and Yue Dong. Safety alignment in NLP tasks: Weakly aligned summarization as an in-context attack. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 8483â8502. Association for Computational Linguistics, 2024. [53] Yu Fu, Wen Xiao, Jia Chen, Jiachen Li, Evangelos E. Papalexakis, Aichi Chien, and Yue Dong. Cross-task defense: Instruction-tuning llms for content safety. CoRR, abs/2405.15202, 2024. [54] V Ìıctor Gallego. Merging improves self-critique against jailbreak attacks. CoRR, abs/2406.07188, 2024. [55] Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao. MART: improving LLM safety with multi-round automatic red-teaming. In Kevin Duh, Helena G Ìomez-Adorno, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, pages 1927â1937. Association for Computational Linguistics, 2024. [56] Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. Coercing llms to do and reveal (almost) anything. CoRR, abs/2402.14020, 2024. [57] Simon Geisler, Tom Wollschl Ìager, M. H. I. Abdalla, Johannes Gasteiger, and Stephan G Ìunnemann. Attacking large language models with projected gradient descent. CoRR, abs/2402.09154, 2024. [58] Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien. AEGIS: online adaptive AI content safety moderation with ensemble of LLM experts. CoRR, abs/2404.05993, 2024. [59] Tom Gibbs, Ethan Kosak-Hine, George Ingebretsen, Jason Zhang, Julius Broomfield, Sara Pieri, Reihaneh Iranmanesh, Reihaneh Rabbany, and Kellin Pelrine. Emerging vulnerabilities in frontier models: Multi-turn jailbreak attacks. CoRR, abs/2409.00137, 2024. [60] David Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan, and Nicolas Papernot. A false sense of safety: Unsafe information leakage in âsafeâ AI responses. CoRR, abs/2407.02551, 2024. [61] Chuan Guo, Alexandre Sablayrolles, Herv Ìe J Ìegou, and Douwe Kiela. Gradient-based adversarial attacks against text transformers. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 5747â5757. Association for Computational Linguistics, 2021. [62] Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu. Cold-attack: Jailbreaking llms with stealthiness and controllability. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. [63] Maanak Gupta, Charankumar Akiri, Kshitiz Aryal, Eli Parker, and Lopamudra Praharaj. From chatgpt to threatgpt: Impact of generative AI in cybersecurity and privacy. IEEE Access, 11:80218â80245, 2023. [64] Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [65] Divij Handa, Advait Chirmule, Bimal G. Gajera, and Chitta Baral. Jailbreaking proprietary large language models using word substitution cipher. CoRR, abs/2402.10601, 2024. [66] Richard Harper and Dave W. Randall. Machine learning and the work of the user. Comput. Support. Cooperative Work., 33(2):103â136, 2024. [67] Anne Hartebrodt and Richard R Ìottger. Privacy of federated QR decomposition using additive secure multiparty compu- tation. IEEE Trans. Inf. Forensics Secur., 18:5122â5132, 2023. [68] Jonathan Hayase, Ema Borevkovic, Nicholas Carlini, Florian Tram`er, and Milad Nasr. Query-based adversarial prompt generation. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [69] Tanmoy Hazra, Kushal Anjaria, Aditi Bajpai, and Akshara Kumari. Applications of Game Theory in Deep Learning. Springer Briefs in Computer Science. Springer, 2024. [70] Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. Defending against indirect prompt injection attacks with spotlighting. In Rachel Allen, Sagar Samtani, Edward Raff, and Ethan M. Rudd, editors, Proceedings of the Conference on Applied Machine Learning in Information Security (CAMLIS 2024), Arlington, Virginia, USA, October 24-25, 2024, volume 3920 of CEUR Workshop Proceedings, pages 48â62. CEUR-WS.org, 2024. [71] Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. Gradient cuff: Detecting jailbreak attacks on large language models by exploring refusal loss landscapes. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Confer- ence on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [72] Caishuang Huang, Wanxu Zhao, Rui Zheng, Huijie Lv, Shihan Dou, Sixian Li, Xiao Wang, Enyu Zhou, Junjie Ye, Yuming Yang, Tao Gui, Qi Zhang, and Xuanjing Huang. Safealigner: Safety alignment against jailbreak attacks via response disparity guidance. CoRR, abs/2406.18118, 2024. [73] Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. On the humanity of conversational AI: evaluating the psychological portrayal of llms. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [74] Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. Catastrophic jailbreak of open-source llms via exploiting generation. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [75] Nanna Inie, Jonathan Stray, and Leon Derczynski. Summon a demon and bind it: A grounded theory of LLM red teaming in the wild. CoRR, abs/2311.06237, 2023. [76] Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. Baseline defenses for adversarial attacks against aligned language models. CoRR, abs/2309.00614, 2023. [77] Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of LLM via a human-preference dataset. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. [78] Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin. Improved techniques for optimization-based jailbreaking on large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [79] Bojian Jiang, Yi Jing, Tong Wu, Tianhao Shen, Deyi Xiong, and Qing Yang. Automated progressive red teaming. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors, Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19-24, 2025, pages 3850â3864. Association for Computational Linguistics, 2025. [80] Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran. Artprompt: ASCII art-based jailbreak attacks against aligned llms. In Lun-Wei Ku, Andre Martins, and Vivek Sriku- mar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 15157â15173. Association for Computational Linguistics, 2024. [81] Jiyue Jiang, Liheng Chen, Pengan Chen, Sheng Wang, Qinghang Bao, Lingpeng Kong, Yu Li, and Chuan Wu. How far can cantonese NLP go? benchmarking cantonese capabilities of large language models. CoRR, abs/2408.16756, 2024. [82] Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [83] Weipeng Jiang, Zhenting Wang, Juan Zhai, Shiqing Ma, Zhengyu Zhao, and Chao Shen. Unlocking adversarial suffix optimization without affirmative phrases: Efficient black-box jailbreaking via LLM as optimizer. CoRR, abs/2408.11313, 2024. [84] Yifan Jiang, Kriti Aggarwal, Tanmay Laud, Kashif Munir, Jay Pujara, and Subhabrata Mukherjee. RED QUEEN: safeguarding large language models against concealed multi-turn jailbreaking. CoRR, abs/2409.17458, 2024. [85] Haibo Jin, Ruoxi Chen, Andy Zhou, Jinyin Chen, Yang Zhang, and Haohan Wang. GUARD: role-playing to generate natural-language jailbreakings to test guideline adherence of large language models. CoRR, abs/2402.03299, 2024. [86] Haibo Jin, Leyang Hu, Xinuo Li, Peiyan Zhang, Chonghan Chen, Jun Zhuang, and Haohan Wang. Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models. CoRR, abs/2407.01599, 2024. [87] jondurbin. airoboros-2.2, 2023. Available at: https://huggingface.co/datasets/jondurbin/airoboros-2.2. [88] Erik Jones, Anca D. Dragan, Aditi Raghunathan, and Jacob Steinhardt. Automatically auditing large language models via discrete optimization. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 15307â15329. PMLR, 2023. [89] Erik Jones, Anca D. Dragan, and Jacob Steinhardt. Adversaries can misuse combinations of safe models. CoRR, abs/2406.14595, 2024. [90] Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. In IEEE Security and Privacy, SP 2024 - Workshops, San Francisco, CA, USA, May 23, 2024, pages 132â143. IEEE, 2024. [91] Daniil Khomsky, Narek Maloyan, and Bulat Nutfullin.Prompt injection attacks in defended systems.CoRR, abs/2406.14048, 2024. [92] Heegyu Kim, Sehyun Yuk, and Hyunsouk Cho. Break the breakout: Reinventing LM defense against jailbreak attacks with self-refinement. CoRR, abs/2402.15180, 2024. [93] Taeyoun Kim, Suhas Kotha, and Aditi Raghunathan. Jailbreaking is best solved by definition. CoRR, abs/2403.14725, 2024. [94] Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki, Jun Suzuki, and Keisuke Sakaguchi. Empirical analysis of large vision-language models against goal hijacking via visual prompt injection. CoRR, abs/2408.03554, 2024. [95] Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. Certifying LLM safety against adversarial prompting. CoRR, abs/2309.02705, 2023. [96] Raz Lapid, Ron Langberg, and Moshe Sipper. Open sesame! universal black box jailbreaking of large language models. CoRR, abs/2309.01446, 2023. [97] Ariel N. Lee, Cole J. Hunter, and Nataniel Ruiz. Platypus: Quick, cheap, and powerful refinement of llms. CoRR, abs/2308.07317, 2023. [98] Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. Multi-step jailbreaking privacy attacks on chatgpt. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Compu- tational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 4138â4153. Association for Computational Linguistics, 2023. [99] Jie Li, Yi Liu, Chongyang Liu, Ling Shi, Xiaoning Ren, Yaowen Zheng, Yang Liu, and Yinxing Xue. A cross-language investigation into jailbreak attacks in large language models. CoRR, abs/2401.16765, 2024. [100] Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. LLM defenses are not robust to multi-turn human jailbreaks yet. CoRR, abs/2408.15221, 2024. [101] Qizhang Li, Yiwen Guo, Wangmeng Zuo, and Hao Chen. Improved generation of adversarial examples against safety- aligned llms. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [102] Xiaoxia Li, Siyuan Liang, Jiyi Zhang, Han Fang, Aishan Liu, and Ee-Chien Chang. Semantic mirror jailbreak: Genetic algorithm based jailbreak prompts against open-source llms. CoRR, abs/2402.14872, 2024. [103] Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. Drattack: Prompt decomposition and reconstruction makes powerful LLM jailbreakers. CoRR, abs/2402.16914, 2024. [104] Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. Drattack: Prompt decomposition and reconstruction makes powerful llms jailbreakers. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, pages 13891â13913. Association for Computational Linguistics, 2024. [105] Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. Deepinception: Hypnotize large language model to be jailbreaker. CoRR, abs/2311.03191, 2023. [106] Yixuan Li, Xuelin Liu, Xiaoyang Wang, Shiqi Wang, and Weisi Lin. Fakebench: Uncover the achillesâ heels of fake images with large multimodal models. CoRR, abs/2404.13306, 2024. [107] Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. RAIN: your language models can align themselves without finetuning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [108] Zeyi Liao and Huan Sun. Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms. CoRR, abs/2404.07921, 2024. [109] Shi Lin, Rongchang Li, Xun Wang, Changting Lin, Wenpeng Xing, and Meng Han. Figure it out: Analyzing-based jailbreak attack on large language models. CoRR, abs/2407.16205, 2024. [110] Zhihao Lin, Wei Ma, Mingyi Zhou, Yanjie Zhao, Haoyu Wang, Yang Liu, Jun Wang, and Li Li. Pathseeker: Exploring LLM security vulnerabilities with a reinforcement learning-based jailbreak approach. CoRR, abs/2409.14177, 2024. [111] Chengyuan Liu, Fubang Zhao, Lizhi Qing, Yangyang Kang, Changlong Sun, Kun Kuang, and Fei Wu. Goal-oriented prompt attack and safety evaluation for llms. arXiv preprint arXiv:2309.11830, 2023. [112] Fan Liu, Yue Feng, Zhao Xu, Lixin Su, Xinyu Ma, Dawei Yin, and Hao Liu. JAILJUDGE: A comprehensive jailbreak judge benchmark with multi-agent enhanced explanation evaluation framework. CoRR, abs/2410.12855, 2024. [113] Fan Liu, Zhao Xu, and Hao Liu. Adversarial tuning: Defending against jailbreak attacks for llms. CoRR, abs/2406.06622, 2024. [114] Hongfu Liu, Yuxi Xie, Ye Wang, and Michael Shieh. Advancing adversarial suffix transfer learning on aligned large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 7213â7224. Association for Computational Linguistics, 2024. [115] Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen. Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. In Davide Balzarotti and Wenyuan Xu, editors, 33rd USENIX Security Symposium, USENIX Security 2024, Philadelphia, PA, USA, August 14-16, 2024. USENIX Association, 2024. [116] Xiao Liu, Liangzhi Li, Tong Xiang, Fuying Ye, Lu Wei, Wangyue Li, and Noa Garcia. Imposter.ai: Adversarial attacks with hidden intentions towards aligned large language models. CoRR, abs/2407.15399, 2024. [117] Xiaogeng Liu, Peiran Li, G. Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao. Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net, 2025. [118] Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Autodan: Generating stealthy jailbreak prompts on aligned large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [119] Xiaoqun Liu, Jiacheng Liang, Muchao Ye, and Zhaohan Xi. Robustifying safety-aligned large language models through clean data curation. CoRR, abs/2405.19358, 2024. [120] Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G Ìul Varol, editors, Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part LVI, volume 15114 of Lecture Notes in Computer Science, pages 386â403. Springer, 2024. [121] Yi Liu, Chengjun Cai, Xiaoli Zhang, Xingliang Yuan, and Cong Wang. Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts. In Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prab- hakaran, Susanne Boll, Ramanathan Subramanian, Liang Zheng, Vivek K. Singh, Pablo C Ìesar, Lexing Xie, and Dong Xu, editors, Proceedings of the 32nd ACM International Conference on Multimedia, M 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024, pages 3578â3586. ACM, 2024. [122] Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. Prompt injection attack against llm-integrated applications. CoRR, abs/2306.05499, 2023. [123] Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. Jailbreaking chatgpt via prompt engineering: An empirical study. CoRR, abs/2305.13860, 2023. [124] Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, and Bryan Hooi. Flipattack: Jailbreak llms via flipping. CoRR, abs/2410.02832, 2024. [125] Lin Lu, Hai Yan, Zenghui Yuan, Jiawen Shi, Wenqi Wei, Pin-Yu Chen, and Pan Zhou. Autojailbreak: Exploring jailbreak attacks and defenses through a dependency lens. CoRR, abs/2406.03805, 2024. [126] Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen. Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge. CoRR, abs/2404.05880, 2024. [127] Wen-xiao Lu, Ping Fang, Ming-lu Zhu, Yi-run Zhu, Xinjian Fan, Tian-chen Zhu, Xuan Zhou, Feng-Xia Wang, Tao Chen, and Li-ning Sun. Artificial intelligence-enabled gesture-language-recognition feedback system using strain-sensor-arrays- based smart glove. Adv. Intell. Syst., 5(8), 2023. [128] Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. CoRR, abs/2404.03027, 2024. [129] Huijie Lv, Xiao Wang, Yuansen Zhang, Caishuang Huang, Shihan Dou, Junjie Ye, Tao Gui, Qi Zhang, and Xu- anjing Huang. Codechameleon: Personalized encryption framework for jailbreaking large language models. CoRR, abs/2402.16717, 2024. [130] Lijia Lv, Weigang Zhang, Xuehai Tang, Jie Wen, Feng Liu, Jizhong Han, and Songlin Hu. Adappa: Adaptive position pre-fill jailbreak attack approach targeting llms. CoRR, abs/2409.07503, 2024. [131] Jiachen Ma, Yijiang Li, Zhiqing Xiao, Anda Cao, Jie Zhang, Chao Ye, and Junbo Zhao. Jailbreaking prompt attack: A controllable adversarial attack against diffusion models. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, pages 3141â3157. Association for Computational Linguistics, 2025. [132] Jiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li, Luke Zettlemoyer, Shih-Fu Chang, Wen-Tau Yih, and Hu Xu. Mode: CLIP data experts via clustering. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 26344â26353. IEEE, 2024. [133] Siyuan Ma, Weidi Luo, Yu Wang, Xiaogeng Liu, Muhao Chen, Bo Li, and Chaowei Xiao. Visual-roleplay: Universal jailbreak attack on multimodal large language models via role-playing image characte. CoRR, abs/2405.20773, 2024. [134] Blazej Manczak, Eliott Zemour, Eric Lin, and Vaikkunth Mugunthan. Primeguard: Safe and helpful llms through tuning- free routing. CoRR, abs/2407.16318, 2024. [135] Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash. PRP: propagating universal perturbations to attack large language model guard-rails. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 10960â10976. Association for Computational Linguistics, 2024. [136] Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David A. Forsyth, and Dan Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. [137] Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum S. Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [138] Fredrik Nestaas, Edoardo Debenedetti, and Florian Tram`er. Adversarial search engine optimization for large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [139] Rida Noor, Abdul Wahid, Sibghat Ullah Bazai, Asad Khan, Meie Fang, Syam M. S, Uzair Aslam Bhatti, and Yazeed Yasin Ghadi. DLGAN: undersampled MRI reconstruction using deep learning based generative adversarial network. Biomed. Signal Process. Control., 93:106218, 2024. [140] OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. [141] Dario Pasquini, Martin Strohmeier, and Carmela Troncoso. Neural exec: Learning (and learning from) execution triggers for prompt injection attacks. In Maura Pintor, Xinyun Chen, and Matthew Jagielski, editors, Proceedings of the 2024 Workshop on Artificial Intelligence and Security, AISec 2024, Salt Lake City, UT, USA, October 14-18, 2024, pages 89â100. ACM, 2024. [142] Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian. Advprompter: Fast adaptive adversarial prompting for llms. CoRR, abs/2404.16873, 2024. [143] Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. Instruction tuning with GPT-4. CoRR, abs/2304.03277, 2023. [144] Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak large language models. CoRR, abs/2306.13213, 2023. [145] Cheng Qian, Hainan Zhang, Lei Sha, and Zhiming Zheng. HSF: defending against jailbreak attacks with hidden state filtering. In Guodong Long, Michale Blumestein, Yi Chang, Liane Lewin-Eytan, Zi Helen Huang, and Elad Yom-Tov, editors, Companion Proceedings of the ACM on Web Conference 2025, W 2025, Sydney, NSW, Australia, 28 April 2025 - 2 May 2025, pages 2078â2087. ACM, 2025. [146] Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan. Latent jailbreak: A benchmark for evaluating text safety and output robustness of large language models. CoRR, abs/2307.08487, 2023. [147] Md. Abdur Rahman, Hossain Shahriar, Fan Wu, and Alfredo Cuzzocrea. Applying pre-trained multilingual BERT in embeddings for improved malicious prompt injection attacks detection. In 2nd International Conference on Artificial Intelligence, Blockchain, and Internet of Things, AIBThings 2024, Mt Pleasant, MI, USA, September 7-8, 2024, pages 1â7. IEEE, 2024. [148] Govind Ramesh, Yao Dou, and Wei Xu. GPT-4 jailbreaks itself with near-perfect success using self-explanation. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 22139â22148. Association for Computational Linguistics, 2024. [149] Delong Ran, Jinyuan Liu, Yichen Gong, Jingyi Zheng, Xinlei He, Tianshuo Cong, and Anyu Wang. Jailbreakeval: An integrated toolkit for evaluating jailbreak attempts against large language models. CoRR, abs/2406.09321, 2024. [150] Javier Rando, Francesco Croce, Krystof Mitka, Stepan Shabalin, Maksym Andriushchenko, Nicolas Flammarion, and Florian Tram`er. Competition report: Finding universal jailbreak backdoors in aligned llms. CoRR, abs/2404.14461, 2024. [151] Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury. Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks. CoRR, abs/2305.14965, 2023. [152] Bin Ren, Yawei Li, Nancy Mehta, Radu Timofte, Hongyuan Yu, Cheng Wan, Yuxin Hong, Bingnan Han, Zhuoyuan Wu, Yajun Zou, Yuqing Liu, Jizhe Li, Keji He, Chao Fan, Heng Zhang, Xiaolin Zhang, Xuanwu Yin, Kunlong Zuo, Bohao Liao, Peizhe Xia, Long Peng, Zhibo Du, Xin Di, Wangkai Li, Yang Wang, Wei Zhai, Renjing Pei, Jiaming Guo, Songcen Xu, Yang Cao, Zhengjun Zha, Yan Wang, Yi Liu, Qing Wang, Gang Zhang, Liou Zhang, Shijie Zhao, Long Sun, Jinshan Pan, Jiangxin Dong, Jinhui Tang, Xin Liu, Min Yan, Qian Wang, Menghan Zhou, Yiqiang Yan, Yixuan Liu, Wensong Chan, Dehua Tang, Dong Zhou, Li Wang, Lu Tian, Emad Barsoum, Bohan Jia, Junbo Qiao, Yunshuai Zhou, Yun Zhang, Wei Li, Shaohui Lin, Shenglong Zhou, Binbin Chen, Jincheng Liao, Suiyi Zhao, Zhao Zhang, Bo Wang, Yan Luo, Yanyan Wei, Feng Li, Mingshen Wang, Yawei Li, Jinhan Guan, Dehua Hu, Jiawei Yu, Qisheng Xu, Tao Sun, Long Lan, Kele Xu, Xin Lin, Jingtong Yue, Lehan Yang, Shiyi Du, Lu Qi, Chao Ren, Zeyu Han, Yuhan Wang, Chaolin Chen, Haobo Li, Mingjun Zheng, Zhongbao Yang, Lianhong Song, Xingzhuo Yan, Minghan Fu, Jingyi Zhang, Baiang Li, Qi Zhu, Xiaogang Xu, Dan Guo, Chunle Guo, Jiadi Chen, Huanhuan Long, Chunjiang Duanmu, Xiaoyan Lei, Jie Liu, Weilin Jia, Weifeng Cao, Wenlong Zhang, Yanyu Mao, Ruilong Guo, Nihao Zhang, Manoj Pandey, Maksym Chernozhukov, Giang Le, Shuli Cheng, Hongyuan Wang, Ziyan Wei, Qingting Tang, Liejun Wang, Yongming Li, Yanhui Guo, Hao Xu, Akram Khatami-Rizi, Ahmad Mahmoudi-Aznaveh, Chih-Chung Hsu, Chia-Ming Lee, Yi-Shiuan Chou, Amogh Joshi, Nikhil Akalwadi, Sampada Malagi, Palani Yashaswini, Chaitra Desai, Ramesh Ashok Tabib, Ujwala Patil, and Uma Mudenagudi. The ninth NTIRE 2024 efficient super-resolution challenge report. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, WA, USA, June 17-18, 2024, pages 6595â6631. IEEE, 2024. [153] Qibing Ren, Chang Gao, Jing Shao, Junchi Yan, Xin Tan, Wai Lam, and Lizhuang Ma. Codeattack: Revealing safety generalization challenges of large language models via code completion. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 11437â11452. Association for Computational Linguistics, 2024. [154] Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao. Derail yourself: Multi-turn LLM jailbreak attack through self-discovered clues. CoRR, abs/2410.10700, 2024. [155] Yupeng Ren. F2A: an innovative approach for prompt injection by utilizing feign security detection agents. CoRR, abs/2410.08776, 2024. [156] Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. Smoothllm: Defending large language models against jailbreaking attacks. Trans. Mach. Learn. Res., 2025, 2025. [157] Mark Russinovich, Ahmed Salem, and Ronen Eldan. Great, now write an article about that: The crescendo multi-turn LLM jailbreak attack. CoRR, abs/2404.01833, 2024. [158] Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Fran ̧cois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan L. Boyd-Graber. Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 4945â4977. Association for Computational Linguistics, 2023. [159] Reshabh K. Sharma, Vinayak Gupta, and Dan Grossman. SPML: A DSL for defending language models against prompt attacks. CoRR, abs/2402.11755, 2024. [160] Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael B. Abu-Ghazaleh. Survey of vulnerabilities in large language models revealed by adversarial attacks. CoRR, abs/2310.10844, 2023. [161] Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. Jailbreak antidote: Runtime safety-utility balance via sparse representation adjustment in large language models. CoRR, abs/2410.02298, 2024. [162] Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. Jailbreak antidote: Runtime safety-utility balance via sparse representation adjustment in large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [163] Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. âdo anything nowâ: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Bo Luo, Xiaojing Liao, Jun Xu, Engin Kirda, and David Lie, editors, Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, pages 1671â1685. ACM, 2024. [164] Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. Optimization-based prompt injection attack to llm-as-a-judge. In Bo Luo, Xiaojing Liao, Jun Xu, Engin Kirda, and David Lie, editors, Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, pages 660â674. ACM, 2024. [165] Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4222â4235. Association for Computational Linguistics, 2020. [166] Sonali Singh, Faranak Abri, and Akbar Siami Namin. Exploiting large language models (llms) through deception tech- niques and persuasion principles. In Jingrui He, Themis Palpanas, Xiaohua Hu, Alfredo Cuzzocrea, Dejing Dou, Dominik Slezak, Wei Wang, Aleksandra Gruca, Jerry Chun-Wei Lin, and Rakesh Agrawal, editors, IEEE International Conference on Big Data, BigData 2023, Sorrento, Italy, December 15-18, 2023, pages 2508â2517. IEEE, 2023. [167] Chawin Sitawarin, Norman Mu, David A. Wagner, and Alexandre Araujo. PAL: proxy-guided black-box attack on large language models. CoRR, abs/2402.09674, 2024. [168] Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer. A strongreject for empty jailbreaks. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [169] Gianluca De Stefano, Lea Sch Ìonherr, and Giancarlo Pellegrino. Rag and roll: An end-to-end evaluation of indirect prompt manipulations in llm-based application frameworks. CoRR, abs/2408.05025, 2024. [170] Penghao Sun, Julong Lan, Yuxiang Hu, Zehua Guo, Chong Wu, and Jiangxing Wu. Realizing the carbon-aware service provision in ICT system. IEEE Trans. Netw. Serv. Manag., 21(4):4090â4103, 2024. [171] Zhifan Sun and Antonio Valerio Miceli Barone. Scaling behavior of machine translation with large language models under prompt injection attacks. CoRR, abs/2403.09832, 2024. [172] ByteDance Seed Team. Seed1.5-vl technical report. arXiv preprint arXiv:2505.07062, 2025. [173] Teknium. Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023. Available at: https: //huggingface.co/datasets/teknium/OpenHermes-2.5. [174] Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su. Evil geniuses: Delving into the safety of llm-based agents. CoRR, abs/2311.11855, 2023. [175] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth Ìe Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aur Ìelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023. [176] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aur Ìelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. [177] Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell. Tensor trust: Interpretable prompt injection attacks from an online game. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [178] Shangqing Tu, Zhuoran Pan, Wenxuan Wang, Zhexin Zhang, Yuliang Sun, Jifan Yu, Hongning Wang, Lei Hou, and Juanzi Li. Knowledge-to-jailbreak: One knowledge point worth one attack. CoRR, abs/2406.11682, 2024. [179] Dmitrii Volkov. Badllama 3: removing safety finetuning from llama 3 in minutes. CoRR, abs/2407.01376, 2024. [180] Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Universal adversarial triggers for attacking and analyzing NLP. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2153â2162. Association for Computational Linguistics, 2019. [181] Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. The instruction hierarchy: Training llms to prioritize privileged instructions. CoRR, abs/2404.13208, 2024. [182] Hao Wang, Hao Li, Minlie Huang, and Lei Sha. From noise to clarity: Unraveling the adversarial suffix of large language model attacks via translation of text embeddings. CoRR, abs/2402.16006, 2024. [183] Haoyu Wang, Bingzhe Wu, Yatao Bian, Yongzhe Chang, Xueqian Wang, and Peilin Zhao. Probing the safety response boundary of large language models via unsafe decoding path generation. CoRR, abs/2408.10668, 2024. [184] Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Junjie Hu, Yixuan Li, Patrick McDaniel, Muhao Chen, Bo Li, and Chaowei Xiao. Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment. CoRR, abs/2402.14968, 2024. [185] Jiongxiao Wang, Zichen Liu, Keun Hee Park, Muhao Chen, and Chaowei Xiao. Adversarial demonstration attacks on large language models. CoRR, abs/2305.14950, 2023. [186] Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. Detoxifying large language models via knowledge editing. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 3093â3118. Association for Computational Linguistics, 2024. [187] Peiran Wang, Xiaogeng Liu, and Chaowei Xiao. Repd: Defending jailbreak attack through a retrieval-based prompt decomposition process. CoRR, abs/2410.08660, 2024. [188] Peiran Wang, Xiaogeng Liu, and Chaowei Xiao. Repd: Defending jailbreak attack through a retrieval-based prompt decomposition process. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computa- tional Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, pages 283â294. Association for Computational Linguistics, 2025. [189] Wenxuan Wang, Kuiyi Gao, Zihan Jia, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Shuai Wang, Wenxiang Jiao, and Zhaopeng Tu. Chain-of-jailbreak attack for image generation models via editing step by step. CoRR, abs/2410.03869, 2024. [190] Xinyuan Wang, Victor Shea-Jay Huang, Renmiao Chen, Hao Wang, Chengwei Pan, Lei Sha, and Minlie Huang. Black- dan: A black-box multi-objective approach for effective and contextual jailbreaking of large language models. CoRR, abs/2410.09804, 2024. [191] Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma, Shuai Wang, Yingjiu Li, Yang Liu, Ning Liu, and Juergen Rahmel. Selfdefend: Llms can defend themselves against jailbreaking in a practical manner. CoRR, abs/2406.05498, 2024. [192] Yihan Wang, Zhouxing Shi, Andrew Bai, and Cho-Jui Hsieh. Defending llms against jailbreaking attacks via backtrans- lation. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 16031â16046. Association for Computational Linguistics, 2024. [193] Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. CoRR, abs/2403.09513, 2024. [194] Zi Wang, Divyam Anshumaan, Ashish Hooda, Yudong Chen, and Somesh Jha. Functional homotopy: Smoothing discrete optimization via continuous parameters for LLM jailbreak attacks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [195] Ziqiu Wang, Jun Liu, Shengkai Zhang, and Yang Yang. Poisoned langchain: Jailbreak llms by langchain. CoRR, abs/2406.18122, 2024. [196] Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does LLM safety training fail?In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. [197] Zeming Wei, Yifei Wang, and Yisen Wang. Jailbreak and guard aligned language models with only few in-context demonstrations. CoRR, abs/2310.06387, 2023. [198] Fenghua Weng, Yue Xu, Chengyan Fu, and Wenjie Wang. Mmj-bench: A comprehensive study on jailbreak attacks and defenses for vision language models. In Toby Walsh, Julie Shah, and Zico Kolter, editors, AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, USA, pages 27689â27697. AAAI Press, 2025. [199] WizardLMTeam. Wizardlmevolinstruct70k, 2023. Available at: https://huggingface.co/datasets/WizardLMTeam/ WizardLM_evol_instruct_70k. [200] Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun. Jailbreaking GPT-4V via self-adversarial attacks with system prompts. CoRR, abs/2311.09127, 2023. [201] Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen. Tastle: Distract large language models for automatic jailbreak attack. CoRR, abs/2403.08424, 2024. [202] Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Zhenqiang Gong. Gradsafe: Detecting unsafe prompts for llms via safety-critical gradient analysis. CoRR, abs/2402.13494, 2024. [203] Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho. Defensive prompt patch: A robust and interpretable defense of llms against jailbreak attacks. CoRR, abs/2405.20099, 2024. [204] Huiyu Xu, Wenhui Zhang, Zhibo Wang, Feng Xiao, Rui Zheng, Yunhe Feng, Zhongjie Ba, and Kui Ren. Redagent: Red teaming large language models with context-aware autonomous language agent. CoRR, abs/2407.16667, 2024. [205] Rongwu Xu, Zehan Qi, and Wei Xu. Preemptive answer âattacksâ on chain-of-thought reasoning. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 14708â14726. Association for Computational Linguistics, 2024. [206] Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. Safedecoding: Defend- ing against jailbreak attacks via safety-aware decoding. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 5587â5605. Association for Computational Linguistics, 2024. [207] Zhao Xu, Fan Liu, and Hao Liu. Bag of tricks: Benchmarking of jailbreak attacks on llms. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [208] Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. A comprehensive study of jailbreak attack versus defense for large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 7432â7449. Association for Computational Linguistics, 2024. [209] An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. Qwen2 technical report. CoRR, abs/2407.10671, 2024. [210] An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report. CoRR, abs/2412.15115, 2024. [211] Sin-Han Yang, Tuomas P. Oikarinen, and Tsui-Wei Weng. Concept-driven continual learning. Trans. Mach. Learn. Res., 2024, 2024. [212] Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han. Chain of attack: a semantic-driven contextual multi-turn attacker for LLM. CoRR, abs/2405.05610, 2024. [213] Yan Yang, Zeguan Xiao, Xin Lu, Hongru Wang, Hailiang Huang, Guanhua Chen, and Yun Chen. Sop: Unlock the power of social facilitation for automatic jailbreak attack. CoRR, abs/2407.01902, 2024. [214] Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. Sneakyprompt: Jailbreaking text-to-image generative models. In IEEE Symposium on Security and Privacy, SP 2024, San Francisco, CA, USA, May 19-23, 2024, pages 897â912. IEEE, 2024. [215] Dongyu Yao, Jianshu Zhang, Ian G. Harris, and Marcel Carlsson. Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. CoRR, abs/2309.05274, 2023. [216] Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang. Toolsword: Unveiling safety issues of large language models in tool learning across three stages. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 2181â2211. Association for Computational Linguistics, 2024. [217] Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li. Jailbreak attacks and defenses against large language models: A survey. CoRR, abs/2407.04295, 2024. [218] Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xianglong Liu, and Dacheng Tao. Jailbreak vision language models via bi-modal adversarial prompt. CoRR, abs/2406.04031, 2024. [219] Zheng Xin Yong, Cristina Menghini, and Stephen H. Bach.Low-resource languages jailbreak GPT-4.CoRR, abs/2310.02446, 2023. [220] Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. GPTFUZZER: red teaming large language models with auto- generated jailbreak prompts. CoRR, abs/2309.10253, 2023. [221] Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [222] Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang. Donât listen to me: Un- derstanding and exploring jailbreak prompts of large language models. In Davide Balzarotti and Wenyuan Xu, editors, 33rd USENIX Security Symposium, USENIX Security 2024, Philadelphia, PA, USA, August 14-16, 2024. USENIX Association, 2024. [223] Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. GPT-4 is too smart to be safe: Stealthy chat with llms via cipher. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [224] Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Song, and Bo Li. Rigorllm: Resilient guardrails for large language models against undesired content. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. [225] Xinyi Zeng, Yuying Shang, Yutao Zhu, Jiawei Chen, and Yu Tian. Root defence strategies: Ensuring safety of LLM at the decoding level. CoRR, abs/2410.06809, 2024. [226] Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing llms. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 14322â14350. Association for Computational Linguistics, 2024. [227] Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li. Air-bench 2024: A safety benchmark based on risk categories from regulations and policies. CoRR, abs/2407.17436, 2024. [228] Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu. Autodefense: Multi-agent LLM defense against jailbreak attacks. CoRR, abs/2403.04783, 2024. [229] Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 10471â10506. Association for Computational Linguistics, 2024. [230] Chong Zhang, Mingyu Jin, Qinkai Yu, Chengzhi Liu, Haochen Xue, and Xiaobo Jin. Goal-guided generative prompt injection attack on large language models. In Elena Baralis, Kun Zhang, Ernesto Damiani, M Ìerouane Debbah, Panos Kalnis, and Xindong Wu, editors, IEEE International Conference on Data Mining, ICDM 2024, Abu Dhabi, United Arab Emirates, December 9-12, 2024, pages 941â946. IEEE, 2024. [231] Hangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao, Lu Lin, Jinyuan Jia, Jinghui Chen, and Dinghao Wu. Jail- break open-sourced large language models via enforced decoding. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 5475â5493. Association for Computational Linguistics, 2024. [232] Xiaoyu Zhang, Cen Zhang, Tianlin Li, Yihao Huang, Xiaojun Jia, Xiaofei Xie, Yang Liu, and Chao Shen. A mutation- based method for multi-modal jailbreaking attack detection. CoRR, abs/2312.10766, 2023. [233] Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao. Intention analysis prompting makes large language models A good jailbreak defender. CoRR, abs/2401.06561, 2024. [234] Zaibin Zhang, Yongting Zhang, Lijun Li, Jing Shao, Hongzhi Gao, Yu Qiao, Lijun Wang, Huchuan Lu, and Feng Zhao. Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 15202â15231. Association for Computational Linguistics, 2024. [235] Zhexin Zhang, Junxiao Yang, Pei Ke, Shiyao Cui, Chujie Zheng, Hongning Wang, and Minlie Huang. Safe unlearning: A surprisingly effective and generalizable solution to defend against jailbreak attacks. CoRR, abs/2407.02855, 2024. [236] Zhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng, and Xiangyu Zhang. On large language modelsâ resilience to coercive interrogation. In IEEE Symposium on Security and Privacy, SP 2024, San Francisco, CA, USA, May 19-23, 2024, pages 826â844. IEEE, 2024. [237] Bowen Zhao, Wei-Neng Chen, Feng-Feng Wei, Ximeng Liu, Qingqi Pei, and Jun Zhang. PEGA: A privacy-preserving genetic algorithm for combinatorial optimization. IEEE Trans. Cybern., 54(6):3638â3651, 2024. [238] Jiawei Zhao, Kejiang Chen, Xiaojian Yuan, and Weiming Zhang. Prefix guidance: A steering wheel for large language models to defend against jailbreak attacks. CoRR, abs/2408.08924, 2024. [239] Wei Zhao, Zhe Li, Yige Li, and Jun Sun. Adversarial suffixes may be features too! CoRR, abs/2410.00451, 2024. [240] Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun. Defending large language models against jailbreak attacks via layer-specific editing. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, pages 5094â5109. Association for Computational Linguistics, 2024. [241] Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang. Weak-to-strong jailbreaking on large language models. CoRR, abs/2401.17256, 2024. [242] Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. Prompt- driven LLM safeguarding via directed representation optimization. CoRR, abs/2401.18018, 2024. [243] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. Lmsys-chat-1m: A large-scale real-world LLM conversation dataset. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. [244] Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. Improved few-shot jailbreaking can circumvent aligned language models and their defenses. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [245] Andy Zhou, Bo Li, and Haohan Wang. Robust prompt optimization for defending language models against jailbreaking attacks. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [246] Yukai Zhou and Wenjie Wang. Donât say no: Jailbreaking LLM by suppressing refusal. CoRR, abs/2404.16369, 2024. [247] Yuqi Zhou, Lin Lu, Ryan Sun, Pan Zhou, and Lichao Sun. Virtual context enhancing jailbreak attacks with special token injection. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, pages 11843â11857. Association for Computational Linguistics, 2024. [248] Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. Autodan: interpretable gradient-based adversarial attacks on large language models. arXiv preprint arXiv:2310.15140, 2023. [249] Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy M. Hospedales. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. [250] Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J. Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. [251] Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. CoRR, abs/2307.15043, 2023. [252] Xiaotian Zou and Yongkang Chen. Image-to-text logic jailbreak: Your imagination can help you do anything. CoRR, abs/2407.02534, 2024. [253] Xiaotian Zou, Yongkang Chen, and Ke Li. Is the system message really important to jailbreaks in large language models? CoRR, abs/2402.14857, 2024.