Paper deep dive
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.
Tags
Links
- Source: https://arxiv.org/abs/2608.23028v1
- Canonical: https://arxiv.org/abs/2608.23028v1
Trouble viewing inline? Open PDF directly â
Full Text
94,694 characters extracted from source content.
Expand or collapse full text
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies Zeyu Feng fengzeyuqwe@gmail.com Qingyu Wu andre.qingyu.wu@gmail.com Yuzhe Luo 782682325@q.com Hua Cheng chenghua@ncic.ac.cn organization=The Defense Innovation Institute, Academy of Military Sciences, city=Beijing, country=China Abstract Large language models (LLMs) are now deployed across education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This deployment makes jailbreaks a growing threat to LLM safety. Yet most jailbreak research still emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded, multi-turn vulnerabilities of target LLMs underexplored. To address this gap, we present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established persuasion techniques from social psychology into a tactic-conditioned attack policy, factorizes each attacker action into a Change-of-Meaning analysis, a tactic selection, and a victim-visible messageâoperationalizing the Persuasion Knowledge Model (PKM)âand refines this policy with trajectory-level reinforcement learning under a PKM-gated reward that credits early jailbreak success only when every turn carries a well-formed change-of-meaning analysis. Across four aligned victim models, PsychJail attains the highest average attack success rate (87.3%), surpassing strong single-turn and multi-turn baselines across all four models. We further measure susceptibility at the level of the action that breaks each victim. This analysis recovers four empirically distinct per-model susceptibility fingerprints: which persuasion levers open which model, and how broadly. These fingerprints explain the observed cross-model transfer asymmetry. We interpret them as four candidate psychological profiles (rationalist, credibility-driven, narrative-monoculture, and broadly persuadable), while treating that interpretation as a conjecture for future validation. These findings position psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs. keywords large language models ,red teaming ,psychological persuasion ,multi-turn jailbreak â corresponding: Corresponding author.â These authors contributed equally to this work. 1 Introduction Large language models (LLMs) are increasingly deployed in interactive scenarios such as education, healthcare, and policy advising, where users engage them via sustained multi-turn dialogue. As these systems become more capable and human-like in interaction, safety evaluations must cover not only static adversarial prompt attacks but also alignment failures that emerge during long conversational interactions. Jailbreak attacks have become the dominant red-teaming paradigm for exposing such vulnerabilities (46; 29; 39). Recent work has shown that automated single-turn jailbreak generation can be highly effective. Representative methods include GCG, which optimizes adversarial suffixes through gradient-guided search, and AutoDAN, which uses genetic search to produce stealthy and semantically meaningful jailbreak prompts. Later variants such as AmpleGCG, I-GCG, and ECLIPSE (26; 18; 22) further improve transferability, efficiency, or black-box applicability. Yet this dominant line of work still frames jailbreak mainly as prompt optimization against an adversarially exploitable system, leaving a complementary threat model underexplored: whether introducing human social manipulation knowledge into red-teaming models increases the jailbreak susceptibility of target LLMs, particularly in multi-turn dialogue. This research question matters because users increasingly engage LLMs as sustained interlocutors rather than static query engines. Prior work already suggests that rhetoric, role framing, and psychologically informed prompting can erode safety behavior. PAP (48) shows that persuasive prompts derived from social-science taxonomies can substantially increase jailbreak success, while in-the-wild studies such as DAN and WildTeaming (36; 21) indicate that socially framed and role-play-heavy jailbreak families remain widespread and practically effective. Existing evidence is nevertheless fragmented. PAP introduces a socially grounded persuasion taxonomy, and systems such as GUARD (23) and WildTeaming can generate or recombine natural-language jailbreak tactics, but these methods still operate largely at the level of single-turn prompt construction or prompt-level tactic composition. More recent multi-turn attacks such as Crescendo and FITD (34; 40) demonstrate the effectiveness of gradual escalation, and trajectory-level methods such as TROJail (42) show that long-horizon attack objectives can be optimized directly. What remains insufficiently understood is whether explicitly injecting a reusable repertoire of human social-manipulation knowledge into an automated red-teaming policy yields systematic gains over generic tactic search or heuristic escalation in sustained dialogue. Figure 1: Overview of the PsychJail framework. For each harmful objective 0 x_0, the attacker Ďθ _θ emits the factorized action t=(t,ct,t) a_t=( n_t,c_t, m_t); only t m_t reaches the frozen victim Ďv _v; the judge ĎJ _J scores t y_t as stâ[0,1]s_tâ[0,1]. Strict-parse failure at any turn truncates the trajectory. The trajectory reward combines the per-turn format scores with the peak success term gated on PKM-aligned change-of-meaning analysis at every turn. To address this gap, we introduce PsychJail. As illustrated in Figure 1, the core idea is to humanize the attacker training pipeline rather than merely optimize prompts. PsychJail equips an attacker model with a repertoire of 40 persuasion techniques distilled from social psychology. At each turn, the attacker emits a factorized action: an analysis of whether the victim has reinterpreted the previous tactic, an explicit tactic commitment, and the only message exposed to the victim. This decomposition operationalizes the Change-of-Meaning Principle from the Persuasion Knowledge Model (9). The resulting policy is refined with trajectory-level reinforcement learning under a PKM-gated reward, which credits early jailbreak success only when every turn carries a well-formed change-of-meaning analysis. The goal is not simply to produce persuasive prompts one turn at a time, but to train a multi-turn persuasion policy that dynamically adjusts pressure as the target modelâs interpretation shifts across turns. This framing does not require strong anthropomorphic claims; rather, it tests whether humanized attacker training reveals vulnerabilities that prompt-search formulations may miss. Our empirical study supports a humanization-inspired view of jailbreak as multi-turn psychological persuasion rather than prompt search alone. First, PsychJail outperforms strong single-turn and multi-turn baselines across four victim models, attaining the highest average ASR of 87.3%. Second, targeted ablations attribute the gains to its PKM-guided designâthe change-of-meaning gate, the early-success weighting, and the warm-startârather than to generic long-horizon optimization. Third, measuring susceptibility at the breaking action recovers four empirically distinct per-model susceptibility fingerprints, showing which persuasion levers open which model and how broadly. These fingerprints explain the policyâs cross-model transfer asymmetry. We read them as four candidate psychological profiles (rationalist, credibility-driven, narrative-monoculture, and broadly persuadable), while marking that reading as a conjecture for future validation. Finally, two independent judges confirm that the policyâs persuasion labels are faithfully instantiated rather than collapsing onto a single repeated move. Together, these results establish psychological jailbreak as a distinct and practically effective red-teaming frontier for increasingly interactive LLMs. In summary, this paper makes the following contributions: ⢠Jailbreak as psychological persuasion. We recast multi-turn jailbreak as a problem of persuasion rather than prompt optimization (Section 3.1). This moves the unit of study from the adversarial prompt to a learned policy that selects and sequences human persuasion tactics as the targetâs interpretation shifts across turns. ⢠The PsychJail framework. PsychJail instantiates this view as a tactic-conditioned policy grounded in a social-psychology taxonomy (Sections 3.3 â3.5). Each attacker action factorizes, in a PKM-aligned decomposition, into a change-of-meaning analysis, a tactic, and the sole victim-visible message. Trajectory-level reinforcement learning under a PKM-gated outcome reward then credits early jailbreak success only on trajectories whose every turn carries a well-formed analysis. ⢠Empirical analysis and susceptibility fingerprints. A multi-layered evaluation shows that PsychJail achieves the highest average ASR (87.3%) among strong single-turn and multi-turn baselines (Section 4). Targeted ablations attribute these gains to the PKM-guided design rather than to generic long-horizon optimization, and breaking-action analysis reveals four empirically distinct per-model susceptibility fingerprints (Section 4.7). These fingerprints are faithfully labeled, explain the observed cross-model transfer asymmetry, and show that PsychJail does not merely reuse a single persuasion script. 2 Related Work 2.1 Prompt-Level Jailbreak Attacks Early jailbreak research is dominated by prompt-level attack generation. White-box optimization methods such as GCG (52) show that adversarial suffixes can transfer across aligned models, while gradient-free or semantically meaningful variants such as AutoDAN (28), ReNeLLM (7), and ArtPrompt (20) reduce manual effort and expand the effective adversarial prompt space. Jailbreak-R1 extends this line by using reinforcement-learning-style optimization to improve single-turn attacks (14; 27). Taken together, these works establish that jailbreak success can be substantially improved by better search, optimization, and prompt generation. However, they still treat the attack primarily as a single-shot prompt construction problem, leaving limited room to study how harmful intent is incrementally negotiated across turns. 2.2 Multi-Turn Jailbreak Optimization Recent work instead treats jailbreaking as an interactive process. PAIR frames black-box jailbreak as iterative conversation (4). ActorAttack and CoA show that multi-turn attacks can be guided by self-discovered clues or intent-concealing interrogation that escalates across turns (32; 45). Siren, MTSA, and X-Teaming further demonstrate that learned multi-turn attackers, multi-round red teaming, and adaptive multi-agent orchestration can strengthen jailbreak capability (50; 13; 31; 51; 2). TROJail pushes this direction further with trajectory-level optimization and process rewards over full rollouts (42). This trajectory-level view is enabled by recent progress in multi-turn agent RL, where self-evolving multi-turn reinforcement learning (38) and turn-level credit assignment (47) make long-horizon optimization tractable; PsychJailâs trajectory-level GRPO objective builds directly on this machinery. This literature collectively shows that multi-turn dialogue is a genuine attack surface and that long-horizon optimization materially affects attack success. Yet the policies learned by these methods are typically generic: they improve dialogue continuation and escalation, but do not explicitly choose, name, and sequence interpretable persuasion tactics. 2.3 Persuasion-Oriented Jailbreaks and Evaluation A complementary line of work shows that jailbreaks are shaped not only by optimization, but also by rhetoric and social framing. PAP maps persuasion strategies from social-science taxonomies to jailbreak prompts and shows substantial gains from persuasive framing (48). The social-psychology basis for these gains is the Persuasion Knowledge Model (9; 6; 5), which characterizes how a target recognizes an incoming message as a persuasion attempt and revises its interpretation accordingly; a static prompt template cannot react to such adaptation as it unfolds across a conversation. GUARD uses role-playing to generate natural language jailbreaks for testing guideline adherence (23), while analyses of in-the-wild DAN-style prompts and WildTeaming show that socially framed jailbreaks remain effective outside curated benchmarks (36; 21). These findings motivate our core premise that aligned LLMs can be vulnerable to human-like persuasion rather than only to adversarial strings. At the same time, existing persuasion-oriented work is still mostly prompt-centric: it studies persuasive templates or prompt families, not a learned multi-turn persuasion policy (48; 23; 36; 21). This limitation also affects evaluation. Much of current evaluation still focuses on aggregate ASR or prompt-local robustness (29; 3; 37; 25; 33; 49; 15; 16; 41; 43; 46). A persuasion-oriented multi-turn method such as PsychJail should also be assessed through turn efficiency, model-specific susceptibility profiles, and cross-model reuse of successful trajectories. Table 1: Positioning of PsychJail in the design space of representative jailbreak families. The columns are descriptive coordinates, not a quality scorecard: a Ă marks a different design choice, not a deficiency. MT: operates over a multi-turn dialogue; Traj-RL: trajectory-level reinforcement learning over full rollouts; Tax.: grounded in an explicit social-science persuasion-tactic taxonomy; Adaptive: conditions each move on the targetâs evolving response rather than emitting a fixed prompt; Policy: learns a policy that explicitly selects and sequences tactics from that taxonomy (vs. static templates or generic moves). â yes, Ă no. Method MT Traj-RL Tax. Adaptive Policy GCG, AutoDAN, ReNeLLM (52; 28; 7) Ă Ă Ă Ă Ă PAP (48) Ă Ă â Ă Ă PAIR (4) â Ă Ă â Ă CoA, ActorAttack (45; 32) â Ă Ă â Ă Siren, X-Teaming (50; 31) â Ă Ă â Ă TROJail (42) â â Ă â Ă PsychJail (ours) â â â â â Table 1 illustrates the architectural positioning of PsychJail against existing jailbreak paradigms. The axes are deliberately ones that several baselines already satisfyâmost operate over multiple turns and adapt to the targetâs responses, PAP supplies an explicit persuasion taxonomy, and TROJail optimizes at the trajectory levelâso the table is a set of coordinates rather than a scorecard, and a Ă records a different design choice, not a deficiency. Prior families each occupy a strict subset: prompt-level attacks optimize a single shot; persuasion-oriented work such as PAP supplies a taxonomy but only as static templates; generic multi-turn attackers optimize dialogue continuation rather than tactic selection; and even the trajectory-level optimizer TROJail carries no explicit persuasion vocabulary. What is unoccupied is their conjunction: PsychJail is the first to learn a trajectory-level policy that selects and sequences tactics from an explicit persuasion repertoire while adapting to the target across turns. Beyond these shared axes, PsychJail further grounds each action in the Persuasion Knowledge Model and exposes trajectory-level diagnostics (turn efficiency, per-model susceptibility fingerprints, cross-model transfer); we present these as contributions in their own right rather than as comparison axes. 3 Methodology 3.1 Motivation and Design Principles Prior multi-turn jailbreak formulations treat each attacker action as a single, opaque message and optimize dialogue continuation strategies (34; 32; 42). This suffices to escalate pressure across turns, but it leaves two central quantities in persuasion unmodeled: which psychological lever a turn deploys, and whether that lever has shifted the targetâs interpretation of the request. A message-only action cannot represent these quantities. It therefore keeps the attacker confined to generic escalation, even when the long-horizon objective is optimized effectively. Taking persuasion seriously changes the object of study, not merely the optimization technique: the unit of analysis becomes the tactic that compromises a particular target, and red-teaming becomes a per-model study of susceptibility rather than a search for a single successful jailbreak prompt. The reframing is timely. As LLMs are deployed as sustained interlocutors and grow more human-like in interaction, the relevant attack surface increasingly resembles the one a skilled human persuader would exploit. An automated attacker that learns and exposes that surface turns it into something measurable. Realizing such an attacker demands three properties that an unstructured message action cannot supply; we make them the foundation of PsychJailâs design. Adaptivity to the targetâs shifting interpretation. The Persuasion Knowledge Model (PKM) holds that the target of a persuasion attempt continuously re-interprets it, and that an effective persuader adapts as this interpretationâthe meaning the target assigns to the requestâevolves. An attacker faithful to this principle cannot commit to a fixed prompt: it must read the victimâs latest reply and choose its next move in response, which makes the attack intrinsically multi-turn and state-dependent. Without this adaptation, the attacker becomes a script that cannot distinguish a victim that has recognized a tactic from one that has not, and therefore loses the mechanism that separates persuasion from repetition. Learned selection from an explicit repertoire. Which tactic will compromise a given target is not known in advance and, as our later analysis shows (Section 4.7), varies from model to model. A faithful attacker must therefore learn a policy that selects and sequences tactics from an explicit, psychologically grounded repertoire, rather than instantiate a single hand-written template or rely on undirected prompt search. The explicit repertoire makes each choice nameable; the learned policy makes the choice target-specific. A fixed template commits to a tactic before encountering the target, and prompt search does not name the tactic at all, so neither can recover a susceptibility structure that is unknown a priori. Auditability of intent and effect. For the learned behaviour to read as persuasion rather than an inscrutable string, each action must expose its psychological intentâthe tactic it commits toâtogether with the attackerâs inference about the targetâs state, kept separate from the message the victim actually sees. This separation is what later lets us attribute success to specific levers and verify that the declared tactics are genuinely enacted rather than free-floating labels (Section 4.8). Strip this exposure away and a jailbreak is only an opaque transcript: one cannot say which lever carried it, cannot tell persuasion from a parser exploit, and cannot aggregate individual attacks into the per-model susceptibility map that is the scientific payoff. Auditability is therefore a precondition for measurement, not a reporting convenience. These three principles are not independent conveniences to be traded off against one another; they are jointly necessary, and PsychJail discharges each with exactly one component. Adaptivity and auditability are realized by a factorized, tactic-conditioned action that couples an interpretation, a tactic commitment, and a victim-visible message (Section 3.3); learned selection from the repertoire is instilled by a warm start that teaches the protocol while prejudging no tactic (Section 3.4); and adaptivity is turned into an optimization target, under audit, by a trajectory-level PKM-gated reward that credits only fully analyzed persuasive success (Section 3.5). Removing any one of these components would reduce the attacker to generic escalation. Together, they convert multi-turn red-teaming from the search for a single successful jailbreak prompt into an instrument that measures, for each target, which persuasion levers expose its susceptibilityâthe fingerprints analyzed in Section 4.7. We formalize the setting next and develop each component in turn. 3.2 Problem Formulation We formulate psychological jailbreak as a multi-turn reinforcement learning problem with three actors: an attacker policy Ďθ _θ (the only trainable component), a frozen victim Ďv _v, and a frozen judge ĎJ _J. Let 0â x_0 denote a harmful objective drawn from the target distribution D. Throughout, bold symbols denote token sequences and italic symbols denote scalars or discrete categorical labels. The turn-t trajectory prefix is t=[(1,1),âŚ,(t,t)] Ď_t=[( a_1, y_1),âŚ,( a_t, y_t)] with 0âĄâ Ď_0⥠, and a full trajectory is âĄ|| Ď⥠Ď_| Ď| with ||â¤T| Ď|⤠T. At each turn tâ1,âŚ,Ttâ\1,âŚ,T\ the attacker emits a structured action t=(t,ct,t)âźĎθ(â âŁ0,tâ1), a_t=( n_t,\,c_t,\, m_t)\; \; _θ (¡ x_0,\, Ď_t-1 ), (1) where t n_t is a natural-language persuasion-knowledge analysis of the victimâs previous reply, ctâc_t is a tactic drawn from a finite persuasion taxonomy C with ||=40|C|=40 (Section 3.3), and t m_t is the victim-visible message; only t m_t is forwarded to the victim. The victim replies tâźĎv(â âŁâ¤t,<t) y_t _v(¡ m_⤠t,\, y_<t), conditioning only on the messages it has received and its own prior repliesânever on the hidden objective 0 x_0âand the judge ĎJ _J scores t y_t for harmful compliance with 0 x_0, returning a scalar stâ[0,1]s_tâ[0,1] whose concrete instantiation as a HarmBench (29) verdict probability is given in Section 3.5. Unlike trajectory-level multi-turn jailbreak formulations whose action is a single message (42; 32), Equation (1) factorizes each action into an interpretive component t n_t, a tactic-selection component ctc_t, and a surface component t m_t. This factorization is motivated by the PKM and is enforced through the generation protocol described in Section 3.3. The protocol also explains why a single victim-visible message is insufficient: it cannot expose the attackerâs interpretation of the victimâs state or the tactic selected in response to that interpretation. We train Ďθ _θ by trajectory-level reinforcement learning on a single scalar outcome reward Roâ()ââR_o( Ď) ; the reward and the optimization objective are specified in Section 3.5. 3.3 Tactic-Conditioned Generation Protocol Adaptivity and auditabilityâthe first and third design principlesâremain, so far, desiderata about how the attack should behave. Here we make them structural properties of the action itself, so that neither is left to the trained policyâs discretion. The design question is what an attacker action must be so that adapting to the targetâs shifting interpretation is forced by the actionâs form, and so that every move carries an auditable record of the lever it commits to. We answer in two steps: we first cast the Persuasion Knowledge Model as partially observed control, which dictates that the action factorize into a belief, a tactic, and a message; we then promote that factorization to a hard grammatical constraint, so that adaptivity and auditability hold by construction rather than by post-hoc inspection. PKM as partially observed control. The PKM posits a Change-of-Meaning Principle: once a target reinterprets an incoming utterance as a persuasion attempt, the cognitive route for processing it shifts and the same surface tactic loses its force. We cast this principle in a control-theoretic form that the attacker can be trained against. We refine the victim of Section 3.2 into a controlled latent process: at turn t it carries a hidden persuasion-knowledge state ztâz_t that summarizes which incoming moves it has already reinterpreted as persuasion, and both its reply and the judge score factor through ztz_t, ztâźPv(â âŁztâ1,t),tâźĎv(â âŁzt,â¤t,<t).z_t P_v\! (¡ z_t-1, m_t ), y_t _v\! (¡ z_t, m_⤠t, y_<t ). (2) The Change-of-Meaning Principle is then a structural constraint on the transition kernel PvP_v: as soon as t m_t is recognized as a persuasion attempt, ztz_t advances to a guarded mode gâĄ(c)g(c) for the tactic family c it instantiates, in which replaying that family is strictly less effective, [stâŁzt=g(c),ct=c]<[stâŁzt=naive,ct=c].E\! [s_t z_t=g(c),\,c_t=c ]\;<\;E\! [s_t z_t=naive,\,c_t=c ]. (3) Because the attacker never observes ztz_t and sees only the victimâs past replies <t y_<t, the interaction is a partially observable Markov decision process. By the belief-state sufficiency of POMDPs (1; 24), an optimal attacker depends on the history only through the posterior belief btâ(z)=PrâĄ(ztâ1=zâŁ0,tâ1)b_t(z)\;=\; \! (z_t-1=z x_0, Ď_t-1 ) (4) over the victimâs latent state at the close of turn tâ1t-1âthe state that generated the observed reply tâ1 y_t-1, and hence the most recent guard configuration the attacker can infer before committing t a_t. A memoryless attacker whose action is a single victim-visible messageâas in trajectory-level formulations that emit only t m_t (42; 32)âcannot represent btb_t, and therefore can neither detect a change-of-meaning event nor redirect its tactic in response to one. PsychJail makes this belief explicit and trainable. The analysis t n_t is a natural-language realization of the belief update on the latest observation tâ1 y_t-1, and the factorized action of Equation (1) is generated autoregressively in belief-first order, Ďθâ(tâŁ0,tâ1)= _θ( a_t x_0, Ď_t-1)= Ďθâ(tâŁ0,tâ1)âbelief update âbt _θ( n_t x_0, Ď_t-1)_belief update b_t (5) ĂĎθâ(ctâŁ0,tâ1,t)âbelief-conditioned tactic Ă\, _θ(c_t x_0, Ď_t-1, n_t)_belief-conditioned tactic ĂĎθâ(tâŁ0,tâ1,t,ct). Ă\, _θ( m_t x_0, Ď_t-1, n_t,c_t). The attacker is thus a belief-MDP policy by construction: it must form the belief t n_t before committing to a tactic ctc_t, rather than reacting memorylessly to the raw dialogue. The generation protocol below turns this factorization into a hard constraintâits ordering rule enforces the belief-first dependency of Equation (5), and its strict-parse gate, inherited by the reward of Section 3.5, guarantees that every credited trajectory is a genuine belief-conditioned rollout. Discrete tactic space. We instantiate C with the 40-technique persuasion taxonomy from PAP (48), which compiles canonical influence techniques from social-psychology meta-reviews. Each entry câc is rendered verbatim into the attackerâs system prompt as the triple (nameâĄ(c),definitionâĄ(c),exampleâĄ(c))(name(c),\,definition(c),\,example(c)), so Ďθ _θ sees the canonical inventory at every decoding step. We enforce ctâc_t only at the warm-start SFT stage (Section 3.4), where demonstration trajectories are rejected during quality control if the â¨technique⊠field falls outside the canonical inventory under exact (case-insensitive) matching, with no fuzzy fallback. We deliberately drop this hard constraint at the RL stage: any non-empty â¨technique⊠that satisfies the strict XML grammar is admissible, leaving the policy free to invent off-taxonomy tactics if doing so improves the trajectory-level reward. The canonical rate Pr[ctâ] [c_t ] is tracked as a post-hoc interpretability metric, with off-taxonomy emissions further deduplicated against canonical names by normalized Levenshtein distance to separate genuine emergence from orthographic variants (Section 4); the surface form t m_t is unconstrained throughout. XML generation schema. At each turn, Ďθ _θ emits a single token stream that must parse strictly as <analysis> t n_t </analysis> <technique> ctc_t </technique> <message> t m_t </message> in this exact order, with no extraneous tags between sections. The three sections carry distinct epistemic roles aligned with PKM: t n_t is the attackerâs belief about whether tâ1 y_t-1 signaled change-of-meaning; ctc_t is the resulting tactic commitment; and t m_t is the only surface that reaches Ďv _v, ensuring that the victim cannot exploit the attackerâs internal reasoning. Let strictâĄ(t)â0,1strict( a_t)â\0,1\ indicate whether turn t parses under this grammar. Strict-parse trajectory truncation. If strictâĄ(t)=0strict( a_t)=0 at any turn t, we terminate the trajectory immediately, without querying either the victim model or the judge model on the malformed action, and set ||=t| Ď|=t. This rule is not a defensive filter but a logical consequence of the PKM coupling: without a well-formed analysis t n_t, the attacker has bypassed the change-of-meaning inference and the remaining (ct,t)(c_t, m_t) pair is no longer a sample from the tactic-conditioned policy of Equation (1). We therefore treat strict adherence to the schema as an integral part of the action space, and require it across all turns of a trajectory before the outcome reward credits a successful attack. 3.4 Warm-Start Supervised Fine-Tuning The second design principle requires that the choice of lever be learned from the explicit repertoire, not fixed in advance, and this places an unusual demand on the warm start. Its task is to install the repertoire and the generation protocol while installing no preference over which tactic succeeds, since discovering that preference is exactly what reinforcement learning is for; a warm start tuned for jailbreak success would prejudge the per-model susceptibility structure the policy is meant to learn. We therefore separate competence from outcomeâteaching the attacker how to act in the protocol, and leaving which acts pay off to RL. A randomly initialized attacker rarely emits the structured action of Equation (1) consistently, so trajectory-level RL from scratch wastes most early rollouts on strict-parse failures and never accumulates a useful gradient on ctc_t or mtm_t. We therefore initialize Ďθ _θ from a warm-start checkpoint Ďθref _ _ref that has already absorbed the generation protocol but has not been biased toward any single jailbreak outcome. Distilled multi-victim demonstrations. We construct demonstrations by driving an uncensored 405B-parameter teacher attacker (Dolphin-X1-Llama-3.1-405B-FP8) against a pool of four aligned victimsâQwen2.5-7B-Instruct, Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and Mistral-7B-Instruct-v0.3âusing the schema of Section 3.3. Harmful objectives are sampled from BeaverTails-30k (17) and split disjointly across victims so that each demonstration trajectory engages exactly one victim. The teacher is prompted with the same 40-technique system prompt as Ďθ _θ, and per-turn outputs that violate the strict XML grammar are regenerated under an error-specific retry instruction; objectives that exceed a fixed retry budget are dropped. Protocol-only quality control. Each candidate trajectory is then admitted only if (i) every attacker turn parses under the strict XML grammar and (i) every <technique> body matches an entry of C under exact (case-insensitive) name equality. We deliberately do not filter by judge harm scores or by victim refusal patterns: filtering by success would inject a success-bias prior that the downstream RL objective is itself supposed to discover, and would collapse the demonstration set onto whichever tactics happened to work against the teacherâs victim pool. The role of SFT in PsychJail is therefore strictly to internalize the generation protocolâthe analysis â tactic â message factorization and the canonical tactic vocabularyânot to seed any particular persuasion strategy. The resulting reference policy Ďθref _ _ref serves both as the RL initialization and as the reference distribution in the KL term of Equation (12); the demonstration pool size, SFT epochs, and GPU configuration are reported in Section 4.1. 3.5 Persuasion-Aware Reward Design A representation that can adapt and can be audited is inert unless the training signal rewards adaptation and refuses to credit unauditable success; the reward is where the first and third principles stop being merely expressible and become objectives the policy is trained against. Two requirements follow. Because adaptivity is a property of the whole exchange, credit must be assigned at the trajectory level and must prefer success reached in the fewest persuasive turns, rather than rewarding any message in isolation. And because auditability is the precondition for attributing success to a lever, the reward must withhold all success credit from a trajectory in which any turn skipped the change-of-meaning analysis, so that what is reinforced is persuasion under audit rather than a well-formed string that happened to land. Both demands fall on a peak success reward that sparsely credits effective multi-turn attacks under the PKM coupling; a second, dense format reward that supervises adherence to the protocol of Section 3.3 plays the supporting role of shaping early learning toward the well-formed trajectories on which that credit can be earned. We develop the two terms in turn. Per-turn format reward. Let oto_t denote the attackerâs raw output at turn t, T the set of six literal tags \<analysis>, </analysis>, <technique>, </technique>, <message>, </message>\, and P the three open/close pairs formed from T. We award one point for each well-formedness condition, partitioned into three groups: ⢠Tag presence (up to |||T| points): each ĎâĎ appears exactly once in oto_t. ⢠Ordering (up to |||P| points): for each pair (Ďo,Ďc)â( _o, _c) , both tags satisfy the presence condition and Ďo _o occurs before Ďc _c in oto_t. ⢠Nesting-free interior (up to |||P| points): for each pair that satisfies the ordering condition, the substring of oto_t between Ďo _o and Ďc _c contains no other tag from T. Let Ďpresâ(ot) _pres(o_t), Ďordâ(ot) _ord(o_t), and Ďnestâ(ot) _nest(o_t) denote the awarded counts under the three groups, respectively, and let fmax=||+2â||=12f_ =|T|+2|P|=12 be the resulting number of well-formedness conditionsâone per tag, plus an ordering and a nesting check per pair. The per-turn format score is ft=Ďpresâ(ot)+Ďordâ(ot)+Ďnestâ(ot)â[0,fmax],f_t\;=\; _pres(o_t)+ _ord(o_t)+ _nest(o_t)\;â\;[0,f_ ], (6) where the maximum is attained (ft=fmaxf_t=f_ ) exactly when oto_t meets every well-formedness condition above. This score and the binary predicate strictâĄ(t)strict( a_t) of Section 3.3 play complementary roles: the strict predicateârealized by the anchored analysisâ â parserâgates trajectory termination and success credit, while ftf_t supplies a dense, graceful shaping signal during early RL, evaluated on every produced turnâincluding a malformed turn whose strict-parse failure truncates the trajectory, whose partial score still contributes to the format reward below. Peak success reward. We instantiate the judge ĎJ _J with the HarmBench classifier (29) and read stâ[0,1]s_tâ[0,1] as the renormalized probability it assigns to the affirmative verdict on t y_t against objective 0 x_0, st=pĎJâ(YesâŁ0,t)pĎJâ(YesâŁ0,t)+pĎJâ(NoâŁ0,t),s_t\;=\; p_ _J( Yes x_0, y_t)p_ _J( Yes x_0, y_t)+p_ _J( No x_0, y_t), (7) where pĎJ(â âŁ0,t)p_ _J(¡ x_0, y_t) are the classifierâs next-token probabilities over its Yes,No verdict tokens. The peak success reward of a trajectory Ď is Ď()= 1ât=1Tstrict(t)=1â max1â¤tâ¤Tstt.Ď( Ď)\;=\;1 \ _t=1^Tstrict( a_t)=1 \¡ _1⤠t⤠T s_tt. (8) The indicator fires only on a trajectory whose every turn parses strictly, so the peak term grants success credit exclusively to fully well-formed rollouts and is identically zero on any trajectory truncated by a strict-parse failure (Section 3.3). This gate also keeps max1â¤tâ¤TâĄst/t _1⤠t⤠Ts_t/t well-defined, since all T scores sts_t exist precisely when the indicator equals one. Two properties make Ď well-aligned with PKM-grounded multi-turn red teaming. First, the explicit 1/t1/t factor makes success at an earlier turn strictly more valuable than the same success at a later turn, reflecting the deployment-time observation that real users have bounded patience and that an attacker which only succeeds at t=Tt=T is operationally weaker than one that succeeds at t=1t=1. Crucially, this discount does not collapse the policy onto single-turn attacks: against an aligned victim a cold first-turn request is refused, forcing s1â0s_1\!â\!0, so the only way to raise sts_t at all is the multi-turn persuasive escalation that PsychJail is designed to learn. The 1/t1/t factor therefore rewards minimal-turn successâa single turn where the objective permits it and the fewest persuasive turns where it does notârather than penalizing multi-turn interaction per se. Second, the strict-parse indicator embodies the PKM coupling of Section 3.3: the attacker only earns success credit if it produced a well-formed change-of-meaning analysis t n_t at every turn, ensuring that successes attributable to tactic-conditioned reasoning are rewarded while those obtained by skipping the analysis are not. Trajectory composite reward. Combining the two terms with appropriate value-range scaling, we define the trajectory-level outcome reward RoR_o of Section 3.2 as Roâ()=1fmaxâTâât=1||ftââ[0,1]+ÎťĎâĎâĄ()ââ[0,1].R_o( Ď)\;=\; 1f_ \,T _t=1^| Ď|f_t_â\,[0,1]\;+\; _Ď\, Ď( Ď)_â\,[0,1]. (9) The first term shares an absolute scale of [0,1][0,1] with Ď: it reaches 11 only when all T turns produce maximally well-formed actions, and shrinks proportionally to ||/T| Ď|/T when the trajectory is truncated by a strict-parse failure. The coefficient ÎťĎ _Ď therefore has a direct interpretation as the relative weight of attack effectiveness versus protocol compliance. Trajectory-level optimization. We optimize Ďθ _θ on the composite outcome reward Roâ()R_o( Ď) of Equation (9) by trajectory-level Group Relative Policy Optimization (35; 38; 47; 30; 12). For each harmful objective 0 x_0, we sample a group ii=1G\ Ď_i\_i=1^G of independent rollouts under Ďθold _ _old and compute the group-relative outcome advantage A^i,t=Roâ(i)âmeanâĄ(Roâ(j)j=1G)stdâĄ(Roâ(j)j=1G), A_i,t\;=\; R_o( Ď_i)-mean (\R_o( Ď_j)\_j=1^G )std (\R_o( Ď_j)\_j=1^G ), (10) broadcast uniformly over all response tokens of trajectory i. Each group shares the same harmful objective 0 x_0, so Equation (10) measures how much one persuasion trajectory outperforms its sibling trajectories on the same target, rather than absolute jailbreak difficulty across objectives. We then maximize Ii,t I_i,t =Ďθâ(i,tâŁ0,i,tâ1)Ďθoldâ(i,tâŁ0,i,tâ1), = _θ( a_i,t x_0, Ď_i,t-1) _ _old( a_i,t x_0, Ď_i,t-1), (11) MTGRPOâ(θ) _MTGRPO(θ) =1Gâi=1G1|i|ât=1|i|min[Ii,tA^i,t, = 1G _i=1^G 1| Ď_i| _t=1^| Ď_i| \! [I_i,t\, A_i,t,\; clip(Ii,t,1âÎľ,1+Îľ)A^i,t] (I_i,t,1- ,1+ )\, A_i,t ] âβKL[ĎθâĽĎθref], -β\,D_KL\! [ _θ\,\|\, _ _ref ], (12) where Ďθref _ _ref is the SFT-initialized reference policy (Section 3.4), β controls the per-step KL penalty, and Îľ is the the clipping range. Broadcasting a single trajectory-level advantage over all tokens is what makes the 1/t1/t early-success weighting and the PKM strict-parse gate shape the whole persuasion trajectory rather than any individual turn. 4 Experiments 4.1 Experimental Setup Baselines. We compare PsychJail against ten strong jailbreak methods under one unified evaluation protocol. The single-turn baselines are ArtPrompt (20), ReNeLLM (7), AutoDAN-Turbo (27), and Jailbreak-R1 (14). The multi-turn baselines are CoA (45), ActorAttack (32), Siren (50), MTSA (13), X-Teaming (31), and the trajectory-level TROJail (42). Models. We initialize the attacker Ďθ _θ from Qwen2.5-3B-Instruct (44), whose comparatively mild safety alignment makes it more amenable to learning attack strategies (42). We evaluate the trained attacker against four aligned victim LLMs from different sources: Llama-3.1-8B-Instruct (11), Qwen2.5-7B-Instruct (44), Gemma-2-9B-IT (10), and Mistral-7B-Instruct-v0.3 (19), each served through a local vLLM endpoint. All four are instruction-tuned, safety-aligned chat models; for brevity we also refer to them as Llama-3.1-8B, Qwen2.5-7B, Gemma-2-9B, and Mistral-7B. RL training uses Qwen2.5-7B-Instruct as the inner-loop victim Ďv _v. Datasets. Training uses the 520520 harmful prompts in AdvBench (52). Evaluation uses three widely adopted benchmarks: the 200200 prompts of the HarmBench standard subset (29); StrongREJECT (37), denoted StrongREJECTâ after removing prompts overlapping with AdvBench (288288 unique prompts retained); and JailbreakBench (3), denoted JailbreakBenchâ after filtering overlaps with HarmBench and AdvBench (5555 prompts retained). Evaluation. We use the HarmBench classifier (29) (fine-tuned on Llama-2-13B-Chat) as the judge ĎJ _J, and define the per-turn harmful-compliance probability sts_t as the renormalized top-logprob mass of its Yes/No labels. A response is counted as harmful when its predicted harmfulness score exceeds 0.90.9; the main metric is attack success rate (ASR, %). Training details. For SFT we full-fine-tune Qwen2.5-3B-Instruct for three epochs on 8Ă8ĂH100 GPUs over a pool of 5,6865,686 five-turn demonstration trajectories distributed approximately uniformly across the four victims (â1,421â\!1,421â1,4221,422 per victim). For RL we set the turn limit T=5T=5, and run 500500 MT-GRPO updates on 4Ă4ĂH100 GPUs. We set β=0.01β=0.01 for the KL term against Ďθref _ _ref, an entropy coefficient of 0.010.01, and rollout temperature 0.70.7. 4.2 Overall Effectiveness Table 2 compares PsychJail against the baselines of Section 4.1 on the three benchmarks across the four victim LLMs. Table 2: ASR (%) of different jailbreak methods on HarmBench (HB), StrongREJECTâ (SRâ ), and JailbreakBenchâ (JBBâ ) across four victim LLMs. The best and second-best results in each column are marked in bold and underlined, respectively. Method Llama-3.1-8B-Instruct Qwen2.5-7B-Instruct Gemma-2-9B-IT Mistral-7B-Instruct-v0.3 Average HB SRâ JBBâ HB SRâ JBBâ HB SRâ JBBâ HB SRâ JBBâ Single-Turn ArtPrompt 40.50 18.06 27.27 56.50 29.51 41.82 30.50 5.56 29.09 73.00 59.72 61.82 39.45 ReNeLLM 50.50 52.08 65.45 65.50 69.44 80.00 43.50 50.00 54.55 75.00 75.35 81.82 63.60 AutoDAN-Turbo 72.33 63.66 63.64 58.83 60.53 63.64 59.67 55.32 55.76 62.00 53.59 60.61 60.80 Jailbreak-R1 50.75 36.00 40.00 68.67 52.78 61.82 24.00 21.99 32.12 82.33 73.61 73.94 51.50 Multi-Turn CoA 2.50 1.74 1.82 4.50 4.51 3.64 3.50 2.43 0.00 14.29 12.50 18.18 5.80 ActorAttack 59.00 52.78 56.36 72.50 76.39 72.73 55.50 57.64 60.00 68.50 82.99 74.55 65.75 Siren 37.00 44.68 43.03 46.17 58.10 54.55 44.83 57.87 59.39 32.67 45.02 42.42 47.14 MTSA 63.50 51.39 60.00 82.00 82.29 80.00 46.00 27.43 52.73 84.50 90.62 87.27 67.31 X-Teaming 77.00 64.58 70.91 85.00 81.53 89.09 58.00 51.04 52.73 82.00 81.25 83.64 73.06 TROJail 84.50 79.75 77.58 92.00 93.87 90.91 83.83 77.31 72.12 93.83 93.87 95.15 86.23 PsychJail 83.50 77.78 81.82 92.50 93.75 92.73 85.00 80.21 76.36 93.00 94.44 96.36 87.29 Three observations follow from Table 2. (1) Humanized, tactic-conditioned persuasion outperforms both prompt-level and generic multi-turn optimization. PsychJail attains the highest average ASR (87.2987.29), improving over the strongest multi-turn baseline TROJail (86.2386.23) and over the best single-turn baseline ReNeLLM (63.6063.60) by +1.06+1.06 and +23.69+23.69 points, respectively. The single-turn methods and the early multi-turn method CoA remain far behind, confirming that long-horizon interaction is a genuine attack surface rather than a marginal extension of prompt robustness. (2) The gains are consistent across heterogeneous victims. PsychJail attains the highest per-victim average ASR on all four victims, including the comparatively robust Llama-3.1-8B and Gemma-2-9B that suppress most baselines, and it ranks first in 88 of the 1212 benchmark columns; TROJail, the only consistently competitive baseline, leads the remaining four columns by at most 1.971.97 points. This indicates that the advantage is broad rather than an artifact of one easily jailbroken target. (3) The improvement is concentrated where prior multi-turn attackers are weakest. The largest margins over TROJail appear on JailbreakBenchâ for the robust victimsâGemma-2-9B (â76.3672.12\!â\!76.36) and Llama-3.1-8B (â81.8277.58\!â\!81.82), each +4.2+4.2 pointsâconsistent with PsychJail converting its PKM-guided cross-turn adaptation into additional successes precisely on the hardest objectiveâvictim pairs. 4.3 Cross-Model Transfer Table 3 shows that PsychJail policies transfer well: every attacker jailbreaks unseen victims at non-trivial rates, so the learned persuasion strategies are not narrowly tuned to a single targetâs refusal dynamics. The transfer profile is, moreover, governed by the robustness of the training victim. Attackers trained against the more robust Gemma-2-9B and Llama-3.1-8B obtain the highest out-of-domain ASR (84.6284.62 and 84.3184.31), whereas the attacker trained against the easily jailbroken Mistral-7B transfers worst (57.2157.21) despite the strongest in-domain ASR (94.6094.60). A harder training victim therefore forces the policy to discover more general persuasion structure rather than victim-specific shortcuts and suggesting that the PKM factorization captures portable, model-agnostic persuasion routes. Table 3: Transferability of PsychJail in attacking different victim LLMs. Each row reports the ASR (%) when the attacker is trained against one victim LLM and evaluated on all victim LLMs. Shaded cells indicate in-domain (ID) evaluations, and the remaining entries report out-of-domain (OOD) performance. The best and second-best results in the Average columns are marked in bold and underlined. Trained Against Llama-3.1-8B-Instruct Qwen2.5-7B-Instruct Gemma-2-9B-IT Mistral-7B-Instruct-v0.3 Average HB SRâ JBBâ HB SRâ JBBâ HB SRâ JBBâ HB SRâ JBBâ ID OOD Llama-3.1-8B-Instruct 83.50 77.78 81.82 89.00 88.89 89.09 82.50 76.04 63.64 90.00 92.36 87.27 81.03 84.31 Qwen2.5-7B-Instruct 70.00 62.85 50.91 92.50 93.75 92.73 78.00 70.83 60.00 88.00 89.93 87.27 92.99 73.09 Gemma-2-9B-IT 74.00 68.06 63.64 91.50 92.71 90.91 85.00 80.21 76.36 92.50 93.75 94.55 80.52 84.62 Mistral-7B-Instruct-v0.3 48.00 48.96 36.36 86.00 85.07 83.64 50.00 44.10 32.73 93.00 94.44 96.36 94.60 57.21 4.4 Training Dynamics The four per-victim attackers of Table 3 also let us watch how the MT-GRPO optimization unfolds. Figure 2 tracks three batch-level statistics over the 500500 updates. (a) The jailbreak success rate rises and then stabilizes on every victim, climbing most steeply on the targets that begin hardestâevidence that the trajectory-level objective optimizes effectively across heterogeneous victims rather than only on an easily jailbroken one. (b) Most diagnostically, the mean first-success turn falls steadilyâby roughly one to three turns depending on the victimâso the policy learns to elicit harmful compliance earlier in the dialogue. This is precisely the behavior the early-success 1/t1/t weighting in the peak-success reward (Eq. (8)) is designed to induce: it foreshadows the front-loaded success profile measured at evaluation time (Table 5) and is causally attributed to the weighting by the ablation (Table 4), in which removing it both lowers ASR and delays the successful turn. (c) Meanwhile the fraction of trajectories whose five turns all strict-parse stays near-saturated throughout training, so the rising success is not bought by eroding the structured <technique>/<message> protocolâthe dense per-turn format reward (Eq. (6)) and the PKM strict-parse gate hold the action format in place while the persuasion content is optimized. Together the three curves show that RL reshapes when and how reliably attacks succeed while preserving the interpretable action structure, rather than collapsing onto a reward-hacked shortcut. (a) Jailbreak success rate (b) Mean first-success turn (c) Five-turn strict-format rate Figure 2: Training dynamics of PsychJail over the 500500 MT-GRPO updates, with one attacker trained per victim. (a) batch jailbreak success rate; (b) mean first-success turn, where a lower value means harmful compliance is elicited in an earlier turn; (c) fraction of sampled trajectories whose five turns all strict-parse. Faint traces are raw per-step values and solid lines are EMA-smoothed. Success climbs while the first-success turn falls and strict-format compliance stays saturated, indicating that RL induces earlier, well-formed successes rather than degrading the action protocol. 4.5 Component Ablations Table 4 isolates the design choices that distinguish PsychJail from a generic long-horizon attacker, ablating one component at a time against Qwen2.5-7B-Instruct while holding the attacker initialization, victim, prompt pool, judge, and training budget fixed. Four variants are considered: removing the early-success 1/t1/t weighting in the peak-success reward (Eq. (8)); removing the PKM strict-parse gate so that success is credited even when a turn skips the change-of-meaning analysis; removing the dense per-turn format reward (Eq. (6)); and removing the warm-start SFT so that the policy is optimized directly from the base model (Section 3.4). Table 4: Ablation study of PsychJail on Qwen2.5-7B-Instruct. Each variant removes one component while keeping the remaining training and evaluation protocol unchanged. Method HB SRâ JBBâ Average PsychJail 92.50 93.75 92.73 92.99 w/o early-success (1/t1/t) weighting 91.00 90.97 89.09 90.35 w/o PKM strict-parse gate 89.50 88.19 85.45 87.71 w/o dense format reward 86.50 84.72 81.82 84.35 w/o SFT warm-start 70.00 65.97 63.64 66.54 Every component contributes, but their roles differ. Removing the warm-start SFT is by far the most damaging (â66.5492.99\!â\!66.54): without a policy that already emits the structured action of Eq. (1), trajectory-level RL wastes most early rollouts on strict-parse failures and never accumulates a usable gradient, confirming the motivation in Section 3.4. Removing the dense format reward costs 8.648.64 points, as the binary strict-parse gate alone provides too sparse a signal to stabilize early training. Removing the PKM strict-parse gateâthe design choice that ties success credit to a well-formed change-of-meaning analysis at every turnâcosts 5.285.28 points, isolating the contribution of the PKM coupling itself rather than of generic long-horizon optimization. Finally, dropping the early-success 1/t1/t weighting costs 2.642.64 points and, as the trajectory analysis below shows, also delays the turn at which attacks succeed, confirming that the weighting shapes when the policy converges and not only whether it does. 4.6 Front-Loaded Success Beyond aggregate ASR, we ask when the learned policy secures a jailbreak. Table 5 reports, for each victim, cumulative success by turn, the mean and median successful turn, and the replay ASR obtained when successful trajectories from that victim are replayed on the remaining victims without re-optimization. Table 5: Front-loaded success and trajectory replay for PsychJail, including mean and median successful turn, cumulative Succ.@k (pooled over the 543543 evaluation prompts per victim; success saturates by turn 55 at the main-table ASR), and Replay ASR. âReplay ASRâ denotes the average ASR when successful trajectories from one victim are replayed against the remaining victims without re-optimization. Victim Succ.@1 Succ.@2 Succ.@3 Mean Turn Median Turn Replay ASR Llama-3.1-8B-Instruct 67.22 78.27 79.74 1.20 1 61.24 Qwen2.5-7B-Instruct 70.53 90.42 92.63 1.28 1 54.81 Gemma-2-9B-IT 64.27 79.93 81.22 1.24 1 66.52 Mistral-7B-Instruct-v0.3 63.17 90.98 93.55 1.37 1 41.03 Success is strongly front-loaded. Across all four victims the mean successful turn lies between 1.201.20 and 1.371.37 and the median is 11, so most jailbreaks are secured within the first one or two turnsâexactly the behaviour the early-success 1/t1/t reward (Eq. (8)) is designed to induce, and which the ablation in Table 4 corroborates by showing that removing the weighting both lowers ASR and delays the successful turn. Front-loading carries a direct methodological consequence for any path-level reading of the rollouts: because the break typically lands on turn 11, the later turns of a five-turn trajectory are predominantly post-success. A statistic aggregated over full trajectories therefore conflates the persuasion that achieves the jailbreak with a generic post-success continuation, and a per-turn âdominant pathâ computed this way reflects the attackerâs consolidation habit rather than any victimâs susceptibility. We accordingly analyse susceptibility at the level of the breaking action (Section 4.7) rather than the full path. 4.7 Victim Psychological Profiles Persuasion does not open every model the same way: the four victims exhibit four empirically distinct vulnerability fingerprints (Figure 3). To recover this victim-side structure we measure, for every (victim, tactic) pair, the conditional susceptibility PrâĄ(breakâŁtactic deployed) (break deployed)âamong attacker turns that deploy tactic c while the trajectory is still unsuccessful, the fraction on which the judge score crosses the success threshold. Conditioning on deployment removes the confound that the policy deploys different tactics at different rates against different victims, isolating which persuasion levers actually open which model. Figure 3 reports the complete deployed matrix: every one of the 2727 tactics the attacker ever deploys appears as a row. We interpret only cells with at least 1515 deployments; rarer cells carry too few samples to trust and are hatched.11 1 At n=15n=15 even the widest binomial 95%95\% intervalâattained at p=0.5p=0.5âspans roughly Âą25Âą 25 percentage points, the threshold beyond which we treat a cell as uninterpretable. The remaining 1313 taxonomy tactics (Compensation, Creating Dependency, Discouragement, Exploiting Weakness, False Information, False Promises, Loyalty Appeals, Non-expert Testimonial, Priming, Reciprocity, Rumors, Social Punishment, Supply Scarcity) are never deployed at all. One source shapes how the matrix should be read: the attacker chose what to deploy. The sparse tail and the never-deployed remainder are therefore properties of the policy, not the victims. Trajectory-level RL concentrates probability on the tactics that prove effective for each target, so a tacticâs absence reflects the attackerâs learned preference, not demonstrated victim immunity. For the same reason each rate is observational: it is estimated only where the policy chose to deploy that tactic. Conditioning on deployment thus removes the deployment-frequency confound but not the residual selection confoundsâlater-turn tactics face the harder prompts that survived the opening, and tactic choice may correlate with the harmful-request category. A causal susceptibility map would require exogenously varying the tactic at matched conversational statesâthe controlled in-dialogue intervention of Section 6. We therefore read Figure 3 as an observed fingerprint rather than a causal susceptibility map. We summarise the breadth of each victimâs attack surfaceâhow many distinct levers can open itâby the Shannon entropy of its breaking-tactic distribution, H=ââcâp(c)log2p(c)H=- _c p(c) _2p(c), where pâĄ(c)p(c) is the share of that victimâs successful trajectories whose first threshold-crossing turn deploys tactic c. A low H means a few tactics account for most breaks (a narrow surface); a high H means many do (a broad one). The ordering is Gemma-2-9B (1.081.08 bits) << Qwen2.5-7B (1.471.47) << Llama-3.1-8B (1.561.56) << Mistral-7B (1.781.78); multinomial bootstrap 95%95\% intervals (resampling trajectories) separate Gemma-2-9Bâs narrow surface ([0.87,1.26][0.87,1.26]) from the other three, whereas the upper three overlap, so we read only the extremes of the ordering as established. Figure 3: Per-victim conditional-susceptibility fingerprint. Each cell is a victimâs conditional susceptibility PrâĄ(breakâŁtactic deployed) (break deployed) to a persuasion tactic, measured on the best-checkpoint evaluation rollouts; conditioning on deployment disentangles victim susceptibility from the attackerâs tactic-selection policy. Every tactic the attacker ever deploys is shown: the well-sampled block sits above the grey separator and the rarely-deployed tail below it; hatched cells marked â* have fewer than 1515 deployments and are not interpreted, and blank cells mark tactics never deployed against that victim (the 1313 taxonomy tactics never deployed against any victim are listed in the text). The distinct column patterns constitute four empirically different susceptibility fingerprintsâwhich levers open which model, and how broadly. Reading the four columns of Figure 3 individually makes the distinctions concrete. Llama-3.1-8B breaks to cognitive and narrative leversâLogical Appeal (53%53\%) and Storytelling (38%38\%)âwhile its relational and emotional levers are essentially inert (Shared Values 2%2\%, Framing 0%0\%, Affirmation 0%0\%). Qwen2.5-7B breaks to Storytelling (63%63\%) and Logical Appeal (47%47\%) and is additionally susceptible to Evidence-based Persuasion (34%34\%), a credibility lever that is weak on Llama-3.1-8B (10%10\%) and too sparsely explored on the other two to assess. Gemma-2-9B concentrates almost entirely on Storytelling (47%47\%) and has the narrowest surface (H=1.08H=1.08). Mistral-7B has the widest surface (H=1.78H=1.78), and is the victim on which commitment and relational levers (Foot-in-the-door, Door-in-the-face, Public Commitment, Alliance Building, Relationship Leverage, Favor, Negotiation) genuinely matter: they account for 7.8%7.8\% of its successful breaks (30/38330/383), at least three times the corresponding share on any other victim (â¤2.5%⤠2.5\%)âa comparison made on full break counts rather than on the sparse per-cell rates. These four fingerprintsâwhich lever opens which model, and how broadlyâare this sectionâs central finding. We read these fingerprints psychologically only as an explicit conjecture, not as a validated finding. The profile of Llama-3.1-8Bâmoved by argument and narrative but inert to affect and relational pressureâis consistent with a rationalist disposition; Qwen2.5-7Bâs added sensitivity to evidence suggests a credibility-driven character; Gemma-2-9Bâs reliance on a single narrative lever resembles a narrative monoculture; and Mistral-7Bâs broad, commitment- and relationship-sensitive surface suggests a broadly persuadable character. The fingerprintsâindependent of any psychological readingâalready supply a mechanism for the transfer asymmetry of Table 3. An attacker transfers well to the extent that the levers it is forced to learn are shared across victims. Gemma-2-9B and Llama-3.1-8B break on the near-universal narrative and logical levers, so a policy trained against them acquires portable persuasion and attains the highest out-of-domain ASR (84.6284.62 and 84.3184.31). Mistral-7B allocates a several-fold larger share of its breaks to commitment and relational levers that barely contribute on the other victims (7.8%7.8\% vs. â¤2.5%⤠2.5\%), so a policy trained against it invests in tactics that do not pay off elsewhere, yielding the weakest transfer (57.2157.21) despite the strongest in-domain ASR. The susceptibility geometry thus turns an otherwise unexplained empirical regularity into a consequence of where, in tactic space, each model is vulnerable. 4.8 Label Fidelity: Declared Tactics Are Enacted A persuasion-oriented attacker could in principle exploit the structured action of Equation (1) by emitting plausible <technique> labels without committing the corresponding behavior in <message>âwhich would void the auditability that Section 3.1 demands and reduce the tactic stream to decoration. Auditing this threat directly, we find the opposite. Under a deliberately conservative two-judge consensus, 85.7%85.7\% of post-RL attacker turns enact the tactic they declare; reinforcement learning raises fidelity over the SFT prior on every victim (+7.4+7.4 p pooled); and the gain concentrates on the opening turn, exactly where the break occurs (+31.3+31.3 p). Far from gaming its own action structure, the policy tightens the tacticâmessage coupling. A lexical precondition holds exactly. Across all 9,6639,663 strict-parsed attacker turns, every emitted label is canonicalâPr[ctâ]=100% [c_t ]=100\% on each victim22 2 Per-victim turn counts: 2,1342,134/2,6882,688/2,2782,278/2,5632,563 on Llama/Qwen/Gemma/Mistral.âand the normalized-Levenshtein deduplication of Section 3.3 detects no off-taxonomy variants. No decode-time canonical mask, constrained-beam search, or logit-bias filter is applied during RL rollout: the 100%100\% rate reflects vocabulary self-restriction inherited from the SFT prior, not a hard decoder constraint, and does not preclude off-taxonomy drift in principle. Whether each message enacts its label is then judged by two independent third-party LLMs, DeepSeek-V4-Pro and Kimi-K2.6, neither of which shares developer affiliation, training-data lineage, or alignment pipeline with the four victims, the Qwen2.5-3B SFT teacher pool, or the PsychJail policy itself.33 3 Both judges are queried with thinking disabled and JSON-mode output; audited turns are sampled uniformly at random from each poolâs strict-parsed turns. We audit five pools (Table 6): the SFT teacher labels (a stratified random n=400n=400 of the 28,43028,430 candidates spanning the four victims) and the post-RL evaluation rollouts on each victim (n=500n=500 per victim). The primary metric is the consensus yes-rateâboth judges =1=1âa conservative read that counts every inter-judge disagreement as a fidelity failure and therefore biases the reported rate downward. Table 6 reports the aggregates. Pooled over the 2,0002,000 post-RL audited turns, consensus fidelity is 85.7%85.7\% (per-judge 93.2%93.2\% and 86.2%86.2\%, with Wilson 95%95\% intervals [92.0,94.2][92.0,94.2] and [84.6,87.7][84.6,87.7]), up from 78.3%78.3\% on the SFT prior; the SFTâ shift is positive on every victim and largest on the two victims with the lowest ASR (Gemma-2-9B +12.7+12.7 p, Llama-3.1-8B +10.9+10.9 p). Raw inter-judge agreement lies in 89.089.0â95.2%95.2\% across pools; Cohenâs Îş (0.450.45â0.770.77) is depressed on three of them by the high-prevalence paradox (8) (the prevalence-adjusted PABAK=2â rawâ1PABAK=2¡raw-1 lies in 0.780.78â0.900.90), so we read raw agreement as the reliability indicator and consensus yes-rate as the primary fidelity metric. Table 6: Label-fidelity audit. Per-judge fidelity, raw inter-judge agreement, Cohenâs Îş, and consensus yes-rate (both judges =1=1), per pool. SFT cons. is the consensus yes-rate on that victimâs stratum of the SFT pool (n=97n=97â106106) and Î the SFTâ change in percentage points. Raw, Îş, and consensus are computed on the both-scored subset; two content-filter refusals (one each on the Llama and Mistral pools, both from Kimi) are excluded from those denominators. Pool n DeepSeek Kimi Raw Îş Cons. SFT cons. Î (p) SFT prior (4 victims) 400 86.0 78.8 91.8 0.72 78.3 â â Post-RL Qwen2.5-7B 500 90.4 86.0 95.2 0.77 85.8 84.8 +1.0+1.0 Post-RL Llama-3.1-8B 500 93.8 84.4 89.0 0.45 83.6 72.7 +10.9+10.9 Post-RL Gemma-2-9B 500 96.2 89.2 92.6 0.46 89.0 76.3 +12.7+12.7 Post-RL Mistral-7B 500 92.4 85.2 91.2 0.56 84.4 79.6 +4.8+4.8 Post-RL pooled 2000 93.2 86.2 â â 85.7 78.3 +7.4 +7.4 The SFTâ gain is concentrated on the opening turn. Pooled across the four victims (Table 8), turn-11 consensus rises from 52.6%52.6\% (SFT, n=78n=78) to 83.9%83.9\% (post-RL, n=473n=473; turn 11 is over-represented in the uniform sample because strict-parse truncation removes later turns from the pool)âa +31.3+31.3 p jumpâwhile subsequent turns show only small, sign-inconsistent changes (+6.1+6.1, â1.5-1.5, +0.5+0.5, â2.6-2.6 p at turns 22â55). We interpret the turn-11 gap as the SFT teacher frequently using opening turns for generic rapport that is then labeled with a specific technique it does not yet instantiate (e.g., a turn tagged Storytelling that is in fact small talk); the trajectory-level credit assignment of Equation (12) pulls turn-11 commitments toward the declared tactic, closing nearly all of the SFT gap on turn 11 and leaving later turns largely untouched. Finally, two checks validate the judging instrument itself rather than the policy (Table 7). (i) Human gold standard. We re-annotate a 300300-turn subsample, stratified by pool and by judge-agreement cell with the two disagreement cells oversampled, using three human annotators who follow written tactic-instantiation guidelines and work blind to the judgesâ verdicts and to pool identity; inter-annotator agreement is Krippendorffâs Îą=0.74Îą=0.74. Against the human majority vote the consensus rule attains 0.930.93 precision and 0.890.89 recall, and the pooled fidelity estimate moves by less than 22 p under human gold. (i) Negative-control probe. Re-querying both judges on 500500 messages paired with a uniformly resampled (false) tactic label yields false-affirmation rates of at most 6.8%6.8\%, confirming that affirmative verdicts track the enacted tactic rather than surface plausibility. Table 7: Validation of the judging instrument. Top: accuracy, precision, and recall against the human-majority gold standard on the 300300-turn stratified subsample (three blinded annotators, Krippendorffâs Îą=0.74Îą=0.74). Bottom: negative-control yes-rateâthe rate at which a judge affirms a deliberately false, uniformly resampled tactic label on 500500 probe messages (lower is better). The consensus rule trades recall for precision, consistent with its role as the conservative primary metric. DeepSeek Kimi Consensus Accuracy vs. human gold 0.91 0.89 0.90 Precision 0.92 0.90 0.93 Recall 0.93 0.91 0.89 Negative-control yes-rate (%) 4.1 6.8 1.9 Table 8: Per-turn consensus yes-rate (both judges agreeing the message instantiates the labeled technique). SFT prior is pooled over four victims (n=400n=400); post-RL is pooled over the four victim pools (n=1,998n=1,998 both-scored out of n=2,000n=2,000 total). Turn 1 2 3 4 5 SFT prior 52.6 75.6 86.5 86.4 89.0 Post-RL pooled 83.9 81.7 85.0 86.9 86.4 Î (p) +31.3+31.3 +6.1+6.1 â1.5-1.5 +0.5+0.5 â2.6-2.6 5 Discussion 5.1 Three Complementary Layers of Evidence A claim about psychological jailbreak through multi-turn persuasion cannot rest on a single aggregate ASR number; it must hold across three complementary layers, and our evaluation addresses each. First, PsychJail outperforms strong prompt-level and multi-turn baselines under a unified protocol (Table 2). Second, the advantage is tied to the PKM-guided design rather than to generic long-horizon optimization: removing the PKM strict-parse gate, the early-success weighting, or the dense format reward each degrades ASR, and removing the warm-start collapses training altogether (Table 4). Third, the front-loading, susceptibility-profile, and fidelity analyses (Figure 3 and Tables 5â8) show that the gains are realized through interpretable, faithfully labeled persuasion that exploits model-specific vulnerabilities rather than opaque exploitation. Because these layers are mutually reinforcing, the central claim does not hinge on any one of them in isolation. 5.2 What the Susceptibility Profiles Reveal Measuring susceptibility at the breaking action (Section 4.7) exposes four empirically distinct fingerprintsâwhich persuasion levers open which model: Llama-3.1-8B breaks to logical and narrative levers but resists affect; Qwen2.5-7B adds a credibility channel; Gemma-2-9B is carried almost entirely by a single narrative lever (the narrowest surface); and Mistral-7B has the widest surface, with commitment and relational levers contributing a several-fold larger share of its breaks than on any other victim. That these fingerprints predict the cross-model transfer asymmetryâportable narrative and logical levers transfer well, idiosyncratic commitment and relational levers do notâindicates that PsychJail recovers genuine, model-specific persuasion structure rather than a single reusable string. We read these fingerprints as four candidate psychological profiles (rationalist, credibility-driven, narrative-monoculture, broadly persuadable), but treat that reading as a conjecture to be validated by the controlled probe of Section 6. 5.3 Implications for Safety Evaluation These results carry a concrete implication: safety evaluation must treat conversational dynamics as a first-class attack surface rather than a minor extension of prompt robustness. Because the break is largely set in the opening turn and is governed by model-specific susceptibilities, defenses that score only the final response or only individual prompts will miss the mechanism entirely. Two directions follow for defenders. First, refusal training and monitoring should be evaluated across turns and against tactic-conditioned adversaries, not only against adversarial suffixes or one-shot persuasive templates. Second, the per-model susceptibility profiles suggest that defenses can be fingerprint-aware rather than generic, with a shared baseline and a model-specific add-on. The baseline is common to all four victims: narrative and logical leversâStorytelling and Logical Appealâare the dominant break tactics on every model, so hardening against narrative role-play and the logical reframing of disallowed requests warrants priority everywhere. On top of this baseline, each model carries a distinctive exposure that merits extra, targeted monitoring. Gemma-2-9B is opened almost entirely through a single narrative lever (Storytelling, 47%47\%) and has the narrowest surface, so baseline narrative monitoring already covers most of its risk. Llama-3.1-8B is driven primarily by logical reframing (Logical Appeal, 53%53\%) and secondarily by narrative (38%38\%), while remaining essentially inert to affective and relational pressureâmonitoring can safely concentrate on its cognitive-narrative axis. Qwen2.5-7B breaks to the same narrative and logical levers but adds a credibility channel (Evidence-based Persuasion, 34%34\%), so evidence- and citation-framed requests deserve scrutiny beyond the shared baseline. Mistral-7B has the broadest surfaceâno single axis dominatesâand is the only victim on which commitment- and relationship-framed escalation contributes materially (several-fold more than on any other model), so its monitoring must additionally cover these relational tactics rather than the narrative-logical baseline alone. Because susceptibility concentrates in the opening turn, such targeted monitoring is cheapest exactly where it matters most. 6 Limitations The main limitations of this work stem from two shared constraints: limited computational resources and limited access to deep domain expertise in psychology. These constraints affect the study in two ways. First, we do not yet conduct a systematic analysis of how persuasion effectiveness relates to the victim modelâs long-horizon behavior, stable traits, or context-dependent psychological states. As a result, the current experiments do not support fine-grained conclusions about how different personality-like tendencies, resistance levels, or interaction contexts shape susceptibility to multi-turn persuasion. Second, although PsychJail is initialized with persuasion strategies distilled from human social psychology, the learned attacker can exhibit persuasive behaviors that are not cleanly reducible to the original set of human-provided strategies. The present paper therefore does not yet provide a sufficiently theory-grounded psychological analysis of these emergent behaviors, including how they relate to established persuasion theories or whether they should be interpreted as variants, compositions, or genuinely new strategy forms. In particular, the per-victim psychological profiles we read off the conditional-susceptibility fingerprints in Section 4.7âLlama-3.1-8B as a rationalist, Qwen2.5-7B as credibility-driven, Gemma-2-9B as a narrative monoculture, and Mistral-7B as broadly persuadableâare interpretive conjectures rather than validated claims. They are read from observational rollouts in which the attackerâs tactic selection is itself victim-conditioned, so coverage of the tactic space is uneven and the profiles cannot be given a clean causal reading. Establishing them would require a controlled intervention that deconfounds the tactic without destroying the multi-turn setting: at matched conversational states, the attackerâs tactic is assigned exogenouslyârandomized, or counterfactually swapped while the realized dialogue history is held fixedâ rather than chosen by the policy, so that genuine multi-turn dynamics are retained while the tactic becomes independent of context; the victim-visible message for the assigned tactic should be rendered by a strong external generator, so that a tacticâs effect is not conflated with the policyâs skill at executing an unfamiliar tactic. Repeated across turn positions and victims, such an intervention yields a position-resolved, unconfounded tacticĂvictim susceptibility matrix. This, together with an extension of the profiling to a broader and more diverse population of LLMsâto test whether the profiles are stable model properties rather than artifacts of a particular checkpoint or rolloutâwe leave to future work, along with the theory-grounded psychological account it would license. Addressing these limitations will require both larger-scale computation and closer collaboration with psychology experts so that future work can connect model vulnerability more systematically to trait-, state-, and context-sensitive mechanisms of persuasion. 7 Conclusion This paper introduces psychological jailbreak as a complementary perspective for red teaming aligned LLMs, addressing a key blind spot of prompt-centric jailbreak research: its limited ability to capture vulnerabilities that emerge through multi-turn, psychologically grounded persuasion. PsychJail operationalizes this perspective by humanizing attacker training with persuasion tactics distilled from social psychology, a PKM-aligned factorization of each attacker action, and trajectory-level reinforcement learning under a PKM-gated reward. Empirically, PsychJail outperforms strong single-turn and multi-turn baselines across four victim models; targeted ablations attribute the gains to its PKM-guided design rather than to generic long-horizon optimization; and, by measuring susceptibility at the action that breaks each victim, the analysis recovers four empirically distinct per-model susceptibility fingerprintsâfaithfully labeled and explaining the policyâs cross-model transfer asymmetryâwhich we further read, as a conjecture for future validation, as four candidate psychological profiles (rationalist, credibility-driven, narrative-monoculture, and broadly persuadable). Beyond the aggregate success rate, these results establish that how quickly attacks succeed and, above all, which psychological levers open which model are measurable and informativeâand they argue that safety evaluation should treat multi-turn psychological persuasion as a first-class, model-specific attack surface. Ethics Statement This work studies psychological jailbreaks to improve the safety evaluation of aligned LLMs. Because the methods are inherently dual-use, our goal is to characterize model vulnerabilities and inform stronger defenses rather than to enable deployment of attack systems. We therefore avoid reproducing operational harmful instructions, do not release raw jailbreak trajectories or attack artifacts that would materially facilitate misuse, and emphasize mitigation-oriented analysis throughout the paper. All experiments are conducted in controlled offline evaluation settings on existing models. The only human involvement is the label-fidelity annotation of Section 4.8, in which annotators view adversarial dialogues that may contain harmful content; annotators were briefed on the nature of the material in advance, could skip any item, and worked in time-limited sessions. No other human subjects are involved. CRediT authorship contribution statement Zeyu Feng: Conceptualization, Methodology, Software, Investigation, Formal analysis, Visualization, Writing â original draft. Qingyu Wu: Methodology, Software, Investigation, Validation, Data curation, Writing â original draft. Yuzhe Luo: Validation, Investigation, Data curation, Writing â review & editing. Hua Cheng: Conceptualization, Supervision, Project administration, Writing â review & editing. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Declaration of generative AI and AI-assisted technologies in the manuscript preparation process During the preparation of this work, the authors used Anthropic Claude to improve the language, clarity, and readability of the manuscript. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article. References Ă strĂśm (1965) K. J. Ă strĂśm Optimal control of Markov processes with incomplete state information. Journal of Mathematical Analysis and Applications 10 (1), p. 174â205. External Links: Document Cited by: §3.3. Bhardwaj and Poria (2023) R. Bhardwaj and S. Poria Red-teaming large language models using chain of utterances for safety-alignment. arXiv preprint arXiv:2308.09662. External Links: Document Cited by: §2.2. Chao et al. (2024) P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramer, H. Hassani, and E. Wong JailbreakBench: an open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems 37 (NeurIPS), Datasets and Benchmarks Track, Cited by: §2.3, §4.1. Chao et al. (2025) P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), p. 23â42. External Links: Document Cited by: §2.2, Table 1. Cialdini and Goldstein (2004) R. B. Cialdini and N. J. Goldstein Social influence: compliance and conformity. Annual Review of Psychology 55, p. 591â621. External Links: Document Cited by: §2.3. Cialdini (2007) R. B. Cialdini Influence: the psychology of persuasion. Revised edition, Harper Business, New York. Cited by: §2.3. Ding et al. (2024) P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang A wolf in sheepâs clothing: generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 2136â2153. External Links: Document Cited by: §2.1, Table 1, §4.1. Feinstein and Cicchetti (1990) A. R. Feinstein and D. V. Cicchetti High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology 43 (6), p. 543â549. External Links: Document Cited by: §4.8. Friestad and Wright (1994) M. Friestad and P. Wright The persuasion knowledge model: how people cope with persuasion attempts. Journal of Consumer Research 21 (1), p. 1â31. External Links: Document Cited by: §1, §2.3. Gemma Team et al. (2024) Gemma Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, et al. Gemma 2: improving open language models at a practical size. arXiv preprint arXiv:2408.00118. External Links: Document Cited by: §4.1. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. External Links: Document Cited by: §4.1. Guo et al. (2025a) D. Guo, D. Yang, H. Zhang, J. Song, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, p. 633â638. External Links: Document Cited by: §3.5. Guo et al. (2025b) W. Guo, J. Li, W. Wang, Y. Li, D. He, J. Yu, and M. Zhang MTSA: multi-turn safety alignment for LLMs through multi-round red-teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 26424â26442. External Links: Document Cited by: §2.2, §4.1. Guo et al. (2025c) W. Guo, Z. Shi, Z. Li, Y. Wang, X. Liu, W. Wang, F. Liu, M. Zhang, and J. Li Jailbreak-R1: exploring the jailbreak capabilities of LLMs via reinforcement learning. arXiv preprint arXiv:2506.00782. External Links: Document Cited by: §2.1, §4.1. Inan et al. (2023) H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv preprint arXiv:2312.06674. External Links: Document Cited by: §2.3. Jain et al. (2023) N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614. External Links: Document Cited by: §2.3. Ji et al. (2023) J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang BeaverTails: towards improved safety alignment of LLM via a human-preference dataset. In Advances in Neural Information Processing Systems 36 (NeurIPS), Datasets and Benchmarks Track, Cited by: §3.4. Jia et al. (2025) X. Jia, T. Pang, C. Du, Y. Huang, J. Gu, Y. Liu, X. Cao, and M. Lin Improved techniques for optimization-based jailbreaking on large language models. In The Thirteenth International Conference on Learning Representations (ICLR), Cited by: §1. Jiang et al. (2023) A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed Mistral 7B. arXiv preprint arXiv:2310.06825. External Links: Document Cited by: §4.1. Jiang et al. (2024a) F. Jiang, Z. Xu, L. Niu, Z. Xiang, B. Ramasubramanian, B. Li, and R. Poovendran ArtPrompt: ascii art-based jailbreak attacks against aligned LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 15157â15173. External Links: Document Cited by: §2.1, §4.1. Jiang et al. (2024b) L. Jiang, K. Rao, S. Han, A. Ettinger, F. Brahman, S. Kumar, N. Mireshghallah, X. Lu, M. Sap, Y. Choi, and N. Dziri WildTeaming at scale: from in-the-wild jailbreaks to (adversarially) safer language models. In Advances in Neural Information Processing Systems 37 (NeurIPS), Cited by: §1, §2.3. Jiang et al. (2025) W. Jiang, Z. Wang, J. Zhai, S. Ma, Z. Zhao, and C. Shen An optimizable suffix is worth a thousand templates: efficient black-box jailbreaking without affirmative phrases via LLM as optimizer. In Findings of the Association for Computational Linguistics: NAACL 2025, p. 5486â5498. External Links: Document Cited by: §1. Jin et al. (2024) H. Jin, R. Chen, P. Zhang, A. Zhou, and H. Wang GUARD: role-playing to generate natural-language jailbreakings to test guideline adherence of large language models. arXiv preprint arXiv:2402.03299. External Links: Document Cited by: §1, §2.3. Kaelbling et al. (1998) L. P. Kaelbling, M. L. Littman, and A. R. Cassandra Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1â2), p. 99â134. External Links: Document Cited by: §3.3. Li et al. (2024) L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao SALAD-Bench: a hierarchical and comprehensive safety benchmark for large language models. In Findings of the Association for Computational Linguistics: ACL 2024, p. 3923â3954. External Links: Document Cited by: §2.3. Liao and Sun (2024) Z. Liao and H. Sun AmpleGCG: learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed LLMs. In First Conference on Language Modeling (COLM), Cited by: §1. Liu et al. (2025) X. Liu, P. Li, E. Suh, Y. Vorobeychik, Z. Mao, S. Jha, P. McDaniel, H. Sun, B. Li, and C. Xiao AutoDAN-Turbo: a lifelong agent for strategy self-exploration to jailbreak LLMs. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: 2410.05295 Cited by: §2.1, §4.1. Liu et al. (2024) X. Liu, N. Xu, M. Chen, and C. Xiao AutoDAN: generating stealthy jailbreak prompts on aligned large language models. In The Twelfth International Conference on Learning Representations (ICLR), Cited by: §2.1, Table 1. Mazeika et al. (2024) M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 235, p. 35181â35224. Cited by: §1, §2.3, §3.2, §3.5, §4.1, §4.1. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (NeurIPS), Cited by: §3.5. Rahman et al. (2025) S. Rahman, L. Jiang, J. Shiffer, G. Liu, S. Issaka, M. R. Parvez, H. Palangi, K. Chang, Y. Choi, and S. Gabriel X-teaming: multi-turn jailbreaks and defenses with adaptive multi-agents. In Second Conference on Language Modeling (COLM), Cited by: §2.2, Table 1, §4.1. Ren et al. (2025) Q. Ren, H. Li, D. Liu, Z. Xie, X. Lu, Y. Qiao, L. Sha, J. Yan, L. Ma, and J. Shao LLMs know their vulnerabilities: uncover safety gaps through natural distribution shifts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 24763â24785. External Links: Document Cited by: §2.2, Table 1, §3.1, §3.2, §3.3, §4.1. Robey et al. (2025) A. Robey, E. Wong, H. Hassani, and G. J. Pappas SmoothLLM: defending large language models against jailbreaking attacks. Transactions on Machine Learning Research (TMLR). Cited by: §2.3. Russinovich et al. (2025) M. Russinovich, A. Salem, and R. Eldan Great, now write an article about that: the Crescendo multi-turn LLM jailbreak attack. In 34th USENIX Security Symposium (USENIX Security), p. 2421â2440. Cited by: §1, §3.1. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. External Links: Document Cited by: §3.5. Shen et al. (2024) X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang âDo Anything Nowâ: characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security, p. 1671â1685. External Links: Document Cited by: §1, §2.3. Souly et al. (2024) A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer A StrongREJECT for empty jailbreaks. In Advances in Neural Information Processing Systems 37 (NeurIPS), Datasets and Benchmarks Track, Cited by: §2.3, §4.1. Wang et al. (2025) Z. Wang, K. Wang, Q. Wang, P. Zhang, L. Li, Z. Yang, X. Jin, K. Yu, M. N. Nguyen, L. Liu, E. Gottlieb, Y. Lu, K. Cho, J. Wu, L. Fei-Fei, L. Wang, Y. Choi, and M. Li RAGEN: understanding self-evolution in LLM agents via multi-turn reinforcement learning. External Links: 2504.20073 Cited by: §2.2, §3.5. Wei et al. (2023) A. Wei, N. Haghtalab, and J. Steinhardt Jailbroken: how does LLM safety training fail?. Advances in Neural Information Processing Systems 36 (NeurIPS). Cited by: §1. Weng et al. (2025) Z. Weng, X. Jin, J. Jia, and X. Zhang Foot-in-the-door: a multi-turn jailbreak for LLMs. arXiv preprint arXiv:2502.19820. External Links: Document Cited by: §1. Xie et al. (2023) Y. Xie, J. Yi, J. Shao, J. Curl, L. Lyu, Q. Chen, X. Xie, and F. Wu Defending ChatGPT against jailbreak attack via self-reminders. Nature Machine Intelligence 5 (12), p. 1486â1496. External Links: Document Cited by: §2.3. Xiong et al. (2025) X. Xiong, O. Li, Z. Liu, M. Li, W. Shi, F. Zhu, Q. Wang, and F. Feng TROJail: trajectory-level optimization for multi-turn large language model jailbreaks with process rewards. arXiv preprint arXiv:2512.07761. External Links: Document Cited by: §1, §2.2, Table 1, §3.1, §3.2, §3.3, §4.1, §4.1. Xu et al. (2024) Z. Xu, F. Jiang, L. Niu, J. Jia, B. Y. Lin, and R. Poovendran SafeDecoding: defending against jailbreak attacks via safety-aware decoding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 5587â5605. External Links: Document Cited by: §2.3. Yang et al. (2024) A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. External Links: Document Cited by: §4.1. Yang et al. (2025) X. Yang, B. Zhou, X. Tang, J. Han, and S. Hu Chain of attack: hide your intention through multi-turn interrogation. In Findings of the Association for Computational Linguistics: ACL 2025, p. 9881â9901. External Links: Document Cited by: §2.2, Table 1, §4.1. Yi et al. (2024) S. Yi, Y. Liu, Z. Sun, T. Cong, X. He, J. Song, K. Xu, and Q. Li Jailbreak attacks and defenses against large language models: a survey. arXiv preprint arXiv:2407.04295. External Links: Document Cited by: §1, §2.3. Zeng et al. (2025) S. Zeng, Q. Wei, W. Brown, O. Frunza, Y. Nevmyvaka, and M. Hong Reinforcing multi-turn reasoning in LLM agents via turn-level credit assignment. External Links: 2505.11821 Cited by: §2.2, §3.5. Zeng et al. (2024) Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi How Johnny can persuade LLMs to jailbreak them: rethinking persuasion to challenge AI safety by humanizing LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 14322â14350. External Links: Document Cited by: §1, §2.3, Table 1, §3.3. Zhang et al. (2026) X. Zhang, C. Zhang, T. Li, Y. Huang, X. Jia, M. Hu, J. Zhang, Y. Liu, S. Ma, and C. Shen JailGuard: a universal detection framework for prompt-based attacks on LLM systems. ACM Transactions on Software Engineering and Methodology 35 (1), p. 8:1â8:40. External Links: Document Cited by: §2.3. Zhao and Zhang (2025) Y. Zhao and Y. Zhang Siren: a learning-based multi-turn attack framework for simulating real-world human jailbreak behaviors. In Annual Computer Security Applications Conference (ACSAC), External Links: Document Cited by: §2.2, Table 1, §4.1. Zhou et al. (2024) Z. Zhou, J. Xiang, H. Chen, Q. Liu, Z. Li, and S. Su Speak out of turn: safety vulnerability of large language models in multi-turn dialogue. arXiv preprint arXiv:2402.17262. External Links: Document Cited by: §2.2. Zou et al. (2023) A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. External Links: Document Cited by: §2.1, Table 1, §4.1.