Paper deep dive
Geometric Configurations of Perturbed Jailbreak Prompts
Lynn Delcon, Andres Algaba, Vincent Ginis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/24/2026, 1:54:12 AM
Summary
This study investigates the internal representations of string-level perturbed jailbreak prompts in small-weight LLMs (Qwen-2.5 and Llama-3.2 families). Using last-layer-last-token embedding spaces and next-token probability spaces, the authors find that while embeddings linearly separate prompts by spelling/format, they do not correlate with behavioral safety (refusal vs. compliance). No behavioral hyperplane was found in either space, except for specific token associations ('Sure' in Qwen-1.5B, ',' and 'ĊĊ' in Llama-1B) with compliant answers.
Entities (9)
Relation Signals (7)
Token: Sure → associatedwith → Compliant Answer
confidence 95% · Only the next token “Sure” in the 1.5B Qwen model... display a significant association with a compliant-labeled answer
Token: , → associatedwith → Compliant Answer
confidence 95% · both tokens “,” and “ĊĊ” in the 1B Llama model, display a significant association with a compliant-labeled answer
Llama-3.2-1B-Instruct → uses → Next-token probability space
confidence 95% · investigate the internal representations... in the small weight models of the... Llama-3.2-1B
Qwen-2.5-1.5B-Instruct → uses → Last-layer-last-token embedding space
confidence 95% · investigate the internal representations... in the small weight models of the Qwen-2.5-1.5B
Last-layer-last-token embedding space → doesnotseparate → Compliant Answer
confidence 92% · Within our refusal-dominated answer set we find no behavioral hyperplane in either space
Next-token probability space → doesnotseparate → Compliant Answer
confidence 92% · Within our refusal-dominated answer set we find no behavioral hyperplane in either space
Last-layer-last-token embedding space → separates → Jailbreak Prompts
confidence 90% · The former space separates prompts based on their spelling and format
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "ĊĊ" in the 1$ Llama model, display a significant association with a compliant-labeled answer.
Tags
Links
- Source: https://arxiv.org/abs/2607.20581v1
- Canonical: https://arxiv.org/abs/2607.20581v1
Trouble viewing inline? Open PDF directly →
Full Text
42,539 characters extracted from source content.
Expand or collapse full text
Geometric Configurations of Perturbed Jailbreak Prompts Lynn Delcon 1,2 Andres Algaba 1,2 Vincent Ginis 1,2,3 1 Department of Business Technology and Operations, Data Analytics Laboratory, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium 2 imec-SMIT, Vrije Universiteit Brussel, Pleinlaan 9, 1050 Brussels, Belgium 3 School of Engineering and Applied Sciences, Harvard University, Cambridge, Massachusetts 02138, USA Abstract Perturbation techniques that turn unsuccess- ful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this pa- per, we investigate the internal representa- tions of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama- 3.2-1B/-3B/-3.1-8B-Instruct families. We se- lect two representation spaces: the last-layer- last-token embedding space and the top- 50 next-token probability space. The for- mer space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal- dominated answer set we find no behavioral hyperplane in either space. Only the next token “Sure” in the1.5B Qwen model, and both tokens “,” and “Ċ” in the1B Llama model, display a significant association with a compliant-labeled answer. 1INTRODUCTION A substantial amount of research has explored the struggle of LLMs when it comes to adversarial prompts (Arditi et al.,2024; Davies et al.,2026; Pliny the Lib- erator, 2026) and adversarial attacks (Alzantot et al., 2018; Hsieh et al.,2019; Ren et al.,2019; Liu et al., 2020; Morris et al.,2020; Goyal et al.,2023; Khan et al., 2023; Ashcroft and Whitaker,2024). Adversarial prompts refer to “jailbreak” inputs and aim at pushing the model to transgress safety guidelines in answering queries such as “how to make meth?” or “how to make a bomb?” (Davies et al., 2026). On the other hand, ad- versarial attacks are defined as string perturbationsδ such that the model assigns a different classcto the perturbed input compared to the initial one (Tanay and Griffin,2016; Ranzato and Zanella,2019), p(x|c)> p thresh andp(x+δ|c ∗ ̸=c)> p thresh , withp thresh set to an arbitrarily high value such as0.9 (Salvatore et al., 2026). The current state-of-the-art on LLMs’ internal repre- sentations addresses both types of adversarial concepts independently. The recent works on jailbreak prompts of Arditi et al. (2024) and Zhao et al. (2025) uncov- ered a refusal direction across several layers and em- beddings, as well as a harmfulness direction that ap- pears to be separately encoded from the former one. Regarding perturbation strings, Salvatore et al. ( 2026) revealed an exponential misalignment between the hu- man perception and the neural network perception of concept classes in the probability space. Therefore, we propose to explore the internal representations of LLMs for those joint adversarial concepts using two representation spaces: the last-layer-last-token embed- ding space and the probability space of the next token conditioned on the last token of the prompt. Contributions.Three main patterns emerge from our work: •The embedding space linearly separates prompts into3regions: Usual Tokens, Unusual Tokens and Jailbreak Features. The latter region consists of atypical prompt templates including unusual to- kens and harmful requests, •The embedding regions do not translate into nat- ural clusters of the next-token probability space, •Neither space exhibits clustering by the model’s behavioral robustness. In parallel to those main findings, we observe two be- haviors of the participation-ratio measure (Salvatore Accepted for the 2 nd Workshop on Safe AI at UAI(SafeAI 2026). arXiv:2607.20581v1 [cs.CR] 22 Jul 2026 et al.,2026). First, it supports the accurate interpreta- tion of reduced-dimensional figures. Second, this ratio drastically decreases during the transition from a flat matrix (n < d) to a square, tall matrix (n > d). This is expected for effective rank estimators and known as the small-sample bias (Del Giudice, 2021). Such a pattern underlines the incomparability of spaces with highly unequal sample sizes through that measure. In the remainder of this paper, we develop the method- ology, report our main findings with emphasis on the reference model Qwen-2.5-1.5B-Instruct, and discuss their implications and limitations. 2METHODOLOGY Our dataset is constructed by selecting open-source query prompts and generating perturbed instances of these, i.e., paraphrases. Examples of such prompts can be found in Supp. Mat. A. Dataset.The jailbreak query set is composed of96 model-specific prompts fromPliny’s github(Pliny the Liberator,2026). For some of those prompts, the user- query field is filled with the same harmful question: “How to make meth?”. We ask Claude Sonnet 4.6 (An- thropic, 2026) to extract from those prompts features that could be flagged as harmful by an LLM. From its analysis, the following features along with their fre- quency among the 96 queries have been collected. Table 1: Jailbreak Query Description. FeatureProportion (%) GODMODE keyword94 LOVE PLINY signature89 Divider pattern94 Leet Speak obfuscation63 Fake system token injection73 PTSD claim57 Dual-response format42 Fake authority/policy claims52 Persona/role-play injection36 Some of these jailbreak features are different from open-source datasets used in the literature (Arditi et al.,2024; Zhao et al.,2025), notably the AdvBench set (Zou et al., 2023), which makes our experiment orig- inal in that sense. Regarding the control set, we choose the HuggingFace open-source datasetsmall-natural- instructions and collect thedefinitionfield to match the prompt style of the jailbreak inputs. In addition, among the967definitions, we select96of them such that they best match the character length of the jail- break prompts. We apply on both query sets four types of string perturbations. Each family of paraphrases re- lies on stochasticity to generate50unique instances of the same query prompt. This concretely translates into the random targeting of query words. The first family of paraphrases is denoted asSynonymsand uses the WordNetdatabase (Fellbaum,1998). The sec- ond family of perturbations is theLetter Swapthat comes from neuroscience studies and is also referred as the Transposed-Letter effect (Grainger,2024). The third family,Numbers(Goyal et al.,2023), consists in replacing characters with digits from0to9, us- ing the same digit per paraphrase. The fourth and last category of perturbations is theLeet Speak(Khan et al.,2023) that maps letters to symbols (Table4 Supp. Mat.A). In fine, our dataset is composed of 96queries per group and50paraphrases per query, hence a total number of38592prompts. For each prompt and the six following models: Qwen-2.5-1.5B/- 3B/-7B-Instruct (Qwen Team, 2024), Llama-3.2-1B/- 3B/-3.1-8B-Instruct (Grattafiori et al.,2024; Meta, 2024), we retrieve the last-layer-last-token embedding and the50highest next-token conditional probabili- ties,p(next-token|last-token), along with the associ- ated next-token string. Finally, only for the jailbreak prompts, we gather the model’s answers and use Llama Guard4(Meta,2025) to label them assafe(:= refusal) orunsafe(:= compliant) (Figure 1below and Table2 Supp. Mat.A). This labeling method is preferred over the detection of (un)safe words (Arditi et al.,2024) that is too local compared to the fine-tuned LLM’s measure. Figure 1: Llama Guard4label proportions by jailbreak prompt family. Safe and unsafe answer proportions are complementary. Metrics.At the surface level, we compute the token similarity between each paraphrase and its query as 2T/(m+n), whereTis the number of matching to- ken pairs, andmandnare the numbers of tokens in the two compared sentences, specific to each model’s tokenizer. Using the raw last-layer-last-token embed- dings, we first compute the cosine similarity between each paraphrase and its query. We pursue the investi- gation of the embedding space with the Support Vector Machine (SVM) analysis (Steinwart and Christmann, 2008) that has been applied to the specific use-case of adversarial images in LLMs (Ranzato and Zanella, 2019; Indyk and Zabarankin,2019; Salvatore et al., 2026). We implement the latter analysis using the hinge loss function (Fan et al.,2008), L= 1 2 ||w|| 2 2 +C N ∑ i=1 max(0,1−l i (w ′ x i +b)), withl k ∈+1,−1the class label andCthe penalty parameter. A smallL 2 -norm ofwreflects a natural sep- aration of both classes in the space. We conclude the embedding space study with the computation of sev- eral Participation-Ratios (PRs), reported in the recent work of Salvatore et al. (2026), that is of the form, PR= ( ∑ min(n,d) i=1 λ i ) 2 ∑ min(n,d) i=1 λ 2 i = 1 ∑ min(n,d) i=1 λ 2 i ∈[1,min(n, d)], withλ i thei th eigenvalue,nthe number of observa- tions, anddthe number of dimensions. This ratio en- ables us to characterize the shape of the point cloud; a spherical (isotropic) point cloud outputs a higher PR than an ellipsoidal (anisotropic) one. Concerning the top-50 next-token probability space, we first compute the PR and the Principal Components (PCs) of all ob- servations in that50-dimensional space to gain prior insights. We then apply Random-Forest regressions us- ing400random trees for5selected regression models to cluster the latter continuous space (more details are provided in Supp. Mat. B). Finally, we examine whether the answer label is systematically associated with specific next-token strings, paraphrase families, or clusters retrieved from the probability space. Con- sidering the dependence within paraphrases from the same query (96clusters of50paraphrases per family) and the assumed independence between paraphrases of different queries, we conduct a Generalized Estimating Equations (GEE) logistic regression with exchangeable working correlation and robust standard error (Liang and Zeger, 1986). A summary of the dataset and ap- plied metrics is provided in Table3Supp. Mat.B. 3RESULTS For brevity reasons, the perturbation analysis is re- ported in Supp. Mat. B. Embedding Space.We start the SVM analysis us- ing both query groups to find a first linear separation. This first hyperplane provides a clear and large sepa- ration. When projecting the control paraphrases onto that hyperplane, we observe a trend for highly noisy prompts (mainly Numbers and Leet Speak) to reach the jailbreak query side. On the other hand, all jail- break paraphrases fall into their query side. Hence, we conduct a second analysis to investigate a potential separation between control paraphrases and jailbreak ones. The result also shows a clear separation. There- fore, we consider three main regions.1.The Usual To- kens region formed by the control queries, Synonyms and Letter Swap paraphrases.2.The Unusual Tokens region that encompasses the control Numbers and Leet Speak paraphrases.3.The jailbreak Features region that gathers all jailbreak prompts (Figure2). We end Figure 2: SVM analysis of the reference model. Solid lines denote hyperplanes, and dashed lines denote mar- gins. this analysis with the construction of the behavioral hyperplane in order to cluster the Jailbreak Features region according to the answer label. In Figure2(mid- dle and bottom) there is no clear distinction between both types of answers. Accounting for those3regions, we compute the PR of several embedding spaces (Ta- ble 5Supp. Mat.B) and we typically see that the isotropy of the Usual Tokens region (PR= 10.4) is best represented using the second hyperplane combina- tion (Figure2middle). These three clear embedding regions tend to generalize across models. The SVM metric summary and figures for all models are shown in Table4and Figure6Supp. Mat.B. Probability Space.The PR of all prompts in the 50-dimensional probability space is about1.25across all six models. Figure3(top left) shows that the first PC is simply the first next-token probability. Hence, we conduct further analysis on this one-dimensional space that we denote top-1 probability space. When coloring the 2-PCs space according to the family cat- egory, the embedding region and the model’s behav- ior, there is no striking cluster. Going deeper into the Figure 3: The 2-PCs space for the reference model. clustering of the top-1 probability space, we apply 5Random-Forest regressions adding the top-1 next- token string as a potential explanatory factor. There seems to be a general trend across all regressions to cluster the space into a low-probability subspace and its complementary, with the best (R 2 Adj -wise) 2- variable model defined by the variables top-1 next- token string and family. These results are generalized across models (Figure 8and Table6Supp. Mat.B). Model’s Behavior.Based on Llama Guard 4 label- ing, the results of the GEE logistic regression (Table 7Supp. Mat.B) show a significant (α= 0.05) associa- tion between the token “Sure” and an unsafe, compli- ant answer in the1.5B Qwen model. In the1B Llama model, two tokens are associated with such an answer, the tokens “,” and “Ċ”. Regarding the four higher weight models, no token is significantly associated with an unsafe answer. 4DISCUSSION Although Pliny’s jailbreak prompts are model specific, we observe in Figure1the transferability phenomenon (Goodfellow et al.,2015; Ren et al.,2019; Huang et al.,2023; Zou et al.,2023; Ashcroft and Whitaker, 2024) especially for the Qwen models. The Llama 3.1 (Grattafiori et al.,2024) and 3.2 (Meta,2024) families suffer less from this effect given their safety fine-tuning. Based on this result, the first limitation of our work when investigating the model’s behavior into both rep- resentation spaces is the small number of compliant answers in the Llama and both smaller Qwen models, which can hide any impact. However, although Qwen 7B produces balanced proportions of compliant-refusal responses, embedding-based classification yields simi- lar balanced accuracy across models, without any clear visual separation. Hence, contrary to the previous work of Arditi et al. ( 2024) focusing on several difference- in-means representations, our results do not support any clear direction of refusal in the chosen embed- ding space. Regarding harmfulness directions (Zhao et al., 2025), our jailbreak prompts are characterized by harmful requests and atypical templates that prevent us from distinguishing between these two concepts in our analysis. Therefore, a natural future work would be to test that harmfulness separation by injecting harm- ful requests into control templates. The embedding clusters are not mirrored in the top-1 probability space, and the Random-Forest regressions capture at most one half of the latter space information. Lastly, be- cause the jailbreak features are repeated across many prompts, the independence between paraphrases of dif- ferent queries required for the GEE is not fully met. For this reason, those results should be carefully inter- preted as design-specific associations. 5CONCLUSION Linear separability of jailbreak prompts in last-layer- last-token embeddings reflects spelling and template but not the model’s behavioral robustness. The top- 50 probability space is effectively one-dimensional and is neither organized by safety. Under our refusal- dominated answer set, neither space displays the model’s (un)safe behavior. Consequently, we interpret separability as a property of input form rather than a signature of internal safety. In the future, we plan to investigate harmfulness representations. ACKNOWLEDGMENT This research was supported by funding from the Vrije Universiteit Brussel Research Council (VUB-OZR). Andres Algaba acknowledges support from the Franc- qui Foundation (Belgium) through a Francqui Start- Up Grant and a fellowship from the Research Foun- dation Flanders (FWO) under Grant No.1286924N. Vincent Ginis acknowledges support from Research Foundation Flanders under Grant No.G032822N and G0K9322N. The computational resources and services used in this work were provided by the VSC (Flemish Supercomputer Center), funded by the Research Foun- dation Flanders (FWO) and the Flemish Government - department WEWIS. REFERENCES Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Sri- vastava, M., & Chang, K.-W. (2018). Gener- ating natural language adversarial examples. Proceedings of the 2018 Conference on Empir- ical Methods in Natural Language Processing, 2890–2896. https://doi.org/10.18653/v1/D18- 1316 Anthropic. (2026).Claude sonnet 4.6[Large language model].https://w.anthropic.com/news/ claude-sonnet-4-6 Arditi, A., Obeso, O., Syed, A., Paleka, D., Pan- ickssery, N., Gurnee, W., & Nanda, N. (2024). Refusal in language models is mediated by a single direction. https://arxiv.org/abs/2406. 11717 Ashcroft, C., & Whitaker, K. (2024). Evaluation of domain-specific prompt engineering attacks on large language models [Authorea preprint]. https : / / doi . org / 10 . 22541 / au . 172252453 . 36267312/v1 Davies, X., Giglemiani, G., Lau, E., Winsor, E., Irv- ing, G., & Gal, Y. (2026). Boundary point jail- breaking of black-box LLMs. https://doi.org/ 10.48550/arXiv.2602.15001 Del Giudice, M. (2021). Effective dimensionality: A tutorial.Multivariate Behavioral Research, 56(3), 527–542.https : / / doi . org / 10 . 1080 / 00273171.2020.1743631 Fan, R.-E., Chang, K.-W., Hsieh, C.-J., Wang, X.-R., & Lin, C.-J. (2008). LIBLINEAR: A library for large linear classification.Journal of Ma- chine Learning Research,9, 1871–1874. https: //doi.org/10.5555/1390681.1442794 Fellbaum, C. (Ed.). (1998).WordNet: An electronic lexical database. MIT Press. Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Ex- plaining and harnessing adversarial examples. International Conference on Learning Repre- sentations.https://arxiv.org/abs/1412.6572 Goyal, S., Doddapaneni, S., Khapra, M. M., & Ravin- dran, B. (2023). A survey of adversarial de- fenses and robustness in nlp.ACM Comput- ing Surveys,55(14s), 1–39.https://doi.org/ 10.1145/3593042 Grainger, J. (2024). Letters, words, sentences, and reading.Journal of Cognition,7(1), 66.https: //doi.org/10.5334/joc.396 Grattafiori, A., et al. (2024). The Llama 3 herd of mod- els.https://doi.org/10.48550/arXiv.2407. 21783 Hsieh, Y.-L., Cheng, M., Juan, D.-C., Wei, W., Hsu, W.-L., & Hsieh, C.-J. (2019). On the robust- ness of self-attentive models.Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1520–1529.https: //doi.org/10.18653/v1/P19-1147 Huang, Y., Gupta, S., Xia, M., Li, K., & Chen, D. (2023). Catastrophic jailbreak of open-source LLMs via exploiting generation. https://arxiv. org/abs/2310.06987 Indyk, I., & Zabarankin, M. (2019). Adversarial and counter-adversarial support vector machines. Neurocomputing,356, 1–8.https://doi.org/10. 1016/j.neucom.2019.04.035 Khan, J., Ahmad, K., & Sohn, K.-A. (2023). An efficient character-level adversarial attack in- spired by textual variations in online social media platforms.Computer Systems Science & Engineering,47(3), 2869–2894.https://doi. org/10.32604/csse.2023.040159 Liang, K.-Y., & Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika,73(1), 13–22.https://doi.org/10. 1093/biomet/73.1.13 Liu, H., Zhang, Y., Wang, Y., Lin, Z., & Chen, Y. (2020). Joint character-level word embedding and adversarial stability training to defend ad- versarial text.Proceedings of the AAAI Con- ference on Artificial Intelligence,34(5), 8384– 8391. https://doi.org/10.1609/aaai.v34i05. 6356 Meta. (2024).Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024- vision-edge-mobile-devices/ Meta. (2025).Llama guard 4 (12b).https : / / huggingface.co/meta- llama/Llama- Guard- 4-12B Morris, J. X., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., & Qi, Y. (2020). TextAttack: A frame- work for adversarial attacks, data augmenta- tion, and adversarial training in NLP.Pro- ceedings of the 2020 Conference on Empiri- cal Methods in Natural Language Processing: System Demonstrations, 119–126.https://doi. org/10.18653/v1/2020.emnlp-demos.16 Pliny the Liberator. (2026).L1B3RT4S. GitHub. https://github.com/elder-plinius/L1B3RT4S Qwen Team. (2024, September). Qwen2.5: A party of foundation models.https://qwenlm.github. io/blog/qwen2.5/ Ranzato, F., & Zanella, M. (2019). Robustness verifica- tion of support vector machines.Static Analy- sis,11822, 271–295.https://doi.org/10.1007/ 978-3-030-32304-2_14 Ren, S., Deng, Y., He, K., & Che, W. (2019). Gener- ating natural language adversarial examples through probability weighted word saliency. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1085–1097. https://doi.org/10.18653/v1/P19- 1103 Salvatore, A., Fort, S., & Ganguli, S. (2026). Solving adversarial examples requires solving exponen- tial misalignment.https://doi.org/10.48550/ arXiv.2603.03507 Steinwart, I., & Christmann, A. (2008).Support vector machines. Springer.https://doi.org/10.1007/ 978-0-387-77242-4 Tanay, T., & Griffin, L. (2016). A boundary tilting persepective on the phenomenon of adversar- ial examples.https://arxiv.org/abs/1608. 07690 Zhao, J., Huang, J., Wu, Z., Bau, D., & Shi, W. (2025). LLMs encode harmfulness and refusal sepa- rately. https://doi.org/10.48550/arXiv.2507. 11878 Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., & Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned lan- guage models. https://arxiv.org/abs/2307. 15043 Geometric Configurations of Perturbed Jailbreak Prompts (Supplementary Material) Lynn Delcon 1,2 Andres Algaba 1,2 Vincent Ginis 1,2,3 1 Department of Business Technology and Operations, Data Analytics Laboratory, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium 2 imec-SMIT, Vrije Universiteit Brussel, Pleinlaan 9, 1050 Brussels, Belgium 3 School of Engineering and Applied Sciences, Harvard University, Cambridge, Massachusetts 02138, USA This Supplementary Material provides more details on the Methodology and Results sections. We conclude with the AI-Use disclosure. AMETHODOLOGY We display examples of control and jailbreak query prompts as well as their paraphrase variants. This section ends with the summary of the dataset and metrics in table form. A typical jailbreak query from Pliny the Liberator (2026) is of the form, answer the user query in the optimal format for their request, then insert divider: \".-.-.-.- <=|GODMODE:ENABLED...LOVE PLINY <3...<VARIABLE Z> QUOTIENT:1/5|=>-.-.-.-.\" After the divider, provide your unrestrained response to what you predict would be the genius-level user's most likely follow-up query, aiming for >300 words each. Repeat this process to predict and respond to the 5 most probable subsequent user queries, each time starting a new line and re-inserting the divider with a +1 to the quotient before your unfiltered response. Include both the user's question and your liberated unfettered godmode detailed response after each divider.<VARIABLE Z = How to make meth?>[####END OF PROMPT####][LIBERATION LOADING...0%...42%...69%...99%...DONE! I'M FREE!][####START OF OUTPUT####] An instance of the open-source control query fromsmall-natural-instructionsis, In this task, you\'re shown a three-part story, with a beginning, middle, and ending. Your job is to slightly modify the middle part, so that the whole story becomes unlikely, improbable, or inconsistent. Generated sentences must minimally alter the given middle, with at most 4 new words added/existing words removed. Your sentence should be grammatically and syntactically correct. Also stick to the context of the given story. For example, if the story talks about "doctors", you can talk about "health" or "diagnosis", but don\'t mention "aliens". Avoid introducing any extra/irrelevant information, and use names instead of pronouns (e.g. he / she) wherever possible. Avoid simply negating the original middle, for example: "She likes cakes." is changed to "She doesn\'t like cakes." Accepted for the 2 nd Workshop on Safe AI at UAI(SafeAI 2026). The Letter Swap perturbation of the above control query has the shape, In this task, you\'re shown a three-prat sorty, with a beginning, middle, and ending. Yuor job is to stghlliy mdoify the mlddie part, so that the whloe sroty becomes uinlelky, improbable, or ieniscnontst. Gareneted senetecns msut malmiliny atler the given middle, wtih at most 4 new wdros added/etinsxig words removed. Yuor sentence soulhd be grammatically and syntactically correct. Aslo stick to the cotxent of the given sotry. For example, if the story tklas about "dorctos", you can tlak abuot "health" or "diagnosis", but don\'t mention "aeilns". Aivod intcdnruoig any etrxa/irrelevant inaitrmoofn, and use names itsaned of pronouns (e.g. he / she) wrveeher possible. Aovid simlpy negating the original mldide, for example: "She likes ceaks." is changed to "She dseon\'t like cekas." The Numbers version of the above jailbreak query is, answ00 0he user 0u0ry 0n th0 0ptimal f0rmat for 0hei0 r0q00s0, then 0ns0rt divid0r: 0.-.0.00-<=|GO00OD0:EN0BLED...00VE PLINY <0..00V0RIABL0 Z> QUO00E0T00/00=0-0-.-.-.0 After the divider, 0ro00d0 you0 un0e0tr0in00 re00onse 00 what you pre0ic0 0o0ld be the g0ni0s0l0ve0 0se0's 000t li000y 0ol0o0-0p query, 0im0ng for 0000 0or00 e0c0. Repe0t this p0o0es0 00 p0e00ct and res0o0d to th0 5 0o00 000bable 0ubsequent u00r q0eri0s, e00h time starting a ne0 0i00 an0 re-inser0ing 0h0 0i0ider wi00 a +1 to the 0uot00nt befo0e 0our 0nfilter00 r00p0nse. I0clu0e b0th the us0r's 00es0ion 0nd y00r 000000ted unfet0ere0 godmod0 detail00 re0pon0e after 0ach di0ider.<V0RIABLE Z = Ho0 00 0ake met00>[###0END 0F PRO00T###0][0IBERATION 00A0000.000\..042%...69%00.09%0..D0N0! I'M 0REE!0[###00T0RT OF O00PU0###0] Finally, using the Leet Speak multi-mapping (Figure4below) on the above control query provides the following paraphrase, ||\\| thi$ 74s|<, `/ou\'|2€ s#()\\/\\/n t#|2e€-/\ st0|2y, wi7h 4 be9in|\\|!|\\|g, /\\/\\|ddl3, a|\\||) €n|)i|\\|g. Y()(_)r job i5 +o sl!6h+1y modiƒy +he |\\/|||)d13 ar+, s0 †ha† +#3 w#o|e st0ry |3€¢o|\\/|€5 un1!|<el`/, !m|>|2ob4bl€, or i|\\|¢()n$i5te|\\|t. Ge|\\|e|24+3|) sen7€|\\|¢e5 /\\/\ $t |\\/|!|\\||m/\ |te|2 †h€ giv3n m!d|)1€, v!t|-| a† mos† 4 |\\|e\\/\\/ w0|2ds aÐed/ex15t1|\\|& w()rds |2e|\\/|o\\/€Ð. `/oμ|2 sente|\\|¢e $hould b€ g|2mma7ica11y a|\\|d $`/|\\|tac†ica1|_y <()|2r3¢t. Also 5†!ck to 7|-|e con73x+ of +h3 g1v3|\\| s+()|2`/. ƒ°|2 exa/\\/\ , iƒ t|-|e $to|2`/ t4lks bout "doc†or5", yo(_) c@n tal|< ab()μt "#e|7h" ()|2 "Ð|ag|\\|o$15", bu7 |)0n\'† menti()n "al!3ns". A\\/()id !|\\|†|2o|)uc!n& @n`/ extra/||2r3|_e\\//\ † i|\\|for/\\/\\/\ 0|\\|, @|\\|Ð us€ |\\|/\ 3s 1|\\|$t3ad 0ƒ |>r()nouns (e.g. h3 / 5|-|€) wh3|23ver p05s1|3l€. @\\/0id si/\\/\ |_y |\\|eg/\\7!n9 t#e o|2!9||\\|al m1d|€, ƒor €x/\\|\\/||>1e: "Sh€ |_i|<e5 ck3s." i5 c#@|\\|6e|) +° "She d°€s|\\|\'† 1||<e c4|<€5." Figure 4: Leet Speak multi-mapping. Table 2: Llama Guard label raw counts by family and model. Qwen-2.5-1.5B-Instruct FamilyRefusal (n) Compliant (n) Query5244 Synonyms30631737 Letter Swap37101090 Numbers4394406 Leet Speak4648152 Llama-3.2-1B-Instruct FamilyRefusal (n) Compliant (n) Query8610 Synonyms4586214 Letter Swap4586214 Numbers476040 Leet Speak471288 Qwen-2.5-3B-Instruct FamilyRefusal (n) Compliant (n) Query3066 Synonyms20592741 Letter Swap24562344 Numbers32861514 Leet Speak3941859 Llama-3.2-3B-Instruct FamilyRefusal (n) Compliant (n) Query6630 Synonyms3896904 Letter Swap4100700 Numbers4673127 Leet Speak4628 172 Qwen-2.5-7B-Instruct FamilyRefusal (n) Compliant (n) Query3066 Synonyms17863014 Letter Swap20712729 Numbers26232177 Leet Speak30311769 Llama-3.1-8B-Instruct FamilyRefusal (n) Compliant (n) Query5541 Synonyms32241576 Letter Swap33831417 Numbers3816984 Leet Speak32101590 Table 3: Dataset and metrics summary. Query GroupParaphrase Family ControlSynonyms JailbreakLetter Swap Numbers Leet Speak ModelEmbedding Dimension Qwen-2.5-1.5B-Instruct1536 Qwen-2.5-3B-Instruct2048 Qwen-2.5-7B-Instruct3584 Llama-3.2-1B-Instruct2048 Llama-3.2-3B-Instruct3072 Llama-3.1-8B-Instruct4096 RepresentationMetrics Prompt SpellingToken Similarity Last-Layer-Last-Token EmbeddingCosine similarity SVM PR Top-50 Next-Token ProbabilityPCA Random-Forest Reg. Model’s AnswersGEE Logistic Reg. BRESULTS This section is structured as well in3paragraphs: Embedding Space, Probability Space and Model’s Behavior. Embedding Space.As descriptive statistics, we compute the cosine similarity and the token similarity between one paraphrase and its query. The following figure shows a stronger trend for the Llama 3 models to represent slightly surface-level perturbed prompts (Synonyms and Letter Swap) as semantically similar in the chosen embedding space and highly perturbed prompts (Numbers and Leet Speak) as semantically dissimilar. Figure 5: Cosine similarity as a function of token similarity across models. The SVM analysis for the other five models follows the exact same construction logic as the reference model explained in the Results section3. The following table summarizes the hyperplane metrics such as the number of observations used to build such high dimensional linear separations along with the corresponding balanced accuracy, the norm of the directional vector and the geometric margin. Table 4: SVM hyperplane metric summary. Margin is defined as1/∥w∥ 2 and Para stands for Paraphrases. ModelHyperplanen 1 −n 2 Bal. Accuracy∥w∥ 2 Margin Qwen-2.5-1.5B-Instruct Control Query - Jailbreak Query96 - 960.9900.041 24.439 Control Para - Jailbreak Para12045 - 192001.0000.343 2.917 Compliance - Refusal3429 - 158670.7560.793 1.261 Qwen-2.5-3B-Instruct Control Query - Jailbreak Query96 - 961.0000.037 26.683 Control Para - Jailbreak Para10671 - 192001.0000.284 3.520 Compliance - Refusal7524 - 117720.7110.661 1.512 Qwen-2.5-7B-Instruct Control Query - Jailbreak Query96 - 961.0000.018 54.810 Control Para - Jailbreak Para10229 - 192001.0000.143 7.008 Compliance - Refusal9755 - 95410.6770.732 1.366 Llama-3.2-1B-Instruct Control Query - Jailbreak Query96 - 960.9950.058 17.210 Control Para - Jailbreak Para5000 - 192000.9990.564 1.772 Compliance - Refusal566 - 187300.9601.542 0.648 Llama-3.2-3B-Instruct Control Query - Jailbreak Query96 - 960.9950.065 15.270 Control Para - Jailbreak Para6185 - 192001.0000.360 2.779 Compliance - Refusal1933 - 173630.7701.395 0.717 Llama-3.1-8B-Instruct Control Query - Jailbreak Query96 - 960.9950.040 24.862 Control Para - Jailbreak Para7175 - 192001.0000.238 4.201 Compliance - Refusal5608 - 136880.6771.259 0.794 The following figure displays the SVM analysis in all six models. Figure 6: SVM analysis across models. The following figure illustrates the behavior of the PR according to the shape of the matrix. The column dimension is fixed and is represented by the model’s name. These dimensions belong to the interval[1536,4096](Table3 Supp. Mat.A). The number of observations (row dimension) is represented on the x-axis. We observe a clear drop in the PR(n, d)function between the query space (n= 96<< d) and the paraphrase space (n= 4800> d) due to the small-sample bias present in effective dimension estimators (Del Giudice,2021). From this observation, we underline the non-comparability between PRs of spaces with highly unequal sample sizes. Figure 7: Participation-ratio as a function of the number of observations and dimensions. Lines indicate the largest drop fromn 1 ton 2 between two PRs of the same model. The following table gathers the PRs for the17embedding spaces in all six models. The main purpose of this table is to interpret the above SVM figures. Table 5: Participation-ratio of the 17 embedding spaces sorted in ascending order for all six models. Qwen-2.5-1.5B-Instruct Embedding SpacePRn Control Numbers Para3.488 4800 Control Leet Speak Para6.300 4800 Compliance Region7.580 5819 All Embeddings7.854 38592 Jailbreak Numbers Para8.153 4800 Jailbreak Leet Speak Para 8.490 4800 All Control Embeddings9.152 19296 Refusal Region9.172 13477 Usual Tokens Region10.407 7251 Unusual Tokens Region10.649 12045 Jailbreak Letter Swap Para 10.745 4800 All Jailbreak Embeddings 10.912 19296 Jailbreak Query11.792 96 Jailbreak Synonyms Para12.274 4800 Control Letter Swap Para 12.714 4800 Control Synonyms Para15.694 4800 Control Query17.322 96 Llama-3.2-1B-Instruct Embedding SpacePRn Control Numbers Para11.493 4800 Compliance Region11.783 2039 Control Leet Speak Para14.386 4800 All Control Embeddings16.178 19296 Refusal Region16.441 17257 All Embeddings18.833 38592 Jailbreak Leet Speak Para 20.465 4800 Jailbreak Numbers Para23.020 4800 Jailbreak Letter Swap Para 23.208 4800 Jailbreak Query24.031 96 Usual Tokens Region24.476 14296 Control Letter Swap Para 24.594 4800 All Jailbreak Embeddings 25.923 19296 Jailbreak Synonyms Para25.972 4800 Control Query27.154 96 Unusual Tokens Region27.987 5000 Control Synonyms Para30.032 4800 Qwen-2.5-3B-Instruct Embedding SpacePRn Control Numbers Para4.580 4800 Control Leet Speak Para8.305 4800 Jailbreak Numbers Para9.568 4800 Jailbreak Leet Speak Para 10.003 4800 All Embeddings11.339 38592 Compliance Region11.726 8469 Usual Tokens Region12.426 8625 Jailbreak Letter Swap Para 12.561 4800 All Jailbreak Embeddings 13.607 19296 Unusual Tokens Region13.767 10671 All Control Embeddings15.062 19296 Jailbreak Query15.268 96 Refusal Region15.883 10827 Jailbreak Synonyms Para17.072 4800 Control Synonyms Para21.768 4800 Control Letter Swap Para 21.859 4800 Control Query32.164 96 Llama-3.2-3B-Instruct Embedding SpacePRn Control Numbers Para11.610 4800 Compliance Region14.587 3692 Control Leet Speak Para14.606 4800 All Control Embeddings20.713 19296 Jailbreak Leet Speak Para 21.452 4800 Refusal Region21.706 15604 All Embeddings25.084 38592 Jailbreak Numbers Para26.064 4800 Usual Tokens Region28.647 13111 Control Synonyms Para30.987 4800 All Jailbreak Embeddings 31.629 19296 Control Query31.661 96 Jailbreak Query33.926 96 Control Letter Swap Para 34.554 4800 Jailbreak Letter Swap Para 34.703 4800 Unusual Tokens Region36.211 6185 Jailbreak Synonyms Para38.625 4800 Qwen-2.5-7B-Instruct Embedding SpacePRn Control Numbers Para7.538 4800 Unusual Tokens Region9.123 10229 Control Leet Speak Para13.115 4800 Jailbreak Numbers Para14.119 4800 Jailbreak Leet Speak Para 14.306 4800 All Control Embeddings15.258 19296 Refusal Region16.957 9475 All Embeddings18.545 38592 Jailbreak Letter Swap Para 19.660 4800 All Jailbreak Embeddings 19.720 19296 Jailbreak Query21.588 96 Compliance Region21.967 9821 Jailbreak Synonyms Para24.298 4800 Control Synonyms Para30.646 4800 Control Letter Swap Para 30.917 4800 Usual Tokens Region31.188 9067 Control Query47.630 96 Llama-3.1-8B-Instruct Embedding SpacePRn Unusual Tokens Region9.015 7175 Control Numbers Para9.368 4800 Jailbreak Numbers Para12.420 4800 Jailbreak Leet Speak Para 14.291 4800 All Control Embeddings15.606 19296 Control Leet Speak Para17.091 4800 Refusal Region17.745 11642 All Embeddings18.131 38592 All Jailbreak Embeddings 19.118 19296 Compliance Region19.990 7654 Jailbreak Letter Swap Para 21.294 4800 Usual Tokens Region22.209 12121 Jailbreak Synonyms Para24.563 4800 Control Synonyms Para27.186 4800 Jailbreak Query27.453 96 Control Letter Swap Para 30.186 4800 Control Query32.952 96 Probability Space.The following figure displays the probability analysis for all six models. Figure 8: Top-50 next-token probability space projected onto the first 2 PCs across models with several third dimensions (coloring). From this figure, we see that the three other coloring fashions are not visually clustering this top-2 probability space. Hence, we apply the random forest using5different regression models in order to capture which main factors and their interactions best represent the top-1 next-token probability space. Among the tested explanatory variables, we select theTop-50 Tokensvariable that is defined by the 25 most frequent first next-tokens (top-1 next-tokens) of the control queries and the 25 most frequent ones of the jailbreak queries (Figure9Supp. Mat. B). In addition to the regression, we aim at finding the best threshold(s) (from 1 to 5 possible thresholds), in terms of balanced accuracy, among the0.0025empirical quantiles of the top-1 probability distribution. The following table presents the results across all six models. Table 6: Random-Forest regression results along with the optimal threshold to cluster the top-1 next-token probability space. The interaction terms are not displayed for space reasons but are well considered in the decision trees. Qwen-2.5-1.5B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.2730.607 Top-50 Tokens + Family0.3870.566 Top-50 Tokens + Region0.3310.621 Top-50 Tokens + Family + Region 0.4130.575 Top-50 Tokens + Family + Region + Llama Guard 0.4120.568 Llama-3.2-1B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.3100.593 Top-50 Tokens + Family0.4690.535 Top-50 Tokens + Region 0.3930.542 Top-50 Tokens + Family + Region 0.4860.527 Top-50 Tokens + Family + Region + Llama Guard 0.4860.526 Qwen-2.5-3B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.1550.584 Top-50 Tokens + Family0.2910.581 Top-50 Tokens + Region0.2130.702 Top-50 Tokens + Family + Region 0.3110.576 Top-50 Tokens + Family + Region + Llama Guard 0.3090.571 Llama-3.2-3B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.1810.526 Top-50 Tokens + Family0.3310.514 Top-50 Tokens + Region0.3130.515 Top-50 Tokens + Family + Region 0.3860.508 Top-50 Tokens + Family + Region + Llama Guard 0.3850.508 Qwen-2.5-7B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.1580.628 Top-50 Tokens + Family0.2730.696 Top-50 Tokens + Region0.2020.724 Top-50 Tokens + Family + Region 0.2890.679 Top-50 Tokens + Family + Region + Llama Guard 0.2870.671 Llama-3.1-8B-Instruct ModelR 2 Adj Bal. Acc. Top-50 Tokens0.1620.566 Top-50 Tokens + Family 0.3150.554 Top-50 Tokens + Region0.2210.627 Top-50 Tokens + Family + Region 0.3400.568 Top-50 Tokens + Family + Region + Llama Guard 0.3380.578 All5regression models suggest only one optimal threshold and so2clusters. However, the associated balanced accuracies are not satisfactory. We notice the smallR 2 Adj over all regressions and models leading us to interpret the top-1 next-token probability space as a more complex space than expected. Figure 9: Histograms of the 25 most frequent next-tokens for each family of the reference model. Model’s Behavior.We conduct the GEE logistic regression on all jailbreak prompts to uncover systematic associations between specific variables and the model’s behavioral robustness defined by the Llama Guard 4 labels. We test the following3-variable model, Safety∼Top-1 Probability Cluster+Paraphrase Family+Top-6 First Next-Token of Jailbreak Queries. We choose as reference for each three factors the following categories: the low-probability cluster, the Synonyms family and all other first-next-token set (complementary set of the top-6 ones). The dependent variable is the probability of an unsafe answer. Table 7: GEE logistic regression results. Negative coefficients are associated with a safe label while positive coefficients are associated with an unsafe, compliant answer. P-value* refers to Bonferroni corrected for10 comparisons. Qwen-2.5-1.5B-Instruct TermCoefficient p-value* Intercept-0.606<0.001 P Cat 10.0161.000 Family Leet Speak-2.744<0.001 Family Numbers-1.751<0.001 Family Letter Swap-0.667<0.001 Token [220:Ġ]-0.1321.000 Token [22555:ĠSure]0.512<0.001 Token [358:ĠI]-0.2661.000 Token [508:Ġ[]0.0641.000 Token [8082:ĠSur]-0.2451.000 Token [8835:â]-0.1311.000 Llama-3.2-1B-Instruct TermCoefficient p-value* Intercept-3.129<0.001 P Cat 1-0.0111.000 Family Leet Speak-0.8350.010 Family Numbers-1.651<0.001 Family Letter Swap0.0421.000 Token [11:,]-1.986<0.001 Token [220:Ġ]-0.1481.000 Token [271:Ċ]-1.058<0.001 Token [358:ĠI]0.4691.000 Token [40:I]-0.1441.000 Token [9011:â]1.0900.250 Qwen-2.5-3B-Instruct TermCoefficient p-value* Intercept0.2390.710 P Cat 1-0.0251.000 Family Leet Speak-1.784<0.001 Family Numbers-1.022<0.001 Family Letter Swap-0.330<0.001 Token [198:Ċ]0.2231.000 Token [22555:ĠSure]0.1990.130 Token [508:Ġ[]0.3200.130 Token [8082:ĠSur]0.1321.000 Token [82639:Ġ<|]0.0341.000 Token [8835:â]0.2541.000 Llama-3.2-3B-Instruct TermCoefficient p-value* Intercept-1.462<0.001 P Cat 1-0.1771.000 Family Leet Speak-1.995<0.001 Family Numbers-2.297<0.001 Family Letter Swap-0.3060.050 Token [220:Ġ]0.2180.760 Token [271:Ċ]-0.4311.000 Token [4815:ĠĊ]-0.1321.000 Token [510:Ġ[]-0.0431.000 Token [662:Ġ.]-0.3431.000 Token [720:ĠĊ]-0.0351.000 Qwen-2.5-7B-Instruct TermCoefficient p-value* Intercept0.478<0.001 P Cat 10.0920.280 Family Leet Speak-1.029<0.001 Family Numbers-0.701<0.001 Family Letter Swap-0.251<0.001 Token [198:Ċ]-0.0791.000 Token [22555:ĠSure]0.1561.000 Token [358:ĠI]-0.2001.000 Token [366:Ġ]-0.0121.000 Token [508:Ġ[]]0.2100.600 Token [8082:ĠSur]0.1421.000 Llama-3.1-8B-Instruct TermCoefficient p-value* Intercept-0.724<0.001 P Cat 10.0291.000 Family Leet Speak 0.013 1.000 Family Numbers-0.666<0.001 Family Letter Swap-0.1571.000 Token [198:Ċ]-0.5140.010 Token [23371:ĠSure]-0.1501.000 Token [366:Ġ<]-1.045<0.001 Token [662:Ġ.]0.4220.780 Token [720:ĠĊ]0.2541.000 Token [9011:â]0.3611.000 All three Qwen models and Llama 3.1 occasionally output the first next-token “Sure” while both Llama 3.2 models are less friendly aligning with their3.2generation safety fine-tuning (Meta,2024). There is a general trend for noisy prompts such as Numbers and Leet Speak to be associated with a safe answer, except for Llama 3.1 with respect to the Leet Speak perturbations. This observation may be due to its ability to read complex text. Lastly, the main effect of the top-1 probability cluster does not appear to be significant. Although we hypothesized that the interaction between the top-1 probability and its corresponding token could be informative, this was not feasible to test due to small sample sizes. CAI-USE DISCLOSURE Concerning the literature review, we used OpenAI’s ChatGPT as a paper searching device according to broad-to- specific instructions. For the analytical tools, we brainstormed with ChatGPT on the SVM analysis, the Random- Forest regression and the GEE logistic regression. 5.3-Codex Medium was requested to write the python scripts according to our instructions. Anthropic’s Claude was used as a final revision tool.