Paper deep dive
The Dark Side of AI: a Systematization of Knowledge on Jailbreaking and Prompt Injection in LLMs
Nayana M, S Rupashree Reshma, Y Lakshmi Raj Harsha, Swaminadhan Rajula
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/11/2026, 1:14:37 AM
Summary
This paper provides a systematization of knowledge regarding security vulnerabilities in Large Language Models (LLMs), specifically focusing on jailbreaking and prompt injection. It introduces a taxonomy based on attacker objectives and technical execution methods, while also evaluating existing defense strategies categorized into input defense, model resilience, and output monitoring.
Entities (6)
Relation Signals (3)
Jailbreaking โ targets โ Large Language Models
confidence 95% ยท The proliferation of Large Language Models (LLMs) has introduced a new frontier of security vulnerabilities... focusing on two dominant attack classes: jailbreaking and prompt injection.
Prompt Injection โ targets โ Large Language Models
confidence 95% ยท The proliferation of Large Language Models (LLMs) has introduced a new frontier of security vulnerabilities... focusing on two dominant attack classes: jailbreaking and prompt injection.
Input Defense โ protects โ Large Language Models
confidence 85% ยท survey the current state of proposed defenses, categorizing them into a multi-layered framework of input defense, model resilience, and output monitoring
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
The proliferation of Large Language Models (LLMs) has introduced a new frontier of security vulnerabilities that diverge significantly from traditional software exploits. This paper provides a systematization of knowledge on the emerging threat vectors targeting LLMs, focusing on two dominant attack classes: jailbreaking and prompt injection. Jailbreaking aims to bypass the safety and ethical alignments of a model, whereas prompt injection manipulates model outputs by embedding adversarial instructions within its context. We structure this landscape by proposing a comprehensive taxonomy that classifies attacks along two primary axes: attacker objective (policy bypass, output hijacking, and unauthorized execution) and technical execution (direct, indirect, and cross-modal injection). Through a structured literature review, we synthesize foundational and recent works to analyze the mechanics of these attacks, from simple role-playing scenarios to sophisticated, multi-stage indirect injections. Furthermore, we survey the current state of proposed defenses, categorizing them into a multi-layered framework of input defense, model resilience, and output monitoring, and we critically evaluate their respective limitations. This work clarifies the current threat landscape, highlights critical research gaps, and provides a foundational reference for researchers and practitioners working to build more secure and robust AI systems.
Tags
Links
Full Text
845 characters extracted from source content.
Expand or collapse full text
The Dark Side of AI: a Systematization of Knowledge on Jailbreaking and Prompt Injection in LLMs | IEEE Conference Publication | IEEE Xplore IEEE Account Change Username/Password Update Address Purchase Details Payment Options Order History View Purchased Documents Profile Information Communications Preferences Profession and Education Technical Interests Need Help? US & Canada: +1 800 678 4333 Worldwide: +1 732 981 0060 Contact & Support About IEEE Xplore Contact Us Help Accessibility Terms of Use Nondiscrimination Policy Sitemap Privacy & Opting Out of Cookies A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.ยฉ Copyright 2026 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.