Paper deep dive
DEO: Jailbreak a Black-box Multimodal Large Language Model with Dual-Embedding Alignment
Lijie Zhang, Mingsi Wang, Yue Zhao, Zijin Lin, Kai Chen
Models: LLaVA, MiniGPT4
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/11/2026, 12:38:39 AM
Summary
The paper introduces Dual-Embedding Optimization (DEO), a black-box jailbreak attack for Multimodal Large Language Models (MLLMs). DEO optimizes visual adversarial perturbations by aligning input image embeddings and output text embeddings with a target harmful text in a shared embedding space, achieving high attack success rates against models like MiniGPT4 and LLaVa.
Entities (4)
Relation Signals (3)
DEO → attacks → MiniGPT4
confidence 95% · significantly improves attack success rates of existing black-box attack methods by up to 30% against two MLLM families, including MiniGPT4
DEO → attacks → LLaVa
confidence 95% · significantly improves attack success rates of existing black-box attack methods by up to 30% against two MLLM families, including... LLaVa
DEO → targets → MLLMs
confidence 95% · we propose a novel dual-embedding optimization (DEO) attack approach to generate visual adversarial perturbations that induce the MLLMs to produce harmful responses
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Multimodal Large Language Models (MLLMs), which integrate textual and visual modalities, have demonstrated unparalleled capabilities in diverse multimodal tasks. However, the inclusion of visual inputs exposes MLLMs to security risks, one of which is jailbreak attacks. Although various methods have been proposed to jailbreak MLLMs via the visual modality, attacks in black-box settings have some limitations. Existing black-box attacks either fail to generate precise harmful outputs in practical scenarios or require substantial preparatory work in constructing adversarial images. In this work, we propose a novel dual-embedding optimization (DEO) attack approach to generate visual adversarial perturbations that induce the MLLMs to produce harmful responses that violate common AI safety policies. Specifically, DEO iteratively optimizes the visual input by enforcing alignment objectives across both the input and output embedding spaces: the image embedding of the input and the text embedding generated by the MLLM are both required to align with a harmful target text within a shared embedding space, which is defined by a frozen pretrained encoder. This alignment is conducted entirely under a black-box setting using a query-based strategy, where the attacker issues queries and observes only the model’s outputs, without access to its internal parameters or gradients. By optimizing in the dual-embedding space, our method can generate an adversarial perturbation to elicit more harmful and precise responses, overcoming the limitations of existing approaches. Experimental results demonstrate that our method significantly improves attack success rates of existing black-box attack methods by up to 30% against two MLLM families, including MiniGPT4 and LLaVa, achieving an average attack success rate of 87% across different models and eight scenarios, demonstrating its superior attack effectiveness. These findings highlight the urgent need for systematic robustness evaluations and improved safety mechanisms in MLLMs.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Content Warning: This paper contains harmful model responses.
Tags
Links
Full Text
837 characters extracted from source content.
Expand or collapse full text
DEO: Jailbreak a Black-box Multimodal Large Language Model with Dual-Embedding Alignment | IEEE Conference Publication | IEEE Xplore IEEE Account Change Username/Password Update Address Purchase Details Payment Options Order History View Purchased Documents Profile Information Communications Preferences Profession and Education Technical Interests Need Help? US & Canada: +1 800 678 4333 Worldwide: +1 732 981 0060 Contact & Support About IEEE Xplore Contact Us Help Accessibility Terms of Use Nondiscrimination Policy Sitemap Privacy & Opting Out of Cookies A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.© Copyright 2026 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.