Paper deep dive
GAIA-UL: Surgical Unlearning of Visual Knowledge via Causally-Guided Orthogonalization
Jinghan Xu, Xiulong Liu, Xin Xie, Kaixuan Zhang, Qixuan Cai, Xinyu Tong, Zheng Gong, Wenyu Qu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/11/2026, 1:16:50 AM
Summary
GAIA-UL is a three-stage machine unlearning framework designed for Multimodal Large Language Models (MLLMs) to remove sensitive visual knowledge. It utilizes Causal Hotspot Diagnosis, Targeted Adapter Intervention, and Semantically Orthogonal Fine-tuning to erase specific visual-semantic links while preserving general model utility.
Entities (5)
Relation Signals (3)
GAIA-UL โ evaluatedon โ MLLMU-Bench
confidence 98% ยท Extensive experiments on the MLLMU-Bench benchmark demonstrate that GAIA-UL significantly outperforms existing baselines.
GAIA-UL โ includesmethod โ Causal Hotspot Diagnosis
confidence 95% ยท Our approach first conducts a Causal Hotspot Diagnosis
GAIA-UL โ targets โ MLLMs
confidence 95% ยท GAIA-UL is a novel three-stage framework that performs Surgical Unlearning of visual knowledge in MLLMs.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Multimodal Large Language Models (MLLMs), while powerful, pose significant privacy risks by memorizing and potentially exposing sensitive information linked to individuals' visual appearances. Existing machine unlearning techniques, developed primarily for text-based models, are ill-equipped to handle the deeply entangled nature of visual and semantic knowledge. To address this challenge, we introduce GAIA-UL, a novel three-stage framework that performs Surgical Unlearning of visual knowledge. Our approach first conducts a Causal Hotspot Diagnosis, using gradient-based analysis to precisely identify influential parameters within the visual-semantic pathway. Second, it performs a Targeted Adapter Intervention, surgically injecting lightweight, trainable adapters only at these hotspots while freezing the base model. Finally, it employs Semantically Orthogonal Fine-tuning, a novel objective that forces the model's internal representation of a target face to become orthogonal to embeddings of associated sensitive concepts, thereby erasing the link at a deep representational level. Extensive experiments on the MLLMU-Bench benchmark demonstrate that GAIA-UL significantly outperforms existing baselines, achieving superior visual knowledge ablation while robustly preserving general model utility and text-only knowledge.
Tags
Links
Full Text
835 characters extracted from source content.
Expand or collapse full text
GAIA-UL: Surgical Unlearning of Visual Knowledge via Causally-Guided Orthogonalization | IEEE Conference Publication | IEEE Xplore IEEE Account Change Username/Password Update Address Purchase Details Payment Options Order History View Purchased Documents Profile Information Communications Preferences Profession and Education Technical Interests Need Help? US & Canada: +1 800 678 4333 Worldwide: +1 732 981 0060 Contact & Support About IEEE Xplore Contact Us Help Accessibility Terms of Use Nondiscrimination Policy Sitemap Privacy & Opting Out of Cookies A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.ยฉ Copyright 2026 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.