Paper deep dive
Selective LLM Unlearning via SAE-Based Token Importance Score
So Yeon Kim, Jung Hun Lim, Dong Woo Lee, Seong Hyun Min, In Jae Kim, Ma Il Jo, Ji Young Shin, Won Gyum Kim
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 5:49:46 PM
Summary
The paper introduces a selective LLM unlearning framework using Sparse Autoencoders (SAE) to identify and remove semantically critical tokens, addressing the need for efficient data removal under GDPR/CCPA without degrading model performance.
Entities (5)
Relation Signals (2)
Sparse Autoencoder → enables → LLM Unlearning
confidence 95% · we propose an unlearning framework that leverages a sparse autoencoder to selectively identify and remove only the most semantically critical tokens
GDPR → motivates → LLM Unlearning
confidence 90% · Regulatory frameworks, such as the EU’s General Data Protection Regulation (GDPR) ... mandate the immediate removal of such data
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Large language models (LLMs) trained on extensive web and dialogue datasets risk inadvertently memorizing sensitive personal or copyrighted information. Regulatory frameworks, such as the EU’s General Data Protection Regulation (GDPR) “right to be forgotten” and the California Consumer Privacy Act (CCPA) mandate the immediate removal of such data, yet retraining the entire model incurs prohibitive computational and financial costs. Conventional gradient ascent-based unlearning methods apply uniform gradient updates across entire input sequences, significantly degrading overall language capabilities. In this study, we propose an unlearning framework that leverages a sparse autoencoder to selectively identify and remove only the most semantically critical tokens. On the TOFU $1 \%$ forget set, our method achieves a Forget Quality of 0.99 and a Model Utility of 0.61, improving performance by over $50 \%$ and $9 \%$, respectively, compared to standard approaches. This method minimizes unnecessary updates to preserve overall performance and enable more robust and efficient unlearning.
Tags
Links
Full Text
810 characters extracted from source content.
Expand or collapse full text
Selective LLM Unlearning via SAE-Based Token Importance Score | IEEE Conference Publication | IEEE Xplore IEEE Account Change Username/Password Update Address Purchase Details Payment Options Order History View Purchased Documents Profile Information Communications Preferences Profession and Education Technical Interests Need Help? US & Canada: +1 800 678 4333 Worldwide: +1 732 981 0060 Contact & Support About IEEE Xplore Contact Us Help Accessibility Terms of Use Nondiscrimination Policy Sitemap Privacy & Opting Out of Cookies A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.© Copyright 2026 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.