Paper deep dive
Machine Unlearning Across Scales: Evaluation of Optimization Methods on Language Models
Amartya Hatua, Trung T. Nguyen
Models: Llama-3-8B-bnb-4bit, SmolLM-135M-bnb-4bit
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/11/2026, 1:03:23 AM
Summary
This paper presents a comparative study of machine unlearning optimization methods (Gradient Ascent, Scaled Gradient Ascent, and WMDP-style Weighted Loss) across different language model scales (Llama-3-8B and SmolLM-135M) using the TOFU dataset, finding that model scale significantly impacts unlearning performance.
Entities (6)
Relation Signals (3)
TOFU โ usedforevaluationof โ Machine Unlearning
confidence 95% ยท Using the Task of Fictitious Unlearning (TOFU) dataset for evaluation
SmolLM-135M โ usedwith โ Scaled Gradient Ascent
confidence 95% ยท SmolLM-135M combined with Scaled Gradient Ascent demonstrating superior unlearning performance
Gradient Ascent โ evaluatedon โ Llama-3-8B
confidence 90% ยท comparative study of three optimization approaches... applied to two models... Llama-3-8B
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Machine unlearning has emerged as a critical capability for removing specific knowledge from trained language models while preserving general performance. However, the effectiveness of different unlearning optimization methods across varying model scales remains underexplored. In this work, we conduct a systematic comparative study of three optimization approaches, such as Gradient Ascent, Scaled Gradient Ascent, and Weapons of Mass Destruction Proxy (WMDP)-style Weighted Loss unlearning, applied to two models of dramatically different scales: Llama-3-8B-bnb-4bit (8 billion parameters) and SmolLM-135M-bnb-4bit (135 million parameters). Using the Task of Fictitious Unlearning (TOFU) dataset for evaluation, we assess the effectiveness of unlearning through both in-domain robustness testing on holdout forget samples and out-of-domain evaluation on general knowledge tasks using the MuskumPillerum/GeneralKnowledge dataset. Our evaluation employs ROUGE scores for answer quality comparison, Kullback-Leibler (KL) divergence for distribution analysis, and a novel retain-model comparison methodology that establishes objective benchmarks for successful unlearning. Results reveal significant differences in unlearning behavior across model scales, with SmolLM-135M combined with Scaled Gradient Ascent demonstrating superior unlearning performance compared to all other model-method combinations, achieving high behavioral alignment with retain-only models. Our work contributes to understanding how model capacity influences unlearning dynamics, introduces a rigorous evaluation framework based on retain-model comparison, and offers empirical evidence for method selection in real-world deployment scenarios that require privacy protection and knowledge management.
Tags
Links
Full Text
836 characters extracted from source content.
Expand or collapse full text
Machine Unlearning Across Scales: Evaluation of Optimization Methods on Language Models | IEEE Conference Publication | IEEE Xplore IEEE Account Change Username/Password Update Address Purchase Details Payment Options Order History View Purchased Documents Profile Information Communications Preferences Profession and Education Technical Interests Need Help? US & Canada: +1 800 678 4333 Worldwide: +1 732 981 0060 Contact & Support About IEEE Xplore Contact Us Help Accessibility Terms of Use Nondiscrimination Policy Sitemap Privacy & Opting Out of Cookies A not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.ยฉ Copyright 2026 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.