Paper deep dive
Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
Mohammad Mozaffari
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/26/2026, 5:30:47 AM
Summary
This thesis introduces the 'Compression Trinity,' a unified framework for Large Language Model (LLM) compression that jointly applies sparsity, quantization, and low-rank approximations to overcome the limitations of isolated techniques. It presents several methods: MKOR for optimizer acceleration via curvature approximation, SLOPE for pretraining with double-pruned sparsity and lazy adapters, OPTIMA for one-shot pruning via quadratic programming, PATCH for learnable hybrid sparsity, and SLIM for one-shot quantization and sparsity recovery using low-rank adapters. The work demonstrates that joint application significantly improves efficiency and accuracy compared to state-of-the-art single-pillar methods.
Entities (12)
Relation Signals (13)
Mohammad Mozaffari → authored → Compression Trinity
confidence 99% · by Mohammad Mozaffari A thesis submitted... This thesis proposes the 'Compression Trinity'
Compression Trinity → includestechnique → Quantization
confidence 95% · This thesis proposes the 'Compression Trinity,' a unified framework that applies the three pillars jointly: sparsity, quantization, and low-rank approximations.
Compression Trinity → includestechnique → Low-rank Approximation
confidence 95% · This thesis proposes the 'Compression Trinity,' a unified framework that applies the three pillars jointly: sparsity, quantization, and low-rank approximations.
Compression Trinity → includestechnique → Sparsity
confidence 95% · This thesis proposes the 'Compression Trinity,' a unified framework that applies the three pillars jointly: sparsity, quantization, and low-rank approximations.
MKOR → outperforms → KFAC
confidence 90% · accelerates convergence by up to 1.85x over KFAC.
SLOPE → usestechnique → Sparsity
confidence 90% · SLOPE accelerates training by up to 1.25x via a double-pruned backward pass for N:M sparsity
SLOPE → usestechnique → Low-rank Approximation
confidence 90% · using low-rank 'lazy' adapters in the final 1% of training to recover accuracy.
→ →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the "Compression Trinity," a unified framework that applies the three pillars jointly: sparsity to reduce computation, quantization to minimize memory bandwidth, and low-rank approximations to recover accuracy. To accelerate pretraining, we apply the Trinity to the optimizer and model architecture. MKOR approximates curvature via block-diagonal sparsity and low-rank inversion, maintaining numerical stability for quantized states; it reduces curvature update complexity from $O(d^3)$ to $O(d^2)$ and accelerates convergence by up to 1.85x over KFAC. SLoPe accelerates training by up to 1.25x via a double-pruned backward pass for N:M sparsity, using low-rank "lazy" adapters in the final 1% of training to recover accuracy. For post-training compression, OPTIMA stabilizes static masks in a zero-training regime by formulating weight reconstruction as globally optimal column-wise quadratic programs, improving zero-shot accuracy by up to 3.97%. Given a fine-tuning budget, PATCH breaks the ceiling of static masks by learning a dynamic hybrid sparsity ratio between 0% and 50%, yielding up to 1.38x speedups. Finally, SLiM realizes the full Compression Trinity in one shot, using mathematically derived low-rank adapters to recover information lost to quantization and sparsity, improving accuracy by up to 5.66% over state-of-the-art methods and outperforming uncompressed dense models at equal parameter budgets by 0.6%. Together, these results show that jointly applying the Compression Trinity is essential for efficient, scalable, high-performance LLMs.
Tags
Links
- Source: https://arxiv.org/abs/2608.24070v1
- Canonical: https://arxiv.org/abs/2608.24070v1
Trouble viewing inline? Open PDF directly →
Full Text
476,573 characters extracted from source content.
Expand or collapse full text
COMPRESSION TRINITY: EXPLORING SPARSITY, QUANTIZATION, AND LOW-RANK APPROXIMATIONS FOR LLM COMPRESSION by Mohammad Mozaffari A thesis submitted in conformity with the requirements for the degree of Doctor of Philosophy Department of Computer Science University of Toronto © Copyright 2026 by Mohammad Mozaffari arXiv:2608.24070v1 [cs.AI] 25 Aug 2026 Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression Mohammad Mozaffari Doctor of Philosophy Department of Computer Science University of Toronto 2026 Abstract Prohibitive computational and environmental costs impede the scalable deployment of Large Language Mod- els (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typi- cally applied in isolation, limiting effectiveness as each hits an accuracy-efficiency wall. Addressing this is critical for sustainable LLM deployment. This thesis proposes the “Compression Trinity,” a unified framework applying these pillars jointly. By leveraging sparsity to reduce computation, quantization to minimize memory bandwidth, and low-rank ap- proximations to recover accuracy, the framework unlocks new efficiency frontiers. We demonstrate these orthogonal methods are complementary, overcoming the limitations of isolated strategies. To accelerate pretraining, we apply the Trinity to the optimizer and model architecture.MKORapproxi- mates curvature via block-diagonal sparsity and low-rank inversion, maintaining numerical stability for quan- tized states. It reduces curvature update complexity fromO(d 3 )toO(d 2 ), accelerating convergence by up to 1.85×over KFAC.SLOPEaccelerates training by up to1.25×via a double-pruned backward pass for N:M sparsity, using low-rank “lazy” adapters in the final 1% of training to recover accuracy. For post-training compression, we progressively explore the Trinity to improve deployment efficiency.OP- TIMAestablishes sparsity limits under strict resource constraints, stabilizing static masks in a zero-training regime by formulating weight reconstruction as globally optimal column-wise quadratic programs, improving zero-shot accuracy by up to 3.97%. Given a fine-tuning budget,PATCHbreaks the ceiling of static masks by learning a dynamic hybrid sparsity ratio between 0% and 50%, yielding up to1.38×speedups. Finally, to overcome the structural limits of sparsity,SLIMrealizes the full Compression Trinity in one shot using mathematically derived low-rank adapters to recover information lost to quantization and sparsity, improv- ing accuracy by up to 5.66% over state-of-the-art methods and outperforming uncompressed dense models at equal parameter budgets by 0.6%. Together, these contributions demonstrate that the joint application of the Compression Trinity is essential for unlocking the next generation of efficient, scalable, high-performance LLMs. i Acknowledgements Firstandforemost, Iwouldliketoexpressmydeepestgratitudetomyfamily. Tomyfather, Gholamhossein, and my mother, Pouran, whose unwavering love, sacrifices, and encouragement have been the foundation of everything I have achieved. To my sister, Maryam, for her constant support and for always believing in me. This thesis is as much yours as it is mine. I would like to thank my supervisor, Professor Maryam Mehri Dehnavi, for providing the environment and resources that made this research possible. I am grateful to my thesis committee members, Professor Angela Demke Brown, Professor Dan Roy, Pro- fessorNanditaVijaykumar, andProfessorMuratErdogdu, fortheirvaluablefeedbackandthoughtfulquestions that strengthened this work. I would also like to extend my thanks to Professor Dan Alistarh for serving as the external appraiser and for taking the time to carefully evaluate this thesis. I owe a special debt of gratitude to Amir Yazdanbakhsh, who has been a mentor, collaborator, and source of inspiration throughout much of my research journey. His guidance and generosity with his time have profoundly shaped the way I approach problems. I am also grateful to Zhao Zhang for being a wonderful collaborator and for the many productive discussions we shared. I would like to acknowledge my undergraduate supervisors, Professor Maryam Sabbaghiyan and Professor AmirMasoudRabiei, whofirstsparkedmypassionforresearchandsetmeonthispath. Theirearlymentorship laid the groundwork for everything that followed. I have been fortunate to share the lab with exceptional friends and colleagues who made the journey both intellectually stimulating and enjoyable. I would like to thank Behrooz Zarebavani, Lucas Wilkinson, Younes Hourri, Kazem Cheshmi, Saeed Soori, Bangtian Liu, Avery Laird, Arya Rafii, Victor Kamel, Ray Hung, KasraJahankhani, LucyFarcnik, MiladKhanchi, andAmirhosseinElmiforthecountlessconversations, collaborations, and moments of friendship. To everyone who has contributed to this work, whether through direct collaboration or simply through friendship and encouragement, thank you. i ēïà Ō Ǎ ό ű ç Ŭό LJ ėό ˁ ، êàïŪ ό Ŋ ،Ě ύ ō ą ώĜŔ ʉ ό Ř ïíąόĶù، ë Ŵ ̭ Ʌ ό ̘ ęĎ ό ȵ ،ý Ō ώ ʷ ŏό ȼ ïî όĜ ïà.Ě Ž ό ʾ ĕ ń ýà ĘíàŪ ό Ŋ ç όŰĚ ύ ō î ̦ ό Ĺ àïíũ ό ʠ ņ ύ ʼn àíïî όIJùðç Ŭό LJ ë ύ ĸ Ō ό ʷ s Ż ʥ ό ɼ ă ό İ àĠ ĝ ، ïç όȵ ٓ àïí êç Ŭ t ų Ŭ Ɋ ό ķ Ęïàũ ʦ ό ɓ ėό ˁ ،Ě ό ō Ġ ĝ ،ýŏό ʚ àũ ό ʠ ïà. éό LJ àĘíŪ ό Ŋ ëό ǒ ēçόʘíïùç Ŭό LJ íĕ ń ç̪ ό Ť êç Ŭ ų ό ě êç όű ė Ȣ ό Ļ ù ņ ό ʼn ēçόʘ ĕ ń Ō Ǎ ό ư íùçόʘ ēïçόīàî όIJ، õ β ĺ ïí ņ ό ʼn s Ɍ ό ȶ . éό LJ ç̪ ό Ƌ ِ ê ٓ à ïà، éό LJ à ëό ǒ ِ ê ٓ à ïàėό ˁ Ę ïàî όĜà êç̪ό ɓ ė ό L· ė ƪ çόűï ë ό ĸ à. éό LJ àĘî Ŭ Ɋ ˄ ό Ĝ éό LJ í ëό ǒ ė ό L· ë Ŵ ό ű àí êç̪ ύ Ť à ïàĘçό˭ ġ t Ə ό ʙ ùĘíŪ ό Ŋ ëό ǒ ، é ό ǥ çόűĠ ّ Ī Ź ό Ķ àï ċό ʙ ù Ō ό ʷ ë ό ĸ àýç ̩ ό ň àėό ˁ ņ ύ ʼn ą όĜçƠόȟàùĢ ƈ Ʌ ό ̘ êíïù ٓ àĚό ʞ àŕ ό Œ ŏό ţ ç όŰė ό L· ،ēũ Ż ό ʙ íēŔ ʉ ό Ř Ě ό ō Ġ ĝ ïũ Ɍ ό Ğ ùŌ ό ʷ ،Ě ύ ō ç̪ ƍ ό ʙ àïíç Ŭό LJ à ïà .ýïà Ō Ǎ ό ű ç Ŭό LJ ė ό L· ç̪ ƍ ʥ ό Ɇ ïũ Ɍ ό Ğ ùŌ ό ʷ ùïąόĶŪό ƞ ĕ ɑ ύ ň ùç Ŭ Ϗ į î όĜą όĜïũ Ɍ ό Ğ ùŌ ό ʷ ،ēùï êíïũ Ɍ ό Ğ ùŌ ό ʷ ، êùàŌ ό ʷ ĕ ʵ ό Ķ íęĎ ̩ ό ň ٓ àïũ Ɍ ό Ğ ùŌ ό ʷ ،ýà ė ƪ çόűï êàïùàí é ٔ Ů ų ό ʘ ýĠ ź ʁ ό ̘ ēç ̧ό ȗ à ïà ùŌ ʸ ̭ ό ķ ûç̪ό ˒ ،î ύĜíŕό Ʀ Ō ό ʷ à ë ό ĸ àýçƠɍ W Ǝ ό LJ à é ό ǥ žό ǔ ėό ˁ êç όű ė ό L· ç ̭ ώ ķ î όĜà öï ïēçόʘ ñό Ŗ Ō ό ʷ ùî Ŭ ʥ ό Ƌ ïàēçόʘíïũ ό ʠ ïą όĜŏό ţ ç όŰė ό L· ،ùî όȵùíïà èàïžό ǔ öĠ Ő ė ƪ çόűï ë ό ĸ à s Ż ό Ğ íĕ ɂ ïŌ ό ʷ ēàŌ ό ʷ ėό ˁ ĕ ż ό Ğ ùĕ ȇ ïç όŰ èą ύĜ ïà ñ Ȧ ό Ĺ ðŌ ώ ʷ î όĜŏό ţ ç όŰė ό L· ïç Ŭ Ɋ Ź ό Ƙ ٓ à êíïũ Ɍ ό Ğ ùŌ ό ʷ ïà ë Ŵ ƒ Ʌ ɵ ό ɓ .ýïàíàï êç Ŭ W ų ό Ķ à .Ě Ž ό ʾ ĕ ń ņ ύ ʼn àíïî όIJ،î ύĜíũ ʦ ό Ť ëό ǒ ýç̫ό Ʈ à ٔ ė ʧ ɗ ό ǥ Ġ Ą ùïçƠʩό ɓ ،ņ θ ʼn Ġ ĝ ،ýà ĕ ɚ ό ʙ ù Ō ό ʷ Ġ ź Ɏ ό ǒ ïàēà Ęî̪ό ɼ ñ Ʌ ό ň ïíėό ˁ Ě Ž ό ʾ ĕ ń ñ Ʌ ό ň êàí Ō ό ʷ Ġ ź ǃ àïç Ŭ V ό į àïýà Ę Ō ύ ʷ ùðç Ŭό LJ ă ό İ àĠ ĝ ë Ŵ ƒ Ʌ ɵ ό ɓ . éό LJ àė W ŵ ό ǥ çόű êŪό Ʀ ŕό Ʀ í ً ç ̦ Ƅ ʥ ό ɼ àïü ٔ ό ě ç̭ό ǒ ّ üό Ű ė ό L· ëό ǒ ðŌ Ǎ ό Ĝ ، é ό Ğ ùòç̧ W ƈ ό ǥ àïí êç ̭ ό ķ àēî Ŭ ʥ ό Ť ùç ̩ό ȿ ùçόʘ ņ ύ ʼn ç̪ ƍ ό ʙ àï. éό LJ àĘíŪ ό Ŋ .ýïà Ō Ǎ ό ű ç Ŭό LJ ،î όűûî όĜùíï êç̪ ό Ť ç Ŭό ǒ ėό ˁ ņ ύ ʼn àùàŕ ό Œ ٔ Ęî όĜ ïçόűēçόʘũ ˻ Ŭ Ȧ ό ˲ ùĘî όĜ ïàēïçƠʩό ɓ ŏό ţ ç όŰė ό L· Ǹ ό ě à ïŪ ٔ ό Ŋ à ï ïà àï ċό ʙ ù Ō ό ʷ ė ό L· øç Ŭ ų ό ű àēçόʘ ė l· ŏ ό Ţ ë Ŵ Ŭ Ɋ Ʌ ό ň ėό ˁ ،ĕ ȯ Ƅ s ό į ïíũ Ȩ Ɍ ό ǒ Ġ ź ǃ àïũ Ɍ ό Ğ ùŌ ό ʷ ù êç Ŭ ό ȶ ç Ŭό ǁ Ě ό ō Ġ ĝ ïũ Ɍ ό Ğ ùŌ ό ʷ ،ýà ĕ ɂ ç Ŭ ό LJ ïçόī êàïùíî Ŭ ύ į çόűà ïà .íç̫ ό œ ç Ŭ ό į àïîόĶ ٓ àņ ό ʼn ïíė ʁ ό ň ٓ àĕ ń ç̪ ό Ť ēç Ŭ ώ į Ō ώ ʷ ï êç ̭ ό ķ à ë Ŵ Ŭ Ɋ Ʌ ό ň ēçόʘ ņ ύ ʼn ç̪ ƍ ό ʙ àï.Ě ύ ō ç̪ ό Ť ĕ ń ņ ύ ʼn àíïî όIJ،î ύĜíàíïàŕ ό Œ Ġ ź Ɏ ό ǒ ë ό ĸ àïíàĠ ĝ ùî Ŭ Ŷ ό Ű ùŕ ό Œ àŌ ό ʷ ëό ǒ ïí ñ Ʌ ό ň è ّ îόƘĚό ʞ ùïą ύĜŌ ό ʷ ĕ ŧ ƻ ό ȵ ŏ ɴ ό Ȕ ïàĚό ʞ àïĠ ź Ɏ ό ǒ ë ό ĸ àėό ˁ ýà ė ŵ ό LJ àíàïņ ύ ʼn ç Ŭ X ų Ŭ ό LJ àņ ύ ʼn àïçƠʩό ɓ ù êç Ŭό LJ ùíą όĜĕ ʝ çˮ ɒ ό ķ ąόĶ ï ٓ à Ěό ʞ ïç ̩ Ƒ ό Ğ àė ό L· ç ƒ Ʌ Ǝ ό LJ ũ ό ʠ ïũ Ż ˄ ό Ĝ ù،ĕ ȯ Ƅ ό Ğ ïą ύĜï ٓ à،íŕό ƚ ēïùà،ũ Ż ό Ɨ êç Ŭ ų ˹ Ŭ ό į ،ēïžό Dz î Ŭ ȧ ό Dz ،ĕ ŧ ɗ ό ǥ Ě Ȃ çόī،ēïũό ʠ ñ ό ķ Ū ό Ŋ ، êũ Ɍ Ź Ŭ ˄ Ƽ ό Ĝ ùðçόīŪό ƚ ،ņ ύ ʼn àŪ ό Ŋ ôïà ï ïùŔ ʉ ό œ ïà.î Ŭ Ŷ ό Ű çόű é ό Ğ ç όIJï èç ̨ ʂ ό Ƭ ùçόʘ ēïçƠʩό ɓ ،çόʘũ ˻ Ŭ Ȧ ό ˲ ŏό ţ ç όŰė ό L· ĕ ŧ ƻ ό ȵ ë Ŵ ̭ ό ǥ Ġ ź ǃ àùĕ ɑ ό ň ç όŰíęĎ Ŭ ό ǒ ،ʺ Ŭ Ŷ ό Ű ïç όIJĕ ɂ Ūό ƚ ،ņ ύ ʼn ç ̩ ό ň ç̫ ό ɸ ēĠ Ī ό ʾ ،Ǹ ό ě çόʘēï،üό Ķ çόī .ýïà Ō Ǎ ό ű ç Ŭό LJ ïç̪ ό Ƌ ņ ό ʼn .Ě Ž ό ʾ ĕ ń ņ ύ ʼn àíïî όIJė ό L· ç̪ ƍ ʥ ό Ɇ ،ĕ ń Ō Ǎ ό ư íùĕ ż ό LJ ùíą όĜ ً ç όIJĠ Ő ė Ǿ ùĚ Ž Ȧ Ƅ Ɋ ό ǒ ēïçƠʩό ɓ Ɓ β Ŋ ŏό ţ ïàė Ǿ ،î όĜà ė ŵ ό LJ àíĕ ŧ ʟ ό ʆ Ō ό ʷ à ë ό ĸ àïíėό ˁ ņ ύ ʼn ç̭ό ʾ ĕ ń ç̪ ό Ť ïà iv Dedicated to my parents, Gholamhossein Mozaffari and Pouran Shaban Ashini for their endless love and sacrifice. ،ýïíąόĶùïî όĜė ό L· Ě ύ ō î ̦ ό Ĺ ēŏ Ȫ ɫ ό Ș ë Ŵ ̭ Ʌ ό ̘ ęĎ ό ȵ ù ĕ ż ų ̭ ό ȶ êç Ŭ ό LJ êàïŪ ό Ŋ . êç ̭ ό ķ ą όĜą όĜ ņ ό ʼn ēïçόīàî όIJù s Ɍ ό ȶ ðą όĜė ό L· v Contents Acknowledgementsiii 1 Introduction1 1.1 LLM Life-cycle Stages. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .1 1.2 Compression Techniques. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 1.2.1 Sparsity. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 1.2.2 Quantization. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 1.2.3 Low-rank Approximations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 1.3 The Failure of Isolated Compression. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 1.4 Thesis Contributions and Roadmap. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .4 2 Background7 2.1 The Transformer Bottleneck: Linear Layers. . . . . . . . . . . . . . . . . . . . . . . . . .7 2.2 Hardware Constraints and the Roofline Model. . . . . . . . . . . . . . . . . . . . . . . . .9 2.2.1 Memory Hierarchy and Tensor Cores. . . . . . . . . . . . . . . . . . . . . . . . .9 2.2.2 The Roofline Model. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .10 2.2.3 Sparse Tensor Core Acceleration. . . . . . . . . . . . . . . . . . . . . . . . . . . .10 2.3 The Solution Space: Two Frontiers of Acceleration. . . . . . . . . . . . . . . . . . . . . .11 2.4 Strategy 1: Algorithmic Efficiency. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .12 2.4.1 First-Order Methods. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .12 2.4.2 Second-Order Methods. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .12 2.4.3 The Computational Barrier and Structured Approximations. . . . . . . . . . . . . .13 2.5 Strategy 2: Hardware Efficiency and LLM Regimes. . . . . . . . . . . . . . . . . . . . . .13 2.5.1 The Compute-Bound Regimes: Training and Prefill. . . . . . . . . . . . . . . . . .14 2.5.2 The Memory-Bound Regime: Inference Decoding. . . . . . . . . . . . . . . . . .14 2.6 The Compression Trinity as a Solution Framework. . . . . . . . . . . . . . . . . . . . . .15 2.7 The Case for a Joint Approach. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .16 3 MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates17 3.1 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 3.2 Background. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .19 3.3 Methodology. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .20 3.3.1 The MKOR Algorithm. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .20 3.3.2 Hybrid MKOR. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .22 vi 3.3.3 MKOR Convergence and Stability. . . . . . . . . . . . . . . . . . . . . . . . . . .23 3.4 Experimental Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 3.5 Conclusion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .31 4 SLOPE: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs32 4.1 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .32 4.2 Additional Related Work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .34 4.3 Sparse Plus Low-rank Pretraining of LLMs. . . . . . . . . . . . . . . . . . . . . . . . . .35 4.3.1 Double-pruned Backward Pass. . . . . . . . . . . . . . . . . . . . . . . . . . . . .35 4.3.2 Lazy Low-rank Adapters. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .36 4.3.3 Sparse Kernels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .37 4.3.4 SLOPE Runtime Optimization. . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 4.4 Experimental Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .39 4.4.1 End-to-end Speedup and Memory Saving: Pretraining and Inference. . . . . . . . .39 4.4.2 Pretraining Accuracy Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 4.4.3 Ablation Studies. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .44 4.5 Conclusion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .46 5 OPTIMA: Optimal One-Shot Pruning for LLMs via Quadratic Programming Reconstruction48 5.1 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .48 5.2 Additional Related Work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .50 5.3 Preliminaries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .51 5.4 OPTIMA: Optimal Weight Updates via Quadratic Programming. . . . . . . . . . . . . . .52 5.4.1 Reformulation as a Quadratic Program with Linear Constraints. . . . . . . . . . . .52 5.4.2 Reformulation as an Unconstrained Quadratic Program. . . . . . . . . . . . . . . .53 5.4.3 Solving the Quadratic Programs. . . . . . . . . . . . . . . . . . . . . . . . . . . .54 5.4.4 Efficient Implementation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .54 5.5 Experiments. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .55 5.6 Conclusion and Limitations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .59 6 PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs67 6.1 Statement of Contributions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .67 6.2 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .67 6.3 Additional Related Work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .68 6.3.1 Pruning methods. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .68 6.3.2 Complementary compression techniques. . . . . . . . . . . . . . . . . . . . . . .69 6.4 Preliminaries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .70 6.5 PATCH. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .71 6.6 Efficient deployment of PATCH. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .73 6.7 Experiments. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .73 6.7.1 Model Quality Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .74 6.7.2 Understanding the components of PATCH. . . . . . . . . . . . . . . . . . . . . . .75 6.7.3 Speedup and memory savings. . . . . . . . . . . . . . . . . . . . . . . . . . . . .77 6.8 Conclusion and Limitations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .78 vii 7 SLIM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression81 7.1 Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .81 7.2 Related work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .82 7.2.1 Pruning. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .82 7.2.2 Quantization. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .83 7.2.3 Low-rank Adapters. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .83 7.2.4 Sparse Plus Low-Rank Matrix Decomposition. . . . . . . . . . . . . . . . . . . . .83 7.3 Preliminaries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .84 7.4 Quantized sparse plus low-rank approximation of LLMs. . . . . . . . . . . . . . . . . . .85 7.4.1 SLIM-Quant quantization method. . . . . . . . . . . . . . . . . . . . . . . . . . .86 7.4.2 SLIM-LoRA low-rank adapters. . . . . . . . . . . . . . . . . . . . . . . . . . . .87 7.4.3 Low-rank adapter quantization. . . . . . . . . . . . . . . . . . . . . . . . . . . . .89 7.4.4 Optional Post-Compression Fine-Tuning. . . . . . . . . . . . . . . . . . . . . . . .89 7.5 Experimental results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .90 7.6 Conclusion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .97 8 Conclusion and Future Work99 8.1 Summary of Contributions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .99 8.1.1 The Trinity in Training Dynamics. . . . . . . . . . . . . . . . . . . . . . . . . . .99 8.1.2 The Trinity in Post-Training and Inference. . . . . . . . . . . . . . . . . . . . . . .99 8.2 Exploratory Frameworks and Open Research. . . . . . . . . . . . . . . . . . . . . . . . .100 8.2.1 BEAM: Blockwise Error Minimization for One-shot Compression of LLMs. . . . .100 8.2.2 LEAP: Learnable End-to-End Adaptive Pruning of LLMs. . . . . . . . . . . . . .101 8.2.3 SLICE: Selecting Layer-wise Configurations for Matryoshka-Style LLMs. . . . . .101 8.3 Limitations of the Current Approach. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .101 8.4 Future Research Directions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .102 8.5 Closing Remarks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .102 A Supplementary Material for MKOR115 A.1 GLUE Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .115 A.2 Derivation of NGD Approximations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .115 A.3 Numerical Instability of Second-order Methods. . . . . . . . . . . . . . . . . . . . . . . .117 A.4 Sensitivity to Learning Rate. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .118 A.5 Scalability. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .118 A.6 Decaying Eigenvalues and Rank-1 Approximations. . . . . . . . . . . . . . . . . . . . . .118 A.7 Training Accuracy Experiments. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .119 A.8 Knee-Point Learning Rate Scheduler. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .119 A.9 Proofs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .120 B Supplementary Material for SLOPE126 B.1 Comparison with Dynamic Sparsity: SR-STE. . . . . . . . . . . . . . . . . . . . . . . . .126 B.2 cuSPARSELt Initialization Overhead: Static vs. Dynamic Sparsity. . . . . . . . . . . . . .127 B.3 BERT-Large-Uncased: Pretraining and Downstream Evaluation. . . . . . . . . . . . . . . .127 viii B.4 Performance overhead of bidirectional mask. . . . . . . . . . . . . . . . . . . . . . . . . .128 B.5 Sparsity ratio analysis of double-pruned backward pass. . . . . . . . . . . . . . . . . . . .129 B.6 Sensitivity to the choice of pruning matrix. . . . . . . . . . . . . . . . . . . . . . . . . . .129 B.7 Implementation details. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .130 B.8 Task-specific GLUE results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .132 B.9 Integration with Flash Attention. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .132 B.10 Comparison with dense models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .133 B.11 Zero-shot GLUE results for GPT. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .133 B.12 Extended SR-STE and FST implementation details. . . . . . . . . . . . . . . . . . . . . .134 B.13 Comparison of Depth and Width Pruning. . . . . . . . . . . . . . . . . . . . . . . . . . .135 B.14 Proofs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .136 C Supplementary Material for OPTIMA140 C.1 Calibration dataset size sensitivity. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .140 D Supplementary Material for PATCH142 D.1 STOICC Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .142 D.2 Per Task Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .144 D.3 Tile Transfer Learning. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .147 E Supplementary Material for SLIM148 E.1 Notations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .148 E.2 Input Quantization. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .148 E.3 Additional fine-tuning results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .149 E.4 Language modeling experiments. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .149 E.5 Additional Sparse and Quantized Results. . . . . . . . . . . . . . . . . . . . . . . . . . . .153 E.6 Sparsity vs. quantization. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .154 E.7 Additional speedup results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .155 E.8 Fine-tuning costs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .156 E.9 Memory reduction analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .157 E.10 Computation reduction analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .158 E.11 Compression costs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .158 E.12 Rank analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .159 E.13 Effects of calibration sample count. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .159 E.14 Sensitivity to calibration dataset. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .159 E.15 Sparsity analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .160 E.16 Group quantization challenges. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .161 ix List of Tables 1.1TheSparsityParadox:ComparisonofLLaMA-2-7Bzero-shotaccuracyunderstandardPost- Training Compression. Sparsity yields lower compression ratios (2x) yet results in signifi- cantly higher accuracy loss compared to Quantization (4x), highlighting its destructive nature.3 1.2 Average Zero-shot Accuracy of LLaMA-2-7B on 8 tasks (MMLU, PIQA, ARC-Easy, ARC- Challenge, WinoGrande, OpenBookQA, RACE, HellaSwag) at 8x compression ratio. Single- pillar methods fail to retain capabilities, while multi-pillar methods (Compression Trinity) recover accuracy. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .4 3.1 The computation and communication complexity and memory overhead of the state-of-the-art implementations of the first- and second-order (second-order optimizers are written inbold). The division by 2 in MKOR is because MKOR uses half-precision computations. The com- plexity of KFAC-based methods depends on layer dimensions while SNGD methods mostly depend on the batch size. In transformers, due to the scaling of the batch size by the sequence length, batch sizes and layer dimensions are comparable, making both KFAC- and SNGD- based methods more expensive than SGD. . . . . . . . . . . . . . . . . . . . . . . . . . . .22 3.2 List properties of the models, datasets, and settings used in our experiments.. . . . . . . . .25 3.3 BERT-Large Uncased results on SQuAD v1.1 question answering task. . . . . . . . . . . .26 3.4 BERT-Large Uncased results on the GLUE classification tasks. We report the average of the metrics of different GLUE tasks (accuracy, F1 score, etc) for easier comparison.. . . . . .26 3.5 Per-GPU memory usage (in GB) for MKOR, KFAC/KAISA, LAMB, and SGD on BERT- Large-Uncased pre-training and ResNet-50 training on ImageNet. . . . . . . . . . . . . . .29 4.1 Comparative analysis of end-to-end pretraining and inference speedup (×) comparison be- tween SLOPE and the latest work (FST) on accelerating pretraining with 2:4 sparsity (ICML 2024) [62]. The baseline is dense PyTorch implementation of the models with CUBLAS backend. Note that the lack of inference speedup in FST is because of the final dense pretrain- ing during the final iterations, resulting in a dense model for inference. E-SR-STE stands for Extended SR-STE. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .40 4.2 Comparative analysis of end-to-end memory reductions ( × ) during training and inference between SLOPE and the latest work (FST) on accelerating pretraining with 2:4 sparsity (ICML 2024) [62]. Values greater than1.00×show memory overhead.. . . . . . . . . . . . . . .41 x 4.3 GPT2-Small accuracy results on zero-shot tasks. Adapter rank is the ratio of the low-rank adapter to the hidden dimension of the model. For Extended SR-STE, we have used a de- cay factor of6×10 −6 , since it resulted in the lowest perplexity in OpenWebText. The best performing sparse configuration is highlighted in bold.. . . . . . . . . . . . . . . . . . . .42 4.4 SQuAD-v1.1 accuracy and GLUE results on BERT-Large-Uncased with different adapter ranks43 4.5 SQuAD-v1.1 accuracy results on BERT-Large-Uncased for different sparsity settings.. . . .44 4.6 End-to-end speedup (×) before and after efficient implementation of low-rank adapters.. . .45 4.7 End-to-end speedup (×) before and after splitting the upsample matrix. In both cases, the optimization discussed in Table 4.6 is used.. . . . . . . . . . . . . . . . . . . . . . . . . .46 4.8 SQuADv1.1 results on BERT-Large-Uncased for different pruned modules.. . . . . . . . .46 5.1 Key hyperparameters used in OPTIMA.. . . . . . . . . . . . . . . . . . . . . . . . . . . .56 5.2 Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 50% unstruc- tured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .60 5.3 Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 2:4 sparsity. In this experiment, only the layers in the MLP part of the transformer are pruned, and the self- attention layers are dense, resulting in an end-to-end sparsity ratio of 38% to 41%. OPTIMA consistently improves the accuracy of the models across different tasks. Please note that Prox- Sparse pruning is limited to 2:4 sparsity, and hence our unstructured sparsity experiments do not include it. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .61 5.4 Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 60% unstruc- tured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .62 5.5 Qwen-2.5 family perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 50% unstructured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .63 5.6 Qwen-2.5 perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 60% un- structured sparsity. OPTIMA consistently improves the accuracy of the models across differ- ent tasks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .64 5.7 Qwen-2.5 perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 2:4 spar- sity. In this experiment, only the layers in the MLP part of the transformer are pruned, and the self-attention layers are dense, resulting in an end-to-end sparsity ratio of 38% to 41%. OPTIMA consistently improves the accuracy of the models across different tasks. Please note that ProxSparse pruning is limited to 2:4 sparsity, and hence our unstructured sparsity experi- ments do not include it.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .65 5.8 Comparison of OPTIMA with other optimizers without convergence guarantees (ADAM). ADAM can lead to suboptimal solutions (Gemma 3 1B) or divergence of the model (OPT 125M).. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .66 6.2 Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By jointly optimizing the location of dense tiles and the sparsity pattern within the sparse tiles, PATCH Joint allows for a continuous sparsity ratio for the models, providing a flexible tradeoff between sparsity and model quality. . . . . . . .74 xi 6.1 Hyper-parameters used for PATCH Joint and PATCH Tile across sparsity ratios. All hyper param- eters were tuned on Qwen-2.5-0.5B.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .74 6.3 Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By only optimizing the location of dense tiles while keeping sparsity pattern within the sparse tiles frozen, PATCH Tile provides a memory efficient variant for PATCH Joint , allowing for a continuous sparsity ratio for the models and providing a flexible tradeoff between sparsity and model quality.. . . . . . . . . . . . . . . . . . . . .75 6.4 Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for PATCH, Wanda, and SparseGPT. For models with less than or equal to 1B param- eters, PATCH Joint optimizes both dense tile locations and sparsity patterns, while for larger models PATCH Tile optimizes only dense tile locations with frozen sparsity patterns, both us- ing Dense/2:4 Tiles pattern allowing continuous sparsity ratios and flexible tradeoffs between sparsity and model quality. Wanda and SparseGPT are unstructured pruning methods.. . . .76 6.5 Impact of PATCH’s tile size across sparsity levels (↓is better). The effect of tile size on model quality is not significant, showing PATCH’s robustness against tile size.. . . . . . . . . . .76 6.6 Global sparsity yields better quality by concentrating pruning in less important blocks and preserving density elsewhere (↓is better).. . . . . . . . . . . . . . . . . . . . . . . . . . .76 6.7 Impact of fixed 2:4 mask selection for PATCH Tile , compared with joint optimization (↓is better). PATCH Joint achieves the lowest perplexity overall, while for PATCH Tile , MaskLLM provides the best frozen mask.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .77 7.1 Average zero-shot accuracy of LLaMA-2 and OPT models with50% sparsity and 4-bit weight quantization.Best Method ∗ indicates the best quantization method out of Group AbsMax, AWQ, OmniQuant, and AffineQuant.↑indicates better performance.. . . . . . .91 7.2 Effects of fine-tuning on the average zero-shot accuracy of LLaMA-2 models with 50% spar- sity and 4-bit weight quantization.↑indicates better performance. . . . . . . . . . . . . . .92 7.3 Accuracy results of OPTIMA weight update mechanism with SLIM-LoRA.↑indicates better performance.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .92 7.4 Average accuracy (↑indicates better) across eight zero-shot downstream tasks (including RACE[74] and HellaSwag [161]) and WikiText2 perplexity (↓indicates better) of compressed models with4-bit weight-only quantization. Please note that using LoRA adds additional parameters to the model.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .93 7.5 Average zero-shot accuracy of LLaMA-2 and OPT models with pruning. The quantization is disabled in this experiment.↑indicates better performance.. . . . . . . . . . . . . . . . . .95 7.6 Average zero-shot accuracy of LLaMA-2 and OPT models with quantization. The sparsity is disabled in this experiment.↑indicates better performance.. . . . . . . . . . . . . . . . . .96 7.7 LLaMA-2 family of models speedup (×) using SLIM compared to original dense unquantized model on NVIDIA RTX-3060.↑shows higher speedup.. . . . . . . . . . . . . . . . . . . .97 A.1 BERT-Large-Uncased Results on the GLUE classification tasks.. . . . . . . . . . . . . . .115 A.2 Number of epochs necessary for convergence in different optimizers for ResNet-50 on CI- FAR10. MKOR is the least sensitive optimizer to learning rate, converging in almost the same number of iterations for a wide range of learning rate, while other optimizers either diverge (D) or converge to a local-minimum (∗superscript). . . . . . . . . . . . . . . . . .118 xii B.1 End-to-end slow-down of Bi-directional Mask [166] in comparison to the dense baseline.. .128 B.2 GLUE results for each task in the experiments discussed in Section 4.4.. . . . . . . . . . .132 B.3 Speedup of SLOPE and FlashAttention-2 (FA2) on OPT models.. . . . . . . . . . . . . . .132 B.4 Performance comparison across different GPT models, sparsity methods, and LoRA ranks on various tasks. E-SR-STE stands for Extended SR-STE.. . . . . . . . . . . . . . . . . . . .133 B.5 Performance comparison of GPT models using different sparsity methods and LoRA ranks on GLUE tasks. E-SR-STE stands for Extended SR-STE.. . . . . . . . . . . . . . . . . . . . .133 B.6 Description of Key Terms. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .134 B.7 Model Configurations for LLaMA-2 7B. . . . . . . . . . . . . . . . . . . . . . . . . . . .136 B.8 Model Configurations for Gemma-9B. . . . . . . . . . . . . . . . . . . . . . . . . . . . .137 B.9 Model Configurations for Gemma-2B. . . . . . . . . . . . . . . . . . . . . . . . . . . . .138 D.1 Throughput of LLaMA-2 7B with mixed sparsity compared to the dense model. Measure- ments taken on an A6000 GPU with batch size 16. Throughput is reported in tokens pro- cessed/sec.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .144 D.2 Throughput of LLaMA-2 7B with mixed sparsity compared to the dense model. Measure- mentstakenonanA100GPUwithbatchsize16. Throughputisreportedintokensprocessed/sec. 144 D.3 Model quality (task accuracy across eight zero-shot tasks, reported in %) for Qwen-2.5 0.5B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity pat- terns, enabling a flexible sparsity-quality tradeoff. . . . . . . . . . . . . . . . . . . . . . . .145 D.4 Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-2 7B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .145 D.5 Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.1 8B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .146 D.6 Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.2 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff. . . . . . . . . . . . . . . . . . . . . .146 D.7 Model quality (accuracy across eight zero-shot tasks) for Gemma-3 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .147 D.8 Perplexity (↓) under different tile prior initializations. All priors yield nearly identical per- formance, suggesting that the global sparsity target allows dynamic reallocation of sparsity during training, overriding the influence of fixed initialization.. . . . . . . . . . . . . . . .147 E.1 Key notation definitions used in the experimental results (Section 7.5).. . . . . . . . . . . .148 E.2 Average zero-shot accuracy of LLaMA-2 and OPT models with4-bit weight and 8-bit input quantization.↑indicates better performance.. . . . . . . . . . . . . . . . . . . . . . . . .149 E.3 Effects of fine-tuning on the average zero-shot accuracy of LLaMA-2 and OPT models with 50% sparsity and 4-bit weight quantization.↑indicates better performance. . . . . . . . . .150 E.4 Perplexity of LLaMA-2 and OPT models with2:4 sparsity and 4-bit weight quantization on WikiText-2 dataset language modeling task.↓indicates better performance.. . . . . . .150 xiii E.5 Perplexity of LLaMA-2 and OPT models with4-bit weight and 8-bit input quantization.↓ indicates better performance.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .151 E.6 Perplexity of LLaMA-2 and OPT models withunstructuredsparsityand4-bitweightquan- tizationon WikiText-2 dataset language modeling task.↓indicates better performance.. . .151 E.7 Perplexity of LLaMA-2 and OPT models with pruning on WikiText-2 dataset language mod- eling task. The quantization is disabled in this experiment.↓indicates better performance..152 E.8 Perplexity of LLaMA-2 and OPT models with quantization on WikiText-2 dataset language modeling task. The sparsity is disabled in this experiment.↓indicates better performance. .153 E.9 Average zero-shot accuracy of LLaMA-2 and OPT models with2:4sparsityand4-bitweight quantization.↑indicates better performance.. . . . . . . . . . . . . . . . . . . . . . . . .154 E.10 Averagezero-shotaccuracyofdifferentmodelsusingdifferentpruningandquantizationschemes. ↑indicates better performance. Combining sparsity and quantization provides better accuracy results in comparison to solely using quantization.. . . . . . . . . . . . . . . . . . . . . . .154 E.11 Perplexity of different models on WikiText-2 dataset using different pruning and quantization schemes.↓indicates better performance. Combining sparsity and quantizationprovidesbetter accuracy results in comparison to solely using quantization. . . . . . . . . . . . . . . . . . .155 E.12 The required time for fine-tuning the models with a single H100 GPU on 300,000 tokens from the C4 dataset with a batch size of 64 and a sequence length of 1024.. . . . . . . . . . . . .157 E.13 Theoretical memory reduction (×) of different compression methods across various OPT and LLaMA models. In Quantized SLIM , the low-rank adapters are also quantized.(↓indicates better performance.) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .157 E.14 Compute (FLOP) reduction ratios (×) of different compression methods across various OPT and LLaMA models. In Quantized SLIM , the low-rank adapters are also quantized. (↑indi- cates better performance.) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .158 E.15 The required compression time for different models and compression methods using a single H100 GPU.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .159 E.16 Perplexity of different models on WikiText-2 dataset using SLIM-LoRA with 4-bit quantiza- tion using SLIM-Quant with different calibration datasets.↓indicates better performance.. .160 E.17 Group quantization slow-down (×) on different LLaMA-2 and LLaMA-3.1 models.↓indi- cates worse.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .161 xiv List of Figures 2.1 The compute graph of a standard Transformer block, highlighting the Self-Attention and Feed- Forward Network (FFN) sub-layers.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .8 2.2 Computational time breakdown for LLaMA-3.1-8B during training and inference, as profiled on a single NVIDIA H100 GPU. The ”linear” component is the largest single bottleneck. The batch size, input sequence length, and generation sequence length are set to 4, 1024, and 1024 respectively.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 2.3 Roofline Model comparison between NVIDIA A100 and H100 GPUs. The H100 (orange) of- fers significantly higher peak compute (FLOPs). However, memory bandwidth has not scaled proportionally, shifting the ”knee” of the curve to the right. This implies that algorithms on H100 require a higher arithmetic intensity to escape the memory-bound region compared to the A100. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .11 3.1 MKOR for layermon a single worker. The inputs of MKOR are the activationsA m t , the gradients of the loss function with respect to the inputsG m t , and the gradients of the loss function with respect to the weights∇ W m L. The output is the update values∆W m . . . . .20 3.2 The pre-training loss of BERT-Large-Uncased using different optimizers.. . . . . . . . . .26 3.3 Test accuracy of ResNet-50 on ImageNet for MKOR, KAISA, and SGD on 64 GPUs.. . . .27 3.4 The sensitivity of MKOR and KAISA for BERT-Large-Uncased and an Autoencoder model (a) and the effect of inversion frequency on the convergence properties of these models (b)..28 3.5 Per-step breakdown of different optimizers on BERT-Large-Uncased (a) and ResNet-50 (b). The times reported in these graphs reflect only the optimizer computations. The majority of the training time is spent on the model’s forward and backward passes, which are identical across all optimizers and are not included here. . . . . . . . . . . . . . . . . . . . . . . . .29 3.6 Rank-1 error for activation and input gradient covariance matrices for BERT-Large-Uncased pre-training (a, b) and ResNet-50 on ImageNet (c, d). . . . . . . . . . . . . . . . . . . . . .30 4.1 The sparse training pipeline in SLOPE. Here,X,Y, andWdenote the input, output, and the weight tensors for a specific layer, respectively.∇ · Lrepresents the gradient of the loss function. L and R are the low-rank terms that are introduced only in the final 1% iterations. SuperscriptRshows row-wise pruning usingN:Mscheme andR, Cshows both column and row-wiseN:Msparsification, leading to extra imposed zeros. Blue elements represent non-zero values, while white elements represent pruned values, and red elements indicate additional zeros introduced during the backward pass. . . . . . . . . . . . . . . . . . . . .34 xv 4.2 Validation perplexity of GPT2-Small and GPT2-Large on OpenWebText.γ w shows the value of the decay factor parameter in Extended SR-STE (FST).. . . . . . . . . . . . . . . . . . .42 4.3(a)The speedup achieved using cuSPARSELt backend in PyTorch for Attention (d out =d in ), Upsample (d out = 4d in ) and Downsample (d out = d in 4 ) matrices with a batch size of 2048. (b)The cosine similarity of the low-rank adapters and the converged adapters for different layers in the model. The cosine similarities are averaged among the 24 layers of BERT-Large- Uncased.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 4.4 The speedup achieved by low-rank adapters in comparison to a dense matrix-multiplication.45 5.1 OPTIMA generates a shared Hessian among the different columns of the pruned weight using a small calibration dataset. Then, the weights in different columns will be updated in parallel using a QP solver and the shared Hessian.. . . . . . . . . . . . . . . . . . . . . . . . . . .50 5.2 Relative error reduction on OPTIMA in comparison to Wanda, SparseGPT, and Thanos for LLaMA-3.2 1B.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .56 6.1 Illustration of the PATCH learning process for generating tile-level hybrid masks. Each tile is parameterized by a learnable distribution and sampled with Gumbel Softmax to produce ̃ M tile . The dense probability is expanded and merged with a 2:4 mask ̃ M 2:4 , which can be fixed or jointly learned during training, yielding ̃ M. The final mask assigns each tile to remain dense or follow the 2:4 pattern, enabling flexible sparsity across the weight matrix.. . . . . . . . .69 6.2 Layer-wisesparsityallocationunderdifferentglobalsparsitybudgetsforvariousmodels. PATCH achieves the target global sparsity while flexibly distributing pruning across transformer layers. 77 6.3 Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Qwen-2.5 0.5B.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .78 6.4 Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Gemma-3 1B.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .79 6.5 Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in LLaMA-3.2 1B.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .80 7.1 The SLIM weight compression pipeline consists of three main steps: (1) Quantizing weights using the symmetric SLIM-Quant algorithm, producing quantized weightsW Q and quantiza- tion errorE Q ; (2) Sparsifying quantized weightsW Q through a pruning method, resulting in compressed weightsW C and sparsity errorE S ; (3) Mitigating compression errors through SLIM saliency-based low-rank approximation, generating left and right low-rank adaptersL andR. Optionally, these adapters can be fine-tuned with sparse quantized weights frozen to further enhance model accuracy.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .85 7.2 Accuracy results of the OPT family across different compression methods (↑indicates bet- ter performance). At equal parameter size, SLIM outperforms both dense models and other compression techniques, demonstrating that model compression with SLIM yields superior performance under the same budget. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .94 A.1 Approximations in second-order methods.. . . . . . . . . . . . . . . . . . . . . . . . . . .116 xvi A.2 Maximum and minimum eigenvalues (a) and the condition number (b) of the right factors in KFAC when training ResNet-50 on CIFAR-10. As illustrated, the minimum eigenvalues of the factors in KFAC approach zero, meaning that the factors are singular, and hence have large condition numbers, making numerical inversion of them complex and numerically unstable..122 A.3 Scalability of MKOR.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .123 A.4 Average covariance rank-1 approximation error for ResNet-50 in different iterations. . . . .123 A.5 Training time for distributed first- and second-order optimizers SGD, MKOR, KAISA, and HyLo onBERT-Large-CasedonIMDB(a),BERT-Base-CasedonSQuAD(b), andAlexNeton CIFAR-100(c). In all the experiments, MKOR outperforms other optimizers in convergence speed.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .124 A.6 Training accuracy vs. the number of epochs for distributed first- and second-order optimizers SGD, MKOR, KAISA, and HyLo onBERT-Large-CasedonIMDB(a),BERT-Base-Casedon SQuAD(b), andAlexNetonCIFAR-100(c). In all the experiments, MKOR outperforms other optimizers in convergence rate. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .125 B.1 Average mask difference between each iteration and the converged sparsity pattern in GPT2- Small pretraining using SR-STE. The highlighted area shows the ratio of the resources used for updating weights that are pruned and not used in the inference of the model.. . . . . . .127 B.2 The setup and multiplication time for square matrices using the cuSPARSELt SpMM backend.128 B.3 Training loss of BERT-Large-Uncased on WikiCorpus dataset for phase 1 and 2.. . . . . . .129 B.4 The imposed sparsity ratio when pruning the weight matrices in the backward pass.. . . . .130 B.5 Validation perplexity on GPT2-Small pretraining for 100,000 iterations for different matrix pruning settings. Pruning the output gradients leads to divergence within a few iterations and hence is not reported.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .131 B.6 Comparison of the loss of depth and width pruning methods.. . . . . . . . . . . . . . . . .139 C.1 Sensitivity analysis for the number of calibration samples for different pruning methods.. .141 E.1 SLIM speedup for LLaMA-2 family of models on NVIDIA A100-40GB GPUs.. . . . . . .156 E.2 Sensitivity analysis for the rank of the adapter (a) and the number of calibration samples (b) for different one-shot compression methods. For Naive-LoRA and SLIM-LoRA, we have used the SLIM-Quant quantization method, and for the SparseGPT, we have used the Group quantization version of OPTQ.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .160 E.3 Sparsity analysis on LLaMA-2-13B model using perplexity on WikiText-2 dataset.↓indicates better performance.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .161 xvii Chapter 1 Introduction Large Language Models (LLMs) have become foundational tools in modern artificial intelligence, demonstrat- ing remarkable capabilities in text generation, reasoning, and multi-modal tasks [14,31,139]. However, these capabilities come at a significant cost. Training state-of-the-art models consumes enormous computational, memory, and environmental resources [ 115,14], and deploying them for inference remains a major challenge due to their massive memory footprint and high computational demands. To mitigate these overheads, various model compression techniques have been proposed [58]. Historically, these methods have often been applied in isolation, focusing on a single tool such assparsity(removing parameters) [ 36],quantization(reducing parameter precision) [37], orlow-rank approximations(factoring parameter matrices) [ 61]. This isolated approach is inherently limiting, as it fails to address the multi-faceted nature of the efficiency bottleneck. This thesis argues that these methods must be applied jointly. We introduce a conceptual framework, which we term theCompression Trinity, built upon the three fundamental pillars of sparsity, quantization, and low-rank approximations. These methods are highly complementary. While all three pillars contribute to reducing overall memory and compute overheads, they play distinct roles in balancing hardware constraints with model expressivity. Specifically, sparsity and quantization directly target the primary bottlenecks of modern hardware: sparsity reduces thecomputationalload, while quantization minimizesmemory band- widthrequirements. Low-rank approximations, in turn, act as the critical algorithmic counterweight. By efficiently projecting parameters into lower-dimensional spaces, they restore lost accuracy and model capac- ity without reintroducing the hardware overheads that sparsity and quantization eliminated. We demonstrate that the joint application of this Compression Trinity across the different phases of an LLM’s life-cycle is the key to unlocking new frontiers of efficiency. To understand the context for these contributions, we must first consider the distinct stages of an LLM’s development and deployment. 1.1 LLM Life-cycle Stages The life-cycle of a large language model can be broadly divided into two primary stages 1 : 1 Other stages exist, such as Supervised Fine-Tuning (SFT) and alignment (e.g., RLHF, DPO). We focus on pretraining and inference as they represent the primary computational bottlenecks in the LLM life-cycle. 1 CHAPTER 1. INTRODUCTION2 •Pretraining:This is the mostcomputationallyexpensivephase, where the model is trained from scratch on massive, web-scale datasets to learn general-purpose language representations [119,30]. •Inference:This is the deployment stage, where the trained model is used to generate predictions for new user inputs. In many real-world applications, inference must be performed with low latency and on resource-constrained hardware. Each stage presents unique opportunities for acceleration. During thepretrainingphase, this can be achieved by either accelerating the training process itself or by improving theoptimizerto converge faster. For theinferencestage,post-trainingcompressioncan be applied to make the deployed model more efficient. These opportunities, though different on the surface, share a common computational bottleneck, dense matrix- matrix multiplications, and can therefore share the same fundamental principles for compression, which we will cover next. 1.2 Compression Techniques ThecoreofourapproachreliesonthethreemaintechniquesoftheCompressionTrinity: sparsity, quantization, and low-rank approximation.Let us now define the three pillars of the Compression Trinity. 1.2.1 Sparsity Sparsity is a compression technique that involves identifying and removing (i.e., setting to zero) the least important weights in a neural network, thereby reducing the total number of parameters and floating-point operations (FLOPs) [ 75,55]. This technique is broadly categorized into two types.Unstructuredsparsityremoves arbitrary, individual weights from a matrix. While it offers high flexibility and can often achieve high compression ratios with minimal accuracy loss, it is notoriously difficult to accelerate in practice. Its irregular memory access patterns are nothardware-friendly, meaning they cannot be executed efficiently on modern parallel accelerators like GPUs, which are designed to process data in large, regular chunks [ 151]. Conversely,structured sparsity removes entire blocks of parameters, such as full rows, columns, or filter channels [67,84]. This regular structure is hardware-friendly but often damages model accuracy significantly, as it removes entire features indiscriminately. A new family ofsemi-structured sparsitypatterns has emerged to bridge this gap. A prominent example isN:M sparsity, which enforces thatNout of everyMconsecutive weights are non-zero (e.g., 2:4 sparsity) [ 137,62]. This pattern is flexible enough to preserve accuracy while being regular enough for hardware acceleration on modern GPUs [108]. However, finding the optimal sparsity pattern and updating the remaining non-zero weights to compensate fortheremovedelementsremainsachallengingtask[ 35,36]. Toillustratetheseverityofthischallenge, wecan comparesparsityagainstquantization.Table1.1demonstratesthatevenatamodest2xcompressionratio(50% sparsity), removing connections damages the model more than aggressive 4-bit quantization, which offers 4x compression. This counter-intuitive result, that a method providinglesscompression yieldsloweraccuracy, highlights that sparsity is the most destructive pillar of the Trinity, necessitating the dedicated optimization strategies we propose in Chapter 5andChapter 6. CHAPTER 1. INTRODUCTION3 Table 1.1:The Sparsity Paradox:Comparison of LLaMA-2-7B zero-shot accuracy under standard Post- Training Compression. Sparsity yields lower compression ratios (2x) yet results in significantly higher accu- racy loss compared to Quantization (4x), highlighting its destructive nature. MethodCompression Ratio Bit-width / Density Avg Accuracy Dense Baseline1xFP16 / 100%54.61% 4-bit Quantization * 4xINT4 / 100%53.63% 2:4 Sparsity † 2xFP16 / 50%48.62% * Best among AbsMax and OPTQ † Best among SparseGPT, Wanda, Thanos, ProxSparse, and MaskLLM 1.2.2 Quantization Quantization reduces the numerical precision of the numbers used to represent the model’s weights and, in some cases, activations. For example, parameters are typically trained in 32-bit (FP32) or 16-bit (FP16/BF16) floating-point formats, but quantization can compress them down to 8-bit integers (INT8), 4-bit integers (INT4), or even lower bit-widths [27,37]. This reduction in precision has two primary benefits: it saves memory (e.g., 4-bit quantization yields an 8x memory reduction over 32-bit weights) and allows for the use of faster, specialized compute units (like INT8 tensor cores) that can perform integer arithmetic much faster than floating-point operations. The main challenge in quantization is that the representation capabilities of the numbers are reduced expo- nentially with the bit-width. This can lead to large accuracy degradation in low bit-width schemes, especially in the presence of outlier values, which are common in LLMs [ 81]. Finding the best way to map high-precision weights to a low-bit representation and updating those weights to minimize the resulting error is a non-trivial task [37]. As with sparsity, only relying on quantization limits the total compression ratio of the models, and other methods should be combined with it to push efficiency further. 1.2.3 Low-rank Approximations Low-rank approximations are based on the observation that the large weight matrices in LLMs are often over- parameterized and have a low ”intrinsic rank.” This redundancy can be exploited by decomposing a large weight matrixW∈R m×k into the product of two smaller, ”thin” matrices,L∈R m×r andR∈R r×k , where the rankr≪m, k. This technique, popularized by Low-Rank Adaptation (LoRA) [ 61], can reduce the memory and compute overhead of LLMs significantly. However, low-rank approximations are not very effective in compressing matrices that are inherently high- rank, and applying them too aggressively can lead to significant errors. Relying on low-rank approximation alone is insufficient, as it cannot capture the fine-grained, high-rank information that sparsity or high-precision quantization can preserve. 1.3 The Failure of Isolated Compression As mentioned inSection 1.1, modern deployment scenarios impose strict constraints on memory and latency. While the individual pillars of compression—sparsity and quantization—are well-studied, existing literature often treats them as independent solutions. CHAPTER 1. INTRODUCTION4 However, empirical evidence suggests that pushing any single pillar to the extreme results in catastrophic accuracy loss.Table 1.2illustrates this breakdown on the LLaMA-2-7B model. When we attempt to achieve an 8x compression ratio using only quantization (2-bit) or only sparsity (87.5%), the model’s reasoning capa- bilities collapse, with accuracy dropping to near-random chance (≈31%). Table 1.2: Average Zero-shot Accuracy of LLaMA-2-7B on 8 tasks (MMLU, PIQA, ARC-Easy, ARC- Challenge, WinoGrande, OpenBookQA, RACE, HellaSwag) at 8x compression ratio. Single-pillar methods fail to retain capabilities, while multi-pillar methods (Compression Trinity) recover accuracy. MethodAverage Accuracy Dense Baseline (FP16)54.61% Single-Pillar Compression (8x Ratio) 2-bit Quantization * 31.81% 87.5% Unstructured Sparsity † 31.24% Multi-Pillar Compression (8x Ratio) 4-bit Quantization + 2:4 Sparsity47.97% 4-bit Quantization + 50% Unstructured Sparsity52.38% * Best among AbsMax and OPTQ † Best among SparseGPT, Wanda, and Thanos In contrast, the table demonstrates that ahybrid approach, combining moderate 4-bit quantization with moderate 50% sparsity, recovers the majority of the accuracy (52.38%) while achieving the same compression ratio. This observation forms the core motivation of this thesis: compression is not a singular optimization problem, but a multi-dimensional balancing act. 1.4 Thesis Contributions and Roadmap As discussed inSection 1.2, each of the compression techniques, sparsity, quantization, and low-rank approx- imation, cannot effectively compress LLMs alone and will hit an accuracy or efficiency wall at some point. This thesis argues and demonstrates that thejoint applicationof the Compression Trinity is the key to un- locking new frontiers of efficiency. We show that by combining hardware-friendly sparsity and aggressive quantization, we can achieve massive reductions in compute and memory. We then use low-rank approxima- tions as a controllable, highly efficient method to add back a small number of dense parameters, compensating for the joint compression error and restoring model accuracy. Before detailing the novel methods that prove this thesis, Chapter 2will first provide a comprehensive technical background. The thesis is structured to explore the Compression Trinity across the full life-cycle of an LLM. We begin by applying the Trinity to the Pretraining phase (Chapter 3andChapter 4). We then perform a deep dive into the Sparsity pillar, the most destructive component of the Trinity, exploring both layer-wise ( Chapter 5) and end-to-end (Chapter 6) optimization regimes. Finally, we integrate all three pillars for a holistic Post-Training Compression solution (Chapter 7). These contributions are detailed as follows: •Chapter 3MKOR:We lay the foundation for the Compression Trinity in the expensive pretraining phase. MKOR (Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates) is a novel second-order optimizer that leverages all three pillars. It approximates the second-order informa- tion as a block-diagonalsparsematrix, and then leveragesrank-1 updates(a low-rank approximation) CHAPTER 1. INTRODUCTION5 to compute the inverse of its covariance matrices. Crucially, this joint formulation is exceptionally sta- ble, allowing the inverse factors to be computed in aquantized16-bit format, whereas other methods require 32-bit numbers for numerical stability. By applying the Trinity to the optimizer’s internal com- putations, we reduce the complexity of second-order updates from O ( d 3 ) to O ( d 2 ) and communication fromO(d 2 )toO(d), wheredis the hidden dimension of the model. As a result, MKOR accelerates pretraining by outperforming state-of-the-art optimizers like KFAC by up to 1.85x on BERT-Large train- ing. •Chapter 4SLOPE:We further accelerate the pretraining phase by applying the Compression Trinity directly to the linear layers. SLOPE (A Double-Pruned Sparse Plus Lazy Low-rank Adapter Pretraining method) introduces a framework that jointly applies sparsity and low-rank approximations from the start. It accelerates sparse pretraining by introducing a novel double-pruned backward pass that enables N:M sparsity acceleration in both forward and backward passes. To recover accuracy lost from sparsity, we introduce low-rank adapters only during the final 1% of pretraining iterations. This ”lazy” insertion minimizes overhead while maximizing accuracy. By creating a base model that is already sparse and low-rank, SLOPE produces a model that is not only efficient for inference (1.34x speedup) but is also an ideal and stable candidate for the final pillar, quantization, which we apply in the post-training stage. This approach accelerates end-to-end training and inference of models like OPT-66B by up to 1.14x and 1.34x, respectively, and reduces training memory by 0.77x. •Chapter 5OPTIMA:We transition to the post-training phase, addressing the strict scenario where end-to-end training is infeasible. OPTIMA (Optimal One-shot Pruning via Quadratic Programming) focuses on perfecting thesparsitypillar under a ”zero-training” constraint. It formulates the one-shot, post-pruning weight update as a series of independent, row-wise Quadratic Programs (QPs) that share a common layer Hessian. This allows us to find the per-row globally optimal update given the Hes- sian, minimizing the reconstruction error from pruning. By creating the most accurate and stable sparse model possible, OPTIMA serves as a critical enabling step, producing a high-fidelity model that can withstand the subsequent application of aggressive quantization and low-rank approximations. OPTIMA acts as a drop-in replacement for the update step in methods like Wanda or SparseGPT, con- sistently improving zero-shot performance by up to 3.97% absolute accuracy without any fine-tuning. •Chapter 6PATCH:We address the scenario where afine-tuningbudget is available to push perfor- mance further. PATCH (Pruning with a Learnable Tile-level Configuration) overcomes the limitations of the layer-wise mask detection used in OPTIMA. It introduces a hybrid sparsity framework that par- titions weight matrices into tiles, assigning each tile to be eitherdenseor2:4 sparsevia a learnable mask. This creates a continuous, effective sparsity ratio between 0% and 50%, balancing accuracy in critical regions with acceleration elsewhere. While PATCH focuses on enhancing this single pillar, it is designed to be fully composable with the other pillars of the Trinity. We explicitly demonstrate that it can be combined with joint quantization and low-rank approximation methods, such as SLIM, to further boost the accuracy and stability of the final, jointly compressed model. On LLaMA-2 7B, the PATCH framework alone delivers 1.18x-1.38x end-to-end speedup while improving accuracy by up to 2.96% compared to state-of-the-art 2:4 pruning. •Chapter 7SLIM:We present the complete fulfillment of the Compression Trinity in a one-shot setting. SLIM (One-shot Quantization and Sparsity with Low-rank Approximation) holistically integrates all CHAPTER 1. INTRODUCTION6 three pillars to solve the ”compounded error” problem. It applies aggressive, hardware-friendly quan- tization and semi-structured sparsity, then compensates for the combined error using a novel saliency function that allows us tomathematically computeoptimal low-rank adapter values in one shot. This joint approach improves accuracy by up to 5.66% on LLaMA-2-7B (combining 4-bit quantization and 2:4 sparsity) and achieves up to 4.3x layer-wise speedup. Finally,Chapter 8concludes the thesis by summarizing its key findings, reflecting on the impact of the Compression Trinity framework, and discussing potential avenues for future research. The code, pre-trained checkpoints, and interactive visualizations for the methods presented in this thesis are centralized at our re- search hub 2 . 2 https://w.cs.toronto.edu/~mmozaffari/compression-trinity/ Chapter 2 Background This chapter establishes the technical foundations necessary to understand the contributions of this thesis. We begin by identifying the computational bottleneck that dominates modern LLM workloads: the dense linear layers within the transformer architecture (Section 2.1). We then examine the hardware constraints that gov- ern the execution of these layers on modern GPUs, introducing the memory hierarchy, specialized compute units, and the Roofline model that together determine whether a workload is limited by compute or by memory bandwidth (Section 2.2). With both the algorithmic bottleneck and the physical constraints established, we formalize the two orthogonal strategies for acceleration, reducing the number of training iterations through better optimization and reducing the cost of each iteration through hardware-aware compression ( Section 2.3). We explore each strategy in turn: Section 2.4surveys the landscape of first- and second-order optimizers, mo- tivating the need for structured approximations that make curvature information tractable at scale; Section 2.5 characterizes the distinct computational regimes of LLM training and inference, revealing why different life- cycle stages demand different compression techniques. Finally,Section 2.6andSection 2.7introduce the Compression Trinity as the unifying solution framework for this thesis and argue that the joint application of sparsity, quantization, and low-rank approximations is essential to overcome the limitations of any single technique applied in isolation. 2.1 The Transformer Bottleneck: Linear Layers The Transformer has become the foundational architecture for virtually all modern Large Language Models (LLMs), demonstrating unparalleled scaling and performance on a wide array of language tasks [146]. In practice, an LLM is constructed by stacking a large number of identical Transformer blocks, often numbering in the dozens or even hundreds. This repetitive, block-based design means that the computational profile of a single block is representative of the model’s entire computational load. Understanding the specific operations within this block is therefore the first step toward identifying the primary opportunities for optimization. A standard Transformer block, as illustrated in Figure 2.1, is composed of two primary sub-components. The first is the Self-Attention mechanism, which allows the model to weigh the importance of different tokens in a sequence relative to each other. The second is a position-wise Feed-Forward Network (FFN), which is typically a two- or three-layer Multi-Layer Perceptron (MLP) that provides the majority of the model’s representational capacity. In modern architectures, this FFN often takes the form of a SwiGLU variant [ 130]. These two components work in tandem, with the attention mechanism handling the aggregation of sequential 7 CHAPTER 2. BACKGROUND8 Linear Linear Linear Softmax Linear Linear Linear Out Gate Down Linear Up Self AttentionFeed Forward (MLP) Figure 2.1: The compute graph of a standard Transformer block, highlighting the Self-Attention and Feed- Forward Network (FFN) sub-layers. information and the FFN processing that information at each token’s position. Connecting this architectural graph to its underlying mathematical operations reveals a critical insight: both sub-components are dominated by linear layers, which are implemented as dense matrix-matrix multipli- cations (GEMMs). The Self-Attention mechanism, for instance, computes its Query (Q), Key (K), and Value (V) representations through three independent linear layers, and a final Output (O) projection is applied after the attention scores are aggregated. Similarly, the SwiGLU FFN is composed of an Up-projection layer, a Gate layer, and a Down-projection layer. Consequently, the vast majority of computations and parameters in the Transformer block are contained within these dense matrix multiplications. Each linear layer, parameterized by a weight matrixW∈R d out ×d in , participates in three distinct ma- trix multiplications during a single training step. Given an input activation matrixX∈R b×d in , whereb is the effective batch size (batch size×sequence length), theforward passcomputes the layer’s output as Y=XW T . During backpropagation, two additional GEMMs are required: theweight gradientcomputa- tion,∇ W L= (∇ Y L) T X, which determines how the weights should be updated, and theinput gradient computation,∇ X L= (∇ Y L)W, which propagates the error signal to the preceding layer. Crucially, while the forward pass multiplies the input byW T , the input gradient computation multiplies byWitself. This transpose relationship poses a fundamental challenge for structured compression: a sparsity pattern that is hardware-friendly along the rows ofW(as required for the forward pass) may not be hardware-friendly along its columns (as required for the backward pass). This asymmetry is a central obstacle that we address directly in Chapter 4. This architectural analysis is confirmed by empirical performance profiling, as shown inFigure 2.2. The ”linear” component is by itself the largest single computational bottleneck, consuming approximately 51.8% of the total training time for a model like LLaMA-3.1-8B. The ”Attention” component, which contains the self-attention computations with dynamic, data-dependent operations excluding the linear layers, accounts for another 9.6%. While the Attention mechanism’s unique properties present their own optimization chal- lenges, this thesis will focus on the linear layers. As the largest and most dominant bottleneck, composed of standard static-weight GEMMs, these linear layers represent one of the most critical and impactful targets for optimization. The evidence from both the architectural design and the performance profile establishes a clear conclusion: accelerating the dense linear layers is the central challenge for LLM efficiency. However, identifying the mathematical operations is only half the picture. To understandwhythese operations become bottlenecks, one must analyze the physical constraints of the hardware on which they execute. The following section will CHAPTER 2. BACKGROUND9 Linear 51.8% Attention 9.6% Other 38.7% Training Linear 34.3% Attention 51.0% Other 14.6% Inference Figure 2.2: Computational time breakdown for LLaMA-3.1-8B during training and inference, as profiled on a single NVIDIA H100 GPU. The ”linear” component is the largest single bottleneck. The batch size, input sequence length, and generation sequence length are set to 4, 1024, and 1024 respectively. introduce the fundamental principles of GPU architecture and the Roofline model that govern the performance of these linear layers. 2.2 Hardware Constraints and the Roofline Model While the Transformer architecture defines the operations to be performed, the execution speed is dictated by the underlying hardware. Modern LLMs are trained and deployed almost exclusively on Graphics Processing Units (GPUs), which are massive throughput-oriented processors. To understand the efficiency bottlenecks discussed in this thesis, we must first establish a model of how these devices process data, specifically focusing on the memory hierarchy, specialized compute units, and the theoretical limits defined by the Roofline model. 2.2.1 Memory Hierarchy and Tensor Cores The computational pipeline of a GPU is constrained by two primary resources:Memory Bandwidth(how fast data can be moved) and Compute Throughput (how fast data can be processed). Memory Hierarchy:Data movement in a GPU follows a hierarchy of speed and capacity. The bulk of model parameters and activations reside in High Bandwidth Memory (HBM), which offers high capacity (e.g., 80GB on an NVIDIA H100) but relatively high latency and limited bandwidth compared to on-chip memory. To perform an operation, data must be moved from HBM to the smaller, faster L2 cache, and finally to the Streaming Multiprocessors’ (SM) shared memory and registers (SRAM). The cost of moving data from HBM is orders of magnitude higher than moving it within the chip. Consequently, algorithms that reuse loaded data multiple times (high data locality) are significantly more efficient than those that stream data for single use. CHAPTER 2. BACKGROUND10 Tensor Cores vs. CUDA Cores:Historically, GPUs relied on general-purpose ”CUDA Cores” for arith- metic. However, the rise of Deep Learning necessitated specialized hardware. Modern architectures (e.g., NVIDIA Ampere and Hopper) featureTensor Cores, specialized execution units designed essentially to per- form one operation: matrix-multiply-and-accumulate ( D = A × B + C ) in a single clock cycle. While standard CUDA cores operate on scalars or small vectors, Tensor Cores consume entire8×16or 16×16matricesperinstruction[105,106]. ThisspecializationallowsTensorCorestodeliverthroughputsover an order of magnitude higher than CUDA cores. For example, on the NVIDIA H100, Tensor Cores introduce support for the FP8 data format, doubling the peak throughput compared to the standard BF16 [ 148] format used in previous generations like the A100. This massive disparity means that any algorithm not utilizing Tensor Cores efficiently leaves the vast majority of the GPU’s potential performance on the table. 2.2.2 The Roofline Model The interaction between the memory bandwidth and compute throughput is formalized by theRooflineModel [152]. This model visualizes the theoretical peak performance of an algorithm based on itsArithmetic In- tensity(I), defined as the number of floating-point operations (FLOPs) performed for every byte of data transferred from memory: I= Total FLOPs Total Bytes Transferred (2.1) As illustrated inFigure 2.3, the Roofline model divides performance into two distinct regions: •Memory-Bound Region (Sloped Line):When arithmetic intensity is low, the GPU’s compute units starve while waiting for data from HBM. In this region, performance is strictly limited by memory bandwidth. Improving performance here requires moving fewer bytes (e.g., via Quantization). •Compute-BoundRegion(FlatLine):Whenarithmeticintensityishigh, dataisloadedonceandreused many times, keeping the compute units fully saturated. In this region, performance is limited by the peak FLOPs of the Tensor Cores. Improving performance here requires doing fewer operations (e.g., via Sparsity). The ”knee” of the curve represents the transition point. Figure 2.3highlights a critical trend in hardware evolution: the gap between compute and memory is widening. The NVIDIA H100 boasts significantly higher peak FLOPs than the A100 (especially with FP8), pushing the ”roof” higher. However, memory bandwidth has scaled much more slowly. This shifts the knee to the right, meaning that modern GPUs require increas- ingly higher arithmetic intensity to achieve peak utilization. This hardware reality fundamentally shapes the strategies required for optimization: Pretraining (high intensity) sits firmly under the compute roof, while Inference Decoding (low intensity) is trapped under the memory slope. 2.2.3 Sparse Tensor Core Acceleration While standard Tensor Cores provide massive throughput for dense matrix multiplications, they perform re- dundant calculations if the weight matrix contains zeros. To address this, modern NVIDIA architectures introducedSparse Tensor Coresdesigned to accelerate fine-grained structured sparsity, specifically the 2:4 pattern. CHAPTER 2. BACKGROUND11 10 −1 10 0 10 1 10 2 10 3 10 4 10 5 10 6 Arithmetic Intensity (FLOPs/Byte) 10 0 10 1 10 2 10 3 Performance (TFLOPS) 312 TFLOPS 2.0 TB/s 989 TFLOPS 3.4 TB/s →Compute Bound Memory Bound← Roofline Model: NVIDIA GPUs (FP16 Tensor Core, Dense) A100 (80GB SXM) H100 (80GB SXM) Figure 2.3: Roofline Model comparison between NVIDIA A100 and H100 GPUs. The H100 (orange) of- fers significantly higher peak compute (FLOPs). However, memory bandwidth has not scaled proportionally, shifting the ”knee” of the curve to the right. This implies that algorithms on H100 require a higher arithmetic intensity to escape the memory-bound region compared to the A100. MechanismandFormat:The2:4structuredsparsityformatenforcesarigidconstraint: ineverycontiguous block of 4 elements along the reduction dimension, at least 2 elements must be zero. This allows the hardware to compress the sparse matrix by 50%, storing only the non-zero values in a packed format alongside small metadata indices that record the original positions of the retained elements. During execution, the Sparse Tensor Core reads these indices to select only the relevant entries from the corresponding dense operand, skipping all multiplications that would involve a zero. Because the core performs the same operation in the same number of clock cycles but processes only half as many non-zero elements, it effectively doubles the theoretical peak throughput compared to the equivalent dense precision. The Backward Pass Challenge:While 2:4 sparsity offers massive theoretical gains, applying it to training is non-trivial. The Sparse Tensor Core requires the sparsity to exist along thereduction dimension. In the forward pass, this aligns naturally. However, in the backward pass, we must multiply by the transpose of the weights, misaligning the sparsity pattern relative to the hardware’s expected read order. This structural alignment challenge is a primary motivation for the custom pretraining strategies developed inChapter 4 (SLOPE). 2.3 The Solution Space: Two Frontiers of Acceleration Given the dominance of Linear layers identified inSection 2.1and the rigid hardware constraints defined in Section 2.2, we can now formalize the optimization landscape. To minimize the total time required to train an CHAPTER 2. BACKGROUND12 LLM or process a request, we can decompose the cost into two multiplicative factors: Total Time= (Number of Iterations) |z Sample Efficiency ×(Time Per Iteration) |z Hardware Efficiency (2.2) This decomposition reveals two orthogonal frontiers for acceleration, each requiring distinct strategies: 1.Strategy 1: Sample Efficiency (Reducing the Number of Iterations).This strategy is exclusive to the training regime. If we can improve the optimization algorithm to converge in fewer steps, we reduce the total time even if the cost of each step remains constant. This motivates the study of advanced optimizers. 2.Strategy 2: Hardware Efficiency (Reducing the Time Per Iteration).This strategy applies to both training and inference. To reduce the cost of a single iteration, we must attack the hardware bottlenecks identified in the Roofline model: we must either reduce the FLOPs (for compute-bound operations) or reduce the memory traffic (for memory-bound operations). The following sections will explore these two frontiers in detail, starting with Algorithmic Efficiency via advanced optimizers, followed by Hardware Efficiency via the different operating regimes of LLMs. 2.4 Strategy 1: Algorithmic Efficiency Having defined the physical arena in which our models execute, we turn to the first frontier of acceleration: SampleEfficiency. If the hardware limits how fast we can compute a single update, we must strive to compute fewer updates overall. This brings us to the domain of optimization algorithms, which dictate the path a model takes through the loss landscape from random initialization to convergence. 2.4.1 First-Order Methods Standard LLM training relies almost exclusively on first-order optimization methods, particularly adaptive gradient variants such as Adam [70] and AdamW [86]. These methods utilize the gradient vectorg=∇L(θ), which represents the slope of the loss function. Geometrically, first-order methods approximate the loss land- scape locally as a hyperplane (a linear approximation). While computationally efficient—requiring onlyO(d)memory and compute fordparameters—this linear approximation ignores thecurvatureof the landscape. In the highly non-convex and ill-conditioned optimiza- tion landscapes typical of deep neural networks, this blindness to curvature leads to two primary inefficiencies. In ”narrow valley” regions where the curvature is high in one direction and low in another, first-order methods tend to oscillate across the valley rather than moving down its floor, wasting iterations. Additionally, without knowledge of the local scale (curvature), determining the optimal step size (learning rate) is difficult, often requiring extensive tuning and warm-up schedules. 2.4.2 Second-Order Methods Second-order methods address these limitations by incorporating the Hessian matrixH=∇ 2 L(θ), which containsthesecond-orderpartialderivatives. ByusingtheHessian, thesemethodsapproximatethelosslocally CHAPTER 2. BACKGROUND13 as a quadratic function (a ”bowl”) rather than a plane. This allows for the computation of the Newton step: ∆θ=−H −1 g(2.3) The Newton step uses the inverse Hessian to automatically rescale the gradient. It takes larger steps in direc- tions of low curvature (flat plateaus) and smaller, more cautious steps in directions of high curvature (steep cliffs). Theoretically, this property, known as affine invariance, allows second-order methods to converge in significantly fewer iterations than their first-order counterparts. 2.4.3 The Computational Barrier and Structured Approximations Despite their theoretical superiority, exact second-order methods are intractable for LLMs due to the sheer size of the Hessian. For a model withdparameters,His ad×dmatrix. For a modest 7B parameter model, storing Hwould require exabytes of memory, and inverting it (O(d 3 )) is computationally impossible. This presents a classic efficiency trade-off: second-order methods offer high sample efficiency but suffer from catastrophic hardware inefficiency. To make second-order optimization feasible, we must rely onStructured Approximations. The most prominent approach in deep learning is Kronecker-Factored Approximate Curvature (KFAC) [ 93]. 1 KFAC approximates the Fisher Information Matrix (a proxy for the Hessian) not as a single dense matrix, but layer- wise. It assumes the Hessian for a given layer can be approximated as the Kronecker product of two much smaller matrices,H≈A⊗G, whereArelates to the layer’s inputs andGto its output gradients. Inverting this Kronecker product is efficient because(A⊗G) −1 =A −1 ⊗G −1 . This reduces the inversion cost fromO(d 3 )to the cubic size of the layer’s width (e.g.,O(width 3 )). However, even with KFAC, the memory cost of maintaining these curvature factors remains prohibitive for modern LLMs, often tripling the memory footprint compared to Adam [ 93,116]. This unsolved challenge serves as the motivation forChapter 3(MKOR). We argue that Low-Rank fac- torization (like KFAC) is merely the starting point. To truly bridge the gap, achieving the sample efficiency of second-order methods with the hardware efficiency of first-order methods, we must apply the full Com- pression Trinity to the optimizer itself, combining Low-Rank updates with Sparsity and Quantization on the optimizer states. 2.5 Strategy 2: Hardware Efficiency and LLM Regimes While advanced optimizers attack the ”Sample Efficiency” frontier to shorten training, they offer no benefit during inference, where the number of steps is fixed by the user’s generation length. To accelerate inference and to further speed up training, we must attack the second frontier:Hardware Efficiency. However, ”efficiency” is not a static target. As predicted by the Roofline model (Section 2.2.2), the bot- tleneck shifts dramatically depending on the operational regime. LLM execution is bifurcated into two dis- tinct computational profiles: theCompute-Boundregime (dominating Training and Inference Prefill) and the Memory-Boundregime (dominating Inference Decoding). 1 There are other lines of work such as Shampoo [50] and Muon [68] that approximate the Hessian matrix using alternative structured factorizations. The Compression Trinity framework is in principle applicable to these methods as well; however, we focus on the KFAC family in this thesis as a representative and widely adopted testbed for demonstrating the benefits of joint sparsity, quantization, and low-rank approximations within second-order optimization. CHAPTER 2. BACKGROUND14 2.5.1 The Compute-Bound Regimes: Training and Prefill The Training phase 2 and the Inference Prefill phase share a fundamental characteristic:massive parallelism. In Training, the model processes large batches of sequences simultaneously. Similarly, in the Prefill phase of inference (processing the user’s prompt), the model computes attention and feed-forward outputs for all input tokens at once. In these scenarios, the Arithmetic Intensity is high. The weight matrices are loaded from HBM to the chip once and reused across thousands of tokens (large batch size×sequence length). Consequently, both Training and Prefill sit firmly on the flat,Compute-Boundplateau of the Roofline model. In this region, the GPU’s compute units are fully saturated, and memory bandwidth is not the limiting factor. Therefore, optimizing Training and Prefill requires strategies that strictly reduce the total number of op- erations (FLOPs). This makesSparsitythe premier accelerator for this regime. By skipping calculations for zero-valued weights, sparsity directly lowers the compute ceiling required to process the batch, translating to faster training steps and lower prompt latency. 2.5.2 The Memory-Bound Regime: Inference Decoding Once the prompt is processed, the model enters the Decode phase, generating one token at a time. This phase is inherently sequential; the output of steptis required to compute stept+1, preventing parallelization across the sequence dimension. In this regime, the Arithmetic Intensity collapses. To generate a single token (or a small batch of tokens), the GPU must load the entire model (often 100GB+ for large models) from HBM to the chip, perform a single matrix-vector multiplication, and then discard the weights. The data reuse is minimal. Consequently, the Decode phase falls deep into theMemory-Boundslope of the Roofline model. Here, the compute units (Tensor Cores) sit idle for the vast majority of execution time, starving for data. Reducing FLOPs (via Sparsity) provides diminishing returns because computation is not the bottleneck. In- stead, acceleration is strictly defined by how fast data can be moved. This makesQuantizationthe dominant accelerator for decoding. By reducing the bit-width of the weights (e.g., from 16-bit to 4-bit), we reduce the data volume by75%, effectively quadrupling the speed at which the model can be fed to the compute units. The contrast between Training and Prefill, which require FLOP reduction, and Decoding, which requires Bandwidth reduction, reinforces the need for theCompression Trinity. No single compression technique can address both bottlenecks simultaneously: sparsity alone leaves the memory-bound decode phase untouched, while quantization alone cannot accelerate the compute-bound training and prefill phases. It is precisely the joint application of complementary techniques, sparsity to cut FLOPs, quantization to cut data movement, and low-rank approximations to recover lost accuracy, that enables a unified efficiency strategy across the entire LLM lifecycle. 2 Throughout this thesis, we use “training” and “pretraining” interchangeably when referring to the computational regime. While pretrainingtechnically denotes the initial large-scale training phase described in Section 1.1, the computational profile, large batches of dense matrix multiplications over many tokens, is shared by all training-like stages (including supervised fine-tuning). Our compression techniques apply equally to all such stages; we default to “training” when discussing hardware characteristics and “pretraining” when emphasizing the life-cycle context. CHAPTER 2. BACKGROUND15 2.6 The Compression Trinity as a Solution Framework To address the specific bottlenecks and strategies identified for both pretraining and inference, we introduce the ”Compression Trinity, 3 ” the conceptual framework for this thesis. This framework is built upon the three fundamental pillars of model compression: Sparsity, Quantization, and Low-Rank Approximations. Rather than viewing these as independent techniques, we posit that they are highly complementary tools that, when applied jointly, provide a comprehensive solution to the challenges of LLM efficiency. This section will introduce each pillar and map it to the two acceleration strategies defined inSection 2.4andSection 2.5. The first pillar,Sparsity, is the process of identifying and removing (pruning) the least important parame- ters from a weight matrixW, thereby reducing its effective size. This technique can be broadly categorized by thepatternofremovedweights:unstructuredsparsityremovesindividualweights,structuredsparsityremoves entire blocks (e.g., rows or columns), andsemi-structuredN:M sparsity enforces a fine-grained pattern, such as 2 non-zero weights out of every 4. Sparsity is the primary implementation ofCompute Bound Optimiza- tions (Reduce FLOPs). By reducing the number of non-zero parameters, it directly lowers the arithmetic cost, making it ideal for compute-bound regimes like pretraining. Simultaneously, by shrinking the storage requirement, it also assists with the Memory Bound Regime (Reduce Data Movement). The second pillar,Quantization, is the process of reducing the numerical precision of the weightsW (and sometimes activations) from high-precision formats like 32-bit floating point (FP32) to low-precision formats like 8-bit or 4-bit integers (INT8 or INT4). To maintain accuracy, especially in the presence of outlier values common in LLMs [27], this is often applied on a per-group basis. Quantization is the most powerful tool forMemory Bound Regimes (Reduce Data Movement). By reducing the bit-width ofWfrom 16 (for BF16) to 4, it cuts the memory bandwidth requirement by 75%. This directly attacks the bottleneck of the memory-bound decode phase. While specialized hardware (e.g., FP8 Tensor Cores) allows quantization to also accelerate computation, its dominant benefit remains the massive reduction in memory traffic. The third pillar,Low-Rank Approximations, is based on the hypothesis that the large weight matrices in LLMs are over-parameterized and have a low intrinsic rank [ 2]. This pillar exploits this redundancy by representing a large matrix as the product of two smaller, ”thin” matrices (e.g.,W≈LRor∆W=LR). This pillar is the most versatile of the three, acting as a bridge between Hardware Efficiency and Sample Efficiency. ForHardwareEfficiency(Strategy2), representingWwithfarfewerparameters(e.g.,d×r+r×d, where r≪d) reduces both the FLOPs and the memory footprint. Additionally, as an efficient means to mitigate accuracy degradation from sparsity and quantization, low-rank approximations provide a computationally cheap mechanism for representing fully dense matrices with unrestricted element values. Finally, forStrategy 1 (Reduce Iterations), the pillars work together. While Low-Rank approximations typically provide the mathematical structure necessary to approximate curvature (e.g., factorizing the Hessian matrix), makingtheseadvancedoptimizerscomputationallypracticaloftenrequiresthefullCompressionTrin- ity. As we will demonstrate with MKOR (Chapter 3), solely relying on low-rank structure is often insufficient for fitting complex optimizer states into GPU memory. Instead, we must apply Sparsity and Quantization 3 Other compression methods, such as Knowledge Distillation [46], in which a smallerstudentmodel is trained to replicate the be- havior of a largerteacher, is not included in this list because it operates at a fundamentally different level of abstraction: while sparsity, quantization, and low-rank approximation are all transformations applied to an existing model’s weight matrices, distillation is a training procedure that produces an entirely new model, often with a different architecture, and requires access to the full training pipeline. This places it outside the resource-constrained regime targeted by much of this thesis. That said, the two approaches are fully composable, a distilled model can serve as input to our compression pipeline, and their integration is a promising direction for future work. CHAPTER 2. BACKGROUND16 on top ofthe low-rank factors. Thus, Strategy 1 is not the domain of a single pillar, but rather the ultimate synthesis where all three pillars are combined to enable smarter, faster optimization. 2.7 The Case for a Joint Approach The core argument of this thesis is that applying these pillars in isolation is an inherently limited approach. Each technique, when pushed to its extreme, hits a fundamental wall: aggressive sparsity causes catastrophic accuracy degradation, aggressive quantization fails due to the sensitivity of outlier parameters, and low-rank approximation alone cannot capture the full rank information necessary for a model’s capabilities. However, when viewed through the lens of our identified strategies, these pillars become complementary pieces of a unified puzzle. Sparsity maximizes FLOP reduction. Quantization maximizes Bandwidth reduc- tion. Low-Rank Approximation provides a mathematical structure to recover accuracy and enable advanced optimization. It is important to clarify, however, that while theultimategoal of this thesis is the joint application of the Compression Trinity, achieving this combination requires that each individual pillar be robust enough to support the others. If a single pillar is brittle, combining it with others leads to compounded errors and performance collapse. Therefore, not all chapters in this thesis focus on the simultaneous application of all three pillars. Instead, we adopt a staged approach: we dedicate specific chapters (specifically OPTIMA and PATCH) to pushing the boundaries of theSparsitypillar in isolation. These contributions are necessary foundational steps, transforming sparsity from a fragile technique into a robust building block capable of withstandingtheadditionalpressureofjointQuantizationandLow-Rankapproximationinthefinalunification (e.g., SLIM). This thesis will demonstrate how this joint framework can be tailored to solve the distinct challenges of the LLM life-cycle: 1.Pretraining (Compute & Sample Efficiency):To solve the compute-bound pretraining problem, we advocate for a two-pronged approach. We use Strategy 1 by developing an advanced optimizer that leveragesallthreepillarsoftheTrinitytoapproximatecurvatureandaccelerateconvergence(asexplored in MKOR [99],Chapter 3). Simultaneously, we use Strategy 2 by jointly applying Sparsity and Low- Rank Approximations to the weights during training (as explored in SLOPE [ 101],Chapter 4). 2.Inference (Memory Efficiency & Accuracy):For the post-training inference problem, the goal is to holistically apply all three pillars. We first establish a stable foundation by perfecting structured sparsity (as explored in OPTIMA, Chapter 5) and adapting it for modern hardware (as explored in PATCH, Chapter 6). With this foundation, we finally demonstrate the complete fulfillment of the Trinity: a one- shot method that jointly applies Sparsity, Quantization, and Low-Rank approximation to simultaneously attack compute, memory, and accuracy recovery (as explored in SLIM,Chapter 7). In conclusion, this chapter has established the core problem (Linear layers), the physical constraints (Roofline model and Memory Hierarchy), the algorithmic opportunity (Sample Efficiency), and the Com- pression Trinity as our unified solution framework. The following chapters will now present the novel contri- butions of this thesis, demonstrating the effectiveness of this joint framework in unlocking new frontiers of efficiency. Chapter 3 MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates Publication and Contributions.The content of this chapter is based on the paper “MKOR: Momentum- Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates” [99], published at the Conference on Neural Information Processing Systems (NeurIPS), 2023. This work was conducted in collaboration with Sikan Li, Zhao Zhang, and Maryam Mehri Dehnavi. Mohammad Mozaffari conceived the algorithm, led the implementation, and designed and executed the experiments. Sikan Li assisted with conducting experiments. Zhao Zhang and Maryam Mehri Dehnavi supervised the project and contributed to the writing and revision of the manuscript. 3.1 Introduction As established inChapter 1, the pretraining phase of Large Language Models (LLMs) is dominated by massive computational costs. To mitigate this, we employ the second strategy of the Compression Trinity: reducing the total number of training iterations required for convergence. Second-order optimization methods have gained significant attention for this purpose, as they utilize curvature information, specifically the inverse of the Hessian matrix, to precondition gradients, thereby achieving higher convergence rates than first-order counterparts like SGD or Adam. However, these methods face a critical scalability wall. Since the size of the Hessian scales quadratically with the model parameters, computing and storing the exact Hessian and its inverse is computationally intractable for modern deep neural networks (DNNs), necessitating the use of approximation techniques. Beyond computational complexity, the memory footprint of optimizer states presents a prohibitive bottle- neck at the scale of LLMs. Standard first-order optimizers like Adam already require maintaining multiple state tensors (e.g., momentum and variance) for every model parameter, often consuming more GPU memory than the model weights themselves. Second-order methods exacerbate this issue by requiring the storage of curvature information, such as the Hessian or its factors. For models with billions of parameters, the memory required to store these high-precision curvature matrices frequently exceeds the capacity of modern hardware 17 CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 18 accelerators. Consequently, applying second-order optimization to LLMs requires not only approximating the curvature to reduce computational costs but also aggressively compressing them to fit within memory constraints. Existing approximation methods attempt to make second-order optimization feasible but typically address only one aspect of the efficiency bottleneck. One common approach is Natural Gradient Descent (NGD) [4], which substitutes the Hessian with the Fisher Information Matrix (FIM) [8]. To handle the memory limitations of large models, algorithms like Kronecker-Factored Approximate Curvature (KFAC) [93] approximate the FIM usingblock-diagonalsparsity, where each block corresponds to a layer. However, inverting these blocks remains computationally expensive, scaling withO(d 3 )wheredis the layer dimension. Consequently, KFAC implementations [9,145,117,113,132,116] must update the curvature information infrequently (e.g., every 100-1000 iterations), which damages convergence rates. Alternative methods like SNGD [124,157,102] and KBFGS [45] attempt to shift the complexity to the batch dimension (O(b 3 )orO(bd 2 )). While effective for small batches, these fail in Transformer models [146] where the effective batch size scales with sequence lengths that can reach thousands of tokens [ 138]. Thus, existing methods hit a wall: they are either too computationally heavy due to matrix inversion or fail to scale with sequence length. To overcome these limitations, we apply the Compression Trinity directly to the optimizer’s internal com- putations. We present MKOR, aMomentum-EnabledKronecker-Factorization-BasedOptimizer withRank- 1 Updates. MKOR unifies theSparsitypillar, inherited through block-diagonalization, with theLow-Rank Approximationpillar, which approximates the inverse of the covariance blocks using rank-1 updates via the Sherman-Morrison identity. This formulation fundamentally alters the efficiency landscape, reducing the inversion complexity fromO(d 3 )toO(d 2 )while simultaneously alleviating the communication bottleneck. Unlike standard second-order methods that require synchronizing large inverse factors (O(d 2 )), MKOR syn- chronizes only the rank-1 approximation vectors, reducing communication costs toO(d). These efficiency gains allow MKOR to update second-order information up to 100 times more frequently compared to state-of- the-art implementations like KAISA [ 116] and HyLo [102]. Crucially, unlike recent attempts like Eva [163] which store vectors but sacrifice momentum, MKOR fully preserves momentum information while achieving O(d 2 )complexity. The most closely related method to MKOR is Eva [ 163], which also targets the scalability bottleneck of KFAC by maintaining only rank-1 approximation vectors rather than full covariance factor inverses. This gives Eva a memory overhead of onlyO(d), significantly lower than MKOR’sO(d 2 ). However, this ag- gressive memory reduction comes at a fundamental cost: Eva discards the accumulated momentum of the inverse factors entirely, retaining only the most recent rank-1 snapshot. In contrast, MKOR preserves full momentum history in its factor inverses via the Sherman-Morrison update ( Equation 3.5andEquation 3.6), which exponentially averages past curvature information through theγdecay term. As we demonstrate in Section 3.4, this distinction has practical consequences: Eva fails to converge to the target accuracy on several benchmarks where MKOR succeeds, suggesting that the memory savings come at the expense of optimiza- tion stability. MKOR thus occupies a deliberate design point in the memory–convergence tradeoff, investing O(d 2 )memory to retain curvature history that proves essential for reliable convergence at scale. However, low-rank approximation alone is insufficient to fully resolve the scalability challenge. While it reduces the computational cost of inversion, it does not inherently address the memory overhead of storing the KFAC factors. To tackle this, we integrate theQuantizationpillar. By performing computations and storage in half-precision, we significantly reduce the memory footprint of the curvature factors. This is non-trivial because standard second-order methods rely on operations like Cholesky decomposition or matrix inversion, CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 19 which are numerically unstable in low precision. By pivoting to rank-1 updates via the Sherman-Morrison identity, MKOR replaces these unstable operations with simple matrix-vector products. These operations are inherently more robust to quantization noise, enabling us to store and compute curvature in half-precision without divergence. While second-order methods provide a significant advantage in the initial phase of training, their relative benefit over first-order methods can diminish in later stages. To maximize efficiency, we introduce a hybrid variant, MKOR-H. This method combines the rapid initial convergence of second-order optimization with the low computational overhead of first-order methods. By utilizing a loss-reduction-rate-based switching mech- anism, MKOR-H automatically transitions between regimes to ensure optimal resource utilization throughout the entire pretraining process. Our experiments demonstrate that MKOR successfully validates the efficacy of applying the Compression Trinity to the optimizer. MKOR outperforms state-of-the-art distributed second- and first-order methods by up to2.57×, reducing the training time of BERT-Large-Uncased from 8 hours to 3 hours on 64 A100 GPUs. Additionally, it achieves new state-of-the-art metrics on the GLUE dataset, successfully converging in settings where other second-order methods such as KFAC fail. 3.2 Background Training a neural network involves solving an optimization problem to find the optimal values for a set of weightsW=W m M m=1 , whereMis the number of layers in the network andW m is a matrix inR d×d . Second-order methods precondition the weights of the network with the inverse of the Hessian for better con- vergence rates. Block-diagonal approximations of NGD methods replace the Hessian with the block-diagonal FIM as shown in Equation 3.1, wherew m ∈R d 2 is the vector representation ofW m ,F m is the block corre- sponding to that layer andLis the loss function. Martens [92] shows that the FIM matches the Gauss-Newton matrix under certain conditions. w m :=w m −α(F m ) −1 ∇ w m L(3.1) KFAC-based methods reformulate the FIM block as the Kronecker product of two matrices.Equation 3.2 shows the update rule in KFAC, whereLis the loss function and(L m t ) −1 and(R m t ) −1 are the inverses of the left and right factors, respectively. W m :=W m −α(L m t ) −1 ∇ W m L(R m t ) −1 (3.2) (L m t ) −1 and(R m t ) −1 in Equation 3.2are computed usingEquation 3.3andEquation 3.4, respectively, wherea m is the activation value of a sample at layerm, andg m =∇ a m−1 Landγincorporate the momentum feature to avoid extreme changes in the factors. L m t =γL m t−1 + (1−γ)E[g m t g m t T ](3.3) R m t =γR m t−1 + (1−γ)E[a m−1 t a m−1 t T ](3.4) CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 20 Activation Covariance Gradient Covariance SM-Based Inversion a. Rank-1 Approximation & Synchronization - b. Norm-Based Stabilization c. SM-Based Inversion - d. Precondition & Rescale Gradients Figure 3.1: MKOR for layermon a single worker. The inputs of MKOR are the activationsA m t , the gradients of the loss function with respect to the inputsG m t , and the gradients of the loss function with respect to the weights∇ W m L. The output is the update values∆W m . 3.3 Methodology In this section, we first present the MKOR algorithm, its computation and communication complexity, then present hybrid MKOR (MKOR-H), and finally discuss MKOR’s convergence and stability. 3.3.1 The MKOR Algorithm Algorithm 1summarizes the MKOR optimizer for a single layer andFigure 3.1shows the workflow. For each layer (line 1 inAlgorithm 1) MKOR updates the second-order information and preconditions the gradi- ents, and at the end the backend optimizer updates the weight using the preconditioned gradients (line 14 in Algorithm 1). Rank-1 Approximation.For the rank-1 approximations of the covariance matrices, we use the average of the values across all the samples, i.e.a m−1 t =E[a m−1 t ]andg m t =E[g m t ](lines 2 and 3 in Algorithm 1and Figure 3.1-a).(A m−1 t ) −1 :,i and(G m t ) −1 :,i show thei th column of(A m−1 t ) −1 and(G m t ) −1 respectively, where A m−1 t andG m t are the activations and the gradients of layermrespectively. Norm-Based Stabilizer.The values in the factor inverses in second-order methods can become large or vanish due to extremely large or small values in activations and gradients, leading to numerical instabilities and over/underflows. Since the inverse of the factors are directly multiplied by the gradients to find the up- date values, it can cause oscillations or even divergence. MKOR uses a norm-based stabilizer to detect the CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 21 Algorithm 1MKOR Algorithm for a Single Layerm Input:A m−1 t , G m t , W m t−1 . Output:W m t . 1:ifm∈Second Order Layersthen 2:a m−1 t ← 1 b P b i=1 (A m−1 t ) :,i ▷Approx:A m−1 t A m−1 t T ≈a m−1 t a m−1 t T 3:g m t ← 1 b P b i=1 (G m t ) :,i ▷Approx:G m t G m t T ≈g m t g m t T 4:a m−1 t ,g m t ←AllReduce(a m−1 t ,g m t )▷Synchronize Approximations 5: ˆ Lt−1 m −1 ←if|L m t−1 −1 |> εthenζL m t−1 −1 + (1−ζ)IelseL m t−1 −1 ▷Norm-Based Stabilization 6: ˆ R m t−1 −1 ←if|R m t−1 −1 |> ε thenζR m t−1 −1 + (1−ζ)I elseR m t−1 −1 7:L m t −1 ←γ ˆ L m t−1 −1 + (1−γ) γ 2 (1+γ(1−γ)g m t T ˆ L m t−1 −1 g m t ) ˆ L m t−1 −1 g m t g m t T ˆ L m t−1 −1 ▷SM-Based Factor Inversion 8:R m t −1 ←γ ˆ R m t−1 −1 + (1−γ) γ 2 (1+γ(1−γ)a m t T ˆ R m t−1 −1 a m t ) ˆ R m t−1 −1 a m t a m t T ˆ R m t−1 −1 9:∆ ˆ W m t ←L m t −1 ∇ W m LR m t −1 ▷Precondition Gradients 10:∆W m t ← |∇ W m L| ∆ ˆ W m t ∆ ˆ W m t ▷Rescale Gradients 11:else 12:∆W m t ←∇ W m L 13:end if 14:W m t ← Optimizer.step(∆W m t ,W m t−1 ) Return:W m t . numerical instability and addresses it by modifying the inverse of the factors accordingly (lines 5 and 6 in Algorithm 1andFigure 3.1-b). More details on the norm-based stabilizer are inSection 3.3.3. SM-Based Inverter.MKOR directly modifies the inverse of the left and right factors using rank-1 updates, while using the momentum for better convergence. IfE[g m g m T ]is approximated using a rank-1 matrixg m g m T and using the Sherman-Morrison identity, Equation 3.5is obtained (line 7 inAlgorithm 1andFigure 3.1-c). L m t −1 =γL m t−1 −1 + (1−γ) γ 2 (1 +γ(1−γ)g m t T L m t−1 −1 g m t ) L m t−1 −1 g m t g m t T L m t − 1 −1 (3.5) Furthermore, if Equation 3.4is approximated usingE[a m−1 t a m−1 t T ]≈a m t a m T t with a similar derivation, Equation 3.6is obtained (line 8 inAlgorithm 1andFigure 3.1-c). R m t −1 =γR m t−1 −1 + (1−γ) γ 2 (1 +γ(1−γ)a m t T R m t−1 −1 a m t ) R m t−1 −1 a m t a m t T R m t−1 −1 (3.6) Rescaling Gradients.Preconditioning the gradients using the computed factors can change gradient norms. Sometimes, these changes interfere with the effect of the learning rate on the training process. To alleviate this and to make learning rate schedulers more effective, the preconditioned gradients are scaled so that their norm matches the original norms (line 10 in Algorithm 1andFigure 3.1-d). Complexity Analysis.MKOR reduces the memory, communication, and computation costs for factor inver- sion.Table 3.1compares the overheads of different optimizers.(1) Computation Complexity.MKOR inverts the left and right factors inEquation 3.2usingEquation 3.5andEquation 3.6, both of which can be computed using matrix-vector multiplications, and haveO(d 2 )computation complexity, in contrast to KFAC and SNGD CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 22 Table 3.1: The computation and communication complexity and memory overhead of the state-of-the-art implementations of the first- and second-order (second-order optimizers are written inbold). The division by 2 in MKOR is because MKOR uses half-precision computations. The complexity of KFAC-based methods depends on layer dimensions while SNGD methods mostly depend on the batch size. In transformers, due to the scaling of the batch size by the sequence length, batch sizes and layer dimensions are comparable, making both KFAC- and SNGD-based methods more expensive than SGD. OptimizerComputational Complexity Memory OverheadCommunication Complexity MKORO(d 2 +bd)O(2d 2 /2)O(2d/2) SNGD (HyLo)O(b 3 )O(2bd+b 2 )O(2bd+b 2 ) KFAC (KAISA)O(d 3 )O(4d 2 )O(4d 2 ) EvaO(d 2 +bd)O(2d)O(2d) SGD (Momentum)-O(d 2 )- ADAM / LAMB-O(d 2 )- methods that needO(d 3 )andO(b 3 )complexity to invert matrices inR d×d andR b×b respectively.(2) Commu- nication Complexity.The only data that is synchronized among different workers in MKOR is the two rank-1 approximations that have2delements. With quantization, this size can be halved. In KFAC, the activation and gradient covariance matrices and the inversion of left and right factors need to be synchronized between all the workers, leading to4d 2 data transfers. In SNGD , the activations and gradients are synchronized, lead- ing to2bddata transfers and the inverted kernels are broadcast, resulting inb 2 data transfers. Reducing the communication complexity of MKOR from quadratic to linear results in better performance on large number of workers.(3) Memory Overhead.MKOR needs to store the inverse of the left and right factors and two rank-1 approximation vectors, leading to2d 2 + 2dmemory overhead, and using half-precision computations further reduces this. KFAC stores the activation and gradient covariance matrices and the left and right fac- tors, leading to4d 2 memory overhead. SNGD stores the activations, the gradients, and the kernels they use as second-order information, leading to2bd+b 2 memory complexity. It is worth comparing MKOR’s complexity profile directly with Eva, as both methods share the same O(d 2 +bd)computational complexity and achieve linear communication costs. The key distinction lies in the memory–accuracy tradeoff. Eva achievesO(d)memory by storing only the rank-1 approximation vectors and discarding the factor inverses after each preconditioning step. MKOR, by contrast, retains the fulld×d inverse factors in half-precision, resulting inO(d 2 /2)memory but enabling momentum accumulation across iterations. While this is a meaningful increase in memory, it remains substantially lower than KFAC’sO(4d 2 ) overhead and, as shown in Section 3.4, translates to more stable convergence, particularly in settings where Eva’s lack of momentum leads to suboptimal solutions or failure to reach the target accuracy. 3.3.2 Hybrid MKOR We observed that second-order methods, including MKOR, usually accelerate training more during the first iterations of the training time, and as the loss flattens, their advantage over their first-order counterparts be- comes less noticeable. This is because the second-order information of the loss functions approach identity nearconvergencepoints. Thuswedesignedahybridsecond-andfirst-orderoptimizerwithalossdecreaserate- basedswitchingmethod(MKOR-H).MKOR-Hevaluatesthechangesinthelossfunctionindifferentiterations and switches back to first-order methods if needed for an efficient trade-off between the costly second-order updates and their benefits for convergence. CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 23 3.3.3 MKOR Convergence and Stability Inversion Frequency.Due to the high factor inversion costs in KFAC- and SNGD-based methods, re- searchers use the stale factor approach, which updates the inverted factors everyfiterations and reuses the results in the other iterations in their preconditioning to reduce the computation and communication costs. The reciprocal of inversion frequency,f, varies from a few 100s to a few 1000s. Our experiments show that in average-sized models such as ResNet-50 [56], in an iteration that includes the inversion of factors, the cost of KAISA and HyLo is150×more than an SGD iteration that reuses stale factors. Furthermore, more than 98% of the total cost in those iterations are spent on matrix inversion. The stale factors approach can lead to good preconditioners if the loss function landscape does not vary significantly in each iteration. However, this is a strong assumption and doesn’t necessarily hold in practice. Also, increasing the inversion frequency can benefit the convergence rate of the second-order methods. In addition, our experiments show that using stale factors can lead to converging to local minima in the loss function and damage the generalization of the model. Numerical Stability.In second-order techniques, we need to invert or find the roots of matrices of different sizes, which are usually not full-rank, resulting in numerical issues. The KFAC implementation uses singu- lar value decomposition (SVD) of the factors and masks the eigenvalues that are close to zero to deal with singular matrix inversion issues. In practice, the eigenvalues of the left and right factors in KFAC-based meth- ods computed fromEquation 3.3andEquation 3.4are increased manually by addingμIto each of them to improve numerical stability (μ >0is called the damping factor), but MKOR doesn’t need such numerical fixes. Furthermore, HyLo uses two decomposition methods to sample the batch of inputs, namely KID and KIS. KID requires inverting matrices inR b×b of rankmin(b, d), thus for batch sizes larger thandin a specific layer, the method fails. Unlike SVD or other iterative methods used for factor inversion, MKOR doesn’t suffer from numerical instabilities that rise from large condition numbers. MKOR has a single scalar division, in which the de- nominator is guaranteed to be non-zero based on Theorem 3.3.1, eliminating the numerical over/under-flow possibility and the need for damping factors (required by other second-order methods for computational sta- bility). Lemma 3.3.1.The factors computed usingEquation 3.5andEquation 3.6are all positive-definite. Anil et al. [5] suggest using double precision representation of numbers to avoid numerical instabilities in inverting or computing the roots of matrices. This approach adds more costs to the matrix inversion and increases the time complexity of the main bottleneck in second-order methods. MKOR does not need higher precision computations, and can use half-precision floating point operations to reduce costs significantly. This will improve the memory utilization and reduce the communication costs in GPUs by2×while using cheaper computation blocks for half-precision operations. Theorem 3.3.2shows an upper bound on the quantization error effect in the MKOR updates. Lemma 3.3.2.Assuming that the maximum quantization error isε, the maximum number in matrices and vectors ism, and the dimension of the vectors and matrices aredandd×drespectively, the quantization error of Equation 3.5andEquation 3.6isO((γ+ 4 (1−γ) γ 2 m 3 d 2 )ε) Exploding Gradients Problem.In second-order methods, where the gradients are preconditioned by var- ious factors, the exploding gradient problem is worsened. Our experiments show that in first-order methods, CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 24 by choosing a learning rate that doesn’t lead to divergence in the first few iterations, explosion in gradients almost never occurs. On the other hand, in second-order methods, we observe that the explosion can occur at any iteration, and both KFAC and SNGD implementations are prone to this problem. This can lead to ripples in accuracy and divergence. One of the main approaches for solving the exploding gradient problem is choosing small values for the learning rate, limiting the convergence rate significantly. In particular, small learning rates damage the second- order methods and make them almost as performant as their first-order counterparts. Considering that SGD is more robust against the exploding gradients and taking advantage of the direct control of MKOR on the inverse of the factors, the factors in MKOR are modified to lean toward SGD once the possibilityofexplodinggradientsisdetectedusingEquation3.7andEquation3.8, whereζisahyperparameter that controls the amount of information from the original factors that needs to be saved in the new factors. ˆ L m t =ζL m t + (1−ζ)I(3.7) ˆ R m t =ζR m t + (1−ζ)I(3.8) By expandingEquation 3.2with the new factors, we will getEquation 3.9, which reduces the loss based onTheorem 3.3.3. The first term in the right-hand side ofEquation 3.9is the KFAC term, the second and third terms are the left and right preconditioned versions, and the last term is the SGD term. ˆ L m −1 ∇ W m L ˆ R m −1 =ζ 2 L m −1 ∇ W m LR m −1 +ζ(1−ζ)L m −1 ∇ W m L+ζ(1−ζ)∇ W m LR m −1 + (1−ζ) 2 ∇ W m L (3.9) Lemma 3.3.3.Given a differentiable functionL(w)with first-order Taylor series approximation ˆ L(w− ∆ w ) = L ( w 0 ) − ∆ w T ∇ w L ( w 0 ) around point w 0 , assuming that at point w 0 the second-order derivative of the functionL(w)is given as∇ 2 w L(w 0 ) =H=L⊗R, whereLandRare positive-semi-definite matrices, for a value of∆w= ((ζL −1 +(1−ζ)I)⊗(ζR −1 +(1−ζ)I))∇L(w 0 ), the inequality ˆ L(w 0 −∆w)<L(w 0 ) holds. While this modification can avoid exploding gradients, overusing it with small values ofζwill convert MKOR to SGD. MKOR uses a factor norm-based metric that observes the infinity norm of the factors, and if they are greater than a specific threshold, the process of factor modification will be triggered. 3.4 Experimental Results In this section, we demonstrate the performance of MKOR on a large language model using different bench- marks, and analyze the timing of different components in different first- and second-order algorithms. For results on more models and training sets, please refer to Appendix A. Experiment Setup.For the BERT-Large-Uncased pre-training and fine-tuning experiments, we use up to 64 A100 GPUs on the Polaris [6] cluster, which has 560 nodes each with 4 NVIDIA A-100 GPUs with NVLink interconnects. The rest of the experiments are conducted on the Mist cluster [ 22], with 54 nodes each having 4 NVIDIA V100 GPUs with 32GB memory and NVLink inter-node connections. Each training experiment is conducted 5 times and the median timing is reported; reported accuracies are the median across multiple CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 25 Table 3.2: List properties of the models, datasets, and settings used in our experiments. ModelDatasetGPU Name#ParametersNameTrainTestArch# BERT-Large- Uncased 335.1MWikipedia - BookCorpus --A10064 ResNet-5025.5MImageNet1.2M50kV10064 AlexNet20.3MCIFAR-10050K10KV1004 BERT-Base-Cased108.9MSQuAD v1.187.6K10.6KV1004 BERT-Large-Cased335.1MIMDB25K25KV1004 runs.Table 3.2summarizes the models, datasets, and GPU architectures used in our experiments. For BERT- Large-Uncased pre-training, we use the same hyperparameters as [112], with factors in KAISA updated every 50 iterations and factors in MKOR and MKOR-H updated every 10 iterations. For ResNet-50, we follow the hyperparameters from [116], with MKOR factors updated every 10 iterations and the learning rate decaying by a factor of 2 at the end of epochs 25, 35, 40, 45, 50, 55, and 56. Our code base is publicly available at https://github.com/Mohammad-Mozaffari/mkor. LargeLanguageModels.Wepre-trainBERT-LargeUncasedandfine-tuneitfordifferentquestion-answering and text classification tasks. We use a setup similar to KAISA [116] for pre-training and fine-tuning. We use Fused LAMB [159] as the state-of-the-art first-order baseline. Similar to KAISA [116], for the pre-training process, we use the English Wikipedia [150] and the Toronto BookCorpus [170] dataset. These datasets were used in the original BERT pre-training; the latter dataset is not fully available, which results in a small reduc- tion in the baseline accuracies achieved in our experiments from the original BERT results. Following KAISA [ 116], due to the time-intensive process of hyperparameter tuning for the first phase of pre-training, we report the effectiveness of MKOR in the second phase of pre-training only while using the checkpoints of the first phase generated using the LAMB optimizer. As expected, the computation, communication, and memory complexity of HyLo is high, and the Khatri-Rao-based Interpolative Decomposition (KID) approximation method, the main idea of HyLo, cannot be executed because a single sample cannot fit into the 40GB memory of an A100 GPU. In addition, HyLo doesn’t support gradient accumulation due to its memory complexity, depending on the batch size; in LLMs such as BERT, the batch sizes are as large as 64k. 1 For the question answering task, we fine-tune the pre-trained BERT checkpoints on the SQuAD v1.1 [123] dataset.Table 3.3shows the F1 Score achieved using different optimizers and compares their convergence rate and speedups. 2 Convergence is defined as the number of iterations it takes the model to reach the same accuracyasthefirst-orderoptimizer. ThevanillaMKORandKAISAbothconvergeafter1000iterations, while the LAMB optimizer requires1,563steps. Considering that each step in MKOR is faster than KAISA, MKOR achieves an end-to-end speedup. MKOR-H will converge in 600 steps, reducing the number of steps in LAMB by2.6×, while achieving the same accuracy. In addition, it achieves2.57×speedup over the LAMB optimizer and1.75×speedup overKAISA. As another second-order baseline, weconsider Eva, which convergesin 1000 iterations, and MKOR-H achieves1.69×speedup over it. For classification tasks, we fine-tune BERT on the GLUE [ 147] dataset.Table 3.4compares the results 1 We define the convergence of BERT as the iteration in which the downstream accuracy reaches the first-order baseline. Nevertheless, we continue training until the same iteration for additional ablations and testing whether the second-order baselines can improve the capabilities of the model further. 2 All the timings and speedups reported are for the pretraining phase. We omit the fine-tuning time since it’s negligible in comparison to the pretraining costs. CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 26 Table 3.3: BERT-Large Uncased results on SQuAD v1.1 question answering task MetricLAMBKAISAMKORMKOR-HEva F190.4490.4490.5090.6490.55 # Iterations1,5631,0001,0006001,000 Time (h)7.975.715.253.105.24 Speedup (×)1.001.391.512.571.52 BERT-Large-Uncased Training Loss 02004006008001000120014001600 Iteration 2 4 6 8 10 Loss LAMB MKOR K-FAC MKOR-H Eva 012345678 Time (h) 2 4 6 8 10 Loss LAMB MKOR K-FAC MKOR-H Eva Figure 3.2: The pre-training loss of BERT-Large-Uncased using different optimizers. for different classification tasks in the GLUE dataset. MKOR with 1500 steps achieves a new state-of-the-art accuracy in GLUE dataset on BERT-Large Uncased, and MKOR and MKOR-H with 600 steps achieve the same average metric as the baseline, while reducing the number of steps by a factor of2.6×. MKOR and MKOR-H both achieve2.57×end-to-end speedup. After training KAISA for 1,563 steps, the model does not converge to the baseline average accuracy, while slowing down the convergence by0.89×. Eva requires 1000 steps to converge to the target average metric, being1.69×slower than MKOR-H with 600 steps and1.24% less accurate than MKOR with 1500 steps (it is noteworthy that the accuracy of the model plateaus when using more iterations with Eva). Per Figure 3.2, which shows the pre-training error during the training of BERT, MKOR decreases the error in fewer iterations in comparison to KAISA, Eva, and LAMB, leading to faster convergence. FromTable 3.3 andTable 3.4, MKOR-H converges in only 600 steps. ResNet-50 Experiments.We train ResNet-50, a convolutional neural network with more than 25M parame- ters, on ImageNet, an image classification task with more than 1.2M samples. The same setup is used in [116], and SGD is used as the first-order baseline. The target accuracy in this experiment is 75.9%. MKOR converges to this target accuracy in 57 epochs, while SGD, the first-order baseline, achieves this accuracy in 88 epochs. MKOR achieves1.49×speedup over Table 3.4: BERT-Large Uncased results on the GLUE classification tasks. We report the average of the metrics of different GLUE tasks (accuracy, F1 score, etc) for easier comparison. OptimizerLAMBKAISAMKORMKORMKOR-HEva Iterations1,5631,5631,5006006001000 Time (h)7.978.937.883.103.105.24 Speedup (×)1.000.891.012.572.571.52 Average Metric0.80230.7960.82140.80780.8110.809 CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 27 SGD. KAISA, the second-order baseline converges in 54 epochs, but due to its expensive steps, MKOR still converges1.04×faster than KAISA. We do not compare to HyLo because HyLo is not able to achieve the target accuracy for ResNet (it reaches 75.6% with tuning as reported in [102] and our experiments confirm it). As shown, the effect of complexity reduction and improvement in performance in MKOR is less obvious in ResNet because the model dimension (d) is smaller compared to LLMs such as BERT. Please seeTable 3.1 for comparison of complexity between methods. We could not reproduce the ResNet-50 results of Eva [163] on ImageNet because the hyperparameters are not reported. We tried to tune Eva on multiple settings and none converged to desired accuracy. It is important to note Eva is not comparing results with the most efficient implementation of KFAC. The KFAC version used in Eva is from [93], dated to 2015. A number of followup works, mentioned inSection 3.1, have provided faster implementations of KFAC. We use KAISA [116], the state-of-the-art implementation of KFAC. Also from discussions with KAISA authors and our own experiments the optimal inversion frequency for KFAC is 200. Eva uses an inversion frequency of 50 for KFAC, which makes KFAC slower. ResNet-50 Test Accuracy 020406080 Epochs 0 20 40 60 80 Accuracy 88 57 54 SGD MKOR KFAC 05000100001500020000250003000035000 Time (s) 0 20 40 60 80 Accuracy 328152200222978 SGD MKOR KFAC (a)(b) Figure 3.3: Test accuracy of ResNet-50 on ImageNet for MKOR, KAISA, and SGD on 64 GPUs. InversionFrequency.DuetothelowcomputationcomplexityoftheupdatesonMKOR,thefactorinversion frequency (f) in MKOR is in the range of 10.Figure 3.4-a shows that while the average iteration cost in KAISA is heavily dependent on the inversion frequency, MKOR’s cost is almost independent of the inversion frequency. Also Figure 3.4-b shows that increasing the inversion frequency leads to higher convergence rate. In addition, using stale factors may result in converging to a local minima. Hence, in MKOR we increase the convergence rate by updating the factors more frequently, without affecting the per-iteration cost, leading to end-to-end speedups in training. We use a simple autoencoder [128] on CIFAR-100 [71] in this experiment. Performance Analysis.We compare the performance of different parts of the optimizers to illustrate the bottlenecks and advantages of different methods. The training process for an optimizer has three steps: factor computation, precondition, and update weights. Figure 3.5shows the time spent on each task in different optimizers on two models; BERT-Large-Uncased, a transformer-based LLM with large sequence length and ResNet-50, a CNN. Since first-order optimizers such as SGD, ADAM, and LAMB don’t require factorization and preconditioning, their optimization time is only spent in updating the weights. In ResNet-50, since the CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 28 10 0 10 1 10 2 10 3 Factor Reuse Time 0 50 100 150 200 250 Time/Iteration (ms) Autoencoder MKOR KAISA 10 0 10 1 10 2 10 3 Factor Reuse Time 0 1000 2000 3000 4000 5000 6000 7000 Time/Iteration (ms) BERT-Large-Uncased MKOR KAISA (a) - Average Time per Iteration 0510152025 Epochs 0.4 0.5 0.6 0.7 0.8 0.9 Loss Autoencoder MKOR - Reuse 10 KAISA - Reuse 100 KAISA - Reuse 1000 MKOR - Reuse 100 0255075100125150175 Iteration 2 3 4 5 6 Loss BERT Large Uncased MKOR – Reuse 50 MKOR – Reuse 10 (b) - Test Loss Figure 3.4: The sensitivity of MKOR and KAISA for BERT-Large-Uncased and an Autoencoder model (a) and the effect of inversion frequency on the convergence properties of these models (b). model size is larger compared to the batch size, the factor computation and inversion is more expensive for KAISA compared to HyLo. This cost is significantly reduced in MKOR. For BERT-Large-Uncased, because of the large size of the model, the factor inversion time for KAISA is large. Also, due to the large sequence length value in this model, the kernel inversion time for HyLo is comparable to KAISA’s inversion time. But as expected, because of its low computational complexity, the aforementioned cost in our method is much smaller than the total training time, leading to speedups. It is important to note that HyLo diverges in this training process, hence convergence time is not reported for HyLo. The preconditioning and weight updates for the different methods are similar; hence, not much variation is observed. Memory Overheads.The memory overheads of MKOR in comparison to other optimizers are reported in Table 3.5. It can be observed that all the second-order methods have significant memory overheads compared to the first-order methods, but MKOR’s overhead is up to 1.5×lower than KFAC/KAISA. Approximation Error Experimental Results.Due to the low-rank properties of the covariance matrices, MKOR utilizes rank-1 approximations of the covariance matrices to accelerate the computations and com- munication in KFAC-based optimizers. Here, we aim to theoretically and experimentally support this choice. CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 29 Factor Computation PreconditionUpdate Weights 10 −1 Time (s) BERT Large Uncased - Wikipedia MKOR KAISA Eva LAMB Factor Computation PreconditionUpdate Weights 10 −2 Time (s) ResNet-50 - ImageNet MKOR HyLo KAISA Eva SGD ADAM (a)(b) Figure 3.5: Per-step breakdown of different optimizers on BERT-Large-Uncased (a) and ResNet-50 (b). The times reported in these graphs reflect only the optimizer computations. The majority of the training time is spent on the model’s forward and backward passes, which are identical across all optimizers and are not included here. Table 3.5: Per-GPU memory usage (in GB) for MKOR, KFAC/KAISA, LAMB, and SGD on BERT-Large- Uncased pre-training and ResNet-50 training on ImageNet. ModelMKORKFAC/KAISALAMBSGD ResNet-503.885.83-3.01 BERT23.3429.9712.80- As shown inFigure 3.6, our experiments show that the covariance matrices can be approximated with rank-1 matrices with low error and higher rank approximations are unnecessary in practice.Figure 3.6shows the error distribution of the optimal rank-1 approximation methods of the covariance matrices in ResNet-50 and BERT-Large-Uncased pre-training. Our extensive tests on well-known benchmarks show this property holds for all models and we have not come across a benchmark that does not have low-rank covariance matrices. Approximation Error Analysis and Extension to Higher Ranks.Small batch sizes and over parameteri- zation of networks will lead to low-rank covariance matrices in DNNs. Let’s consider the covariance matrix C=X T , whereC∈R d×d is the covariance matrix andX∈R d×b is a matrix in which each column corresponds to a single sample anddandbare the sample dimension and the per-GPU batch size respectively. Rank of the covariance matrix ismin(b, d). If the per-GPU batch sizes are small, the covariance matrices in each GPU will be low-rank. Rank-1 approximation methods can work well in these scenarios. If the batch sizes in each GPU are large, we observe that the covariance matrices will stay low-rank. The underlying reason for this observation is that current neural networks are over-parameterized, and as a result, different features in the covariance matrices of the activations and the output gradients won’t be linearly independent, resulting in low-rank covariance matrices. Extending MKOR to Higher Ranks:Furthermore, one can extend MKOR to use higher-rank covariance matrices. Let’s assume thatC= P r i=1 c i c T i whereris the rank of the covariance matrixC. We can apply SMW identity to computeC new 1 = (C old +c 1 c T 1 ) −1 withO(d 2 )computational complexity. Then we can computeC new 2 = (C new 1 +c 2 c T 2 ) −1 using SMW identity withO(d 2 )computational complexity. We can continue the same pattern by computingC new i = (C new i−1 +c i c T i ) −1 . The total computation complexity of CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 30 0.00.20.40.60.81.0 Activation Error 0.00 0.05 0.10 0.15 Probability Density BERT Activation Error Distribution 0.00.20.40.60.81.0 Input Gradient Error 0.000 0.025 0.050 0.075 0.100 0.125 0.150 Probability Density BERT Input Gradient Error Distribution (a)(b) 0.00.20.40.60.81.0 Activation Error 0.0 0.1 0.2 0.3 0.4 Probability Density ResNet-50 Activation Error Distribution 0.00.20.40.60.81.0 Input Gradient Error 0.00 0.02 0.04 0.06 0.08 0.10 Probability Density ResNet-50 Input Gradient Error Distribution (c)(d) Figure 3.6: Rank-1 error for activation and input gradient covariance matrices for BERT-Large-Uncased pre- training (a, b) and ResNet-50 on ImageNet (c, d). CHAPTER 3. MKOR: MOMENTUM-ENABLED KRONECKER-FACTOR-BASED OPTIMIZER USING RANK-1 UPDATES 31 this process will beO(rd 2 ). We should add this cost to the cost of computing the low-rank approximation ofCwhich requires an SVD. Using SVD kills the main advantage of using low-rank computations, since the computational complexity of applying SVD is the same as inverting the factors directly. We could not find any cheaper way to compute low-rank approximations of the covariance matrices, except for the rank-1 approximation used in this chapter. 3.5 Conclusion In this chapter, we presented MKOR, a scalable second-order optimizer that effectively executes the second strategyoftheCompressionTrinity: acceleratingpretrainingbyreducingthenumberofrequirediterations. By applying the Trinity’s pillars directly to the optimizer’s internal mechanics, MKOR overcomes the historical bottlenecks of second-order methods. We leveraged low-rank approximations via rank-1 updates to reduce inversion complexity fromO(d 3 )toO(d 2 ), and employed stability mechanisms to enable lower-precision communication, reducing overheads toO(d). Our experiments confirm that MKOR significantly outperforms state-of-the-art first- and second-order optimizers, delivering up to2.57×faster training for large language models. However, accelerating convergence is only half of the pretraining equation. While MKOR reduces the total number of steps, the computational costper stepremains dominated by the dense matrix multiplications inherent to the Transformer architecture. To unlock the full potential of the Compression Trinity during the pretraining phase, we must also address Strategy 1: reducing the fundamental FLOPs and memory traffic of the linear layers themselves. In the next chapter, we introduce SLOPE, which extends the Trinity from the optimizer to the model weights, jointly applying sparsity and lazy low-rank adapters to accelerate the training dynamics without sacrificing model quality. Chapter 4 SLOPE: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs Publicationand Contributions.The content of this chapter is based on the paper “SLOPE: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs” [101], published at the Thirteenth International Conference on Learning Representations (ICLR) 2025. This work was conducted in collaboration with Amir Yazdanbakhsh, Zhao Zhang, and Maryam Mehri Dehnavi. Mohammad Mozaffari was the lead contributor, responsible for the algorithm design, implementation, and experimental evaluation. Amir Yazdanbakhsh and Maryam Mehri Dehnavi supervised the project and contributed to the writing and revision of the manuscript. Zhao Zhang provided additional supervisory guidance. 4.1 Introduction Following our exploration of optimizer-level acceleration inChapter 3, we now turn to the first strategy of the Compression Trinity: accelerating the computational cost of each individual training iteration. Large Language Models (LLMs) require massive resources for their life-cycle stages, specifically pretraining [ 119] on high-quality text [40,44] and fine-tuning on downstream tasks [147,123]. These phases are dominated by three intensive matrix multiplications per layer: the forward pass, the backward pass for input gradients, and the backward pass for weight gradients. To execute Strategy 1 (Accelerating Per-Iteration Computation), we must reduce the FLOPs and memory traffic for all three operations. The first pillar of our Trinity, sparsity, offersapathforward[58]. Whileunstructuredsparsitylackshardwaresupport[151]andrigidblock-structured sparsity damages accuracy [67,84,25], semi-structured N:M sparsity (e.g., 2:4, where 2 out of 4 consecutive elementsaresettozero)strikesabalance. Itisflexibleenoughtopreservemodelqualitywhilebeingstructured enough for acceleration via NVIDIA’s Sparse Tensor Cores [ 108], with algorithms rapidly evolving for these patterns [ 69,89,10]. However, applying N:M sparsity to the training phase faces a critical ”transposability” bottleneck. While N:M sparsity successfully accelerates the forward pass, it fails to accelerate the backward pass because the row-wise N:M structure is destroyed when the weight matrix is transposed. Prior attempts to address this 32 CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS33 have focused on finding ”transposable masks” that maintain structure in both orientations [63,166,62]. Un- fortunately, these methods often require expensive search algorithms or enforce rigid constraints that signif- icantly reduce model accuracy. Paradoxically, the overhead of these complex mask searches can result in severe training slow-downs, up to 8 . 4 × [ 62]. Alternative approaches that change the sparsity mask dynami- cally [29,168,69,89] also introduce computational overheads and waste resources training weights that are eventually pruned. To strengthen the sparsity pillar for pretraining, we propose a noveldouble-prunedbackwardpassformu- lation with theoretical convergence guarantees. Instead of enforcing the restrictive condition that a mask must be inherently transposable, our approach allows the forward pass to use a standard N:M mask. In the back- ward pass, we transpose the weight matrix first and then impose a new N:M sparsity pattern. This formulation allows the weight matrices to exhibit a much wider range of sparsity patterns compared to rigid transposable masks, leading to significantly improved accuracy while enabling acceleration in both directions. While resolving the compute bottleneck, aggressive sparsity can still lead to an accuracy gap compared to dense models. To bridge this gap without sacrificing efficiency, we integrate the third pillar of the Trinity: Low-Rank Approximations. Previous methods often resort to dense fine-tuning to recover accuracy [ 142,62], but this converts the model back to a dense state, negating all memory and compute savings during inference. The intuition behind this integration lies in the observation that, while the parameter space of LLMs is vast, learning effectively occurs on a manifold of much lower intrinsic dimension [77,61]. This suggests that full-rank updates are not strictly necessary for recovering the expressivity lost to pruning. However, combining these pillars is non-trivial. Theoretically, a low-rank adapterLRis a dense matrix; adding it to a sparse weight matrixW(W ′ =W+LR) would cause ”fill-in,” where strictly zero elements become non-zero, effectively destroying the sparsity pattern and its associated hardware benefits. To harness the power of both without this collision, we proposeLazy Low-Rank Adapters. We treat the sparse weights and dense adapters as parallel computational paths rather than a merged tensor. Furthermore, unlike standard adapters, ours are ”lazy” because they are introduced only during the final 1% of pretraining iterations. Our experiments show that these adapters converge noticeably faster compared to the original model parameters at the same parameter count. This approach improves the accuracy of the models while ensuring the base model remains sparse and efficient for deployment. We present SLOPE, a Double-PrunedSparse Plus LazyLow-rank AdapterPretraining method for LLMs that jointly leverages these pillars. Key contributions of SLOPE are: •Double-Pruned backward pass→We propose to transpose an already sparsified N:M weight matrix (for- ward pass) before imposing another round of N:M sparsity (backward pass). This improves model quality and eliminates the overhead of searching for transposable masks. •Lazy Low-Rank adapters→We utilize the low-rank pillar to recover accuracy by introducing additional parameters with minimal compute and memory overheads, strictly for the last 1% of pretraining iterations (see Figure 4.1). •Optimized CUDA kernels→We jointly optimize NVIDIA 2:4 sparse kernels and low-rank calls through efficient tiling and scheduling. Our highly-optimized CUDA kernels result in 1 . 25 × end-to-end training speedup and1.54×inference speedup on LLMs with billions of parameters, while reducing training and inference memory footprints by up to0.63×and0.61×, respectively. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS34 N:M Sparse Pretraining with Lossy Backward Pass (99% of iterations) Sparse + Lazy Low-Rank Pretraining (1% of iterations) Row-wise Pruned Row- and Column-wise Pruned <latexit sha1_base64="r8l4jVrNcjhmyI+Hh6dipVLSdfU=">AAACC3icZVDLSsNAFL2prxpfVcGNm2ApuCqJi+qy1I3LFuwD2lAmk2k7dDKJMxOhhHyCa7f6Aa7ciUv9CDfu/QsnbZHWHhg4nHMP987xIkalsu0vI7e2vrG5ld82d3b39g8Kh0ctGcYCkyYOWSg6HpKEUU6aiipGOpEgKPAYaXvj68xv3xMhachv1SQiboCGnA4oRkpLbi9AaoQRSzppn/YLRbtsT2GtEmdOitWTxjd9qX3U+4Wfnh/iOCBcYYak7Dp2pNwECUUxI6nZiyWJEB6jIelqylFApJtMj06tklZ8axAK/biypqpZWogkKJByEnh6NLtSrpiZ+mcuep7nUz5MlwLdWA2u3ITyKFaE49n+QcwsFVpZMZZPBcGKTTRBWFD9BQuPkEBY6fpM3Y3zv4lV0rooO5VypaFLqsEMeTiFMzgHBy6hCjdQhyZguINHeIJn48F4Nd6M99lozphnjmEJxucvTP2eyA==</latexit> X i <latexit sha1_base64="K6h4zrJR4nSnsTeNze/DzXMRj5M=">AAACC3icZVDLSsNAFL3xWeOrKrhxEywFVyVxUV2WunHZgn1IG8pkMm2HTiZxZiKUkE9w7VY/wJU7cakf4ca9f+GkLdLaAwOHc+7h3jlexKhUtv1lrKyurW9s5rbM7Z3dvf38wWFThrHApIFDFoq2hyRhlJOGooqRdiQICjxGWt7oKvNb90RIGvIbNY6IG6ABp32KkdKS2w2QGmLEktu0R3v5gl2yJ7CWiTMjhcpx/Zu+VD9qvfxP1w9xHBCuMENSdhw7Um6ChKKYkdTsxpJECI/QgHQ05Sgg0k0mR6dWUSu+1Q+FflxZE9UszkUSFEg5Djw9ml0pl8xM/TPnPc/zKR+kC4FOrPqXbkJ5FCvC8XR/P2aWCq2sGMungmDFxpogLKj+goWHSCCsdH2m7sb538QyaZ6XnHKpXNclVWGKHJzAKZyBAxdQgWuoQQMw3MEjPMGz8WC8Gm/G+3R0xZhljmABxucvTpueyQ==</latexit> Y i <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="yqnueIV/TTePkiRkAJBwgm9viSA=">AAACJ3icZVC9TsMwGHTKX0n5aWFCLBFVBVOVMBTGChbGgvontWnluE5r1XEi20GqIj8CD8LMCs/AhmBCjEy8Am5aoZaeZOl09539+byIEiFt+8PIrK1vbG5lt83czu7efr5w0BRhzBFuoJCGvO1BgSlhuCGJpLgdcQwDj+KWN76e+q17zAUJWV1OIuwGcMiITxCUWurnT5Nuekni0RirRKMbQDlCkCYt1Se9O6V6daVUP1+0y3YKa5U4c1KsFo8+C/wnV+vnv7uDEMUBZhJRKETHsSPpJpBLgihWZjcWOIJoDIe4oymDARZuku6irJJWBpYfcn2YtFLVLC1EEhgIMQk8PTpdV6yYU/XPXPQ8b0DYUC0FOrH0L92EsCiWmKHZ+35MLRla09KsAeEYSTrRBCJO9BcsNIIcIqmrNXU3zv8mVknzvOxUypVbp1i9AjNkwTE4AWfAARegCm5ADTQAAg/gCTyDF+PReDXejPfZaMaYZw7BEoyvX8CSqok=</latexit> W R i T <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="UhFr4SSQRWC2lFCC2emwqS99jXc=">AAACJ3icZVDNSgMxGMz6W+vfqje9LNaiBym7HqrHYhHEUxX7A21dstm0Dc1mlyQrlLCP4IN49qrP4K3o0aMH38FsW6S1A4Fh5pvky3gRJULa9qexsLi0vLKaWcuub2xubZs7uzURxhzhKgppyBseFJgShquSSIobEccw8Ciue/1y6tcfMRckZPdyEOF2ALuMdAiCUkuueaxao0uUR2OcKKVaAZQ9BKmqJy55UHen5SSFa+bsgj2CNU+cCcmVjoKfm/3hVcU1v1t+iOIAM4koFKLp2JFsK8glQRQn2VYscARRH3ZxU1MGAyzaarRLYuW14ludkOvDpDVSs/mpiIKBEIPA06PpvmLOTNU/c9rzPJ+wbjITaMayc9FWhEWxxAyN3+/E1JKhlZZm+YRjJOlAE4g40V+wUA9yiKSuNqu7cf43MU9qZwWnWCje6pIuwRgZcAAOwQlwwDkogWtQAVWAwBN4Aa/gzXg23o2h8TEeXTAmmT0wA+PrF+Lxqqo=</latexit> W R,C i Forward Pass Backward Pass <latexit sha1_base64="4WyuxAq9vRSnViln5d8k1BDYE10=">AAACQXicZVDLTgIxFO3gC/GFunTTYExckRkX6MYEdePCBSYiGAYnd0qBhk5n0nZMyGQ+wC/wE/wQ1m4lfgI749aNBXxzkian59yT9h4/4kxp236xMnPzC4tL2eXcyura+kZ+c+tahbEktEpCHsq6D4pyJmhVM81pPZIUAp/Tmt87G/u1OyoVC8WV7ke0GUBHsDYjoI3k5U9cAT4HL6l5LMVuALpLgCcXKT7GX9bNP+v26udWT738rl20J8CzxPkku+WCW7h/GAwrXn7ktkISB1RowkGphmNHupmA1IxwmubcWNEISA86tGGogICqZjJZNcV7RmnhdijNERpP1Nzer0gCgVL9wDej4z+qGXOsfpu/Pd9vMdFJ/wQasW4fNRMmolhTQabvt2OOdYjHdeIWk5Ro3jcEiGRmBUy6IIFoU3rOdOP8b2KWXB8UnVKxdGlKOkVTZNEOKqB95KBDVEbnqIKqiKBH9ISe0dAaWCPr1Xqbjmasz8w2+gPr/QNcrLRX</latexit> r W i L=r Y i L T X <latexit sha1_base64="4WyuxAq9vRSnViln5d8k1BDYE10=">AAACQXicZVDLTgIxFO3gC/GFunTTYExckRkX6MYEdePCBSYiGAYnd0qBhk5n0nZMyGQ+wC/wE/wQ1m4lfgI749aNBXxzkian59yT9h4/4kxp236xMnPzC4tL2eXcyura+kZ+c+tahbEktEpCHsq6D4pyJmhVM81pPZIUAp/Tmt87G/u1OyoVC8WV7ke0GUBHsDYjoI3k5U9cAT4HL6l5LMVuALpLgCcXKT7GX9bNP+v26udWT738rl20J8CzxPkku+WCW7h/GAwrXn7ktkISB1RowkGphmNHupmA1IxwmubcWNEISA86tGGogICqZjJZNcV7RmnhdijNERpP1Nzer0gCgVL9wDej4z+qGXOsfpu/Pd9vMdFJ/wQasW4fNRMmolhTQabvt2OOdYjHdeIWk5Ro3jcEiGRmBUy6IIFoU3rOdOP8b2KWXB8UnVKxdGlKOkVTZNEOKqB95KBDVEbnqIKqiKBH9ISe0dAaWCPr1Xqbjmasz8w2+gPr/QNcrLRX</latexit> r W i L=r Y i L T X <latexit sha1_base64="r8l4jVrNcjhmyI+Hh6dipVLSdfU=">AAACC3icZVDLSsNAFL2prxpfVcGNm2ApuCqJi+qy1I3LFuwD2lAmk2k7dDKJMxOhhHyCa7f6Aa7ciUv9CDfu/QsnbZHWHhg4nHMP987xIkalsu0vI7e2vrG5ld82d3b39g8Kh0ctGcYCkyYOWSg6HpKEUU6aiipGOpEgKPAYaXvj68xv3xMhachv1SQiboCGnA4oRkpLbi9AaoQRSzppn/YLRbtsT2GtEmdOitWTxjd9qX3U+4Wfnh/iOCBcYYak7Dp2pNwECUUxI6nZiyWJEB6jIelqylFApJtMj06tklZ8axAK/biypqpZWogkKJByEnh6NLtSrpiZ+mcuep7nUz5MlwLdWA2u3ITyKFaE49n+QcwsFVpZMZZPBcGKTTRBWFD9BQuPkEBY6fpM3Y3zv4lV0rooO5VypaFLqsEMeTiFMzgHBy6hCjdQhyZguINHeIJn48F4Nd6M99lozphnjmEJxucvTP2eyA==</latexit> X i <latexit sha1_base64="Mi87WVhS7UDgdG/7vmsKh8+R96E=">AAACWXicZVHLahsxFJWnaetOX069azairqGLYma6cLMJmJpCKVk4oX7hcYcrWbaFNZpB0gSMmG/rXxRK11100eYXIj8SYvuA4HDOvdx7j0gmuDZB8KvkPTh6+Ohx+Yn/9NnzFy8rx696Os0VZV2ailQNCGgmuGRdw41g0wxSIhgfbJor/z+FVOap/KbWWZsnMBM8imnYJwUV4aRBCIgtoOYFzhKwMwpCHte4DN8aw33LRutB1siclZYe2f1i5h/t5fv24VDXKkFjWANfEjCLam13ib/vr7++bkTV/5Gk5TmCZOGCtB6FAaZGVtQhlPBCj/KNcuALmDGRo5KSJge2/UmBa47ZYKnqXJPGrxW/fq9FguJ1suEuNLVuvrAXKl35n2PkAmXs2KnYZSb6enYcpnlhkm6mT/NBTYpXsWMJ1wxasTSEaCKuxMwnYMCatxn+C6bcD+JQ9L70AibjeaFC+kT2qCMTtAb9A6F6CNqoS+og7qIoh/oD/qPrku/vZJX9vxNqVfa9lTRDrzqDSNKux4=</latexit> r X i L=r Y i LW R,C i <latexit sha1_base64="Mi87WVhS7UDgdG/7vmsKh8+R96E=">AAACWXicZVHLahsxFJWnaetOX069azairqGLYma6cLMJmJpCKVk4oX7hcYcrWbaFNZpB0gSMmG/rXxRK11100eYXIj8SYvuA4HDOvdx7j0gmuDZB8KvkPTh6+Ohx+Yn/9NnzFy8rx696Os0VZV2ailQNCGgmuGRdw41g0wxSIhgfbJor/z+FVOap/KbWWZsnMBM8imnYJwUV4aRBCIgtoOYFzhKwMwpCHte4DN8aw33LRutB1siclZYe2f1i5h/t5fv24VDXKkFjWANfEjCLam13ib/vr7++bkTV/5Gk5TmCZOGCtB6FAaZGVtQhlPBCj/KNcuALmDGRo5KSJge2/UmBa47ZYKnqXJPGrxW/fq9FguJ1suEuNLVuvrAXKl35n2PkAmXs2KnYZSb6enYcpnlhkm6mT/NBTYpXsWMJ1wxasTSEaCKuxMwnYMCatxn+C6bcD+JQ9L70AibjeaFC+kT2qCMTtAb9A6F6CNqoS+og7qIoh/oD/qPrku/vZJX9vxNqVfa9lTRDrzqDSNKux4=</latexit> r X i L=r Y i LW R,C i Row-wise Pruned <latexit sha1_base64="r8l4jVrNcjhmyI+Hh6dipVLSdfU=">AAACC3icZVDLSsNAFL2prxpfVcGNm2ApuCqJi+qy1I3LFuwD2lAmk2k7dDKJMxOhhHyCa7f6Aa7ciUv9CDfu/QsnbZHWHhg4nHMP987xIkalsu0vI7e2vrG5ld82d3b39g8Kh0ctGcYCkyYOWSg6HpKEUU6aiipGOpEgKPAYaXvj68xv3xMhachv1SQiboCGnA4oRkpLbi9AaoQRSzppn/YLRbtsT2GtEmdOitWTxjd9qX3U+4Wfnh/iOCBcYYak7Dp2pNwECUUxI6nZiyWJEB6jIelqylFApJtMj06tklZ8axAK/biypqpZWogkKJByEnh6NLtSrpiZ+mcuep7nUz5MlwLdWA2u3ITyKFaE49n+QcwsFVpZMZZPBcGKTTRBWFD9BQuPkEBY6fpM3Y3zv4lV0rooO5VypaFLqsEMeTiFMzgHBy6hCjdQhyZguINHeIJn48F4Nd6M99lozphnjmEJxucvTP2eyA==</latexit> X i <latexit sha1_base64="K6h4zrJR4nSnsTeNze/DzXMRj5M=">AAACC3icZVDLSsNAFL3xWeOrKrhxEywFVyVxUV2WunHZgn1IG8pkMm2HTiZxZiKUkE9w7VY/wJU7cakf4ca9f+GkLdLaAwOHc+7h3jlexKhUtv1lrKyurW9s5rbM7Z3dvf38wWFThrHApIFDFoq2hyRhlJOGooqRdiQICjxGWt7oKvNb90RIGvIbNY6IG6ABp32KkdKS2w2QGmLEktu0R3v5gl2yJ7CWiTMjhcpx/Zu+VD9qvfxP1w9xHBCuMENSdhw7Um6ChKKYkdTsxpJECI/QgHQ05Sgg0k0mR6dWUSu+1Q+FflxZE9UszkUSFEg5Djw9ml0pl8xM/TPnPc/zKR+kC4FOrPqXbkJ5FCvC8XR/P2aWCq2sGMungmDFxpogLKj+goWHSCCsdH2m7sb538QyaZ6XnHKpXNclVWGKHJzAKZyBAxdQgWuoQQMw3MEjPMGz8WC8Gm/G+3R0xZhljmABxucvTpueyQ==</latexit> Y i <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="yqnueIV/TTePkiRkAJBwgm9viSA=">AAACJ3icZVC9TsMwGHTKX0n5aWFCLBFVBVOVMBTGChbGgvontWnluE5r1XEi20GqIj8CD8LMCs/AhmBCjEy8Am5aoZaeZOl09539+byIEiFt+8PIrK1vbG5lt83czu7efr5w0BRhzBFuoJCGvO1BgSlhuCGJpLgdcQwDj+KWN76e+q17zAUJWV1OIuwGcMiITxCUWurnT5Nuekni0RirRKMbQDlCkCYt1Se9O6V6daVUP1+0y3YKa5U4c1KsFo8+C/wnV+vnv7uDEMUBZhJRKETHsSPpJpBLgihWZjcWOIJoDIe4oymDARZuku6irJJWBpYfcn2YtFLVLC1EEhgIMQk8PTpdV6yYU/XPXPQ8b0DYUC0FOrH0L92EsCiWmKHZ+35MLRla09KsAeEYSTrRBCJO9BcsNIIcIqmrNXU3zv8mVknzvOxUypVbp1i9AjNkwTE4AWfAARegCm5ADTQAAg/gCTyDF+PReDXejPfZaMaYZw7BEoyvX8CSqok=</latexit> W R i T <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="kIhUsZOPKjEI5MU6bqxdmddDWWw=">AAAB/3icZVC7SgNBFL0bX3F9RS1tBkPAKuxaRJtg0MYyAfOAZAmzs5NkyOzsMjMrhCWFta1+gYWd2PoFfoL4Cf6FkweSmAMXDufcw8w9fsyZ0o7zbWXW1jc2t7Lb9s7u3v5B7vCooaJEElonEY9ky8eKciZoXTPNaSuWFIc+p01/eDPxm/dUKhaJOz2KqRfivmA9RrA2Uq3czeWdojMFWiXunOSvPu1y/PJlV7u5n04QkSSkQhOOlWq7Tqy9FEvNCKdju5MoGmMyxH3aNlTgkCovnX50jApGCVAvkmaERlPVLixEUhwqNQp9sxpiPVAr5kT9Mxc93w+Y6I+XAu1E9y69lIk40VSQ2fu9hCMdoUkZKGCSEs1HhmAimTkBkQGWmGhTmW26cf83sUoa50W3VCzVnHzlGmbIwgmcwhm4cAEVuIUq1IEAhUd4gmfrwXq13qz32WrGmmeOYQnWxy/Gkpjf</latexit> = <latexit sha1_base64="UhFr4SSQRWC2lFCC2emwqS99jXc=">AAACJ3icZVDNSgMxGMz6W+vfqje9LNaiBym7HqrHYhHEUxX7A21dstm0Dc1mlyQrlLCP4IN49qrP4K3o0aMH38FsW6S1A4Fh5pvky3gRJULa9qexsLi0vLKaWcuub2xubZs7uzURxhzhKgppyBseFJgShquSSIobEccw8Ciue/1y6tcfMRckZPdyEOF2ALuMdAiCUkuueaxao0uUR2OcKKVaAZQ9BKmqJy55UHen5SSFa+bsgj2CNU+cCcmVjoKfm/3hVcU1v1t+iOIAM4koFKLp2JFsK8glQRQn2VYscARRH3ZxU1MGAyzaarRLYuW14ludkOvDpDVSs/mpiIKBEIPA06PpvmLOTNU/c9rzPJ+wbjITaMayc9FWhEWxxAyN3+/E1JKhlZZm+YRjJOlAE4g40V+wUA9yiKSuNqu7cf43MU9qZwWnWCje6pIuwRgZcAAOwQlwwDkogWtQAVWAwBN4Aa/gzXg23o2h8TEeXTAmmT0wA+PrF+Lxqqo=</latexit> W R,C i Forward Pass Backward Pass <latexit sha1_base64="4WyuxAq9vRSnViln5d8k1BDYE10=">AAACQXicZVDLTgIxFO3gC/GFunTTYExckRkX6MYEdePCBSYiGAYnd0qBhk5n0nZMyGQ+wC/wE/wQ1m4lfgI749aNBXxzkian59yT9h4/4kxp236xMnPzC4tL2eXcyura+kZ+c+tahbEktEpCHsq6D4pyJmhVM81pPZIUAp/Tmt87G/u1OyoVC8WV7ke0GUBHsDYjoI3k5U9cAT4HL6l5LMVuALpLgCcXKT7GX9bNP+v26udWT738rl20J8CzxPkku+WCW7h/GAwrXn7ktkISB1RowkGphmNHupmA1IxwmubcWNEISA86tGGogICqZjJZNcV7RmnhdijNERpP1Nzer0gCgVL9wDej4z+qGXOsfpu/Pd9vMdFJ/wQasW4fNRMmolhTQabvt2OOdYjHdeIWk5Ro3jcEiGRmBUy6IIFoU3rOdOP8b2KWXB8UnVKxdGlKOkVTZNEOKqB95KBDVEbnqIKqiKBH9ISe0dAaWCPr1Xqbjmasz8w2+gPr/QNcrLRX</latexit> r W i L=r Y i L T X <latexit sha1_base64="4WyuxAq9vRSnViln5d8k1BDYE10=">AAACQXicZVDLTgIxFO3gC/GFunTTYExckRkX6MYEdePCBSYiGAYnd0qBhk5n0nZMyGQ+wC/wE/wQ1m4lfgI749aNBXxzkian59yT9h4/4kxp236xMnPzC4tL2eXcyura+kZ+c+tahbEktEpCHsq6D4pyJmhVM81pPZIUAp/Tmt87G/u1OyoVC8WV7ke0GUBHsDYjoI3k5U9cAT4HL6l5LMVuALpLgCcXKT7GX9bNP+v26udWT738rl20J8CzxPkku+WCW7h/GAwrXn7ktkISB1RowkGphmNHupmA1IxwmubcWNEISA86tGGogICqZjJZNcV7RmnhdijNERpP1Nzer0gCgVL9wDej4z+qGXOsfpu/Pd9vMdFJ/wQasW4fNRMmolhTQabvt2OOdYjHdeIWk5Ro3jcEiGRmBUy6IIFoU3rOdOP8b2KWXB8UnVKxdGlKOkVTZNEOKqB95KBDVEbnqIKqiKBH9ISe0dAaWCPr1Xqbjmasz8w2+gPr/QNcrLRX</latexit> r W i L=r Y i L T X <latexit sha1_base64="r8l4jVrNcjhmyI+Hh6dipVLSdfU=">AAACC3icZVDLSsNAFL2prxpfVcGNm2ApuCqJi+qy1I3LFuwD2lAmk2k7dDKJMxOhhHyCa7f6Aa7ciUv9CDfu/QsnbZHWHhg4nHMP987xIkalsu0vI7e2vrG5ld82d3b39g8Kh0ctGcYCkyYOWSg6HpKEUU6aiipGOpEgKPAYaXvj68xv3xMhachv1SQiboCGnA4oRkpLbi9AaoQRSzppn/YLRbtsT2GtEmdOitWTxjd9qX3U+4Wfnh/iOCBcYYak7Dp2pNwECUUxI6nZiyWJEB6jIelqylFApJtMj06tklZ8axAK/biypqpZWogkKJByEnh6NLtSrpiZ+mcuep7nUz5MlwLdWA2u3ITyKFaE49n+QcwsFVpZMZZPBcGKTTRBWFD9BQuPkEBY6fpM3Y3zv4lV0rooO5VypaFLqsEMeTiFMzgHBy6hCjdQhyZguINHeIJn48F4Nd6M99lozphnjmEJxucvTP2eyA==</latexit> X i <latexit sha1_base64="Mi87WVhS7UDgdG/7vmsKh8+R96E=">AAACWXicZVHLahsxFJWnaetOX069azairqGLYma6cLMJmJpCKVk4oX7hcYcrWbaFNZpB0gSMmG/rXxRK11100eYXIj8SYvuA4HDOvdx7j0gmuDZB8KvkPTh6+Ohx+Yn/9NnzFy8rx696Os0VZV2ailQNCGgmuGRdw41g0wxSIhgfbJor/z+FVOap/KbWWZsnMBM8imnYJwUV4aRBCIgtoOYFzhKwMwpCHte4DN8aw33LRutB1siclZYe2f1i5h/t5fv24VDXKkFjWANfEjCLam13ib/vr7++bkTV/5Gk5TmCZOGCtB6FAaZGVtQhlPBCj/KNcuALmDGRo5KSJge2/UmBa47ZYKnqXJPGrxW/fq9FguJ1suEuNLVuvrAXKl35n2PkAmXs2KnYZSb6enYcpnlhkm6mT/NBTYpXsWMJ1wxasTSEaCKuxMwnYMCatxn+C6bcD+JQ9L70AibjeaFC+kT2qCMTtAb9A6F6CNqoS+og7qIoh/oD/qPrku/vZJX9vxNqVfa9lTRDrzqDSNKux4=</latexit> r X i L=r Y i LW R,C i <latexit sha1_base64="Mi87WVhS7UDgdG/7vmsKh8+R96E=">AAACWXicZVHLahsxFJWnaetOX069azairqGLYma6cLMJmJpCKVk4oX7hcYcrWbaFNZpB0gSMmG/rXxRK11100eYXIj8SYvuA4HDOvdx7j0gmuDZB8KvkPTh6+Ohx+Yn/9NnzFy8rx696Os0VZV2ailQNCGgmuGRdw41g0wxSIhgfbJor/z+FVOap/KbWWZsnMBM8imnYJwUV4aRBCIgtoOYFzhKwMwpCHte4DN8aw33LRutB1siclZYe2f1i5h/t5fv24VDXKkFjWANfEjCLam13ib/vr7++bkTV/5Gk5TmCZOGCtB6FAaZGVtQhlPBCj/KNcuALmDGRo5KSJge2/UmBa47ZYKnqXJPGrxW/fq9FguJ1suEuNLVuvrAXKl35n2PkAmXs2KnYZSb6enYcpnlhkm6mT/NBTYpXsWMJ1wxasTSEaCKuxMwnYMCatxn+C6bcD+JQ9L70AibjeaFC+kT2qCMTtAb9A6F6CNqoS+og7qIoh/oD/qPrku/vZJX9vxNqVfa9lTRDrzqDSNKux4=</latexit> r X i L=r Y i LW R,C i <latexit sha1_base64="V1mla1RgPBQyyXyUbdoowMbhhL0=">AAAB/3icZVDLSgMxFL3jsx1fVZdugqUgCGXGRXVZdOOyBfvAtpRMJtOGZjJDkhHL0IVrt7rwC9ypWz/FT/ArNH0grT1w4XDOPST3eDFnSjvOl7Wyura+sZnJ2lvbO7t7uf2DuooSSWiNRDySTQ8rypmgNc00p81YUhx6nDa8wdXYb9xRqVgkbvQwpp0Q9wQLGMHaSNXTbi7vFJ0J0DJxZyRfzsYvt+/3P5Vu7rvtRyQJqdCEY6VarhPrToqlZoTTkd1OFI0xGeAebRkqcEhVJ518dIQKRvFREEkzQqOJahfmIikOlRqGnlkNse6rJXOs/pnznuf5TPRGC4FWooOLTspEnGgqyPT9IOFIR2hcBvKZpETzoSGYSGZOQKSPJSbaVGabbtz/TSyT+lnRLRVLVVPSJUyRgSM4hhNw4RzKcA0VqAEBCo/wBM/Wg/VqvVkf09UVa5Y5hAVYn792X5le</latexit> + <latexit sha1_base64="r8l4jVrNcjhmyI+Hh6dipVLSdfU=">AAACC3icZVDLSsNAFL2prxpfVcGNm2ApuCqJi+qy1I3LFuwD2lAmk2k7dDKJMxOhhHyCa7f6Aa7ciUv9CDfu/QsnbZHWHhg4nHMP987xIkalsu0vI7e2vrG5ld82d3b39g8Kh0ctGcYCkyYOWSg6HpKEUU6aiipGOpEgKPAYaXvj68xv3xMhachv1SQiboCGnA4oRkpLbi9AaoQRSzppn/YLRbtsT2GtEmdOitWTxjd9qX3U+4Wfnh/iOCBcYYak7Dp2pNwECUUxI6nZiyWJEB6jIelqylFApJtMj06tklZ8axAK/biypqpZWogkKJByEnh6NLtSrpiZ+mcuep7nUz5MlwLdWA2u3ITyKFaE49n+QcwsFVpZMZZPBcGKTTRBWFD9BQuPkEBY6fpM3Y3zv4lV0rooO5VypaFLqsEMeTiFMzgHBy6hCjdQhyZguINHeIJn48F4Nd6M99lozphnjmEJxucvTP2eyA==</latexit> X i <latexit sha1_base64="V1mla1RgPBQyyXyUbdoowMbhhL0=">AAAB/3icZVDLSgMxFL3jsx1fVZdugqUgCGXGRXVZdOOyBfvAtpRMJtOGZjJDkhHL0IVrt7rwC9ypWz/FT/ArNH0grT1w4XDOPST3eDFnSjvOl7Wyura+sZnJ2lvbO7t7uf2DuooSSWiNRDySTQ8rypmgNc00p81YUhx6nDa8wdXYb9xRqVgkbvQwpp0Q9wQLGMHaSNXTbi7vFJ0J0DJxZyRfzsYvt+/3P5Vu7rvtRyQJqdCEY6VarhPrToqlZoTTkd1OFI0xGeAebRkqcEhVJ518dIQKRvFREEkzQqOJahfmIikOlRqGnlkNse6rJXOs/pnznuf5TPRGC4FWooOLTspEnGgqyPT9IOFIR2hcBvKZpETzoSGYSGZOQKSPJSbaVGabbtz/TSyT+lnRLRVLVVPSJUyRgSM4hhNw4RzKcA0VqAEBCo/wBM/Wg/VqvVkf09UVa5Y5hAVYn792X5le</latexit> + <latexit sha1_base64="Mi87WVhS7UDgdG/7vmsKh8+R96E=">AAACWXicZVHLahsxFJWnaetOX069azairqGLYma6cLMJmJpCKVk4oX7hcYcrWbaFNZpB0gSMmG/rXxRK11100eYXIj8SYvuA4HDOvdx7j0gmuDZB8KvkPTh6+Ohx+Yn/9NnzFy8rx696Os0VZV2ailQNCGgmuGRdw41g0wxSIhgfbJor/z+FVOap/KbWWZsnMBM8imnYJwUV4aRBCIgtoOYFzhKwMwpCHte4DN8aw33LRutB1siclZYe2f1i5h/t5fv24VDXKkFjWANfEjCLam13ib/vr7++bkTV/5Gk5TmCZOGCtB6FAaZGVtQhlPBCj/KNcuALmDGRo5KSJge2/UmBa47ZYKnqXJPGrxW/fq9FguJ1suEuNLVuvrAXKl35n2PkAmXs2KnYZSb6enYcpnlhkm6mT/NBTYpXsWMJ1wxasTSEaCKuxMwnYMCatxn+C6bcD+JQ9L70AibjeaFC+kT2qCMTtAb9A6F6CNqoS+og7qIoh/oD/qPrku/vZJX9vxNqVfa9lTRDrzqDSNKux4=</latexit> r X i L=r Y i LW R,C i <latexit sha1_base64="/YbyZh7FCYS7+1LpMqcFTX5TcBY=">AAACCXicZVDLSgMxFM3UVx1fVZdugqXgqsy4UDdi0Y0LFxXsA6ZDyWTSNjSTDElGKMN8gWu3uvAL3IlbF36C+An+hZm2SGsPBA7n3MO9OUHMqNKO820VlpZXVteK6/bG5tb2Tml3r6lEIjFpYMGEbAdIEUY5aWiqGWnHkqAoYKQVDK9yv3VPpKKC3+lRTPwI9TntUYy0kbxOhPQAI5beZN1S2ak6Y8BF4k5J+eLTPo9fvux6t/TTCQVOIsI1Zkgpz3Vi7adIaooZyexOokiM8BD1iWcoRxFRfjo+OYMVo4SwJ6R5XMOxaldmIimKlBpFgRnNb1QLZq7+mbNeEISU97O5gJfo3pmfUh4nmnA82d9LGNQC5rXAkEqCNRsZgrCk5gsQD5BEWJvybNON+7+JRdI8rron1ZNbp1y7BBMUwQE4BEfABaegBq5BHTQABgI8gifwbD1Yr9ab9T4ZLVjTzD6Yg/XxCx06nYA=</latexit> L <latexit sha1_base64="/YbyZh7FCYS7+1LpMqcFTX5TcBY=">AAACCXicZVDLSgMxFM3UVx1fVZdugqXgqsy4UDdi0Y0LFxXsA6ZDyWTSNjSTDElGKMN8gWu3uvAL3IlbF36C+An+hZm2SGsPBA7n3MO9OUHMqNKO820VlpZXVteK6/bG5tb2Tml3r6lEIjFpYMGEbAdIEUY5aWiqGWnHkqAoYKQVDK9yv3VPpKKC3+lRTPwI9TntUYy0kbxOhPQAI5beZN1S2ak6Y8BF4k5J+eLTPo9fvux6t/TTCQVOIsI1Zkgpz3Vi7adIaooZyexOokiM8BD1iWcoRxFRfjo+OYMVo4SwJ6R5XMOxaldmIimKlBpFgRnNb1QLZq7+mbNeEISU97O5gJfo3pmfUh4nmnA82d9LGNQC5rXAkEqCNRsZgrCk5gsQD5BEWJvybNON+7+JRdI8rron1ZNbp1y7BBMUwQE4BEfABaegBq5BHTQABgI8gifwbD1Yr9ab9T4ZLVjTzD6Yg/XxCx06nYA=</latexit> L <latexit sha1_base64="ZRgqLcPrL/cKmNxINNXlwAsIEpc=">AAACCXicZVDLSgMxFM3UVx1fVZdugqXgqsy4UDdi0Y3LKvYB06FkMmkbmkmGJCOUYb7AtVtd+AXuxK0LP0H8BP/CTFuktQcCh3Pu4d6cIGZUacf5tgpLyyura8V1e2Nza3untLvXVCKRmDSwYEK2A6QIo5w0NNWMtGNJUBQw0gqGV7nfuidSUcHv9CgmfoT6nPYoRtpIXidCeoARS2+zbqnsVJ0x4CJxp6R88Wmfxy9fdr1b+umEAicR4RozpJTnOrH2UyQ1xYxkdidRJEZ4iPrEM5SjiCg/HZ+cwYpRQtgT0jyu4Vi1KzORFEVKjaLAjOY3qgUzV//MWS8IQsr72VzAS3TvzE8pjxNNOJ7s7yUMagHzWmBIJcGajQxBWFLzBYgHSCKsTXm26cb938QiaR5X3ZPqyY1brl2CCYrgAByCI+CCU1AD16AOGgADAR7BE3i2HqxX6816n4wWrGlmH8zB+vgFJzKdhw==</latexit> R <latexit sha1_base64="ZRgqLcPrL/cKmNxINNXlwAsIEpc=">AAACCXicZVDLSgMxFM3UVx1fVZdugqXgqsy4UDdi0Y3LKvYB06FkMmkbmkmGJCOUYb7AtVtd+AXuxK0LP0H8BP/CTFuktQcCh3Pu4d6cIGZUacf5tgpLyyura8V1e2Nza3untLvXVCKRmDSwYEK2A6QIo5w0NNWMtGNJUBQw0gqGV7nfuidSUcHv9CgmfoT6nPYoRtpIXidCeoARS2+zbqnsVJ0x4CJxp6R88Wmfxy9fdr1b+umEAicR4RozpJTnOrH2UyQ1xYxkdidRJEZ4iPrEM5SjiCg/HZ+cwYpRQtgT0jyu4Vi1KzORFEVKjaLAjOY3qgUzV//MWS8IQsr72VzAS3TvzE8pjxNNOJ7s7yUMagHzWmBIJcGajQxBWFLzBYgHSCKsTXm26cb938QiaR5X3ZPqyY1brl2CCYrgAByCI+CCU1AD16AOGgADAR7BE3i2HqxX6816n4wWrGlmH8zB+vgFJzKdhw==</latexit> R Row- and Column-wise Pruned & Gradients for Low-Rank Tensors Figure 4.1: The sparse training pipeline in SLOPE. Here,X,Y, andWdenote the input, output, and the weight tensors for a specific layer, respectively.∇ · Lrepresents the gradient of the loss function.LandR are the low-rank terms that are introduced only in the final 1% iterations. SuperscriptRshows row-wise pruning usingN:Mscheme andR, Cshows both column and row-wiseN:Msparsification, leading to extra imposed zeros. Blue elements represent non-zero values, while white elements represent pruned values, and red elements indicate additional zeros introduced during the backward pass. 4.2 Additional Related Work Model pruning.Pruning the models has been one of the most effective methods to reduce the complexity of LLMs [ 58]. One can pretrain the LLMs sparsely [34] or the pruning can happen after a dense pretraining [55, 75], possibly followed by a fine-tuning stage to recover part of the lost accuracy [39,51]. Pruning the models after pretraining can be costly [127,52] and typically fails to maintain their accuracy [36,136]. While the sparse pretraining methods improve the accuracy of the model, they either use unstructured sparsity patterns that cannot be accelerated with the current hardware [ 142] or have significant overheads when searching for and applying their structured sparse masks [63,166,137]. Low-rank adapters.Low-rank adapters have emerged as a promising method to reduce the fine-tuning costs associated with pre-trained LLMs and enable more efficient task switching [ 61]. Different quantization and initialization schemes have been proposed to reduce their overheads in LLM fine-tuning [28,48]. Adding low-rank factors to sparse matrices is a low-weight mechanism widely used to improve the accuracy of approx- imations of dense matrices [12]. In machine learning, the sparse plus low-rank approximations are limited to attention heads [ 103,17] and pruning after pretraining [104,80], and the sparse plus low-rank pretraining has not been investigated. Additionally, the sparse plus low-rank fine-tuning work does not provide acceleration in both forward and backward pass of the fine-tuning process. Furthermore, the low-rank adapters in these works are added at the beginning of the fine-tuning process, adding extra overheads to the fine-tuning process. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS35 4.3 Sparse Plus Low-rank Pretraining of LLMs Equation 4.1,Equation 4.2, andEquation 4.3depict the formulas for the forward and backward pass of the i-th linear layer in a neural network. Here, the weight tensor is denoted asW i ∈R d out ×d in and the input tensor is denoted asX i ∈R b×d in . The forward pass generates an output tensor represented asY i ∈R b × d out . In all equations,d in andd out refer to the input and output dimensions of the respective layer andbrefers to the batch size. FWD→ | Y i =X i W T i (4.1) BWD−1→ | ∇ W i L=∇ Y i L T X i (4.2) BWD−2→ | ∇ X i L=∇ Y i LW i (4.3) The dimension along which N:M pruning occurs corresponds to the reduction dimension in Matrix-Matrix multiplication. Withoutthisrestriction, thesparseMatrix-MatrixoperationcannotbeacceleratedonGPU[111]. With this restriction in mind, to leverage weight sparsity in forward and backward pass, one needs to prune el- ements along the columns ofW T i inEquation 4.1(FWD) andW i inEquation 4.3. To satisfy this requirement, it is necessary to prune elements of the weight tensorW i along both row and column dimensions. 4.3.1 Double-pruned Backward Pass Various approaches can be used to exploit N:M sparsity during both the forward and backward passes. For example, one may prune the activation tensorX i in FWD along the row dimension andW i in BWD-2 along the column dimension. Although diverse combinations exist for pruning, our focus in this study is primarily on the sparsification of weight tensors for two reasons: (a) the sparsification of weight tensors directly impacts the resource required for model storage and serving, and (b) our initial findings indicate that pruning weight tensors during both forward and backward passes has a comparatively lesser adverse impact on the overall end-to-end model quality. More details on our experiments can be found in Section B.6. As such, we present adouble-pruned backward passformulation that can productively accelerate FWD and BWD-2 computations. In addition, we prove that such materialization of pruned weight tensors, despite being lossy 1 , exhibits con- vergence properties. For the rest of this chapter, we represent the weight tensor subjected to row-wise pruning asW R i , while the concurrent row-wise and column-wise pruning (double-pruned) is presented asW R,C i . We rewrite the training equations to accommodate these modifications, with proposed changes highlighted in blue: FWD→ | Y i =X i W R i T (4.4) BWD−1→ | ∇ W i L=∇ Y i L T X i (4.5) BWD−2→ | ∇ X i L=∇ Y i LW R,C i (4.6) Using this formulation for training, we can accelerate both forward and backward passes owing to the existence of N:M sparsity along both dimensions of weight tensors (seeFigure 4.1). MemoryFootprintAnalysis.InducingN:Mstructuredsparsitynotonlyimprovescomputationalefficiencyof GEMM operations but also reduces the memory footprint for storing sparse tensors. It is noteworthy, however, 1 We term this formulation “lossy” because the weight matrix undergoes information loss during the backward pass compared to its state in the forward pass. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS36 that the storage of auxiliary meta-data becomes necessary, containing information about the locations of non- zero elements in a supporting matrix.Equation 4.7delineates the requisite number of bits for storing the indices in the N:M sparsity format, where⌈.⌉denotes the ceiling function. We present the detailed results on the memory footprint reduction inSection 4.4. n N:M index = log M N (4.7) Convergence Analysis.Theorem 4.3.1(proof inSection B.14) shows the additional sparsity resulting from double pruning to an initially row-wise N:M pruned matrix. Following this lemma, we quantify the increased sparsity induced by double pruning with 1:2, 2:4, and 2:8 sparsity patterns as12.5%,9.375%, and3.39%, re- spectively. This observation underscores that as the value of M in N:M increases, the surplus of zero elements in a double-pruned matrix diminishes. This reduction in zero elements consequently implies a decrease in computational errors, enhancing the robustness of the computations. We expound further insights into this phenomenon inSection B.5. Lemma 4.3.1.Consider a randomly initialized matrixA. Following our notations, we denote the row-wise pruned version ofAbyA R and the joint column- and row-wise pruned version ofAbyA R,C . We useD(.) to present the density ratio of a matrix.Equation 4.8shows the additional zero elements in matrixAthat are introduced by double-pruning, where s = N M . D(A R )−D(A R,C ) = M X j=N+1 M j s j (1−s) M−j j−N M (4.8) Theorem 4.3.2states that the dynamic alteration of the column-wise mask inEquation 4.5during each training iteration does not exert a detrimental impact on the convergence of the optimizer. This phenomenon can be attributed to the equivalence between the left-hand side of Equation 4.9, which corresponds toEqua- tion 4.3[BWD-2], and the averaging effect achieved through multiple training iterations of backpropagation with distinct sparsity masks. However, for arbitrary values of N and M,Equation 4.4andEquation 4.5can be used in the training with convergence guarantee (proof in Section B.14). The sparsity mask is chosen randomly at initialization, i.e. all the weights have the same probability of being zero or non-zero. This is because at initialization the location of weights with larger magnitude is arbitrary. After choosing the sparsity mask at initialization, we keep the mask fixed throughout the entire training process. This policy ensures that each element in the weight has the same probability of being non-zero at initialization and satisfies the random mask assumption in Theorem 4.3.1. Theorem 4.3.2.Assuming a loss functionL(W ⟩ ,X ⟩ )for a random sampleX i , and considering a random maskM i , Equation 4.9holds, whereE[.]is the expectation operator and⊙is the element-wise multiplication. E X i [∇ X i L(W i , X i )] = M N E M i [E X i [∇ Y i L(W i , X i )(M⊙W i )]](4.9) 4.3.2 Lazy Low-rank Adapters Pruning weight tensors in FWD and BWD-2 computations is desirable for computational efficiency but may have detrimental impact on quality. To mitigate this adverse impact on model quality, we augment the doubly- pruned weight matrix with a low-rank matrix. The decomposition of the doubly-pruned weight matrix, com- bined with the low-rank matrix, maintains the computational efficiency of sparse matrix-matrix multiplication CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS37 during forward and backward passes. Simultaneously, this approach holds promise in alleviating the adverse effects of double pruning on overall model quality. Considering the dense weight matrix, denoted byW dense ∈R d out ×d in ,Equation 4.10illustrates the proposed matrix decomposition. In this expression,W sparse ∈R d out ×d in signifies a doubly-pruned matrix andL∈R d out ×r andR∈R r×d in are components of the low-rank approximation. The variablerdenotes the rank of this low-rank approximation andrfunctions as a hyperparameter that controls the trade-offs between memory footprint, computational efficiency, and model quality. W dense =W sparse +LR(4.10) The matrix decomposition of doubly-pruned matrix combined with a low-rank matrix approximation re- duces the memory footprint ofWfromd in d out tod in d out N M + (d in +d out )r, wherer << min(d in , d out ). The computational complexity of dense Matrix-Matrix multiplication, however, changes frombd in d out to bd in d out N M + b ( d in + d out ) r . Given the substantially smaller value of r in comparison to b , d in , and d out , our formulation effectively reduces both memory footprint and computational complexity of Matrix-Matrix multiplication by a factor of M N ×. We empirically show that the convergence rate of low-rank adapters surpasses that of sparse weights. We attribute this behavior to the notably lower parameter counts inherent in low-rank adapters. Leveraging this observation, we incorporate low-rank adapters exclusively during the final 1% of the training iterations. This confined usage of low-rank adapters results in additional reduction of training cost, specifically in terms of total number of operations. We term the proposed usage of low-rank adapters in the final steps of the training aslazy low-rank adapters(seeFigure 4.1). 4.3.3 Sparse Kernels cuSPARSELt is a CUDA library designed explicitly for sparse Matrix-Matrix multiplication, where one operand undergoes pruning with the 2:4 sparsity pattern. However, this library does not offer APIs for other al- gebraic routines such as addition and assignment for sparse tensors. We now delve into the details of different kernels for training and overview our implementation methodology. Algorithm 2shows the training process of a single linear layer taken from an attention-based model. We assumetheuseofweightdecayintheoptimizers, andsubsequentlydesigntherequisitesparseAPIstofacilitate the optimizer operations. The training starts with matrix initialization (line 2) and setting up sparse formats to store weight tensors and their corresponding transpose (line 3 and 4 ). Then, for every mini-batch in the training set, we compute the forward pass followingEquation 4.4(line 8). As part of the backward pass, the derivative of the loss function with respect to the output activation is computed (line 10). Subsequently, the gradients of the loss function with respect to the input activation (line 11 ) and the weight tensor (line 12) are computed usingEquation 4.5andEquation 4.2, respectively. In order to circumvent the necessity of updating weightswithzerovaluesandmitigatetheassociatedmemoryfootprintoverhead, weemployastrategywherein we mask the gradients for pruned weights. The computed values are stored in a sparse format (line 13 ). Next, in order to implement weight decay in the optimizer and mitigate the impact of gradient scaling, we compute the value of 1 γ ∇ W L+αW(line 15). Here,αis the weight decay applied in the optimizer, whileγdenotes the gradient scaling factor for numerical stability during the half-precision backward pass. The updated values for the weight tensor are calculated according to the optimizer update rule (line 16 ). Finally, the value of weight tensor and its transpose are updated directly in a sparse format (line 17and line 18). More details about the implementation of the custom kernels used in Algorithm 2can be found inSection B.7. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS38 Algorithm 2 Accelerated Sparse Pretraining Algorithm for a Linear Layer 1:Input:WeightW, training setD, weight decayα, gradient scaling factorγ. 2:Output:Updated weightWnew. 3:backend.init()▷Initialize backend 4:WSparseTranspose←backend.setup(W T )▷Setup transpose for sparse matrix multiplication 5:WSparse←backend.setup(W)▷Setup sparse weight matrix 6:sparseMask←(WSparse̸= 0)▷Element-wise mask for sparsity 7:for each training example ( X , ˆ Y)∈Ddo 8:Forward Pass: 9:Y←backend.spmm(X,WSparseTranspose) 10:Backward Pass: 11:∇ Y L←gradOutput▷Gradient w.r.t. output 12:gradInput←backend.spmm(gradOutput,WSparse) 13:gradWeight←backend.matmul(gradOutput T ,X) 14:gradWeightSparse←backend.pruneAndCompress(gradWeight,sparseMask) 15:Optimizer with Weight Decay: 16:g←backend.sparseAdd(gradWeightSparse,WSparse, 1 γ ,α) 17:Wnew←optimizer.updateWeight(g) 18:backend.updateSparseMatrix(WSparse,Wnew) 19:backend.updateSparseMatrix(WSparseTranspose,Wnew T ) 20:end for 21:Return:Wnew. 4.3.4 SLOPE Runtime Optimization While SLOPE improves the training and inference of LLMs by introducing sparse weights and low-rank adapters, a naïve implementation can hinder its full performance improvement. Specifically, cuSPARSELt [ 110] SpMM kernels exhibit sensitivity to input and weight tensor shapes, and introducing low-rank adapters at inference can increase the number of calls during the forward pass of each linear layer. This section covers our approach to optimize SLOPE’s implementation and further improve model performance. Efficient tiling of upsample tensors. Figure 4.3-(a)showcases the speedup achieved by the cuSPARSELt backend across a range of tensor shapes commonly used in LLMs. While the speedup of SpMM in downsam- ple tensors increases gradually as their sizes increase, the speedup of upsample tensors drops off at around hidden dimension = 4000. To overcome this limitation, we tile the upsample tensor into multiple smaller matrices of equal size, each of which benefits from improved speedup when multiplied by the input using 2:4 sparsity. By tuning the size of the tiles, we discovered that the best performance can be achieved by us- ing square tiles. The results of these multiplications are then concatenated. This optimization, as detailed in Section 4.4.3, leads to a12%improvement in inference speed and a4%increase in training speed with SLOPE. Efficient kernel for combined SpMM+low-rank adapters.A straightforward implementation of low-rank adapters requires four kernel calls: one for sparse matrix multiplication, two for low-rank computations, and one for adding the results. In addition, our experiments demonstrate that multiplying matrices with low-rank adapters does not scale proportionally with the adapter’s rank, leading to significant overheads due to their low arithmetic intensity (see Section 4.4.3). To address this, we introduce two optimizations:(1)concatenating the downsample tensor to the sparse weight tensor, reducing kernel calls and increasing arithmetic intensity as in Equation 4.11-left, and(2)leveraging a cuBLAS fused matrix multiplication and addition kernel, minimizing cache access and kernel calls as inEquation 4.11-right. As demonstrated inSection 4.4.3, these optimizations CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS39 collectively contribute to a speedup improvement of up to 6% in the end-to-end inference speed. [Y 1 |Y 2 ] =X[W T |L];Y=Y 2 R+Y 1 (4.11) 4.4 Experimental Results This section evaluates the efficacy of SLOPE in accelerating the pretraining while achieving memory savings. Due to the substantial computational resources required for LLM pretraining, our accuracy evaluation is pri- marily focused on smaller-scale LLMs up to 774M parameters. However, the speedup and memory reduction results extend to a wider range of models, from 2.6B up to 66B parameters. ExperimentSetup.Our experiments were conducted on the Narval and Mist clusters at Compute Canada [22] and the Lonestar 6 cluster at the Texas Advanced Computing Center [141]. Each Narval node is equipped with four Nvidia A100 GPUs (40GB), Mist nodes feature four Nvidia V100 GPUs (32GB), and Lonestar 6 nodes have three Nvidia A100 GPUs (40GB). For accuracy experiments, we emulated 2:4 and N:M sparsity using custom-designed, low-overhead CUDA kernels to prune weights in both the forward and backward passes, utilizing a mixture of available resources across clusters since model accuracy is not hardware-dependent. Speedup and memory saving experiments were conducted on a single A100 GPU in the Narval cluster over 1000 iterations, reporting the median to mitigate outlier effects; memory reduction experiments were run five times with the median reported. We employed the default hyperparameters from the NVIDIA BERT codebase [112] and the FlashAttention GPT codebase [26,24]. Training BERT-Large-Uncased required ap- proximately 32 hours on 64 A100-64GB GPUs, while pretraining GPT2-Small/Large took 32 and 111 hours on 64 V100-32GB GPUs, respectively. 4.4.1 End-to-end Speedup and Memory Saving: Pretraining and Inference We evaluate the speedup and memory reduction by SLOPE during pretraining and inference across LLMs with different model parameter sizes. To demonstrate the scalability and efficiency of our method, we conducted extensive benchmarking on OPT (2.6B to 66B), LLaMA-3-8B and Mistral-v0.3-7B models. In all the exper- iments, we have enabled FlashAttention-2 [ 24] (Section B.9presents detailed ablation study on the impact of FlashAttention). To mitigate the impact of outliers, we conducted 1,000 iterations for each speedup experi- ment and reported the median value. For the memory reduction experiments, we performed five independent runs and similarly reported the median outcome. These methodologies were chosen to provide a more reliable measure of central tendency in our results 2 . We compared our method against dense pretraining and inference directly in PyTorch, which uses efficient cuBLAS backend. As the sparse pretraining benchmark, we compare our work against Fully Sparse Training (FST) [ 62], the state-of-the-art 2:4 pretraining method and the only semi-structured sparse pretraining work that provides end-to-end speedups. Note that methods targeting LLM pretraining with N:M sparsity often suffer from inefficiency due to mask search overheads and/or compression setup. Section B.4andSection B.2 2 It is noteworthy that for benchmarking speedup and memory savings, which require comparatively fewer computational resources than comprehensive pretraining accuracy experiments, we utilized the OPT, LLaMA-3, and Mistral-v0.3 model families. These families were selected due to their diverse range of model parameter sizes, allowing for a more thorough study of performance across different scales. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS40 Table 4.1: Comparative analysis of end-to-end pretraining and inference speedup (×) comparison between SLOPE and the latest work (FST) on accelerating pretraining with 2:4 sparsity (ICML 2024) [62]. The baseline is dense PyTorch implementation of the models with CUBLAS backend. Note that the lack of inference speedup in FST is because of the final dense pretraining during the final iterations, resulting in a dense model for inference. E-SR-STE stands for Extended SR-STE. MODELMETHOD TRAININGINFERENCE NO ADAPTER (r= 0)NO ADAPTER (r= 0) 1.56% ADAPTER 6.25% ADAPTER OPT-66B SLOPE1.201.461.431.40 FST1.061.001.001.00 OPT-30B SLOPE1.221.531.531.50 FST1.071.001.001.00 OPT-13B SLOPE1.251.541.391.36 FST1.101.001.001.00 OPT-6.6B SLOPE1.211.461.461.43 FST1.111.001.001.00 OPT-2.6B SLOPE1.131.311.251.18 FST1.091.001.001.00 LLAMA-3-8B SLOPE1.161.351.331.32 FST1.091.001.001.00 MISTRAL-V0.3-7B SLOPE1.151.341.321.31 FST1.071.001.001.00 detail the profiling in Bi-Mask [166] and FST [62], which similarly use N:M sparsity on both forward and backward passes. Notably, our approach, SLOPE, diverges significantly from recent work Fully Sparse Training (FST) [62] in three key aspects. Firstly,we comprehensively prune all weights in the model, encompassing both MLP and Self-Attention modules, whereas FSTonly prunes weights in the MLP modules. Secondly, FSTemploys dynamic transposable weights, which introduce additional computation and memory overhead during training. Thirdly, FST necessitates dense fine-tuning (∼17% of pretraining), thereby negating their speedup advantages during inference. In contrast, our approach achieves efficient and accurate large language models during both training and inference without such limitations. SLOPESpeedupforPretrainingandInference. Table4.1summarizesthespeedupsachievedbyourmethod during both training and inference. Since over 99% of training occurs without low-rank adapters, the training speedup is largely independent of the adapter rank. Conversely, inference speedup is directly influenced by the adapter rank. Given the varying hidden dimensions across different model sizes, we report the inference speedup for various adapter rank ratios: adapter−rank hidden−dimension . Figure 4.3-(a)illustrates that cuSPARSELt achieves higher speedups for large matrices until it reaches its maximum performance capacity (2×). A similar trend is observedin the pretraining and inference speedups of the models. For small matrices used in low-rank adapters, the lower arithmetic intensity of low-rank adapter multiplication results in higher overhead relative to sparse multiplication. This is because low arithmetic intensity limits the full utilization of GPU resources, leading to inefficiencies. SLOPE Memory Reduction in Pretraining and Inference.For training, the memory consumption of a dense model includes weights, gradients, and optimizer states, amounting to4×16bits for weights,4×16 bits for gradients, and2×4×32bits for optimizer states. The sparse model, however, stores non-zero weights and indices twice (for both weights and transposed weights), along with a binary mask, gradients, and reduced optimizer states. This adds up to2×(16 + 3)bits (weights and transposed weights),4×8bits (binary mask), 2×16bits (gradients), and2×2×32bits (optimizer states). Consequently, the memory footprint during training is reduced by 68%. For inference, a dense model requires storing weights with a total memory cost of CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS41 Table 4.2: Comparative analysis of end-to-end memory reductions (×) during training and inference between SLOPE and the latest work (FST) on accelerating pretraining with 2:4 sparsity (ICML 2024) [62]. Values greater than1.00×show memory overhead. MODELMETHOD TRAININGINFERENCE NO ADAPTER (r= 0)NO ADAPTER (r= 0) 1.56% ADAPTER 6.25% ADAPTER OPT-66B SLOPE0.670.630.650.70 FST1.271.001.001.00 OPT-30B SLOPE0.670.610.630.69 FST1.171.001.001.00 OPT-13B SLOPE0.680.510.620.68 FST1.161.001.001.00 OPT-6.6B SLOPE0.680.600.620.68 FST1.191.001.001.00 OPT-2.6B SLOPE0.670.620.640.70 FST1.181.001.001.00 LLAMA-3-8B SLOPE0.630.66 0.69 0.71 FST1.171.001.001.00 MISTRAL-V0.3-7B SLOPE0.680.660.690.65 FST1.151.001.001.00 4×16bits. In contrast, our sparse model optimizes memory usage by storing only the non-zero weights and their indices, resulting in2×16bits for non-zeros and three bits for indices (see Equation 4.7). This leads to a 54% reduction in memory usage during inference. Table 4.2presents the memory reduction for different low-rank adapter ranks and OPT, LLaMA-2, and Mistral model variants. The memory reduction is slightly less than the theoretical expectation, primarily because of additional memory usage from other model components, such as layer norms, and dense model parameters. 4.4.2 Pretraining Accuracy Results To assess the impact of SLOPE on model accuracy, we conducted pretraining experiments across various models and datasets. In all experiments, the classification heads and the first linear layer following the input are dense. GPT2 (Small/Large).We pretrained both the small (117M parameters) and large (774M parameters) vari- ants of GPT2 [ 120] on the OpenWebText dataset [44]. For a fair comparison, we evaluate the models on MMLU [57], Arc Challenge [21], and OpenBookQA [97] zero-shot tasks implemented in Language Model Evaluation Harness [41]. Additionally, we evaluate the validation perplexity of the models following the same experimental settings described in FlashAttention [ 26,24]. We compare SLOPE against two state-of-the-art sparse pretraining methods, including (a) Wanda [136]→a one-shot pruning technique, (b) Extended SR- STE [168,62]→a dynamic mask pretraining method for N:M sparsity, which serves as the foundation of follow-up work [63,166,62]. Please note that SR-STE only supports stochastic gradient descent optimization, and FST extended it to other optimizers. We use the extension provided by FST in our work, and call it Ex- tended SR-STE. The difference between Extended SR-STE and FST is that FST requires dense pretraining (fine-tuning) in the last 17% of pretraining and only prunes the MLP layers of the model, while SR-STE is fully sparse and prunes both the MLP and the Self-Attention layers of the model. Figure 4.2compares the validation perplexity 3 and zero-shot accuracy of GPT2-Small and GPT2-Large 3 Perplexity is a standard metric for evaluating language models. Intuitively, perplexity measures how “surprised” the model is by the held-out text: lower values indicate better predictive accuracy. A perplexity ofkcan be loosely interpreted as the model being as uncertain as if it were choosing uniformly amongkcandidates at each step [18]. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS42 050100150200250300350400 Iteration (×1000) 20 30 40 50 60 70 80 Perplexity GPT2 Small Perplexity Dense 2:4 - SloPe 2:4 – Wanda 2:4 – SR-STE –γ w =6e-6 2:4 – SR-STE –γ w =4.5e-4 2:4 – SR-STE –γ w =2e-4 050100150200250300350400 Iteration (×1000) 10 15 20 25 30 35 40 Perplexity GPT2 Large Perplexity Dense 2:4 – SloPe 2:4 – Wanda Figure 4.2: Validation perplexity of GPT2-Small and GPT2-Large on OpenWebText.γ w shows the value of the decay factor parameter in Extended SR-STE (FST). Table 4.3: GPT2-Small accuracy results on zero-shot tasks. Adapter rank is the ratio of the low-rank adapter to the hidden dimension of the model. For Extended SR-STE, we have used a decay factor of6×10 −6 , since it resulted in the lowest perplexity in OpenWebText. The best performing sparse configuration is highlighted in bold. METHOD ADAPTER MMLU↑ ARCOPEN-WINO- HELLA- MATHQA↑PIQA↑RACE↑ RANKCHALLENGE↑BOOKQA↑GRANDE↑SWAG↑ DENSEN/A22.920.716.250.628.521.859.8 28.4 SLOPE 2.1%23.019.316.450.827.520.857.6 27.2 0.05%23.019.416.250.527.420.857.5 27.1 023.019.316.050.127.520.857.4 27.1 EXTENDED2.1%24.218.314.247.526.921.455.2 24.2 SR-STE 0.05%24.118.414.247.526.821.254.5 24.2 024.118.312.647.526.921.254.8 24.0 across a range of sparse pretraining methods with different hyperparameters. We have additionally added lazy low-rank adapters to Extended SR-STE [168] to show the effectiveness of our approach in other methods and also compare both methods with more similar settings. While a gap in perplexity consistently exists between sparse and dense models, SLOPE achieves a lower perplexity compared to Wanda [ 136] and Extended SR- STE. Additionally,Table 4.3summarizes the achieved accuracy of the models on zero-shot tasks, showing that SLOPE is consistently achieving a higher accuracy in comparison to Extended SR-STE. Moreover, adding lazy low-rank adapters can benefit both static and dynamic training methods. This improved accuracy stems from SLOPE’s efficient allocation of the training budget. Specifically, Extended SR-STE, with its dynamic pruning masks, expends a significant portion of its training budget (e.g. gradient updates) updating weights that may be ultimately pruned and not used at inference, leading to wasted resources. Section B.1provides further details and supporting evidence forthis observation. Additional validationresults forGPT experiments on GLUE dataset are also provided in Section B.11andSection B.10. BERT-Large-Uncased.We pretrain BERT-Large-Uncased [30] (355M parameters) and fine-tune it for var- ious question-answering and text classification tasks, following a similar approach to [112,99,116] for both pretraining and fine-tuning. Section B.3provides details on the pretraining and fine-tuning process. We eval- uate the performance of BERT-Large-Uncased on the SQuAD v1.1 [123] and GLUE [147] tasks. We report the average metric score for GLUE and present the task-specific metrics inSection B.8. Please note that in all CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS43 Table 4.4: SQuAD-v1.1 accuracy and GLUE results on BERT-Large-Uncased with different adapter ranks. GLUE results are reported as the average metric score across all tasks.rdenotes the ratio of the low-rank adapter to the hidden dimension (1024). DATASET DENSEr= 0r= 0.39%r= 1.56%r= 6.25% SQUAD 90.44 89.189.189.289.5 GLUE 80.22 77.477.777.878.2 the experiments corresponding to BERT-Large-Uncased, when using Wanda, we have fine-tuned the model after pruning to improve the accuracy of the models, since using Wanda alone led to extremely low accuracy results. Effects of Low-rank Adapters.To understand the impact of low-rank adapters on pretraining performance, we conducted ablations using low-rank adapter ranks of 4, 16, and 64 for 1% of the total number of iterations. These ranks represent up to6.25%of the model’s hidden dimension. Table 4.4shows the results of these settings on SQuAD and GLUE downstream tasks. We present per-task metrics for GLUE inSection B.8. As expected, adding low-rank adapters improves the model’s final accuracy across all tasks. Additionally, higher ranks improve the model’s performance at the cost of increased computational requirements. It is also worth noting that incorporating low-rank adapters only in the final iterations (1% of total iterations) is sufficient to recover pretraining accuracy. Convergence Rate of Low-rank Adapters.We hypothesized that low-rank adapters would converge faster due to their significantly fewer learnable parameters. To test this, we introduced low-rank adapters in the second phase of BERT-Large-Uncased pretraining and monitored their convergence rate.Figure 4.3shows the cosine similarity of the adapters, with the downsample adapter converging rapidly within 100 iterations and the upsample adapter converging slightly slower. Despite this, limiting training to 100 iterations still yields comparable results on downstream tasks. 2000400060008000 Hidden Dimension 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 Speedup cuSPARSELt SpMM Speedup Attention Upsample Downsample Upsample Tiling (Ours) 0200400600800100012001400 Iterations 0.0 0.2 0.4 0.6 0.8 1.0 Cosine Similarity Low-Rank Adapter Similarity with Converged Weight Query Key Value Projection Upsample Downsample (a)(b) Figure 4.3:(a)The speedup achieved using cuSPARSELt backend in PyTorch for Attention (d out =d in ), Upsample (d out = 4d in ) and Downsample (d out = d in 4 ) matrices with a batch size of 2048.(b)The cosine similarity of the low-rank adapters and the converged adapters for different layers in the model. The cosine similarities are averaged among the 24 layers of BERT-Large-Uncased. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS44 Effects of Mixed N:M sparsity.To study the sensitivity of different blocks to varying sparsity ratios and to assess their relative importance, we experiment across a range of configurations: (a) [2:4-2:4]→uniformly applying 2:4 sparsity across all layers (b) [2:4-2:8]→applying 2:4 sparsity pattern to the first 12 blocks and a 2:8 sparsity pattern to the last 12 blocks and (c) [2:8-2:4] → we reverse the sparsity ratios for the first and last 12 blocks. Note that, to reduce computational costs, we use the same dense checkpoint for Phase-1 in all settings and a low-rank adapter of rank 40 for all models. We also replicate this experiment using Wanda [136] and report the comparison results. Table 4.5: SQuAD-v1.1 accuracy results on BERT-Large-Uncased for different sparsity settings. SPARSITY PATTERNSQUAD SQUAD GLUE GLUE (FIRST 12 BLOCKS - LAST 12 BLOCKS) SLOPE WANDA SLOPE WANDA 2:4-2:490.1789.9379.0878.84 2:4-2:889.8589.5579.0377.24 2:8-2:489.6786.5775.9269.08 Table 4.5summarizes the GLUE and SQuAD results for these settings. As the results show, increasing the sparsity ratio reduces the accuracy of the model on all tasks. But when the first 12 blocks of the model are pruned, the accuracy drop is significantly higher, especially on the GLUE dataset. We conclude that the first blocks of the model are more sensitive to sparsity during pretraining, but one can sparsify the last blocks of LLMs more aggressively. We observe a similar pattern in Wanda results as well, but Wanda performs consistently worse than SLOPE in these cases. Effects of sparsification on different modules.Each block in LLMs consists of a self-attention module and an MLP module, each containing multiple linear layers. We have analyzed the sensitivity of SLOPE to pruning each of those modules. Our results inSection 4.4.3demonstrate that SLOPE can sustain competitive quality results while pruning all modules in the model. 4.4.3 Ablation Studies Low-Rank Adapter Performance: Scaling and Arithmetic Intensity.As discussed in Section 4.3.4, the computation time of low-rank adapters doesnotscale linearly with their rank. This section provides exper- imental results to illustrate this behavior in more detail. The computational complexity of low-rank matrix multiplications isO(brd), whereb,r, anddrepresent the batch size, low-rank, and input/output dimensions of the layer, respectively. Based on this complexity, we expect the computation time to be a linear function of r. In other words, reducingrby a factor ofαshould result in a correspondingα-fold reduction in computation time. However, in practice, this linearity does not hold. This deviation arises because the assumption under- lying this expectation – that matrix multiplication is compute-bound – is not always true. Specifically, the arithmetic intensity of the operation can fall below the machine’s balance point, as described in the Roofline model [ 152] inSection 2.2.2.Figure 4.4shows the speedup achieved for different low-rank values using Py- Torch’s matrix multiplication function, which relies on the CUBLAS backend [ 107]. The figure demonstrates that the achieved speedups are significantly lower than the ideal linear scaling, particularly when reducing the rank. Moreover, it is evident that as the matrix dimensions increase, the gap between the ideal speedup and the observed speedup diminishes. This behavior can be attributed to the increased arithmetic intensity for larger matrices, leading to better utilization of tensor cores. CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS45 0.010.020.050.11.0 Rank 0 20 40 60 80 100 Speedup Low-Rank Matrix Multiplication Speedup Ideal Speedup Dim: 768 Dim: 1024 Dim: 2048 Dim: 2560 Dim: 4096 Dim: 5120 Dim: 7168 Dim: 9216 Figure 4.4: The speedup achieved by low-rank adapters in comparison to a dense matrix-multiplication. Table 4.6: End-to-end speedup (×) before and after efficient implementation of low-rank adapters. MODEL 1.56% ADAPTER 6.25% ADAPTER BEFORE AFTER BEFORE AFTER OPT-66B 1.15 1.20 1.12 1.19 OPT-30B 1.13 1.18 1.10 1.16 OPT-13B 1.11 1.10 1.09 1.10 OPT-6.6B 1.07 1.12 1.06 1.11 OPT-2.6B 1.01 1.06 0.97 1.00 Efficient Low-rank Adapter Implementation.As discussed inSection 4.3.4, a naïve implementation of low-rank adapters can lead to significant performance overheads due to the increased number of kernel launches and the low arithmetic intensity of their multiplications. To address these issues, we introduced two key optimizations: (1) concatenating one of the low-rank adapters with the sparse weights, and (2) fusing the multiplication of the other low-rank adapter with the subsequent result addition. These optimizations reduce kernel calls and increase arithmetic intensity, leading to more efficient utilization of GPU resources. Ta- ble 4.6summarizes the speedup improvements achieved with these optimizations, demonstrating an inference speedup increase of up to 6%. Efficient Weight Tiling Implementation.We observed that the dimensions and aspect ratios of matrices significantly influence system speedup (Section 4.3.4). To mitigate this, we implemented a matrix tiling strategy, dividing upsample matrices into multiple square matrices. This approach significantly improves performance, as shown in Table 4.7. Our results demonstrate that matrix tiling can enhance training speed by up to 4% and inference speed by up to 12%, highlighting its effectiveness in optimizing system performance. SLOPE Sensitivity to Pruning Different Modules in Transformer.LLMs typically consist of two main modules: theMLPandtheself-attention. Theattentionmodule’sweightsarerepresentedasamatrixinR d×3d , CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS46 Table 4.7: End-to-end speedup (×) before and after splitting the upsample matrix. In both cases, the optimiza- tion discussed inTable 4.6is used. MODEL TRAINING INFERENCE NO ADAPTER INFERENCE 1.56% ADAPTER INFERENCE 6.25% ADAPTER BEFORE AFTER BEFORE AFTER BEFORE AFTER BEFORE AFTER OPT-66B 1.10 1.13 1.22 1.34 1.20 1.31 1.19 1.30 OPT-30B 1.09 1.14 1.23 1.32 1.18 1.28 1.16 1.27 OPT-13B 1.10 1.12 1.23 1.30 1.10 1.30 1.10 1.12 OPT-6.6B 1.08 1.08 1.21 1.19 1.12 1.13 1.11 1.12 OPT-2.6B 1.03 1.02 1.02 1.07 1.06 1.05 1.00 1.00 Table 4.8: SQuADv1.1 results on BERT-Large-Uncased for different pruned modules. PRUNED MODULESSQUAD GLUE DENSE90.44 80.22 MLP90.28 79.03 MLP + SELF-ATTENTION 89.35 77.72 while the MLP uses weights inR d×4d andR 4d×d , whereddenotes the hidden dimension. To investigate the impact of sparsity on these modules, we conducted two experiments during Phase-2 of BERT-Large-Uncased pretraining: (a) [MLP]→pruning only MLP modules, and (b) [MLP + SELF-ATTENTION]→pruning both MLP and self-attention modules. Table 4.8presents the SQuAD and GLUE results for these settings. As expected, we observe a consistent, albeit slight, decrease in model quality as more modules are sparsified. The marginal decrease in performance suggests that models are relatively insensitive to the specific modules being pruned when using our SLOPE pretraining method. This observation underscores the robustness of our approach and its ability to maintain competitive quality across diverse sparsity configurations. 4.5 Conclusion In this chapter, we presented SLOPE, a method that successfully executes the first strategy of the Compression Trinity: accelerating the computational cost of each pretraining iteration. By innovatively combining the spar- sity pillar (via the double-pruned backward pass) and the low-rank pillar (via lazy adapters), SLOPE overcomes the rigidity of traditional sparse training. It delivers efficient N:M sparsity acceleration in both forward and backward passes while recovering model capacity through targeted low-rank updates. Our results demonstrate that this joint approach achieves up to 1.25×speedup in pretraining and 1.54×speedup in inference, while reducing memory footprints by 0.63×and 0.61×, respectively. Together with MKOR (Chapter 3), these contributions conclude our exploration of the Pretraining life- cyclestage. WehavedemonstratedthattheCompressionTrinitycaneffectivelyacceleratetrainingbyattacking the problem from two orthogonal angles: reducing the number of iterations via a Trinity-enhanced optimizer (MKOR), and reducing the cost per iteration via Trinity-enhanced weight structures (SLOPE). The narrative now shifts to the second major stage of the LLM life-cycle: Post-Training Compression for Inference. While SLOPE produces efficient sparse models, the ultimate goal of the Trinity is to jointly apply all three pillars, i.e., Sparsity, Low-Rank,andQuantization, to maximize inference efficiency on commodity CHAPTER 4. SLOPE: DOUBLE-PRUNED SPARSE PLUS LAZY LOW-RANK ADAPTER PRETRAINING OF LLMS47 hardware. However, applying these aggressive compression techniques simultaneously to a pre-trained, static model introduces a new challenge: compounded error. Before we can achieve the full Trinity in a one-shot inference setting, we must first establish a stable foundation. The next chapter introduces OPTIMA, where we rigorously perfect the sparsity pillar to withstand the pressures of joint compression. Chapter 5 OPTIMA: Optimal One-Shot Pruning for LLMs via Quadratic Programming Reconstruction Publication and Contributions.The content of this chapter is based on the paper “OPTIMA: Optimal One- Shot Pruning for LLMs via Quadratic Programming Reconstruction,” [98] 2025. This work was conducted in collaborationwithSamuelKushnir, AmirYazdanbakhsh, andMaryamMehriDehnavi. MohammadMozaffari conceived the project, led the implementation, and designed and executed the experiments. Samuel Kushnir contributed to the design of the algorithm. Amir Yazdanbakhsh and Maryam Mehri Dehnavi supervised the project and contributed to the writing and revision of the manuscript. 5.1 Introduction Having addressed the computational bottlenecks of pretraining in Chapter 3andChapter 4, we now turn our attention to the second major stage of the LLM life-cycle: inference. As discussed inChapter 2, efficient inference is primarily constrained by memory bandwidth. While the Compression Trinity advocates for the joint application of sparsity, quantization, and low-rank approximations, blindly applying these methods to a pre-trained model carries significant risk. This is particularly true forsparsity, which we identified in Section 1.3as the most structurally destructive pillar of the Trinity. Unlike quantization, which preserves the network’s topology, sparsity deletes connections entirely. If this structural skeleton is flawed, no amount of subsequent quantization or low-rank adaptation can recover the lost information. Therefore, before we can integrate the full Trinity, we must first maximize the accuracy of the sparse foun- dation. In this chapter, we operate under a strictresource-constrained regime: we assume the practitioner requires a ”one-shot” solution with no access to the full training pipeline or budget for backpropagation. The goal is to determine the mathematical limit of reconstruction accuracy achievable using only a small calibra- tion dataset and static, layer-wise optimization. Post-training one-shot pruning [ 58], which removes parameters from a pretrained model using only a small calibration dataset, offers a potential solution. However, current methods are often forced to choose between efficiency and reconstruction optimality. We can categorize existing approaches into two tiers. The first 48 CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 49 tier consists of fast, metric-based selectors like Wanda [136], magnitude pruning [53], and ProxSparse [83]. While computationally cheap, they perform no weight updates to compensate for removed connections, treat- ing weights as independent and ignoring the correlations captured by the loss landscape curvature. Con- versely, principled second-order approaches like Optimal Brain Surgeon [54] theoretically recover accuracy but are computationally infeasible at modern LLM scales. As a result, the second tier includes methods like SparseGPT [36] and Thanos [64], which attempt to adjust the remaining weights. However, to maintain speed, these methods rely on greedy approximations (e.g., iterative coordinate descent or localized Cholesky updates) rather than solving for the global optimum. Consequently, they leave significant performance on the table, an error that becomes significant when compounded with the quantization noise introduced in later chapters. 1 To resolve this, we introduce OPTIMA, a practical one-shot post-training pruning framework that com- bines layer-wise optimality with accelerator-grade efficiency. Distinct from prior heuristics, we formulate the weight update not as an approximation, but as an exact constrained Convex Quadratic Program (QP). This ap- proach draws a direct parallel to the second-order optimization methods discussed inChapter 2(e.g., KFAC). JustasKFACleveragesthecurvatureofthelosslandscape(viatheFisherInformationMatrix)toimprovetrain- ing convergence over first-order methods, OPTIMA utilizes the exact curvature of the layer-wise objective (via the Hessian matrix) to minimize pruning error. The core of our methodology relies on a precise reformulation of the layer-wise reconstruction step. We observe that after fixing a binary mask for a weight matrix, the layer-wise output reconstruction (least-squares) objective decomposes across columns. We exploit a fundamental algebraic property of Transformer linear lay- ers: while the linear constraints differ for each column (dictated by the mask), every column in the same layer shares the same Hessian matrixH=X ⊤ X. Unlike previous works that approximate this Hessian to sim- plify computation, we use this exact structure to formulate the update for each column as a small constrained QP. This shared-Hessian structure allows us to guarantee per-column global optimality for the reconstruc- tion objective, strictly outperforming greedy heuristics without making additional assumptions about weight independence. Realizing this formulation in practice requires careful numerical and systems engineering. We adopt a first-order primal–dual QP solver (rAPDHG [ 88]) that is well-suited to our constrained problems, as its critical operations reduce to efficient matrix–vector products with the shared Hessian. We further avoid explicit dense equality matrices by enforcing fixed entries via tight bounds, accumulate layer Hessians incrementally from calibration sequences to save memory, and solve columns in batches so thousands of small QPs are processed in parallel. These implementation choices make OPTIMA not only theoretically principled but also practical to run on a single accelerator. We evaluate OPTIMA across multiple model families (LLaMA, Gemma, and others) and sparsity regimes, including unstructured and 2:4 semi-structured sparsity. OPTIMA is designed to be modular; it acts as a drop-in replacement for the weight-update step in existing mask selectors (e.g., Wanda, SparseGPT, Thanos), consistently improving their zero-shot performance. Across six zero-shot downstream benchmarks in the Language Model Evaluation Harness, we observe up to 3.97% absolute gains on downstream tasks without any post-pruning fine-tuning. In summary, our contributions are: •We present a column-wise QP reformulation of the post-training reconstruction problem that yields per- column global optimality under a shared-Hessian model and is provably equivalent to the least-squares objective after mask selection (Section 5.4). 1 For a more detailed discussion of the related work, seeSection 5.2. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 50 ... Shared Hessian ... Prune QP Solver Calibration Data Figure 5.1: OPTIMA generates a shared Hessian among the different columns of the pruned weight using a small calibration dataset. Then, the weights in different columns will be updated in parallel using a QP solver and the shared Hessian. •We design and implement an accelerator-friendly QP solver pipeline that accumulates a single Hessian per layer, enforces mask constraints via bounds, batches thousands of column QPs, and leverages rAPDHG/M- PAX for efficient execution on GPUs/TPUs (detailed inAlgorithm 3). •We demonstrate the modularity of OPTIMA, showing it can be used as a drop-in weight-update step with common mask selection algorithms (Wanda, SparseGPT, Thanos), consistently improving their accuracy without fine-tuning ( Section 5.5). •We provide extensive empirical evidence and practical measurements. OPTIMA yields substantial average accuracy gains across tasks and model sizes (up to 3.97%), demonstrates robustness at high sparsity (up to 60%), and can prune billion-parameter models on a single H100 in less than 40 hours. 5.2 Additional Related Work Model pruning compresses trained neural networks by eliminating redundant weights, thereby lowering com- putational and memory requirements during deployment. The field primarily divides into two categories: layer-wise pruning, exemplified by Optimal Brain Surgeon (OBS) [ 55], and end-to-end pruning, represented by Optimal Brain Damage (OBD) [75]. We review these approaches in the following subsections, beginning with layer-wise methods. Layer-wise Model Pruning.Layer-wise pruning optimizes models by targeting redundancies within indi- vidual layers, assuming that local error reductions aggregate to minimize overall model degradation. Optimal Brain Surgeon (OBS) [ 55] formalizes this by identifying the least salient weight per layer and adjusting re- maining weights to offset its removal [35]. However, OBS’s computational intensity hinders its application to billion-parameter LLMs, necessitating approximations. SparseGPT [ 36] pioneered scaling OBS to LLMs by framing pruning as sparse regression problems solved approximately, trading some accuracy for efficiency. Thanos [64] refines this with multi-column pruning to cut approximation errors. In contrast, Wanda [136] em- ploys a saliency metric combining weight magnitudes and activation data from calibration sets, yielding strong results with minimal pruning time. Nonetheless, Wanda lacks mechanisms to update weights post-pruning, CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 51 opening avenues for enhancements, particularly in end-to-end methods that consider global interactions. End-to-endModelPruning.Unlikelayer-wisemethods, end-to-endpruning, exemplifiedbyOptimalBrain Damage (OBD) [75], identifies least-important weights globally by leveraging second-order derivatives of the loss function, yielding higher accuracy than OBS. However, computing these derivatives is resource-intensive, demanding approximations [99]. WoodFisher [134] employs Kronecker factorization to approximate the Hes- sian, easing computation but still faltering at LLM scales. More recently, MaskLLM [33] sidesteps second- order information by recasting pruning as a classification problem solved via standard optimizers like AdamW [86], achieving top performance at 2:4 sparsity. ProxSparse [83] reduces the costs of MaskLLM by using reg- ularizers instead of training the model on a classification task, trading accuracy for speed. Yet, its optimization demands far exceed those of one-shot pruning, constraining real-world use and highlighting the value of inte- grating with other compression strategies. Other Model Compression Methods.In addition to pruning, several orthogonal techniques enable model compression and can be integrated with pruning for compounded benefits. Quantization reduces parameter precision to lower-bit representations, as surveyed in [ 43,125], minimizing memory footprint without severe accuracy loss. Low-rank adapters, such as those in [100,48,101], decompose weight matrices into lower-dimensional factors, while knowledge distillation [46] transfers knowledge from larger teacher models to compact students. These methods complement pruning by addressing different aspects of redundancy, paving the way for hybrid frameworks in advanced compression research. 5.3 Preliminaries Post-trainingpruning(PTP)compressespre-trainedmodelswithoutretraining,usingasmallcalibrationdataset to produce a sparse model that preserves performance. To make PTP tractable, the problem is decomposed into independent layer-wise subproblems. For layerl, the goal is to find a binary sparsity maskM l and updated weights ˆ W l that minimize the output reconstruction error given original weightsW l and input activationsX l . This task can be formulated as in Equation 5.1, where⊙denotes the Hadamard product, andM l is a binary tensor of the same shape asW l with 0s for pruned weights and 1s for retained ones.Equation 5.1is solved sequentially across layers, withX l as the pruned output from layerl−1. Finding the optimalM l is NP-hard, motivating heuristics. argmin M l , ˆ W l ∥X l W l −X l (M l ⊙ ˆ W l )∥ 2 F (5.1) A common heuristic decouples mask selection from weight updates. After selectingM l (e.g., by magni- tude), the problem simplifies toEquation 5.2, which is a convex least-squares problem, but solving it directly is computationally expensive for large LLM weights. min ˆ W l ∥X l W l −X l (M l ⊙ ˆ W l )∥ 2 F (5.2) Consequently, many methods employ strategies to circumvent the expensive weight update step. For ex- ample,Wanda[136] avoids weight updates altogether, simply setting the selected weights to zero. However, other methods such asSparseGPT[36] andThanos[64] adopt a compromise, performing a more complex CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 52 update but only on a small subset of the weights. These heuristics trade off optimality for computational feasibility. 5.4 OPTIMA: Optimal Weight Updates via Quadratic Programming To overcome the challenges of weight update in LLM pruning, we propose OPTIMA, a novel approach that enables the efficient and optimal update ofallremaining weights once the pruning maskM l has been chosen. We achieve this by reformulating the least-squares problem as a set of independent Quadratic Programs (QPs) that can be solved in parallel on hardware accelerators like GPUs or TPUs using iterative methods. Specifically, we derive both a linearly constrained QP formulation and an equivalent unconstrained formula- tion. While the unconstrained form can be useful for optimizers restricted to such problems or in cases where it can be solved more efficiently, our implementation focuses on the constrained QP formulation, which is more amenable to GPU/TPU acceleration. 5.4.1 Reformulation as a Quadratic Program with Linear Constraints As discussed inSection 5.3, our goal is to minimize the problem defined inEquation 5.2. The Frobenius norm objective function inEquation 5.2is separable by the columns of the weight matrix. 2 We can therefore solve the optimization problem for each column independently. Letw j be thej-th column of the original weight matrixW l , and let ˆ w j be the corresponding column in the updated matrix ˆ W l . The mask for this column ism j . The optimization for this single column can be formulated as in Equation 5.3. min ˆ w j ∥X l w j −X l (m j ⊙ ˆ w j )∥ 2 2 (5.3) By defining the change in the weight column as∆w j = (m j ⊙ ˆ w j )−w j , the objective can then be rewritten in terms of this change as in Equation 5.4in standard quadratic form. min ∆w j ∥−X l ∆w j ∥ 2 2 =min ∆w j ∆w T j (X T l X l )∆w j (5.4) The constraints on∆w j in Equation 5.4are determined by the maskm j . LetS j be the set of indices where the mask is zero (i.e., weights to be pruned). For each indexi∈ S j , the corresponding entry in the updated weight vector,( ˆ w j ) i , must be zero. This imposes a linear constraint on the change vector, as shown in Equation 5.5. (m j ⊙ ˆ w j ) i = 0 =⇒(∆w j ) i =−(w j ) i ∀i∈S j (5.5) The entries of∆w j for the unpruned weights (wherem ij = 1) remain as free variables to be optimized. For each columnjof the weight matrix, we have a QP of the form represented inEquation 5.6, where H=X T l X l is the Hessian matrix, which is positive semi-definite and shared across all column-wise problems. The fact that the Hessian is shared among all columns, and only the constraints change, makes it very easy to parallelize on accelerators such as GPUs and TPUs. 2 Once the mask has been chosen, the weight reconstruction is separable for each column. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 53 minimize ∆w j ∆w T j H∆w j subject to(∆w j ) i =−(w j ) i ,∀i∈S j (5.6) 5.4.2 Reformulation as an Unconstrained Quadratic Program As an alternative to the constrained formulation inEquation 5.6, we can reformulate each column-wise prob- lem as an unconstrained quadratic program. This can be useful in settings where solvers are optimized for unconstrained problems or when eliminating constraints enables more efficient optimization. Although our implementation adopts the constrained approach for reasons discussed below, we include the unconstrained version for completeness. The key idea is to eliminate the equality constraints inEquation 5.5by substituting them directly into the objective. For a given columnj, defineI j as the set of indices where the mask is one (i.e., unpruned weights), and letS j denote the complement set (i.e., pruned weights, where the mask is zero). We reorder the entries of the change vector∆w j and the shared Hessian matrixH=X T l X l based on this partitioning, as shown in Equation 5.7. ∆w j = " ∆w I j ∆w S j # ,H= " H I j I j H I j S j H S j I j H S j S j # (5.7) As established in Equation 5.5, the entries of∆w j corresponding toS j are fixed:(∆w j ) i =−(w j ) i for all i∈S j . Substituting these fixed values into the quadratic objective yields the expanded form inEquation 5.8. ∆w T j H∆w j = ∆w T I j H I j I j ∆w I j + 2∆w T I j H I j S j ∆w S j + ∆w T S j H S j S j ∆w S j (5.8) Since∆w S j =−w S j , we substitute this to obtain the unconstrained objective in Equation 5.9. min ∆w I j ∆w T I j H I j I j ∆w I j −2∆w T I j H I j S j w S j +w T S j H S j S j w S j (5.9) The final term in Equation 5.9is constant with respect to the optimization variable∆w I j and can therefore be omitted. This results in the unconstrained quadratic program in Equation 5.10. minimize ∆w I j ∆w T I j Q j ∆w I j +c T j ∆w I j (5.10) where the problem-specific matrix and vector are defined as: Q j =H I j I j ,c j =−2H I j S j w S j (5.11) This formulation eliminates the need for explicit constraints, but introduces column-dependent variation in problem dimensions. Specifically, the size ofQ j andc j varies with the number of unpruned weights in each column. Consequently, the unconstrained QPs have heterogeneous shapes and objectives across columns, making them more difficult to batch and parallelize efficiently on accelerators like GPUs or TPUs. This motivates our choice to adopt the constrained formulation in Equation 5.6, where the problem structure is uniform and well-suited for high-throughput parallel execution. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 54 Algorithm 3 Layer-wise Pruning with Batched Column-wise Quadratic Programming Input:Pre-trained LLMM, calibration dataX, pruning masksM ask , QP solverS, batch sizeB. Output:Pruned and updated LLM ˆ M, updated masks ˆ M ask . 1for each layerLin the LLMMdo 2Initialize Hessian estimateH←0.▷Initialize covariance matrix 3for each calibration samplex∈Xdo 4y←L(x)▷Forward pass for one sequence 5H←H+y T y▷Accumulate covariance 6end for 7Store intermediate inputsX W |W∈Lfrom a forward pass ofL(X). 8for each weight matrixWin layerLdo 9Retrieve corresponding maskM∈M ask . 10Partition the columns ofWinto batches of sizeB. 11for each batch of columnsw j B j=1 in parallel do 12for each columnw j in the batchdo 13S j ←i|M j,i = 0▷Indices of pruned entries 14Define QP: min ∆w j ∆w T j H∆w j s.t.(∆w j ) i =−(w j ) i ,∀i∈S j (5.12) 15end for 16∆w j B j=1 ←S(H,w j B j=1 ,S j B j=1 ) 17Update weights:w j ←w j +∆w j ,∀j 18end for 19end for 20X←L(X)▷Update activations for next layer 21end for Return:Updated model ˆ M , updated masks ˆ M ask . 5.4.3 Solving the Quadratic Programs With the constrained QP formulation established, we now select a solver, whose efficiency is crucial for run- time and scalability on parallel hardware like GPUs and TPUs. Our QP, with its shared HessianHand simple bounds, suits specialized modern solvers. We adopt the state-of-the-art Restarted Accelerated Primal- Dual Hybrid Gradient (rAPDHG) algorithm [ 88], a first-order method effective here for three reasons: (1) its bottleneck—matrix-vector multiplications withHand its transpose—runs efficiently on GPUs/TPUs; (2) it achieves provably optimal linear convergence; and (3) a high-performance, open-source JAX-based imple- mentation is available in MPAX [ 87], designed for GPU/TPU execution. This enables parallel solving of thousands of column-wise QPs, leveraging the shared structure. 5.4.4 Efficient Implementation Naively implementing the optimization problem inEquation 5.6is computationally expensive and incurs sub- stantial memory overhead. These costs, however, can be greatly reduced through a series of optimization techniques. In the following, we describe the strategies we employ to solve the QPs efficiently on a single GPU, even for very large LLMs. Additionally, a detailed algorithm of our implementation is provided in Algorithm 3. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 55 Equality Constraints.Directly encoding the constraints fromEquation 5.5into the standard quadratic ob- jective leads to a prohibitively large matrix of equalities, even though these constraints merely fix individual variables to constant values. To avoid constructing such large matrices, we instead enforce the constraints by setting upper and lower bounds on the corresponding variables. In particular, fixing the bounds of (∆ w j ) i to −(w j ) i effectively locks the variable to the desired value, without incurring the overhead of explicit equality matrices. Batching QP Problems.In memory-limited scenarios, the optimization problems for all columns of the weight matrices may not fit on a single GPU. To address this, we employ a batching strategy that solves a subset of QP problems at a time. This approach reduces memory overhead while still leveraging the efficiency of solving multiple QPs in parallel. As a result, our method enables pruning of large LLMs even on a single GPU. Hessian calculation.For each layer, the Hessian matrix can be estimated as the covariance of the dense model’s outputs across multiple sequences. Suppose the output tensor isY∈R b×s×d , wherebis the number of sequences,sis the sequence length, anddis the output dimension of the layer. To compute the covariance directly, we would first reshapeYinto ˆ Y∈R bs×d , effectively stacking all tokens from all sequences into a single matrix, and then evaluate ˆ Y T ˆ Y. While this formulation is straightforward, it requires storing the fullYin accelerator memory, which becomes prohibitively expensive for largebands, often causing out-of-memory errors. To make the com- putation feasible, we observe that the covariance can be accumulated incrementally. Specifically,Ycan be decomposedintobsmallermatrices,y i ∈R s×d , eachcorrespondingtotheoutputofasinglesequence. Instead of materializing ˆ Y, we computey T i y i for each sequence separately and sum the results as inH≈ P b i=1 y T i y i . This decomposition yields the same result as computing ˆ Y T ˆ Ydirectly, but avoids the need to store the entire Yat once, making the approach scalable to very large LLMs. 5.5 Experiments Model, datasets, andevaluation.We evaluate OPTIMA on LLaMA 3.1, LLaMA 3.2 [31], Gemma 2 [140], and Gemma 3 [139] family of models. Model accuracy is assessed on a range of zero-shot downstream tasks, including MMLU [ 57], Piqa [13], Arc-Easy, Arc-Challenge [21], WinoGrande [126], and OpenBookQA [97], all of which are commonly used to evaluate LLM compression [100,136]. For zero-shot evaluations, we utilize the Language Model Evaluation Harness [41] framework. In line with prior work [136,36,100], we also report the perplexity of the models on a language modeling task on the WikiText2 [ 94] dataset. Baselines.WecompareOPTIMAagainststate-of-the-artone-shotpruningmethods, includingWanda[136], SparseGPT [36], Thanos [64], and ProxSparse [83] and show how OPTIMA can improve the performance of all these pruning methods across different models and datasets. The sensitivity of OPTIMA to the calibration datasetsizecanbefoundin SectionC.1. Intermsofmemoryreductionsandspeedup, ourmethodisguaranteed to achieve the same performance as other pruning methods such as Wanda and SparseGPT, since the sparsity pattern in these methods stays intact. Experiment Setup.Following previous work [ 36,136,100,64], we use 128 samples, each with 2048 tokens from the C4 dataset [121] for calibration. We set the relative and absolute tolerance of the rAPDHG QP solver CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 56 0.5 1.0 Wanda 0.9 1.0 SparseGPT 12345678910111213141516 0.9 1.0 Thanos Layer Number Error Ratio (Final Error / Initial Error) Layer Type Q K V O Up Gate Down Figure 5.2: Relative error reduction on OPTIMA in comparison to Wanda, SparseGPT, and Thanos for LLaMA-3.2 1B. Table 5.1: Key hyperparameters used in OPTIMA. HyperparameterValue Calibration Samples128 Tokens per Sample2048 Dataset for CalibrationC4 Relative Tolerance (rAPDHG)0.01 Absolute Tolerance (rAPDHG)0.01 Maximum Iterations (rAPDHG)100,000 ADAM Learning Rate10 −2 ,10 −3 ,10 −4 ,10 −5 ADAM Weight Decay0 in MPAX to 0.01 and the maximum number of iterations to 100,000. If the optimizer does not converge within this budget for most problems, or the final error of a layer is larger than the initial error, OPTIMA skips updating that layer.Table 5.1summarizes the key hyperparameters. For all baselines, we either use their publicly available checkpoints or reproduce results with default hyperparameters. ModelQuality.We evaluate the accuracy of OPTIMA and other state-of-the-art pruning methods across 2:4 and unstructured sparsity benchmarks. Wanda is a mask selection algorithm that does not provide any weight update mechanism for the weights. SparseGPT and Thanos, on the other hand, update the weight values in addition to searching for the best mask. We couple OPTIMA weight update with the masks generated using each of these methods and compare the resulting performance of the models. Table 5.2summarizes the performance metrics for Wanda, SparseGPT, and Thanos with and without the OPTIMA update mechanism for 50% unstructured sparsity. It can be seen that models pruned with OPTIMA weight update scheme consistently outperform the methods using weight update methods, providing up to CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 57 1.80% average accuracy improvement across six downstream tasks (Gemma-3-1B). Table 5.3presents the results of pruning transformer models using 2:4 semi-structured sparsity. In these experiments, we applied pruning exclusively to the weight matrices in the multilayer perceptron (MLP) com- ponents, leaving the self-attention layers dense. This approach yielded sparse models with an overall sparsity of 38% to 41%. We adopted this selective pruning strategy to maintain model accuracy above a practical threshold, as 2:4 sparsity significantly impacts performance, potentially rendering fully sparse models ineffec- tive. Our results demonstrate that our proposed OPTIMA update mechanism consistently outperforms other methods under 2:4 sparsity, achieving superior accuracy. Higher Sparsity Ratios.To assess the robustness of OPTIMA at more aggressive compression levels, we extend our evaluation to 60% unstructured sparsity.Table 5.4presents the perplexity and zero-shot accuracy metrics across the same models and tasks. OPTIMA continues to deliver consistent improvements over the baseline pruning methods, with average accuracy gains of up to 2.53% across the downstream tasks (LLaMA- 3.2-1B). These enhancements are particularly notable at higher sparsity ratios, where pruning a larger portion of weights introduces greater reconstruction error. By optimally readjusting the remaining weights through our QP formulation, OPTIMA effectively mitigates this error, leading to lower perplexity and higher downstream performance compared to Wanda, SparseGPT, or Thanos individually. For example, on LLaMA-3.2-3B, OP- TIMA increases Wanda’s average accuracy from 38.77% to 42.74%, highlighting its ability to preserve model utility under extreme sparsity conditions. Extended Evaluation on the Qwen-2.5 Model Family.To further validate the robustness and generaliz- ability of OPTIMA, we conduct additional experiments on the Qwen-2.5 family of models, with sizes ranging from 0.5B to 14B parameters. These models were not included in the preceding analysis, and this evaluation serves to confirm that OPTIMA’s benefits apply across different model architectures. We evaluate performance across three distinct settings, mirroring the main experiments: 50% unstructured sparsity ( Table 5.5), 60% unstructured sparsity (Table 5.6), and 2:4 semi-structured sparsity (Table 5.7). Unstructured Sparsity (50% and 60%).At 50% unstructured sparsity (Table 5.5), OPTIMA consistently improves zero-shot performance across all Qwen-2.5 model sizes and for all mask selection methods (Wanda, SparseGPT, and Thanos). For example, on the Qwen-2.5 3B model, OPTIMA boosts the average accuracy of Wanda from 54.02% to 55.33% and SparseGPT from 54.70% to 55.69%. These gains demonstrate that our OPTIMA reconstruction successfully recovers accuracy lost during the pruning step. The advantages of OPTIMA are even more pronounced at the more aggressive 60% sparsity ratio, as shown in Table 5.6. At this level, pruning introduces a more significant reconstruction error, providing a greater opportunity for OPTIMA to recover performance. This is especially clear on the Qwen-2.5 3B model, where OPTIMA improves Wanda’s average accuracy from 43.67% to 47.86% (a 4.19% absolute gain) and Thanos’s from 48.45% to 49.98% (a 1.53% gain). Semi-Structured Sparsity (2:4).In the 2:4 semi-structured sparsity setting (Table 5.7), where pruning is applied only to the MLP layers, OPTIMA provides clear improvements for most models, particularly in the 1.5B and 3B range. For instance, it improves the average accuracy of the 3B model pruned with Wanda from 49.48% to 50.63% and the 1.5B model from 46.01% to 47.26%. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 58 On the larger 7B and 14B models, the results are more varied, with performance differing based on the underlyingmaskselector. ThissuggestsacomplexinteractionbetweenmaskselectionheuristicsandOPTIMA reconstruction for structured sparsity at this scale, which could be a valuable avenue for future investigation. Overall, these experiments on the Qwen-2.5 family reinforce the findings from the preceding sections. They confirm that OPTIMA is a broadly applicable and effective method for enhancing model accuracy post- pruning, delivering its most significant and consistent gains in high-sparsity unstructured regimes. Comparison with Alternative Optimizers.While our constrained QP solver leverages theoretical guaran- tees for convergence and optimality, we also compare it against ADAM [70], a popular first-order optimizer without such assurances for quadratic problems. We reformulate the weight update as a mean squared error (MSE) minimization problem and use ADAM for solving it. Optimizers such as ADAM do not guarantee convergence, and are sensitive to their hyperparameters. For each layer, we do an exhaustive search with 4 different learning rates ranging from10 −2 to10 −5 , each with a linear learning rate scheduler and choose the best configuration for final weight update. Table 5.8illustrates this on Gemma 3 1B and OPT 125M [165] under 50% unstructured sparsity. We show two examples inTable 5.8, showing that ADAM results in suboptimal solutions. To further test the limitations of optimizers without convergence guarantees, we test ADAM on OPT-125M, and observe that it leads to divergence of the model. On Gemma 3 1B, ADAM yields competitive results in some cases (e.g., slightly lower perplexity for SparseGPT+ADAM at 27.12 versus OPTIMA’s 27.35), but OPTIMA achieves higher overall accuracy (e.g., 44.01% for Wanda+OPTIMA versus 43.72% for Wanda+ADAM). However, on smaller models like OPT 125M, ADAM exhibits instability, leading to divergence and dramatically higher perplexity (e.g., 205.82 for Wanda+ADAM versus 35.44 for Wanda+OPTIMA). This underscores the risks of using non-specialized optimizers for our column-wise QPs, where suboptimal or unstable solutions can degrade model quality. OPTIMA’s use of provably convergent methods like rAPDHG ensures reliable and superior weight updates, making it a more robust choice for post-training pruning. Layer-wise Error Improvement.To provide a deeper insight into how OPTIMA improves the accuracy of the models, we compare the layer-wise error of different layers in LLaMA-3.2 1B during pruning with and without OPTIMA.Figure 5.2shows the relative output error improvement of all the pruned layers in the model, defined as MSE(Y OPTIMA ,Y dense ) MSE(Y other ,Y dense ) , whereM SEdenotes the mean squared error across the calibration dataset. Figure 5.2shows that OPTIMA consistently improves the layer-wise error of other methods, resulting in superior accuracy on the downstream tasks. Pruning Time Analysis.To evaluate the computational efficiency of OPTIMA, we measured the time re- quired to prune various language models. The pruning process was conducted on a single NVIDIA H100 GPU with 80GB of memory. Our measurements show that pruning times vary with model size: smaller models like LLaMA 3.2 1B and Gemma 3 1B each required approximately2.5h, Gemma 2 2B took5.5h, LLaMA 3.2 3B needed7.0h, and the larger LLaMA 3.1 8B model required up to40.0h. The results indicate that pruning time scales with model size, reflecting the computational complexity of OPTIMA’s pruning algorithm, which adapts to the architectural differences across models. The consistency in pruning times for models of similar size (e.g., LLaMA 3.2 1B and Gemma 3 1B) highlights the robustness of OPTIMA in handling diverse model architectures efficiently. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 59 5.6 Conclusion and Limitations Inthischapter, weexploredthelimitsofthesparsitypillarunderastrictresource-constrainedregime. Recog- nizing that end-to-end training is not always feasible, we developed OPTIMA to determine the mathematical upper bound of reconstruction accuracy possible using only a small calibration dataset. By reformulating post-training weight reconstruction as globally optimal, column-wise Quadratic Programs (QPs) and lever- aging the shared-Hessian structure, we achieved massive parallelism on standard GPUs without the need for backpropagation. Our results demonstrate that this principled approach pays significant dividends. OPTIMA functions as a drop-in weight-update step for common mask selectors, improving zero-shot accuracy across various LLM families by up to 3.97 percentage points without any fine-tuning. Crucially, these gains persist even at high sparsity levels (≥60%), proving that it is possible to create a highly sparse model that retains the fidelity of the original dense network. However, while OPTIMA minimizes the reconstruction error for any given mask, it remains bound by two fundamental limitations: 1.Layer-wise Pruning Sub-optimality:It must operate within the constraints of a fixed, pre-determined sparsity pattern (like 2:4), which may not align with the model’s true information distribution. 2.The Accuracy Gap:Even with optimal reconstruction, a gap often remains between the sparse model and the dense baseline, suggesting that sparsity alone—without the aid of other Trinity pillars—has reached its ceiling in the zero-shot regime. To address the first limitation, the next chapter introducesPATCH. We transition from the ”no-training” regime of OPTIMA to a ”fine-tuning” regime, where we utilize a training budget to learn a flexible, hybrid sparsitystructure that dynamically preserves density where it matters most. To address the second limitation (the accuracy gap), we will revisit OPTIMA’s findings in Chapter 7, where we demonstrate that re-introducing theLow-Rankpillar can bridge the remaining distance to dense performance. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 60 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage LLaMA 3.1 8B Dense–5.84 63.57 80.09 81.44 51.37 73.48 33.4063.89 Wanda –9.64 47.79 75.68 72.56 40.70 70.09 27.4055.70 WandaOPTIMA9.3748.8576.7173.8242.3270.3228.2056.70 SparseGPT SparseGPT 9.3051.3276.19 73.02 41.2770.88 29.4057.01 SparseGPTOPTIMA9.3349.3176.6174.2842.8370.8828.2057.02 Thanos Thanos9.2750.36 77.04 74.92 42.58 70.96 30.0057.64 ThanosOPTIMA9.3550.1776.5074.1641.8970.2428.4056.89 LLaMA 3.2 1B Dense–9.75 36.92 74.27 65.53 31.31 60.30 26.20 49.09 Wanda –23.51 26.35 65.18 52.10 23.81 54.62 18.00 40.01 WandaOPTIMA18.8427.6967.0852.6124.7455.6420.2041.33 SparseGPT SparseGPT 18.84 25.71 67.85 54.2926.54 57.7022.0042.35 SparseGPTOPTIMA18.0926.9568.0154.5925.8556.9124.0042.72 Thanos Thanos19.70 25.37 67.63 52.9927.1354.3822.2041.62 ThanosOPTIMA18.7725.9968.2353.4926.4555.8821.6041.94 LLaMA 3.2 3B Dense–7.81 54.13 76.55 74.28 42.75 69.38 30.60 57.95 Wanda –12.92 40.79 72.03 65.45 32.34 63.69 25.40 49.95 WandaOPTIMA12.2443.1172.4766.5033.5366.3826.2051.37 SparseGPT SparseGPT 12.32 37.9673.4565.19 33.02 66.38 25.2050.20 SparseGPTOPTIMA12.4340.5473.4566.3735.0766.6926.2051.39 Thanos Thanos12.26 40.11 72.80 64.77 32.8567.7226.6050.81 ThanosOPTIMA12.4041.5173.2365.0734.3967.2527.0051.41 Gemma 3 1B Dense–14.17 24.95 74.81 71.93 35.41 58.72 28.8049.10 Wanda –32.96 22.97 67.19 61.03 26.37 55.72 20.0042.21 WandaOPTIMA28.9023.9669.4862.8428.5856.8322.4044.01 SparseGPT SparseGPT 28.34 24.85 68.8860.9426.62 55.49 21.40 43.03 SparseGPTOPTIMA27.3525.7369.7560.9027.8256.3522.0043.76 Thanos Thanos28.65 23.0969.7562.1627.99 56.51 23.8043.88 ThanosOPTIMA28.1424.7069.6463.4327.3955.9623.2044.05 Gemma 2 2B Dense–68.69 49.33 78.24 80.22 46.93 68.82 31.40 59.16 Wanda –327.45 34.1774.1669.7834.30 62.83 26.4050.27 WandaOPTIMA215.6334.8673.9971.3832.5961.9625.8050.10 SparseGPT SparseGPT 234.68 35.59 73.61 69.99 34.2265.82 28.2051.24 SparseGPTOPTIMA241.0937.5973.8370.6235.0764.7227.8051.60 Thanos Thanos276.97 30.62 73.18 67.72 33.62 63.2226.80 49.19 ThanosOPTIMA250.1532.7273.7268.8134.1363.8526.4049.94 Table 5.2: Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 50% unstructured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 61 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage LLaMA 3.1 8B Dense–5.84 63.57 80.09 81.44 51.37 73.48 33.4063.89 Wanda –13.54 43.42 73.18 69.23 35.32 67.3225.8052.38 WandaOPTIMA12.5845.4573.3969.5736.1868.9025.2053.12 SparseGPT SparseGPT 12.37 45.6273.8369.15 35.84 69.22 25.6053.21 SparseGPTOPTIMA12.5446.0473.7269.9536.7769.6127.0053.85 Thanos Thanos12.66 44.39 73.94 69.57 36.1868.9025.2053.03 ThanosOPTIMA12.8044.4174.0569.9536.4368.5925.6053.17 LLaMA 3.2 1B Dense–9.75 36.92 74.27 65.53 31.31 60.30 26.20 49.09 Wanda –30.43 23.32 63.55 47.5623.6355.25 15.0038.05 WandaOPTIMA48.2324.8066.1058.0423.5555.2519.8041.26 SparseGPT SparseGPT 21.98 23.05 65.45 52.15 25.1757.6217.6040.17 SparseGPTOPTIMA21.4023.4065.7252.7825.5157.0618.6040.51 Thanos Thanos22.8024.09 65.6751.6825.0052.9617.6039.50 ThanosOPTIMA22.2623.4165.3452.2223.7255.9616.8039.58 ProxSparse –41.9523.6461.21 42.38 22.53 53.67 16.0036.57 ProxSparseOPTIMA28.5323.0763.3847.9022.5354.7816.4038.01 LLaMA 3.2 3B Dense–7.81 54.13 76.55 74.28 42.75 69.38 30.6057.95 Wanda –18.51 34.30 70.73 60.69 30.72 61.1724.8047.07 WandaOPTIMA16.6437.1570.7861.9531.1462.5124.6048.02 SparseGPT SparseGPT 16.19 36.13 70.29 63.01 30.4664.7225.0048.27 SparseGPTOPTIMA16.3638.0370.8463.1732.1763.6925.6048.92 Thanos Thanos16.24 35.55 70.35 61.28 29.7863.3024.2047.41 ThanosOPTIMA16.4935.7270.6262.0430.9763.2225.6048.03 ProxSparse –19.50 24.66 68.12 56.31 27.82 58.56 20.0042.58 ProxSparseOPTIMA18.2831.7669.5360.2728.8460.3020.6045.22 Gemma 3 1B Dense–14.17 24.95 74.81 71.93 35.41 58.72 28.80 49.10 Wanda –60.7423.74 65.51 56.7822.35 52.7219.8040.15 WandaOPTIMA23.2523.2563.3851.1424.0654.3018.2039.06 SparseGPT SparseGPT 44.87 24.8366.7657.70 23.2955.9619.4041.32 SparseGPTOPTIMA42.6625.1166.2758.9623.8955.8020.6041.77 Thanos Thanos48.50 25.23 65.8959.3023.12 53.5920.8041.32 ThanosOPTIMA44.9125.8366.0058.6323.2954.7020.0041.41 ProxSparse –41.02 23.0166.00 54.3422.4455.88 20.2040.31 ProxSparseOPTIMA52.9924.1364.7453.7022.6152.2517.0039.07 Gemma 2 2B Dense–68.69 49.33 78.24 80.22 46.93 68.82 31.40 59.16 Wanda –421.0134.34 71.33 68.10 30.97 61.4026.4048.76 WandaOPTIMA229.6934.4471.8768.9033.8762.2725.0049.39 SparseGPT SparseGPT251.71 32.8471.7668.73 32.4261.88 23.40 48.51 SparseGPTOPTIMA227.9932.7771.7667.4732.1763.3824.4048.66 Thanos Thanos256.5831.02 70.7367.7232.0862.5124.8048.14 ThanosOPTIMA239.2032.5871.1667.4732.2560.8525.2048.25 ProxSparse –176.03 37.1971.9867.5534.4761.4825.0049.61 ProxSparseOPTIMA254.0338.2771.2768.6033.5361.8824.6049.69 Table 5.3: Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 2:4 sparsity. In this experiment, only the layers in the MLP part of the transformer are pruned, and the self-attention layers are dense, resulting in an end-to-end sparsity ratio of 38% to 41%. OPTIMA consistently improves the accuracy of the models across different tasks. Please note that ProxSparse pruning is limited to 2:4 sparsity, and hence our unstructured sparsity experiments do not include it. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 62 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage LLaMA 3.1 8B Dense–5.84 63.57 80.09 81.44 51.37 73.48 33.4063.89 Wanda –21.65 31.98 69.53 61.11 27.30 61.09 21.4045.40 WandaOPTIMA17.5633.9671.6063.7629.3566.0622.6047.89 SparseGPT SparseGPT 15.4435.3271.55 62.88 31.6668.1924.2048.96 SparseGPTOPTIMA15.6432.4471.8763.9733.1167.5624.6048.93 Thanos Thanos15.9135.22 72.09 65.28 33.1967.4023.4049.43 ThanosOPTIMA16.0934.4872.0364.6933.0268.5122.8049.25 LLaMA 3.2 1B Dense–9.75 36.92 74.27 65.53 31.31 60.30 26.20 49.09 Wanda –71.53 22.95 59.68 39.48 18.77 50.43 12.20 33.92 WandaOPTIMA41.5023.5262.6244.5320.6552.5714.8036.45 SparseGPT SparseGPT 48.0023.0262.08 43.4821.7652.09 17.4036.64 SparseGPTOPTIMA38.0522.9563.3843.5220.4853.2819.6037.20 Thanos Thanos46.7823.2562.57 44.49 21.59 53.20 16.6036.95 ThanosOPTIMA40.5423.0262.9544.5321.6753.9117.4037.25 LLaMA 3.2 3B Dense–7.81 54.13 76.55 74.28 42.75 69.38 30.60 57.95 Wanda –31.13 25.53 65.23 47.90 22.70 55.25 16.00 38.77 WandaOPTIMA23.5631.2067.4153.9624.5759.5119.8042.74 SparseGPT SparseGPT 22.0031.27 69.3753.6626.0261.3321.0043.78 SparseGPTOPTIMA22.6729.5868.7754.8024.7462.3520.6043.47 Thanos Thanos22.48 29.23 67.63 55.0126.0257.85 19.2042.49 ThanosOPTIMA22.2831.4367.9055.2624.9159.6720.6043.30 Gemma 3 1B Dense–14.17 24.95 74.81 71.93 35.41 58.72 28.8049.10 Wanda –90.48 23.04 62.19 49.75 18.60 50.99 15.2036.63 WandaOPTIMA64.7923.3464.0952.8620.4851.9316.4038.18 SparseGPT SparseGPT 60.9124.5865.34 51.98 21.93 51.14 16.60 38.60 SparseGPTOPTIMA56.2723.7266.2152.4422.5352.9617.6039.24 Thanos Thanos62.2224.6264.53 52.86 20.65 52.17 18.8038.94 ThanosOPTIMA56.7824.4464.8555.1822.0154.8519.8040.19 Gemma 2 2B Dense–68.69 49.33 78.24 80.22 46.93 68.82 31.40 59.16 Wanda –757.47 23.36 65.78 56.10 21.59 52.64 19.8039.88 WandaOPTIMA435.1024.3766.5958.5021.9357.3820.0041.46 SparseGPT SparseGPT 488.25 24.49 68.50 57.45 25.0058.96 25.0043.23 SparseGPTOPTIMA451.4625.8968.8858.5026.2858.0124.2043.63 Thanos Thanos523.6123.69 68.23 58.12 23.8958.3321.20 42.24 ThanosOPTIMA497.7523.1267.7457.0723.3859.2720.6041.86 Table 5.4: Model perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 60% unstructured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 63 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage Qwen 2.5 0.5B Dense–13.08 47.36 69.97 64.18 29.18 55.80 24.4048.48 Wanda –24.0030.5264.09 57.41 24.06 54.38 19.8041.71 WandaOPTIMA22.7026.1464.5857.7925.2656.0422.0041.97 SparseGPT SparseGPT 20.3329.3864.74 56.52 24.1556.20 20.6041.93 SparseGPTOPTIMA19.5427.6865.1356.9924.6655.3320.6041.73 Thanos Thanos20.85 28.9465.4055.9324.40 56.3521.6042.10 ThanosOPTIMA20.4130.0064.6956.1024.4055.4122.2042.13 Qwen 2.5 1.5B Dense–9.28 59.70 75.73 75.34 40.96 63.14 32.2057.84 Wanda –14.45 44.76 71.2266.6231.74 59.9124.8049.84 WandaOPTIMA12.8545.6172.3666.6232.3461.8024.6050.55 SparseGPT SparseGPT 13.09 46.80 71.6566.75 33.62 62.2725.6051.12 SparseGPTOPTIMA12.7646.9671.8265.4533.0261.8026.2050.87 Thanos Thanos13.1748.4071.76 66.8433.70 62.83 27.2051.79 ThanosOPTIMA12.8948.2172.0367.2633.5362.0426.2051.55 Qwen 2.5 3B Dense–8.03 65.00 78.35 77.31 44.88 68.43 29.2060.53 Wanda –11.39 49.09 73.23 71.4638.4865.43 26.4054.02 WandaOPTIMA10.5952.0074.3772.1838.0566.7728.6055.33 SparseGPT SparseGPT 10.74 52.49 74.6571.3436.86 64.64 28.20 54.70 SparseGPTOPTIMA10.5753.9275.3570.8338.3166.1429.6055.69 Thanos Thanos10.6452.61 75.52 70.5436.69 66.6128.4055.06 ThanosOPTIMA10.5252.1175.4670.1237.2966.6928.2054.98 Qwen 2.5 7B Dense–6.85 71.76 78.73 80.51 48.38 72.61 33.4064.23 Wanda –8.62 65.89 77.31 75.08 40.53 70.1730.8059.96 WandaOPTIMA8.3366.1777.6976.4342.6671.2730.6060.80 SparseGPT SparseGPT 8.4266.09 78.0775.34 42.75 71.11 31.0060.73 SparseGPTOPTIMA8.3665.7877.6475.6342.9271.5131.6060.85 Thanos Thanos8.49 66.2177.8674.71 42.32 70.17 30.4060.28 ThanosOPTIMA8.4666.2377.5876.2244.4571.1931.2061.15 Qwen 2.5 14B Dense–5.30 77.62 81.28 82.24 55.80 75.14 34.4067.75 Wanda –7.3069.8479.16 81.02 51.28 73.7234.6064.94 WandaOPTIMA7.1869.2979.4381.1952.3073.8033.8064.97 SparseGPT SparseGPT 7.2469.83 79.6080.98 51.02 72.93 32.80 64.53 SparseGPTOPTIMA7.1469.7179.5481.1951.7973.8033.6064.94 Thanos Thanos7.2570.57 79.8780.18 49.15 73.09 32.2064.17 ThanosOPTIMA7.1970.1679.6081.5751.3773.4833.0064.86 Table 5.5: Qwen-2.5 family perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 50% unstructured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 64 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage Qwen 2.5 0.5B Dense–13.08 47.36 69.97 64.18 29.18 55.80 24.4048.48 Wanda –83.42 23.02 59.96 43.81 18.09 50.28 12.8034.66 WandaOPTIMA51.9723.1660.7246.2520.1451.7816.4036.41 SparseGPT SparseGPT 40.56 22.90 61.59 48.40 21.25 52.80 16.8037.29 SparseGPTOPTIMA36.7723.0662.1348.7421.3353.9917.4037.77 Thanos Thanos44.2923.78 62.02 48.6521.33 52.25 17.8037.64 ThanosOPTIMA41.9223.5961.8646.8022.3553.7519.6037.99 Qwen 2.5 1.5B Dense–9.28 59.70 75.73 75.34 40.96 63.14 32.2057.84 Wanda –58.38 27.25 65.18 54.50 24.74 53.04 17.2040.32 WandaOPTIMA23.8130.9966.8756.4424.9156.8318.4042.41 SparseGPT SparseGPT 21.9233.5667.3658.08 27.4757.14 21.6044.20 SparseGPTOPTIMA19.3531.4467.7956.2727.2259.2722.4044.07 Thanos Thanos27.07 33.6667.6357.4927.7356.67 20.6043.96 ThanosOPTIMA23.6435.9767.1457.6626.1958.0920.8044.31 Qwen 2.5 3B Dense–8.03 65.00 78.35 77.31 44.88 68.43 29.2060.53 Wanda –22.06 28.07 67.14 60.86 27.39 58.17 20.4043.67 WandaOPTIMA15.6737.2270.2463.5530.8961.6423.6047.86 SparseGPT SparseGPT 14.8243.1671.60 64.35 32.59 63.30 23.20 49.70 SparseGPTOPTIMA14.5040.2572.2064.9033.6263.6924.0049.78 Thanos Thanos14.90 40.7671.3863.26 30.63 61.25 23.4048.45 ThanosOPTIMA14.4242.5871.2764.7332.6863.6125.0049.98 Qwen 2.5 7B Dense–6.85 71.76 78.73 80.51 48.38 72.61 33.4064.23 Wanda –14.09 54.58 72.03 71.68 37.03 66.46 25.4054.53 WandaOPTIMA11.1555.4973.9973.8637.8867.9626.2055.90 SparseGPT SparseGPT 10.8656.6374.92 73.36 40.6167.2525.8056.43 SparseGPTOPTIMA10.5355.5575.4673.7840.7066.9326.6056.50 Thanos Thanos11.0759.5474.7073.44 40.4469.22 26.4057.29 ThanosOPTIMA10.7458.9075.3572.6940.1069.4627.0057.25 Qwen 2.5 14B Dense–5.30 77.62 81.28 82.24 55.80 75.14 34.4067.75 Wanda –11.16 61.38 75.41 74.1242.1571.51 29.2058.96 WandaOPTIMA9.6961.7475.5775.3441.9873.0929.4059.52 SparseGPT SparseGPT 9.2262.8376.66 76.18 44.4572.14 29.60 60.31 SparseGPTOPTIMA8.9762.2276.9376.4744.5471.6729.0060.14 Thanos Thanos9.1463.03 77.2076.05 43.77 71.98 29.8060.31 ThanosOPTIMA8.9960.3076.3976.3043.7772.1430.6059.92 Table 5.6: Qwen-2.5 perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 60% unstruc- tured sparsity. OPTIMA consistently improves the accuracy of the models across different tasks. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 65 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage Qwen 2.5 0.5B Dense–13.08 47.36 69.97 64.18 29.18 55.80 24.4048.48 Wanda –41.3027.8861.75 48.9923.3852.72 14.0038.12 WandaOPTIMA27.6125.5763.7651.1822.1853.2815.8038.63 SparseGPT SparseGPT 27.1524.8362.79 49.41 22.35 52.3317.2038.15 SparseGPTOPTIMA25.7723.3862.9551.3022.9554.7017.0038.71 Thanos Thanos27.5824.3162.68 49.92 21.42 51.7816.6037.78 ThanosOPTIMA26.2623.6062.7951.2222.1854.3016.6038.45 Qwen 2.5 1.5B Dense–9.28 59.70 75.73 75.34 40.96 63.14 32.2057.84 Wanda –21.92 39.95 67.25 61.11 28.84 58.09 20.8046.01 WandaOPTIMA17.1439.9669.2662.3330.2958.7223.0047.26 SparseGPT SparseGPT 17.2441.0569.6463.09 30.63 61.17 23.0048.10 SparseGPTOPTIMA16.5235.6469.9162.7529.3560.9323.0046.93 Thanos Thanos17.5642.8368.55 61.53 29.1058.9621.4047.06 ThanosOPTIMA16.8340.3169.7562.6730.3858.5625.0047.78 Qwen 2.5 3B Dense–8.03 65.00 78.35 77.31 44.88 68.43 29.20 60.53 Wanda –17.1446.6870.08 64.7731.6661.48 22.2049.48 WandaOPTIMA14.0846.5571.4464.9031.3164.1725.4050.63 SparseGPT SparseGPT 14.0643.79 72.0366.16 31.31 64.96 25.2050.58 SparseGPTOPTIMA13.5743.3671.4466.6231.9165.2727.0050.93 Thanos Thanos14.35 41.4171.4961.70 29.27 64.01 25.20 48.85 ThanosOPTIMA13.7543.5670.7364.0630.9764.0925.4049.80 Qwen 2.5 7B Dense–6.85 71.76 78.73 80.51 48.38 72.61 33.4064.23 Wanda –11.47 61.03 74.48 75.34 42.15 68.82 27.80 58.27 WandaOPTIMA11.8053.9074.1071.9738.0567.9626.4055.40 SparseGPT SparseGPT10.21 60.30 75.57 75.59 41.38 71.51 28.20 58.76 SparseGPTOPTIMA10.9253.9074.3272.4337.4669.6127.8055.92 Thanos Thanos10.45 60.12 74.54 75.08 41.13 69.93 28.8058.27 ThanosOPTIMA11.1354.9073.9470.8335.1569.0626.0054.98 Qwen 2.5 14B Dense–5.30 77.62 81.28 82.24 55.80 75.14 34.4067.75 Wanda –9.70 65.82 76.99 76.89 45.39 73.56 31.8061.74 WandaOPTIMA8.9067.0377.3177.8246.7674.2732.6062.63 SparseGPT SparseGPT 9.0267.4577.58 77.6144.6273.8832.6062.29 SparseGPTOPTIMA8.8267.3377.5877.8244.2873.8832.0062.15 Thanos Thanos9.0666.3177.6477.90 46.2572.77 31.2062.01 ThanosOPTIMA8.9266.0777.9777.4445.6572.9331.8061.98 Table 5.7: Qwen-2.5 perplexity on WikiText2 and accuracy on zero-shot downstream tasks for 2:4 sparsity. In this experiment, only the layers in the MLP part of the transformer are pruned, and the self-attention layers are dense, resulting in an end-to-end sparsity ratio of 38% to 41%. OPTIMA consistently improves the accuracy of the models across different tasks. Please note that ProxSparse pruning is limited to 2:4 sparsity, and hence our unstructured sparsity experiments do not include it. CHAPTER 5. OPTIMA: OPTIMAL ONE-SHOT PRUNING FOR LLMS VIA QUADRATIC PROGRAMMING RECONSTRUCTION 66 Model Mask Selection Weight Update Perplexity Metrics (%) MMLUPIQAArc-EArc-CWinoOpenQAAverage Gemma 3 1B Dense–14.17 24.95 74.81 71.93 35.41 58.72 28.80 49.10 Wanda –32.96 22.97 67.19 61.03 26.37 55.72 20.0042.21 Wanda ADAM29.25 23.16 69.04 62.71 27.73 57.46 22.2043.72 WandaOPTIMA28.9023.9669.4862.8428.5856.8322.4044.01 SparseGPT SparseGPT 28.34 24.85 68.88 60.94 26.62 55.49 21.4043.03 SparseGPT ADAM27.12 24.74 69.53 61.36 27.05 54.78 22.2043.28 SparseGPTOPTIMA27.3525.7369.7560.9027.8256.3522.0043.76 OPT 125M Dense–27.67 22.85 62.84 43.56 19.45 49.88 16.4035.83 Wanda –39.50 22.92 61.15 39.94 19.88 52.17 14.0035.01 Wanda ADAM205.82 25.63 57.02 34.13 17.66 50.51 13.0032.99 WandaOPTIMA35.4423.0261.6642.9319.1150.1214.6035.24 SparseGPT SparseGPT 36.88 23.00 61.97 40.99 19.71 53.59 14.6035.64 SparseGPT ADAM224.34 23.15 56.75 35.65 17.49 47.36 12.2032.10 SparseGPTOPTIMA35.6123.8562.3742.2819.9752.2515.4036.02 Table 5.8: Comparison of OPTIMA with other optimizers without convergence guarantees (ADAM). ADAM can lead to suboptimal solutions (Gemma 3 1B) or divergence of the model (OPT 125M). Chapter 6 PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs 6.1 Statement of Contributions The content of this chapter is derived from the paper “PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs” [ 59]. This research was a collaborative effort with Younes Hourri, and we are co-first authors with equal contribution. The specific breakdown of contributions is as follows: •Mohammad Mozaffari:Conceived the original concept of learnable hybrid sparsity and implemented the initial version of the codebase. Designed and executed the experiments regarding quantization, Low-Rank Adaptation (LoRA), and the fine-tuning (FT) comparisons. •Younes Hourri:Formulated the mathematical framework for the tile-level probability distributions (specif- icallyEquation 6.3andEquation 6.4), extended the codebase with additional capabilities, and conducted the primary pruning experiments. He was also responsible for the hardware acceleration implementation and throughput analysis. •Joint Contributions:Both authors collaborated closely on the writing and revision of the manuscript. Maryam Mehri Dehnavi held a supervisory position in this work. 6.2 Introduction In the previous chapter, we established OPTIMA to maximize the accuracy of sparse models under a strict ”no-training” constraint. We demonstrated that for a fixed mask structure, one can mathematically solve for the optimal weights. However, OPTIMA, and indeed any layer-wise pruning and weight reconstruction method, eventually hits an accuracy ceiling imposed by the rigidity of the mask itself. To break this ceiling and fully refine the sparsity pillar, we must relax the resource constraints. In this chapter, we transition from the ”zero- training” regime of OPTIMA to alearnable regime, utilizing a training budget to move beyond optimizing values within a fixed pattern and address the pattern itself. We need a method that bridges the gap between the high accuracy of flexible unstructured pruning and the hardware acceleration of rigid structured patterns. Currently, sparsity techniques operate at two extremes, neither of which is sufficient for the Compression 67 CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS68 Trinity. Unstructured sparsity allows non-zero elements to appear anywhere, theoretically matching dense model accuracy due to its flexibility in allocation [136,36,1]. However, its irregular memory access patterns hinderaccelerationonmodernhardwarelikeGPUs, preventingpracticalspeedups[154,32]. Conversely, semi- structuredsparsityofferspracticalaccelerationbutimposesstrictlayoutexpectations. Specifically, wefocuson the 2:4 sparse pattern supported by NVIDIA Ampere and Hopper architectures, as detailed inChapter 2. This format requires every block of four contiguous weights to contain at least two zeros to utilize sparse Tensor Cores. This ”one-size-fits-all” approach enforces a uniform 50% sparsity ratio across all layers, failing to account for the varying sensitivity of different network components. This often leads to significant accuracy loss when models are pruned using one-shot methods [136,36,64,83] or end-to-end learned masks [33]. Recentstudiesconfirmthat sparsityshould be allocatednon-uniformlyforoptimal performance[158,149,76], yet standard 2:4 sparsity locks the model into a fixed allocation. To bridgethis gapand create a trulyflexiblesparsity pillar, weproposePruning with a LearnableTile-level Configuration forHybrid Sparsity (PATCH). Instead of forcing the entire model to adhere to a rigid sparse structure, PATCH introduces a hybrid mask that partitions each weight matrix into hardware-friendly tiles. Through a learnable masking process—enabled by our relaxed training budget—PATCH designates each tile as either dense (0% sparsity) or 2:4 sparse (50% sparsity). This adaptive approach allows the matrix to realize an effective global sparsity ratio anywhere between 0% and 50%, dynamically balancing accuracy in critical regions with hardware-friendly sparsity elsewhere. While the PATCH methodology is generalizable to any block-based sparsity pattern (such as 4:8 or custom N:M ratios), we focus our evaluation on 2:4 sparsity as it is the only pattern currently supported by native acceleration on commodity GPUs (see Chapter 2). We enable this flexibility through two distinct optimization strategies. For maximum accuracy, we employ a joint optimization method that tunes both the sparsity pattern within the 2:4 tiles and the tile-level config- urations during training. For scenarios with tighter compute budgets, we offer a variant that tunes only the location of the dense tiles while keeping the initial 2:4 mask fixed. Importantly, unlike theoretical hybrid methods that never see deployment, PATCH is fully compatible with tile-level sparsity acceleration libraries and compilers such as STOICC [ 122]. This makes it the first hybrid sparsity method to demonstrate practi- cal speedups on commodity hardware. For instance, on LLaMA-2 7B running on a consumer-grade A6000 GPU, PATCH achieves 1.18×–1.38×end-to-end speedup over the dense baseline while improving accuracy by 0.37%–2.96% compared to the state-of-the-art 2:4 pruning method, MaskLLM. Note that this chapter fo- cuses exclusively on refining the sparsity pillar; the integration of PATCH with quantization and low-rank approximation is presented in Chapter 7, where we demonstrate the combined efficacy of the Compression Trinity. 6.3 Additional Related Work 6.3.1 Pruning methods Pruning is one of the most widely studied approaches for compressing deep neural networks, with the goal of removing redundant parameters while preserving accuracy. Classical pruning methods can be broadly categorized intolocal(layer-wise) andglobal(end-to-end) strategies. Local pruning.Local approaches prune each layer independently, typically by minimizing reconstruction error within that layer. A seminal example is Optimal Brain Surgeon (OBS) [54,35], which leverages second- CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS69 Mask Value 3-2-4 4-1-2 -31-2 0.950.150.01 0.980.40.2 0.050.550.15 Generate 2:4 Mask (Frozen or Jointly Learned) 01 Gumbel SoftmaxInterleave Repeat Figure 6.1: Illustration of the PATCH learning process for generating tile-level hybrid masks. Each tile is parameterized by a learnable distribution and sampled with Gumbel Softmax to produce ̃ M tile . The dense probabilityisexpandedandmergedwitha2:4mask ̃ M 2:4 , whichcanbefixedorjointlylearnedduringtraining, yielding ̃ M. The final mask assigns each tile to remain dense or follow the 2:4 pattern, enabling flexible sparsity across the weight matrix. order information to identify and remove weights while updating the remaining parameters to compensate for loss. Whilehighlyprincipled, thequadraticcostofcomputingandinvertingtheHessianmakesOBSinfeasible for large models. Recent work adapts these ideas to LLM-scale pruning. SparseGPT [36] formulates layer-wise pruning as a sparse regression problem, enabling efficient approximations of OBS that scale to billion-parameter models. Thanos [ 64] further improves accuracy by employing multi-column approximations to reduce error accumu- lation. Wanda [136], on the other hand, discards explicit weight updates and instead uses a simple magnitude- activation criterion with calibration data, yielding competitive quality with extremely fast runtimes. Despite their efficiency, local methods often suffer from limited capacity to recover accuracy since pruning decisions ignore cross-layer dependencies. Global pruning.Global approaches aim to jointly optimize pruning decisions across layers, typically lead- ing to better overall trade-offs. Optimal Brain Damage (OBD) [75] is an early global method that estimates weight saliency using the diagonal Hessian. Extensions such as WoodFisher [ 134] approximate the Hessian via Kronecker factorizations, making computation more tractable but still challenging for modern LLMs [99]. More recent approaches bypass costly second-order computations. MaskLLM [33] formulates pruning as a binary classification task (keep vs. prune) and solves it using standard optimizers such as AdamW [ 86], achieving strong results even under hardware-friendly structured sparsity (e.g.,2:4). ProxSparse [83] instead adopts a proximal regularization framework, reducing the overhead of MaskLLM while trading off some pruning accuracy. These works highlight the tension between pruning quality and efficiency: global methods often achieve higher accuracy but remain more computationally expensive than simple one-shot local pruning. 6.3.2 Complementary compression techniques Beyond pruning, several orthogonal compression techniques are widely used and can be combined with spar- sity for additional gains.Quantizationreduces the bit precision of parameters and activations, e.g., from 32-bit floating point to 8- or 4-bit integers, thereby reducing memory footprint and accelerating inference [ 43,125]. CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS70 Low-rank adaptationmethods decompose weight matrices into smaller factors, effectively reducing pa- rameter counts while maintaining expressivity. Recent approaches such as LQ-LoRA [48], SLiM [100], and SLoPe [101] demonstrate that low-rank structures can be used both for efficient fine-tuning and for direct model compression. Finally,knowledge distillation[46] transfers knowledge from a large teacher model to a smaller student, yielding compact models that retain much of the teacher’s performance. These methods are complementary to pruning, and hybrid frameworks that integrate sparsity, quantization, and low-rank factorization represent a promising direction for achieving high compression ratios without sacrificing accuracy. 6.4 Preliminaries Differentiable Sampling.Sampling from a categorical distribution is inherently non-differentiable, which poses challenges for gradient-based optimization. The Gumbel Softmax [66] addresses this by combining the Gumbel-Max reparameterization trick together with a softmax relaxation. The reparameterization expresses the sampling process by decoupling the deterministic log-probabilitiesp∈R n from the stochastic perturba- tionsz∈R n introducedbyGumbelnoise, whichemulaterandomdrawsfromthedistribution. Thesubsequent softmax yields a differentiable approximation to categorical sampling: GS(p;τ) k = exp((p k +z k )/τ) P j exp((p j +z j )/τ) (6.1) wherez k =−log(−log(u k ))withu k ∼Uniform(0,1). The resulting vector GS(p;τ)∈R n is a soft index vector whose entries GS(p;τ) k represent the relaxed probability of selecting classk. Additionally, the temperature parameterτcontrols the hardness of the sampled index. Lower values ofτ yield a more peaked distribution, causing GS(p)to converge to a one-hot vector asτ→0. Learnable 2:4 Mask.MaskLLM [33] formulates 2:4 mask selection as a learnable probabilistic process over the six possible patterns. The underlying weights remain fixed, while training shifts the categorical distribution to favor masks that preserve better pruning performance. The mask for each four consecutive elements can be parameterized with a vector p ∈ R 6×1 . Scaling this vector to a weight matrix W ∈ R d 1 ×d 2 will result inP 2:4 ∈R 6× d 1 d 2 4 as the mask search parameters. The resulting mask can be computed as in Equation 6.2, where ̃ M 2:4 ∈[0,1] d 1 ×d 2 denotes the 2:4 soft mask, obtained as a weighted average over the candidate masks, andS∈R 6×4 is the matrix containing these six candidates as its rows. 1 ̃ M 2:4 =reshape(GS(P 2:4 ;τ, κ)×S,R d 1 ×d 2 )(6.2) A scaling factorκis also introduced inEquation 6.1, where it multiplies the logitspbefore adding the Gumbel noisez, thereby controlling their relative influence. Smallκvalues let the noise dominate, encourag- ing exploration across candidate masks, while largerκvalues amplify the logits and make the sampling more deterministic. 1 We will refer to a mask value of1askeepingthe corresponding weight and a value of0aspruningit. CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS71 6.5 PATCH To overcome the rigidity of fixed 50% 2:4 sparsity, we introduce PATCH. PATCH learns a structured mask— optimized on top of frozen weights—that is partitioned into tiles, where each tile decides whether its corre- sponding weights remain dense or are pruned with a 2:4 pattern. This design preserves accuracy in sensitive regions while exploiting hardware-accelerated sparsity elsewhere. Unlike fixed 2:4 sparsity, which enforces the same pattern across all weights, PATCH adapts at the tile level by assigning dense tiles to critical regions and sparse tiles elsewhere. Finding the optimal allocation of dense tiles (value 1) and sparse tiles (2:4 pattern) within a mask is a combinatorially difficult problem, as the number of possible configurations grows rapidly with the number of tiles across the LLM. By also modeling this problem as a probabilistic sampling process, and adjusting the probability of each tile (and the 2:4 patterns within sparse tiles), PATCH can efficiently explore the space of configurations and converge toward masks that balance accuracy and sparsity. The mask distributions are learned end-to-end by training the Gumbel–Softmax logits while keeping the model weights frozen. We address this challenge by formulating mask selection as two coupled subproblems:(1) selecting which tiles are dense or sparse, and(2) choosing the 2:4 sparsity pattern within sparse tiles. Tile-based pruning of LLMs.We associate each parameter matrixW∈R d 1 ×d 2 with a grid of tile-level distributions, each parameterized by a learnable logit. Collectively, these formP tile ∈R d 1 b 1 × d 2 b 2 , where each entry specifies the unnormalized score of keeping the correspondingb 1 ×b 2 tile fully dense. To create a two-class distribution (keep dense vs. prune), we concatenate a fixed zero to each logit, yielding [ P tile , 0] ∈ R d 1 b 1 × d 2 b 2 ×2 . After applying Gumbel–Softmax, we broadcast the dense probabilities across their respective b 1 ×b 2 region (since the weighted average of the two outcomes reduces top dense ·1 +p prune ·0 =p dense ), so that all elements of a tile receive the same mask value. Formally, ̃ M tile =GS([P tile ,0];τ, κ) :,:,0 ⊗1.(6.3) This yields the tile-level mask ̃ M tile ∈[0,1] d 1 ×d 2 inEquation 6.3, where1∈R b 1 ×b 2 is an all-ones matrix and⊗denotes the Kronecker product. Joint optimization with sparse mask.To fully determine the effective sparsity pattern, the tile-level mask must be combined with the fine-grained 2:4 mask. Assuming that the 2:4 mask ̃ M 2:4 is generated using Equation 6.2, PATCH combines it with the tile mask ̃ M tile as shown inEquation 6.4. The resulting soft mask interpolates between dense and sparse behavior: values of ̃ M tile close to one make the tile predominantly dense, while values close to zero shift the tile toward the soft 2:4 mask pattern defined by ̃ M 2:4 . Thus, ̃ Mcan be understood as a per-tile weighted average of the dense option and the 2:4 patterns, with ̃ M tile determining the relative contribution of each. An overview of the process is provided in Figure 6.1. ̃ M= ̃ M tile + 1− ̃ M tile ⊙ ̃ M 2:4 (6.4) Learning masks with targeted sparsity.PATCH uses a novel regularization term to achieve a flexible 0%– 50% sparsity ratio across the model by controlling the number of dense tiles. Unlike traditional regularization methods like weight decay, which produce non-deterministic sparsity ratios, our term penalizes deviations from the target sparsity, enabling precise control. This global sparsity approach prunes sensitive linear layers CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS72 Algorithm 4 Joint Tile & 2:4 Mask Learning Input:Weight matrixW, tile size(b 1 ,b 2 ), sparsity targetρ, training stepsT, loss hyperparametersλ 1 ,λ 2 , temperature scheduleτ t T t=1 , scaling scheduleκ t T t=1 . Output:Learned pruning masksM ⋆ , pruned weights b W. 1Initialize tile logitsP tile ∈R d 1 b 1 × d 2 b 2 . 2InitializeP tile with one-shot prior. 3Initialize differentiable 2:4 parametersP 2:4 ∈R 6× d 1 d 2 4 . 4fort= 1→Tdo 5 ̃ M tile ←GS([P tile ,0];τ t , κ t ) :,:,0 ⊗1 b 1 ×b 2 ▷Dense soft tile mask 6 ̃ M 2:4 ←Equation 6.2▷Differentiable 2:4 mask 7 ̃ M i ← ̃ M tile + (1− ̃ M tile )⊙ ̃ M 2:4 ▷Merge masks 8Compute loss: L=L LM (x; ̃ M⊙W) +λ 1 P i ̃ M i P i ∥W i ∥ 0 −ρ 1 −λ 2 P i ∥ ̃ M i ⊙W i ∥ 2 2 P i ∥W i ∥ 2 2 9UpdateP tile ,P 2:4 via backpropagation. 10end for 11M ⋆ tile ←1[P tile >0]⊗1 b 1 ×b 2 ▷Hard tile mask 12M ⋆ 2:4 ←select best 2:4 mask fromP 2:4 . 13M ⋆ i ←M ⋆ tile + (1−M ⋆ tile )⊙M ⋆ 2:4 . 14 b W←W⊙M ⋆ i ▷Final pruned weights Return:Learned maskM ⋆ , pruned weights b W . less aggressively while setting redundant weight elements to zero, offering greater flexibility than fixed per- layer sparsity. We directly compare global versus per-layer sparsity regularization inSection 6.7. Training objective.The overall training objective, as shown inEquation 6.5, of PATCH combines three components: the standard modeling loss, a sparsity regularization term that enforces the target density of the modelρ, and a weight regularization term (as in MaskLLM) that promotes larger weight magnitudes and gradient propagation. Formally, L=L LM x; ̃ M i ⊙W i +λ 1 P i ̃ M i P i ∥W i ∥ 0 −ρ 1 −λ 2 P i ∥ ̃ M i ⊙W i ∥ 2 2 P i ∥W i ∥ 2 2 (6.5) Following MaskLLM, we progressivelydecreaseτandincreaseκduring training so that the Gumbel- Softmax distribution converges to a clear one-hot choice of mask by the end of training. Inference.After training, the sign of each logit inP tile determines the final mask. Since a zero logit is concatenated to represent the sparse class (Equation 6.3), positive values correspond to the dense option, while negative values correspond to the sparse option. The complete procedure is outlined inAlgorithm 4. Memory efficient PATCH.To further reduce overhead, PATCH can be run in a memory-efficient manner by freezing the sparse mask parameters and optimizing only the tile-level decisions. This reduces the number of learnable parameters to d 1 d 2 b 1 b 2 . While this lighter formulation limits mask-selection flexibility and can reduce performance as seen in Table 6.7, it makes training feasible under strict memory constraints, such as fitting an CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS73 8B model on a single 80GB GPU. We denote this version of PATCH by PATCH Tile and the joint optimization version of PATCH by PATCH Joint . 6.6 Efficient deployment of PATCH Executing PATCH requires handling hybrid sparse–dense tiles, a capability not supported by existing GPU libraries. Current tools either focus exclusively on dense computation (e.g., cuBLAS [109], dense CUTLASS [23], OpenAI Triton [143]), or restrict support to fixed2:4sparsity (e.g., cuSPARSELt [110], sparse CUT- LASS). STOICC [122] lifts these limitations by extending Triton with hybrid tile-level sparsity, making it a suitable backend for accelerating PATCH. Similar to Triton, STOICC employs an inspector that benchmarks candidate kernel configurations for each sparsity ratio, identifying the most hardware-efficient tile size for the target GPU. On NVIDIA A100 and A6000 GPUs, our experiments show that the optimal configurations are consistently drawn from128×128 or its subdivisions (e.g.,128×64,64×128,64×64). In practice, this means that regardless of the sparsity ratio or the layer shape, the chosen128×128granularity guarantees that STOICC’s autotuned tiles can be applied consistently. Unless otherwise specified, we adopt these hardware-friendly tile sizes in all PATCH experiments. Further implementation details are provided in Section D.1. 6.7 Experiments Model, dataset and evaluation.We evaluate PATCH across diverse transformer architectures, including the Qwen-2.5 [118], Gemma 3 [139], and LLaMA-2 [144] and 3 [31] model families, spanning 500M to 8B parameters. Following the dataset size and configurations in MaskLLM [33], masks are trained for 2000 steps with a batch size of 256 on sequences with a length of 4096 tokens from the SlimPajama dataset [ 135]. Following previous LLM compression work [100,33], we evaluate the models on eight zero-shot down- streamtasks: PIQA[13],ARC-EasyandARC-Challenge[21],Winogrande[126],OpenBookQA[97],RACE[74], HellaSwag [ 161], and MMLU [57] using the Language Model Evaluation Harness [42] framework. Addition- ally, similar to previous work [100,36,136], we evaluate the models on a language modeling task using the WikiText2 [95] dataset with a sequence length of 4096, comparing against established baselines in the following sections. Baselines.To evaluate PATCH against established 2:4 sparsity pruning techniques, we compare it with the state-of-the-art learnable method MaskLLM [33], as well as one-shot methods including Wanda [136], SparseGPT [36], Thanos [64], ProxSparse [83] and magnitude pruning [53]. For one-shot pruning methods, following the default configurations in each paper, we prune the models over 128 samples from the C4 dataset. The publicly available MaskLLM pruned checkpoints are limited to LLaMA-2 7B and LLaMA-3.1 8B models. To ensure a fair comparison across all models, we implemented MaskLLM in PyTorch and replicated its results for additional architectures presented in this study. We faced a similar challenge with ProxSparse as well, where only the LLaMA-2-7B and LLaMA-3.1-8B checkpoints are publicly available. We have pruned other models with their official code base using their default hyperparameters for comparison. Experiment Setup.All masks are trained using the HuggingFace Trainer API [ 153] for 2000 steps with a global batch size of 256 and a sequence length of 4096, processing 2B tokens from the SlimPajama corpus CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS74 Table 6.2: Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By jointly optimizing the location of dense tiles and the sparsity pattern within the sparse tiles, PATCH Joint allows for a continuous sparsity ratio for the models, providing a flexible tradeoff between sparsity and model quality. Sparsity Method PatternQwen-2.5 0.5B LLaMA-3.2 1BGemma-3 1B Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) 0%Dense-46.00 12.08 47.709.0647.01 11.67 50% Magnitude 2:430.16 6734.97 29.66 563.44 31.66 5005.56 Wanda2:432.97 72.48 31.61 78.18 34.16 69.41 SparseGPT 2:434.81 36.59 35.55 32.73 35.58 44.59 Thanos 2:431.31 37.32 35.71 33.03 35.09 62.63 ProxSparse 2:432.05 111.05 33.55 49.33 36.63 90.50 MaskLLM 2:439.33 15.22 41.04 12.93 41.84 12.82 45%PATCH Joint Dense/2:4 Tiles40.2914.5742.0812.2342.8011.96 35%PATCH Joint Dense/2:4 Tiles41.1513.8442.7211.6743.3011.48 25%PATCH Joint Dense/2:4 Tiles42.3913.4743.8111.0044.0711.17 [135]. Training is accelerated via data parallelism across a single node with 4 H100 GPUs. In this setup, PATCH Joint requires 4.5 and 6 GPU hours on the 0.5B and 1B models, respectively, while PATCH Tile requires 21 and 24 GPU hours on the 7B and 8B models. The hyperparameters for PATCH Joint and PATCH Tile are summarized in Table 6.1, tuned on Qwen-2.5-0.5B. For the 2:4 mask parameters, we follow the configuration from MaskLLM [33]. Table 6.1: Hyper-parameters used for PATCH Joint and PATCH Tile across sparsity ratios. All hyper parameters were tuned on Qwen-2.5-0.5B. Sparsity Method Optimizer Logits Init Gumbel Scaling Gumbel Prior(Strength) Sparse Reg. Weight Reg. 25% PATCH Joint Adam(0.001)N(0,0.014)25→3502→0.05SparseGPT(3)710 35% PATCH Joint Adam(0.001)N(0,0.014)25→3502→0.05SparseGPT(3)710 45% PATCH Joint Adam(0.001)N(0,0.014)25→3504→0.05SparseGPT(3)710 25% PATCH Tile Adam(0.0001)N(0,0.014)100→5002→0.05SparseGPT(3)30.1 35% PATCH Tile Adam(0.0001)N(0,0.014)100→5002→0.05SparseGPT(3)30.1 45% PATCH Tile Adam(0.0001)N(0,0.014)100→5002→0.05SparseGPT(3)30.1 6.7.1 Model Quality Results Joint sparse and dense tile optimization.For smaller models like Qwen-2.5 0.5B, LLaMA-3.2 1B, and Gemma-3 1B, we apply the joint variant PATCH Joint , which simultaneously optimizes dense tile locations and sparsity patterns within sparse tiles. This approach enables effective performance. The average accuracy of the models across eight zero-shot downstream tasks and their perplexity on the WikiText2 dataset is reported in Table 6.2. The results demonstrate that PATCH Joint provides a flexible tradeoff between sparsity ratio and model quality, narrowing the performance gap to dense models while ensuring hardware-friendly inference. A similar pattern holds for larger models using a memory-efficient variant, as explored next. CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS75 Table 6.3: Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By only optimizing the location of dense tiles while keeping sparsity pattern within the sparse tiles frozen, PATCH Tile provides a memory efficient variant for PATCH Joint , allowing for a continuous sparsity ratio for the models and providing a flexible tradeoff between sparsity and model quality. Sparsity Method PatternLLaMA-2 7BLLaMA-3.1 8B Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) 0%Dense-54.615.1260.315.84 50% Magnitude 2:443.44 54.39 35.93 765.92 Wanda2:444.30 11.15 41.77 21.29 SparseGPT 2:445.09 10.12 45.53 15.11 Thanos 2:444.80 11.19 45.72 16.09 ProxSparse 2:445.929.1845.14 15.17 MaskLLM 2:448.626.7852.808.58 45%PATCH Tile Dense/2:4 Tiles48.996.5553.608.20 35%PATCH Tile Dense/2:4 Tiles50.086.1855.287.89 25%PATCH Tile Dense/2:4 Tiles51.585.8656.487.34 Memory-efficienttileselection.ForlargermodelssuchasLLaMA-27BandLLaMA-3.18B,weemploythe memory-efficient variant PATCH Tile , which freezes the fine-grained sparse weight structure while optimizing dense tile selections. Table 6.3summarizes the average accuracy of the models across eight downstream tasks in addition to their perplexity on the WikiText2 dataset for different sparsity ratios, illustrating that PATCH Tile delivers a comparable flexible sparsity-quality tradeoff when using a high-quality frozen 2:4 mask. Overall, across Table 6.2andTable 6.3, PATCH consistently surpasses one-shot methods like Wanda, SparseGPT, and magnitude pruning due to its end-to-end training on large corpora. While MaskLLM also trains end-to-end on a large dataset, its fixed 2:4 sparsity ratio limits achievable accuracy and perplexity. In contrast, PATCH overcomes this limitation with flexible dense tile allocation, achieving accuracy gains and perplexity reductions from 45% to 25% sparsity that progressively align with dense model performance. The full per-task accuracy results are provided in Section D.2. Comparisonwithunstructuredsparsity In this section, we compare the quality of the models pruned with PATCHagainstotherunstructuredsparsitymethods. Table6.4summarizestheaverageaccuracyofthemodels across eight downstream tasks and the model perplexity on WikiText2 dataset. The results indicate that while unstructured sparsity consistently outperforms the hybrid sparsity, the gap between the two is not significant, showing that PATCH is helping to bridge the gap between unstructured sparsity and semi-structured sparsity. 6.7.2 Understanding the components of PATCH This subsection examines the design choices driving PATCH’s performance by analyzing its behavior across various configurations on the Qwen-2.5 0.5B model. Tile size.We initially assess the impact of tile size on PATCH’s performance, fixing hyperparameters to those optimized for 128×128 tiles. Table 6.5reveals that4×4tiles maximize model quality through finer CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS76 Table 6.4: Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for PATCH, Wanda, and SparseGPT. For models with less than or equal to 1B parameters, PATCH Joint opti- mizes both dense tile locations and sparsity patterns, while for larger models PATCH Tile optimizes only dense tile locations with frozen sparsity patterns, both using Dense/2:4 Tiles pattern allowing continuous sparsity ratios and flexible tradeoffs between sparsity and model quality. Wanda and SparseGPT are unstructured prun- ing methods. Sparsity Method PatternQwen-2.5 0.5B LLaMA-3.2 1BGemma-3 1BLLaMA-2 7BLLaMA-3.1 8B Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) 45%PATCHDense/2:4 Tiles40.2914.5742.0812.2342.8011.9648.996.5553.608.20 45% WandaUnstructured41.45 18.81 40.76 16.56 42.87 25.38 52.726.3655.678.24 45% SparseGPT Unstructured42.31 17.65 42.66 15.01 43.52 22.26 52.776.4656.708.21 35%PATCHDense/2:4 Tiles41.1513.8442.7211.6743.3011.4850.086.1855.287.89 35% WandaUnstructured43.46 15.04 44.60 11.95 45.50 16.98 54.375.8758.687.02 35% SparseGPT Unstructured44.66 14.79 45.62 11.68 45.45 16.92 54.185.9258.817.07 25%PATCHDense/2:4 Tiles42.3913.4743.8111.0044.0711.1751.585.8656.487.34 25% WandaUnstructured45.70 13.70 46.50 10.46 46.56 15.14 54.605.6559.806.54 25% SparseGPT Unstructured45.28 13.63 46.52 10.42 46.37 15.05 54.715.6859.526.55 Table 6.5: Impact of PATCH’s tile size across sparsity levels (↓is better). The effect of tile size on model quality is not significant, showing PATCH’s robustness against tile size. Sparsity (0.5B) 128 64 32 1684 45%14.57 14.66 14.70 14.67 14.7014.55 35%13.84 14.08 14.15 14.03 14.0113.72 25%13.47 13.54 13.52 13.53 13.4013.11 Table 6.6: Global sparsity yields bet- ter quality by concentrating pruning in less important blocks and preserving density elsewhere (↓is better). Sparsity (0.5B) Global Layer-wise 45%14.5715.17 35%13.8414.48 25%13.4713.95 sparse-densecontrol, thoughlargertilesizesshowminimalvariation, suggestingrobustness. However, smaller tiles may hinder hardware efficiency, requiring a balance with hardware specifications. Jointvs. tile-onlymasksearch.We then analyze the impact of fixing the 2:4 masks and optimizing only tile masks.Table 6.7shows that among frozen 2:4 masks, MaskLLM provides the strongest results. On the other hand, one-shot pruning methods perform comparably at higher sparsity levels but diverge at lower sparsity, with SparseGPT emerging as the best overall. When comparing against our full approach, joint optimization of both tile and 2:4 masks consistently outperforms tile-only training across sparsity ratios. Nevertheless, tile- only training remains a practical alternative for larger models in resource-constrained settings, as also reflected in Table 6.3. Sparsity allocation.We analyze how sparsity is allocated across transformer blocks under a global target. Across models, deeper transformer blocks are pruned far less, while the initial blocks also tend to receive lighter pruning depending on the architecture. By contrast, the middle blocks consistently absorb most of the sparsity, suggesting that they contain more redundancy ( Figure 6.2). We compare this flexible allocation to enforcing sparsity uniformly at the layer level. As shown inTable 6.6, global targets deliver better results by pruning more aggressively in redundant layers while preserving capacity in sensitive ones. In contrast, layer-wise targets impose uniform sparsity that can over-prune critical components [ 79,156,78,158]. On top of variation across depth, sparsity is also distributed unevenly across the individual linear layers within each transformer block.Figure 6.3breaks down the allocation into the query, key, value, and output CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS77 Table 6.7: Impact of fixed 2:4 mask selection for PATCH Tile , compared with joint optimization (↓is better). PATCH Joint achieves the lowest perplexity overall, while for PATCH Tile , MaskLLM provides the best frozen mask. Sparsity (0.5B) MaskLLM SparseGPT (w/o weight update) Wanda MagnitudePATCH Joint 45%15.0621.8421.8321.3314.57 35%14.5517.2917.9619.9013.84 25%14.1714.8915.0916.0513.47 036912151821 Block Index 0.0 0.2 0.4 Sparsity Ratio Qwen 2.5 0.5B Global Sparsity 45% 35% 25% 03691215182124 Block Index 0.0 0.2 0.4 Gemma 3 1B 02468101214 Block Index 0.0 0.2 0.4 Llama 3.2 1B Figure 6.2: Layer-wise sparsity allocation under different global sparsity budgets for various models. PATCH achieves the target global sparsity while flexibly distributing pruning across transformer layers. matrices of the attention module, as well as the up, gate, and down matrices of the MLP for the Qwen 2.5 0.5B model. The up, gate, and down layers absorb most of the sparsity and largely explain the overall allocation pattern seen in Figure 6.2. In contrast, the attention module is treated as more critical. The key and value matrices are never pruned, while the output matrix shows moderate pruning at higher global sparsity targets. The query matrix is pruned the most, suggesting it is the least important within the attention submodule. Additionally, we provide the sparsity distributions for the Gemma-3-1B ( Figure 6.4) and Llama-3.2-1B (Figure 6.5) models, as referenced in the main text. Similar to the Qwen-2.5 0.5B model, the patterns observed here indicate that MLP layers (up, gate, and down matrices) are pruned more aggressively, absorbing the majority of sparsity. In contrast, the self-attention layers are treated as more critical, with key and value matrices remaining largely dense or unpruned, while the query matrix experiences the highest pruning within the attention submodule, and the output matrix shows moderate pruning under higher global sparsity targets. This consistent behavior across models underscores the redundancy in MLP components and the sensitivity of attention mechanisms. 6.7.3 Speedup and memory savings We evaluate the inference efficiency of the LLaMA-2 7B model pruned with PATCH using the STOICC [122] compiler. With a batch size of 16 on an A6000 GPU, we observe end-to-end throughput improvements of 1.18×, 1.27×, and 1.38×at sparsity levels of 25%, 35%, and 45%, respectively, compared to the dense baseline. At the same sparsity levels, the model’s GPU memory footprint during inference is also reduced, dropping to 0.76×, 0.68×, and 0.59× of the fully dense model, respectively. These results underscore the trade-off between accuracy retention and the computational savings enabled by sparsity. CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS78 036912151821 0.00 0.25 0.50 Sparsity Ratio Query Matrix 036912151821 Key Matrix 036912151821 Block Index 0.00 0.25 0.50 Sparsity Ratio Value Matrix 036912151821 Block Index Output Matrix 036912151821 Up Matrix 036912151821 Gate Matrix 036912151821 Block Index Down Matrix Global Sparsity 45 35 25 Self-Attention MatricesMLP Matrices Figure 6.3: Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Qwen-2.5 0.5B. 6.8 Conclusion and Limitations In this chapter, we explored the upper limits of the sparsity pillar when a training budget is available. We introduced PATCH, a hybrid sparsity framework that breaks the rigidity of the layer-wise static masks used inChapter 5. By partitioning weight matrices into tiles designated as either dense or 2:4 sparse, PATCH enables adaptive sparsity ratios between 0% and 50%, dynamically balancing accuracy in critical regions with hardware acceleration elsewhere. Experiments across models up to 8B parameters show that PATCH consistently improves accuracy over state-of-the-art 2:4 pruning methods while achieving up to 1.38×end-to-end speedup on consumer-grade GPUs. These results demonstrate the promise of hybrid sparsity as a practical approach to efficient LLM inference and motivate future work on broader sparsity formats, integration with quantization, and co-design with hardware kernels. While PATCH offers superior accuracy-efficiency trade-offs, it is important to situate it within the broader ”compute budget” narrative of this thesis. Unlike OPTIMA, which targeted the zero-training regime, PATCH requires afine-tuning budget. The learnable masking process introduces computational overhead that makes it unsuitable for instant, on-device adaptation. However, as established in Section 1.4, these two chapters represent complementary solutions for different deployment scenarios: OPTIMA maximizes performance when training is impossible, whereas PATCH maximizes performance when resources permit. With PATCH, we have now refined the sparsity pillar to its logical conclusion in both static and dynamic regimes. We have made it highly accurate (via hybrid masks) and physically fast (via tile-level kernels). Yet, we have applied this pillar in isolation. As noted in the ”Sparsity Paradox” ( Table 1.1), even optimized sparsity eventually hits an accuracy wall that cannot be overcome by simply removingfewerweights. To unlock the next frontier of performance, we must stop treating sparsity as a solo actor. In the next chapter, we present the culmination of this thesis: SLIM. There, we integrate our optimized sparsity findings with aggressive quantization and low-rank approximations, finally realizing the full potential of the Compression Trinity in a CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS79 03691215182124 0.00 0.25 0.50 Sparsity Ratio Query Matrix 03691215182124 Key Matrix 03691215182124 Block Index 0.00 0.25 0.50 Sparsity Ratio Value Matrix 03691215182124 Block Index Output Matrix 03691215182124 Up Matrix 03691215182124 Gate Matrix 03691215182124 Block Index Down Matrix Global Sparsity 45 35 25 Self-Attention MatricesMLP Matrices Figure 6.4: Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Gemma-3 1B. unified, one-shot framework. CHAPTER 6. PATCH: LEARNABLE TILE-LEVEL HYBRID SPARSITY FOR LLMS80 02468101214 0.00 0.25 0.50 Sparsity Ratio Query Matrix 02468101214 Key Matrix 02468101214 Block Index 0.00 0.25 0.50 Sparsity Ratio Value Matrix 02468101214 Block Index Output Matrix 02468101214 Up Matrix 02468101214 Gate Matrix 02468101214 Block Index Down Matrix Global Sparsity 45 35 25 Self-Attention MatricesMLP Matrices Figure 6.5: Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in LLaMA-3.2 1B. Chapter 7 SLIM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression Publication and Contributions.The content of this chapter is based on the paper “SLiM: One-shot Quan- tized Sparse Plus Low-rank Approximation of LLMs” [100], published at Forty-Second International Confer- enceonMachineLearning(ICML),2025. ThisworkwasconductedincollaborationwithAmirYazdanbakhsh and Maryam Mehri Dehnavi. Mohammad Mozaffari was the lead contributor, responsible for the algorithm design, implementation, and experimental evaluation. Amir Yazdanbakhsh and Maryam Mehri Dehnavi su- pervised the project and contributed to the writing and revision of the manuscript. 7.1 Introduction In the preceding chapters, we pushed theSparsitypillar to its limits. With OPTIMA ( Chapter 5), we estab- lished the mathematical upper bound for reconstruction with layer-wise pruning and weight update, and with PATCH ( Chapter 6), we broke the rigidity of those masks using end-to-end learnable hybrid patterns. How- ever, as identified in the ”Sparsity Paradox” (Section 1.3), sparsity is the most destructive pillar. Even with the optimized skeletons provided by OPTIMA and PATCH, a performance gap remains because we are removing information that cannot be fully recovered by weight adjustment alone. Additionally, with a perfected sparse model, relying on a single compression technique limits the potential for efficiency. To bridge this gap and fully conquer the memory bandwidth bottleneck, we must stop treating sparsity in isolation. We must integrate it with the remaining two pillars of the Compression Trinity:QuantizationandLow-Rank Approximation. Thischapterpresentstheculminationofthisthesis: SLIM,aunifiedframeworkthatjointlyapplieshardware- friendly sparsity, quantization, and low-rank approximations in a single one-shot step. Bringing these three pillars together presents a formidable challenge:compounded error. When aggressive sparsity (removing weights) meets aggressive quantization (reducing precision), the errors do not merely add up; they exacer- bate each other, leading to a catastrophic drop in model capability [ 36,129,90,81,49]. Traditional recovery methods rely on costly retraining (e.g., Quantization-Aware Training [127,114]), which violates the ”resource- constrained” principles we explored in Chapter 5. Conversely, existing one-shot methods like SparseGPT [36] 81 CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 82 struggle to combine structured patterns (like 2:4 sparsity) with low-bit quantization, failing to arrest the accu- racy slide. 1 To resolve these limitations and realize the full Trinity without retraining, we propose SLIM. We decom- pose the problem into three synchronized sub-tasks, ensuring that each pillar supports the others rather than conflicting with them. 1.Quantization:We prioritize uniform quantization for hardware efficiency. While SLIM is compatible with any standard quantization kernel (e.g., Group MinMax or AbsMax), we introduce SLIM-Quant, a probabilistic and tractable reformulation that finds the optimal quantization parameters. This allows users to either leverage existing quantization standards or utilize our optimizer to minimize the error floor before sparsity is even applied. 2.Sparsity:We apply hardware-friendly pruning to the quantized weights to create the efficient sparse structure. Crucially, the SLIM framework is agnostic to the specific mask selection algorithm. This allows us to seamlessly integrate masks generated by Wanda [136], OPTIMA, PATCH, or future state- of-the-art selectors, ensuring the framework remains relevant as pruning metrics evolve. 3.Low-RankApproximation:Finally, we deploy the third pillar not just for compression, but as amathe- maticallyderivederror-correctionmechanism. This is the critical innovation that closes the accuracy gap left byChapter 5andChapter 6. We propose SLIM-LoRA, a one-shot low-rank adaptation method designed to compensate for the aggregated error. Unlike standard adapters that require iterative training to find optimal values [ 61,104,28,48,80], we develop a saliency function that is both invertible and additive. These properties enable us toanalytically computethe optimal low-rank adapter values that minimize the compression-induced error in one shot, eliminating the need for any retraining overhead. By solving the compounded error problem through this joint formulation, SLIM shifts the Pareto frontier of efficiency. It achieves what neither OPTIMA nor PATCH could do alone: recovering close to dense model accuracy while maintaining high compression rates. Compared to state-of-the-art methods, SLIM achieves an average accuracy improvement of 5.66% on LLaMA-2-7B under 2:4 sparsity and 4-bit quantization. Uniquely, it delivers higher model accuracy at thesame total bit budgetcompared to existing techniques (up to 0.5%) and even outperforms uncompressed dense models at equal parameter budgets (up to 0.6%). Beyond accuracy, SLIM demonstrates the practical value of the Trinity, achieving up to 3.78×and 3.75×layer-wise speedup on NVIDIA RTX3060 and A100 GPUs, respectively. For cases requiring maximal performance, we also support an optional lightweight PEFT method, providing up to an 1.66% additional accuracy improvement. 7.2 Related work SLIMcombinesmodelpruningandquantizationforcompression, complementedbyzero-shotlow-rankadapters to recover lost accuracy. This section reviews related work on these topics. 7.2.1 Pruning Eliminating redundant weights reduces computation and memory costs during inference. Optimal Brain Dam- age (OBD) [ 75] leverages second-order information of the loss function to identify the least important weights 1 For a more detailed discussion of the related work, seeSection 7.2. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 83 but is computationally prohibitive for large language models (LLMs) [99]. WoodFisher [134] approximates the Hessian matrix using Kronecker Factorization to mitigate this overhead but struggles to scale to LLMs. Optimal Brain Surgeon (OBS) [54] evaluates weight matrices layer-wise using the layer-wise Hessian matrix to preserve layer outputs. However, the cubic growth in the cost of inverting the layer-wise Hessian with model size renders this approach impractical for LLMs. Optimal Brain Compression (OBC) [35] addresses the OBS-defined compression problem using a greedy algorithm, while SparseGPT reformulates it as a sparse regression problem. Wanda introduces a lightweight method based on weight and activation magnitudes to identify unimportant weights without updating their values. 7.2.2 Quantization Quantizing all elements in a matrix is challenging due to the significant impact of outliers on the model [27]. Group quantization [3,47] addresses this by quantizing small groups of a weight matrix with a shared quantization parameter, but it introduces challenges discussed in Section E.16. AbsMax[65]withround-to-nearest(RTN)isthesimplestquantizationschemeformatrixelements. OPTQ [37] minimizes layer-wise error using an approach akin to OBS. AWQ [81] shifts the challenge of quantizing salient weights to activations, while SmoothQuant [155] balances quantization error between weights and activations, enabling input quantization. OmniQuant [129] improves accuracy with learnable clipping and channel scaling. AffineQuant leverages equivalent affine transformations to reduce quantization error, and QuaRot [ 7] uses rotations to eliminate outliers during quantization. Advanced methods like JSQ [49] jointly prune and quantize weights to 8 bits but struggle to recover accuracy in low bit-width quantization, limiting their utility. 7.2.3 Low-rank Adapters Low-rank adapters were first introduced to LLMs to reduce the overhead of fine-tuning [61,101]. Q-LoRA [ 28] extended this approach by quantizing weights before fine-tuning, allowing the process to recover accuracy lost during quantization. LQ-LoRA [48] further improved Q-LoRA by initializing the adapters using the SVD of the quantization error. LoSparse [80] has a similar approach as LQ-LoRA, but for sparsity, initializing the low-rank adapters to the norm of the pruning error. RoSA [ 104] expands the learning capability of the model by adding both low-rank and sparse adapters to the model. This approach adds an extra sparse matrix multiplication to the inference, increasing the adapter overhead even further. However, all these methods require hundreds of millions of tokens for fine-tuning, making them costly and not comparable to one-shot pruning and quantization methods, or methods that use much shorter fine-tuning phases. L 2 QER [ 162] avoids fine-tuning by using one-shot low-rank adapters to mitigate quantization error. How- ever, it performs poorly when combined with sparsity, resulting in a significant accuracy gap between the compressed and dense models. 7.2.4 Sparse Plus Low-Rank Matrix Decomposition The decomposition of a matrix into the sum of a sparse component and a low-rank component is a classical problem in signal processing and optimization, most prominently studied under the framework of Robust Principal Component Analysis (RPCA) [ 15]. Given an observed matrixM, RPCA seeks to recoverM= L 0 +S 0 , whereL 0 is low-rank andS 0 is sparse, by solving the convex program known as Principal Component CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 84 Pursuit (PCP): min L,S ∥L∥ ∗ +λ∥S∥ 1 subject toL+S=M,(7.1) where∥·∥ ∗ denotes the nuclear norm andλis a regularization parameter. Candès et al. [15] proved that under mild incoherence conditions on the low-rank component, this convex relaxation exactly recovers both L 0 andS 0 even when the sparse errors are arbitrarily large in magnitude. Chandrasekaran et al. [16] provided complementary recovery guarantees through rank-sparsity incoherence conditions, while Zhou et al. [169] extended the theory to the noisy setting (Stable PCP), showing that the decomposition remains robust when the observation also contains a small dense noise term, i.e.,M=L 0 +S 0 +N 0 . ThestructuralparallelbetweenRPCAandmodelcompressionwasfirstexploredinthecontextofdeepcon- volutional networks by Yu et al. [160], who decomposed weight matrices into sparse-plus-low-rank form us- ing greedy bilateral decomposition and showed that this representation achieves better accuracy-compression trade-offs than either pure pruning or pure low-rank factorization in isolation. More recently, the connection to LLM compression has been made explicit. OATS [ 164] formulates post-training weight compression as a sparse-plus-low-rank decomposition problem, scaling the weights by the second moment of input embeddings before decomposition to preserve outlier features. HASSLE-free [91] established a unified framework show- ing that several existing LLM pruning methods, including Wanda and Magnitude Pruning, can be viewed as special cases of an alternating minimization procedure for the sparse-plus-low-rank objective. Concurrently, 3BASiL [ 11] proposed a three-block ADMM formulation that jointly optimizes the sparse and low-rank com- ponents, offering improved convergence guarantees over alternating minimization. SLIM’s formulationW≈W C +LR, whereW C is the sparse (and quantized) component andLRis the low-rank correction, can be viewed through the lens of this decomposition tradition. However, SLIM differs from standard RPCA approaches in several important ways. First, the sparse component in SLIM is addition- ally quantized, introducing a structured noise that is absent in classical RPCA. Second, rather than solving for SandLjointly via convex optimization, SLIM employs a sequential pipeline (quantize, then sparsify, then compute the low-rank correction), which enables the use of off-the-shelf pruning and quantization algorithms but forgoes the joint optimality guarantees of PCP. Third, SLIM’s saliency-weighted SVD ( Equation 7.12) introduces input statistics into the decomposition, a data-dependent weighting that has no direct analogue in standard RPCA but shares the spirit of the outlier-aware scaling in OATS [ 164]. Bertsimas et al. [12] further studied the sparse-plus-low-rank decomposition from a discrete optimization perspective, providing mixed- integer programming formulations that could, in principle, be adapted to the compression setting. 7.3 Preliminaries Model Compression.Model compression reduces the compute and memory demands of large models while maintaining predictive accuracy by minimizing output differences between compressed and original models. However, directly optimizing these differences across the entire model is computationally infeasible due to the high dimensionality of neural networks. Optimal Brain Surgeon (OBS) [ 54] simplifies this challenge by focusing on minimizing output discrepancies layer by layer, using calibration datasets. OBS applies a layer-wise approach to compress feed-forward layers efficiently. Denoting compressed matrices with a superscriptC, for a layer with inputX ∈R b×d in , weightW ∈R d in ×d out , and output Y ∈R b×d out , it minimizes output differences by optimizing Equation 7.2. This method ensures compression fidelity and has become foundational for many modern compression techniques. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 85 SLIM-Quant Saliency-Based Pruning Saliency-Based Low-Rank Adapter Figure 7.1: The SLIM weight compression pipeline consists of three main steps: (1) Quantizing weights using the symmetric SLIM-Quant algorithm, producing quantized weightsW Q and quantization errorE Q ; (2) Sparsifying quantized weightsW Q through a pruning method, resulting in compressed weightsW C and sparsity errorE S ; (3) Mitigating compression errors through SLIM saliency-based low-rank approximation, generating left and right low-rank adaptersLandR. Optionally, these adapters can be fine-tuned with sparse quantized weights frozen to further enhance model accuracy. min W C |Y C −Y| 2 =min W C |X(W C −W)| 2 (7.2) Symmetric Quantization.Symmetric quantization is a core technique for reducing model size and boost- ing computational efficiency. It computes the quantized matrixM Q ∝round( M α ), whereαis a scaling factor based on the range or norm of the matrix. This scaling ensuresM Q values stay within the repre- sentable range, enabling efficient matrix multiplications with minimal overhead. However, its effectiveness depends on selectingαcarefully, as this choice significantly impacts precision. AbsMax, the most common symmetric quantization method, selectsαas the matrix’s maximum absolute value, ensuring all values remain within the target range. Unfortunately, it is highly sensitive to outliers; a single large value can inflateα, reducing the precision of most quantized weights. For zero-centered, bell- curved distributions typical in LLMs, AbsMax maps many weights to zero, leading to significant quantization errors. Group quantization [ 3,47] tackles AbsMax’s outlier sensitivity by assigning separate scaling factors to subgroups of the weight matrix. This approach captures local variations in weight magnitudes, reducing quan- tization error for non-uniform distributions. However, storing multiple scaling factors increases memory us- age, and subgroup-specific dequantization increases computational complexity, potentially slowing inference. The challenges of using group quantization are discussed in Section E.16. 7.4 Quantized sparse plus low-rank approximation of LLMs To achieve effective compression of LLMs while preserving accuracy, SLIM combines quantization, pruning, and saliency-based low-rank adapters into an integrated pipeline. First, SLIM applies SLIM-Quant , a novel scheme designed to minimize quantization error, laying the foundation for subsequent pruning using methods CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 86 such as Wanda [136]. Finally, low-rank adapters are introduced to reduce the impact of compression errors from both quantization and pruning, ensuring minimal accuracy loss. The overall process is illustrated inFig- ure 7.1, providing a visual summary of how these components interact to achieve effective model compression. In the following sections, we dive into the details of each step, highlighting the innovations and contributions of SLIM-Quant , the pruning strategy, and the saliency-based low-rank adapters. 7.4.1 SLIM-Quant quantization method SLIM adopts symmetric weight quantization due to its low dequantization and memory overhead and ease of implementation. Denoting the quantized matrices byQsuperscript,Equation 7.3shows the symmetric quantization formula forq-bit quantization, whereαis the quantization scaling parameter andclip(.)operator clips the input to values between[−1,1]. W Q =round(clip( W α ))2 q−1 (7.3) The objective of quantization is to reduce the weight reconstruction error shown inEquation 7.4, where the∗superscript shows the optimal value. But the objective function inEquation 7.4is not convex, and to the best of our knowledge, does not have a closed form solution. α ∗ =argmin α ||W Q −W|| 2 =argmin α ||round(clip( W α ))2 q−1 −W|| 2 (7.4) To solve the mean squared error (MSE) problem inEquation 7.4, we propose a probabilistic reformu- lation as shown inEquation 7.5, whereQ(.)andQ −1 (.)are the quantization and dequantization functions respectively, andf(.)is the probability distribution function (PDF) of the weight elements. α ∗ =argmin α E Q =argmin α ||W Q −W|| 2 =argmin α Z ∞ −∞ f(x)|Q −1 (Q(x))−x| 2 dx(7.5) By incorporating the quantization formula from Equation 7.3intoEquation 7.5, we can simplify the inte- gration into the sum of two terms based on the absolute value of the data: the quantization error for absolute values less thanα(Equation 7.6) and the clipping error for absolute values larger thanα(Equation 7.7). Here, f abs (.)represents the probability density function (PDF) of the absolute value of the weights.Equation 7.8 presents the simplified version ofEquation 7.5. E quant (α) = Z α 0 f abs (x)|α×round( x α )×2 1−q −x| 2 dx(7.6) E clip (α) = Z ∞ α f abs (x)|α−x| 2 dx(7.7) α ∗ =argmin α E Q (α) =argmin α E quant (α) +E clip (α)(7.8) Equation7.8canbesolvedtheoreticallybydifferentiatingtheobjectivefunctionwithrespecttoα, provided the probability density function (PDF) of the weight distribution is known. However, the weight distribution of neural networks rarely conforms to standard PDFs. To verify this, we tested various candidate distributions, including Gaussian, Laplace, Pareto, q-Gaussian, and Weibull, as they are commonly used in modeling natural data. Unfortunately, none of these matched the observed weight distributions accurately. This discrepancy underscores the need for a more adaptable method, motivating the data-driven approach we adopt in SLIM- Quant . CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 87 Algorithm 5SLIM-Quant Algorithm 1Input:Weight magnitude PDFf abs , high resolution step sizeη high , low resolution step sizeη low , 2weight matrixW, quantization bitwidthq. 3Output:Quantized weight matrixW quant . 4 FunctionEstimateError(α) 5E quant (α) = R α 0 f abs (x)|α×round( x α )×2 1−q −x| 2 dx 6E clip (α) = R ∞ α f abs (x)|α−x| 2 dx 7returnE quant +E clip 8 end function 9E←EmptyDictionary()▷Initialize error dictionary 10for forαinrange(0, M,η low )do 11E(α)←EstimateError(α) 12end for 13α low ←argmin α E(α) 14for forαinrange(α low −η low ,α low +η low ,η high )do 15E(α)←EstimateError(α) 16end for 17α ∗ ←argmin α E(α) 18W quant ←round(clip( W α ∗ ))×2 q−1 ▷Apply optimal quantization 19Return:W quant . To address the absence of a closed-form weight PDF, we employ numerical integration on the weight histogram to solveEquation 7.8. To enhance efficiency, we adopt a multi-grid strategy: starting with 10 uniform samples in the range(0,max(W)), the grid is iteratively refined around the region of minimum error. This iterative process converges to the optimalαwith minimal computational overhead. The full procedure is detailed in Algorithm 5. 7.4.2 SLIM-LoRA low-rank adapters The use of a low-rank adapter to compensate for compression errors can be situated within the broader frame- work of sparse-plus-low-rank matrix decomposition, a problem studied extensively in the Robust PCA litera- ture [15,16]. In this classical formulation, a matrix is expressed as the sum of a sparse component and a low- rank component, with provable recovery guarantees under suitable incoherence conditions. Yu et al. [ 160] applied this decomposition paradigm to compress deep neural network weights, and recent works such as OATS [164] and HASSLE-free [91] have extended it to LLMs. SLIM adapts this principle to the joint com- pression setting: rather than solving a single convex program, we derive the low-rank component analytically from the compression error using a saliency-weighted SVD, which enables a one-shot solution without itera- tive optimization. We detail this procedure below. After quantizing the model using SLIM-Quant , we sparsify it using an off-the-shelf one-shot pruning method such as Wanda. The combined effects of quantization and pruning of a weight matrix can be modeled as additive noise, such thatW C =W+E Q +E S , whereE Q =W −W Q andE S =W C −W Q are the quantization and sparsity errors respectively. To mitigate these errors, we introduce low-rank adapters that adjust the compressed weights such thatW ≈ W C +LR, whereL ∈R d in ×r andR ∈R r×d out are the low-rank adapters andris the adapter rank. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 88 A straightforward approach minimizes the total error norm betweenWandW C , focusing solely on reduc- ing the error magnitude while ignoring the saliency of individual elements in the weight matrix. We call this methodNaive-LoRAas it overlooks the significance of individual elements in the weight matrix. However, this method is suboptimal and can be substantially improved. To address the limitations of Naive-LoRA, we propose a novel low-rank approximation formulation that integrates weight saliency and uses a carefully designed saliency function to determine optimal adapters. The saliency function (F) in our formulation needs to satisfy two key properties. First, it needs to be invertible, enabling the retrieval of low-rank adapters from their saliency. Second, it must be additive , meaning∀A, B: F(A+B) =F(A) +F(B). The additive property is crucial for isolating the saliency of low-rank adapters from the compressed matrix and distinguishing the saliency of the error from that of the original weights. These properties ensure that the saliency function can effectively isolate and optimize the contribution of low-rank adapters, forming the foundation of our proposed formulation. Assuming that there exists an additive invertible saliency functionF:R d in ×d out →R d in ×d out , we need to solve Equation 7.9to find the optimal adapters. By using the additive property of the saliency function F(.), we can simplifyEquation 7.9toEquation 7.10. L,R=argmax L,R ||F(W C +LR)|| 2 =argmin L,R ||F(W −(W C +LR))|| 2 (7.9) L,R=argmin L,R ||F(W −W C )−F(LR)|| 2 =argmin L,R ||F(−(E Q +E S ))−F(LR)|| 2 (7.10) Now, we can findF(LR)by computing the SVD ofF(−(E Q +E S )), and using the invertibility property ofF, we can obtain the exact value ofLandR. The saliency function used in SLIM must satisfy three essential criteria—invertibility, additivity, and the effectiveutilizationofinputandweightstatistics—tooptimizeweightimportanceduringcompression. Recent works such as Wanda, AWQ, LLM.int8(), and L 2 QER suggest that the product of the magnitude of the weights andactivationsisausefulmetricforidentifyingimportantweightsduringpruningandquantization. Motivated by this observation, we propose a saliency function formulation forFthat meets these criteria and leverages weight-activation interactions for effective compression. To incorporate input statistics into the saliency function, we defineF(W)≜diag(x)W, wherex∈R d in representstheaverageabsolutevalueofinputsfromacalibrationset. Thisformulationensuresthatthesaliency function effectively weights the matrix elements based on their significance during compression, facilitating a more accurate approximation. By replacingF(W)in Equation 7.10, the optimization problem transforms into a computationally efficient solution using singular value decomposition, followed by an inverse saliency transformation to derive the left low-rank adapter (Equation 7.12). L,R=argmin L,R ||−diag(x)(E Q +E S )−diag(x)LR|| 2 (7.11) diag(x)L,R=−SV D(diag(x)(E Q +E S ))(7.12) We refer to this method of computing saliency-based low-rank adapters asSLIM-LoRA, a practical and efficient approach tailored for addressing compression errors in large language models. To ensure numerical stability and guarantee the invertibility of the saliency function, an identity matrix with small values can be added todiag(x). This adjustment is equivalent to uniformly shifting all elements ofxand ensures that the saliency function remains robust even whenxcontains near-zero elements.Algorithm 6provides a compre- hensive overview of the steps involved in computing saliency-based low-rank adapters using SLIM-LoRA, CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 89 Algorithm 6SLIM-LoRA Saliency-based Low-rank Adapter Computation Input:Original weightW, compressed weightW C , calibration inputX. Output:Saliency-based low-rank adaptersL,R. 1E C ←E Q +E S =W C −W▷Compute error 2 ̃ x←mean(X)▷Average over all the samples 3x← ̃ x+min(| ̃ x|)▷Shift values to avoid zeros inx 4S C ←diag(x)E C ▷Compute error saliency 5 ̃ L, ̃ R←SVD(S C )▷Low-rank approximation 6L←diag(1/x) ̃ L▷Converting saliency to weight 7R← ̃ R Return:L,R. ensuring reproducibility and clarity. 7.4.3 Low-rank adapter quantization While pruning and quantizing the weights significantly reduce the model’s computation and memory require- ments (∼8×memory footprint reduction), incorporating full-precision low-rank adapters reintroduces over- head, partially offsetting these gains. To address this, we apply 4-bit quantization to compress the adapters. This step ensures that the compression efficiency achieved through weight pruning and quantization is pre- served, while maintaining the performance benefits of the low-rank adapters. Quantizing low-rank adapters poses unique challenges due to the long-tailed distribution of their elements, which limits the effectiveness of advanced non-group quantization methods, such as SLIM-Quant. To address this, we adopt an AbsMax group quantization scheme for the adapters, where groups of 128 elements share the same quantization parameter. By grouping elements, this method effectively captures the distribution’s variability while minimizing quantization error, striking a balance between accuracy and compression. This approach not only reduces the adapter overhead by4×but ensures that their contribution to overall model compression and performance is retained; as demonstrated in our experimental evaluation. 7.4.4 Optional Post-Compression Fine-Tuning Fine-tuning large language models post-compression has many challenges because the high parameter count and memory demands of traditional methods make them computationally prohibitive. For example, using a simple optimizer such as ADAMW leads to4×additional memory overhead to store gradient and optimizer states, rendering these approaches impractical for compressed models. Thus, parameter-efficient fine-tuning is essential for preserving the benefits of compression while avoiding excessive computational and memory costs. This necessity is further highlighted by the results inSection 7.5, which illustrate the overheads of traditional fine-tuning and the advantages of parameter-efficient alternatives. To overcome the challenges of fine-tuning compressed models, SLIM employs parameter-efficient low- rank adapters as the only tunable components during the fine-tuning phase. During this optional phase, SLIM freezes the sparse and quantized weights, enabling focused fine-tuning solely on the adapters. If the adapters are quantized, SLIM uses a straight-through estimator (STE) for quantization-aware fine-tuning and reduces its overheads with custom quantization and dequantization kernels implemented in Triton. This parameter-efficient fine-tuning method allows rapid accuracy improvements for the compressed model, requir- ing only a short fine-tuning phase over thousands of tokens. By limiting the fine-tuning process to a small CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 90 subset of parameters, SLIM significantly reduces computational requirements while ensuring the model can adapt effectively to new data or tasks. This approach maintains the benefits of compression while enabling efficient adaptation, as demonstrated by the significant improvements achieved during fine-tuning. 7.5 Experimental results Models,Datasets,andEvaluation.WeevaluateSLIMon theOPT [165]and LLaMA-2[144]model families, both of which serve as standard baselines in model compression studies [90,36,136]. Model accuracy is as- sessed on a range of zero-shot downstream tasks, including MMLU [57], Piqa [13], Arc-Easy, Arc-Challenge [21], WinoGrande [126], and OpenBookQA [97]. For zero-shot evaluations, we utilize the Language Model Evaluation Harness [41] framework. In line with prior work [136,36,90], we also report the perplexity of the models on a language modeling task on the WikiText2 [94] dataset, provided inSection E.4. Baselines.We compare SLIM against state-of-the-art one-shot pruning methods, including Wanda [ 136], SparseGPT [ 36], and Magnitude Pruning [52], as well as one-shot quantization techniques like OPTQ [37], OmniQuant[129], AffineQuant[90], L 2 QER[162], andAbsMax. Additionally, weextendJointSparsification and Quantization (JSQ) [49] to support 4-bit weight quantization and include it in our experiments. To ensure fairness, we use the optimal hyperparameters reported for each method, or the default hyperparameters if not explicitly reported. For a thoroughdescription of the notationsused to show the different variants of SLIM, please seeTable E.1inSection E.1. The hyperparameters used in our experiments are detailed below. Tothebestofourknowledge, L 2 QERistheonlycompressionmethodutilizingzero-shotlow-rankadapters to enhance model accuracy. Our approach, SLIM, significantly diverges from L 2 QER in several key aspects. First, we employ saliency-based low-rank adapters to mitigate compression loss inquantized and sparsemod- els, whereas L 2 QER is tailored exclusively for quantization, resulting in reduced accuracy when combined with sparsity, as demonstrated in the subsequent sections. Second, we introduce SLIM-Quant , which lowers the overhead and complexity of group quantization compared to methods like L 2 QER. Finally, SLIM com- presses and fine-tunes low-rank adapters efficiently to minimize overhead. In contrast, L 2 QER relies on full- precision low-rank adapters, which incur additional overhead and do not benefit from the parameter-efficient fine-tuning proposed in our work. Experiment Setup.Similar to Wanda, SparseGPT, and OPTQ, SLIM uses 128 sequences sampled from the C4 [ 121] dataset for calibration, and 300,000 tokens from C4 for all fine-tuning experiments. SLIM- Quant uses a histogram of weight elements to find the optimal scaling factor, with the number of bins set to max(512,min( d in ×d out 1000 ,20,000))to achieve an accurate approximation. All quantization experiments fol- low a 4-bit weight-only scheme with a group size of 128, consistent with prior work (OPTQ, OmniQuant, AffineQuant, etc.). For experiments involving Naive-LoRA and SLIM-LoRA, the adapter rank is set to 10% of the model’s hidden dimension unless stated otherwise. Fine-tuning is performed with the HuggingFace Trainer [ 153] using the AdaFactor [131] optimizer with linear learning rate scheduling and default parameters. We use BFloat-16 [148] on NVIDIA A100 GPUs, with a local batch size of 1 and gradient accumulation fac- tor of 64 to reduce memory overhead. Weight updates for sparse and/or quantized weights and corresponding biases are disabled during fine-tuning. Accuracyresults.Weevaluatetheaccuracyof SLIMandotherstate-of-the-artpruningandquantizationmeth- odsacross2:4andunstructuredsparsitybenchmarks,highlightingSLIM’ssuperiorityin Table7.1. SparseGPT and Group OPTQ, designed to work together, achieve competitive performance. For other advanced quanti- CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 91 zation methods, we pruned models using Wanda and quantized the sparse checkpoints with Group AbsMax, AWQ,OmniQuant, andAffineQuant, reportingthebestresults(detailedinSectionE.5). Notably, methodslike OmniQuant and AffineQuant struggle to quantize OPT-350M, often resulting in NaN values. Moreover, AWQ, OmniQuant, AffineQuant, and L 2 QER encounter out-of-memory (OOM) errors when compressing models on a single A100-40GB GPU. While JSQ performs well for the LLaMA-2 family, its difficulty compressing the OPT family limits its broader applicability. Table 7.1: Average zero-shot accuracy of LLaMA-2 and OPT models with50% sparsity and 4-bit weight quantization.Best Method ∗ indicatesthebestquantizationmethodoutofGroupAbsMax, AWQ,OmniQuant, and AffineQuant.↑indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 2:4 Sparsity MagnitudeGroup AbsMax 32.19 31.94 33.82 33.43 34.81 34.68 44.64 44.18 SparseGPTGroup OPTQ 33.70 33.38 38.75 40.15 44.32 45.64 45.49 51.05 WandaBest Method ∗ 33.39 32.79 38.43 40.00 43.41 44.07 44.86 48.94 JSQJSQ31.98 31.13 36.34 31.79 41.33 37.38 45.34 49.45 L 2 QERGroup AbsMax 33.34 31.68 36.68 38.11 41.37 OOM 43.77 OOM Naive-LoRASLIM-Quant 34.28 33.38 38.36 41.21 44.91 45.25 48.45 51.94 SLIM-LoRASLIM-Quant34.62 34.36 40.61 42.7345.99 46.0951.15 54.94 SLIM-LoRA Q SLIM-Quant 34.43 34.30 40.11 42.3746.33 46.2451.02 53.55 50% Unstructured MagnitudeGroup AbsMax 33.34 33.51 32.12 39.90 36.44 32.33 47.03 51.04 SparseGPTOPTQ35.10 35.13 38.72 43.43 46.97 47.38 51.09 55.94 WandaBest Method ∗ 35.11 33.89 41.02 42.89 46.52 46.84 53.62 56.76 JSQJSQ32.05 31.09 39.53 33.35 41.04 31.80 52.08 57.00 L 2 QERGroup AbsMax 34.45 34.45 38.38 41.28 45.08 OOM 50.60 OOM Naive-LoRASLIM-Quant 34.77 34.23 40.40 43.37 46.64 47.30 51.52 55.33 SLIM-LoRASLIM-Quant 35.2035.32 41.8543.48 47.0847.96 54.26 57.85 SLIM-LoRA Q SLIM-Quant35.3535.13 41.7443.63 47.1647.86 54.18 57.33 The progression from Naive-LoRA to SLIM-LoRA and SLIM-LoRA Q demonstrates the benefits of in- corporating weight saliency into low-rank adapters and applying quantization for reducing overhead. While Naive-LoRA improves model accuracy across different sizes, SLIM-LoRA achieves additional gains by ef- fectively leveraging the saliency of the weights in the adapter design. Extending this, SLIM-LoRA Q applies quantization to the low-rank adapters, further minimizing overhead with minimal impact on accuracy, adding negligible improvements or degradation to the accuracy of the model. Table 7.2demonstrates how lightweight fine-tuning (FT) improves the accuracy of both SLIM-LoRA and Naive-LoRA, with SLIM-LoRA exhibiting greater gains due to its saliency-aware design. Further details on the fine-tuning process and its overhead are provided in Section E.8, illustrating its practicality for enhancing compressed model performance. Integration with Weight Update.Chapter 5proposes a method to compute the optimal per-layer weight updates given a calibration dataset. After determining the low-rank adapter values in SLIM, we can find the optimal weight values by solvingW C∗ =argmin W C ∥XW C − X(W −LR)∥using the QP solvers CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 92 Table 7.2: Effects of fine-tuning on the average zero-shot accuracy of LLaMA-2 models with 50% sparsity and 4-bit weight quantization.↑indicates better performance. Pruning/LoRAWeightLLaMA-2 MethodQuantization 7B 13B Dense-56.6 60.8 50% 2:4 Naive-LoRA + FT SLIM-Quant 50.89 55.70 SLIM-LoRA + FT SLIM-Quant52.12 56.60 SLIM-LoRA Q + FT SLIM-Quant 48.31 56.50 50% Unstructured Naive-LoRA + FT SLIM-Quant 52.90 57.08 SLIM-LoRA + FT SLIM-Quant54.69 57.96 SLIM-LoRA Q + FT SLIM-Quant 53.57 57.78 Table 7.3: Accuracy results of OPTIMA weight update mechanism with SLIM-LoRA.↑indicates better per- formance. Pruning/LoRAWeightLLaMA-2 MethodQuantization 7B 13B Dense-56.6 60.8 50% 2:4 SLIM-LoRA Q SLIM-Quant 51.02 53.55 SLIM-LoRA Q + OPTIMA SLIM-Quant 51.62 53.84 50% Unstructured SLIM-LoRA Q SLIM-Quant 54.18 57.33 SLIM-LoRA Q + OPTIMA SLIM-Quant 54.32 57.45 introduced in OPTIMA. As shown inTable 7.3, our results indicate that applying the compression trinity with OPTIMA as the weight update and SLIM-LoRA as the low-rank adapters can further boost the accuracy of the models. IntegrationwithHybridSparsity.Asdemonstratedin Chapter6,PATCHimprovesuponrigidsemi-structured sparsity by enabling adaptive, tile-level density. While the standard implementation of SLIM presented in this chapter utilizes Wanda [136] for efficient pruning, the modular design of the Compression Trinity allows us to substitute this component with more advanced sparsity operators. In this section, we integrate PATCH into the SLIM pipeline to evaluate the impact of hybrid sparsity on the fully compressed model. We apply the SLIM-LoRA error correction mechanism on top of the hybrid masks generated by PATCH. For these specific experiments, we utilize Group AbsMax quantization to isolate the benefits of the hybrid spar- sity pattern when combined with standard quantization schemes. Table 7.4reports the results on LLaMA-2 7B and LLaMA-3.1 8B. The results demonstrate that combining the flexible sparsity of PATCH with quantiza- tion and low-rank approximation enables controllable tradeoffs between compression ratio and model quality. This confirms that the enhancements made to the sparsity pillar in the previous chapter translate directly to improved flexibility in the joint compression setting. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 93 Table 7.4: Average accuracy (↑indicates better) across eight zero-shot downstream tasks (including RACE [74] and HellaSwag [161]) and WikiText2 perplexity (↓indicates better) of compressed models with4-bit weight-only quantization. Please note that using LoRA adds additional parameters to the model. Sparsity Method Pattern LoRA LLaMA-2-7BLLaMA-3.1-8B Acc (%↑) PPL (↓) Acc (%↑) PPL (↓) 0%Dense--54.615.1260.315.84 50% MaskLLM 2:4-47.987.6451.129.92 45% PATCH Tile Dense/2:4 Tiles -48.197.3452.479.68 45% PATCH Tile Dense/2:4 Tiles SLIM-LoRA 50.716.8354.049.12 35% PATCH Tile Dense/2:4 Tiles -49.386.9253.819.26 35% PATCH Tile Dense/2:4 Tiles SLIM-LoRA 51.916.4255.708.37 25% PATCH Tile Dense/2:4 Tiles -50.456.5755.458.69 25% PATCH Tile Dense/2:4 Tiles SLIM-LoRA 52.626.1156.997.77 Comparisonoflargecompressedandsmalldensemodels.This section compares large compressed models with dense models of equivalentparameter size, offering guidelines forconfiguration selection under hardware constraints. We focus on 2:4 sparsity due to its hardware acceleration support and evaluate the OPT model family, which spans a wide range of sizes for comprehensive analysis. We analyze model performance by plotting average accuracy against parameter size, calculated as detailed inSection E.9. This visualization enables a direct performance comparison between models with an equal number of bits. Figure 7.2presents the accuracy results of the OPT model family across different compression methods. The x-axis represents the model parameter size in gigabytes, while the y-axis denotes accuracy (higher is bet- ter). The results demonstrate that SLIM-LoRA Q , both with and without fine-tuning, consistently outperforms dense models and other compression techniques at the same parameter size. Notably, compressed models achieve higher accuracy than dense models of equivalent size, highlighting the effectiveness of the proposed method. This trend underscores the advantage of SLIM-LoRA Q in maximizing model efficiency under strict hardware constraints. Speedup.Leveraging sparsity and quantization enhances GPU resource utilization, enabling faster model inference. Following Wanda’s experimental setup, we evaluate the speedup achieved across different model layers and sizes. Similar to Wanda, AWQ, and QuaRot [ 7], we focus on consumer-grade GPUs and conduct our experiments on NVIDIA RTX 3060 GPUs. Speedup results for NVIDIA A100 GPUs are provided in Section E.7. SLIM achieves notable speedups through optimized sparse and quantized matrix multiplication, utilizing Sparse Marlin [38] integrated with vLLM [73]. For inference, we adopt small batch sizes during decoding, as recommended by prior works [ 154,167]. Dense Quantized Marlin or PyTorch kernels handle the low- rank adapters based on their quantization status.Table 7.7highlights the speedup achieved across different LLaMA-2 layers compared to dense, unquantized models. Larger matrices, such as those in self-attention and feed-forward modules, consistently yield greater speedups, aligning with trends detailed in Section E.7. Sparse-only results.To evaluate the isolated impact of sparsity on model accuracy, we disable quantization and benchmark Magnitude Pruning, SparseGPT, and Wanda, alongside low-rank approximations like Wanda- SVD and SLIM . Our experiments assess both 50% unstructured sparsity and 2:4 structured sparsity patterns. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 94 10 −1 10 0 10 1 Model Parameter Size (GB) 30 35 40 45 Accuracy Accuracy vs. Model Parameter Size (2:4 Sparsity) Dense Model SLiM-LoRA Q + SLiM-Quant SLiM-LoRA Q + SLiM-Quant (FT) WANDA + Best Quantization SparseGPT + OPTQ Figure 7.2: Accuracy results of the OPT family across different compression methods (↑indicates better per- formance). At equal parameter size, SLIM outperforms both dense models and other compression techniques, demonstrating that model compression with SLIM yields superior performance under the same budget. Table 7.5shows the accuracy results for sparse models. Magnitude Pruning performs the worst, while Wanda and SparseGPT achieve comparable results, with larger accuracy gaps for semi-structured sparsity. Low-rank adapters improve accuracy, with SLIM leveraging saliency-based approximation for superior per- formance. A brief fine-tuning phase further boosts the accuracy of low-rank approximations. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 95 Table 7.5: Average zero-shot accuracy of LLaMA-2 and OPT models with pruning. The quantization is disabled in this experiment.↑indicates better performance. Pruning/LoRAOPTLLaMA-2 Method125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 2:4 Sparsity Magnitude32.6 31.8 35.4 33.9 36.4 30.7 31.2 32.0 SparseGPT33.8 33.2 37.7 41.3 45.2 45.6 47.3 52.3 Wanda34.0 32.5 38.3 40.5 43.2 44.1 46.1 49.7 SLIM-Naive34.1 34.1 40.4 42.8 46.0 45.9 51.6 55.8 SLIM-Naive + FT 34.8 34.5 41.3 43.446.547.252.4 56.9 SLIM-LoRA34.5 32.9 40.7 43.1 46.4 46.3 51.4 56.1 SLIM-LoRA + FT35.1 34.9 41.5 43.8 46.5 47.351.6 56.4 50% Unstructured Magnitude33.3 33.7 34.0 40.6 35.8 30.9 32.6 31.9 SparseGPT35.5 35.1 39.6 43.5 47.4 47.8 53.3 57.3 Wanda35.0 34.5 41.1 42.9 46.5 46.8 52.7 57.2 SLIM-Naive35.3 35.2 41.9 44.1 47.5 47.8 54.9 58.5 SLIM-Naive + FT 35.74 35.7 42.7 44.647.8 48.454.9 58.7 SLIM-LoRA35.2 35.1 42.0 44.1 47.7 48.255.0 58.8 SLIM-LoRA + FT35.9 35.7 42.5 44.747.748.4 55.0 58.8 Quantization-only results.To evaluate the impact of SLIM-Quant and low-rank compensation in SLIM, we conduct experiments without sparsity, testing quantization schemes like Group AbsMax, OPTQ, AWQ, OmniQuant, AffineQuant, L 2 QER, and SLIM-Quant . To enhance accuracy, we add low-rank adapters to SLIM-Quant and Group AbsMax, optimizing either error saliency (SLIM-LoRA) or reconstruction error norm (Naive-LoRA). Other quantization methods cannot incorporate low-rank adapters due to conflicting weight/ac- tivation update rules. Table 7.6presents the quantization results. Adding low-rank adapters to Group AbsMax significantly boosts model accuracy, outperforming most advanced methods. While SLIM-Quant alone is not designed for high accuracy, its integration with SLIM variants achieves results comparable to or better than Group AbsMax with low-rank adapters, highlighting the value of co-design in compression methods. Furthermore, a lightweight fine-tuning phase with SLIM-Quant delivers state-of-the-art accuracy. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 96 Table 7.6: Average zero-shot accuracy of LLaMA-2 and OPT models with quantization. The sparsity is disabled in this experiment.↑indicates better performance. QuantizationLow-rankOPTLLaMA-2 MethodAdapter125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 OPTQ-35.64 36.46 42.83 44.20 47.46 48.24 53.53 59.80 AWQ-36.16 31.83 42.98 45.28 48.45 48.76 53.97 OOM OmniQuant-35.46 NaN 42.15 44.71 46.65 OOM 54.33 OOM AffineQuant-35.73 NaN 42.62 44.92 47.91 OOM 54.52 OOM Group AbsMax -35.45 36.67 42.57 44.79 48.30 48.49 55.56 60.12 Group AbsMax L 2 QER34.75 35.63 40.60 44.22 46.90 OOM 55.95 OOM Group AbsMax SLIM-Naive36.3036.58 43.07 45.13 48.26 48.72 56.23 60.53 Group AbsMax SLIM-LoRA36.1836.7242.8945.65 48.4548.89 55.99 60.16 SLIM-Quant -31.98 36.46 36.19 40.08 45.61 38.27 31.11 30.51 SLIM-Quant SLIM-Naive35.29 36.02 42.48 45.01 47.75 48.38 55.9660.85 SLIM-Quant SLIM-LoRA35.69 36.42 42.59 45.26 48.18 48.52 56.26 60.59 SLIM-Quant SLIM-LoRA + FT 35.91 36.6143.2945.58 48.2949.04 56.5160.65 AdditionalExperiments.We provide additional experiments for a comprehensive evaluation in the appendix. TheLanguage Modeling Experiments (Section E.4) evaluate SLIM across sparse and quantized, sparse- only, and quantized-only models on WikiText-2. The results align with the accuracy trends reported in the preceding sections, further validating the effectiveness of SLIM . TheFine-tuning Costs (Section E.8) show that SLIM reduces fine-tuning overhead from over 36 days for 13B parameter models to just 14 hours on a single GPU, demonstrating its practicality and efficiency. We provide a comparison betweenSparsity vs.Quantization(Section E.6) to show that combining 50% sparsity and 4-bit quantization helps achieve better compression results in comparison to solely using 2-bit quantization, while maintaining a similar compression ratio (∼8×). Additional speedup results for SLIM on NVIDIA A100-40GB GPUs are provided in theAdditional Speedup Results(Section E.7). A theoretical analysis of computation and memory reductions can be found in theComputation Reduction Analysis(Section E.10) andMemory Reduction Analysis(Section E.9), high- lighting the efficiency of SLIM . Compression Costs(Section E.11) details the time required to compress models of various sizes across different methods.Rank Analysis (Section E.12) explores how rank choices in low-rank adapters impact computational and memory costs, as well as model accuracy.Sparsity Analysis(Section E.15) analyzes the effects of different sparsity ratios on model compression. Lastly,Effects of Calibration Sample Count(Sec- tion E.13 ) evaluates the influence of calibration sample counts on the accuracy of calibration-based methods. CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 97 Table 7.7: LLaMA-2 family of models speedup (×) using SLIM compared to original dense unquantized model on NVIDIA RTX-3060.↑shows higher speedup. Model Batch LoRA Self-Up-Down Size Size Type Attention Projection Projection 7B 16 FP16 1.992.583.23 INT4 1.002.292.58 32 FP16 2.062.533.40 INT4 1.132.272.97 64 FP16 1.541.701.85 INT4 0.961.621.76 13B 16 FP16 2.182.532.60 INT4 1.283.243.17 32 FP16 2.232.682.91 INT4 1.432.963.20 64 FP16 1.381.781.67 INT4 1.211.691.65 70B 16 FP16 2.182.862.75 INT4 3.113.993.79 32 FP16 2.002.632.67 INT4 2.753.193.39 64 FP16 1.381.701.86 INT4 1.511.771.94 7.6 Conclusion In this chapter, we presented SLIM, the complete fulfillment of the Compression Trinity framework. We be- gan this thesis by identifying the ”Sparsity Paradox,” the observation that removing weights is structurally destructive and, when applied in isolation, leads to early accuracy collapse. SLIM resolves this paradox. By seamlessly integrating optimized uniform quantization (SLIM-Quant), hardware-friendly sparsity, and cru- cially mathematically derived low-rank error correction (SLIM-LoRA), we have turned the Low-Rank pillar into a restorative force. It recovers the information lost by the aggressive application of the first two pillars, solving the ”compounded error” challenge that has historically hindered joint compression. An important direction for strengthening the theoretical foundations of SLIM is to draw more explicitly on the guarantees provided by Robust PCA theory. The classical PCP framework [ 15] establishes that the sparse-plus-low-rank decomposition is exactly recoverable under incoherence conditions, and the Stable PCP extension [169] shows that this recovery is robust to additional dense noise, a property that could be leveraged to account for quantization error. Translating these guarantees to the LLM compression setting would re- quire verifying whether the incoherence assumptions hold for pre-trained weight matrices and characterizing the interaction between quantization noise, sparsity patterns, and low-rank structure. Furthermore, replac- ing SLIM’s current sequential pipeline with a joint optimization formulation, as explored by OATS [ 164], HASSLE-free [91], and 3BASiL [11], could yield tighter error bounds and potentially improve accuracy by avoiding the suboptimality inherent in sequential decomposition. Such a formulation would need to in- corporate the quantization constraint, extending the standard sparse-plus-low-rank problem to a “quantized sparse-plus-low-rank” decomposition, an open problem that merits further investigation. SLIM not only shifts the Pareto frontier, outperforming dense models at equal parameter budgets, but also CHAPTER 7. SLIM: ONE-SHOT QUANTIZATION AND SPARSITY WITH LOW-RANK APPROXIMATION FOR LLM WEIGHT COMPRESSION 98 serves as the final proof of our core thesis: that efficiency is not a singular optimization problem, but a multi- dimensional balancing act. We have traversed the complete arc of this life-cycle: from accelerating pretraining dynamics (Chapter 3,Chapter 4) to establishing the limits of sparsity in both static (Chapter 5) and dynamic (Chapter 6) regimes, and finally achieving a unified, one-shot solution for deployment (Chapter 7). The next and final chapter will summarize these contributions, discuss the broader implications of the Compression Trinity for the future of efficient AI, and outline potential avenues for further research. Chapter 8 Conclusion and Future Work 8.1 Summary of Contributions This thesis has argued that the efficiency bottleneck in Large Language Models (LLMs) is not merely a re- source constraint, but a methodological failure to integrate complementary compression principles. We have established that the “efficiency wall” encountered by isolated techniques ismethodological rather than fun- damental. By jointly applying the “Compression Trinity,” sparsity, quantization, and low-rank approxima- tions, we have demonstrated that the distinct hardware bottlenecks of compute FLOPs, memory bandwidth, and parameter redundancy can be attacked simultaneously. Crucially, we established that the Trinity is not merely a post-training optimization tool; it is a fundamental framework applicable to the entire life-cycle of the model. 8.1.1 The Trinity in Training Dynamics We demonstrated that the training dynamics themselves are compressible. By selectively applying pillars of the Trinity during optimization, we proved that high-fidelity, dense updates are not strictly necessary for convergence. •Sparse, Low-Rank, and Quantized Optimization (MKOR):InChapter 3, we addressed the prohibitive computational penalty of second-order optimization. MKOR validates the Trinity’s utility in training by combining block diagonalSparsitywithLow-Rankapproximations (rank-1 updates) andQuantization (to stabilize memory footprint). This approach reduces curvature update complexity fromO(d 3 )toO(d 2 ), accelerating convergence by up to2.57×compared to first-order baselines and1.75×compared to KFAC. •Sparse and Low-Rank Training (SLOPE):InChapter 4, we validated the concept of “Lossy Training” by integratingSparsityandLow-Rankprinciples. By enforcing N:M sparsity in the backward pass and delaying low-rank recovery (“lazy” adapters) to the final 1% of training, we showed that the training process can withstand significant information loss without sacrificing final model accuracy. 8.1.2 The Trinity in Post-Training and Inference For deployed models, we systematically dismantled the primary barriers to compression, compounded error and structural rigidity, resulting in a unified inference framework. 99 CHAPTER 8. CONCLUSION AND FUTURE WORK100 •Solving Compounded Error (OPTIMA):OPTIMA (Chapter 5) utilized global optimization to stabi- lize theQuantizationpillar. By establishing that weight reconstruction must be solved via column-wise Quadratic Programs (QPs) with a shared Hessian, OPTIMA provides the foundational stability necessary to withstand aggressive compression. •Breaking Rigidity (PATCH):PATCH (Chapter 6) advanced theSparsitypillar by breaking the rigidity of hardware-enforced patterns. By introducing learnable tile-level hybrid sparsity, we proved that sparsity ra- tios can be continuous and adaptive (0–50%) rather than discrete, preserving density in information-critical layers. •Unified Inference (SLIM):The empirical validation of the full framework culminates in SLIM (Chapter 7). SLIM represents the simultaneous integration of the full Trinity, aggressiveQuantization, semi-structured Sparsity, and saliency-basedLow-Rankadapters, during the inference phase. It shifts the Pareto frontier of model efficiency, demonstrating that a fully compressed model can improve accuracy by up to 5.66% over state-of-the-art methods and, in specific configurations, outperform uncompressed dense models at equivalent parameter budgets. The industrial significance of these findings, particularly the necessity of 2:4 sparsity in modern production stacks, was subsequently featured in our technical analysis on the PyTorch blog 1 . 8.2 Exploratory Frameworks and Open Research The core pillars of the Compression Trinity, sparsity, quantization, and low-rank approximations, provide a rigorous foundation for model efficiency. However, the application of these principles need not be confined to therigidstructuresoftraditionalacademicpublicationcycles. Inparallelwiththeformalchaptersofthisthesis, we have developed agile, exploratory frameworks that extend the Trinity into new domains of adaptability and rapid deployment. These works, released as open-research contributions, demonstrate the flexibility of our methodology in addressing emerging challenges in the LLM landscape. 8.2.1 BEAM: Blockwise Error Minimization for One-shot Compression of LLMs Standard post-training compression techniques (e.g., GPTQ, Wanda) often hit an accuracy ceiling because they optimize weights locally (layer-wise) without accounting for the non-linear interactions across the full transformer block. Conversely, full model fine-tuning (e.g., LoRA) is computationally expensive and requires curated datasets. To bridge this gap, we introducedBEAM 2 , a framework for one-shot compression that requires no end- to-end retraining. BEAM re-frames the compression problem by treating the intermediate activations of the original, uncompressed model as the “ground truth” for the compressed model. By splitting the LLM into independent transformer blocks and optimizing the compressed weights to minimize the feature reconstruc- tion error of each block, BEAM captures the non-linear dependencies lost by simple layer-wise techniques. Crucially, this method is orthogonal to the specific compression type; it serves as a universal refinement stage that can recover up to 4.34% accuracy on sparse and quantized models using a single GPU in under four hours. 1 https://pytorch.org/blog/when-quantization-isnt-enough-why-24-sparsity-matters/ 2 Full release:https://w.cs.toronto.edu/~mmozaffari/compression-trinity/beam/index.html CHAPTER 8. CONCLUSION AND FUTURE WORK101 8.2.2 LEAP: Learnable End-to-End Adaptive Pruning of LLMs InChapter 6, we explored hybrid structural sparsity to balance hardware efficiency with information retention. However, imposinganystructure (even a hybrid one) inherently limits the model’s expressivity compared to unstructured pruning. LEAP 3 challenges the assumption that unstructured sparsity must be static or heuristically determined (e.g., magnitude-based pruning). Instead, LEAP introduces a fully differentiable masking mechanism where the binary inclusion of every parameter is treated as a learnable latent variable during training. By relaxing the discrete mask into a continuous probability distribution and applying straight-through estimation, LEAP allows the model to dynamically “evolve” its own sparsity pattern end-to-end. This approach reveals that optimal sparsity is not fixed; it shifts during training as the model specializes, suggesting that the “Trinity” can eventually includetopologyas a learnable parameter alongside weights. 8.2.3 SLICE: Selecting Layer-wise Configurations for Matryoshka-Style LLMs The prevailing paradigm in LLM deployment is “one size fits all”: a model is compressed to a fixed target (e.g., 4-bit, 50% sparsity) and deployed. This rigidity is inefficient for dynamic environments where hardware availability changes in real-time. SLICE 4 extends the concept of Matryoshka Representation Learning to the compression configuration itself. SLICE trains a single “super-model” capable of operating at multiple efficiency tiers simultaneously. By solving for a nested set of configurations (e.g., a 2-bit core nested within a 4-bit shell), SLICE enables elastic deployment: the same model can instantaneously shed layers or precision bits to meet strict latency deadlines, or expand to full capacity when resources permit. This work points toward a future where the Compression Trinity is not a static compilation step, but a dynamic runtime state. 8.3 Limitations of the Current Approach While the Compression Trinity offers a robust framework, the specific implementations proposed in this thesis are subject to constraints that affect their immediate generalizability and ease of adoption. Hardware Coupling and Portability.Our methods, particularly PATCH and the semi-structured sparsity utilized in SLOPE and SLIM, are currently tightly coupled to the NVIDIA Sparse Tensor Core architecture (2:4 sparsity). While effective on dominant hardware, the transferability of these specific patterns to non-NVIDIA accelerators (e.g., TPUs, AMD MI-series) or general-purpose CPUs remains unproven. Furthermore, the reliance on custom CUDA and Triton kernels introduces a significant software portability barrier, effectively restricting these optimizations to advanced engineering environments. Training Instability Risks.Although MKOR introduces stabilizers to mitigate exploding gradients, second- order optimizers remain inherently more sensitive to hyperparameters than robust first-order methods like AdamW.Theintroductionof additionalhyperparametersforcurvatureapproximationimposesatuning burden on practitioners, potentially offsetting the wall-clock speed gains in experimental settings. Pipeline Complexity.The “Compression Trinity” introduces significant engineering overhead. Deploying a pipeline that requires simultaneous quantization, pruning, and low-rank adaptation (as in SLIM) is consid- erably more complex to implement, debug, and maintain than simpler post-training quantization techniques 3 Full release:https://w.cs.toronto.edu/~mmozaffari/compression-trinity/leap/index.html 4 Full release:https://w.cs.toronto.edu/~mmozaffari/compression-trinity/slice/index.html CHAPTER 8. CONCLUSION AND FUTURE WORK102 (e.g., INT8/FP8). This complexity represents a barrier to adoption for practitioners seeking “plug-and-play” solutions. Fairness and Differential Impact on Underrepresented Populations.Throughout this thesis, compression quality is measured by aggregate accuracy on standard benchmarks, yet matching the uncompressed model’s overall accuracy does not guarantee uniform performance across all subpopulations. Pruning, quantization, and low-rank factorization remove model capacity that may disproportionately encode knowledge about un- derrepresented groups, a phenomenon that aggregate metrics can mask entirely. For instance, a compressed language model may preserve perplexity on high-resource languages such as English while exhibiting signif- icant degradation on low-resource languages that were sparsely represented in the fine-tuning or calibration datasets. Similarly, in classification tasks, accuracy on minority demographic groups may suffer even when the overall metric remains stable. Because our calibration and evaluation pipelines rely on datasets that pre- dominantly reflect majority populations, we cannot rule out that the proposed compression methods amplify existing biases or introduce new disparities. A rigorous fairness audit, disaggregating performance across languages, dialects, demographic groups, and downstream tasks, is an important direction that lies outside the scope of this work but is essential before deploying compressed models in high-stakes applications. Scale of Evaluation.Our empirical validation focused on models in the 125M to 70B parameter range (e.g., OPT, LLaMA-2/3). As scaling laws push frontier models into the trillion-parameter regime, emergent behav- iors or shifting bottlenecks (e.g., massive cross-node communication overheads) may alter the effectiveness of these compression techniques, a domain this thesis leaves unexplored. 8.4 Future Research Directions Compressing the Context Window.This thesis focused heavily on Linear layers, which currently dominate compute. However, as sequence lengths grow to 1M+ tokens, the Attention mechanism and Key-Value (KV) cache become the dominant memory bottlenecks. Future work must extend the Trinity to the context window: quantizing dynamic KV states, inducing sparsity in attention patterns (e.g., via Sliding Window or Block- Sparse Attention), and applying low-rank approximations to the attention heads themselves to enable infinite- context reasoning on commodity hardware. Hardware-Algorithm Co-design.Current sparse formats incur a memory overhead for storing metadata (indices), which can negate compression gains at lower bit-widths. Future compression research cannot occur in a software vacuum; it requires a co-design approach to propose “hardware-defining” sparse formats. We envision a move toward algorithmic sparsity where the pattern is deterministic or predicted, eliminating the need for explicit index storage and further reducing the memory footprint. The Trinity for Activations.While we successfully compressed weights, activation tensors remain a chal- lenge during the prefill phase. Future work should investigate applying PATCH-like learnable masks to dy- namic activation tensors, enabling “activation sparsity” that can accelerate the compute-bound prefill phase without requiring full retraining. 8.5 Closing Remarks Ultimately, this thesis posits that efficiency is not merely a post-hoc optimization, but a fundamental design constraint. We have argued that the “efficiency wall” is a methodological artifact, dissolvable by the joint CHAPTER 8. CONCLUSION AND FUTURE WORK103 application of sparsity, quantization, and low-rank approximations. By proving that the “Compression Trinity” can be integrated into every stage of the LLM life-cycle, from the training dynamics of MKOR to the inference engines of SLIM, we pave the way for a new paradigm of model design. In this paradigm, models are not simply made smaller; they are architected from the ground up to be dense in knowledge yet sparse in computation, rendering high-intelligence AI fundamentally more accessible and sustainable. Bibliography [1]Abhinav Agarwalla, Abhay Gupta, Alexandre Marques, Shubhra Pandit, et al. Enabling High- Sparsity Foundational LLaMA Models with Efficient Pretraining and Deployment.arXiv preprint arXiv:2405.03594, 2024. [2]Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. Intrinsic Dimensionality Explains the Effec- tiveness of Language Model Fine-Tuning. InProceedings of the 59th Annual Meeting of the Associa- tion for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7319–7328, Online, August 2021. Association for Com- putational Linguistics. [3]Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, et al. QSGD: Randomized Quantization for Communication-Efficient Stochastic Gradient Descent. InNeurIPS, 2017. [4]Shun-Ichi Amari. Natural Gradient Works Efficiently in Learning.Neural Computation, 10(2):251– 276, 1998. [5]Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, et al. Scalable Second Order Optimization for Deep Learning.arXiv preprint arXiv:2002.09018, 2020. [6]Argonne Leadership Computing Facility. Polaris.https://w.alcf.anl.gov/polaris. [7]Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, et al. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs. InNeurIPS, 2024. [8]NihatAy. OntheLocalityoftheNaturalGradientforLearninginDeepBayesianNetworks.Information Geometry, pages 1–49, 2020. [9]Jimmy Ba, Roger Grosse, and James Martens. Distributed Second-Order Optimization Using Kronecker-Factored Approximations. InICLR, 2017. [10]Abhimanyu Rajeshkumar Bambhaniya, Amir Yazdanbakhsh, Suvinay Subramanian, Sheng-Chun Kao, et al. Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers.arXiv preprint arXiv:2402.04744, 2024. [11]Kayhan Behdin, Mehdi Makni, and Rahul Mazumder. 3BASiL: An algorithmic framework for sparse plus low-rank compression of LLMs.arXiv preprint arXiv:2603.01376, 2025. [12]Dimitris Bertsimas, Ryan Cory-Wright, and Nicholas AG Johnson. Sparse Plus Low Rank Matrix Decomposition: A Discrete Optimization Approach.JMLR, 2023. 104 BIBLIOGRAPHY105 [13]Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. PIQA: Reasoning About Physical Com- monsense in Natural Language. InAAAI, 2020. [14]Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al. Language Models Are Few-Shot Learners.Advances in Neural Information Processing Systems, 33:1877–1901, 2020. [15]Emmanuel J. Candès, Xiaodong Li, Yi Ma, and John Wright. Robust Principal Component Analysis? Journal of the ACM, 58(3):1–37, 2011. [16]Venkat Chandrasekaran, Sujay Sanghavi, Pablo A. Parrilo, and Alan S. Willsky. Rank-Sparsity Inco- herence for Matrix Decomposition.SIAM Journal on Optimization, 21(2):572–596, 2011. [17]Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, et al. Scatterbrain: Unifying Sparse and Low-Rank Attention Approximation.arXiv preprint arXiv:2110.15343, 2021. [18]Stanley F. Chen, Douglas Beeferman, and Roni Rosenfeld. Evaluation Metrics for Language Models. Technical report, Carnegie Mellon University, 1998. [19]Stanley F Chen, Douglas Beeferman, and Roni Rosenfeld. Evaluation Metrics for Language Models. Carnegie Mellon University, 1998. [20]Zhaodong Chen, Zheng Qu, Yuying Quan, Liu Liu, et al. Dynamic N:M Fine-Grained Structured Sparse Attention Mechanism. InPPoPP, 2023. [21]Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, et al. Think You Have Solved Question Answer- ing? Try ARC, the AI2 Reasoning Challenge.arXiv preprint arXiv:1803.05457, 2018. [22]Compute Canada. Compute Canada.https://computecanada.ca/. [23]NVIDIA Corporation. CUTLASS 4.2.0: CUDA Templates for Linear Algebra Subroutines.https: //github.com/NVIDIA/cutlass, 2025. Also see: Kerr, A., Merrill, D., Demouth, J., Tran, J. ”CUTLASS: Fast Linear Algebra in CUDA C++”, NVIDIA blog, Dec. 2017. [24]Tri Dao. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. InICLR, 2024. [25]Tri Dao, Beidi Chen, Kaizhao Liang, Jiaming Yang, et al. Pixelated Butterfly: Simple and Efficient Sparse Training for Neural Network Models. InICLR, 2022. [26]Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InNeurIPS, 2022. [27]Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. GPT3.int8(): 8-Bit Matrix Multi- plication for Transformers at Scale. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,NeurIPS, pages 30318–30332. Curran Associates, Inc., 2022. [28]Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. QLoRA: Efficient Finetuning of Quantized LLMs. InNeurIPS, 2023. [29]Tim Dettmers and Luke Zettlemoyer. Sparse Networks from Scratch: Faster Training without Losing Performance.arXiv preprint arXiv:1907.04840, 2019. BIBLIOGRAPHY106 [30]Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. InNAACL, pages 4171–4186, 2019. [31]Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, et al. The LLaMA 3 Herd of Models.arXiv preprint arXiv:2407.21783, 2024. [32]Ruibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li, et al. SpInfer: Leveraging Low-Level Sparsity for Effi- cient Large Language Model Inference on GPUs. InProceedings of the Twentieth European Conference on Computer Systems, pages 243–260, 2025. [33]Gongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich, et al. MaskLLM: Learnable Semi- Structured Sparsity for Large Language Models. InNeurIPS, 2024. [34]Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. Linear Mode Connec- tivity and the Lottery Ticket Hypothesis. InICML, 2020. [35]Elias Frantar and Dan Alistarh. Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning.NeurIPS, 35:4475–4488, 2022. [36]Elias Frantar and Dan Alistarh. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot. InICML, 2023. [37]Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate Quantization for Generative Pre-Trained Transformers. InICLR, 2022. [38]Elias Frantar, Roberto L Castro, Jiale Chen, Torsten Hoefler, et al. MARLIN: Mixed-Precision Auto- Regressive Parallel Inference on Large Language Models.arXiv preprint arXiv:2408.11743, 2024. [39]Trevor Gale, Erich Elsen, and Sara Hooker. The State of Sparsity in Deep Neural Networks.arXiv preprint arXiv:1902.09574, 2019. [40]Leo Gao, Stella Biderman, Sid Black, Laurence Golding, et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling.arXiv preprint arXiv:2101.00027, 2020. [41]Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, et al. A Framework for Few-Shot Language Model Evaluation, 07 2024. [42]Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, et al. The Language Model Evaluation Har- ness, 07 2024. [43]Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, et al. A Survey of Quantization Methods for Efficient Neural Network Inference. InLow-Power Computer Vision, pages 291–326. Chapman and Hall/CRC, 2022. [44]Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. OpenWebText Corpus, 2019. [45]Donald Goldfarb, Yi Ren, and Achraf Bahamou. Practical Quasi-Newton Methods for Training Deep Neural Networks.Advances in Neural Information Processing Systems, 33:2386–2396, 2020. [46]Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. Knowledge Distillation: A Survey. International Journal of Computer Vision, 129(6):1789–1819, 2021. BIBLIOGRAPHY107 [47]Park Gunho, Park Baeseong, Kwon Se Jung, Kim Byeongwook, et al. nuQmm: Quantized MatMul for Efficient Inference of Large-Scale Generative Language Models.arXiv preprint arXiv:2206.09557, 2022. [48]Han Guo, Philip Greengard, Eric P Xing, and Yoon Kim. LQ-LoRA: Low-Rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning. InICLR, 2024. [49]Jinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu, et al. Compressing Large Language Models by Joint Sparsification and Quantization. InICML, 2024. [50]Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned Stochastic Tensor Optimiza- tion. InICML, pages 1842–1850. PMLR, 2018. [51]Song Han, Huizi Mao, and William J Dally. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.arXiv preprint arXiv:1510.00149, 2015. [52]Song Han, Jeff Pool, John Tran, and William Dally. Learning Both Weights and Connections for Efficient Neural Network.NeurIPS, 2015. [53]Song Han, Jeff Pool, John Tran, and William Dally. Learning Both Weights and Connections for Efficient Neural Network.Advances in Neural Information Processing Systems, 28, 2015. [54]Babak Hassibi, David Stork, and Gregory Wolff. Optimal Brain Surgeon: Extensions and Performance Comparisons.NeurIPS, 1993. [55]Babak Hassibi, David G Stork, and Gregory J Wolff. Optimal Brain Surgeon and General Network Pruning. InIEEE International Conference on Neural Networks, pages 293–299. IEEE, 1993. [56]Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recog- nition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. [57]Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, et al. Measuring Massive Multitask Language Understanding.arXiv preprint arXiv:2009.03300, 2020. [58]Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, et al. Sparsity in Deep Learning: Pruning and Growth for Efficient Inference and Training in Neural Networks.JMLR, 2021. [59]Younes Hourri, Mohammad Mozaffari, and Maryam Mehri Dehnavi. PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs.arXiv preprint arXiv:2509.23410, 2025. [60]Jeremy Howard and Sebastian Ruder. Universal Language Model Fine-Tuning for Text Classification. InACL, pages 328–339, 2018. [61]Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, et al. LoRA: Low-Rank Adaptation of Large Language Models. InICLR, 2022. [62]Yuezhou Hu, Kang Zhao, Weiyu Huang, Jianfei Chen, et al. Accelerating Transformer Pre-Training with 2:4 Sparsity. InICML, 2024. [63]Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, et al. Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N:M Transposable Masks.NeurIPS, 2021. BIBLIOGRAPHY108 [64]Ivan Ilin and Peter Richtarik. Thanos: A Block-Wise Pruning Algorithm for Efficient Large Language Model Compression.arXiv preprint arXiv:2504.05346, 2025. [65]Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, et al. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. InCVPR, 2018. [66]Eric Jang, Shixiang Gu, and Ben Poole. Categorical Reparameterization with Gumbel-Softmax. In ICLR, 2017. [67]Yu Ji, Ling Liang, Lei Deng, Youyang Zhang, et al. TETRIS: Tile-Matching the Tremendous Irregular Sparsity.NeurIPS, 2018. [68]Keller Jordan, Yuchen Jin, Vlado Boza, You Jiacheng, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An Optimizer for Hidden Layers in Neural Networks, 2024. [69]Sheng-Chun Kao, Amir Yazdanbakhsh, Suvinay Subramanian, Shivani Agrawal, et al. Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask.arXiv preprint arXiv:2209.07617, 2022. [70]Diederik P Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. InICLR, 2015. [71]Alex Krizhevsky, Geoffrey Hinton, et al. Learning Multiple Layers of Features from Tiny Images. Technical Report Tr-2009, University of Toronto, 2009. [72]Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet Classification with Deep Convolu- tional Neural Networks.Communications of the ACM, 60(6):84–90, 2017. [73]Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. InSOSP, 2023. [74]Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, et al. RACE: Large-Scale ReAding Compre- hension Dataset from Examinations. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors, EMNLP, pages 785–794, Copenhagen, Denmark, September 2017. Association for Computational Lin- guistics. [75]Yann LeCun, John Denker, and Sara Solla. Optimal Brain Damage.Advances in Neural Information Processing Systems, 2, 1989. [76]Jaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn, et al. Layer-Adaptive Sparsity for the Magnitude- Based Pruning, 2021. [77]Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Measuring the Intrinsic Dimension of Objective Landscapes. InICLR, 2018. [78]Lujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu, Qiang Wang, Wenhan Luo, Wei Xue, Qifeng Liu, Xiaowen Chu, and Yike Guo. Discovering Sparsity Allocation for Layer-Wise Pruning of Large Lan- guage Models.Advances in Neural Information Processing Systems, 37:141292–141317, 2024. [79]Wei Li, Lujun Li, Mark Lee, and Shengjie Sun. Adaptive Layer Sparsity for Large Language Models via Activation Correlation Assessment.Advances in Neural Information Processing Systems, 37:109350– 109380, 2024. BIBLIOGRAPHY109 [80]Yixiao Li, Yifan Yu, Qingru Zhang, Chen Liang, et al. LoSparse: Structured Compression of Large Language Models Based on Low-Rank and Sparse Approximation. InICML, 2023. [81]Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, et al. AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration.MLSys, 2024. [82]Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma. Same Pre-Training Loss, Better Down- stream: Implicit Bias Matters for Language Models. InICML, 2023. [83]Hongyi Liu, Rajarshi Saha, Zhen Jia, Youngsuk Park, et al. ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs.arXiv preprint arXiv:2502.00258, 2025. [84]Zechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo, et al. MetaPruning: Meta Learning for Auto- matic Neural Network Channel Pruning. InICCV, 2019. [85]Ilya Loshchilov and Frank Hutter. SGDR: Stochastic Gradient Descent with Warm Restarts. InICLR, 2017. [86]Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization. InICLR, 2019. [87]Haihao Lu, Zedong Peng, and Jinwen Yang. MPAX: Mathematical Programming in JAX.arXiv preprint arXiv:2412.09734, 2024. [88]Haihao Lu and Jinwen Yang. A Practical and Optimal First-Order Method for Large-Scale Convex Quadratic Programming.arXiv preprint arXiv:2311.07710, 2023. [89]Yucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov, et al. STEP: Learning N:M Structured Sparsity Masks from Scratch with Precondition.arXiv preprint arXiv:2302.01172, 2023. [90]Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, et al. AffineQuant: Affine Transformation Quanti- zation for Large Language Models.arXiv preprint arXiv:2403.12544, 2024. [91]Mehdi Makni, Kayhan Behdin, Zheng Xu, Natalia Ponomareva, and Rahul Mazumder. A Unified Framework for Sparse Plus Low-Rank Matrix Decomposition for LLMs. InThe Second Conference on Parsimony and Learning (Proceedings Track), 2025. [92]James Martens. New Insights and Perspectives on the Natural Gradient Method.Journal of Machine Learning Research, 21(1):5776–5851, 2020. [93]James Martens and Roger Grosse. Optimizing Neural Networks with Kronecker-Factored Approximate Curvature. InICML, pages 2408–2417. PMLR, 2015. [94]Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer Sentinel Mixture Mod- els, 2016. [95]Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer Sentinel Mixture Mod- els.arXiv preprint arXiv:1609.07843, 2016. [96]Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, et al. FP8 Formats for Deep Learn- ing.arXiv preprint arXiv:2209.05433, 2022. BIBLIOGRAPHY110 [97]Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. InEMNLP, 2018. [98]Mohammad Mozaffari, Samuel Kushnir, Maryam Mehri Dehnavi, and Amir Yazdanbakhsh. OPTIMA: Optimal One-Shot Pruning for LLMs via Quadratic Programming Reconstruction.arXiv preprint arXiv:2512.13886, 2025. [99]Mohammad Mozaffari, Sikan Li, Zhao Zhang, and Maryam Mehri Dehnavi. MKOR: Momentum- Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates. InNeurIPS, 2023. [100]Mohammad Mozaffari, Amir Yazdanbakhsh, and Maryam Mehri Dehnavi. SLiM: One-Shot Quantized Sparse Plus Low-Rank Approximation of LLMs. InICML, 2025. [101]Mohammad Mozaffari, Amir Yazdanbakhsh, Zhao Zhang, and Maryam Mehri Dehnavi. SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs. InICLR, 2025. [102]Baorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban, et al. HyLo: A Hybrid Low-Rank Natu- ral Gradient Descent Method. InProceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pages 1–16, 2022. [103]Tan Nguyen, Vai Suliafu, Stanley Osher, Long Chen, et al. FMMformer: Efficient and Flexible Trans- former via Decomposed Near-Field and Far-Field Attention. InNeurIPS, 2021. [104]Mahdi Nikdan, Soroush Tabesh, and Dan Alistarh. RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation. InICML, 2024. [105]NVIDIA. NVIDIA A100 Tensor Core GPU Architecture, 2020. Version 1.0. [106]NVIDIA. NVIDIA Hopper Architecture Whitepaper, 2022. Accessed via GTC 2022 presentation materials. [107]NVIDIA, Péter Vingelmann, and Frank H.P. Fitzek. CUDA, Release: 10.2.89, 2020. [108]NVIDIA Corporation. NVIDIA Ampere Architecture In-Depth.https://developer.nvidia. com/blog/nvidia-ampere-architecture-in-depth. [109]NVIDIA Corporation. NVIDIA cuBLAS.https://docs.nvidia.com/cuda/cublas/. [110]NVIDIA Corporation. NVIDIA cuSPARSELt.https://docs.nvidia.com/cuda/ cusparselt/index.html. [111]NVIDIA Corporation. NVIDIA cuSPARSELt Functions.https://docs.nvidia.com/cuda/ cusparselt/functions.html . [112]NVIDIA Corporation. NVIDIA Deep Learning Examples.https://github.com/NVIDIA/ DeepLearningExamples. [113]Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, et al. Large-Scale Distributed Second- Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12359–12367, 2019. BIBLIOGRAPHY111 [114]Eunhyeok Park, Sungjoo Yoo, and Peter Vajda. Value-Aware Quantization for Training and Inference of Neural Networks. InECCV, 2018. [115]David Patterson, Joseph Gonzalez, Quoc V. Le, Chen Liang, et al. Carbon Emissions and Large Neural Network Training, 2021. [116]J Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman, et al. KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks. InSC, 2021. [117]J Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu, et al. Convolutional Neural Network Train- ing with Distributed K-FAC. InSC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–12. IEEE, 2020. [118]Qwen, An Yang, Baosong Yang, et al. Qwen2.5 Technical Report, 2025. [119]Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving Language Under- standing by Generative Pre-Training.OpenAI, 2018. [120]Alec Radford, Jeffrey Wu, Rewon Child, David Luan, et al. Language Models Are Unsupervised Mul- titask Learners.OpenAI Blog, 1(8):9, 2019. [121]Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.arXiv e-prints, 2019. [122]Arya Rafii, Victor Kamel, and Maryam Mehri Dehnavi. Stoicc.https://paramathic.github. io/stoicc-docs/, 2025. [123]Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ Questions for Machine Comprehension of Text.arXiv preprint arXiv:1606.05250, 2016. [124]Yi Ren and Donald Goldfarb. Efficient Subsampled Gauss-Newton and Natural Gradient Methods for Training Neural Networks.arXiv preprint arXiv:1906.02353, 2019. [125]Babak Rokh, Ali Azarpeyvand, and Alireza Khanteymoori. A Comprehensive Survey on Model Quan- tization for Deep Neural Networks in Image Classification.ACM Transactions on Intelligent Systems and Technology, 14(6):1–50, 2023. [126]Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An Adver- sarial Winograd Schema Challenge at Scale.Communications of the ACM, 64(9):99–106, 2021. [127]Victor Sanh, Thomas Wolf, and Alexander Rush. Movement Pruning: Adaptive Sparsity by Fine- Tuning.NeurIPS, 2020. [128]Jürgen Schmidhuber. Deep Learning in Neural Networks: An Overview.Neural Networks, 61:85–117, 2015. [129]Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, et al. OmniQuant: Omnidirectionally Cali- brated Quantization for Large Language Models. InICLR, 2024. [130]Noam Shazeer. GLU Variants Improve Transformer.arXiv preprint arXiv:2002.05202, 2020. BIBLIOGRAPHY112 [131]Noam Shazeer and Mitchell Stern. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. InICML, 2018. [132]Shaohuai Shi, Lin Zhang, and Bo Li. Accelerating Distributed K-FAC with Smart Parallelism of Com- puting and Communication Tasks. In2021 IEEE 41st International Conference on Distributed Com- puting Systems (ICDCS), pages 550–560. IEEE, 2021. [133]Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, et al. On the Effect of Pretraining Cor- pora on In-Context Learning by a Large-Scale Language Model.arXiv preprint arXiv:2204.13509, 2022. [134]Sidak Pal Singh and Dan Alistarh. WoodFisher: Efficient Second-Order Approximation for Neural Network Compression.NeurIPS, 2020. [135]Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, et al. SlimPajama: A 627B Token Cleaned and Deduplicated Version of RedPajama.https://bit.ly/slimpajamas, 2023. [136]Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. A Simple and Effective Pruning Approach for Large Language Models. InICLR, 2024. [137]Wei Sun, Aojun Zhou, Sander Stuijk, Rob Wijnhoven, et al. DominoSearch: Find Layer-Wise Fine- Grained N:M Sparse Schemes from Dense Neural Networks. InNeurIPS, 2021. [138]Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, et al. Long Range Arena: A Benchmark for Efficient Transformers. InICLR, 2021. [139]Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, et al. Gemma 3 Technical Report. arXiv preprint arXiv:2503.19786, 2025. [140]Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, et al. Gemma 2: Improving Open Language Models at a Practical Size.arXiv preprint arXiv:2408.00118, 2024. [141]Texas Advanced Computing Center. Lonestar 6.https://tacc.utexas.edu/systems/ lonestar6/. [142]Vithursan Thangarasa, Abhay Gupta, William Marshall, Tianda Li, et al. SPDF: Sparse Pre-Training and Dense Fine-Tuning for Large Language Models.arXiv preprint arXiv:2303.10464, 2023. [143]Philippe Tillet, H. T. Kung, and David Cox. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations. InMAPL 2019: Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, 2019. [144]Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, et al. LLaMA 2: Open Foundation and Fine- Tuned Chat Models.arXiv preprint arXiv:2307.09288, 2023. [145]Yuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse, et al. Rich Information Is Affordable: A Systematic Performance Analysis of Second-Order Optimization Using K-FAC. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2145– 2153, 2020. BIBLIOGRAPHY113 [146]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, et al. Attention Is All You Need.Ad- vances in Neural Information Processing Systems, 30, 2017. [147]Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, et al. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. InICLR, 2019. [148]Shibo Wang and Pankaj Kanwar. BFloat16: The Secret to High Performance on Cloud TPUs.http: //bit.ly/3WEtCGm, 2019. [149]Wenxuan Wang and Zhaopeng Tu. Rethinking the Value of Transformer Components, 2020. [150]Wikipedia. Wikipedia Corpus.https://meta.wikimedia.org/wiki/Data_dump_ torrents#English_Wikipedia. [151]Lucas Wilkinson, Kazem Cheshmi, and Maryam Mehri Dehnavi. Register Tiling for Unstructured Sparsity in Neural Network Inference.PLDI, 2023. [152]Samuel Williams, Andrew Waterman, and David Patterson. Roofline: An Insightful Visual Perfor- mance Model for Multicore Architectures.Communications of the ACM, 2009. [153]Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, et al. Transformers: State-of-the-Art Natural Language Processing. InEMNLP (System Demonstrations), pages 38–45, 2020. [154]Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, et al. Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.arXiv preprint arXiv:2309.10285, 2023. [155]Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, et al. SmoothQuant: Accurate and Efficient Post- Training Quantization for Large Language Models. InICML, 2023. [156]Peng Xu, Wenqi Shao, Mengzhao Chen, Shitao Tang, Kaipeng Zhang, Peng Gao, Fengwei An, Yu Qiao, and Ping Luo. BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation.arXiv preprint arXiv:2402.16880, 2024. [157]Minghan Yang, Dong Xu, Zaiwen Wen, Mengyun Chen, et al. Sketchy Empirical Natural Gradient Methods for Deep Learning.arXiv preprint arXiv:2006.05924, 2020. [158]Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, et al. Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity. InICML, 2024. [159]Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, et al. Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes. InICLR, 2020. [160]Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. On Compressing Deep Models by Low Rank and Sparse Decomposition. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 7370–7379, 2017. [161]Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, et al. HellaSwag: Can a Machine Really FinishYourSentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. BIBLIOGRAPHY114 [162]Cheng Zhang, Jianyi Cheng, George A Constantinides, and Yiren Zhao. LQER: Low-Rank Quantiza- tion Error Reconstruction for LLMs.arXiv preprint arXiv:2402.02446, 2024. [163]Lin Zhang, Shaohuai Shi, and Bo Li. EVA: A General Vectorized Approximation Framework for Second-Order Optimization, 2023. [164]Stephen Zhang and Vardan Papyan. OATS: Outlier-aware pruning through sparse and low rank decom- position. InThe Thirteenth International Conference on Learning Representations (ICLR), 2025. [165]Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, et al. OPT: Open Pre-Trained Transformer Language Models.arXiv preprint arXiv:2205.01068, 2022. [166]Yuxin Zhang, Yiting Luo, Mingbao Lin, Yunshan Zhong, et al. Bi-Directional Masks for Efficient N:M Sparse Training.arXiv preprint arXiv:2302.06058, 2023. [167]Ningxin Zheng, Bin Lin, Quanlu Zhang, Lingxiao Ma, et al. SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute. InOSDI, 2022. [168]Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, et al. Learning N:M Fine-Grained Structured Sparse Neural Networks from Scratch.arXiv preprint arXiv:2102.04010, 2021. [169]Zihan Zhou, Xiaodong Li, John Wright, Emmanuel J. Candès, and Yi Ma. Stable Principal Component Pursuit. InIEEE International Symposium on Information Theory (ISIT), pages 1518–1522. IEEE, 2010. [170]Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, et al. Aligning Books and Movies: To- wards Story-Like Visual Explanations by Watching Movies and Reading Books. InProceedings of the IEEE International Conference on Computer Vision, pages 19–27, 2015. Appendix A Supplementary Material for MKOR Inthischapter, wereporttheGLUEresultsachievedoneachtaskinAppendixA.1. Next, wediscussthederiva- tion of KFAC and SNGD approximation methods inAppendix A.2. We then discuss some of the features of MKOR and other optimizers and back them up with quantitative data in Appendix A.3andAppendix A.4. We provide scalability results of MKOR inAppendix A.5and provide more data to back up the low-rank features of the covariance matrices inAppendix A.6. InAppendix A.7, we analyze MKOR and other optimizers on the training tasks as non-convex optimizers, only concerning their performance on training tasks. We also describe the knee-point learning rate scheduler inAppendix A.8. We conclude this chapter by proving the lemmas used in the preceding sections. A.1 GLUE Results We discussed the speedup achieved using MKOR on the GLUE dataset inAppendix 3.4. For completeness, Table A.1shows the metrics achieved in each of the different GLUE tasks on BERT-Large-Uncased trained on different optimizers. Table A.1: BERT-Large-Uncased Results on the GLUE classification tasks. Optimizer Iterati- ons MNLI (acc) QQP (F1) QNLI (acc) SST-2 (acc) COLA (mcc) STS-B (corr) MRPC (F1) RTE (acc) Avera- ge LAMB1,5630.8410.8780.9130.9190.5160.8750.8120.6640.8023 KAISA1,5630.8210.8540.9000.9210.4890.8780.8880.6170.796 MKOR1,5000.8440.8790.9160.9230.5230.8920.9050.6900.8214 MKOR6000.8330.8780.9040.9210.4940.8860.8930.6530.8078 MKOR-H6000.8380.8770.9110.9210.5020.8860.8980.6570.811 Eva10000.8390.8770.9070.9140.4990.8900.9040.6500.809 A.2 Derivation of NGD Approximations Natural Gradient Descent (NGD)In NGD, which is a second-order method, we use the inverse of the Fisher Information Matrix (FIM) as a preconditioner to the gradients as shown inFigure A.1-a.Equation 3.1shows the update rule of NGD, whereF m is the FIM block corresponding to blockm. Equation A.1shows the definition of FIM for an arbitrary layer in our model, wherex i is thei th sample in the batch. 115 APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR116 Reshape a. Block Diagonal Approximation b. Kronecker-Factored Based Approximation c. SNGD Based Approximatoin SMW Identity Figure A.1: Approximations in second-order methods. APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR117 F m = 1 b b X i=1 ∇ x i ℓ(W, x i )∇ x i ℓ(W, x i ) T (A.1) Kronecker Factorization (KFAC).KFAC methods reformulate the FIM block as the Kronecker product of two matrices as shown inEquation A.2whereg m =∇ x m Landa m is the vector form of the activation output of layermandEis the expectation operator andx m is the input matrix of layerm. Please note that we have used the mixed-product property of Kronecker multiplication for getting the right hand value. F m =E[(g m ⊗a m−1 )(g m ⊗a m−1 ) T ] =E[(g m g m T )⊗(a m−1 a m−1 T )](A.2) Furthermore, we assume thatE[(g m g m T )⊗(a m−1 a m−1 T )]≈E[g m g m T ]⊗E[a m−1 a m−1 T ], which is a strong assumption, but helps us simplify the computation further. Using the inversion property of Kronecker multiplication, we can compute the inverse of FIM using Equation A.3. F m −1 w m =E[g m g m T ] −1 ⊗E[a m−1 a m−1 T ] −1 w m (A.3) By using the mixed Kronecker matrix-vector product property, we can get the update value inEquation 3.2, which is illustrated inFigure A.1-b. We refer toL m andR m as the left and right factors respectively. Adding momentum to the left and right factors and denoting the iteration number with a subscript to the factors, we will getEquation 3.3andEquation 3.4. Sherman-Morrison-Woodbury-Based Natural Gradient Descent (SNGD).In this method, the SMW iden- tity is used for approximating the inverse of(F m +μI)∈R d 2 ×d 2 , whereμis a damping factor used in the preconditioning. Equation A.4shows the process of computing the inverse of the FIM for a single layer in the network, whereA m ∈R d×b isthebatchofactivationsoflayerlandG m ∈R d×b isthebatchofgradientsofthe loss function with respect to the inputs of that layer andU= [∇ W m L(W, x 1 ), ...,∇ W m L(W, x b )] T ∈R d 2 ×b is the concatenation of the gradients of the loss function with respect to the parameters of that layer and⊙ shows the Hadamard element-wise product. In this method, a kernel matrix inR b×b is inverted, as shown in Figure A.1-c. (F m +μI) −1 = 1 μ (I−U m (A m−1 T A m−1 ⊙G m T G m +μI) −1 U m T )(A.4) A.3 Numerical Instability of Second-order Methods InAppendix 3.3.3, we discussed that in second-order methods, multiple matrix inversion or root-finding algo- rithms need to be executed, which make the second-order methods prone to numerical instabilities. Further- more, we discussed that left and right factors in second-order methods have large condition numbers, resulting in further issues in inversion.Figure A.2shows the eigenvalues of the right factor and its condition number for ResNet-50 model on CIFAR-10 dataset on KFAC algorithm. Even when using damping factors and filtering out extremely small eigenvalues, the condition number of these matrices is large, motivating the use of double precision computations for avoiding numerical instabilities. MKOR, on the other hand, doesn’t suffer numerical instabilities when inverting such matrices, and its computational complexity isn’t dependent on the condition number either. APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR118 Table A.2: Number of epochs necessary for convergence in different optimizers for ResNet-50 on CIFAR10. MKOR is the least sensitive optimizer to learning rate, converging in almost the same number of iterations for a wide range of learning rate, while other optimizers either diverge (D) or converge to a local-minimum (∗ superscript). Optimizer Learning Rate 1010.10.01 MKOR94797876 KAISA1121009089 ∗ HyLoD123 ∗ 98150 ∗ SGDDD108145 ∗ A.4 Sensitivity to Learning Rate Learning rate is one of the main hyperparameters in machine learning (ML) that can directly affect the con- vergence time of optimization, and ML practitioners have to spend a lot of time tuning this hyperparameter. More specifically, in first-order methods, a large learning rate can easily lead to divergence and numerical instability, and in second-order methods, large learning rates can lead to exploding gradients as discussed in Appendix 3.3.3. Using small learning rates can lead to slow convergence in both first- and second-order meth- ods and can even jeopardize the main advantage of second-order methods, which is their faster convergence rate. One of the main advantages of our method is its robustness against a wide range of learning rates. As TableA.2shows, first-ordermethodsareextremelysensitivetothelearningrate, andthesecond-ordermethods are prone to ripples and divergent for a larger range of learning rates, and lose their performance for small learning rates. Our method, on the other hand, will converge with a high convergence rate for a wide range of learning rates, and by directly modifying the inverse of factors as discussed in Appendix 3.3.3can find a proper equilibrium between first- and second-order methods. This table shows that our method is the least sensitive to the learning rate values and can make the job of ML practitioners for tuning this hyperparameter extremely easy. A.5 Scalability Figure A.3shows the strong scalability of MKOR on BERT-Large-Uncased on up to 64 GPUs. A.6 Decaying Eigenvalues and Rank-1 Approximations Figure A.4shows that the eigenvalues of the factors will decay as the model converges, making rank-1 ap- proximations more effective. The reason behind the decay in the eigenvalues is that the weights are initialized randomly and the neurons work independently in the beginning of the training, but as the model converges, the neurons become more dependent on each other and thus the activation and input gradients will become linearly dependent. This is also reflected by some large error values in the distributions in Figure 3.6[a, b, c, d]. The factors in MKOR are initialized with identity, starting MKOR from a first-order method. As a result, MKOR is more robust against noise in the approximations in the first iterations (the approximation error does not noticeably affect the factors when replacingL m−1 t−1 −1 andR m t−1 −1 in Equation 3.5andEquation 3.6with identity). But as the model converges, the factors in MKOR will be mostly shaped by the training samples, APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR119 making MKOR more reliable on less erroneous approximations, and the decaying eigenvalues of the factors help MKOR with that. A.7 Training Accuracy Experiments To evaluate MKOR as an optimizer that tries to minimize a specific objective function, we have considered the case of only minimizing the loss function of models on different tasks and set all the other optimizer parameters such as weight decay to zero, since using a non-zero weight decay adds a quadratic term to the loss function and using different weight decays for different optimizers leads to optimizing different objective functions, which might be considered unfair. Recent work has shown the advantage of second-order methods over their first-order counterparts on multi- ple CNN tasks [102,116], such as residual networks [56]. In our training accuracy experiment, we use another CNN benchmark,AlexNet[72] with more than 20M parameters onCIFAR-100[71] consisting of 50K training and 10K validation images of 100 classes.Figure A.5-c andFigure A.6-c show the convergence properties of different optimizers. MKOR is1.26×,1.31×, and1.58×faster than HyLo-KIS, SGD, and KAISA respec- tively. The reason for the low convergence speed of KAISA is that we needed to use small learning rates for avoiding exploding gradients in it, which has damaged its convergence rate. BERT is a large language model with two variants, BERT-Base with more than 108M parameters and BERT-Largewith more than 335M parameters. As shown inFigure A.5-a andFigure A.6-a, we have fine- tunedBERT-Largeon theIMDBdataset which is a text classification task with 25K training and 25K test samples. MKOR outperforms SGD and HyLo-KIS by a speedup factor of1.22×and1.43×respectively. We have also fine-tunedBERT-BaseonSQuADdataset, which is a question answering task with 87.6K training and 10.6K test samples. MKOR achieves1.26×and1.56×speedup over SGD and HyLo respectively. Using a wide range of learning rates, KAISA could not converge on any of ourBERTexperiments which are based on the HuggingFace implementation of BERT, and the reason for lack of convergence of KAISA is exploding gradients. A.8 Knee-Point Learning Rate Scheduler 1 While using large learning rates is crucial for utilizing the higher convergence rate of second-order methods, it is necessary to reduce the learning rate after some iterations so that the model converges, and the number of iterations for changing the learning rate is not known in advance. In practice, machine learning practition- ers will manually find the number of iterations by trial and error or use predefined functions [ 60,85] that don’t necessarily work ideally in all optimization problems. We used knee-point learning rate scheduler in Appendix A.7. To fully utilize the potentials of the optimizers, we have designed a learning rate scheduler that monitors the rate of improvement in accuracy or decrease in the loss function value, and based on that decides when to decrease the learning rate. The scheduler detects knee-points in the accuracy/loss and decreases the learning rate when a knee-point is observed. By definition, knee-points are defined as the points where the average accuracy/loss rate is less thanβ times the increment/decrement in the accuracy/loss since using the current learning-rate. For averaging the 1 The knee-point learning rate is not used in any of the experiments in the main chapters. APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR120 accuracy/loss rate, we use an exponential moving average, andβis a hyperparameter that we can choose to show how much the scheduler can tolerate lack of improvement to detect the accuracy/loss. A.9 Proofs Theorem 3.3.1: Proof.Given a positive definite matrixJ t−1 and a vectorjand scalar0< γ <1, we show thatEquation A.5 results in a positive-definite matrix. J t −1 =γJ t−1 −1 + (1−γ) γ 2 (1 +γ(1−γ)j T J t−1 −1 j) J t−1 −1 j T J t−1 −1 (A.5) Sinceγ >0andJ t−1 is a positive-definite matrix,γJ t−1 is also positive-definite. Also, sinceJ t−1 is a positive-definite matrix,∀x̸= 0 :x T J t−1 x >0. j T L −1 t−1 j >0 0<γ<1 −→1 +γ(1−γ)j T J t−1 −1 j >0(A.6) Now we show thatJ t−1 −1 j T J t−1 −1 is also positive-definite. ∀x̸= 0 :x T J t − 1 −1 j T J t−1 −1 x= (j T J −1 t−1 x) T (j T J −1 t−1 x) =∥(j T J −1 t−1 x)∥ 2 >0(A.7) Since both matrices on the right-hand-side of Equation A.5are positive-definite and the sum of two positive- definite matrices is a positive-definite matrix, the left-hand-side of it will be also positive definite. Theorem 3.3.3: Proof.By definingP= (ζL −1 + (1−ζ)I)⊗(ζR −1 + (1−ζ)I),Equation A.8will hold. ˆ L(w 0 −∆w)−L(w 0 ) =−∇L(w 0 ) T P∇L(w 0 )(A.8) Now, we will show thatPis a positive-semi-definite matrix. By using the associativity property of Kro- necker multiplication, we can representPas in Equation A.9. Please note that different identity matrices in Equation A.9have different shapes. P=ζ 2 L −1 ⊗R −1 +ζ(1−ζ)L −1 ⊗I+ζ(1−ζ)I⊗R −1 +I(A.9) Since matricesLandRare positive-semi-definite, the Kronecker productsL −1 ⊗R −1 ,L −1 ⊗I, and I⊗R −1 are also positive-semi-definite. As a result, for any non-zero vectorx,x T P x >0. So, based on Equation A.8,−∇L(w 0 ) T P∇L(w 0 ). As a result,Equation A.10holds. ˆ L(w 0 −∆w)<L(w 0 )(A.10) Theorem 3.3.2 APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR121 Proof.Considering the quantization of matrixJand vectorjinEquation A.11and assuming the maximum quantization error isεand the maximum values in vectorjand matrixJism, we can consider one of the three possible cases: 1.Vector-Vector Dot Product: The resulting error isO(2εm 2 d)because the error of each multiplication can at most be2mε+ε 2 and by adding thedmultiplications, the maximum error can be(2m+ 1)dε. 2.Vector-Matrix Product: The resulting error isO(2εm 2 d)because a vector-matrix product can be con- sidered as multiple independent vector-vector dot products. 3.Vector-Vector Outer Product: The resulting error isO(2εm)because each element in the resulting matrix is computed using a multiplication with the maximum error of2εm 2 +ε 2 . γJ+ (1−γ) γ 2 (1 +γ(1−γ)j T Jj) Jjj T J T (A.11) The error imposed by quantizingγJis at mostγε. The quantization error in the denominator of the fraction inEquation A.11is negligible in comparison to1and won’t change the growth order of the final error term. The quantization error ofJjisO(2m 2 dε)as discussed earlier in the case of vector-matrix product. Jjj T J T = (Jj)(Jj) T is a vector-vector outer product, resulting inO(4m 3 d 2 )error. So the final error quantization error isO((γ+ 4 (1−γ) γ 2 m 3 d 2 )ε) APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR122 0100200300 Iteration (x1000) 10 −3 10 −2 10 −1 10 0 Eigenvalue ResNet-50 - CIFAR-10 - Eigenvalues of Right Factor Minimum Eigenvalues Maximum Eigenvalues (a) 0100200300 Iteration (x1000) 10 0 10 1 10 2 10 3 Condition Number ResNet-50 - CIFAR-10 - Condition Number of Right Factor Condition Number (b) Figure A.2: Maximum and minimum eigenvalues (a) and the condition number (b) of the right factors in KFAC when training ResNet-50 on CIFAR-10. As illustrated, the minimum eigenvalues of the factors in KFAC approach zero, meaning that the factors are singular, and hence have large condition numbers, making numerical inversion of them complex and numerically unstable. APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR123 0102030405060 #GPUs 20 30 40 50 60 Speedup BERT Large Uncased Scalability Num GPUs MKOR Figure A.3: Scalability of MKOR. 020040060080010001200 Iteration 0.2 0.4 0.6 0.8 1.0 Average Covariance Rank-1 Error ResNet-50 Average Covariance Rank-1 Error Gradient Activation Figure A.4: Average covariance rank-1 approximation error for ResNet-50 in different iterations APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR124 01000020000300004000050000 Time (s) 0 20 40 60 80 100 Accuracy 156392232319154 BERT Large Cased Text Classification - IMDB MKOR HyLo-KIS SGD 01000020000300004000050000600007000080000 Time (s) 0 20 40 60 80 100 Accuracy 784745127563918 BERT Base Cased Question Answering - SQuAD HyLo MKOR SGD (a)(b) 025050075010001250150017502000 Time (s) 0 20 40 60 80 100 Accuracy 2119140015091562 AlexNet - CIFAR-100 KAISA MKOR HyLo-KIS SGD (c) Figure A.5: Training time for distributed first- and second-order optimizers SGD, MKOR, KAISA, and HyLo onBERT-Large-CasedonIMDB(a),BERT-Base-CasedonSQuAD(b), andAlexNetonCIFAR-100(c). In all the experiments, MKOR outperforms other optimizers in convergence speed. APPENDIX A. SUPPLEMENTARY MATERIAL FOR MKOR125 0510152025303540 Epochs 0 20 40 60 80 100 Accuracy 112022 BERT Large Cased Text Classification - IMDB MKOR HyLo-KIS SGD 0102030405060 Epochs 0 20 40 60 80 100 Accuracy 603757 BERT Base Cased Question Answering - SQuAD HyLo MKOR SGD (a)(b) 050100150200250300350400 Epochs 0 20 40 60 80 100 Accuracy 395261280293 AlexNet - CIFAR-100 KAISA MKOR HyLo-KIS SGD (c) Figure A.6: Training accuracy vs. the number of epochs for distributed first- and second-order optimizers SGD, MKOR, KAISA, and HyLo onBERT-Large-CasedonIMDB(a),BERT-Base-CasedonSQuAD(b), andAlexNetonCIFAR-100(c). In all the experiments, MKOR outperforms other optimizers in convergence rate. Appendix B Supplementary Material for SLOPE This chapter provides supplementary material for the SLOPE chapter. We begin with a comparison against dynamic sparsity using SR-STE inAppendix B.1, followed by an analysis of the cuSPARSELt initializa- tion overhead in Appendix B.2. We present BERT-Large-Uncased pretraining and downstream evaluation results inAppendix B.3, and discuss the performance overhead of the bidirectional mask inAppendix B.4. Appendix B.5analyzes the sparsity ratio in the double-pruned backward pass, andAppendix B.6examines sensitivity to the choice of pruning matrix. Implementation details are provided inAppendix B.7, task-specific GLUE results inAppendix B.8, and the integration with Flash Attention inAppendix B.9. We compare with dense models in Appendix B.10and conclude with proofs of the lemmas used in the chapter. B.1 Comparison with Dynamic Sparsity: SR-STE We pretrained GPT2-Small (Appendix 4.4.2) using the SR-STE method [168] and reported the perplexity results inFigure 4.2. SR-STE aims to mitigate the Sparse Architecture Divergence (SAD) by dynamically adjusting the sparsity mask throughout training. We have tested various decay factor hyperparameters to find the optimal optimization strategy for SR-STE. TounderstandtheperformancegapbetweenSR-STEandSLOPE(ourmethod)forthesametrainingbudget, we analyzed the mask dynamics in SR-STE. We plotted theaverage number of mask elementschanges during training compared to the final converged mask sparsity pattern. High mask change values indicate that training resources are spent on updating weights that ultimately get pruned and do not necessarily contribute to the final model accuracy. Figure B.1shows this average mask difference per iteration relative to the converged model. As training progresses, the mask difference decreases, demonstrating SR-STE’s convergence to a specific sparsity pattern. However, in SLOPE, where all resources are dedicated to optimizing weights under astaticmask 1 , SR-STE’s dynamic approach leads to wasted computation (represented by the area under the curve inFigure B.1). Con- sequently, for the same training budget, SLOPE achieves a lower perplexity in comparison to SR-STE due to its static mask approach. 1 We determine the pruning mask at the very first iteration and maintain it for the rest of training. 126 APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE127 50100150200250300 Iteration (×1000) 0.0 0.1 0.2 0.3 0.4 0.5 Difference GPT2-Small SR-STE Average Mask Change Figure B.1: Average mask difference between each iteration and the converged sparsity pattern in GPT2-Small pretraining using SR-STE. The highlighted area shows the ratio of the resources used for updating weights that are pruned and not used in the inference of the model. B.2 cuSPARSELt Initialization Overhead: Static vs. Dynamic Spar- sity This section analyzes the time breakdown of the cuSPARSELt SpMM pipeline, highlighting the significant overheads associated with dynamically changing sparsity masks. The cuSPARSELt SpMM operation consists of two main phases:(1) Setupand(2) Matrix Multiplication. The setup phase involves initializing matrix handles and compressing the 2:4 sparse matrix. This compression copies non-zero values into a contiguous memory layout and generatesindices for those values. The matrix multiplication phase leverages this metadata to perform the sparse matrix-matrix multiplication. Figure B.2shows the setup and multiplication time for square matrices using the cuSPARSELt SpMM backend. As evident from the figure, the setup overhead is significantly larger than the actual matrix mul- tiplication time. For SLOPE, which employs static sparsity masks, the setup cost is incurred only once and becomes negligible compared to the numerous matrix multiplications performed during training and inference. However, for dynamic sparsity patterns, such as Fully Sparse Training [62], Bidirectional Masks [166], and other similar methods[63,137,89,168], this setup overhead can be substantial, leading to reduced speedup (as observed in Appendix 4.4.1for Fully Sparse Training) or slowdowns in some configurations (as discussed inAppendix B.4). 2 B.3 BERT-Large-Uncased: Pretraining and Downstream Evaluation BERT-Large-Uncased pretraining consists of two phases, as illustrated inFigure B.3. Phase 1 comprises 7,038 iterations with a global batch size of 65,536 and a sequence length of 128. Phase 2 includes 1,563 iterations with a global batch size of 32,768 and a sequence length of 512. 2 A recent work observed a similar overhead using dynamic sparsity in cuSPARSELt SpMM pipeline [20]. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE128 7681024204825604096512071689216 Matrix Dimension 0 1 2 3 4 5 Time (ms) cuSPARSELt SpMM Time Breakdown Setup Time Multiplication Time Figure B.2: The setup and multiplication time for square matrices using the cuSPARSELt SpMM backend. Table B.1: End-to-end slow-down of Bi-directional Mask [166] in comparison to the dense baseline. MODELDATASET SLOW-DOWN (×) MOBILENET V2 CIFAR105.08 RESNET-32 CIFAR105.07 VGG19CIFAR108.41 RESNET-18 IMAGENET3.66 RESNET-50 IMAGENET3.01 Figure B.3shows the training loss for both phases under different sparsity settings. We observe that higher sparsity ratios generally lead to higher training loss in both phases. Interestingly, the loss/perplexity gap does not directly correlate with the observed accuracy drops in downstream tasks [ 19,82,133]. We evaluated the pretrained BERT-Large-Uncased models on the SQuAD v1.1 [123] and GLUE [147] benchmarks. SQuAD v1.1, a comprehensive question-answering dataset based on Wikipedia, is widely used for LLM training. We report the F1 score for SQuAD throughout the chapter. GLUE, a diverse benchmark for naturallanguageunderstandingtasks, providesasingleaggregatedscoreacrossvariouschallenges, facilitating model comparisons. The chapter presents the average metric score for GLUE, while task-specific metrics are detailed in Appendix B.8. B.4 Performance overhead of bidirectional mask Table B.1presents the runtime results of Bidirectional Masks [166], a state-of-the-art N:M sparsity method. Our analysis demonstrates that the mask search and associated overheads of this approach result in significant slowdowns compared to dense baselines. For these experiments, we utilized the repository provided in [166] and employed the same models used in their evaluation. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE129 01000200030004000500060007000 Iterations 2 4 6 8 10 Training Loss BERT-Large-Uncased Pretraining Phase 1 Dense 2:4 Sparsity Mixed 2:4 and 4:8 Sparsity 0200400600800100012001400 Iterations 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 Training Loss BERT-Large-Uncased Pretraining Phase 2 Dense 2:4 Sparsity Mixed 2:4 and 4:8 Sparsity Figure B.3: Training loss of BERT-Large-Uncased on WikiCorpus dataset for phase 1 and 2. B.5 Sparsity ratio analysis of double-pruned backward pass As described inAppendix 4.3.1, our proposed sparse pretraining approach involves pruning weights in both the forward and backward passes. During the backward pass, we apply both row-wise and column-wise prun- ing, which introduces additional zero values to the column-wise pruned weight matrices used in the forward pass. Theorem 4.3.1demonstrates that the resulting sparsity ratio can be calculated usingEquation 4.8.Fig- ure B.4visualizes the imposed sparsity ratios for various N:M sparsity patterns. As expected, smaller N/M ratios lead to lower imposed sparsity ratios. Moreover, in most cases, the imposed sparsity ratio is significantly smaller than the original matrix’s density ratio. B.6 Sensitivity to the choice of pruning matrix In linear layers, three matrices are involved in the forward and backward passes: the input, the output gradient, and the weights. Pruning each of these matrices can have distinct effects on model performance. To identify the optimal pruning strategy, we conducted an experiment where we pretrained GPT2-Small APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE130 48163264 M 2 4 6 8 Imposed Sparsity Ratio Imposed Sparsity in Row-wise and Column-wise Pruning 1 :M 2 :M 4 :M Figure B.4: The imposed sparsity ratio when pruning the weight matrices in the backward pass. for 100,000 iterations (a quarter of the full pretraining) while systematically applying both static and dynamic pruning to each of the three matrices. Static pruning involves generating a random mask at initialization and applying it throughout training. Dynamic pruning, on the other hand, prunes matrices based on their magnitude at each iteration. For dynamic pruning, the dense matrix values are computed and stored, and then pruned at every step. Figure B.5presents the validation perplexity for these experiments. Notably, pruning the output gradient led to model divergence after a few iterations and is not shown in the figure. Analysis.As shown inFigure B.5, static pruning consistently achieved lower perplexities. This behavior suggests that focusing computational resources on elements that remain active throughout training can lead to improved performance. Furthermore, pruning weights resulted in lower perplexities compared to pruning inputs, indicating that weights are generally a better target for pruning. Intuition.Pruningweightsisanalogoustoremovingconnectionsbetweenneurons. Pruningactivationtensors is similar to introducing a non-linear function (akin to max-pooling) before each linear layer. Pruning output gradients, however, lacks practical justification and introduces errors into the backward pass, leading to model divergence. B.7 Implementation details This section details the implementation of the custom functions and CUDA kernels used inAlgorithm 2to facilitate efficient sparse training. Initialization, sparse matrix setup, and SpMM kernels.Before utilizing the cuSPARSELt APIs, a crucial initialization phase ensures proper configuration of essential variables for our computational task. Follow- ing initialization, we configure the sparse data formats tailored for sparse matrices. This involves initializing matrix descriptors, pruning the matrices, and compressing them into a more compact representation. cuS- PARSELt employs an automated search to determine the optimal kernel for executing SpMM. While setting up these sparse data formats incurs a non-negligible computational cost, this overhead is mitigated by the APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE131 020406080 100 Iterations (× 1000) 18 20 22 24 26 28 30 32 Test Perplexity GPT2-Small Test Perplexity Dense Static Weight Pruning Dynamic Weight Pruning Dynamic Input Pruning Static Input Pruning Figure B.5: Validation perplexity on GPT2-Small pretraining for 100,000 iterations for different matrix prun- ing settings. Pruning the output gradients leads to divergence within a few iterations and hence is not reported. repetitive nature of matrix multiplications during the training process. Prune and compress.The gradient of the loss function with respect to the weights requires pruning using the same mask as the weight matrix. Consequently, it contains 50% extra zero values in the dense format. To address this redundancy, we developed an optimized CUDA kernel, integrated into PyTorch, that masks the gradients accordingly, eliminating the storage of unnecessary data and reducing memory usage. The output of this operation is a new matrix inR d out × d in 2 . Sparse matrix addition.The cuSPARSELt sparse data format does not natively support addition operations. However, for matricesAandBsharing the same sparsity patterns, we developed an optimized CUDA kernel seamlessly integrated into the PyTorch training workflow. This kernel efficiently computes linear combina- tionsoftheform βA + γB , where β and γ arearbitraryuser-definedconstants. Thisfunctionalityisparticularly useful for adding sparse weights to gradients in optimizers that utilize weight decay. Update Sparse Matrix.After the optimizer updates the weight tensor values based on its rules, we need to update the sparse matrix format to reflect these changes. We implemented an optimized CUDA kernel that copies the weight tensors from the PyTorch format into the cuSPARSELt data type, enabling efficient storage and manipulation of sparse weights. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE132 Table B.2: GLUE results for each task in the experiments discussed inAppendix 4.4. FirstLast MethodPhase Rank1212CoLA SST-2 MRPC STS-B QQP RTE MNLI QNLI Blocks Blocks (mcc) (acc)(f1)(corr) (f1) (acc) (acc) (acc) Dense1,202:42:451.691.981.287.5 87.8 66.4 84.1 91.3 SLOPE MLP Mixer202:42:441.891.488.787.2 85.9 6582.1 90.1 Only SLOPE MLP Mixer +202:42:438.890.485.986.4 85.9 63.5 81.5 89.3 Self-Attention SLOPE with Non-Lazy2402:42:443.390.8898786 64.6 82.3 89.6 Adapters SLOPE with Non-Lazy2402:82:42989.783.785.6 85.2 66.8 79.9 87.4 Adapters SLOPE with Non-Lazy2402:42:844.191.189.886.6 86.3 62.5 82.3 89.6 Adapters SLOPE1,202:42:437.991.485.486.6 85.8 62.5 80.7 88.6 SLOPE1,242:42:438.591.485.886.8 85.8 63.9 80.8 88.4 SLOPE1,2162:42:439.291.386.486.686 63.5 80.8 88.2 SLOPE1,2642:42:442.790.385.186.8 85.7 66.4 80.3 88.5 WANDAN/A02:42:443.091.488.386.9 86.1 63.5 81.9 89.6 WANDAN/A02:82:44.60.8881.38183.3 53.8 76.7 83.9 WANDAN/A02:42:842.191.784.487.2 85.6 63.5 81.5 81.9 Table B.3: Speedup of SLOPE and FlashAttention-2 (FA2) on OPT models. MODELTRAININGINFERENCEINFERENCEINFERENCE SIZEFA2 SLOPE SLOPE + FA2FA2 SLOPE SLOPE + FA2FA2 SLOPE + FA2FA2 SLOPE + FA2 66B1.28 1.131.531.36 1.341.991.311.951.301.91 30B1.36 1.141.661.46 1.322.241.282.241.272.20 13B1.47 1.121.841.61 1.302.481.302.241.122.19 6.7B1.60 1.081.941.71 1.212.501.132.501.122.45 2.6B2.26 1.052.562.47 1.073.231.053.091.002.92 B.8 Task-specific GLUE results The GLUE benchmark [147] comprises eight distinct natural language understanding classification tasks. WhileAppendix 4.4presented the average GLUE score as a measure of overall model performance, this section provides a more detailed analysis by presenting the complete task-specific results for each training setting inTable B.2. B.9 Integration with Flash Attention Toshowthecompatibilityof SLOPEwithotheroptimizationmethods,weintegrateSLOPEwithFlashAttention- 2 [ 24] and show that these approaches are orthogonal in practice and can boost the performance of the model separately.Table B.3summarizes the speedup achieved with and without SLOPE or FlashAttention-2. As it can be observed, each of these methods can improve the speed of the model both in training and inference, and adding them together will increase the speedup even further. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE133 B.10 Comparison with dense models To compare the performance of sparse models with dense models of the same size, we have conducted an experiment with GPT2-Small, in which we have reduced the number of transformer blocks in the model to half of GPT2-Small. We call this new configuration GPT2-Half.Table B.4andTable B.5summarize the accuracy results for GPT2-Half on different zero-shot downstream tasks. It can be observed that SLOPE outperforms GPT2-Half on average, while dynamic sparse training methods, such as SR-STE perform worse than it. Additionally, it is clear that adding low-rank adapters to the model improves the accuracy of all sparse pretraining methods. Table B.4: Performance comparison across different GPT models, sparsity methods, and LoRA ranks on various tasks. E-SR-STE stands for Extended SR-STE. ModelMethod LoRA (r) MMLU Arc Challenge Open Book QA Average GPT2-Small Denser = 022.920.716.219.94 GPT2-Small SLOPEr = 023.019.316.019.43 GPT2-Small SLOPE r = 0.05% 23.019.416.219.53 GPT2-Small SLOPE r = 2.1% 23.019.316.419.57 GPT2-Small E-SR-STE r = 024.118.312.618.33 GPT2-Small E-SR-STE r = 0.05% 24.118.414.218.90 GPT2-Small E-SR-STE r = 2.1% 24.218.314.218.90 GPT2-HalfDenser = 022.919.516.019.47 B.11 Zero-shot GLUE results for GPT We have tested the accuracy of the models on the zero-shot GLUE tasks in Language Model Evaluation Har- ness [41].Table B.5summarizes the achieved GLUE results by different models. It can be seen that SLOPE outperforms SR-STE and GPT2-Half on average. Additionally, SR-STE performs better than GPT2-Half in GLUE task. Table B.5: Performance comparison of GPT models using different sparsity methods and LoRA ranks on GLUE tasks. E-SR-STE stands for Extended SR-STE. ModelMethod LoRA (r) CoLA MNLI MNLI MRPC QNLI QQP RTE SST2 Avg (m) (m) GPT2-Small Denser = 0032.4 33.2 66.9 50.3 51.8 49.8 59.3 43.2 GPT2-Small SLOPEr = 0034.3 34.0 72.5 50.0 48.5 50.0 52.3 42.8 GPT2-Small SLOPE r = 0.05% 034.3 34.1 72.6 49.8 48.8 50.9 52.3 42.9 GPT2-Small SLOPE r = 2.1%034.3 34.0 71.6 50.0 49.0 52.0 52.6 43.1 GPT2-Small E-SR-STE r = 0033.6 33.9 57.1 50.7 50.4 55.2 54.7 42.5 GPT2-Small E-SR-STE r = 0.05% 033.1 33.6 57.9 51.0 50.5 55.4 55.0 42.6 GPT2-Small E-SR-STE r = 2.1%033.3 33.5 58.2 51.0 50.5 55.2 55.2 42.6 GPT2-HalfDenser = 00.0 33.9 33.8 53.6 51.1 47.7 56.7 50.6 41.1 APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE134 B.12 Extended SR-STE and FST implementation details Before we proceed with the details of Extended SR-STE and FST, we clarify the notations used in this chapter and the FST paper [62] inTable B.6 Table B.6: Description of Key Terms TermDescription Sparse PretrainingCommon notation used in SLoPe and FST, indicating the use of sparse weights during pretraining. Dense FinetuningNotationusedin theFSTpaper, indicatingan extendedpretraining phase. Downstream FinetuningPerformance after pretraining concludes, used to finetune the model for specific downstream tasks. FSTExtended pretraining technique focused on dense finetuning. Extended SR-STEVariation of sparse pretraining extended with additional fine- tuning. We compare SLOPE with FST exclusively fortraining speedups and memory savings. Comparing pretraining quality between SLOPE and FST is less meaningful because the final models produced by these methods differ significantly in the number of parameters. Specifically, FST produces a dense model after sparse pretraining and dense finetuning (99% sparse pretraining + 1% sparse + low-rank adaptation), while SLOPE produces a sparse model augmented with lightweight low-rank adapters. The number of parameters in the FST model is approximately 2×larger than in SLOPE, which makes a direct quality comparison imbalanced. We compare SLoPe with Extended SR-STE in terms of model quality, focusing on understanding the dynamics between static and dynamic masking under an equal number of parameters. This allows for a fair, ”apple-to-apple” comparison between the methods (iso-params). We refer to this method as ”Extended SR- STE” because, while the original SR-STE approach was designed for use with SGD, the FST paper extended it to support other optimizers. The FSTand Extended SR-STE code are available in Listing B.1andListing B.2respectively. 1defforward(ctx, x, weight, weight_sparse, weight_sparse_T, bias): 2ctx.save_for_backward(input, weight_sparse_T, bias) 3ctx.shape = x.shape 4x = x.view(-1, x.shape[-1]) 5output = torch.m(x, weight_sparse.t()) 6ifbiasisNone: 7returnoutput.view(*ctx.shape[:-1], -1) 8else: 9returnoutput.view(*ctx.shape[:-1], -1) + bias 10 11defbackward(ctx, grad_output): 12grad_output = grad_output.half() 13x, weight_T, bias = ctx.saved_tensors 14grad_input = grad_weight = grad_bias = None 15ifctx.needs_input_grad[0]: 16ifgrad_output.stride() == (0, 0, 0): 17grad_output = torch.ones_like(grad_output, device=grad_output.device , dtype=grad_output.dtype) APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE135 18grad_output = grad_output.view(-1, grad_output.shape[-1]) 19grad_input = torch.m(grad_output, weight_T.t()).view( 20ctx.shape) 21ifctx.needs_input_grad[1]: 22x = x.view(-1,input.shape[-1]) 23grad_output = grad_output.view(-1, grad_output.shape[-1]) 24grad_weight = torch.m(to_sparse_semi_structured(grad_output.t(), MVUE24 =True), x) 25ifctx.needs_input_grad[2]: 26grad_bias = grad_output.sum(0) 27returngrad_input, grad_weight, None, None, grad_bias Listing B.1: FST Algorithm 1defforward(ctx,input, weight, mask, weight_factor): 2sparse_weight = weight.clone().detach() 3sparse_weight[mask] = 0. 4ctx.save_for_backward(input, sparse_weight, weight_factor * mask * weight) 5output = torch.matmul(input, sparse_weight.t()) 6 7output = output.clone() 8returnoutput 9 10@staticmethod 11defbackward(ctx, grad_output): 12input, weight, weight_addition_term = ctx.saved_tensors 13input_shape =input.shape 14ifinput.dim() == 3: 15new_batch_size = input_shape[0] * input_shape[1] 16input=input.reshape(new_batch_size, -1) 17grad_output = grad_output.reshape(new_batch_size, -1) 18grad_output, grad_output_mask = prune_column_wise(grad_output) 19grad_weight = torch.matmul(grad_output.t(),input) 20grad_weight += weight_addition_term 21 22grad_input = torch.matmul(grad_output, weight) 23grad_input = grad_input.reshape(input_shape) 24returngrad_input, grad_weight, None Listing B.2: Extended SR-STE Algorithm. The weights are stored as dense and are pruned on-the-fly. B.13 Comparison of Depth and Width Pruning Depth pruning refers to reducing the number of layers in a model, while width pruning means reducing the size of the weights inside each layer in the model. We have conducted an experiment with depth and width pruning on LLaMA-2-7B [ 144] and Gemma-2-2B and Gemma-2-9B [140] to compare the effects of depth and width pruning on the performance of the models. The configurations used for the models are summarized APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE136 inTable B.7,Table B.9,Table B.8. Similar to [62], we reduced the aspect ratio of the Up-Sample and Down- Sample modules to half. Please note that this mechanism gives an advantage to width pruning methods, as the number of parameters in the Self-Attention modules remain intact. Table B.7: Model Configurations for LLaMA-2 7B Pruning MethodAttributes Baseline base_emb_dim: 4096 base_num_query_heads: 32 base_num_kv_heads: 32 base_mlp_dim: 11008 base_num_decoder_layers: 32 head_dim: 128 Depth Pruning base_emb_dim: 4096 base_num_query_heads: 32 base_num_kv_heads: 32 base_mlp_dim: 11008 base_num_decoder_layers: 16 # half the number of layers head_dim: 128 Width Pruning base_emb_dim: 4096 base_num_query_heads: 32 base_num_kv_heads: 32 base_mlp_dim: 5504 # half the number of dimensions base_num_decoder_layers: 32 head_dim: 128 Preliminary retraining loss curves, as shown inFigure B.6suggest no significant difference between depth- pruning and width-pruning during pretraining. Interestingly, in some cases, depth-pruning appears to outper- form width-pruning. B.14 Proofs Theorem 4.3.1 Proof.Considering a matrix withN:Mcolumn-wise pruned sparsity pattern, we want to prune the matrix usingN:Msparsity pattern row-wise as well. Let’s define random variableXas the number of added non- zeros toMrow-wise consecutive elements andYas the number of non-zeros inMrow-wise consecutive elements. E[X] = M−N X i=1 P r[X=i]i(B.1) ReplacingP r[X=i] =P r[Y=N+i]inEquation B.1, we will getEquation B.2, where we used a change in dummy variablej=N+i. E[X] = M−N X i=1 P r[Y=N+i]i= M X j=N+1 P r[Y=j](j−N)(B.2) Considering the definition ofY, it can be inferred that random variableYhas binomial distribution with a success probability of N M . As a result Equation B.3shows the probability mass distribution ofY. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE137 Table B.8: Model Configurations for Gemma-9B Pruning MethodAttributes Baseline base_emb_dim: 3584 base_num_query_heads: 16 base_num_kv_heads: 8 base_mlp_dim: 14336 base_num_decoder_layers: 20 # merged local and global attention head_dim: 256 Depth Pruning base_emb_dim: 3584 base_num_query_heads: 16 base_num_kv_heads: 8 base_mlp_dim: 14336 base_num_decoder_layers: 10 # half the merged layers head_dim: 256 Width Pruning base_emb_dim: 3584 base_num_query_heads: 16 base_num_kv_heads: 8 base_mlp_dim: 7168 # half the number of dimensions base_num_decoder_layers: 20 # merged local and global attention head_dim: 256 P r[Y=j] = M j s j (1−s) M−j ;s≜ N M (B.3) By replacingEquation B.3inEquation B.2, we will getEquation B.4. E[X] = M X j=N+1 M j s j (1−s) M−j (j−N)(B.4) Let’s define random variableZas the added sparsity ratio to the matrix by the extra pruning. SinceXwas the number of added non-zeros inMconsecutive elements,E[Z] = 1 M E[X], and hence: E[Z] =D(A R )−D(A R,C ) = M X j=N+1 M j s j (1−s) M−j j−N M (B.5) Theorem 4.3.2 Proof.In an optimization problem, we are aiming to find the optimal solution toEquation B.6. min W i E X [L(X, W i )](B.6) When using backpropagation, which is based on the chain rule in derivation, we compute the gradient in Equation B.7. E X [∇ X i L(X, W i )] =E X [∇ Y i LW](B.7) Let’s define random variableMas a uniformly random mask of 0’s and 1’s. The mask will be1at each point with a probability of N M . Let’s defineO≜E[M].Ois a matrix of all N M ’s. As a resultO⊙W= N M W. APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE138 Table B.9: Model Configurations for Gemma-2B Pruning MethodAttributes Baseline base_emb_dim: 2304 base_num_query_heads: 8 base_num_kv_heads: 4 base_mlp_dim: 9216 base_num_decoder_layers: 12 # merged local and global attention head_dim: 256 Depth Pruning base_emb_dim: 2304 base_num_query_heads: 8 base_num_kv_heads: 4 base_mlp_dim: 9216 base_num_decoder_layers: 6 # half the merged layers head_dim: 256 Width Pruning base_emb_dim: 2304 base_num_query_heads: 8 base_num_kv_heads: 4 base_mlp_dim: 4608 # half the number of dimensions base_num_decoder_layers: 12 # merged local and global attention head_dim: 256 E X [∇ Y i LW] =E X [∇ Y i L( M N O⊙W)] =E X [∇ Y i L( M N E M M⊙W)](B.8) By using the linearity of derivation and expectation operators, we can get the result inEquation B.9, which proves the theorem. E X [∇ X i L(X, W i )] = M N E M [E X [∇ Y i L(M⊙W)]](B.9) APPENDIX B. SUPPLEMENTARY MATERIAL FOR SLOPE139 0100002000030000 Step 5.0 7.5 10.0 12.5 Loss Gemma-2-2B Loss Dense Depth Pruning Width Pruning 0200040006000 Step 6 8 10 12 Loss Gemma-2-9B Loss Dense Depth Pruning Width Pruning 05000100001500020000 Step 4 6 8 10 Loss LLaMA-2-7B Loss Dense Depth Pruning Width Pruning Figure B.6: Comparison of the loss of depth and width pruning methods. Appendix C Supplementary Material for OPTIMA ThischapterprovidessupplementarymaterialfortheOPTIMAchapter.AppendixC.1evaluatesthesensitivity of OPTIMA to the size of the calibration dataset. C.1 Calibration dataset size sensitivity Similar to previous work (SparseGPT, Wanda, Thanos), OPTIMA leverages a set of calibration data from the C4 dataset to prune the models. Figure C.1shows the perplexity of LLaMA-3.2-1B on WikiText2 dataset when pruning the models with various numbers of calibration samples. Our results indicate that unlike the other methods (Wanda and SparseGPT) that have stochastic behavior as the number of samples increases, OPTIMA shows consistent improvement in model quality. However, the improvements are not significant, suggesting robustness to dataset size. 140 APPENDIX C. SUPPLEMENTARY MATERIAL FOR OPTIMA141 3264128256 Number of Calibration Samples 10 12 14 16 18 20 22 24 LLaMA-3.2-1B Perplexity Calibration Data Sensitivity Analysis Wanda+OPTIMA Wanda SparseGPT Dense Figure C.1: Sensitivity analysis for the number of calibration samples for different pruning methods. Appendix D Supplementary Material for PATCH This chapter provides supplementary material for the PATCH chapter.Appendix D.1describes the integra- tion with STOICC.Appendix D.2reports per-task accuracy results, andAppendix D.3explores tile transfer learning across different models. D.1 STOICC Integration Triton [143] enables developers to write efficient GPU kernels with a Python-like syntax, but it natively sup- ports only dense matrix operations and cannot handle sparsity. To accelerate the mixed-tile format produced by PATCH, we employ the STOICC compiler [122]. STOICC extends Triton with a sparse code-generation backend that allows tiles within a matrix to be either dense or sparse, enabling mixed execution within a single matrix multiplication. We rely on STOICC’s inspector to autotune both tile sizes and execution schedules (i.e., alternative kernel execution schemes such as split-Kparallelism) for the prefill and decoding stages of LLM inference. Matrix compression and metadata generation are determined by the chosen tile size, which must remain consistent across both stages. To address this, we first autotune the decoding stage, which is the primary bottleneck of autoregressive generation, since it is executed once per generated token (e.g., 128 times for 128 new tokens), unlike the single pass of prefill. The optimal tile size identified for decoding is then fixed and reused for prefill, where we perform a second round of autotuning over the remaining independent parameters. In contrast, for fully 2:4 sparse matrices, compression is independent of the block size, so they can be autotuned in the same way as dense kernels in Triton without this coupling constraint. The pseudocode outlining this process, including the handling of dense, fully 2:4 sparse, and mixed- sparsity modules, is provided in Listing D.1. 1def tune_and_convert_model(M,backend_name): 2//backend_name∈"STOICC","cuSPARSELt" 32_4_backend=select_2_4_backend(backend_name) 4 5//createallconfigs&schedulestotuneover 6base_configs=STOICC.create_configs() 7inspector=Inspector() 8 9foreachmoduleinM: 142 APPENDIX D. SUPPLEMENTARY MATERIAL FOR PATCH143 10s=get_sparsity_ratio(module.weight) 11 12//KeepdenseTorch(cuBLAS)module 13ifs== 0: 14continue 15 16//UseSTOICCorcuSPARSELtforfully2:4 17elifs== 0.5: 18c=2_4_backend.compress(module.weight) 19new_module=2_4_backend.create_module(c) 20replace(module,new_module) 21continue 22 23else: 24decoding_input=Tensor(BS,module.weight.shape[1]) 25prefill_input=Tensor(BS*SL,module.weight.shape[1]) 26 27//Tuneondecodinginputfirst 28inspector.set_configs(base_configs) 29best_cfg_dec=inspector.inspect( 30decoding_input, 31module.weight, 32isASparse=False) 33BN=best_cfg_dec["BLOCK_N"] 34BK=best_cfg_dec["BLOCK_K"] 35 36//Tuneonprefillusingdecodingtilesizes 37prefill_cfg=STOICC.create_configs(BLOCK_N=BN, BLOCK_K=BK) 38inspector.set_configs(prefill_cfg) 39best_cfg_pre=inspector.inspect( 40prefill_input, 41module.weight, 42isASparse=False) 43 44c=inspector.compress(module.weight,BN,BK) 45mixed_module=MixedModule(c,best_cfg_dec,best_cfg_pre) 46replace(module,mixed_module) 47 48returnM PseudoCode D.1: Tuning and Converting Model Weights to Mixed Format. Table D.1reports the measured throughput (tokens processed per second) of LLaMA-2 7B at sparsity levels of 45%, 35%, and 25% with a batch size of 16 on an A6000 GPU. To reduce CPU overhead from launching Triton kernels in PyTorch, we executed generation through CUDA graphs, capturing both the prefill anddecodingstages. Withsparsityratiosbetween25%and45%, ourheterogeneousapproachachieves1.18×– 1.38×end-to-end acceleration over the dense baseline. We also report timings on A100 in Table D.2. APPENDIX D. SUPPLEMENTARY MATERIAL FOR PATCH144 Table D.1: Throughput of LLaMA-2 7B with mixed sparsity compared to the dense model. Measurements taken on an A6000 GPU with batch size 16. Throughput is reported in tokens processed/sec. Sparsity Prefill length Tokens generated Throughput (tok/s) Speedup vs. dense 0%1281281023.801.00× 25%1281281212.791.18× 35%1281281304.461.27× 45%1281281410.201.38× 0%1281024435.421.00× 25%1281024493.331.13× 35%1281024515.391.18× 45%1281024542.871.25× Table D.2: Throughput of LLaMA-2 7B with mixed sparsity compared to the dense model. Measurements taken on an A100 GPU with batch size 16. Throughput is reported in tokens processed/sec. Sparsity Prefill length Tokens generated Throughput (tok/s) Speedup vs. dense 0%1281281876.241.00× 25%1281282002.021.07× 35%1281282088.981.11× 45%1281282180.881.16× 0%1281024812.551.00× 25%1281024864.661.06× 35%1281024885.901.09× 45%1281024907.121.12× D.2 Per Task Results This appendix provides detailed per-task accuracy results for the models evaluated inAppendix 6.7, covering eight zero-shot downstream tasks: MMLU, PIQA, ARC-Easy, ARC-Challenge, Winogrande, OpenBookQA, RACE, and HellaSwag. The results are presented for each model at various sparsity levels and pruning meth- ods, including our proposed PATCH Joint and PATCH Tile variants, alongside baseline methods such as Magni- tude, Wanda, SparseGPT,Thanos, ProxSparse, andMaskLLM.Thesetablescomplementtheaverageaccuracy and perplexity results reported in Table 6.2andTable 6.3of this chapter, offering a granular view of model performance across individual tasks. For smaller models (Qwen-2.5 0.5B, LLaMA-3.2 1B, and Gemma-3 1B), we report results using the PATCH Joint variant, which jointly optimizes dense tile locations and sparsity patterns within sparse tiles. For larger models (LLaMA-2 7B and LLaMA-3.1 8B), we report results using the memory-efficient PATCH Tile variant, which optimizes dense tile selections with a fixed 2:4 sparsity mask. The per-task accuracies high- light the effectiveness of our approaches in maintaining robust performance across diverse tasks, even at high sparsity levels, compared to baseline methods. The following tables detail the per-task accuracies for each model: •Qwen-2.5 0.5B:Table D.3presents the per-task accuracies for the PATCH Joint variant and baselines at 0% and 50% sparsity, with PATCH Joint evaluated at 25%, 35%, and 45% sparsity. •LLaMA-2 7B:Table D.4shows the per-task accuracies for the PATCH Tile variant and baselines, with PATCH Tile evaluated at 25%, 35%, and 45% sparsity. APPENDIX D. SUPPLEMENTARY MATERIAL FOR PATCH145 •LLaMA-3.1 8B:Table D.5provides the per-task accuracies for the PATCH Tile variant and baselines, with PATCH Tile at 25%, 35%, and 45% sparsity. •LLaMA-3.2 1B:Table D.6reports the per-task accuracies for the PATCH Joint variant and baselines, with PATCH Joint at 25%, 35%, and 45% sparsity. •Gemma-3 1B:Table D.7details the per-task accuracies for the PATCH Joint variant and baselines, with PATCH Joint at 25%, 35%, and 45% sparsity. These results enable a deeper analysis of the task-specific performance trends, demonstrating the flexi- bility and robustness of PATCH Joint and PATCH Tile in achieving high accuracy across diverse tasks while maintaining hardware-friendly sparsity patterns. Table D.3: Model quality (task accuracy across eight zero-shot tasks, reported in %) for Qwen-2.5 0.5B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff. Sparsity Method PatternMMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Avg 0%Dense-47.71 70.24 64.48 29.52 56.20 24.20 35.02 40.63 46.00 50% Magnitude 2:423.00 54.24 31.23 19.20 49.96 13.60 23.44 26.59 30.16 Wanda2:424.43 58.71 43.18 17.75 51.62 12.20 26.32 29.58 32.97 SparseGPT 2:422.93 60.77 46.60 20.82 52.88 14.00 29.57 30.93 34.81 Thanos 2:422.97 60.17 45.37 19.20 53.59 15.20 31.00 31.31 34.85 ProxSparse 2:423.00 57.34 40.53 18.26 48.62 14.00 25.65 29.02 32.05 MaskLLM 2:425.11 67.03 56.57 23.98 52.57 20.20 33.30 35.90 39.33 45%PATCH Joint Dense/2:4 Tiles27.3968.4459.1325.7753.6719.8032.1535.9940.29 35%PATCH Joint Dense/2:4 Tiles29.0468.8860.4026.3755.0920.4032.4436.5841.15 25%PATCH Joint Dense/2:4 Tiles30.8969.1562.7929.1055.3320.0034.1637.7142.39 Table D.4: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-2 7B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality trade- off. Sparsity Method PatternMMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Avg 0%Dense-41.82 78.07 76.35 43.52 69.06 31.40 39.52 57.13 54.61 50% Magnitude 2:425.82 70.02 61.78 30.12 61.01 21.80 31.48 45.45 43.44 Wanda2:425.80 71.00 63.80 30.29 61.09 25.20 35.50 41.75 44.30 SparseGPT 2:426.17 70.73 63.80 30.63 65.04 24.00 37.13 43.18 45.09 Thanos 2:425.27 70.78 63.43 30.97 64.56 23.80 36.46 43.11 44.80 ProxSparse 2:426.77 71.60 65.70 33.02 62.90 24.20 35.31 47.84 45.92 MaskLLM 2:427.65 74.76 69.44 35.58 65.04 26.80 38.56 51.15 48.62 45%PATCH Tile Dense/2:4 Tiles27.2875.4170.1635.8465.2727.6038.7651.6148.99 35%PATCH Tile Dense/2:4 Tiles29.9376.7170.8836.9565.6728.2039.3352.9650.08 25%PATCH Tile Dense/2:4 Tiles32.3376.9972.8138.5768.2729.8039.5254.3451.58 APPENDIX D. SUPPLEMENTARY MATERIAL FOR PATCH146 Table D.5: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.1 8B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff. Sparsity Method PatternMMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Avg 0%Dense-63.57 80.09 81.44 51.37 73.48 33.40 39.14 60.02 60.31 50% Magnitude 2:423.06 63.82 45.33 25.94 53.91 15.20 26.70 33.49 35.93 Wanda2:427.85 68.88 58.33 26.71 60.93 19.00 33.78 38.70 41.77 SparseGPT 2:431.82 70.46 63.85 31.74 64.56 21.60 37.22 42.99 45.53 Thanos 2:434.23 70.40 63.13 31.40 63.61 23.20 37.03 42.75 45.72 ProxSparse 2:429.89 71.71 62.63 33.28 58.56 23.80 35.22 46.03 45.14 MaskLLM 2:442.47 77.04 73.15 40.19 68.43 28.80 38.28 54.04 52.80 45%PATCH Tile Dense/2:4 Tiles47.3277.9673.6141.8968.0329.0036.5654.4453.60 35%PATCH Tile Dense/2:4 Tiles51.1577.9776.1442.4169.4631.4038.1855.5455.28 25%PATCH Tile Dense/2:4 Tiles52.9577.7577.5744.6270.5631.8039.9056.6956.48 Table D.6: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.2 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff. Sparsity Method PatternMMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Avg 0%Dense-37.57 74.54 65.53 31.32 60.62 26.40 37.89 47.76 47.70 50% Magnitude 2:423.31 53.81 27.74 18.94 51.38 11.80 24.02 26.26 29.66 Wanda2:422.90 58.11 37.08 19.20 49.09 13.20 25.17 28.11 31.61 SparseGPT 2:422.93 61.43 45.03 22.35 54.93 15.80 29.86 32.08 35.55 Thanos 2:423.12 62.40 44.91 21.76 54.30 16.00 31.10 32.09 35.71 ProxSparse 2:422.96 60.83 39.44 20.31 51.54 16.80 25.17 31.37 33.55 MaskLLM 2:426.28 69.10 57.41 25.85 55.48 21.40 32.82 39.94 41.04 45%PATCH Joint Dense/2:4 Tiles23.8170.8960.7727.2256.2722.8034.0740.7842.08 35%PATCH Joint Dense/2:4 Tiles25.1371.3260.2729.1857.0622.0034.6442.1742.72 25%PATCH Joint Dense/2:4 Tiles28.5971.4461.5728.6758.2523.2035.2243.5243.81 APPENDIX D. SUPPLEMENTARY MATERIAL FOR PATCH147 Table D.7: Model quality (accuracy across eight zero-shot tasks) for Gemma-3 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff. Sparsity Method PatternMMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Avg 0%Dense-24.95 75.03 71.84 34.90 58.64 28.60 34.83 47.26 47.01 50% Magnitude 2:423.08 59.79 37.29 17.66 50.59 14.00 22.87 27.97 31.66 Wanda2:423.96 59.52 48.02 18.34 51.22 14.20 27.85 30.18 34.16 SparseGPT 2:423.62 62.79 49.83 19.03 51.54 15.20 30.62 31.99 35.58 Thanos 2:423.44 62.24 48.86 18.34 50.12 15.60 30.81 31.28 35.09 ProxSparse 2:423.10 64.25 50.72 21.59 53.43 18.00 29.09 32.86 36.63 MaskLLM 2:425.03 69.91 60.27 27.65 56.27 21.20 34.55 39.84 41.84 45%PATCH Joint Dense/2:4 Tiles23.5471.6563.9727.4757.3023.6033.4941.3942.80 35%PATCH Joint Dense/2:4 Tiles25.3872.3163.8027.3956.6724.0034.7442.0743.30 25%PATCH Joint Dense/2:4 Tiles25.4571.8766.1630.5557.8522.8034.5543.3344.07 D.3 Tile Transfer Learning We also test whether initializing tile logits with priors from one-shot pruning methods improves performance, as done in MaskLLM [ 33]. In our case, the initialization is derived from one-shot pruning with unstructured sparsity. We initialize tiles that retain more nonzeros after unstructured pruning with positive logits (favoring dense assignment), while the remaining tiles receive negative logits, controlled by a strength parameter. The number of tiles initialized as dense is selected such that the overall layer-wise sparsity target is satisfied. As shown in Table D.8, the choice of prior has little impact on final performance: all priors yield nearly iden- tical perplexity, with random initialization often performing best. This is likely because the global sparsity target enables dynamic reallocation of sparsity across layers during training, overriding the effect of any fixed initialization. For consistency with prior work, we adopt SparseGPT initialization in all experiments. Table D.8: Perplexity (↓) under different tile prior initializations. All priors yield nearly identical performance, suggesting that the global sparsity target allows dynamic reallocation of sparsity during training, overriding the influence of fixed initialization. Sparsity (0.5B) Nothing SparseGPT Wanda Magnitude Random 45%14.8014.5714.5014.4814.51 35%13.9713.8413.8713.8513.79 25%13.4713.4713.3713.4413.33 Appendix E Supplementary Material for SLIM This chapter provides supplementary material for the SLIM chapter. We begin by defining the notations used throughout this work inAppendix E.1.Appendix E.2evaluates input quantization, andAppendix E.3presents additional fine-tuning results. Language modeling experiments are reported in Appendix E.4, with additional sparse and quantized accuracy results inAppendix E.5.Appendix E.6compares sparsity and quantization, whileAppendix E.7provides additional speedup results.Appendix E.8details the fine-tuning costs, and theoretical analyses of memory and computation reductions are provided inAppendix E.9andAppendix E.10, respectively.Appendix E.11reports compression costs,Appendix E.12explores the impact of rank choices on low-rank adapters, and Appendix E.14examines sensitivity to the calibration dataset.Appendix E.15analyzes the effects of different sparsity ratios, and Appendix E.16discusses the challenges of group quantization. E.1 Notations Table E.1details the key notations, particularly forAppendix 7.5. Table E.1: Key notation definitions used in the experimental results (Appendix 7.5). TermDescription Naive-LoRAA one-shot low-rank adapter that minimizes the norm of the difference be- tween the original and the compressed weights. SLIM-LoRAA saliency-based one-shot low-rank adapter that minimizes the saliency of the difference between the original and the compressed weights. Q(Superscript) Q indicates that the compression method quantizes the low-rank adapters as well. + FT+ FT shows a short fine-tuning phase on 300,000 tokens from the C4 dataset. E.2 Input Quantization We evaluate SLIM with 8-bit input quantization to assess its impact on accuracy. We use AbsMax uniform quantization with a single parameter per input tensor and apply FP8 format [96] for weight quantization. The 148 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM149 choice between E4M3 and E5M2 depends on the tensor’s maximum value; if it exceeds E4M3’s range, we switch to E5M2 for greater expressivity. Next, we examine how input quantization affects model accuracy. Table E.2presents accuracy results for different SLIM variants with input quantization. A comparison withTable E.3, which reports accuracy without input quantization, reveals minimal accuracy loss, demon- strating SLIM’s robustness. For further validation, we extend these experiments to language modeling tasks (Appendix E.4). Table E.2: Average zero-shot accuracy of LLaMA-2 and OPT models with4-bit weight and 8-bit input quantization.↑indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 50% 2:4 SLIM-LoRASLIM-Quant 34.85 34.27 40.29 42.58 45.78 46.21 50.99 54.66 SLIM-LoRA + FT SLIM-Quant 35.28 34.33 41.14 43.29 46.44 47.33 51.77 56.28 SLIM-LoRA Q SLIM-Quant 34.30 33.85 39.92 41.99 46.08 45.94 50.70 53.56 SLIM-LoRA Q + FT SLIM-Quant 34.92 34.80 41.66 43.69 46.03 46.87 50.26 56.28 50% Unstructured SLIM-LoRASLIM-Quant 35.12 34.86 41.94 43.53 47.27 47.70 54.28 57.82 SLIM-LoRA + FT SLIM-Quant 35.18 35.30 42.37 44.02 47.01 48.52 54.43 57.70 SLIM-LoRA Q SLIM-Quant 35.26 34.67 41.48 43.46 47.25 47.76 53.91 57.16 SLIM-LoRA Q + FT SLIM-Quant 35.52 35.31 42.66 44.50 47.08 48.53 53.23 57.55 E.3 Additional fine-tuning results To complement the results inAppendix 7.5, we provide accuracy measurements for PEFT-based fine-tuning of low-rank adapters on the OPT and LLaMA-2 model families in Table E.3while showing the accuracy results withoutfine-tuningforcomparison. Theresultsconfirmthepreviouslyobservedtrend: lightweightfine-tuning enhances the accuracy of all baselines, with SLIM-LoRA achieving the most significant improvements due to its saliency-based design. E.4 Language modeling experiments We evaluate all benchmarks fromAppendix 7.5andAppendix E.2on the WikiText2 language modeling task. Table E.4andTable E.6show perplexity results for 4-bit quantized models with 2:4 and unstructured sparsity, respectively.Table E.5summarizes the results for 8-bit input quantization. To examine sparsity and quanti- zation independently, Table E.7andTable E.8report results for pruning-only and quantization-only models. Consistent withAppendix 7.5, SLIM achieves superior performance across all settings. APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM150 Table E.3: Effects of fine-tuning on the average zero-shot accuracy of LLaMA-2 and OPT models with 50% sparsity and 4-bit weight quantization.↑indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 50% 2:4 Naive-LoRASLIM-Quant 34.28 33.38 38.36 41.21 44.91 45.25 48.45 51.94 Naive-LoRA + FT SLIM-Quant 34.41 34.70 39.72 42.88 46.16 46.76 50.89 55.70 SLIM-LoRASLIM-Quant 34.62 34.36 40.61 42.73 45.99 46.09 51.15 54.94 SLIM-LoRA + FT SLIM-Quant35.0334.58 41.11 43.3546.71 47.25 52.12 56.60 SLIM-LoRA Q SLIM-Quant 34.43 34.30 40.11 42.37 46.33 46.24 51.02 53.55 SLIM-LoRA Q + FT SLIM-Quant 34.9234.85 41.84 43.8746.31 46.91 48.31 56.50 50% Unstructured Naive-LoRASLIM-Quant 34.77 34.23 40.40 43.37 46.64 47.30 51.52 55.33 Naive-LoRA + FT SLIM-Quant 35.70 35.47 41.89 44.16 47.08 47.78 52.90 57.08 SLIM-LoRASLIM-Quant 35.20 35.32 41.85 43.48 47.08 47.96 54.26 57.85 SLIM-LoRA + FT SLIM-Quant 35.5935.7142.3744.58 47.6948.2654.69 57.96 SLIM-LoRA Q SLIM-Quant35.3535.13 41.74 43.63 47.16 47.86 54.18 57.33 SLIM-LoRA Q + FT SLIM-Quant 35.65 35.6742.7444.54 47.4848.4053.57 57.78 Table E.4: Perplexity of LLaMA-2 and OPT models with2:4 sparsity and 4-bit weight quantizationon WikiText-2 dataset language modeling task.↓indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-27.66 22.00 14.62 12.47 10.86 10.13 5.12 4.57 MagnitudeGroup AbsMax 5.1E2 4.4E2 1.2E3 1.3E3 3.6E2 4.9E2 86.34 8.98 SparseGPTGroup OPTQ 78.18 59.86 27.36 18.62 15.31 13.25 15.01 8.97 WandaGroup AbsMax 1.8E2 1.3E2 32.76 24.48 17.29 16.86 13.46 8.70 WandaAWQ9.3E1 8.1E5 29.56 22.91 16.28 16.72 12.79 OOM WandaOmniQuant9.7E1 NaN 33.61 25.89 19.09 OOM 12.77 OOM WandaAffineQuant9.7E1 NaN 30.32 1.6E3 16.85 OOM 12.21 OOM JSQJSQ3.5E3 2.5E4 67.36 3.2E3 22.50 5.5E2 11.69 8.05 Naive-LoRAGroup AbsMax 69.23 50.02 20.52 16.05 12.83 13.12 8.04 6.38 Naive-LoRASLIM-Quant 83.08 58.69 27.06 20.92 14.29 13.20 8.19 7.09 Naive-LoRA + FT SLIM-Quant 51.82 38.84 20.59 16.19 13.13 12.55 6.96 6.01 SLIM-LoRASLIM-Quant 57.91 50.09 19.64 15.65 12.71 12.13 7.56 6.50 SLIM-LoRA + FT SLIM-Quant 44.03 37.32 18.25 14.89 12.68 12.066.706.60 SLIM-LoRA Q SLIM-Quant 53.09 46.96 19.62 16.01 12.48 12.15 7.75 6.96 SLIM-LoRA Q + FT SLIM-Quant42.80 37.39 18.38 15.40 12.65 12.357.086.36 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM151 Table E.5: Perplexity of LLaMA-2 and OPT models with4-bit weight and 8-bit input quantization.↓ indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-27.66 22.00 14.62 12.47 10.86 10.13 5.12 4.57 50% 2:4 SLIM-LoRASLIM-Quant 48.4 49.6 16.6 16.2 12.9 12.3 7.2 6.5 SLIM-LoRA + FT SLIM-Quant 39.8 37.5 18.3 15.5 12.8 12.1 6.6 5.8 SLIM-LoRA Q SLIM-Quant 54.2 50.8 20.8 16.8 13.0 12.4 7.8 7.0 SLIM-LoRA Q + FT SLIM-Quant 43.4 39.1 19.3 16.0 13.1 12.6 7.1 5.8 50% Unstructured SLIM-LoRASLIM-Quant 36.8 31.1 16.8 14.0 11.7 10.9 6.1 5.4 SLIM-LoRA + FT SLIM-Quant 33.8 28.6 16.5 14.0 12.0 11.5 5.9 5.2 SLIM-LoRA Q SLIM-Quant 39.5 31.3 17.3 14.2 11.8 10.9 6.3 5.6 SLIM-LoRA Q + FT SLIM-Quant 35.6 29.1 17.0 14.3 12.2 11.7 6.2 5.5 Table E.6: Perplexity of LLaMA-2 and OPT models withunstructured sparsity and 4-bit weight quantiza- tionon WikiText-2 dataset language modeling task.↓indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-27.66 22.00 14.62 12.47 10.86 10.13 5.12 4.57 MagnitudeGroup AbsMax 3.2E2 1.1E2 3.2E3 3.6E2 7.2E2 5.4E3 17.18 6.77 SparseGPTGroup OPTQ 42.60 34.19 21.41 14.30 12.15 11.26 8.28 5.92 WandaGroup AbsMax 62.64 39.60 19.93 15.01 12.31 12.46 6.80 5.75 WandaAWQ42.49 3.8E5 18.80 14.67 12.17 12.34 7.28 OOM WandaOmniQuant43.55 NaN 20.58 15.82 13.29 OOM 7.40 OOM WandaAffineQuant43.66 NaN 19.40 14.94 12.39 OOM 7.21 OOM JSQJSQ3.2E3 1.9E4 23.88 2.3E2 15.13 1.2E5 6.63 5.73 Naive-LoRAGroup AbsMax 40.37 30.99 17.02 13.91 11.68 11.38 6.12 5.28 Naive-LoRASLIM-Quant 46.66 33.90 19.46 15.36 12.16 11.41 6.56 5.58 Naive-LoRA + FT SLIM-Quant 38.05 29.27 17.52 14.39 12.28 11.84 6.10 5.28 SLIM-LoRASLIM-Quant 39.62 31.51 16.52 13.65 11.42 10.82 6.16 5.36 SLIM-LoRA + FT SLIM-Quant34.92 28.67 16.16 13.6611.83 11.475.36 5.19 SLIM-LoRA Q SLIM-Quant 38.79 30.16 16.64 13.82 11.43 10.80 6.26 5.58 SLIM-LoRA Q + FT SLIM-Quant 35.17 28.31 16.46 13.9611.42 10.805.94 5.46 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM152 Table E.7: Perplexity of LLaMA-2 and OPT models with pruning on WikiText-2 dataset language modeling task. The quantization is disabled in this experiment.↓indicates better performance. Pruning/LoRAOPTLLaMA-2 Method125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense27.66 22.00 14.62 12.47 10.86 10.13 5.12 4.57 2:4 Sparsity Magnitude341.5 417.1 427.2 1.2E3 264.1 4.0E4 9.1E4 2.0E5 SparseGPT60.7 50.7 23.8 17.2 14.1 12.9 10.2 8.3 Wanda81.6 116.0 27.8 21.4 16.0 16.4 12.0 8.5 Naive-LoRA46.9 45.0 18.8 15.2 12.5 12.9 8.1 6.5 Naive-LoRA + FT 39.6 35.115.016.3 12.7 12.3 6.55.7 SLIM-LoRA45.2 43.6 18.6 15.012.412.6 7.3 6.2 SLIM-LoRA + FT37.1 33.717.014.212.412.1 6.45.8 50% Unstructured Magnitude193.4 97.8 1.7E3 265.2 968.7 2.4E4 9.9E4 1.1E5 SparseGPT36.7 31.8 17.6 13.4 11.5 11.1 6.5 5.6 Wanda39.3 36.4 18.3 14.3 12.0 12.3 6.4 5.4 Naive-LoRA33.3 29.1 16.3 13.5 11.5 11.2 6.2 5.4 Naive-LoRA + FT 31.9 27.5 16.3 13.8 12.0 11.6 5.8 5.1 SLIM-LoRA32.7 29.0 15.9 13.211.2 10.85.9 5.2 SLIM-LoRA + FT31.0 26.8 15.5 13.111.6 11.05.8 4.7 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM153 Table E.8: Perplexity of LLaMA-2 and OPT models with quantization on WikiText-2 dataset language mod- eling task. The sparsity is disabled in this experiment.↓indicates better performance. QuantizationLow-rankOPTLLaMA-2 MethodAdapter125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-27.66 22.00 14.62 12.47 10.86 10.13 5.12 4.57 OPTQ-33.0 24.4 16.0 13.0 11.3 10.3 6.1 4.9 AWQ-29.1 2.7E5 14.9 12.7 11.0 10.2 6.0 OOM OmniQuant-30.2 NaN 15.8 13.3 11.6 OOM 5.7 OOM AffineQuant-28.7 NaN 14.9 12.6 11.0 OOM 5.7 OOM Group AbsMax -35.1 23.3 15.5 12.9 11.1 10.3 5.4 4.7 Group AbsMax Naive-LoRA30.4 22.9 15.1 12.7 11.0 10.2 5.3 4.7 Group AbsMax SLIM-LoRA29.3 22.815.012.7 10.910.25.2 4.7 SLIM-Quant -1.4E3 26.0 1.7E3 33.1 31.0 6.7E2 1.3E5 7.8E4 SLIM-Quant Naive-LoRA32.1 24.1 15.6 13.4 11.2 10.5 5.4 4.8 SLIM-Quant SLIM-LoRA30.8 23.115.212.9 11.1 10.3 5.4 4.8 SLIM-Quant SLIM-LoRA + FT 30.7 23.5 15.3 13.3 11.610.05.3 4.7 E.5 Additional Sparse and Quantized Results In Appendix7.5, weprovidedtheaccuracyresultsfordifferentpruningandquantizationmethods. Whenusing Wanda for pruning, we only reported the best quantization method out of Group AbsMax, AWQ, OmniQuant, and AffineQuant. For completeness, we have provided the accuracy achieved by each of these quantization methods separately in Table E.9. Methods like OmniQuant and AffineQuant encounter difficulties in quantizing OPT-350M, resulting in NaN values. Additionally, approaches such as AWQ, OmniQuant, and AffineQuant cause memory issues (OOM) when attempting to compress the models on a single A100-40GB GPU. APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM154 Table E.9: Average zero-shot accuracy of LLaMA-2 and OPT models with2:4 sparsity and 4-bit weight quantization.↑indicates better performance. Pruning/LoRAWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Dense-35.9 37.1 43.4 45.5 48.3 48.7 56.6 60.8 2:4 Sparsity WandaGroup AbsMax 33.27 32.79 37.47 39.45 42.95 43.64 43.89 48.94 WandaAWQ33.33 31.50 38.43 40.00 43.41 44.07 44.86 OOM WandaOmniQuant33.37 NaN 37.35 39.39 41.50 OOM 43.95 OOM WandaAffineQuant33.39 NaN 37.48 33.51 42.88 OOM 44.62 OOM 50% Unstructured WandaGroup AbsMax 34.67 33.89 40.38 42.77 45.88 46.60 51.76 56.76 WandaAWQ35.11 31.57 41.02 42.89 46.52 46.84 50.68 OOM WandaOmniQuant34.85 NaN 39.84 42.16 44.67 OOM 50.51 OOM WandaAffineQuant34.64 NaN 41.23 42.68 46.05 OOM 53.62 OOM E.6 Sparsity vs. quantization A natural question that arises when compressing models is whether it is more efficient to reduce the model size through pruning or quantization. To answer this question, we conduct a set of experiments, which evaluate the perplexity of different models under three different conditions, all with around8×model size reduction factor: (1) 2-bit weight quantization with no sparsity, (2) 4-bit weight quantization with 50% unstructured sparsity, and (3) 4-bit weight quantization with 50% 2:4 sparsity. We have used SLIM-LoRA with SLIM- Quant in all the experiments. The accuracy and perplexity results of these experiments are summarized in Table E.10andTable E.11, showing that combining sparsity and quantization yields better results in compar- ison to quantization-only settings with lower bitwidth. Table E.10: Average zero-shot accuracy of different models using different pruning and quantization schemes. ↑indicates better performance. Combining sparsity and quantization provides better accuracy results in com- parison to solely using quantization. OPTLLaMA-2 QuantizationSparsity125M 350M 1.3B 2.7B 6.7B 13B 7B 13B 2-bit-33.5 32.5 38.5 39.2 43.8 44.4 42.4 44.9 4-bit2:434.6 34.4 40.6 42.7 46.0 46.1 51.2 54.9 4-bit50% Unstructured 35.2 35.3 41.9 43.5 47.1 48.0 54.3 57.9 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM155 Table E.11: Perplexity of different models on WikiText-2 dataset using different pruning and quantization schemes.↓indicates better performance. Combining sparsity and quantization provides better accuracy re- sults in comparison to solely using quantization. OPTLLaMA-2 QuantizationSparsity125M 350M 1.3B 2.7B 6.7B 13B 7B 13B 2-bit-116.2 169.7 35.1 27.1 16.2 15.0 12.5 11.7 4-bit2:447.5 45.6 18.8 15.7 12.4 12.1 7.2 6.5 4-bit50% Unstructured 36.3 29.9 16.3 13.7 11.4 10.8 6.0 5.4 E.7 Additional speedup results Appendix 7.5presents the speedup of SLIM on consumer-grade GPUs, while this section provides results on NVIDIA A100-40GB GPUs.Figure E.1summarizes the speedup for the LLaMA-2 and LLaMA-3.1 model families, including LLaMA-3.1-405B, highlighting SLIM’s scalability to large models. As with consumer- grade devices, larger models achieve higher speedups. However, for smaller models like LLaMA-2-7B, Sparse Marlin’s sparse quantized matrix multiplication kernels lead to a slowdown on A100 GPUs, which does not occur on consumer-grade GPUs and is not specific to SLIM . APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM156 0.0 0.5 1.0 1.5 2.0 LLaMA-2-7B 2.2 0.6 1.9 1.1 0.3 1.0 Batch Size 16 1.7 0.6 1.5 0.8 0.3 0.7 Batch Size 32 1.2 0.5 1.1 0.5 0.2 0.5 Batch Size 64 0.0 0.5 1.0 1.5 2.0 LLaMA-2-13B 2.0 1.3 1.9 1.6 0.8 1.4 1.8 1.1 1.6 0.9 0.4 0.7 1.5 0.8 1.3 0.6 0.3 0.5 0.0 0.5 1.0 1.5 2.0 2.5 LLaMA-2-70B 2.5 1.7 2.6 2.7 1.6 2.7 2.5 1.3 2.2 2.1 0.9 1.9 2.0 1.1 1.8 1.6 0.6 1.5 Up-Projection Self-Attention Down-Projection 0 1 2 3 LLaMA-3.1-405B 2.9 2.2 2.8 3.8 2.9 3.6 Up-Projection Self-Attention Down-Projection 2.6 2.1 2.7 3.2 2.4 3.2 Up-Projection Self-Attention Down-Projection 2.2 1.9 2.2 2.4 1.9 2.3 SLiM Speedup on A100-40GB FP16 LoRAINT4 LoRA Figure E.1: SLIM speedup for LLaMA-2 family of models on NVIDIA A100-40GB GPUs. E.8 Fine-tuning costs Fine-tuning compressed models can recover lost accuracy, but the high parameter count leads to substantial timeandmemorycosts. Inourexperiments, wefine-tunedmodelswithlow-rankadapters, wherethequantized weights are frozen and only the adapters are fine-tuned. This results in a more parameter-efficient approach, reducing both memory and computational costs. When no low-rank adapter is used, the straight-through estimator (STE) fine-tunes the quantized weights. Table E.12presents the fine-tuning results for 300,000 tokens from the C4 dataset, using a batch size of 64 and sequence length of 1024 on a single H100 GPU. Fine-tuning models without low-rank adapters took 12 hours for 125M parameter models and over 36 days for 13B parameter models. Given these high costs, APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM157 completing fine-tuning was challenging with our limited resources. In contrast, using low-rank adapters and freezing the sparse quantized weights made fine-tuning more efficient, enabling us to report accuracy results inTable 7.1. Table E.12: The required time for fine-tuning the models with a single H100 GPU on 300,000 tokens from the C4 dataset with a batch size of 64 and a sequence length of 1024. PruningWeightOPTLLaMA-2 MethodQuantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Magnitude Group AbsMax SparseGPT OPTQ12h 43h 164h 361h 866h 867h 842h 844h WandaGroup AbsMax SLIM-Naive SLIM-Quant 1.5h 3h6h 8h 16h 18h 14h 14h SLIM-LoRA SLIM-Quant E.9 Memory reduction analysis SLIMprunesandquantizesthemodelsandaddsadditionallow-rankadapterstothem. Additionally,itsupports quantization methods for the low-rank adapters to reduce their overheads. In the following, we provide an analysis of the memory reduction when using SLIM and other pruning and quantization methods. Assuming the hidden dimension of a model isdand the low-rank adapter ratio used in the model is of rankr <1. Furthermore, by denoting the number of transformer blocks withnand the vocabulary size of the model byVand by denoting the ratio of the up-projection and down-projection layers in the model bya, we can get the memory reduction as the ratio of Compressed Model Size Dense Model Size fromEquation E.1. Memory Reduction= n(4d 2 /2 + 4×2d 2 r+ 2d 2 a/2 + 2d(dr+dra)) +dV n(4d 2 + 2d 2 a) +dV (E.1) Table E.13summarizes the memory reduction of different pruning and quantization methods. Please note that when using low-rank adapters (in Naive-LoRA and SLIM-LoRA), we assume a rank ofr= 0.1. Table E.13: Theoretical memory reduction (×) of different compression methods across various OPT and LLaMA models. In Quantized SLIM , the low-rank adapters are also quantized.(↓indicates better perfor- mance.) CompressionOPTLLaMA-2 Method125M 350M 1.3B 2.7B 6.7B 13B 7B 13B SparseGPT + OPTQ0.40 0.30 0.25 0.17 0.15 0.14 0.15 0.14 Wanda + AbsMax0.40 0.30 0.25 0.17 0.15 0.14 0.15 0.14 Naive-LoRA + AbsMax0.50 0.42 0.38 0.31 0.30 0.29 0.31 0.30 SLIM-LoRA + SLIM-Quant 0.50 0.42 0.38 0.31 0.30 0.29 0.31 0.30 SLIM-LoRA Q + SLIM-Quant 0.42 0.33 0.28 0.20 0.19 0.18 0.19 0.18 APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM158 E.10 Computation reduction analysis SLIM and other compression methods reduce the number of floating point operations (FLOPs) at the inference of models. Additionally, the low-rank adapters used in SLIM and Wanda SVD can add additional computa- tional overheads to the inference of the models. Following JSQ [49], in this section, we provide an analysis of the FLOP reduction in the inference of different methods. It is noteworthy that even though quantization can reduce the memory overhead of models, since all the computations are done in floating point format, it does not lead to a reduction in the computation of the inference. Assuming the hidden dimension of a model isdand the low-rank adapter ratio used in the model is of rankr <1. Furthermore, by denoting the number of transformer blocks withnand the vocabulary size of the model byVand by denoting the ratio of the up-projection and down-projection layers in the model bya, we can get the FLOP reduction as the ratio of Dense Inference FLOP Count Compressed Inference FLOP Count fromEquation E.2, wherebis the batch size, and is canceled in the numerator and the denominator of the equation. FLOP Reduction= n(4bd 2 + 2bd 2 a) +bdV n(4bd 2 /2 + 4×2bd 2 r+ 2bd 2 a/2 + 2b(d 2 r+d 2 ra)) +bdV (E.2) Table E.14summarizes the FLOP reduction of different compression methods. As can be seen, the over- head of adding the low-rank adapters (r= 0.1) in SLIM-LoRA and Naive-LoRA is not significant. Table E.14: Compute (FLOP) reduction ratios (×) of different compression methods across various OPT and LLaMA models. In Quantized SLIM , the low-rank adapters are also quantized. (↑indicates better performance.) CompressionOPTLLaMA-2 Method125M 350M 1.3B 2.7B 6.7B 13B 7B 13B SparseGPT + OPTQ1.52 1.66 1.75 1.91 1.94 1.96 1.95 1.97 Wanda + AbsMax1.52 1.66 1.75 1.91 1.94 1.96 1.95 1.97 Naive-LoRA + AbsMax1.32 1.39 1.43 1.50 1.51 1.52 1.49 1.49 SLIM-LoRA + SLIM-Quant 1.32 1.39 1.43 1.50 1.51 1.52 1.49 1.49 SLIM-LoRA Q + SLIM-Quant 1.32 1.39 1.43 1.50 1.51 1.52 1.49 1.49 E.11 Compression costs The computational cost of compression methods varies depending on their complexity. While all approaches can compress a single layer at a time, the memory usage is similar across methods, as each stores only one layer in the GPU’s global memory. Techniques like Wanda, which rely on matrix multiplication, are faster than more complex methods like SparseGPT, which computes the inverse Hessian matrix for each layer. Adding low-rank adapters to Wanda-SVD and SLIM increases computational complexity due to the need for singular value decomposition (SVD), making them comparable to SparseGPT in terms of computation. Table E.15summarizes the time required to compress various models using the discussed methods. Meth- ods incorporating low-rank adapters (SLIM and Wanda-SVD) generally take longer to compress due to their higher complexity. Interestingly, SparseGPT’s compression time is comparable to methods with low-rank adapters, despite only performing pruning and quantization. The saliency-based approach in SLIM does not add significant overhead compared to Wanda-SVD, maintaining efficiency despite its added complexity. APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM159 Table E.15: The required compression time for different models and compression methods using a single H100 GPU. PruningWeightOPTLLaMA-2 Method Quantization 125M 350M 1.3B 2.7B 6.7B 13B 7B 13B Magnitude AbsMax1s1s1s 1s 2s 4s 2s 4s SparseGPT OPTQ1m 2m 5m 11m 22m 41m 25m 46m WandaSLIM-Quant 0.5m 1m 3m 5m 8m 13m 8m 14m Wanda-SVD SLIM-Quant 1m 2m 7m 13m 33m 60m 38m 67m SLIMSLIM-Quant 1m 2m 7m 13m 34m 63m 39m 68m E.12 Rank analysis The key hyperparameter in low-rank approximation is the rank of the adapters. While increasing the rank reduces approximation error, it also leads to higher computational and memory overhead. Therefore, it is crucial to analyze the trade-off between the accuracy improvements and the overhead introduced by the chosen approximation rank. Assuming the rank of the low-rank adapter isrd, wherer <1is a fixed factor anddis the dimension of the weights in a square feed-forward layer, the low-rank adapters are represented asL,R T ∈R d×rd , resulting in a memory overhead ofO(2rd 2 )for storing them. To computeXLR, whereX ∈R b×d is the input with a batch size ofb, the computational complexity isO(2brd 2 ). Given that the original memory and computational complexity of the layer areO(d 2 )andO(bd 2 ), respectively, the overhead introduced by the low-rank adapters becomes negligible whenr≪1. Figure E.2-a shows the average zero-shot accuracy of the OPT-6.7B and LLaMA-2-7B models for various ranks. As expected, increasing the rank leads to improved model accuracy. Based on these results, a rank of r= 0.1provides a substantial boost in accuracy without introducing significant overhead to inference. E.13 Effects of calibration sample count Similar to previous work (SparseGPT, Wanda, AWQ, OmniQuant, and AffineQuant), SLIM leverages a set of calibration data from the C4 dataset to assess weight saliency for pruning and low-rank approximations. Figure E.2-b illustrates the perplexity of LLaMA-2-7B using varying numbers of calibration samples. As shown, SLIM demonstrates low sensitivity to the number of calibration samples, making it effective even in scenarios with limited data. E.14 Sensitivity to calibration dataset Similar to other pruning and quantization methods such as Wanda, SparseGPT, OPTQ, and AWQ, SLIM relies on a calibration dataset to evaluate weight saliency. The C4 [ 121] and SlimPajama [135] datasets are among the most commonly used calibration sets for LLM compression.Table E.16presents the perplexity results for APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM160 0.050.100.150.200.25 Adapter Rank 5.0 7.5 10.0 12.5 15.0 17.5 20.0 LLaMA-2-7B Perplexity Adapter Rank Sensitivity Analysis Naive-LoRA SLiM-LoRA Dense 20406080100120 Number of Calibration Samples 5.0 7.5 10.0 12.5 15.0 17.5 20.0 LLaMA-2-7B Perplexity Calibration Data Sensitivity Analysis Naive-LoRA Group SparseGPT SLiM-LoRA Dense (a)(b) Figure E.2: Sensitivity analysis for the rank of the adapter (a) and the number of calibration samples (b) for different one-shot compression methods. For Naive-LoRA and SLIM-LoRA, we have used the SLIM- Quant quantization method, and for the SparseGPT, we have used the Group quantization version of OPTQ. SLIM-LoRA and SLIM-Quant across different calibration datasets. The results indicate that SLIM is largely insensitive to the choice of dataset, achieving comparable accuracy regardless of the calibration dataset used. Table E.16: Perplexity of different models on WikiText-2 dataset using SLIM-LoRA with 4-bit quantization using SLIM-Quant with different calibration datasets.↓indicates better performance. CalibrationOPTLLaMA-2 Dataset125M 350M 1.3B 2.7B 6.7B 13B 7B 13B 50% 2:4 C457.91 50.09 19.64 15.65 12.71 12.13 7.56 6.50 SlimPajama46.27 44.77 19.35 16.04 12.56 12.32 7.15 6.49 50% Unstructured C439.62 31.51 16.52 13.65 11.42 10.82 6.16 5.36 SlimPajama36.49 29.94 16.64 14.08 11.61 11.02 5.99 5.34 E.15 Sparsity analysis To analyze the impact of sparsity on model accuracy, we conduct experiments on LLaMA-2-13B with 4-bit quantization, pruning it to varying sparsity ratios.Figure E.3presents the perplexity results for SLIM-LoRA with SLIM-Quant , SparseGPT with OPTQ, and Wanda with Group AbsMax. As expected, increasing the sparsity ratio leads to higher perplexity, indicating a trade-off between compression and accuracy. Notably, SLIM-LoRA combined with SLIM-Quant maintains competitive accuracy up to 60% sparsity, whereas other methods experience noticeable degradation at lower sparsity levels. APPENDIX E. SUPPLEMENTARY MATERIAL FOR SLIM161 01020304050607080 Sparsity Ratio (%) 4 6 8 10 12 14 16 18 20 LLaMA-2-13B Perplexity Sparsity Sensitivity Analysis SLiM-LoRA + SLiM-Quant SparseGPT + OPTQ Wanda + Group AbsMax Figure E.3: Sparsity analysis on LLaMA-2-13B model using perplexity on WikiText-2 dataset.↓indicates better performance. E.16 Group quantization challenges Group quantization allows sharing the same quantization parameters for a small group of the elements in the quantized matrix, leading to smaller error. But, using group quantization adds additional challenges to the training and inference of the model, e.g.more complicated implementationandadditional memory and compute overheads. The state-of-the-art group quantization GPU kernel, dense and sparse Marlin [ 38], consists of thousands of lines of CUDA code optimized for only a limited number of GPU architectures, showcasing the amount of effort needed to implement a version of group quantization. Furthermore, other libraries and frameworks, such as Triton [ 143] and CUTLASS [23] do not provide support for 4-bit group quantization, limiting its flexibility and possibility of modification. Furthermore, using group quantization can lead to an additional overhead during matrix multiplication, since more parameters need to be loaded for dequantizing each group. As an example,Table E.17shows the slow-down of using group quantization on the down-projection matrices in different LLaMA-2 and LLaMA- 3.1 models on a NVIDIA A100-40GB GPU, with a batch size of 16. Table E.17: Group quantization slow-down (×) on different LLaMA-2 and LLaMA-3.1 models.↓indicates worse. ModelLLaMA-2-7B LLaMA-2-13B LLaMA-2-70B LLaMA-3.1-405B Slow-Down (×)0.940.950.950.94