Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models Binxu Wang, Demba Ba, Ekdeep Singh Lubana, Isabel Papadimitriou Published: 2025-02-18Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 32 |
| Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders Jiayi Yuan, Ninghao Liu, Wenlin Yao, Xiaoming Zhai Published: 2025-02-21Area: Mechanistic Interp.Citations: 20 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-02-21 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R4 (93%) | 20 |
| SAE-V: Interpreting Multimodal Models for Enhanced Alignment Changye Li, Hantao Lou, Jiaming Ji, Yaodong Yang Published: 2025-02-22Area: Mechanistic Interp.Citations: 8 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2025-02-22 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E4 / R3 (96%) | 8 |
| Are Sparse Autoencoders Useful? A Case Study in Sparse Probing Joshua Engels, Max Tegmark, Neel Nanda, Senthooran Rajamanoharan Published: 2025-02-23Area: Mechanistic Interp.Citations: 62 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 62 |
| FADE: Why Bad Descriptions Happen to Good Features Aakriti Jain, Bruno Puri, Elena Golimblevskaia, Patrick Kahardipraja Published: 2025-02-24Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2025-02-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | 6 |
| Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations Conor Houghton, Laurence Aitchison, Lucy Farnik, Tim Lawson Published: 2025-02-25Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 6 |
| Interpreting CLIP with Hierarchical Sparse Autoencoders Hubert Baniecki, Przemyslaw Biecek, Vladimir Zaigrajew Published: 2025-02-27Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 19 |
| Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification Ding Zhu, Mohammad Mahdi Khalili, Vishnu Kabir Chhabra Published: 2025-02-27Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 5 |
| Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? Francois Portet, Maxime Meloux, Maxime Peyrard, Silviu Maniu Published: 2025-02-28Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-02-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (94%) | 14 |
| From superposition to sparse codes: interpretable representations in neural networks Charles O'Neill, David Klindt, Harald Maurer, Nina Miolane Published: 2025-03-03Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, mechanistic-interp, safety-evaluation, theoretical | 2025-03-03 | Mechanistic Interp. | ai-safety, mechanistic-interp, safety-evaluation, theoretical | E5 / R3 (93%) | 7 |
| Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry Demba Ba, Ekdeep Singh Lubana, Sai Sumedh R. Hindupur, Thomas Fel Published: 2025-03-03Area: Mechanistic Interp.Citations: 34 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-03 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | 34 |
| Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation Gal Niv, Jonathan Jacobi Published: 2025-03-03Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-03 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 3 |
| Mixture of Experts Made Intrinsically Interpretable Adel Bibi, Ashkan Khakzar, Christian Schroeder de Witt, Constantin Venhoff Published: 2025-03-05Area: Mechanistic Interp.Citations: 12 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 12 |
| RouteSAE: Route Sparse Autoencoder to Interpret Large Language Models Guojun Ma, Mingyang Wan, Sihang Li, Tao Liang Published: 2025-03-11Area: Mechanistic Interp.Citations: 17 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-03-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (95%) | 17 |
| SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability Adam Karvonen, Arthur Conmy, Callum McDougall, Can Rager Published: 2025-03-12Area: Mechanistic Interp.Citations: 62 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2025-03-12 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (95%) | 62 |
| HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks Atticus Geiger, Christopher Potts, Jing Huang, Jiuding Sun Published: 2025-03-13Area: Mechanistic Interp.Citations: 8 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-03-13 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 8 |
| Combining Causal Models for More Accurate Abstractions of Neural Networks Atticus Geiger, Sara Magliacane, Theodora-Mara P卯slar Published: 2025-03-14Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-14 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research Abir Harrasse, Amir Abdullah, Clement Neo, Dhruv Nathawani Published: 2025-03-17Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, dataset, interpretability, mechanistic-interp | 2025-03-17 | Mechanistic Interp. | ai-safety, dataset, interpretability, mechanistic-interp | E6 / R3 (95%) | 2 |
| Learning Multi-Level Features with Matryoshka Sparse Autoencoders Adam Karvonen, Bart Bussmann, Neel Nanda, Noa Nabeshima Published: 2025-03-21Area: Mechanistic Interp.Citations: 63 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 63 |
| Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need Adam Karvonen Published: 2025-03-21Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 1 |
| I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders Alexey Dontsov, Andrey Galichin, Anton Razzhigaev, Elena Tutubalina Published: 2025-03-24Area: Mechanistic Interp.Citations: 24 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 24 |
| The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction Dian Zhou, Lei Yu, Meng Cao, Yihuai Hong Published: 2025-03-29Area: Mechanistic Interp.Citations: 15 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-29 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 15 |
| Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality Adam Davies, Julia Hockenmaier, Marc E. Canby, Sewoong Lee Published: 2025-03-31Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2025-03-31 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | 2 |
| Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition Brianna Chrisman, Lee Sharkey, Lucius Bushnaq Published: 2025-03-31Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, mechanistic-interp | 2025-03-31 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | 3 |
| Automated Feature Labeling with Token-Space Gradient Descent Julian Schulz, Seamus Fallows Published: 2025-04-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-04-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | - |
| How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence Himabindu Lakkaraju, Hongzhe Du, Karim Saraipour, Min Cai Published: 2025-04-03Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-03 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R4 (94%) | 5 |
| Robustly Identifying Concepts Introduced During Chat Fine-tuning Using Crosscoders Bilal Chughtai, Caden Juang, Clement Dumas, Julian Minder Published: 2025-04-03Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-03 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (97%) | 7 |
| Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models Mateusz Pach, Quentin Bouniot, Serge Belongie, Shyamgopal Karthik Published: 2025-04-03Area: Mechanistic Interp.Citations: 21 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-04-03 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (97%) | 21 |
| Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs Letao Han, Ling Hu, Xiaoyang Gu, Yuemei Xu Published: 2025-04-07Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (94%) | 1 |
| Towards Combinatorial Interpretability of Neural Computation Dan Alistarh, Micah Adler, Nir Shavit Published: 2025-04-10Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-04-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | 7 |