Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Steering CLIP's vision transformer with sparse autoencoders Blake Aaron Richards, Danilo Bzdok, Ethan Goldfarb, Lorenz Hufe Published: 2025-04-11Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (96%) | 14 |
| MIB: A Mechanistic Interpretability Benchmark Aaron Mueller, Adam Belfki, Alessandro Stolfo, Amir Zur Published: 2025-04-17Area: Mechanistic Interp.Citations: 16 Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2025-04-17 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E6 / R3 (97%) | 16 |
| Scaling Sparse Feature Circuits For Studying In-Context Learning Arthur Conmy, Dmitrii Kharlapenko, Fazl Barez, Neel Nanda Published: 2025-04-18Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 4 |
| The Geometry of Self-Verification in a Task-Specific Reasoning Model Andrew Lee, Chris Wendler, Fernanda Vi茅gas, Lihao Sun Published: 2025-04-19Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp | 2025-04-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 6 |
| Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video Blake Aaron Richards, Danilo Bzdok, Edward Stevinson, Lee Sharkey Published: 2025-04-28Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-04-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (98%) | 11 |
| Representation Learning on a Random Lattice Aryeh Brill Published: 2025-04-28Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, mechanistic-interp, theoretical | 2025-04-28 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E6 / R3 (96%) | 1 |
| Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Junping Zhang, Junxuan Wang, Qiong Tang, Rui Lin Published: 2025-04-29Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-04-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 2 |
| Empirical Evaluation of Progressive Coding for Sparse Autoencoders Anders S酶gaard, Hans Peter Published: 2025-04-30Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2025-04-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | - |
| A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i Kola Ayonrinde, Louis Jaburi Published: 2025-05-01Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-05-01 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (95%) | 4 |
| Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kola Ayonrinde, Louis Jaburi Published: 2025-05-02Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-05-02 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | 2 |
| Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders Dong Shu, Haiyan Zhao, Mengnan Du, Ninghao Liu Published: 2025-05-12Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 2 |
| Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates Hang Chen, Jiaying Zhu, Wenya Wang, Xinyu Yang Published: 2025-05-15Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, mechanistic-interp, theoretical | 2025-05-15 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E6 / R3 (94%) | 3 |
| Superposition Yields Robust Neural Scaling Jeff Gore, Yizhou Liu, Ziming Liu Published: 2025-05-15Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-15 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 9 |
| Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders Adria Garriga-Alonso, David Chanin, Tomas Dulka Published: 2025-05-16Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 7 |
| Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence Bofan Gong, Dawn Song, James Evans, Shiyang Lai Published: 2025-05-16Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 2 |
| Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors Christopher Potts, Diyi Yang, Jing Huang, Junyi Tao Published: 2025-05-17Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E8 / R3 (99%) | 4 |
| SplInterp: Improving our Understanding and Training of Sparse Autoencoders Benjamin Macdowall Rynne, Javier Ideami, Jeremy Budd, Keith Duggar Published: 2025-05-17Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2025-05-17 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| Truth Neurons Haohang Li, Jordan W. Suchow, Yangyang Yu, Yupeng Cao Published: 2025-05-18Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (95%) | 1 |
| Towards eliciting latent knowledge from LLMs with mechanistic interpretability Bartosz Cywi艅ski, Emil Ryd, Neel Nanda, Senthooran Rajamanoharan Published: 2025-05-20Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-20 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 6 |
| Ensembling Sparse Autoencoders Chris Lin, Soham Gadgil, Su-In Lee Published: 2025-05-21Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| Interpretability Illusions with Sparse Autoencoders Aaron J. Li, Himabindu Lakkaraju, Suraj Srinivas, Usha Bhalla Published: 2025-05-21Area: Mechanistic Interp.Citations: 5 Tags: adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-21 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 5 |
| GIM: Improved Interpretability for Large Language Models Jing Huang, Joakim Edin, Lars Maal酶e, Maria Maistro Published: 2025-05-23Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E7 / R3 (96%) | 1 |
| Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models Neel Nanda, Noura Al Moubayed, Patrick Leask Published: 2025-05-23Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| Understanding How Value Neurons Shape the Generation of Specified Values in LLMs Di Wang, Jiayi Zhang, Lijie Hu, Shu Yang Published: 2025-05-23Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (95%) | 7 |
| Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Aashiq Muhamed, Kun Zhang, Lingjing Kong, Mona T. Diab Published: 2025-05-26Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, interpretability, mechanistic-interp, position, safety-evaluation | 2025-05-26 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position, safety-evaluation | E5 / R3 (95%) | 6 |
| SAEs Are Good for Steering -- If You Select the Right Features Aaron Mueller, Dana Arad, Yonatan Belinkov Published: 2025-05-26Area: Mechanistic Interp.Citations: 23 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 23 |
| Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders Grigorios G Chrysos, Ioannis Patras, James Oldfield, Mihalis A. Nicolaou Published: 2025-05-27Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 1 |
| Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders Daniil Gavrilov, Daniil Laptev, Nikita Balagansky, Vadim Kurochkin Published: 2025-05-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-05-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| Understanding Refusal in Language Models with Sparse Autoencoders Clement Neo, Erik Cambria, Nirmalendu Prakash, Ranjan Satapathy Published: 2025-05-29Area: Mechanistic Interp.Citations: 7 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-05-29 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | 7 |
| Circuit Stability Characterizes Language Model Generalization Alan Sun Published: 2025-05-30Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-05-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 2 |