Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| ProtoMech: Protein Circuit Tracing via Cross-layer Transcoders Amirali Aghazadeh, Daniel Saeedi, Darin Tsui, Kunal Talreja Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Prototype Transformer: Towards Language Model Architectures Interpretable by Design Amine M'Charrak, Bayar Menzat, Chang Qi, Markus Kaltenberger Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Accelerating Sparse Autoencoder Training via Layer-Wise Transfer Learning in Large Language Models Davide Ghilardi, Federico Belotti, Jaehyuk Lim, Marco Molinari Published: -Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (96%) | 2 |
| A Mathematical Framework for Transformer Circuits Amanda Askell, Andy Jones, Anna Chen, Ben Mann Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | - |
| A Toy Model of Mechanistic (Un)Faithfulness Chris Olah Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | - |
| Attribution Patching: Activation Patching At Industrial Scale Neel Nanda Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | - |
| Causal Scrubbing: A Method for Rigorously Testing Interpretability Hypotheses Adri脙 Garriga-Alonso, Ansh Radhakrishnan, Buck Shlegeris, Jenny Nitishinskaya Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Circuits Updates - April 2025 Adam Jermyn, Brian Chen, Jack Lindsey, Josh Batson Published: -Area: Mechanistic Interp.Citations: - Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Circuit Tracing: Revealing Computational Graphs in Language Models Adam Jermyn, Adam Pearce, Adly Templeton, Andrew Persic Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Cracking the Circuits: Mechanistic Interpretability in Large Language Models Dost Muhammad, Malika Bendechache, Muhammad Salman, Mushtaq Ali Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | - |
| Curve Detectors Chris Olah, Gabriel Goh, Ludwig Schubert, Michael Petrov Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Explaining AI through mechanistic interpretability Barnaby Crook, Lena K盲stner Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | - |
| From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models Etowah Adams, Liam Bai, Minji Lee, Mohammed AlQuraishi Published: -Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 32 |
| Gemma Scope 2: Comprehensive Suite of SAEs and Transcoders for Gemma 3 Arthur Conmy, Callum McDougall, Janos Kramar, Neel Nanda Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (98%) | - |
| Goodfire Ember: Scaling Interpretability for Frontier Model Alignment Atticus Geiger, Curt Tigges, Dan Braun, Daniel Balsam Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, alignment-training, interpretability, mechanistic-interp, tool | E4 / R3 (99%) | - |
| In-context Learning and Induction Heads Amanda Askell, Andy Jones, Anna Chen, Ben Mann Published: -Area: Mechanistic Interp.Citations: 751 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 751 |
| Insights on Crosscoder Model Diffing Adam Jermyn, Christopher Olah, Jack Lindsey, Jonathan Marcus Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Interpreting GPT: The Logit Lens nostalgebraist Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Language Models Can Explain Neurons in Language Models Dan Mossing, Gabriel Goh, Henk Tillman, Ilya Sutskever Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R2 (96%) | - |
| Multimodal Neurons in Artificial Neural Networks Alec Radford, Chelsea Voss, Chris Olah, Gabriel Goh Published: -Area: Mechanistic Interp.Citations: 390 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R4 (96%) | 390 |
| Neuronpedia: Interactive SAE Feature Explorer Johnny Lin Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E6 / R3 (95%) | - |
| On the Biology of a Large Language Model Adam Jermyn, Adam Pearce, Adly Templeton, Andrew Persic Published: -Area: Mechanistic Interp.Citations: - Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Privileged Bases in the Transformer Residual Stream Christopher Olah, Nelson Elhage, Robert Lasenby Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | - |
| Progress on Attention Adam Jermyn, Adam Pearce, Ben Thompson, Callum McDougall Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| SAELens: A Library for Training and Analyzing Sparse Autoencoders Anthony Duong, Curt Tigges, David Chanin, Joseph Bloom Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (97%) | - |
| Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Adam Jermyn, Adam Pearce, Adly Templeton, Alex Tamkin Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models Arman Cohan, Carter Teplica, Tim G. J. Rudner, Yixin Liu Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| SFAL: Semantic-Functional Alignment Scores for Distributional Evaluation of Auto-Interpretability in Sparse Autoencoders Andrea Seveso, Antonio Serino, Daniele Potert矛, Fabio Mercorio Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp, safety-evaluation | - | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | - |
| Softmax Linear Units Amanda Askell, Andy Jones, Anna Chen, Ben Mann Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Sparse Autoencoders Find Partially Interpretable Features in Italian Small Language Models Alessandro Bondielli, Alessandro Lenci, Lucia C. Passaro Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | - |