Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Explorations of Self-Repair in Language Models Cody Rushing, Neel Nanda Published: 2024-02-23Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 20 |
| Information Flow Routes: Automatically Interpreting Language Models at Scale Elena Voita, Javier Ferrando Published: 2024-02-27Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 74 |
| How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning Joykirat Singh, Soumen Chakrabarti, Subhabrata Dutta, Tanmoy Chakraborty Published: 2024-02-28Area: Mechanistic Interp.Citations: 54 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 54 |
| AtP*: An efficient and scalable method for localizing LLM behaviour to components János Kramár, Neel Nanda, Rohin Shah, Tom Lieberum Published: 2024-03-01Area: Mechanistic Interp.Citations: 71 Tags: ai-safety, empirical, mechanistic-interp | 2024-03-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 71 |
| pyvene: A Library for Understanding and Improving PyTorch Models via Interventions Aryaman Arora, Atticus Geiger, Christopher D. Manning, Christopher Potts Published: 2024-03-12Area: Mechanistic Interp.Citations: 44 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-03-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 44 |
| Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms Michael Hanna, Sandro Pezzelle, Yonatan Belinkov Published: 2024-03-26Area: Mechanistic Interp.Citations: 90 Tags: ai-safety, empirical, mechanistic-interp | 2024-03-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 90 |
| Mechanisms of Non-Factual Hallucination in Language Models Jackie Chi Kit Cheung, Lei Yu, Meng Cao, Yue Dong Published: 2024-03-27Area: Mechanistic Interp.Citations: 38 Tags: ai-safety, empirical, mechanistic-interp | 2024-03-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (95%) | 38 |
| Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models Aaron Mueller, Can Rager, David Bau, Eric J. Michaud Published: 2024-03-28Area: Mechanistic Interp.Citations: 270 Tags: ai-safety, empirical, mechanistic-interp | 2024-03-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 270 |
| LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models Elena Voita, Igor Tufanov, Javier Ferrando, Karen Hambardzumyan Published: 2024-04-10Area: Mechanistic Interp.Citations: 15 Tags: ai-safety, mechanistic-interp, tool | 2024-04-10 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (95%) | 15 |
| Automatic Discovery of Visual Circuits Achyuta Rajaram, Antonio Torralba, Jacob Andreas, Neil Chowdhury Published: 2024-04-22Area: Mechanistic Interp.Citations: 10 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2024-04-22 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 10 |
| MAIA: A Multimodal Automated Interpretability Agent Achyuta Rajaram, Antonio Torralba, Evan Hernandez, Franklin Wang Published: 2024-04-22Area: Mechanistic Interp.Citations: 45 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E6 / R4 (95%) | 45 |
| How to Use and Interpret Activation Patching Neel Nanda, Stefan Heimersheim Published: 2024-04-23Area: Mechanistic Interp.Citations: 109 Tags: ai-safety, interpretability, mechanistic-interp, survey | 2024-04-23 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E6 / R4 (95%) | 109 |
| Improving Dictionary Learning with Gated Sparse Autoencoders Arthur Conmy, János Kramár, Lewis Smith, Neel Nanda Published: 2024-04-24Area: Mechanistic Interp.Citations: 138 Tags: ai-safety, empirical, mechanistic-interp | 2024-04-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 138 |
| Let's Think Dot by Dot: Hidden Computation in Transformer Language Models Jacob Pfau, Samuel R. Bowman, William Merrill Published: 2024-04-24Area: Mechanistic Interp.Citations: 145 Tags: ai-safety, empirical, mechanistic-interp | 2024-04-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 145 |
| How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability Alejandro Maté, Jorge García-Carrasco, Juan Trujillo Published: 2024-05-07Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 13 |
| Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2024-05-14Area: Mechanistic Interp.Citations: 66 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-05-14 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | 66 |
| Learnable Privacy Neurons Localization in Language Models Ruizhe Chen, Tianxiang Hu, Yang Feng, Zuozhu Liu Published: 2024-05-16Area: Mechanistic Interp.Citations: 30 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2024-05-16 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 30 |
| Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning Dan Braun, Jordan Taylor, Lee Sharkey, Nicholas Goldowsky-Dill Published: 2024-05-17Area: Mechanistic Interp.Citations: 57 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 57 |
| The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks Avery Griffin, Dan Braun, Jake Mendel, Jörn Stöhler Published: 2024-05-17Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (93%) | 6 |
| Using Degeneracy in the Loss Landscape for Mechanistic Interpretability Cindy Wu, Dan Braun, Jake Mendel, Kaarel Hänni Published: 2024-05-17Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-05-17 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 11 |
| Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models Charles O'Neill, Thang Bui Published: 2024-05-21Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 11 |
| Automatically Identifying Local and Global Circuits with Linear Computation Graphs Fukang Zhu, Junxuan Wang, Wentao Shu, Xipeng Qiu Published: 2024-05-22Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 20 |
| Not All Language Model Features Are Linear Eric J. Michaud, Isaac Liao, Joshua Engels, Max Tegmark Published: 2024-05-23Area: Mechanistic Interp.Citations: 106 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (94%) | 106 |
| From Neurons to Neutrons: A Case Study in Interpretability Mike Williams, Niklas Nolte, Ouail Kitouni, Sokratis Trifinopoulos Published: 2024-05-27Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 4 |
| Knowledge Circuits in Pretrained Transformers Huajun Chen, Mengru Wang, Ningyu Zhang, Shumin Deng Published: 2024-05-28Area: Mechanistic Interp.Citations: 44 Tags: ai-safety, empirical, mechanistic-interp | 2024-05-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 44 |
| Evidence of Learned Look-Ahead in a Chess-Playing Neural Network Cameron Allen, Erik Jenner, Scott Emmons, Shreyas Kapur Published: 2024-06-02Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-02 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 25 |
| From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation Eugene Belilovsky, Geraldin Nanfack, Michael Eickenberg Published: 2024-06-03Area: Mechanistic Interp.Citations: 1 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2024-06-03 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 1 |
| Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience David Poeppel, Federico Adolfi, Gemma Roig, Martina G. Vilas Published: 2024-06-03Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-06-03 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (94%) | 10 |
| Scaling and Evaluating Sparse Autoencoders Alec Radford, Gabriel Goh, Henk Tillman, Ilya Sutskever Published: 2024-06-06Area: Mechanistic Interp.Citations: 334 Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2024-06-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (96%) | 334 |
| The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision Liv Gorton Published: 2024-06-06Area: Mechanistic Interp.Citations: 31 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 31 |