Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 465 papers (page 1 of 16)路 42 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Compositional Explanations of Neurons Jacob Andreas, Jesse Mu Published: 2020-06-24Area: Mechanistic Interp.Citations: 206 Tags: ai-safety, empirical, mechanistic-interp | 2020-06-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | 206 |
| Understanding the Role of Individual Units in a Deep Neural Network Agata Lapedriza, Antonio Torralba, Bolei Zhou, David Bau Published: 2020-09-10Area: Mechanistic Interp.Citations: 505 Tags: ai-safety, empirical, mechanistic-interp | 2020-09-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 505 |
| Towards Falsifiable Interpretability Research Ari S. Morcos, Matthew L. Leavitt Published: 2020-10-22Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, interpretability, mechanistic-interp, position | 2020-10-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E6 / R4 (97%) | 74 |
| Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors Bruno A Olshausen, Yann LeCun, Yubei Chen, Zeyu Yun Published: 2021-03-29Area: Mechanistic Interp.Citations: 113 Tags: ai-safety, empirical, mechanistic-interp | 2021-03-29 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (93%) | 113 |
| Knowledge Neurons in Pretrained Transformers Baobao Chang, Damai Dai, Furu Wei, Li Dong Published: 2021-04-18Area: Mechanistic Interp.Citations: 601 Tags: ai-safety, empirical, mechanistic-interp | 2021-04-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 601 |
| Inducing Causal Structure for Interpretable Neural Networks Atticus Geiger, Christopher Potts, Elisa Kreiss, Hanson Lu Published: 2021-12-01Area: Mechanistic Interp.Citations: 95 Tags: ai-safety, empirical, mechanistic-interp | 2021-12-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (97%) | 95 |
| Causal Distillation for Language Models Atticus Geiger, Christopher Potts, Elisa Kreiss, Hanson Lu Published: 2021-12-05Area: Mechanistic Interp.Citations: 29 Tags: ai-safety, empirical, mechanistic-interp | 2021-12-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (97%) | 29 |
| Sparse Interventions in Language Models with Differentiable Masking Dieuwke Hupkes, Ivan Titov, Leon Schmid, Nicola De Cao Published: 2021-12-13Area: Mechanistic Interp.Citations: 33 Tags: ai-safety, empirical, mechanistic-interp | 2021-12-13 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 33 |
| Natural Language Descriptions of Deep Visual Features Antonio Torralba, David Bau, Evan Hernandez, Jacob Andreas Published: 2022-01-26Area: Mechanistic Interp.Citations: 149 Tags: ai-safety, empirical, mechanistic-interp | 2022-01-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 149 |
| Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space Avi Caciularu, Kevin Ro Wang, Mor Geva, Yoav Goldberg Published: 2022-03-28Area: Mechanistic Interp.Citations: 485 Tags: ai-safety, empirical, mechanistic-interp | 2022-03-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 485 |
| CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks Tsui-Wei Weng, Tuomas Oikarinen Published: 2022-04-23Area: Mechanistic Interp.Citations: 130 Tags: ai-safety, mechanistic-interp, tool | 2022-04-23 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E6 / R3 (96%) | 130 |
| LM-Debugger: An Interactive Tool for Inspection and Intervention in Transformer-Based Language Models Avi Caciularu, Bar Tamir, Guy Dar, Micah Shlain Published: 2022-04-26Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, mechanistic-interp, tool | 2022-04-26 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E4 / R3 (95%) | 32 |
| Analyzing Transformers in Embedding Space Ankit Gupta, Guy Dar, Jonathan Berant, Mor Geva Published: 2022-09-06Area: Mechanistic Interp.Citations: 127 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2022-09-06 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R3 (96%) | 127 |
| Polysemanticity and Capacity in Neural Networks Adam Scherlis, Adam S. Jermyn, Buck Shlegeris, Joe Benton Published: 2022-10-04Area: Mechanistic Interp.Citations: 52 Tags: ai-safety, mechanistic-interp, theoretical | 2022-10-04 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (92%) | 52 |
| Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt Published: 2022-11-01Area: Mechanistic Interp.Citations: 834 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2022-11-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (96%) | 834 |
| Engineering Monosemanticity in Toy Models Adam S. Jermyn, Evan Hubinger, Nicholas Schiefer Published: 2022-11-16Area: Mechanistic Interp.Citations: 15 Tags: ai-safety, empirical, mechanistic-interp | 2022-11-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 15 |
| Interpreting Neural Networks through the Polytope Lens Beren Millidge, Carlos Ram贸n Guevara, Connor Leahy, Dan Braun Published: 2022-11-22Area: Mechanistic Interp.Citations: 36 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2022-11-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 36 |
| Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Amir Zur, Aryaman Arora, Atticus Geiger, Christopher Potts Published: 2023-01-11Area: Mechanistic Interp.Citations: 118 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-01-11 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (95%) | 118 |
| Tracr: Compiled Transformers as a Laboratory for Interpretability David Lindner, Janos Kramar, Matthew Rahtz, Sebastian Farquhar Published: 2023-01-12Area: Mechanistic Interp.Citations: 91 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-01-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R4 (95%) | 91 |
| A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations Bilal Chughtai, Lawrence Chan, Neel Nanda Published: 2023-02-06Area: Mechanistic Interp.Citations: 135 Tags: ai-safety, empirical, mechanistic-interp | 2023-02-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 135 |
| Analyzing And Editing Inner Mechanisms of Backdoored Language Models Anka Reuel, Max Lamparth Published: 2023-02-24Area: Mechanistic Interp.Citations: 15 Tags: ai-safety, empirical, mechanistic-interp | 2023-02-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (95%) | 15 |
| Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations Atticus Geiger, Christopher Potts, Noah D. Goodman, Thomas Icard Published: 2023-03-05Area: Mechanistic Interp.Citations: 147 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2023-03-05 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E4 / R2 (94%) | 147 |
| Localizing Model Behavior with Path Patching Aryaman Arora, Chris MacLeod, Lucas Sato, Nicholas Goldowsky-Dill Published: 2023-04-12Area: Mechanistic Interp.Citations: 130 Tags: ai-safety, empirical, mechanistic-interp | 2023-04-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 130 |
| Disentangling Neuron Representations with Concept Vectors Henning Muller, Laura O'Mahony, Mara Graziani, Vincent Andrearczyk Published: 2023-04-19Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, empirical, mechanistic-interp | 2023-04-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 25 |
| N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models Alex Foote, Esben Kran, Fazl Barez, Ionnis Konstas Published: 2023-04-22Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 4 |
| Dissecting Recall of Factual Associations in Auto-Regressive Language Models Amir Globerson, Jasmijn Bastings, Katja Filippova, Mor Geva Published: 2023-04-28Area: Mechanistic Interp.Citations: 440 Tags: ai-safety, empirical, mechanistic-interp | 2023-04-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 440 |
| Towards Automated Circuit Discovery for Mechanistic Interpretability Adria Garriga-Alonso, Aengus Lynch, Arthur Conmy, Augustine N. Mavor-Parker Published: 2023-04-28Area: Mechanistic Interp.Citations: 485 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-04-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 485 |
| How does GPT-2 Compute Greater-Than?: Interpreting Mathematical Abilities in a Pre-trained Language Model Alexandre Variengien, Michael Hanna, Ollie Liu Published: 2023-04-30Area: Mechanistic Interp.Citations: 190 Tags: ai-safety, empirical, mechanistic-interp | 2023-04-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 190 |
| Seeing is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability Eric Gan, Max Tegmark, Ziming Liu Published: 2023-05-04Area: Mechanistic Interp.Citations: 52 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-05-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 52 |
| A Technical Note on Bilinear Layers for Interpretability Lee Sharkey Published: 2023-05-05Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-05-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 10 |