Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Residual Stream Analysis with Multi-Layer SAEs Conor Houghton, Laurence Aitchison, Lucy Farnik, Tim Lawson Published: 2024-09-06Area: Mechanistic Interp.Citations: 12 Tags: ai-safety, empirical, mechanistic-interp | 2024-09-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 12 |
| Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small Atticus Geiger, Maheep Chaudhary Published: 2024-09-05Area: Mechanistic Interp.Citations: 30 Tags: ai-safety, empirical, mechanistic-interp | 2024-09-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (97%) | 30 |
| On the Complexity of Neural Computation in Superposition Micah Adler, Nir Shavit Published: 2024-09-05Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, mechanistic-interp, theoretical | 2024-09-05 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | 11 |
| Safety Layers in Aligned Large Language Models Lan Zhang, Liuyi Yao, Shen Li, Yaliang Li Published: 2024-08-30Area: Mechanistic Interp.Citations: 88 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2024-08-30 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E4 / R3 (95%) | 88 |
| Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference Andr茅 Freitas, Geonhee Kim, Marco Valentino Published: 2024-08-16Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, empirical, mechanistic-interp | 2024-08-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | 13 |
| Mathematical Models of Computation in Superposition Dmitry Vaintrob, Jake Mendel, Kaarel H盲nni, Lawrence Chan Published: 2024-08-10Area: Mechanistic Interp.Citations: 23 Tags: ai-safety, mechanistic-interp, theoretical | 2024-08-10 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | 23 |
| Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 Anca Dragan, Arthur Conmy, J脙隆nos Kram脙隆r, Lewis Smith Published: 2024-08-09Area: Mechanistic Interp.Citations: 254 Tags: ai-safety, mechanistic-interp, tool | 2024-08-09 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (96%) | 254 |
| Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models Adam Karvonen, Benjamin Wright, Can Rager, Claudio Mayrink Verdun Published: 2024-07-31Area: Mechanistic Interp.Citations: 49 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 49 |
| Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability Alejandro Mat茅, Jorge Garc铆a-Carrasco, Juan Trujillo Published: 2024-07-29Area: Mechanistic Interp.Citations: 6 Tags: adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-29 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 6 |
| Planning in a Recurrent Neural Network that Plays Sokoban Aaron David Tucker, Adam Gleave, Adri脿 Garriga-Alonso, Chris Cundy Published: 2024-07-22Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, empirical, mechanistic-interp | 2024-07-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (95%) | 11 |
| Adversarial Circuit Evaluation Adri脿 Garriga-Alonso, Niels uit de Bos Published: 2024-07-21Area: Mechanistic Interp.Citations: 1 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp, safety-evaluation | 2024-07-21 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | 1 |
| InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques Adri脿 Garriga-Alonso, Iv谩n Arcuschin, Rohan Gupta, Thomas Kwa Published: 2024-07-19Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-07-19 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (98%) | 7 |
| Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders Arthur Conmy, J脙隆nos Kram脙隆r, Neel Nanda, Nicolas Sonnerat Published: 2024-07-19Area: Mechanistic Interp.Citations: 194 Tags: ai-safety, empirical, mechanistic-interp | 2024-07-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 194 |
| Relational Composition in Neural Networks: A Survey and Call to Action Fernanda B. Vi茅gas, Martin Wattenberg Published: 2024-07-19Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, mechanistic-interp, survey | 2024-07-19 | Mechanistic Interp. | ai-safety, mechanistic-interp, survey | E6 / R3 (91%) | 19 |
| Mechanistically Interpreting a Transformer-based 2-SAT Solver: An Axiomatic Approach Corina P膬s膬reanu, Nils Palumbo, Ravi Mangal, Saranya Vijayakumar Published: 2024-07-18Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2024-07-18 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 2 |
| NNsight and NDIF: Democratizing Access to Foundation Model Internals Aaron Mueller, Adam Belfki, Alexander R. Loftus, Arjun Guha Published: 2024-07-18Area: Mechanistic Interp.Citations: 23 Tags: ai-safety, mechanistic-interp, tool | 2024-07-18 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (96%) | 23 |
| Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent Artem Zholus, Blake Richards, George Adamopoulos, Irina Rish Published: 2024-07-16Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (97%) | 4 |
| LLM Circuit Analyses Are Consistent Across Training and Scale Curt Tigges, Michael Hanna, Qinan Yu, Stella Biderman Published: 2024-07-15Area: Mechanistic Interp.Citations: 42 Tags: ai-safety, empirical, mechanistic-interp | 2024-07-15 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 42 |
| Transformer Circuit Faithfulness Metrics are not Robust Bilal Chughtai, Joseph Miller, William Saunders Published: 2024-07-11Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 10 |
| Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Aaron Mueller Published: 2024-07-05Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-07-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 18 |
| Functional Faithfulness in the Wild: Circuit Discovery with Differentiable Computation Graph Pruning Gerald Penn, Jingcheng Niu, Lei Yu, Zining Zhu Published: 2024-07-04Area: Mechanistic Interp.Citations: 8 Tags: ai-safety, empirical, mechanistic-interp | 2024-07-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (95%) | 8 |
| The Remarkable Robustness of LLMs: Stages of Inference? Max Tegmark, Vedang Lad, Wes Gurnee Published: 2024-06-27Area: Mechanistic Interp.Citations: 98 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R4 (94%) | 98 |
| A Closer Look into Mixture-of-Experts in Large Language Models Jie Fu, Ka Man Lo, Zeyu Huang, Zihan Qiu Published: 2024-06-26Area: Mechanistic Interp.Citations: 28 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 28 |
| Interpreting Attention Layer Outputs with Sparse Autoencoders Arthur Conmy, Connor Kissane, Joseph Isaac Bloom, Neel Nanda Published: 2024-06-25Area: Mechanistic Interp.Citations: 40 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 40 |
| Understanding Language Model Circuits through Knowledge Editing Frank Rudzicz, Huaizhi Ge, Zining Zhu Published: 2024-06-25Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 2 |
| Confidence Regulation Neurons in Language Models Alessandro Stolfo, Ben Wu, Mrinmaya Sachan, Neel Nanda Published: 2024-06-24Area: Mechanistic Interp.Citations: 45 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (94%) | 45 |
| Finding Transformer Circuits with Edge Pruning Adithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi Chen Published: 2024-06-24Area: Mechanistic Interp.Citations: 40 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (97%) | 40 |
| Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models Jun Zhao, Kang Liu, Pengfei Cao, Tianyi Men Published: 2024-06-23Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-06-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 19 |
| Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons Jianhui Chen, Juanzi Li, Lei Hou, Xiaozhi Wang Published: 2024-06-20Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2024-06-20 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R3 (94%) | 25 |
| Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries Amir Globerson, Daniela Gottesman, Eden Biran, Mor Geva Published: 2024-06-18Area: Mechanistic Interp.Citations: 75 Tags: ai-safety, empirical, mechanistic-interp | 2024-06-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (97%) | 75 |