Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Function Vectors in Large Language Models Aaron Mueller, Arnab Sen Sharma, Byron C. Wallace, David Bau Published: 2023-10-23Area: Mechanistic Interp.Citations: 201 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 201 |
| In-Context Learning Creates Task Vectors Amir Globerson, Mor Geva, Roee Hendel Published: 2023-10-24Area: Mechanistic Interp.Citations: 258 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 258 |
| Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism Andr茅 Bauer, Arham Khan, Aswathy Ajith, Daniel Grzenda Published: 2023-10-25Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, mechanistic-interp, tool | 2023-10-25 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (95%) | 18 |
| Codebook Features: Sparse and Discrete Interpretability for Neural Networks Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman Published: 2023-10-26Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 41 |
| How Do Language Models Bind Entities in Context? Jacob Steinhardt, Jiahai Feng Published: 2023-10-26Area: Mechanistic Interp.Citations: 70 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (95%) | 70 |
| Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models Fazl Barez, Michael Lan Published: 2023-11-07Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, mechanistic-interp | 2023-11-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | 9 |
| Uncovering Intermediate Variables in Transformers using Circuit Probing Ellie Pavlick, Michael A. Lepori, Thomas Serre Published: 2023-11-07Area: Mechanistic Interp.Citations: 12 Tags: ai-safety, empirical, mechanistic-interp | 2023-11-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 12 |
| Future Lens: Anticipating Subsequent Tokens from a Single Hidden State Andrew Yuan, Byron C. Wallace, David Bau, Jiuding Sun Published: 2023-11-08Area: Mechanistic Interp.Citations: 97 Tags: ai-safety, empirical, mechanistic-interp | 2023-11-08 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (96%) | 97 |
| Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks David Krueger, Edward Grefenstette, Ekdeep Singh Lubana, Hidenori Tanaka Published: 2023-11-21Area: Mechanistic Interp.Citations: 99 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2023-11-21 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R3 (92%) | 99 |
| Localizing Lying in Llama: Understanding Instructed Dishonesty on True-False Questions Through Prompting, Probing, and Patching James Campbell, Phillip Guo, Richard Ren Published: 2023-11-25Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, mechanistic-interp | 2023-11-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 26 |
| Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2023-11-28Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-11-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (93%) | 41 |
| FlexModel: A Framework for Interpretability of Distributed Large Language Models David B. Emerson, John Willes, Matthew Choi, Muhammad Adil Asif Published: 2023-12-05Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (95%) | 1 |
| Generating Interpretable Networks using Hypernetworks Isaac Liao, Max Tegmark, Ziming Liu Published: 2023-12-05Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2023-12-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (95%) | 2 |
| Interpretability Illusions in the Generalization of Simplified Models Andrew Lampinen, Asma Ghandeharioun, Dan Friedman, Danqi Chen Published: 2023-12-06Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-12-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 20 |
| Grokking Group Multiplication with Cosets Dashiell Stander, Honglu Fan, Qinan Yu, Stella Biderman Published: 2023-12-11Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, mechanistic-interp | 2023-12-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 18 |
| Forbidden Facts: An Investigation of Competing Objectives in Llama-2 Kaivalya Hariharan, Miles Wang, Nir Shavit, Tony T. Wang Published: 2023-12-14Area: Mechanistic Interp.Citations: 3 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2023-12-14 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 3 |
| Successor Heads: Recurring, Interpretable Attention Heads In The Wild Arthur Conmy, Euan Ong, George Ogden, Rhys Gould Published: 2023-12-14Area: Mechanistic Interp.Citations: 69 Tags: ai-safety, empirical, mechanistic-interp | 2023-12-14 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (97%) | 69 |
| Observable Propagation: Uncovering Feature Vectors in Transformers Arman Cohan, Jacob Dunefsky Published: 2023-12-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2023-12-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 2 |
| A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity Andrew Lee, Itamar Pres, Jonathan K. Kummerfeld, Martin Wattenberg Published: 2024-01-03Area: Mechanistic Interp.Citations: 165 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2024-01-03 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R3 (94%) | 165 |
| Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability Jatin Nainani Published: 2024-01-08Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-08 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 3 |
| Universal Neurons in GPT2 Language Models Dimitris Bertsimas, Neel Nanda, Qinyi Sun, Tara Rezaei Kheirkhah Published: 2024-01-22Area: Mechanistic Interp.Citations: 83 Tags: ai-safety, empirical, mechanistic-interp | 2024-01-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 83 |
| A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments Aryaman Arora, Atticus Geiger, Christopher Potts, Jing Huang Published: 2024-01-23Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 9 |
| Fluent dreaming for language models Michael Sklar, T. Ben Thompson, Zygimantas Straznickas Published: 2024-01-24Area: Mechanistic Interp.Citations: 4 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2024-01-24 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E4 / R2 (97%) | 4 |
| Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability Alex Grzankowski Published: 2024-01-31Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-01-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E8 / R4 (94%) | 10 |
| Challenges in Mechanistically Interpreting Model Representations James Dao, Satvik Golechha Published: 2024-02-06Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-02-06 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (93%) | 4 |
| Opening the AI black box: program synthesis via mechanistic interpretability Anish Mudide, Chloe Loughridge, Eric J. Michaud, Isaac Liao Published: 2024-02-07Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (96%) | 19 |
| Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs Alan Cooney, Bilal Chughtai, Neel Nanda Published: 2024-02-11Area: Mechanistic Interp.Citations: 31 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (93%) | 31 |
| Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals Alberto Cazzaniga, Bernhard Sch枚lkopf, Diego Doimo, Francesco Ortu Published: 2024-02-18Area: Mechanistic Interp.Citations: 35 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (95%) | 35 |
| A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task Abhay Sheshadri, Christian Bartelt, Jannik Brinkmann, Paul Swoboda Published: 2024-02-19Area: Mechanistic Interp.Citations: 48 Tags: ai-safety, empirical, mechanistic-interp | 2024-02-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 48 |
| Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT Qinyuan Cheng, Qiong Tang, Tianxiang Sun, Xipeng Qiu Published: 2024-02-19Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-19 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 25 |