Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| In-Context Learning Creates Task Vectors Amir Globerson, Mor Geva, Roee Hendel Published: 2023-10-24Area: Mechanistic Interp.Citations: 258 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 258 |
| Function Vectors in Large Language Models Aaron Mueller, Arnab Sen Sharma, Byron C. Wallace, David Bau Published: 2023-10-23Area: Mechanistic Interp.Citations: 201 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 201 |
| Identifying Interpretable Visual Features in Artificial and Biological Neural Systems David Klindt, Francisco Acosta, Fr茅d茅ric Poitevin, Nina Miolane Published: 2023-10-17Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 10 |
| Attribution Patching Outperforms Automated Circuit Discovery Aaquib Syed, Arthur Conmy, Can Rager Published: 2023-10-16Area: Mechanistic Interp.Citations: 108 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 108 |
| Circuit Component Reuse Across Tasks in Transformer Language Models Carsten Eickhoff, Ellie Pavlick, Jack Merullo Published: 2023-10-12Area: Mechanistic Interp.Citations: 99 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (96%) | 99 |
| Interpreting Learned Feedback Patterns in Large Language Models Amir Abdullah, Clement Neo, Fazl Barez, Luke Marks Published: 2023-10-12Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 5 |
| Understanding and Controlling a Maze-Solving Policy Network Alexander Matt Turner, Austin Meek, Monte MacDiarmid, Mrinank Sharma Published: 2023-10-12Area: Mechanistic Interp.Citations: 22 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (94%) | 22 |
| An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L Can Rager, James Dao, Jett Janiak, Yeu-Tong Lau Published: 2023-10-11Area: Mechanistic Interp.Citations: 6 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2023-10-11 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 6 |
| The Importance of Prompt Tuning for Automated Neuron Explanations Arjun Chatha, Justin Lee, Keng-Chi Chang, Tsui-Wei Weng Published: 2023-10-09Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 11 |
| Copy Suppression: Comprehensively Understanding an Attention Head Arthur Conmy, Callum McDougall, Cody Rushing, Neel Nanda Published: 2023-10-06Area: Mechanistic Interp.Citations: 56 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 56 |
| Discovering Knowledge-Critical Subnetworks in Pretrained Language Models Antoine Bosselut, Deniz Bayazit, Gail Weiss, Negar Foroutan Published: 2023-10-04Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, mechanistic-interp | 2023-10-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 20 |
| Efficient Streaming Language Models with Attention Sinks Beidi Chen, Guangxuan Xiao, Mike Lewis, Song Han Published: 2023-09-29Area: Mechanistic Interp.Citations: 1422 Tags: ai-safety, empirical, mechanistic-interp | 2023-09-29 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 1422 |
| Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Fred Zhang, Neel Nanda Published: 2023-09-27Area: Mechanistic Interp.Citations: 193 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-09-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | 193 |
| Rigorously Assessing Natural Language Explanations of Neurons Atticus Geiger, Christopher Potts, Jing Huang, Karel D'Oosterlinck Published: 2023-09-19Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2023-09-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (95%) | 41 |
| Sparse Autoencoders Find Highly Interpretable Features in Language Models Aidan Ewart, Hoagy Cunningham, Lee Sharkey, Logan Riggs Published: 2023-09-15Area: Mechanistic Interp.Citations: 881 Tags: ai-safety, empirical, mechanistic-interp | 2023-09-15 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (94%) | 881 |
| Uncovering Mesa-Optimization Algorithms in Transformers Alexander Meulemans, Blaise Ag眉era y Arcas, Eyvind Niklasson, Jo茫o Sacramento Published: 2023-09-11Area: Mechanistic Interp.Citations: 86 Tags: ai-safety, empirical, mechanistic-interp | 2023-09-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 86 |
| Neurons in Large Language Models: Dead, N-gram, Positional Christoforos Nalmpantis, Elena Voita, Javier Ferrando Published: 2023-09-09Area: Mechanistic Interp.Citations: 75 Tags: ai-safety, empirical, mechanistic-interp | 2023-09-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (94%) | 75 |
| FIND: A Function Description Benchmark for Evaluating Interpretability Methods Antonio Torralba, David Bau, Jacob Andreas, Joanna Materzynska Published: 2023-09-07Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2023-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E4 / R3 (94%) | 32 |
| Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP Aryaman Arora, Paul Pu Liang, Rohan Pandey, Vedant Palit Published: 2023-08-27Area: Mechanistic Interp.Citations: 47 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-08-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 47 |
| The Hydra Effect: Emergent Self-repair in Language Model Computations Janos Kramar, Matthew Rahtz, Shane Legg, Thomas McGrath Published: 2023-07-28Area: Mechanistic Interp.Citations: 96 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 96 |
| On Privileged and Convergent Bases in Neural Network Representations Davis Brown, Nikhil Vyas, Yamini Bansal Published: 2023-07-24Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2023-07-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | - |
| FACADE: A Framework for Adversarial Circuit Anomaly Detection and Evaluation Andres Carranza, Arnuv Tandon, Dhruv Pai, Rylan Schaeffer Published: 2023-07-20Area: Mechanistic Interp.Citations: 2 Tags: adversarial-robustness, ai-safety, mechanistic-interp, safety-evaluation, tool | 2023-07-20 | Mechanistic Interp. | adversarial-robustness, ai-safety, mechanistic-interp, safety-evaluation, tool | E5 / R3 (94%) | 2 |
| Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla Geoffrey Irving, J脙隆nos Kram脙隆r, Matthew Rahtz, Neel Nanda Published: 2023-07-18Area: Mechanistic Interp.Citations: 144 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-18 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 144 |
| Overthinking the Truth: Understanding How Language Models Process False Demonstrations Danny Halawi, Jacob Steinhardt, Jean-Stanislas Denain Published: 2023-07-18Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, empirical, mechanistic-interp | 2023-07-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (92%) | 74 |
| Discovering Variable Binding Circuitry with Desiderata David Bau, Max Nadeau, Nikhil Prakash, Tamar Rott Shaham Published: 2023-07-07Area: Mechanistic Interp.Citations: 22 Tags: ai-safety, empirical, mechanistic-interp | 2023-07-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R2 (94%) | 22 |
| The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks Jacob Andreas, Max Tegmark, Ziming Liu, Ziqian Zhong Published: 2023-06-30Area: Mechanistic Interp.Citations: 145 Tags: ai-safety, empirical, mechanistic-interp | 2023-06-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 145 |
| Learning Transformer Programs Alexander Wettig, Dan Friedman, Danqi Chen Published: 2023-06-01Area: Mechanistic Interp.Citations: 48 Tags: ai-safety, empirical, mechanistic-interp | 2023-06-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 48 |
| Neuron to Graph: Interpreting Language Model Neurons at Scale Alex Foote, Esben Kran, Fazl Barez, Ioannis Konstas Published: 2023-05-31Area: Mechanistic Interp.Citations: 28 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-05-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | 28 |
| Language Models Implement Simple Word2Vec-style Vector Arithmetic Carsten Eickhoff, Ellie Pavlick, Jack Merullo Published: 2023-05-25Area: Mechanistic Interp.Citations: 86 Tags: ai-safety, empirical, mechanistic-interp | 2023-05-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 86 |
| A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis Alessandro Stolfo, Mrinmaya Sachan, Yonatan Belinkov Published: 2023-05-24Area: Mechanistic Interp.Citations: 71 Tags: ai-safety, empirical, mechanistic-interp | 2023-05-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 71 |