Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Mechanistic? Naomi Saphra, Sarah Wiegreffe Published: 2024-10-07Area: Surveys & ReviewsCitations: 38 Tags: ai-safety, interpretability, position, surveys-reviews | 2024-10-07 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (94%) | 38 |
| Gradient Routing: Masking Gradients to Localize Computation in Neural Networks Alexander Matt Turner, Alex Cloud, Evzen Wybitul, Jacob Goldman-Wetzler Published: 2024-10-06Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 18 |
| Optimal Ablation for Interpretability Lucas Janson, Maximilian Li Published: 2024-09-16Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-09-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | 14 |
| Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis Oliver Guest, Oscar Delaney, Zoe Williams Published: 2024-09-12Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R3 (94%) | 3 |
| TracrBench: Generating Interpretability Testbeds with Large Language Models Hannes Thurnherr, Jérémy Scheurer Published: 2024-09-07Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (97%) | 4 |
| Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience Anna Ivanova, Chole Li, Danyal Akarca, George Ogden Published: 2024-08-22Area: Surveys & ReviewsCitations: 7 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (95%) | 7 |
| The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms Adam Davies, Ashkan Khakzar Published: 2024-08-11Area: Surveys & ReviewsCitations: 14 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-11 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (95%) | 14 |
| The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability Aaron Mueller, Arnab Sen Sharma, Aruna Sankaranarayanan, Can Rager Published: 2024-08-02Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models Adam Karvonen, Benjamin Wright, Can Rager, Claudio Mayrink Verdun Published: 2024-07-31Area: Mechanistic Interp.Citations: 49 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 49 |
| Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability Alejandro Maté, Jorge García-Carrasco, Juan Trujillo Published: 2024-07-29Area: Mechanistic Interp.Citations: 6 Tags: adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-29 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 6 |
| InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques Adrià Garriga-Alonso, Iván Arcuschin, Rohan Gupta, Thomas Kwa Published: 2024-07-19Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-07-19 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (98%) | 7 |
| Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent Artem Zholus, Blake Richards, George Adamopoulos, Irina Rish Published: 2024-07-16Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (97%) | 4 |
| Transformer Circuit Faithfulness Metrics are not Robust Bilal Chughtai, Joseph Miller, William Saunders Published: 2024-07-11Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 10 |
| Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Aaron Mueller Published: 2024-07-05Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-07-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 18 |
| A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Abulhair Saparov, Daking Rai, Shi Feng, Yilun Zhou Published: 2024-07-02Area: Surveys & ReviewsCitations: 91 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-07-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 91 |
| Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models Jun Zhao, Kang Liu, Pengfei Cao, Tianyi Men Published: 2024-06-23Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-06-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 19 |
| Compact Proofs of Model Performance via Mechanistic Interpretability Alex Gibson, Chun Hei Yip, Euan Ong, Jason Gross Published: 2024-06-17Area: Formal/TheoreticalCitations: 12 Tags: ai-safety, empirical, formaltheoretical, interpretability | 2024-06-17 | Formal/Theoretical | ai-safety, empirical, formaltheoretical, interpretability | E4 / R3 (92%) | 12 |
| Transcoders Find Interpretable LLM Feature Circuits Jacob Dunefsky, Neel Nanda, Philippe Chlenski Published: 2024-06-17Area: Mechanistic Interp.Citations: 102 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-06-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 102 |
| Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience David Poeppel, Federico Adolfi, Gemma Roig, Martina G. Vilas Published: 2024-06-03Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-06-03 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (94%) | 10 |
| From Neurons to Neutrons: A Case Study in Interpretability Mike Williams, Niklas Nolte, Ouail Kitouni, Sokratis Trifinopoulos Published: 2024-05-27Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 4 |
| Using Degeneracy in the Loss Landscape for Mechanistic Interpretability Cindy Wu, Dan Braun, Jake Mendel, Kaarel Hänni Published: 2024-05-17Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-05-17 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 11 |
| Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2024-05-14Area: Mechanistic Interp.Citations: 66 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-05-14 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | 66 |
| How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability Alejandro Maté, Jorge García-Carrasco, Juan Trujillo Published: 2024-05-07Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 13 |
| A Primer on the Inner Workings of Transformer-based Language Models Arianna Bisazza, Gabriele Sarti, Javier Ferrando, Marta R. Costa-jussà Published: 2024-04-30Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-04-30 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (95%) | 80 |
| How to Use and Interpret Activation Patching Neel Nanda, Stefan Heimersheim Published: 2024-04-23Area: Mechanistic Interp.Citations: 109 Tags: ai-safety, interpretability, mechanistic-interp, survey | 2024-04-23 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E6 / R4 (95%) | 109 |
| MAIA: A Multimodal Automated Interpretability Agent Achyuta Rajaram, Antonio Torralba, Evan Hernandez, Franklin Wang Published: 2024-04-22Area: Mechanistic Interp.Citations: 45 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E6 / R4 (95%) | 45 |
| Mechanistic Interpretability for AI Safety — A Review Efstratios Gavves, Leonard Bereska Published: 2024-04-22Area: Surveys & ReviewsCitations: 335 Tags: ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | 2024-04-22 | Surveys & Reviews | ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 335 |
| pyvene: A Library for Understanding and Improving PyTorch Models via Interventions Aryaman Arora, Atticus Geiger, Christopher D. Manning, Christopher Potts Published: 2024-03-12Area: Mechanistic Interp.Citations: 44 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-03-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 44 |
| RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations Atticus Geiger, Christopher Potts, Jing Huang, Mor Geva Published: 2024-02-27Area: Safety EvaluationCitations: 61 Tags: ai-safety, benchmark, interpretability, safety-evaluation | 2024-02-27 | Safety Evaluation | ai-safety, benchmark, interpretability, safety-evaluation | E4 / R3 (98%) | 61 |
| CausalGym: Benchmarking Causal Interpretability Methods on Linguistic Tasks Aryaman Arora, Christopher Potts, Dan Jurafsky Published: 2024-02-19Area: Safety EvaluationCitations: 36 Tags: ai-safety, benchmark, interpretability, safety-evaluation | 2024-02-19 | Safety Evaluation | ai-safety, benchmark, interpretability, safety-evaluation | E5 / R3 (95%) | 36 |