Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent Artem Zholus, Blake Richards, George Adamopoulos, Irina Rish Published: 2024-07-16Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (97%) | 4 |
| InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques Adri脿 Garriga-Alonso, Iv谩n Arcuschin, Rohan Gupta, Thomas Kwa Published: 2024-07-19Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-07-19 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (98%) | 7 |
| Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability Alejandro Mat茅, Jorge Garc铆a-Carrasco, Juan Trujillo Published: 2024-07-29Area: Mechanistic Interp.Citations: 6 Tags: adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-29 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 6 |
| Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models Adam Karvonen, Benjamin Wright, Can Rager, Claudio Mayrink Verdun Published: 2024-07-31Area: Mechanistic Interp.Citations: 49 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 49 |
| The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability Aaron Mueller, Arnab Sen Sharma, Aruna Sankaranarayanan, Can Rager Published: 2024-08-02Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms Adam Davies, Ashkan Khakzar Published: 2024-08-11Area: Surveys & ReviewsCitations: 14 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-11 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (95%) | 14 |
| Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience Anna Ivanova, Chole Li, Danyal Akarca, George Ogden Published: 2024-08-22Area: Surveys & ReviewsCitations: 7 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (95%) | 7 |
| TracrBench: Generating Interpretability Testbeds with Large Language Models Hannes Thurnherr, J茅r茅my Scheurer Published: 2024-09-07Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (97%) | 4 |
| Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis Oliver Guest, Oscar Delaney, Zoe Williams Published: 2024-09-12Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R3 (94%) | 3 |
| Optimal Ablation for Interpretability Lucas Janson, Maximilian Li Published: 2024-09-16Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-09-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | 14 |
| Gradient Routing: Masking Gradients to Localize Computation in Neural Networks Alexander Matt Turner, Alex Cloud, Evzen Wybitul, Jacob Goldman-Wetzler Published: 2024-10-06Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 18 |
| Mechanistic? Naomi Saphra, Sarah Wiegreffe Published: 2024-10-07Area: Surveys & ReviewsCitations: 38 Tags: ai-safety, interpretability, position, surveys-reviews | 2024-10-07 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (94%) | 38 |
| Bilinear MLPs Enable Weight-Based Mechanistic Interpretability Alice Rigg, Jose M. Oramas, Lee Sharkey, Michael T. Pearce Published: 2024-10-10Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| The Computational Complexity of Circuit Discovery for Inner Interpretability Federico Adolfi, Martina G. Vilas, Todd Wareham Published: 2024-10-10Area: Formal/TheoreticalCitations: 14 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2024-10-10 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (98%) | 14 |
| ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability Han Li, Jun Xu, Kai Zheng, Weijie Yu Published: 2024-10-15Area: Mechanistic Interp.Citations: 68 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-15 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (95%) | 68 |
| Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization Aaquib Syed, Abhay Sheshadri, Aidan Ewart, Gintare Karolina Dziugaite Published: 2024-10-16Area: Model EditingCitations: 19 Tags: ai-safety, empirical, interpretability, model-editing | 2024-10-16 | Model Editing | ai-safety, empirical, interpretability, model-editing | E6 / R3 (96%) | 19 |
| Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness Jingyi Cui, Qi Lei, Qi Zhang, Stefanie Jegelka Published: 2024-10-27Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (94%) | 5 |
| Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders Alasdair Paren, David Krueger, Fazl Barez, Luke Marks Published: 2024-11-02Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2024-11-02 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| Towards Unifying Interpretability and Control: Evaluation via Intervention Asma Ghandeharioun, Himabindu Lakkaraju, Suraj Srinivas, Usha Bhalla Published: 2024-11-07Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2024-11-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (94%) | 20 |
| Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation Erica K. Shimomoto, Kazuhiro Fukui, Lincon S. Souza, Pedro H. V. Valois Published: 2024-12-10Area: Representation AnalysisCitations: 2 Tags: ai-safety, empirical, interpretability, representation-analysis | 2024-12-10 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 2 |
| Large Language Model Safety: A Holistic Survey Bojian Jiang, Chuang Liu, Dan Shi, Deyi Xiong Published: 2024-12-23Area: Surveys & ReviewsCitations: 47 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-12-23 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R5 (99%) | 47 |
| Enhancing Automated Interpretability with Output-Centric Feature Descriptions Atticus Geiger, Chen Agassy, Mor Geva, Roy Mayan Published: 2025-01-14Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-14 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (92%) | 26 |
| Open Problems in Mechanistic Interpretability Adria Garriga-Alonso, Alejandro Ortega, Arthur Conmy, Atticus Geiger Published: 2025-01-27Area: Surveys & ReviewsCitations: 107 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-01-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 107 |
| Propositional Interpretability in Artificial Intelligence David J. Chalmers Published: 2025-01-27Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-01-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (94%) | 13 |
| Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning Anna Soligo, David Boyle, Pietro Ferraro Published: 2025-01-28Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 2 |
| Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers Laurence Aitchison, Lucy Farnik, Thomas Heap, Tim Lawson Published: 2025-01-29Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 26 |
| Building Bridges, Not Walls: Advancing Interpretability by Unifying Feature, Data, and Model Component Attribution Hima Lakkaraju, Shichang Zhang, Tessa Han, Usha Bhalla Published: 2025-01-31Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, position, surveys-reviews | 2025-01-31 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (97%) | 3 |
| Modular Training of Neural Networks aids Interpretability Alessandro Abate, Joan Velja, Maheep Chaudhary, Nandi Schoots Published: 2025-02-04Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-02-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 1 |
| A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models Arman Zarei, Barry Menglong Yao, Hongxuan Li, Keivan Rezaei Published: 2025-02-22Area: Surveys & ReviewsCitations: 20 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-02-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (96%) | 20 |
| Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? Francois Portet, Maxime Meloux, Maxime Peyrard, Silviu Maniu Published: 2025-02-28Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-02-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (94%) | 14 |