Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Exploring the Robustness of Model-Graded Evaluations and Automated Interpretability Ondrej Kvapil, Simon Lermen Published: 2023-11-26Area: Safety EvaluationCitations: 3 Tags: ai-safety, empirical, interpretability, safety-evaluation | 2023-11-26 | Safety Evaluation | ai-safety, empirical, interpretability, safety-evaluation | E5 / R3 (93%) | 3 |
| Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2023-11-28Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-11-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (93%) | 41 |
| FlexModel: A Framework for Interpretability of Distributed Large Language Models David B. Emerson, John Willes, Matthew Choi, Muhammad Adil Asif Published: 2023-12-05Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (95%) | 1 |
| Interpretability Illusions in the Generalization of Simplified Models Andrew Lampinen, Asma Ghandeharioun, Dan Friedman, Danqi Chen Published: 2023-12-06Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-12-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 20 |
| Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability Jatin Nainani Published: 2024-01-08Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-08 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 3 |
| Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models Adam Pearce, Asma Ghandeharioun, Avi Caciularu, Lucas Dixon Published: 2024-01-11Area: Representation AnalysisCitations: 174 Tags: ai-safety, empirical, interpretability, representation-analysis | 2024-01-11 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 174 |
| A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments Aryaman Arora, Atticus Geiger, Christopher Potts, Jing Huang Published: 2024-01-23Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 9 |
| Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability Alex Grzankowski Published: 2024-01-31Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-01-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E8 / R4 (94%) | 10 |
| Challenges in Mechanistically Interpreting Model Representations James Dao, Satvik Golechha Published: 2024-02-06Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-02-06 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (93%) | 4 |
| Opening the AI black box: program synthesis via mechanistic interpretability Anish Mudide, Chloe Loughridge, Eric J. Michaud, Isaac Liao Published: 2024-02-07Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (96%) | 19 |
| Towards Uncovering How Large Language Model Works: An Explainability Perspective Fan Yang, Haiyan Zhao, Himabindu Lakkaraju, Mengnan Du Published: 2024-02-16Area: Surveys & ReviewsCitations: 26 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-02-16 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E5 / R3 (93%) | 26 |
| CausalGym: Benchmarking Causal Interpretability Methods on Linguistic Tasks Aryaman Arora, Christopher Potts, Dan Jurafsky Published: 2024-02-19Area: Safety EvaluationCitations: 36 Tags: ai-safety, benchmark, interpretability, safety-evaluation | 2024-02-19 | Safety Evaluation | ai-safety, benchmark, interpretability, safety-evaluation | E5 / R3 (95%) | 36 |
| Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT Qinyuan Cheng, Qiong Tang, Tianxiang Sun, Xipeng Qiu Published: 2024-02-19Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-19 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 25 |
| RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations Atticus Geiger, Christopher Potts, Jing Huang, Mor Geva Published: 2024-02-27Area: Safety EvaluationCitations: 61 Tags: ai-safety, benchmark, interpretability, safety-evaluation | 2024-02-27 | Safety Evaluation | ai-safety, benchmark, interpretability, safety-evaluation | E4 / R3 (98%) | 61 |
| pyvene: A Library for Understanding and Improving PyTorch Models via Interventions Aryaman Arora, Atticus Geiger, Christopher D. Manning, Christopher Potts Published: 2024-03-12Area: Mechanistic Interp.Citations: 44 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-03-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 44 |
| MAIA: A Multimodal Automated Interpretability Agent Achyuta Rajaram, Antonio Torralba, Evan Hernandez, Franklin Wang Published: 2024-04-22Area: Mechanistic Interp.Citations: 45 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2024-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E6 / R4 (95%) | 45 |
| Mechanistic Interpretability for AI Safety — A Review Efstratios Gavves, Leonard Bereska Published: 2024-04-22Area: Surveys & ReviewsCitations: 335 Tags: ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | 2024-04-22 | Surveys & Reviews | ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 335 |
| How to Use and Interpret Activation Patching Neel Nanda, Stefan Heimersheim Published: 2024-04-23Area: Mechanistic Interp.Citations: 109 Tags: ai-safety, interpretability, mechanistic-interp, survey | 2024-04-23 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E6 / R4 (95%) | 109 |
| A Primer on the Inner Workings of Transformer-based Language Models Arianna Bisazza, Gabriele Sarti, Javier Ferrando, Marta R. Costa-jussà Published: 2024-04-30Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-04-30 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (95%) | 80 |
| How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability Alejandro Maté, Jorge García-Carrasco, Juan Trujillo Published: 2024-05-07Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 13 |
| Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2024-05-14Area: Mechanistic Interp.Citations: 66 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-05-14 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | 66 |
| Using Degeneracy in the Loss Landscape for Mechanistic Interpretability Cindy Wu, Dan Braun, Jake Mendel, Kaarel Hänni Published: 2024-05-17Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-05-17 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 11 |
| From Neurons to Neutrons: A Case Study in Interpretability Mike Williams, Niklas Nolte, Ouail Kitouni, Sokratis Trifinopoulos Published: 2024-05-27Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-05-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 4 |
| Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience David Poeppel, Federico Adolfi, Gemma Roig, Martina G. Vilas Published: 2024-06-03Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-06-03 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (94%) | 10 |
| Compact Proofs of Model Performance via Mechanistic Interpretability Alex Gibson, Chun Hei Yip, Euan Ong, Jason Gross Published: 2024-06-17Area: Formal/TheoreticalCitations: 12 Tags: ai-safety, empirical, formaltheoretical, interpretability | 2024-06-17 | Formal/Theoretical | ai-safety, empirical, formaltheoretical, interpretability | E4 / R3 (92%) | 12 |
| Transcoders Find Interpretable LLM Feature Circuits Jacob Dunefsky, Neel Nanda, Philippe Chlenski Published: 2024-06-17Area: Mechanistic Interp.Citations: 102 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-06-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 102 |
| Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models Jun Zhao, Kang Liu, Pengfei Cao, Tianyi Men Published: 2024-06-23Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-06-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 19 |
| A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Abulhair Saparov, Daking Rai, Shi Feng, Yilun Zhou Published: 2024-07-02Area: Surveys & ReviewsCitations: 91 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-07-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 91 |
| Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Aaron Mueller Published: 2024-07-05Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-07-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 18 |
| Transformer Circuit Faithfulness Metrics are not Robust Bilal Chughtai, Joseph Miller, William Saunders Published: 2024-07-11Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-07-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 10 |