Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT Qinyuan Cheng, Qiong Tang, Tianxiang Sun, Xipeng Qiu Published: 2024-02-19Area: Mechanistic Interp.Citations: 25 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-19 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 25 |
| Towards Uncovering How Large Language Model Works: An Explainability Perspective Fan Yang, Haiyan Zhao, Himabindu Lakkaraju, Mengnan Du Published: 2024-02-16Area: Surveys & ReviewsCitations: 26 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-02-16 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E5 / R3 (93%) | 26 |
| Opening the AI black box: program synthesis via mechanistic interpretability Anish Mudide, Chloe Loughridge, Eric J. Michaud, Isaac Liao Published: 2024-02-07Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-02-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (96%) | 19 |
| Challenges in Mechanistically Interpreting Model Representations James Dao, Satvik Golechha Published: 2024-02-06Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-02-06 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E5 / R3 (93%) | 4 |
| Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability Alex Grzankowski Published: 2024-01-31Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, position | 2024-01-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E8 / R4 (94%) | 10 |
| A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments Aryaman Arora, Atticus Geiger, Christopher Potts, Jing Huang Published: 2024-01-23Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-23 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 9 |
| Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models Adam Pearce, Asma Ghandeharioun, Avi Caciularu, Lucas Dixon Published: 2024-01-11Area: Representation AnalysisCitations: 174 Tags: ai-safety, empirical, interpretability, representation-analysis | 2024-01-11 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 174 |
| Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability Jatin Nainani Published: 2024-01-08Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-01-08 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 3 |
| Interpretability Illusions in the Generalization of Simplified Models Andrew Lampinen, Asma Ghandeharioun, Dan Friedman, Danqi Chen Published: 2023-12-06Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-12-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 20 |
| FlexModel: A Framework for Interpretability of Distributed Large Language Models David B. Emerson, John Willes, Matthew Choi, Muhammad Adil Asif Published: 2023-12-05Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (95%) | 1 |
| Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Aleksandar Makelov, Georg Lange, Neel Nanda Published: 2023-11-28Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-11-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (93%) | 41 |
| Exploring the Robustness of Model-Graded Evaluations and Automated Interpretability Ondrej Kvapil, Simon Lermen Published: 2023-11-26Area: Safety EvaluationCitations: 3 Tags: ai-safety, empirical, interpretability, safety-evaluation | 2023-11-26 | Safety Evaluation | ai-safety, empirical, interpretability, safety-evaluation | E5 / R3 (93%) | 3 |
| AI Alignment: A Comprehensive Survey Aidan O'Gara, Borong Zhang, Boyuan Chen, Brian Tse Published: 2023-10-30Area: Surveys & ReviewsCitations: 320 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2023-10-30 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R4 (97%) | 320 |
| Codebook Features: Sparse and Discrete Interpretability for Neural Networks Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman Published: 2023-10-26Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 41 |
| Identifying Interpretable Visual Features in Artificial and Biological Neural Systems David Klindt, Francisco Acosta, Fr茅d茅ric Poitevin, Nina Miolane Published: 2023-10-17Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 10 |
| Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Fred Zhang, Neel Nanda Published: 2023-09-27Area: Mechanistic Interp.Citations: 193 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-09-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | 193 |
| Large Language Model Alignment: A Survey Chuang Liu, Deyi Xiong, Renren Jin, Tianhao Shen Published: 2023-09-26Area: Surveys & ReviewsCitations: 292 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | 2023-09-26 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 292 |
| FIND: A Function Description Benchmark for Evaluating Interpretability Methods Antonio Torralba, David Bau, Jacob Andreas, Joanna Materzynska Published: 2023-09-07Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2023-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E4 / R3 (94%) | 32 |
| Provably safe systems: the only path to controllable AGI Max Tegmark, Steve Omohundro Published: 2023-09-05Area: Formal/TheoreticalCitations: 38 Tags: ai-safety, formaltheoretical, interpretability, position | 2023-09-05 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, position | E7 / R4 (96%) | 38 |
| Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP Aryaman Arora, Paul Pu Liang, Rohan Pandey, Vedant Palit Published: 2023-08-27Area: Mechanistic Interp.Citations: 47 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-08-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 47 |
| The Hydra Effect: Emergent Self-repair in Language Model Computations Janos Kramar, Matthew Rahtz, Shane Legg, Thomas McGrath Published: 2023-07-28Area: Mechanistic Interp.Citations: 96 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 96 |
| Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla Geoffrey Irving, J脙隆nos Kram脙隆r, Matthew Rahtz, Neel Nanda Published: 2023-07-18Area: Mechanistic Interp.Citations: 144 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-18 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 144 |
| Neuron to Graph: Interpreting Language Model Neurons at Scale Alex Foote, Esben Kran, Fazl Barez, Ioannis Konstas Published: 2023-05-31Area: Mechanistic Interp.Citations: 28 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-05-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | 28 |
| Interpretability at Scale: Identifying Causal Mechanisms in Alpaca Atticus Geiger, Christopher Potts, Noah D. Goodman, Thomas Icard Published: 2023-05-15Area: Mechanistic Interp.Citations: 113 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2023-05-15 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 113 |
| A Technical Note on Bilinear Layers for Interpretability Lee Sharkey Published: 2023-05-05Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-05-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 10 |
| Seeing is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability Eric Gan, Max Tegmark, Ziming Liu Published: 2023-05-04Area: Mechanistic Interp.Citations: 52 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-05-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 52 |
| Towards Automated Circuit Discovery for Mechanistic Interpretability Adria Garriga-Alonso, Aengus Lynch, Arthur Conmy, Augustine N. Mavor-Parker Published: 2023-04-28Area: Mechanistic Interp.Citations: 485 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-04-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 485 |
| N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models Alex Foote, Esben Kran, Fazl Barez, Ionnis Konstas Published: 2023-04-22Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 4 |
| Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Aviya Skowron, Edward Raff, Eric Hallahan, Hailey Schoelkopf Published: 2023-04-03Area: Training DynamicsCitations: 1708 Tags: ai-safety, interpretability, tool, training-dynamics | 2023-04-03 | Training Dynamics | ai-safety, interpretability, tool, training-dynamics | E5 / R3 (99%) | 1708 |
| Red Teaming Deep Neural Networks with Feature Synthesis Tools Dylan Hadfield-Menell, Jiawei Li, Kaivalya Hariharan, Kevin Zhang Published: 2023-02-08Area: Safety EvaluationCitations: 21 Tags: ai-safety, benchmark, interpretability, red-teaming, safety-evaluation | 2023-02-08 | Safety Evaluation | ai-safety, benchmark, interpretability, red-teaming, safety-evaluation | E6 / R3 (96%) | 21 |