Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 236 papers (page 1 of 8)路 244 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Primer in BERTology: What We Know About How BERT Works Anna Rogers, Anna Rumshisky, Olga Kovaleva Published: 2020-02-27Area: Surveys & ReviewsCitations: 1772 Tags: ai-safety, interpretability, survey, surveys-reviews | 2020-02-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 1772 |
| Towards Falsifiable Interpretability Research Ari S. Morcos, Matthew L. Leavitt Published: 2020-10-22Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, interpretability, mechanistic-interp, position | 2020-10-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E6 / R4 (97%) | 74 |
| An Interpretability Illusion for BERT Adam Pearce, Andy Coenen, Ann Yuan, Emily Reif Published: 2021-04-14Area: Representation AnalysisCitations: 96 Tags: ai-safety, empirical, interpretability, representation-analysis | 2021-04-14 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (96%) | 96 |
| Robust Feature-Level Adversaries are Interpretability Tools Dylan Hadfield-Menell, Gabriel Kreiman, Max Nadeau, Stephen Casper Published: 2021-10-07Area: Adversarial RobustnessCitations: 34 Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2021-10-07 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E5 / R3 (94%) | 34 |
| Scaling Laws and Interpretability of Learning from Repeated Data Ben Mann, Catherine Olsson, Chris Olah, Danny Hernandez Published: 2022-05-21Area: Training DynamicsCitations: 148 Tags: ai-safety, empirical, interpretability, training-dynamics | 2022-05-21 | Training Dynamics | ai-safety, empirical, interpretability, training-dynamics | E5 / R3 (94%) | 148 |
| Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Anson Ho, Dylan Hadfield-Menell, Stephen Casper, Tilman R盲uker Published: 2022-07-27Area: Surveys & ReviewsCitations: 174 Tags: ai-safety, interpretability, survey, surveys-reviews | 2022-07-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E8 / R4 (94%) | 174 |
| Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt Published: 2022-11-01Area: Mechanistic Interp.Citations: 834 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2022-11-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (96%) | 834 |
| Interpreting Neural Networks through the Polytope Lens Beren Millidge, Carlos Ram贸n Guevara, Connor Leahy, Dan Braun Published: 2022-11-22Area: Mechanistic Interp.Citations: 36 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2022-11-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 36 |
| Circumventing interpretability: How to defeat mind-readers Lee Sharkey Published: 2022-12-21Area: Deception & FailureCitations: 5 Tags: ai-safety, deception-failure, interpretability, position | 2022-12-21 | Deception & Failure | ai-safety, deception-failure, interpretability, position | E5 / R3 (93%) | 5 |
| Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Amir Zur, Aryaman Arora, Atticus Geiger, Christopher Potts Published: 2023-01-11Area: Mechanistic Interp.Citations: 118 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-01-11 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (95%) | 118 |
| Progress Measures for Grokking via Mechanistic Interpretability Jacob Steinhardt, Jess Smith, Lawrence Chan, Neel Nanda Published: 2023-01-12Area: Training DynamicsCitations: 680 Tags: ai-safety, empirical, interpretability, training-dynamics | 2023-01-12 | Training Dynamics | ai-safety, empirical, interpretability, training-dynamics | E5 / R3 (95%) | 680 |
| Tracr: Compiled Transformers as a Laboratory for Interpretability David Lindner, Janos Kramar, Matthew Rahtz, Sebastian Farquhar Published: 2023-01-12Area: Mechanistic Interp.Citations: 91 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-01-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R4 (95%) | 91 |
| Red Teaming Deep Neural Networks with Feature Synthesis Tools Dylan Hadfield-Menell, Jiawei Li, Kaivalya Hariharan, Kevin Zhang Published: 2023-02-08Area: Safety EvaluationCitations: 21 Tags: ai-safety, benchmark, interpretability, red-teaming, safety-evaluation | 2023-02-08 | Safety Evaluation | ai-safety, benchmark, interpretability, red-teaming, safety-evaluation | E6 / R3 (96%) | 21 |
| Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Aviya Skowron, Edward Raff, Eric Hallahan, Hailey Schoelkopf Published: 2023-04-03Area: Training DynamicsCitations: 1708 Tags: ai-safety, interpretability, tool, training-dynamics | 2023-04-03 | Training Dynamics | ai-safety, interpretability, tool, training-dynamics | E5 / R3 (99%) | 1708 |
| N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models Alex Foote, Esben Kran, Fazl Barez, Ionnis Konstas Published: 2023-04-22Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-04-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 4 |
| Towards Automated Circuit Discovery for Mechanistic Interpretability Adria Garriga-Alonso, Aengus Lynch, Arthur Conmy, Augustine N. Mavor-Parker Published: 2023-04-28Area: Mechanistic Interp.Citations: 485 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-04-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 485 |
| Seeing is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability Eric Gan, Max Tegmark, Ziming Liu Published: 2023-05-04Area: Mechanistic Interp.Citations: 52 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-05-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 52 |
| A Technical Note on Bilinear Layers for Interpretability Lee Sharkey Published: 2023-05-05Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-05-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 10 |
| Interpretability at Scale: Identifying Causal Mechanisms in Alpaca Atticus Geiger, Christopher Potts, Noah D. Goodman, Thomas Icard Published: 2023-05-15Area: Mechanistic Interp.Citations: 113 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2023-05-15 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | 113 |
| Neuron to Graph: Interpreting Language Model Neurons at Scale Alex Foote, Esben Kran, Fazl Barez, Ioannis Konstas Published: 2023-05-31Area: Mechanistic Interp.Citations: 28 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-05-31 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | 28 |
| Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla Geoffrey Irving, J脙隆nos Kram脙隆r, Matthew Rahtz, Neel Nanda Published: 2023-07-18Area: Mechanistic Interp.Citations: 144 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-18 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 144 |
| The Hydra Effect: Emergent Self-repair in Language Model Computations Janos Kramar, Matthew Rahtz, Shane Legg, Thomas McGrath Published: 2023-07-28Area: Mechanistic Interp.Citations: 96 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-07-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 96 |
| Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP Aryaman Arora, Paul Pu Liang, Rohan Pandey, Vedant Palit Published: 2023-08-27Area: Mechanistic Interp.Citations: 47 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-08-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 47 |
| Provably safe systems: the only path to controllable AGI Max Tegmark, Steve Omohundro Published: 2023-09-05Area: Formal/TheoreticalCitations: 38 Tags: ai-safety, formaltheoretical, interpretability, position | 2023-09-05 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, position | E7 / R4 (96%) | 38 |
| FIND: A Function Description Benchmark for Evaluating Interpretability Methods Antonio Torralba, David Bau, Jacob Andreas, Joanna Materzynska Published: 2023-09-07Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2023-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E4 / R3 (94%) | 32 |
| Large Language Model Alignment: A Survey Chuang Liu, Deyi Xiong, Renren Jin, Tianhao Shen Published: 2023-09-26Area: Surveys & ReviewsCitations: 292 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | 2023-09-26 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 292 |
| Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Fred Zhang, Neel Nanda Published: 2023-09-27Area: Mechanistic Interp.Citations: 193 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-09-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | 193 |
| Identifying Interpretable Visual Features in Artificial and Biological Neural Systems David Klindt, Francisco Acosta, Fr茅d茅ric Poitevin, Nina Miolane Published: 2023-10-17Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 10 |
| Codebook Features: Sparse and Discrete Interpretability for Neural Networks Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman Published: 2023-10-26Area: Mechanistic Interp.Citations: 41 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2023-10-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 41 |
| AI Alignment: A Comprehensive Survey Aidan O'Gara, Borong Zhang, Boyuan Chen, Brian Tse Published: 2023-10-30Area: Surveys & ReviewsCitations: 320 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2023-10-30 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R4 (97%) | 320 |