Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 211-236 of 236 papers (page 8 of 8)路 72 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Progress Measures for Grokking via Mechanistic Interpretability Jacob Steinhardt, Jess Smith, Lawrence Chan, Neel Nanda Published: 2023-01-12Area: Training DynamicsCitations: 680 Tags: ai-safety, empirical, interpretability, training-dynamics | 2023-01-12 | Training Dynamics | ai-safety, empirical, interpretability, training-dynamics | E5 / R3 (95%) | 680 |
| Tracr: Compiled Transformers as a Laboratory for Interpretability David Lindner, Janos Kramar, Matthew Rahtz, Sebastian Farquhar Published: 2023-01-12Area: Mechanistic Interp.Citations: 91 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2023-01-12 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R4 (95%) | 91 |
| Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Amir Zur, Aryaman Arora, Atticus Geiger, Christopher Potts Published: 2023-01-11Area: Mechanistic Interp.Citations: 118 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-01-11 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (95%) | 118 |
| Circumventing interpretability: How to defeat mind-readers Lee Sharkey Published: 2022-12-21Area: Deception & FailureCitations: 5 Tags: ai-safety, deception-failure, interpretability, position | 2022-12-21 | Deception & Failure | ai-safety, deception-failure, interpretability, position | E5 / R3 (93%) | 5 |
| Interpreting Neural Networks through the Polytope Lens Beren Millidge, Carlos Ram贸n Guevara, Connor Leahy, Dan Braun Published: 2022-11-22Area: Mechanistic Interp.Citations: 36 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2022-11-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 36 |
| Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt Published: 2022-11-01Area: Mechanistic Interp.Citations: 834 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2022-11-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (96%) | 834 |
| Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Anson Ho, Dylan Hadfield-Menell, Stephen Casper, Tilman R盲uker Published: 2022-07-27Area: Surveys & ReviewsCitations: 174 Tags: ai-safety, interpretability, survey, surveys-reviews | 2022-07-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E8 / R4 (94%) | 174 |
| Scaling Laws and Interpretability of Learning from Repeated Data Ben Mann, Catherine Olsson, Chris Olah, Danny Hernandez Published: 2022-05-21Area: Training DynamicsCitations: 148 Tags: ai-safety, empirical, interpretability, training-dynamics | 2022-05-21 | Training Dynamics | ai-safety, empirical, interpretability, training-dynamics | E5 / R3 (94%) | 148 |
| Robust Feature-Level Adversaries are Interpretability Tools Dylan Hadfield-Menell, Gabriel Kreiman, Max Nadeau, Stephen Casper Published: 2021-10-07Area: Adversarial RobustnessCitations: 34 Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2021-10-07 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E5 / R3 (94%) | 34 |
| An Interpretability Illusion for BERT Adam Pearce, Andy Coenen, Ann Yuan, Emily Reif Published: 2021-04-14Area: Representation AnalysisCitations: 96 Tags: ai-safety, empirical, interpretability, representation-analysis | 2021-04-14 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (96%) | 96 |
| Towards Falsifiable Interpretability Research Ari S. Morcos, Matthew L. Leavitt Published: 2020-10-22Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, interpretability, mechanistic-interp, position | 2020-10-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E6 / R4 (97%) | 74 |
| A Primer in BERTology: What We Know About How BERT Works Anna Rogers, Anna Rumshisky, Olga Kovaleva Published: 2020-02-27Area: Surveys & ReviewsCitations: 1772 Tags: ai-safety, interpretability, survey, surveys-reviews | 2020-02-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 1772 |
| Causal Scrubbing: A Method for Rigorously Testing Interpretability Hypotheses Adri脙 Garriga-Alonso, Ansh Radhakrishnan, Buck Shlegeris, Jenny Nitishinskaya Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Core Views on AI Safety: When, Why, What, and How Anthropic Published: -Area: Surveys & ReviewsCitations: - Tags: ai-safety, interpretability, position, surveys-reviews | - | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E6 / R4 (95%) | - |
| Cracking the Circuits: Mechanistic Interpretability in Large Language Models Dost Muhammad, Malika Bendechache, Muhammad Salman, Mushtaq Ali Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | - |
| Explaining AI through mechanistic interpretability Barnaby Crook, Lena K盲stner Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | - |
| From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models Etowah Adams, Liam Bai, Minji Lee, Mohammed AlQuraishi Published: -Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 32 |
| Gemma Scope 2: Comprehensive Suite of SAEs and Transcoders for Gemma 3 Arthur Conmy, Callum McDougall, Janos Kramar, Neel Nanda Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (98%) | - |
| Goodfire Ember: Scaling Interpretability for Frontier Model Alignment Atticus Geiger, Curt Tigges, Dan Braun, Daniel Balsam Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, alignment-training, interpretability, mechanistic-interp, tool | E4 / R3 (99%) | - |
| Language Models Can Explain Neurons in Language Models Dan Mossing, Gabriel Goh, Henk Tillman, Ilya Sutskever Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R2 (96%) | - |
| Neuronpedia: Interactive SAE Feature Explorer Johnny Lin Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E6 / R3 (95%) | - |
| SFAL: Semantic-Functional Alignment Scores for Distributional Evaluation of Auto-Interpretability in Sparse Autoencoders Andrea Seveso, Antonio Serino, Daniele Potert矛, Fabio Mercorio Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp, safety-evaluation | - | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (95%) | - |
| Sparse Autoencoders Find Partially Interpretable Features in Italian Small Language Models Alessandro Bondielli, Alessandro Lenci, Lucia C. Passaro Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| Survey on the Role of Mechanistic Interpretability in Generative AI Leonardo Ranaldi Published: -Area: Surveys & ReviewsCitations: 4 Tags: adversarial-robustness, ai-safety, interpretability, survey, surveys-reviews | - | Surveys & Reviews | adversarial-robustness, ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 4 |
| TransformerLens Joseph Bloom, Neel Nanda Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | - |
| Understanding RL Vision Chris Olah, Gabriel Goh, Jacob Hilton, Nick Cammarata Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | - |