Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Ablation is Not Enough to Emulate DPO: How Neuron Dynamics Drive Toxicity Reduction Adam Mahdi, Filip Sondej, Harry Mayne, Yushi Yang Published: 2024-11-10Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R2 (95%) | 3 |
| Towards Unifying Interpretability and Control: Evaluation via Intervention Asma Ghandeharioun, Himabindu Lakkaraju, Suraj Srinivas, Usha Bhalla Published: 2024-11-07Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2024-11-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (94%) | 20 |
| A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning Cyrus Rashtchian, Enming Luo, Guan Zhe Hong, Nishanth Dikkala Published: 2024-11-06Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (94%) | 4 |
| Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders Kola Ayonrinde Published: 2024-11-04Area: Mechanistic Interp.Citations: 8 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 8 |
| Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders Alasdair Paren, David Krueger, Fazl Barez, Luke Marks Published: 2024-11-02Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2024-11-02 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models Aashiq Muhamed, Mona Diab, Virginia Smith Published: 2024-11-01Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 10 |
| Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics Aaron Mueller, Anja Reusch, Yaniv Nikankin, Yonatan Belinkov Published: 2024-10-28Area: Mechanistic Interp.Citations: 74 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 74 |
| Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups Davide Ghilardi, Federico Belotti, Marco Molinari Published: 2024-10-28Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 10 |
| One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models Antonio Mari, Caglar Gulcehre, Chris Wendler, David Bau Published: 2024-10-28Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (95%) | 18 |
| Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness Jingyi Cui, Qi Lei, Qi Zhang, Stefanie Jegelka Published: 2024-10-27Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (94%) | 5 |
| Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders Frances Liu, Junxuan Wang, Lingjie Chen, Qipeng Guo Published: 2024-10-27Area: Mechanistic Interp.Citations: 91 Tags: ai-safety, mechanistic-interp, tool | 2024-10-27 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (97%) | 91 |
| Decomposing The Dark Matter of Sparse Autoencoders Joshua Engels, Logan Smith, Max Tegmark Published: 2024-10-18Area: Mechanistic Interp.Citations: 33 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (93%) | 33 |
| Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models Erik Cambria, Ranjan Satapathy, Wei Jie Yeo Published: 2024-10-18Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 4 |
| Automatically Interpreting Millions of Features in Large Language Models Alex Mallen, Caden Juang, Gon脙搂alo Paulo, Nora Belrose Published: 2024-10-17Area: Mechanistic Interp.Citations: 69 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 69 |
| On the Role of Attention Heads in Large Language Model Safety Fei Huang, Haiyang Yu, Junfeng Fang, Kun Wang Published: 2024-10-17Area: Mechanistic Interp.Citations: 46 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 46 |
| Analyzing (In)Abilities of SAEs via Formal Languages Abhinav Menon, David Krueger, Ekdeep Singh Lubana, Manish Shrivastava Published: 2024-10-15Area: Mechanistic Interp.Citations: 15 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-15 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (95%) | 15 |
| ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability Han Li, Jun Xu, Kai Zheng, Weijie Yu Published: 2024-10-15Area: Mechanistic Interp.Citations: 68 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-15 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (95%) | 68 |
| The Persian Rug: solving toy models of superposition using large-scale symmetries Aditya Cowsik, Alex Infanger, Kfir Dolev Published: 2024-10-15Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2024-10-15 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| Bilinear MLPs Enable Weight-Based Mechanistic Interpretability Alice Rigg, Jose M. Oramas, Lee Sharkey, Michael T. Pearce Published: 2024-10-10Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| Efficient Dictionary Learning with Switch Sparse Autoencoders Anish Mudide, Christian Schroeder de Witt, Eric J. Michaud, Joshua Engels Published: 2024-10-10Area: Mechanistic Interp.Citations: 31 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 31 |
| Mechanistic Permutability: Match Features Across Layers Daniil Gavrilov, Ian Maksimov, Nikita Balagansky Published: 2024-10-10Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (98%) | 14 |
| The Geometry of Concepts: Sparse Autoencoder Feature Structure David D. Baek, Eric J. Michaud, Joshua Engels, Max Tegmark Published: 2024-10-10Area: Mechanistic Interp.Citations: 40 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 40 |
| SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders Anisoara Calinescu, Christian Schroeder de Witt, Constantin Venhoff, Philip Torr Published: 2024-10-09Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2024-10-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E6 / R3 (96%) | 1 |
| Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models Ashkan Khakzar, Austin Meek, David Krueger, Fazl Barez Published: 2024-10-09Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 11 |
| Gradient Routing: Masking Gradients to Localize Computation in Neural Networks Alexander Matt Turner, Alex Cloud, Evzen Wybitul, Jacob Goldman-Wetzler Published: 2024-10-06Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-06 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 18 |
| Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language Anthony Costarelli, Mat Allen, Severin Field Published: 2024-10-03Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-03 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 5 |
| Sparse Attention Decomposition Applied to Circuit Tracing Gabriel Franco, Mark Crovella Published: 2024-10-01Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, mechanistic-interp | 2024-10-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (97%) | 3 |
| A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders David Chanin, Hardik Bhatnagar, James Wilken-Smith, Joseph Bloom Published: 2024-09-22Area: Mechanistic Interp.Citations: 83 Tags: ai-safety, empirical, mechanistic-interp | 2024-09-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 83 |
| Optimal Ablation for Interpretability Lucas Janson, Maximilian Li Published: 2024-09-16Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-09-16 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | 14 |
| TracrBench: Generating Interpretability Testbeds with Large Language Models Hannes Thurnherr, J茅r茅my Scheurer Published: 2024-09-07Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2024-09-07 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (97%) | 4 |