Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders Jiayi Yuan, Ninghao Liu, Wenlin Yao, Xiaoming Zhai Published: 2025-02-21Area: Mechanistic Interp.Citations: 20 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-02-21 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R4 (93%) | 20 |
| Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models Binxu Wang, Demba Ba, Ekdeep Singh Lubana, Isabel Papadimitriou Published: 2025-02-18Area: Mechanistic Interp.Citations: 32 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-18 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 32 |
| SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models Ali Payani, Fan Yang, Haiyan Zhao, Jing Ma Published: 2025-02-17Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 18 |
| Sparse Autoencoder Features for Classifications and Transferability Danielle S. Bitterman, Hugo Aerts, Jack Gallifant, Kuleen Sasse Published: 2025-02-17Area: Mechanistic Interp.Citations: 16 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 16 |
| Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis Benyou Wang, Difan Zou, Reynold Cheng, Wenyu Du Published: 2025-02-17Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (94%) | 10 |
| The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Haining Yu, Qiguang Chen, Wenbo Pan, Xiangyang Zhou Published: 2025-02-13Area: Mechanistic Interp.Citations: 8 Tags: adversarial-robustness, ai-safety, alignment-training, empirical, mechanistic-interp | 2025-02-13 | Mechanistic Interp. | adversarial-robustness, ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R4 (94%) | 8 |
| Deciphering Functions of Neurons in Vision-Language Models Cuiling Lan, Jiaqi Xu, Xuejin Chen, Yan Lu Published: 2025-02-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-02-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | - |
| Position-aware Automatic Circuit Discovery Aaron Mueller, David Bau, Hadas Orgad, Tal Haklay Published: 2025-02-07Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 6 |
| Sparse Autoencoders Do Not Find Canonical Units of Analysis Bart Bussmann, Curt Tigges, Joseph Bloom, Lee Sharkey Published: 2025-02-07Area: Mechanistic Interp.Citations: 43 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | 43 |
| Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment Harrish Thasarathan, Julian Forsyth, Konstantinos Derpanis, Matthew Kowal Published: 2025-02-06Area: Mechanistic Interp.Citations: 27 Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2025-02-06 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E4 / R3 (97%) | 27 |
| Analyze Feature Flow to Enhance Interpretation and Steering in Language Models Daniil Gavrilov, Daniil Laptev, Nikita Balagansky, Yaroslav Aksenov Published: 2025-02-05Area: Mechanistic Interp.Citations: 6 Tags: ai-safety, empirical, mechanistic-interp | 2025-02-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 6 |
| Modular Training of Neural Networks aids Interpretability Alessandro Abate, Joan Velja, Maheep Chaudhary, Nandi Schoots Published: 2025-02-04Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-02-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 1 |
| Low-Rank Adapting Models for Sparse Autoencoders Joshua Engels, Matthew Chen, Max Tegmark Published: 2025-01-31Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2025-01-31 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 4 |
| Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers Laurence Aitchison, Lucy Farnik, Thomas Heap, Tim Lawson Published: 2025-01-29Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 26 |
| SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders Bartosz Cywinski, Kamil Deja Published: 2025-01-29Area: Mechanistic Interp.Citations: 37 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-01-29 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 37 |
| Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning Anna Soligo, David Boyle, Pietro Ferraro Published: 2025-01-28Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 2 |
| Sparse Autoencoders Trained on the Same Data Learn Different Features Gon莽alo Paulo, Nora Belrose Published: 2025-01-28Area: Mechanistic Interp.Citations: 40 Tags: ai-safety, empirical, mechanistic-interp | 2025-01-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 40 |
| Propositional Interpretability in Artificial Intelligence David J. Chalmers Published: 2025-01-27Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-01-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (94%) | 13 |
| Enhancing Automated Interpretability with Output-Centric Feature Descriptions Atticus Geiger, Chen Agassy, Mor Geva, Roy Mayan Published: 2025-01-14Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-14 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (92%) | 26 |
| Mechanistic understanding and validation of large AI models with SemanticLens Jim Berend, Johanna Vielhaben, Maximilian Dreyer, Sebastian Lapuschkin Published: 2025-01-09Area: Mechanistic Interp.Citations: 29 Tags: ai-safety, mechanistic-interp, tool | 2025-01-09 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (95%) | 29 |
| Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Gouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo Published: 2025-01-09Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2025-01-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | 10 |
| Transformers Use Causal World Models in Maze-Solving Tasks Adrians Skapars, Alessandra Russo, Alex F. Spies, Katsumi Inoue Published: 2024-12-16Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, mechanistic-interp | 2024-12-16 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | 9 |
| BatchTopK Sparse Autoencoders Bart Bussmann, Neel Nanda, Patrick Leask Published: 2024-12-09Area: Mechanistic Interp.Citations: 64 Tags: ai-safety, empirical, mechanistic-interp | 2024-12-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 64 |
| Monet: Mixture of Monosemantic Experts for Transformers Jaewoo Kang, Jungwoo Park, Kee-Eung Kim, Young Jin Ahn Published: 2024-12-05Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, empirical, mechanistic-interp | 2024-12-05 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 9 |
| Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration Chun Hei Yip, Jason Gross, Lawrence Chan, Rajashree Agrawal Published: 2024-12-04Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, mechanistic-interp, theoretical | 2024-12-04 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 4 |
| Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks Adam Karvonen, Can Rager, Neel Nanda, Samuel Marks Published: 2024-11-28Area: Mechanistic Interp.Citations: 9 Tags: ai-safety, benchmark, mechanistic-interp, safety-evaluation | 2024-11-28 | Mechanistic Interp. | ai-safety, benchmark, mechanistic-interp, safety-evaluation | E5 / R3 (96%) | 9 |
| Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models Javier Ferrando, Neel Nanda, Oscar Obeso, Senthooran Rajamanoharan Published: 2024-11-21Area: Mechanistic Interp.Citations: 88 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 88 |
| Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders Charles O'Neill, David Klindt Published: 2024-11-20Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, mechanistic-interp, theoretical | 2024-11-20 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (97%) | 7 |
| JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit Chun Chen, Huiyu Xu, Kui Ren, Rui Zheng Published: 2024-11-17Area: Mechanistic Interp.Citations: 16 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2024-11-17 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E6 / R3 (93%) | 16 |
| Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning Bryce Hepner, David Wingate, Jared Wilson, Jeffrey Olmo Published: 2024-11-15Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2024-11-15 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 4 |