Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Interpreting Transformers Through Attention Head Intervention Mason Kadem, Rong Zheng Published: 2026-01-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, survey | 2026-01-07 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E5 / R3 (95%) | - |
| Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders Qi Su, Ruikang Zhang, Shuo Wang Published: 2026-01-06Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Attribution-Guided Distillation of Matryoshka Sparse Autoencoders Cristina P. Martin-Linares, Jonathan P. Ling Published: 2025-12-31Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-12-31 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| On the geometry and topology of representations: the manifolds of modular addition Colin Daniels, Doina Precup, Gabriela Moisescu-Pareja, Gavin McCracken Published: 2025-12-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-31 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | - |
| Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization Kerem Zaman, Shashank Srivastava Published: 2025-12-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (97%) | - |
| Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits Amirhosein Ghasemabadi, Di Niu Published: 2025-12-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (96%) | - |
| The Dead Salmons of AI Interpretability Fran莽ois Portet, Giada Dirupo, Maxime M茅loux, Maxime Peyrard Published: 2025-12-21Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, interpretability, mechanistic-interp, position | 2025-12-21 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E7 / R3 (93%) | 3 |
| Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability Ge Yan, Tsui-Wei (Lily) Weng, Tuomas Oikarinen Published: 2025-12-19Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-19 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers Adam Karvonen, Arnab Sen Sharma, Clement Dumas, Daniel Wen Published: 2025-12-17Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 2 |
| From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? Aaron Mueller, Andrew Lee, Dhanya Sridhar, Ekdeep Singh Lubana Published: 2025-12-17Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 1 |
| Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants Dami Choi, Daniel D. Johnson, Jacob Steinhardt, Sarah Schwettmann Published: 2025-12-17Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 3 |
| Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability Efstratios Gavves, Leonard Bereska, Reza Samavi, Zoe Tzifa-Kratira Published: 2025-12-15Area: Mechanistic Interp.Citations: 1 Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-12-15 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 1 |
| Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality Eric P. Xing, Guangyi Chen, Kun Zhang, Lingjing Kong Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2025-12-11 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| Interpretable and Steerable Concept Bottleneck Sparse Autoencoders Akshay Kulkarni, Kowshik Thopalli, Shusen Liu, Tsui-Wei Weng Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (96%) | - |
| Multi-Granular Node Pruning for Circuit Discovery A.B. Siddique, Hammad Rizwan, Hassan Sajjad, Muhammad Umair Haider Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (95%) | - |
| Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders Dianyun Wang, Huijia Wu, Huining Li, Jiaming Lyu Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders Aidong Zhang, Bohan Liu, Guangzhi Xiong, Sanchit Sinha Published: 2025-12-09Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | - |
| A Geometric Unification of Concept Learning with Concept Cones Alexandre Rocchi-Henry, Gianni Franchi, Thomas Fel Published: 2025-12-08Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, mechanistic-interp, theoretical | 2025-12-08 | Mechanistic Interp. | ai-safety, alignment-training, mechanistic-interp, theoretical | E5 / R4 (94%) | - |
| Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis Amartya Hatua Published: 2025-12-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima Dianbo Liu, Harshvardhan Saini, Yiming Tang, Yizhen Liao Published: 2025-12-05Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | - |
| Sparse Attention Post-Training for Mechanistic Interpretability Anson Lei, Bernhard Sch枚lkopf, Florent Draye, Ingmar Posner Published: 2025-12-05Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-05 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| AlignSAE: Concept-Aligned Sparse Autoencoders Jinhe Bi, Liangming Pan, Mihai Surdeanu, Minglai Yang Published: 2025-12-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-12-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 2 |
| Enforcing Orderedness in SAEs to Improve Feature Consistency Alex Quach, John J. Yang, Nithin Parsan, Sophie L. Wang Published: 2025-12-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-12-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Unsupervised decoding of encoded reasoning using language model interpretability Ching Fang, Samuel Marks Published: 2025-12-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 2 |
| SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models Jiaojiao Han, Mengnan Du, Mingyu Jin, Wujiang Xu Published: 2025-11-25Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-11-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 1 |
| Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models Aaron Mueller, Dana Arad, Gabriele Sarti, Hanjie Chen Published: 2025-11-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, mechanistic-interp | 2025-11-23 | Mechanistic Interp. | ai-safety, benchmark, mechanistic-interp | E6 / R3 (95%) | - |
| BlockCert: Certified Blockwise Extraction of Transformer Mechanisms Sandro Andric Published: 2025-11-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, tool | 2025-11-20 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (99%) | - |
| nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers Cl茅ment Dumas Published: 2025-11-18Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-11-18 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 1 |
| Weight-Sparse Transformers Have Interpretable Circuits Achyuta Rajaram, Bowen Baker, Dan Mossing, Jacob Coxon Published: 2025-11-17Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, empirical, mechanistic-interp | 2025-11-17 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | 14 |
| Decomposition of Small Transformer Models Casper L. Christensen, Logan Riggs Smith Published: 2025-11-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-11-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (94%) | - |