Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants Dami Choi, Daniel D. Johnson, Jacob Steinhardt, Sarah Schwettmann Published: 2025-12-17Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 3 |
| Provably Extracting the Features from a General Superposition Allen Liu Published: 2025-12-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-12-17 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E4 / R2 (96%) | - |
| From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda D. Nardi, F. Giarrusso, F. Pierucci, M. Bracale Syrnikov Published: 2025-12-16Area: Adversarial RobustnessCitations: 1 Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2025-12-16 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E5 / R3 (97%) | 1 |
| Interpretable and Steerable Concept Bottleneck Sparse Autoencoders Akshay Kulkarni, Kowshik Thopalli, Shusen Liu, Tsui-Wei Weng Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (96%) | - |
| Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis Amartya Hatua Published: 2025-12-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima Dianbo Liu, Harshvardhan Saini, Yiming Tang, Yizhen Liao Published: 2025-12-05Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | - |
| Sparse Attention Post-Training for Mechanistic Interpretability Anson Lei, Bernhard Schölkopf, Florent Draye, Ingmar Posner Published: 2025-12-05Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-05 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective Anne Lauscher, Jae Hee Lee, Stefano V. Albrecht Published: 2025-12-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | 2025-12-04 | Agent Safety | agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | E5 / R3 (94%) | 1 |
| Unsupervised decoding of encoded reasoning using language model interpretability Ching Fang, Samuel Marks Published: 2025-12-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 2 |
| Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Bianka Kowalska, Halina Kwaśnicka Published: 2025-11-24Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-11-24 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 1 |
| nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers Clément Dumas Published: 2025-11-18Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-11-18 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 1 |
| SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs Andrew Gordon, David Quarel, Peter Lai, Sean P. Fillingham Published: 2025-11-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2025-11-10 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| Atlas-Alignment: Making Interpretability Transferable Across Language Models Bruno Puri, Jim Berend, Sebastian Lapuschkin, Wojciech Samek Published: 2025-10-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | - |
| Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI Anthony Baez, Pat Pataranutaporn, Sheer Karny Published: 2025-10-31Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 1 |
| Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability Alex Oesterling, Claudio Mayrink Verdun, Flavio P. Calmon, Himabindu Lakkaraju Published: 2025-10-30Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-30 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (96%) | - |
| Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention Gustavo Penha, Hugues Bouchard, José Luis Redondo García, J Rosser Published: 2025-10-22Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-22 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Subliminal Corruption: Mechanisms, Thresholds, and Interpretability Reya Vir, Sarvesh Bhatnagar Published: 2025-10-22Area: Deception & FailureCitations: 1 Tags: ai-safety, alignment-training, deception-failure, empirical, interpretability | 2025-10-22 | Deception & Failure | ai-safety, alignment-training, deception-failure, empirical, interpretability | E5 / R3 (95%) | 1 |
| WeightLens and CircuitLens: Circuit Insights Towards Interpretability Beyond Activations Aakriti Jain, Ammar Ibrahim, Bruno Puri, Elena Golimblevskaia Published: 2025-10-16Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-10-16 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | - |
| Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders Benyou Wang, Difan Zou, Xu Wang, Yan Hu Published: 2025-10-04Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E7 / R4 (97%) | 2 |
| Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders Charibeth Cheng, Kriz Tahimic Published: 2025-10-03Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-03 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | - |
| AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features Mohammad Mahdi Khalili, Xudong Zhu, Zhihui Zhu Published: 2025-10-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Mechanistic Interpretability as Statistical Estimation: A Variance Analysis of EAP-IG François Portet, Maxime Méloux, Maxime Peyrard Published: 2025-10-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| Binary Sparse Coding for Interpretability Lucia Quirke, Nora Belrose, Stepan Shabalin Published: 2025-09-29Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (92%) | 1 |
| Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF Jing Liu Published: 2025-09-29Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, interpretability, theoretical | 2025-09-29 | Alignment Training | ai-safety, alignment-training, interpretability, theoretical | E5 / R3 (94%) | - |
| LLM Interpretability with Identifiable Temporal-Instantaneous Representation Jiaqi Sun, Kun Zhang, Xiangchen Song, Yujia Zheng Published: 2025-09-27Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| Analysis of Variational Sparse Autoencoders Yuxiao Li, Zachary Baker Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research Sean Trott Published: 2025-09-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-26 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| Binary Autoencoder for Mechanistic Interpretability of Large Language Models Brian M. Kurkoski, Hakaze Cho, Haolin Yang, Naoya Inoue Published: 2025-09-25Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-25 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Who is In Charge? Dissecting Role Conflicts in Instruction Following Siqi Zeng Published: 2025-09-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-09-23 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | - |
| Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models Katharina Simbeck, Mariam Mahran Published: 2025-09-22Area: Representation AnalysisCitations: 2 Tags: ai-safety, empirical, interpretability, representation-analysis | 2025-09-22 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 2 |