Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs Andrew Gordon, David Quarel, Peter Lai, Sean P. Fillingham Published: 2025-11-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2025-11-10 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers Cl茅ment Dumas Published: 2025-11-18Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-11-18 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | 1 |
| Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Bianka Kowalska, Halina Kwa艣nicka Published: 2025-11-24Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-11-24 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 1 |
| Unsupervised decoding of encoded reasoning using language model interpretability Ching Fang, Samuel Marks Published: 2025-12-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 2 |
| Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective Anne Lauscher, Jae Hee Lee, Stefano V. Albrecht Published: 2025-12-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | 2025-12-04 | Agent Safety | agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | E5 / R3 (94%) | 1 |
| A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima Dianbo Liu, Harshvardhan Saini, Yiming Tang, Yizhen Liao Published: 2025-12-05Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | - |
| Sparse Attention Post-Training for Mechanistic Interpretability Anson Lei, Bernhard Sch枚lkopf, Florent Draye, Ingmar Posner Published: 2025-12-05Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-05 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis Amartya Hatua Published: 2025-12-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Interpretable and Steerable Concept Bottleneck Sparse Autoencoders Akshay Kulkarni, Kowshik Thopalli, Shusen Liu, Tsui-Wei Weng Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (96%) | - |
| From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda D. Nardi, F. Giarrusso, F. Pierucci, M. Bracale Syrnikov Published: 2025-12-16Area: Adversarial RobustnessCitations: 1 Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2025-12-16 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E5 / R3 (97%) | 1 |
| Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers Adam Karvonen, Arnab Sen Sharma, Clement Dumas, Daniel Wen Published: 2025-12-17Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 2 |
| From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? Aaron Mueller, Andrew Lee, Dhanya Sridhar, Ekdeep Singh Lubana Published: 2025-12-17Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (94%) | 1 |
| Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants Dami Choi, Daniel D. Johnson, Jacob Steinhardt, Sarah Schwettmann Published: 2025-12-17Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-12-17 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 3 |
| Provably Extracting the Features from a General Superposition Allen Liu Published: 2025-12-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-12-17 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E4 / R2 (96%) | - |
| Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability Ge Yan, Tsui-Wei (Lily) Weng, Tuomas Oikarinen Published: 2025-12-19Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-19 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| The Dead Salmons of AI Interpretability Fran莽ois Portet, Giada Dirupo, Maxime M茅loux, Maxime Peyrard Published: 2025-12-21Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, interpretability, mechanistic-interp, position | 2025-12-21 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, position | E7 / R3 (93%) | 3 |
| Interpreting Transformers Through Attention Head Intervention Mason Kadem, Rong Zheng Published: 2026-01-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, survey | 2026-01-07 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E5 / R3 (95%) | - |
| SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis Xufeng Duan, Zhenguang G. Cai, Zihao Fu Published: 2026-01-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2026-01-12 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (97%) | - |
| Patterning: The Dual of Interpretability Daniel Murfet, George Wang Published: 2026-01-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-01-20 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R4 (95%) | - |
| Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions Usman Naseem Published: 2026-01-21Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2026-01-21 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E6 / R4 (96%) | 1 |
| Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits Been Kim, Drew Proud, Mani Malek, Neha Kalibhat Published: 2026-01-23Area: Model EditingCitations: - Tags: ai-safety, empirical, interpretability, model-editing | 2026-01-23 | Model Editing | ai-safety, empirical, interpretability, model-editing | E5 / R3 (91%) | - |
| DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders Baosong Yang, Bingqing Jiang, Difan Zou, Lingpeng Kong Published: 2026-02-05Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2026-02-05 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| Towards Worst-Case Guarantees with Scale-Aware Interpretability Alexander Stapleton, Andrew Mack, Anindita Maiti, Artemy Kolchinsky Published: 2026-02-05Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, position | 2026-02-05 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, position | E5 / R3 (94%) | - |
| Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints Andres Saurez, Dongsoo Har, Yousung Lee Published: 2026-02-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-02-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| Interpretability-by-Design with Accurate Locally Additive Models and Conditional Feature Effects Christos Diou, Dimitrios Kyriakopoulos, Dimitrios Rontogiannis, Giuseppe Casalicchio Published: 2026-02-18Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-02-18 | cs.LG | ai-safety, cslg, interpretability, preprint | E6 / R6 (91%) | - |
| Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy Bianca Raimondi, Maurizio Gabbrielli Published: 2026-02-19Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-02-19 | cs.AI | ai-safety, csai, interpretability, preprint | E9 / R6 (93%) | - |
| Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal Ihor Kendiukhov Published: 2026-02-19Area: q-bio.GNCitations: 153 Tags: ai-safety, interpretability, preprint, q-biogn, safety-evaluation | 2026-02-19 | q-bio.GN | ai-safety, interpretability, preprint, q-biogn, safety-evaluation | E10 / R8 (91%) | 153 |
| RamanSeg: Interpretability-driven Deep Learning on Raman Spectra for Cancer Diagnosis Anna M眉hlig, Anna Xylander, Chris Tomy, David Pertzborn Published: 2026-02-20Area: eess.IVCitations: 15 Tags: ai-safety, eessiv, interpretability, preprint | 2026-02-20 | eess.IV | ai-safety, eessiv, interpretability, preprint | E10 / R6 (93%) | 15 |
| Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling Norrakith Srisumrith, Sunantha Sodsee Published: 2026-02-22Area: cs.CRCitations: 28 Tags: ai-safety, cscr, interpretability, preprint | 2026-02-22 | cs.CR | ai-safety, cscr, interpretability, preprint | E14 / R11 (95%) | 28 |
| MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning Davis Brown, Gal Mishne, Helen Jenne, Henry Kvinge Published: 2026-02-24Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-02-24 | cs.LG | ai-safety, cslg, interpretability, preprint | E9 / R7 (91%) | - |