Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework Amrithaa Ashok Kumar, Hao Gu, Jayvart Sharma, Ryan Lagasse Published: 2025-09-28Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (97%) | 2 |
| Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs Atharva Nijasure, James Allan, Tanya Chowdhury, Yair Zick Published: 2025-09-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| Measuring Sparse Autoencoder Feature Sensitivity Claire Tian, Katherine Tian, Nathan Hu Published: 2025-09-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2025-09-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | - |
| Binary Sparse Coding for Interpretability Lucia Quirke, Nora Belrose, Stepan Shabalin Published: 2025-09-29Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (92%) | 1 |
| AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features Mohammad Mahdi Khalili, Xudong Zhu, Zhihui Zhu Published: 2025-10-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Feature Identification via the Empirical NTK Jennifer Lin Published: 2025-10-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 2 |
| Mechanistic Interpretability as Statistical Estimation: A Variance Analysis of EAP-IG François Portet, Maxime Méloux, Maxime Peyrard Published: 2025-10-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders Charibeth Cheng, Kriz Tahimic Published: 2025-10-03Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-03 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | - |
| Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders Benyou Wang, Difan Zou, Xu Wang, Yan Hu Published: 2025-10-04Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E7 / R4 (97%) | 2 |
| Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language Angie Boggust, Arvind Satyanarayan, Dominik Moritz, Donghao Ren Published: 2025-10-07Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods Ahmad Dawar Hakimi, Barbara Plank, Hinrich Schütze, Leonor Veloso Published: 2025-10-08Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-08 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 2 |
| Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training Junyu Ren, T. Ed Li Published: 2025-10-09Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-10-09 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (97%) | - |
| Verifying Chain-of-Thought Reasoning via Its Computational Graph Naila Murray, Nicola Cancedda, Xianjun Yang, Yeskendir Koishekenov Published: 2025-10-10Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (97%) | 2 |
| WeightLens and CircuitLens: Circuit Insights Towards Interpretability Beyond Activations Aakriti Jain, Ammar Ibrahim, Bruno Puri, Elena Golimblevskaia Published: 2025-10-16Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-10-16 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | - |
| Finding Manifolds With Bilinear Autoencoders Thomas Dooms, Ward Gauderis Published: 2025-10-19Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R2 (94%) | 2 |
| ActivationReasoning: Logical Reasoning in Latent Activation Spaces Antonia Wüst, Felix Friedrich, Hikaru Shindo, Kristian Kersting Published: 2025-10-21Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R5 (97%) | 1 |
| DePass: Unified Feature Attributing by Simple Decomposed Forward Pass Biqing Qi, Bowen Zhou, Che Jiang, Kai Tian Published: 2025-10-21Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-10-21 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | - |
| Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention Gustavo Penha, Hugues Bouchard, José Luis Redondo García, J Rosser Published: 2025-10-22Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-22 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Mapping Faithful Reasoning in Language Models Andreas Damianou, Jiazheng Li, José Luis Redondo García, J Rosser Published: 2025-10-25Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-25 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 2 |
| PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization Di Wang, Huanyi Xie, Liangyu Wang, Lijie Hu Published: 2025-10-27Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-10-27 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 2 |
| Automatically Finding Rule-Based Neurons in OthelloGPT Adam Karvonen, Aditya Singh, Can Rager, Srujananjali Medicherla Published: 2025-10-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-10-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers Rabin Adhikari Published: 2025-10-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-10-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability Alex Oesterling, Claudio Mayrink Verdun, Flavio P. Calmon, Himabindu Lakkaraju Published: 2025-10-30Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-30 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (96%) | - |
| Atlas-Alignment: Making Interpretability Transferable Across Language Models Bruno Puri, Jim Berend, Sebastian Lapuschkin, Wojciech Samek Published: 2025-10-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | - |
| Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI Anthony Baez, Pat Pataranutaporn, Sheer Karny Published: 2025-10-31Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 1 |
| Addressing divergent representations from causal interventions on neural networks Alexa R. Tartaglini, Christopher Potts, Satchel Grant, Simon Jerome Han Published: 2025-11-06Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-11-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (96%) | - |
| APP: Accelerated Path Patching with Task-Specific Pruning Carsten Eickhoff, Frauke Andersen, Ruochen Zhang, William Rudman Published: 2025-11-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-11-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder Kaidi Xu, Song Wang, Tianlong Chen, Zhen Tan Published: 2025-11-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-11-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts Bei Wang, Kowshik Thopalli, Shusen Liu, Xinyuan Yan Published: 2025-11-08Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, tool | 2025-11-08 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E6 / R3 (96%) | - |
| SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs Andrew Gordon, David Quarel, Peter Lai, Sean P. Fillingham Published: 2025-11-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2025-11-10 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E5 / R3 (97%) | - |