Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Evaluating SAE interpretability without explanations Gonçalo Paulo, Nora Belrose Published: 2025-07-11Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-07-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | 1 |
| The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability? Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel Published: 2025-07-11Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-07-11 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 11 |
| Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning Adam Karvonen, Caden Juang, Helena Casademunt, Neel Nanda Published: 2025-07-22Area: Model EditingCitations: 13 Tags: ai-safety, alignment-training, empirical, interpretability, model-editing | 2025-07-22 | Model Editing | ai-safety, alignment-training, empirical, interpretability, model-editing | E6 / R3 (93%) | 13 |
| Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Jinku Li, Rui Jiao, Yue Zhang Published: 2025-07-25Area: Deception & FailureCitations: - Tags: ai-safety, deception-failure, empirical, interpretability | 2025-07-25 | Deception & Failure | ai-safety, deception-failure, empirical, interpretability | E5 / R3 (96%) | - |
| A Review of Developmental Interpretability in Large Language Models Ihor Kendiukhov Published: 2025-08-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-08-19 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (94%) | - |
| Mechanistic Exploration of Backdoored Large Language Model Attention Patterns Lakshmi Babu-Saheer, Mohammed Abu Baker Published: 2025-08-19Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-19 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | - |
| Evaluating Sparse Autoencoders for Monosemantic Representation A.B. Siddique, Moghis Fereidouni, Muhammad Umair Haider, Peizhong Ju Published: 2025-08-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-20 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| AdaptiveK Sparse Autoencoders: Dynamic Sparsity Allocation for Interpretable LLM Representations Mengnan Du, Yifei Yao Published: 2025-08-24Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-24 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (94%) | 1 |
| Mechanistic Interpretability for Steering Vision-Language-Action Models Bear HĂ€on, Claire Tomlin, Ian Chuang, Kaylene Stocking Published: 2025-08-30Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-30 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 3 |
| CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders Alex Gulko, Sachin Kumar, Yusen Peng Published: 2025-08-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2025-08-31 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | - |
| Can LLMs Lie? Investigation beyond Hallucination Deepak Pathak, Haoran Huan, Mengning Wu, Mihir Prabhudesai Published: 2025-09-03Area: Deception & FailureCitations: 1 Tags: ai-safety, deception-failure, empirical, interpretability | 2025-09-03 | Deception & Failure | ai-safety, deception-failure, empirical, interpretability | E5 / R3 (94%) | 1 |
| Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Arkaitz Zubiaga, Shaz Furniturewala Published: 2025-09-16Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2025-09-16 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E6 / R3 (96%) | - |
| Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models Katharina Simbeck, Mariam Mahran Published: 2025-09-22Area: Representation AnalysisCitations: 2 Tags: ai-safety, empirical, interpretability, representation-analysis | 2025-09-22 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 2 |
| Who is In Charge? Dissecting Role Conflicts in Instruction Following Siqi Zeng Published: 2025-09-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-09-23 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | - |
| Binary Autoencoder for Mechanistic Interpretability of Large Language Models Brian M. Kurkoski, Hakaze Cho, Haolin Yang, Naoya Inoue Published: 2025-09-25Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-25 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Analysis of Variational Sparse Autoencoders Yuxiao Li, Zachary Baker Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research Sean Trott Published: 2025-09-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-26 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| LLM Interpretability with Identifiable Temporal-Instantaneous Representation Jiaqi Sun, Kun Zhang, Xiangchen Song, Yujia Zheng Published: 2025-09-27Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| Binary Sparse Coding for Interpretability Lucia Quirke, Nora Belrose, Stepan Shabalin Published: 2025-09-29Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (92%) | 1 |
| Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF Jing Liu Published: 2025-09-29Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, interpretability, theoretical | 2025-09-29 | Alignment Training | ai-safety, alignment-training, interpretability, theoretical | E5 / R3 (94%) | - |
| AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features Mohammad Mahdi Khalili, Xudong Zhu, Zhihui Zhu Published: 2025-10-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Mechanistic Interpretability as Statistical Estimation: A Variance Analysis of EAP-IG François Portet, Maxime Méloux, Maxime Peyrard Published: 2025-10-01Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-01 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | 2 |
| Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders Charibeth Cheng, Kriz Tahimic Published: 2025-10-03Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-03 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (95%) | - |
| Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders Benyou Wang, Difan Zou, Xu Wang, Yan Hu Published: 2025-10-04Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E7 / R4 (97%) | 2 |
| WeightLens and CircuitLens: Circuit Insights Towards Interpretability Beyond Activations Aakriti Jain, Ammar Ibrahim, Bruno Puri, Elena Golimblevskaia Published: 2025-10-16Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-10-16 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (94%) | - |
| Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention Gustavo Penha, Hugues Bouchard, JosĂ© Luis Redondo GarcĂa, J Rosser Published: 2025-10-22Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-22 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |
| Subliminal Corruption: Mechanisms, Thresholds, and Interpretability Reya Vir, Sarvesh Bhatnagar Published: 2025-10-22Area: Deception & FailureCitations: 1 Tags: ai-safety, alignment-training, deception-failure, empirical, interpretability | 2025-10-22 | Deception & Failure | ai-safety, alignment-training, deception-failure, empirical, interpretability | E5 / R3 (95%) | 1 |
| Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability Alex Oesterling, Claudio Mayrink Verdun, Flavio P. Calmon, Himabindu Lakkaraju Published: 2025-10-30Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-30 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (96%) | - |
| Atlas-Alignment: Making Interpretability Transferable Across Language Models Bruno Puri, Jim Berend, Sebastian Lapuschkin, Wojciech Samek Published: 2025-10-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E6 / R3 (95%) | - |
| Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI Anthony Baez, Pat Pataranutaporn, Sheer Karny Published: 2025-10-31Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-10-31 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 1 |