Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video Blake Aaron Richards, Danilo Bzdok, Edward Stevinson, Lee Sharkey Published: 2025-04-28Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, tool | 2025-04-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (98%) | 11 |
| MIB: A Mechanistic Interpretability Benchmark Aaron Mueller, Adam Belfki, Alessandro Stolfo, Amir Zur Published: 2025-04-17Area: Mechanistic Interp.Citations: 16 Tags: ai-safety, benchmark, interpretability, mechanistic-interp | 2025-04-17 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp | E6 / R3 (97%) | 16 |
| Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems Mateusz Dziemian, Natalia P茅rez-Campanero Antol铆n, Simon Lermen Published: 2025-04-10Area: Deception & FailureCitations: 4 Tags: ai-safety, deception-failure, empirical, interpretability | 2025-04-10 | Deception & Failure | ai-safety, deception-failure, empirical, interpretability | E6 / R3 (94%) | 4 |
| Towards Combinatorial Interpretability of Neural Computation Dan Alistarh, Micah Adler, Nir Shavit Published: 2025-04-10Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-04-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | 7 |
| AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability Alexander Boyd, Fernando Rosas, Manuel Baltieri Published: 2025-04-06Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-04-06 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (93%) | 2 |
| Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability Mohammad Mahdi Khalili, Vishnu Kabir Chhabra Published: 2025-04-05Area: Model EditingCitations: - Tags: ai-safety, empirical, interpretability, model-editing | 2025-04-05 | Model Editing | ai-safety, empirical, interpretability, model-editing | E7 / R3 (98%) | - |
| Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models Mateusz Pach, Quentin Bouniot, Serge Belongie, Shyamgopal Karthik Published: 2025-04-03Area: Mechanistic Interp.Citations: 21 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-04-03 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (97%) | 21 |
| TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research Abir Harrasse, Amir Abdullah, Clement Neo, Dhruv Nathawani Published: 2025-03-17Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, dataset, interpretability, mechanistic-interp | 2025-03-17 | Mechanistic Interp. | ai-safety, dataset, interpretability, mechanistic-interp | E6 / R3 (95%) | 2 |
| HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks Atticus Geiger, Christopher Potts, Jing Huang, Jiuding Sun Published: 2025-03-13Area: Mechanistic Interp.Citations: 8 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-03-13 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | 8 |
| SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability Adam Karvonen, Arthur Conmy, Callum McDougall, Can Rager Published: 2025-03-12Area: Mechanistic Interp.Citations: 62 Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2025-03-12 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (95%) | 62 |
| RouteSAE: Route Sparse Autoencoder to Interpret Large Language Models Guojun Ma, Mingyang Wan, Sihang Li, Tao Liang Published: 2025-03-11Area: Mechanistic Interp.Citations: 17 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-03-11 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R4 (95%) | 17 |
| Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models Boussad Addad, Katarzyna Kapusta, Thomas Winninger Published: 2025-03-08Area: Adversarial RobustnessCitations: 3 Tags: adversarial-robustness, ai-safety, empirical, interpretability | 2025-03-08 | Adversarial Robustness | adversarial-robustness, ai-safety, empirical, interpretability | E6 / R3 (96%) | 3 |
| Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? Francois Portet, Maxime Meloux, Maxime Peyrard, Silviu Maniu Published: 2025-02-28Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-02-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (94%) | 14 |
| A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models Arman Zarei, Barry Menglong Yao, Hongxuan Li, Keivan Rezaei Published: 2025-02-22Area: Surveys & ReviewsCitations: 20 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-02-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (96%) | 20 |
| Modular Training of Neural Networks aids Interpretability Alessandro Abate, Joan Velja, Maheep Chaudhary, Nandi Schoots Published: 2025-02-04Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-02-04 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 1 |
| Building Bridges, Not Walls: Advancing Interpretability by Unifying Feature, Data, and Model Component Attribution Hima Lakkaraju, Shichang Zhang, Tessa Han, Usha Bhalla Published: 2025-01-31Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, position, surveys-reviews | 2025-01-31 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (97%) | 3 |
| Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers Laurence Aitchison, Lucy Farnik, Thomas Heap, Tim Lawson Published: 2025-01-29Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-29 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 26 |
| Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning Anna Soligo, David Boyle, Pietro Ferraro Published: 2025-01-28Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-28 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 2 |
| Open Problems in Mechanistic Interpretability Adria Garriga-Alonso, Alejandro Ortega, Arthur Conmy, Atticus Geiger Published: 2025-01-27Area: Surveys & ReviewsCitations: 107 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-01-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 107 |
| Propositional Interpretability in Artificial Intelligence David J. Chalmers Published: 2025-01-27Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-01-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (94%) | 13 |
| Enhancing Automated Interpretability with Output-Centric Feature Descriptions Atticus Geiger, Chen Agassy, Mor Geva, Roy Mayan Published: 2025-01-14Area: Mechanistic Interp.Citations: 26 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-01-14 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R3 (92%) | 26 |
| Large Language Model Safety: A Holistic Survey Bojian Jiang, Chuang Liu, Dan Shi, Deyi Xiong Published: 2024-12-23Area: Surveys & ReviewsCitations: 47 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-12-23 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R5 (99%) | 47 |
| Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation Erica K. Shimomoto, Kazuhiro Fukui, Lincon S. Souza, Pedro H. V. Valois Published: 2024-12-10Area: Representation AnalysisCitations: 2 Tags: ai-safety, empirical, interpretability, representation-analysis | 2024-12-10 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 2 |
| Towards Unifying Interpretability and Control: Evaluation via Intervention Asma Ghandeharioun, Himabindu Lakkaraju, Suraj Srinivas, Usha Bhalla Published: 2024-11-07Area: Mechanistic Interp.Citations: 20 Tags: ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | 2024-11-07 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp, safety-evaluation | E6 / R3 (94%) | 20 |
| Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders Alasdair Paren, David Krueger, Fazl Barez, Luke Marks Published: 2024-11-02Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2024-11-02 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness Jingyi Cui, Qi Lei, Qi Zhang, Stefanie Jegelka Published: 2024-10-27Area: Mechanistic Interp.Citations: 5 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-27 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (94%) | 5 |
| Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization Aaquib Syed, Abhay Sheshadri, Aidan Ewart, Gintare Karolina Dziugaite Published: 2024-10-16Area: Model EditingCitations: 19 Tags: ai-safety, empirical, interpretability, model-editing | 2024-10-16 | Model Editing | ai-safety, empirical, interpretability, model-editing | E6 / R3 (96%) | 19 |
| ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability Han Li, Jun Xu, Kai Zheng, Weijie Yu Published: 2024-10-15Area: Mechanistic Interp.Citations: 68 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-15 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E6 / R4 (95%) | 68 |
| Bilinear MLPs Enable Weight-Based Mechanistic Interpretability Alice Rigg, Jose M. Oramas, Lee Sharkey, Michael T. Pearce Published: 2024-10-10Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2024-10-10 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (95%) | 19 |
| The Computational Complexity of Circuit Discovery for Inner Interpretability Federico Adolfi, Martina G. Vilas, Todd Wareham Published: 2024-10-10Area: Formal/TheoreticalCitations: 14 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2024-10-10 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (98%) | 14 |