Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Case Study on Concept Induction for Neuron-Level Interpretability in CNN Moumita Sen Sarma, Pascal Hitzler, Samatha Ereshi Akkamahadevi Published: 2026-02-27Area: cs.CVCitations: 10 Tags: ai-safety, cscv, interpretability, preprint | 2026-02-27 | cs.CV | ai-safety, cscv, interpretability, preprint | E9 / R8 (93%) | 10 |
| Unpacking Interpretability: Human-Centered Criteria for Optimal Combinatorial Solutions David Steyrl, Dominik Pegler, Filip Melinscak, Frank J盲kel Published: 2026-03-09Area: cs.HCCitations: - Tags: ai-safety, cshc, interpretability, preprint | 2026-03-09 | cs.HC | ai-safety, cshc, interpretability, preprint | E5 / R3 (93%) | - |
| Age Predictors Through the Lens of Generalization, Bias Mitigation, and Interpretability: Reflections on Causal Implications Alessandro Cellerino, Debdas Paul, Elisa Ferrari, Irene Gravili Published: 2026-03-17Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-03-17 | cs.LG | ai-safety, cslg, interpretability, preprint | E5 / R3 (92%) | - |
| Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations Aakriti Kinra, Bhairavi Muralidharan, John Morgan, Namrata Elamaran Published: 2026-03-18Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-03-18 | cs.AI | ai-safety, csai, interpretability, preprint | E6 / R4 (98%) | - |
| Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models Benlin Liu, Jacob Feldman, Liwei Che, Michelle Hurst Published: 2026-03-19Area: cs.CVCitations: - Tags: ai-safety, cscv, interpretability, preprint | 2026-03-19 | cs.CV | ai-safety, cscv, interpretability, preprint | E5 / R3 (95%) | - |
| Pitfalls in Evaluating Interpretability Agents Aaron Mueller, Antonio Torralba, Jacob Andreas, Nikhil Prakash Published: 2026-03-20Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-03-20 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (96%) | - |
| Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era Abu Noman Md Sakib, Protik Dey, Taslima Akter, Zijie Zhang Published: 2026-03-31Area: cs.HCCitations: - Tags: ai-safety, cshc, interpretability, preprint | 2026-03-31 | cs.HC | ai-safety, cshc, interpretability, preprint | E5 / R3 (94%) | - |
| From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability Guillaume Corlouer, Max Hennick Published: 2026-03-31Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-03-31 | cs.LG | ai-safety, cslg, interpretability, preprint | E5 / R3 (96%) | - |
| From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability Guillaume Corlouer, Max Hennick Published: 2026-03-31Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-03-31 | cs.LG | ai-safety, cslg, interpretability, preprint | E5 / R3 (95%) | - |
| Detecting Multi-Agent Collusion Through Multi-Agent Interpretability Aaron Rose, Brandon Gary Kaplowitz, Carissa Cullen, Christian Schroeder de Witt Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-01 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (94%) | - |
| Distributed Interpretability and Control for Large Language Models Dev Arpan Desai, Shaoyi Huang, Zining Zhu Published: 2026-04-07Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-04-07 | cs.LG | ai-safety, cslg, interpretability, preprint | E5 / R3 (93%) | - |
| Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? Aashiq Muhamed, Aditi Raghunathan, Mona T. Diab, Virginia Smith Published: 2026-04-13Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-04-13 | cs.LG | ai-safety, cslg, interpretability, preprint | E5 / R3 (95%) | - |
| Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers Bernt Schiele, Jonas Fischer, Nina 呕ukowska, Wolfgang Stammer Published: 2026-04-15Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-15 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (95%) | - |
| How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study Chengzhi Li, Ping Jian, Wenpeng Lu, Xinyue Zhang Published: 2026-04-16Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-16 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (94%) | - |
| Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures Liangming Pan, Qinglin Meng, Yuan Zhou, Yutong Gao Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-17 | cs.CL | ai-safety, cscl, interpretability, preprint | E12 / R5 (99%) | - |
| Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures Liangming Pan, Qinglin Meng, Yuan Zhou, Yutong Gao Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-17 | cs.CL | ai-safety, cscl, interpretability, preprint | E19 / R10 (100%) | - |
| Using Large Language Models and Knowledge Graphs to Improve the Interpretability of Machine Learning Models in Manufacturing Alexander Lohr, Bernd Michelberger, Sarah Wei脽, Thomas Bayer Published: 2026-04-17Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-17 | cs.AI | ai-safety, csai, interpretability, preprint | E8 / R7 (96%) | - |
| Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks Dongmei Zhang, Jue Zhang, Qingwei Lin, Rongyuan Tan Published: 2026-04-20Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-20 | cs.AI | ai-safety, csai, interpretability, preprint | E10 / R4 (95%) | - |
| AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models Michael Keeman Published: 2026-04-26Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-26 | cs.CL | ai-safety, cscl, interpretability, preprint | E7 / R5 (97%) | - |
| Beyond the Black Box: Interpretability of Agentic AI Tool Use Ariye Shater, Hariom Tatsat Published: 2026-05-07Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-05-07 | cs.AI | ai-safety, csai, interpretability, preprint | E10 / R8 (90%) | - |
| ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows Alvaro Lopez Pellicer, Eduardo Soares, Jemma Kerns, Marwan Bukhari Published: 2026-05-13Area: cs.CVCitations: - Tags: ai-safety, cscv, interpretability, preprint | 2026-05-13 | cs.CV | ai-safety, cscv, interpretability, preprint | E11 / R9 (94%) | - |
| Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning Ali Dasdan, Chad Coleman, Kund Meghani, Manan Shah Published: 2026-06-13Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-06-13 | cs.AI | ai-safety, csai, interpretability, preprint | E7 / R4 (93%) | - |
| Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images Daesoo Kim, Daeun Yoo, Eunsu Lee, Ian Choi Published: 2026-06-30Area: cs.LGCitations: - Tags: ai-safety, alignment-training, cslg, interpretability, preprint | 2026-06-30 | cs.LG | ai-safety, alignment-training, cslg, interpretability, preprint | E8 / R4 (97%) | - |
| Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders Christopher Earls, Zihao Qi Published: 2026-07-01Area: quant-phCitations: - Tags: ai-safety, interpretability, preprint, quant-ph | 2026-07-01 | quant-ph | ai-safety, interpretability, preprint, quant-ph | E6 / R4 (96%) | - |
| Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability Rodrigo Mendoza-Smith Published: 2026-07-02Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-02 | cs.LG | ai-safety, cslg, interpretability, preprint | E9 / R5 (97%) | - |
| Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models Franz Motzkus, Sebastian Bernhard Published: 2026-07-07Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-07-07 | cs.AI | ai-safety, csai, interpretability, preprint | E8 / R6 (89%) | - |
| Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs Anupam Wagle, Chaowei Zhang, Ifrat Ikhtear Uddin, Longwei Wang Published: 2026-07-08Area: cs.CRCitations: - Tags: adversarial-robustness, ai-safety, cscr, interpretability, preprint | 2026-07-08 | cs.CR | adversarial-robustness, ai-safety, cscr, interpretability, preprint | E6 / R6 (91%) | - |
| Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Huaxing Liu, Shuai Li, Sixian Li, Xiang Wang Published: 2026-07-13Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-13 | cs.LG | ai-safety, cslg, interpretability, preprint | E16 / R15 (94%) | - |
| AIMO Interpretability Challenge Adam Vawda-Oomerjee, Andreas Waldis, Barbara Plank, Chaoran Liu Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-07-15 | cs.AI | ai-safety, csai, interpretability, preprint | E12 / R8 (89%) | - |
| Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control Alice Chan, Glen Chou, Jihoon Hong, Julian Skifstad Published: 2026-07-16Area: cs.ROCitations: 67 Tags: ai-safety, csro, interpretability, preprint | 2026-07-16 | cs.RO | ai-safety, csro, interpretability, preprint | E8 / R5 (93%) | 67 |