Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 236 papers (page 1 of 8)路 5775 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning Hang Chen, Jiaying Zhu, Wenya Wang Published: 2026-08-25Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-08-25 | cs.LG | ai-safety, cslg, interpretability, preprint | - | - |
| Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B David Bau, Jiahao Liu, Octavia Camps, Pu Zhao Published: 2026-08-19Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-08-19 | cs.LG | ai-safety, cslg, interpretability, preprint | E8 / R6 (93%) | - |
| Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability Vijay Erramilli Published: 2026-08-19Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-08-19 | cs.LG | ai-safety, cslg, interpretability, preprint | E9 / R7 (94%) | - |
| Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation Ajay Pravin Mahale Published: 2026-08-13Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-08-13 | cs.AI | ai-safety, csai, interpretability, preprint | E10 / R7 (93%) | - |
| HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry Binbin Yong, Haoran Li, Haoran Pei, Jun Shen Published: 2026-08-12Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-08-12 | cs.AI | ai-safety, csai, interpretability, preprint | - | - |
| From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das Published: 2026-08-11Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-08-11 | cs.CL | ai-safety, cscl, interpretability, preprint | E12 / R9 (91%) | - |
| MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering Jakub Po膰wiardowski, Mateusz Modrzejewski Published: 2026-08-06Area: cs.SDCitations: - Tags: ai-safety, cssd, interpretability, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, interpretability, preprint | E11 / R13 (91%) | - |
| Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers Jacek Tabor, Kamil Ksi膮偶ek, Micha艂 Jan W艂odarczyk, Piotr Suszy艅ski Published: 2026-07-31Area: cs.CVCitations: 41 Tags: ai-safety, cscv, interpretability, preprint | 2026-07-31 | cs.CV | ai-safety, cscv, interpretability, preprint | E10 / R8 (92%) | 41 |
| ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders Wei Qiu, Yixuan Duan Published: 2026-07-29Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-29 | cs.LG | ai-safety, cslg, interpretability, preprint | E13 / R12 (91%) | - |
| KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability Aditi Anand, Ananya Lakshmi Ravi, Balaraman Ravindran, Gokul S. Krishnan Published: 2026-07-27Area: cs.CVCitations: - Tags: ai-safety, cscv, interpretability, preprint | 2026-07-27 | cs.CV | ai-safety, cscv, interpretability, preprint | - | - |
| Interior interpretability with attention rollout: contraction and propagation profiles in Transformers Enrique Zuazua, Qian Huang, Umberto Biccari Published: 2026-07-24Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-24 | cs.LG | ai-safety, cslg, interpretability, preprint | - | - |
| Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent Soheil Feizi, Sriram Balasubramanian Published: 2026-07-17Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-17 | cs.LG | ai-safety, cslg, interpretability, preprint | E11 / R9 (90%) | - |
| Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control Alice Chan, Glen Chou, Jihoon Hong, Julian Skifstad Published: 2026-07-16Area: cs.ROCitations: 67 Tags: ai-safety, csro, interpretability, preprint | 2026-07-16 | cs.RO | ai-safety, csro, interpretability, preprint | E8 / R5 (93%) | 67 |
| AIMO Interpretability Challenge Adam Vawda-Oomerjee, Andreas Waldis, Barbara Plank, Chaoran Liu Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-07-15 | cs.AI | ai-safety, csai, interpretability, preprint | E12 / R8 (89%) | - |
| Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Huaxing Liu, Shuai Li, Sixian Li, Xiang Wang Published: 2026-07-13Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-13 | cs.LG | ai-safety, cslg, interpretability, preprint | E16 / R15 (94%) | - |
| Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs Anupam Wagle, Chaowei Zhang, Ifrat Ikhtear Uddin, Longwei Wang Published: 2026-07-08Area: cs.CRCitations: - Tags: adversarial-robustness, ai-safety, cscr, interpretability, preprint | 2026-07-08 | cs.CR | adversarial-robustness, ai-safety, cscr, interpretability, preprint | E6 / R6 (91%) | - |
| Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models Franz Motzkus, Sebastian Bernhard Published: 2026-07-07Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-07-07 | cs.AI | ai-safety, csai, interpretability, preprint | E8 / R6 (89%) | - |
| Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability Rodrigo Mendoza-Smith Published: 2026-07-02Area: cs.LGCitations: - Tags: ai-safety, cslg, interpretability, preprint | 2026-07-02 | cs.LG | ai-safety, cslg, interpretability, preprint | E9 / R5 (97%) | - |
| Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders Christopher Earls, Zihao Qi Published: 2026-07-01Area: quant-phCitations: - Tags: ai-safety, interpretability, preprint, quant-ph | 2026-07-01 | quant-ph | ai-safety, interpretability, preprint, quant-ph | E6 / R4 (96%) | - |
| Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images Daesoo Kim, Daeun Yoo, Eunsu Lee, Ian Choi Published: 2026-06-30Area: cs.LGCitations: - Tags: ai-safety, alignment-training, cslg, interpretability, preprint | 2026-06-30 | cs.LG | ai-safety, alignment-training, cslg, interpretability, preprint | E8 / R4 (97%) | - |
| Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning Ali Dasdan, Chad Coleman, Kund Meghani, Manan Shah Published: 2026-06-13Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-06-13 | cs.AI | ai-safety, csai, interpretability, preprint | E7 / R4 (93%) | - |
| ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows Alvaro Lopez Pellicer, Eduardo Soares, Jemma Kerns, Marwan Bukhari Published: 2026-05-13Area: cs.CVCitations: - Tags: ai-safety, cscv, interpretability, preprint | 2026-05-13 | cs.CV | ai-safety, cscv, interpretability, preprint | E11 / R9 (94%) | - |
| Beyond the Black Box: Interpretability of Agentic AI Tool Use Ariye Shater, Hariom Tatsat Published: 2026-05-07Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-05-07 | cs.AI | ai-safety, csai, interpretability, preprint | E10 / R8 (90%) | - |
| AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models Michael Keeman Published: 2026-04-26Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-26 | cs.CL | ai-safety, cscl, interpretability, preprint | E7 / R5 (97%) | - |
| Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks Dongmei Zhang, Jue Zhang, Qingwei Lin, Rongyuan Tan Published: 2026-04-20Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-20 | cs.AI | ai-safety, csai, interpretability, preprint | E10 / R4 (95%) | - |
| Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures Liangming Pan, Qinglin Meng, Yuan Zhou, Yutong Gao Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-17 | cs.CL | ai-safety, cscl, interpretability, preprint | E12 / R5 (99%) | - |
| Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures Liangming Pan, Qinglin Meng, Yuan Zhou, Yutong Gao Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, interpretability, preprint | 2026-04-17 | cs.CL | ai-safety, cscl, interpretability, preprint | E19 / R10 (100%) | - |
| Using Large Language Models and Knowledge Graphs to Improve the Interpretability of Machine Learning Models in Manufacturing Alexander Lohr, Bernd Michelberger, Sarah Wei脽, Thomas Bayer Published: 2026-04-17Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-17 | cs.AI | ai-safety, csai, interpretability, preprint | E8 / R7 (96%) | - |
| How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study Chengzhi Li, Ping Jian, Wenpeng Lu, Xinyue Zhang Published: 2026-04-16Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-16 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (94%) | - |
| Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers Bernt Schiele, Jonas Fischer, Nina 呕ukowska, Wolfgang Stammer Published: 2026-04-15Area: cs.AICitations: - Tags: ai-safety, csai, interpretability, preprint | 2026-04-15 | cs.AI | ai-safety, csai, interpretability, preprint | E5 / R3 (95%) | - |