Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 465 papers (page 1 of 16)路 47 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders Bin Dong, Jiedong Jiang, Mingrui Wu, Tianyang Liu Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| MonoLoss: A Training Objective for Interpretable Monosemantic Representations Ali Nasiri-Sarvi, Anh Tien Nguyen, Dimitris Samaras, Hassan Rivaz Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| ProtoMech: Protein Circuit Tracing via Cross-layer Transcoders Amirali Aghazadeh, Daniel Saeedi, Darin Tsui, Kunal Talreja Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Prototype Transformer: Towards Language Model Architectures Interpretable by Design Amine M'Charrak, Bayar Menzat, Chang Qi, Markus Kaltenberger Published: 2026-02-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-12 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features Adriano Koshiyama, Seonglae Cho, Zekun Wu Published: 2026-02-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (97%) | - |
| How Many Features Can a Language Model Store Under the Linear Representation Hypothesis? Jon Kleinberg, Kenny Peng, Nikhil Garg Published: 2026-02-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2026-02-11 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E4 / R2 (93%) | - |
| Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models Djam茅 Seddah, Francis Kulumba, Th茅o Lasnier, Wissam Antoun Published: 2026-02-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (95%) | - |
| Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path Andres Suarez, Dongsoo Har, Neha Sengar Published: 2026-02-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-10 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints Andres Saurez, Dongsoo Har, Yousung Lee Published: 2026-02-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-02-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery Bangwei Guo, Difei Gu, Dimitris Metaxas, Gerasimos Chatzoudis Published: 2026-02-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, mechanistic-interp | 2026-02-07 | Mechanistic Interp. | ai-safety, alignment-training, empirical, mechanistic-interp | E5 / R4 (94%) | - |
| Learning a Generative Meta-Model of LLM Activations Alec Radford, Grace Luo, Jacob Steinhardt, Jiahai Feng Published: 2026-02-06Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Your Language Model Secretly Contains Personality Subnetworks Bo Hui, Manling Li, Ruimeng Ye, Xiaolong Ma Published: 2026-02-06Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-06 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders Baosong Yang, Bingqing Jiang, Difan Zou, Lingpeng Kong Published: 2026-02-05Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2026-02-05 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | - |
| Decomposing Query-Key Feature Interactions Using Contrastive Covariances Andrew Lee, Fernanda Vi茅gas, Martin Wattenberg, Yonatan Belinkov Published: 2026-02-04Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Identifying Intervenable and Interpretable Features via Orthogonality Regularization Bernhard Sch枚lkopf, Florent Draye, Moritz Miller Published: 2026-02-04Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2026-02-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning Donald Ye, Linus Wong, Max Loffgren, Om Kotadia Published: 2026-02-04Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-04 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R4 (97%) | - |
| A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior Adam Mahdi, Dewi Gould, Harry Mayne, Justin Singh Kang Published: 2026-02-02Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-02 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (96%) | - |
| Spectral Superposition: A Theory of Feature Geometry Amir Abdullah, Georgi Ivanov, Narmeen Oozeer, Shivam Raval Published: 2026-02-02Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2026-02-02 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding Andreas D. Demou, James Oldfield, Mihalis Nicolaou, Panagiotis Koromilas Published: 2026-02-01Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-02-01 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Beyond Activation Patterns: A Weight-Based Out-of-Context Explanation of Sparse Autoencoder Features Yiting Liu, Zhi-Hong Deng Published: 2026-01-30Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (95%) | - |
| Language Model Circuits Are Sparse in the Neuron Basis Aryaman Arora, Jacob Steinhardt, Sarah Schwettmann, Zhengxuan Wu Published: 2026-01-30Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2026-01-30 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | 1 |
| Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units Jianhui Chen, Liangming Pan, Yuzhang Luo Published: 2026-01-29Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2026-01-29 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs Anton van den Hengel, Dong Gong, Erdun Gao, Javen Qinfeng Shi Published: 2026-01-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Sycophancy Hides Linearly in the Attention Heads Hilal Alquabeh, Kentaro Inui, Munachiso Nwadike, Nurdaulet Mukhituly Published: 2026-01-23Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2026-01-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (95%) | 1 |
| Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction Tony Cristofano Published: 2026-01-22Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Patterning: The Dual of Interpretability Daniel Murfet, George Wang Published: 2026-01-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-01-20 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R4 (95%) | - |
| Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition Mohammed Kaif Pasha, Mohammed Mudassir Uddin, Shahnawaz Alam Published: 2026-01-19Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-19 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders Archie Chaudhury, Shikhar Shiromani, Sri Pranav Kunda Published: 2026-01-14Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2026-01-14 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | - |
| SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis Xufeng Duan, Zhenguang G. Cai, Zihao Fu Published: 2026-01-12Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2026-01-12 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (97%) | - |
| When Models Manipulate Manifolds: The Geometry of a Counting Task Adam Pearce, Chris Olah, Emmanuel Ameisen, Isaac Kauvar Published: 2026-01-08Area: Mechanistic Interp.Citations: 12 Tags: ai-safety, empirical, mechanistic-interp | 2026-01-08 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 12 |