Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs Atharva Nijasure, James Allan, Tanya Chowdhury, Yair Zick Published: 2025-09-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| Measuring Sparse Autoencoder Feature Sensitivity Claire Tian, Katherine Tian, Nathan Hu Published: 2025-09-28Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp, safety-evaluation | 2025-09-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | - |
| LLM Interpretability with Identifiable Temporal-Instantaneous Representation Jiaqi Sun, Kun Zhang, Xiangchen Song, Yujia Zheng Published: 2025-09-27Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| Analysis of Variational Sparse Autoencoders Yuxiao Li, Zachary Baker Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models Biwei Huang, Kun Wang, Miao Yu, Moayad Aloqaily Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (98%) | - |
| Can Large Language Models Develop Gambling Addiction? Donghyeon Shin, Seungpil Lee, Sundong Kim, Yunjeong Lee Published: 2025-09-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (93%) | 2 |
| Concept-SAE: Active Causal Probing of Visual Model Behavior Chenchen Zhao, Jianrong Ding, Muxi Chen, Qiang Xu Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: adversarial-robustness, ai-safety, empirical, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | adversarial-robustness, ai-safety, empirical, mechanistic-interp | E6 / R4 (94%) | - |
| OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features Alexey Dontsov, Andrey Galichin, Anton Korznikov, Elena Tutubalina Published: 2025-09-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R4 (95%) | 2 |
| Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research Sean Trott Published: 2025-09-26Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-09-26 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E4 / R3 (94%) | 2 |
| Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing Jun Sun, Wei Zhao, Yige Li, Zhe Li Published: 2025-09-26Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-26 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (98%) | - |
| Binary Autoencoder for Mechanistic Interpretability of Large Language Models Brian M. Kurkoski, Hakaze Cho, Haolin Yang, Naoya Inoue Published: 2025-09-25Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-09-25 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (93%) | - |
| Towards Atoms of Large Language Models Chenhui Hu, Jun Zhao, Kang Liu, Pengfei Cao Published: 2025-09-25Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2025-09-25 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (98%) | - |
| Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI Eduard Kapelko Published: 2025-09-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (96%) | - |
| Who is In Charge? Dissecting Role Conflicts in Instruction Following Siqi Zeng Published: 2025-09-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | 2025-09-23 | Mechanistic Interp. | ai-safety, alignment-training, empirical, interpretability, mechanistic-interp | E5 / R4 (92%) | - |
| ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models Chenxiao Li, Haoxuan Li, Minfeng Zhu, Qiqi Jiang Published: 2025-09-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, tool | 2025-09-20 | Mechanistic Interp. | ai-safety, mechanistic-interp, tool | E5 / R3 (97%) | - |
| Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework Hanyu Zhang, Han Zheng, Hui Xue, Jialing Tao Published: 2025-09-11Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-11 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (97%) | 1 |
| From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers Danilo Bzdok, Jack Stanley, Luca Scimeca, Praneet Suresh Published: 2025-09-08Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-08 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | 1 |
| Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal Amir Abdullah, Erik Cambria, Nirmalendu Prakash, Ranjan Satapathy Published: 2025-09-07Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2025-09-07 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 4 |
| Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function Jason Abohwo, Thomas Mosen Published: 2025-09-02Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-09-02 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (99%) | - |
| Understanding sparse autoencoder scaling in the presence of feature manifolds Eric J. Michaud, Liv Gorton, Tom McGrath Published: 2025-09-02Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2025-09-02 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E4 / R3 (93%) | 2 |
| CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders Alex Gulko, Sachin Kumar, Yusen Peng Published: 2025-08-31Area: Mechanistic Interp.Citations: - Tags: ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | 2025-08-31 | Mechanistic Interp. | ai-safety, benchmark, interpretability, mechanistic-interp, safety-evaluation | E5 / R3 (94%) | - |
| Superposition in Graph Neural Networks Han Xuanyuan, Lukas Pertl, Pietro Li貌 Published: 2025-08-31Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | 2025-08-31 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E7 / R3 (95%) | 1 |
| Mechanistic Interpretability for Steering Vision-Language-Action Models Bear H盲on, Claire Tomlin, Ian Chuang, Kaylene Stocking Published: 2025-08-30Area: Mechanistic Interp.Citations: 3 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-30 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (97%) | 3 |
| Distribution-Aware Feature Selection for SAEs Alice Rigg, Amirali Abdullah, Michael Lan, Narmeen Oozeer Published: 2025-08-29Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-08-29 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching Ashkan Khakzar, Farnoush Rezaei Jafari, Neel Nanda, Oliver Eberle Published: 2025-08-28Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2025-08-28 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (95%) | 4 |
| Activation Transport Operators Andrzej Szablewski, Marek Masiak Published: 2025-08-24Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-08-24 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (94%) | - |
| AdaptiveK Sparse Autoencoders: Dynamic Sparsity Allocation for Interpretable LLM Representations Mengnan Du, Yifei Yao Published: 2025-08-24Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-24 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E4 / R3 (94%) | 1 |
| Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning Junxuan Wang, Wentao Shu, Xipeng Qiu, Xuyang Ge Published: 2025-08-23Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | 2025-08-23 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R4 (94%) | - |
| Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders Adri脿 Garriga-Alonso, David Chanin Published: 2025-08-22Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, empirical, mechanistic-interp | 2025-08-22 | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (94%) | 4 |
| Evaluating Sparse Autoencoders for Monosemantic Representation A.B. Siddique, Moghis Fereidouni, Muhammad Umair Haider, Peizhong Ju Published: 2025-08-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | 2025-08-20 | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (96%) | - |