Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 451-465 of 465 papers (page 16 of 16)路 31 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Sparse Crosscoders for Cross-Layer Features and Model Diffing Adly Templeton, Christopher Olah, Jack Lindsey, Jonathan Marcus Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| Sparse Mixtures of Linear Transforms (MOLT) Adam Pearce, Brian Chen, Jack Lindsey, Sasha Hydrie Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| Stage-Wise Model Diffing Adam Jermyn, Christopher Olah, Jonathan Marcus, Kelley Rivoire Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Superposition, Memorization, and Double Descent Christopher Olah, Nelson Elhage, Nicholas Schiefer, Robert Lasenby Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | - |
| The Circuits Research Landscape: Results and Perspectives Callum McDougall, Connor Watts, Curt Tigges, Elana Simon Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (93%) | - |
| Thread: Circuits Ben Egan, Chelsea Voss, Chris Olah, Gabriel Goh Published: -Area: Mechanistic Interp.Citations: 142 Tags: ai-safety, mechanistic-interp, survey | - | Mechanistic Interp. | ai-safety, mechanistic-interp, survey | E5 / R3 (95%) | 142 |
| Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Adam Jermyn, Adly Templeton, Alex Tamkin, Amanda Askell Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |
| Toy Models of Superposition Carol Chen, Catherine Olsson, Christopher Olah, Dario Amodei Published: -Area: Mechanistic Interp.Citations: 621 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 621 |
| Tracing Attention Computation Through Feature Interactions Adam Pearce, Chris Olah, Emmanuel Ameisen, Harish Kamath Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (93%) | - |
| Transformer Feed-Forward Layers Are Key-Value Memories Jonathan Berant, Mor Geva, Omer Levy, Roei Schuster Published: -Area: Mechanistic Interp.Citations: 1203 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E4 / R3 (94%) | 1203 |
| TransformerLens Joseph Bloom, Neel Nanda Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, tool | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, tool | E5 / R3 (96%) | - |
| Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron Anirudh Goyal, Kenji Kawaguchi, Michael Qizhe Shieh, Wenxuan Zhang Published: -Area: Mechanistic Interp.Citations: 38 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E6 / R3 (96%) | 38 |
| Understanding RL Vision Chris Olah, Gabriel Goh, Jacob Hilton, Nick Cammarata Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, interpretability, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, interpretability, mechanistic-interp | E5 / R3 (94%) | - |
| Where Confabulation Lives: Latent Feature Discovery in LLMs Gerhard Wunder, Thibaud Ardoin, Yi Cai Published: -Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (94%) | 1 |
| Zoom In: An Introduction to Circuits Chris Olah, Gabriel Goh, Ludwig Schubert, Michael Petrov Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, empirical, mechanistic-interp | - | Mechanistic Interp. | ai-safety, empirical, mechanistic-interp | E5 / R3 (95%) | - |