Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 181-199 of 199 papers (page 7 of 7)路 38 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Interpreting Neural Networks through the Polytope Lens Beren Millidge, Carlos Ram贸n Guevara, Connor Leahy, Dan Braun Published: 2022-11-22Area: Mechanistic Interp.Citations: 36 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2022-11-22 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 36 |
| Polysemanticity and Capacity in Neural Networks Adam Scherlis, Adam S. Jermyn, Buck Shlegeris, Joe Benton Published: 2022-10-04Area: Mechanistic Interp.Citations: 52 Tags: ai-safety, mechanistic-interp, theoretical | 2022-10-04 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (92%) | 52 |
| On the Impossible Safety of Large AI Models El-Mahdi El-Mhamdi, John Stephan, L锚-Nguy锚n Hoang, Nirupam Gupta Published: 2022-09-30Area: Formal/TheoreticalCitations: 37 Tags: ai-safety, formaltheoretical, theoretical | 2022-09-30 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R4 (94%) | 37 |
| Defining and Characterizing Reward Hacking David Krueger, Dmitrii Krasheninnikov, Joar Skalse, Nikolaus H. R. Howe Published: 2022-09-27Area: Deception & FailureCitations: 96 Tags: ai-safety, deception-failure, theoretical | 2022-09-27 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (97%) | 96 |
| The Alignment Problem from a Deep Learning Perspective Lawrence Chan, Richard Ngo, S枚ren Mindermann Published: 2022-08-30Area: Deception & FailureCitations: 263 Tags: ai-safety, alignment-training, deception-failure, theoretical | 2022-08-30 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E5 / R3 (93%) | 263 |
| Parametrically Retargetable Decision-Makers Tend to Seek Power Alexander Matt Turner, Prasad Tadepalli Published: 2022-06-27Area: Formal/TheoreticalCitations: 21 Tags: ai-safety, formaltheoretical, theoretical | 2022-06-27 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E6 / R3 (96%) | 21 |
| Is Power-Seeking AI an Existential Risk? Joseph Carlsmith Published: 2022-06-16Area: Deception & FailureCitations: 129 Tags: ai-safety, alignment-training, deception-failure, theoretical | 2022-06-16 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E4 / R3 (94%) | 129 |
| Towards Understanding Grokking: An Effective Theory of Representation Learning Eric J. Michaud, Max Tegmark, Mike Williams, Niklas Nolte Published: 2022-05-20Area: Training DynamicsCitations: 217 Tags: ai-safety, theoretical, training-dynamics | 2022-05-20 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R4 (94%) | 217 |
| Consequences of Misaligned AI Dylan Hadfield-Menell, Simon Zhuang Published: 2021-02-07Area: Formal/TheoreticalCitations: 95 Tags: ai-safety, formaltheoretical, theoretical | 2021-02-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E6 / R3 (96%) | 95 |
| Agent Incentives: A Causal Perspective Eric D. Langlois, Pedro A. Ortega, Ryan Carey, Shane Legg Published: 2021-02-02Area: Formal/TheoreticalCitations: 61 Tags: ai-safety, formaltheoretical, theoretical | 2021-02-02 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E6 / R3 (96%) | 61 |
| Deep Learning is Singular, and That's Good Daniel Murfet, Hui Li, Jesse Gell-Redman, Mingming Gong Published: 2020-10-22Area: Formal/TheoreticalCitations: 38 Tags: ai-safety, formaltheoretical, theoretical | 2020-10-22 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (95%) | 38 |
| Optimal Policies Tend to Seek Power Alexander Matt Turner, Andrew Critch, Logan Smith, Prasad Tadepalli Published: 2019-12-03Area: Formal/TheoreticalCitations: 97 Tags: ai-safety, formaltheoretical, theoretical | 2019-12-03 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (93%) | 97 |
| A Mathematical Framework for Transformer Circuits Amanda Askell, Andy Jones, Anna Chen, Ben Mann Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | - |
| A Toy Model of Mechanistic (Un)Faithfulness Chris Olah Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | - |
| A Two-Step, Multidimensional Account of Deception in Language Models Leonard Dung Published: -Area: Deception & FailureCitations: - Tags: ai-safety, deception-failure, theoretical | - | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (95%) | - |
| Aversion to external feedback suffices to ensure agent alignment Paulo Garcia Published: -Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, theoretical | - | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E5 / R3 (93%) | 1 |
| Eliciting Latent Knowledge (ELK) Ajeya Cotra, Mark Xu, Paul Christiano Published: -Area: Scalable OversightCitations: - Tags: ai-safety, scalable-oversight, theoretical | - | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E6 / R2 (96%) | - |
| Explaining AI through mechanistic interpretability Barnaby Crook, Lena K盲stner Published: -Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | - | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | - |
| Toward Constitutional Autonomy in AI Systems: A Theoretical Framework for Aligned Agentic Intelligence William Torgbi Agbemabiese Published: -Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, theoretical | - | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E4 / R3 (93%) | 1 |