Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking Anca Dragan, Cassidy Laidlaw, Shivam Singhal Published: 2024-03-05Area: Deception & FailureCitations: 26 Tags: ai-safety, deception-failure, theoretical | 2024-03-05 | Deception & Failure | ai-safety, deception-failure, theoretical | E6 / R4 (94%) | 26 |
| When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning Anca Dragan, Davis Foote, Erik Jenner, Leon Lang Published: 2024-02-27Area: Deception & FailureCitations: 13 Tags: ai-safety, deception-failure, theoretical | 2024-02-27 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (94%) | 13 |
| Immunization against harmful fine-tuning attacks Domenic Rosati, Frank Rudzicz, Hassan Sajjad, Jan Batzner Published: 2024-02-26Area: Adversarial RobustnessCitations: 34 Tags: adversarial-robustness, ai-safety, theoretical | 2024-02-26 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E4 / R2 (97%) | 34 |
| Robust agents learn causal world models Jonathan Richens, Tom Everitt Published: 2024-02-16Area: Formal/TheoreticalCitations: 67 Tags: ai-safety, formaltheoretical, theoretical | 2024-02-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 67 |
| Representation Surgery: Theory and Practice of Affine Steering Jonathan Herzig, Ponnurangam Kumaraguru, Roee Aharoni, Ryan Cotterell Published: 2024-02-15Area: Representation AnalysisCitations: 32 Tags: ai-safety, representation-analysis, theoretical | 2024-02-15 | Representation Analysis | ai-safety, representation-analysis, theoretical | E4 / R3 (94%) | 32 |
| Reward Generalization in RLHF: A Topological Perspective Dong Yan, Fanzhi Zeng, Han Yang, Jiaming Ji Published: 2024-02-15Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2024-02-15 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 7 |
| The Reasons that Agents Act: Intention and Instrumental Goals Francesca Toni, Francesco Belardinelli, Francis Rhys Ward, Matt MacDermott Published: 2024-02-11Area: Agent SafetyCitations: 22 Tags: agent-safety, ai-safety, theoretical | 2024-02-11 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R3 (96%) | 22 |
| Tradeoffs Between Alignment and Helpfulness in Language Models Amnon Shashua, Binyamin Rothberg, Dorin Shteyman, Noam Wies Published: 2024-01-29Area: Representation AnalysisCitations: 21 Tags: ai-safety, alignment-training, representation-analysis, theoretical | 2024-01-29 | Representation Analysis | ai-safety, alignment-training, representation-analysis, theoretical | E5 / R3 (94%) | 21 |
| Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback Dahua Lin, Hang Yan, Junjie Ye, Qiming Ge Published: 2024-01-21Area: Alignment TrainingCitations: 22 Tags: ai-safety, alignment-training, theoretical | 2024-01-21 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R4 (96%) | 22 |
| Deception and Manipulation in Generative AI Christian Tarsney Published: 2024-01-20Area: Deception & FailureCitations: 21 Tags: ai-safety, deception-failure, theoretical | 2024-01-20 | Deception & Failure | ai-safety, deception-failure, theoretical | E6 / R3 (93%) | 21 |
| Quantifying stability of non-power-seeking in artificial agents Evan Ryan Gunter, Victoria Krakovna, Yevgeny Liokumovich Published: 2024-01-07Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, theoretical | 2024-01-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (94%) | 2 |
| Honesty Is the Best Policy: Defining and Mitigating AI Deception Francesca Toni, Francesco Belardinelli, Francis Rhys Ward, Tom Everitt Published: 2023-12-03Area: Deception & FailureCitations: 48 Tags: ai-safety, deception-failure, theoretical | 2023-12-03 | Deception & Failure | ai-safety, deception-failure, theoretical | E4 / R2 (93%) | 48 |
| Hashmarks: Privacy-Preserving Benchmarks for High-Stakes AI Evaluation Paul Bricman Published: 2023-12-01Area: Safety EvaluationCitations: - Tags: ai-safety, safety-evaluation, theoretical | 2023-12-01 | Safety Evaluation | ai-safety, safety-evaluation, theoretical | E5 / R3 (95%) | - |
| Scalable AI Safety via Doubly-Efficient Debate Geoffrey Irving, Georgios Piliouras, Jonah Brown-Cohen Published: 2023-11-23Area: Scalable OversightCitations: 40 Tags: ai-safety, scalable-oversight, theoretical | 2023-11-23 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (94%) | 40 |
| Scheming AIs: Will AIs Fake Alignment During Training in Order to Get Power? Joe Carlsmith Published: 2023-11-14Area: Deception & FailureCitations: 59 Tags: ai-safety, alignment-training, deception-failure, theoretical | 2023-11-14 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E6 / R3 (94%) | 59 |
| The Linear Representation Hypothesis and the Geometry of Large Language Models Kiho Park, Victor Veitch, Yo Joong Choe Published: 2023-11-07Area: Representation AnalysisCitations: 371 Tags: ai-safety, representation-analysis, theoretical | 2023-11-07 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (95%) | 371 |
| The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback Nathan Lambert, Roberto Calandra Published: 2023-10-31Area: Alignment TrainingCitations: 42 Tags: ai-safety, alignment-training, safety-evaluation, theoretical | 2023-10-31 | Alignment Training | ai-safety, alignment-training, safety-evaluation, theoretical | E5 / R3 (95%) | 42 |
| A General Theoretical Paradigm to Understand Learning from Human Feedback Bilal Piot, Daniele Calandriello, Daniel Guo, Mark Rowland Published: 2023-10-18Area: Alignment TrainingCitations: 894 Tags: ai-safety, alignment-training, theoretical | 2023-10-18 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (97%) | 894 |
| Dynamical versus Bayesian Phase Transitions in a Toy Model of Superposition Daniel Murfet, Edmund Lau, Jake Mendel, Susan Wei Published: 2023-10-10Area: Training DynamicsCitations: 22 Tags: ai-safety, theoretical, training-dynamics | 2023-10-10 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R3 (95%) | 22 |
| The Local Learning Coefficient: A Singularity-Aware Complexity Measure Daniel Murfet, Edmund Lau, George Wang, Susan Wei Published: 2023-08-23Area: Training DynamicsCitations: 34 Tags: ai-safety, theoretical, training-dynamics | 2023-08-23 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R3 (96%) | 34 |
| A Geometric Notion of Causal Probing Alexander Warstadt, Anej Svete, Cl茅ment Guerner, Ryan Cotterell Published: 2023-07-27Area: Representation AnalysisCitations: 25 Tags: ai-safety, representation-analysis, theoretical | 2023-07-27 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (92%) | 25 |
| LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? David Glukhov, Ilia Shumailov, Nicolas Papernot, Vardan Papyan Published: 2023-07-20Area: Adversarial RobustnessCitations: 72 Tags: adversarial-robustness, ai-safety, theoretical | 2023-07-20 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (94%) | 72 |
| Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety Ben Smith, Catalin Mitelut, Peter Vamplew Published: 2023-05-30Area: Formal/TheoreticalCitations: 7 Tags: ai-safety, formaltheoretical, theoretical | 2023-05-30 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 7 |
| A Technical Note on Bilinear Layers for Interpretability Lee Sharkey Published: 2023-05-05Area: Mechanistic Interp.Citations: 10 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-05-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 10 |
| Fundamental Limitations of Alignment in Large Language Models Amnon Shashua, Noam Wies, Oshri Avnery, Yoav Levine Published: 2023-04-19Area: Formal/TheoreticalCitations: 178 Tags: adversarial-robustness, ai-safety, alignment-training, formaltheoretical, theoretical | 2023-04-19 | Formal/Theoretical | adversarial-robustness, ai-safety, alignment-training, formaltheoretical, theoretical | E4 / R2 (94%) | 178 |
| Power-seeking can be probable and predictive for trained agents Janos Kramar, Victoria Krakovna Published: 2023-04-13Area: Formal/TheoreticalCitations: 22 Tags: ai-safety, formaltheoretical, theoretical | 2023-04-13 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E6 / R3 (96%) | 22 |
| Characterizing Manipulation from AI Systems Alan Chan, David Krueger, Henry Ashton, Micah Carroll Published: 2023-03-16Area: Deception & FailureCitations: 66 Tags: ai-safety, deception-failure, theoretical | 2023-03-16 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R4 (93%) | 66 |
| Conditioning Predictive Models: Risks and Strategies Adam S. Jermyn, Evan Hubinger, Johannes Treutlein, Kate Woolverton Published: 2023-02-02Area: Deception & FailureCitations: 8 Tags: ai-safety, alignment-training, deception-failure, theoretical | 2023-02-02 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E5 / R3 (94%) | 8 |
| Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Amir Zur, Aryaman Arora, Atticus Geiger, Christopher Potts Published: 2023-01-11Area: Mechanistic Interp.Citations: 118 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2023-01-11 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (95%) | 118 |
| Transformers Learn In-Context by Gradient Descent Alexander Mordvintsev, Andrey Zhmoginov, Ettore Randazzo, Eyvind Niklasson Published: 2022-12-15Area: Training DynamicsCitations: 677 Tags: ai-safety, theoretical, training-dynamics | 2022-12-15 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R3 (95%) | 677 |