Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Honesty Is the Best Policy: Defining and Mitigating AI Deception Francesca Toni, Francesco Belardinelli, Francis Rhys Ward, Tom Everitt Published: 2023-12-03Area: Deception & FailureCitations: 48 Tags: ai-safety, deception-failure, theoretical | 2023-12-03 | Deception & Failure | ai-safety, deception-failure, theoretical | E4 / R2 (93%) | 48 |
| Quantifying stability of non-power-seeking in artificial agents Evan Ryan Gunter, Victoria Krakovna, Yevgeny Liokumovich Published: 2024-01-07Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, theoretical | 2024-01-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (94%) | 2 |
| Deception and Manipulation in Generative AI Christian Tarsney Published: 2024-01-20Area: Deception & FailureCitations: 21 Tags: ai-safety, deception-failure, theoretical | 2024-01-20 | Deception & Failure | ai-safety, deception-failure, theoretical | E6 / R3 (93%) | 21 |
| Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback Dahua Lin, Hang Yan, Junjie Ye, Qiming Ge Published: 2024-01-21Area: Alignment TrainingCitations: 22 Tags: ai-safety, alignment-training, theoretical | 2024-01-21 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R4 (96%) | 22 |
| Tradeoffs Between Alignment and Helpfulness in Language Models Amnon Shashua, Binyamin Rothberg, Dorin Shteyman, Noam Wies Published: 2024-01-29Area: Representation AnalysisCitations: 21 Tags: ai-safety, alignment-training, representation-analysis, theoretical | 2024-01-29 | Representation Analysis | ai-safety, alignment-training, representation-analysis, theoretical | E5 / R3 (94%) | 21 |
| The Reasons that Agents Act: Intention and Instrumental Goals Francesca Toni, Francesco Belardinelli, Francis Rhys Ward, Matt MacDermott Published: 2024-02-11Area: Agent SafetyCitations: 22 Tags: agent-safety, ai-safety, theoretical | 2024-02-11 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R3 (96%) | 22 |
| Representation Surgery: Theory and Practice of Affine Steering Jonathan Herzig, Ponnurangam Kumaraguru, Roee Aharoni, Ryan Cotterell Published: 2024-02-15Area: Representation AnalysisCitations: 32 Tags: ai-safety, representation-analysis, theoretical | 2024-02-15 | Representation Analysis | ai-safety, representation-analysis, theoretical | E4 / R3 (94%) | 32 |
| Reward Generalization in RLHF: A Topological Perspective Dong Yan, Fanzhi Zeng, Han Yang, Jiaming Ji Published: 2024-02-15Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2024-02-15 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 7 |
| Robust agents learn causal world models Jonathan Richens, Tom Everitt Published: 2024-02-16Area: Formal/TheoreticalCitations: 67 Tags: ai-safety, formaltheoretical, theoretical | 2024-02-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 67 |
| Immunization against harmful fine-tuning attacks Domenic Rosati, Frank Rudzicz, Hassan Sajjad, Jan Batzner Published: 2024-02-26Area: Adversarial RobustnessCitations: 34 Tags: adversarial-robustness, ai-safety, theoretical | 2024-02-26 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E4 / R2 (97%) | 34 |
| When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning Anca Dragan, Davis Foote, Erik Jenner, Leon Lang Published: 2024-02-27Area: Deception & FailureCitations: 13 Tags: ai-safety, deception-failure, theoretical | 2024-02-27 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (94%) | 13 |
| Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking Anca Dragan, Cassidy Laidlaw, Shivam Singhal Published: 2024-03-05Area: Deception & FailureCitations: 26 Tags: ai-safety, deception-failure, theoretical | 2024-03-05 | Deception & Failure | ai-safety, deception-failure, theoretical | E6 / R4 (94%) | 26 |
| The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists Elliott Thornley Published: 2024-03-07Area: Formal/TheoreticalCitations: 18 Tags: ai-safety, formaltheoretical, theoretical | 2024-03-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | 18 |
| Understanding the Learning Dynamics of Alignment with Human Feedback Shawn Im, Yixuan Li Published: 2024-03-27Area: Alignment TrainingCitations: 18 Tags: ai-safety, alignment-training, theoretical | 2024-03-27 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 18 |
| Human-AI Safety: A Descendant of Generative AI and Control Systems Safety Andrea Bajcsy, Jaime F. Fisac Published: 2024-05-16Area: Formal/TheoreticalCitations: 9 Tags: ai-safety, formaltheoretical, theoretical | 2024-05-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R4 (93%) | 9 |
| Using Degeneracy in the Loss Landscape for Mechanistic Interpretability Cindy Wu, Dan Braun, Jake Mendel, Kaarel H盲nni Published: 2024-05-17Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-05-17 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 11 |
| A statistical framework for weak-to-strong generalization Felipe Maia Polo, Mikhail Yurochkin, Moulinath Banerjee, Seamus Somerstep Published: 2024-05-25Area: Scalable OversightCitations: 8 Tags: ai-safety, scalable-oversight, theoretical | 2024-05-25 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (95%) | 8 |
| Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer Boyi Liu, Hongyi Guo, Jose Blanchet, Miao Lu Published: 2024-05-26Area: Alignment TrainingCitations: 90 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-05-26 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (97%) | 90 |
| A Theoretical Understanding of Self-Correction through In-context Alignment Stefanie Jegelka, Yifei Wang, Yisen Wang, Yuyang Wu Published: 2024-05-28Area: Alignment TrainingCitations: 56 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-05-28 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E6 / R4 (93%) | 56 |
| One-Shot Safety Alignment for Large Language Models via Optimal Dualization Dongsheng Ding, Edgar Dobriban, Hamed Hassani, Osbert Bastani Published: 2024-05-29Area: Alignment TrainingCitations: 18 Tags: ai-safety, alignment-training, theoretical | 2024-05-29 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (94%) | 18 |
| The Geometry of Categorical and Hierarchical Concepts in Large Language Models Kiho Park, Victor Veitch, Yibo Jiang, Yo Joong Choe Published: 2024-06-03Area: Representation AnalysisCitations: 76 Tags: ai-safety, representation-analysis, theoretical | 2024-06-03 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (95%) | 76 |
| Information Theoretic Guarantees For Policy Alignment In Large Language Models Youssef Mroueh Published: 2024-06-09Area: Formal/TheoreticalCitations: 19 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-06-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 19 |
| Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models Alkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Grigoris Velegkas Published: 2024-06-09Area: Deception & FailureCitations: 2 Tags: ai-safety, deception-failure, theoretical | 2024-06-09 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (94%) | 2 |
| Data Shapley in One Training Run Dawn Song, Jiachen T. Wang, Prateek Mittal, Ruoxi Jia Published: 2024-06-16Area: Training DynamicsCitations: 48 Tags: ai-safety, theoretical, training-dynamics | 2024-06-16 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R3 (95%) | 48 |
| Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference Anton Xue, Avishree Khare, Eric Wong, Rajeev Alur Published: 2024-06-21Area: Adversarial RobustnessCitations: 4 Tags: adversarial-robustness, ai-safety, theoretical | 2024-06-21 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (96%) | 4 |
| Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs? Elias Stengel-Eskin, Mohit Bansal, Peter Hase, Thomas Hofweber Published: 2024-06-27Area: Model EditingCitations: 24 Tags: ai-safety, model-editing, theoretical | 2024-06-27 | Model Editing | ai-safety, model-editing, theoretical | E5 / R3 (96%) | 24 |
| A False Sense of Safety: Unsafe Information Leakage in 'Safe' AI Responses David Glukhov, Ilia Shumailov, Nicolas Papernot, Vardan Papyan Published: 2024-07-02Area: Adversarial RobustnessCitations: 11 Tags: adversarial-robustness, ai-safety, theoretical | 2024-07-02 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (94%) | 11 |
| Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Aaron Mueller Published: 2024-07-05Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-07-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 18 |
| Learning Dynamics of LLM Finetuning Danica J. Sutherland, Yi Ren Published: 2024-07-15Area: Training DynamicsCitations: 67 Tags: ai-safety, alignment-training, theoretical, training-dynamics | 2024-07-15 | Training Dynamics | ai-safety, alignment-training, theoretical, training-dynamics | E5 / R3 (94%) | 67 |
| Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization Akshay Krishnamurthy, Audrey Huang, Dylan J. Foster, Jason D. Lee Published: 2024-07-18Area: Alignment TrainingCitations: 54 Tags: ai-safety, alignment-training, theoretical | 2024-07-18 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (96%) | 54 |