Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| The Persian Rug: solving toy models of superposition using large-scale symmetries Aditya Cowsik, Alex Infanger, Kfir Dolev Published: 2024-10-15Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2024-10-15 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| On Goodhart's Law, with an Application to Value Alignment El-Mahdi El-Mhamdi, L锚-Nguy锚n Hoang Published: 2024-10-12Area: Formal/TheoreticalCitations: 6 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-10-12 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 6 |
| The Computational Complexity of Circuit Discovery for Inner Interpretability Federico Adolfi, Martina G. Vilas, Todd Wareham Published: 2024-10-10Area: Formal/TheoreticalCitations: 14 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2024-10-10 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (98%) | 14 |
| RL, but don't do anything I wouldn't do Marcus Hutter, Michael K. Cohen, Stuart Russell, Yoshua Bengio Published: 2024-10-08Area: Alignment TrainingCitations: 2 Tags: ai-safety, alignment-training, theoretical | 2024-10-08 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (94%) | 2 |
| Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment Nicholas A. Caputo Published: 2024-10-07Area: Alignment TrainingCitations: 2 Tags: ai-safety, alignment-training, theoretical | 2024-10-07 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R4 (94%) | 2 |
| Provable Weak-to-Strong Generalization via Benign Overfitting Anant Sahai, David X. Wu Published: 2024-10-06Area: Scalable OversightCitations: 15 Tags: ai-safety, scalable-oversight, theoretical | 2024-10-06 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (95%) | 15 |
| Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols Alessandro Abate, Buck Shlegeris, Charlie Griffin, Louis Thomson Published: 2024-09-12Area: Scalable OversightCitations: 17 Tags: ai-safety, safety-evaluation, scalable-oversight, theoretical | 2024-09-12 | Scalable Oversight | ai-safety, safety-evaluation, scalable-oversight, theoretical | E5 / R3 (94%) | 17 |
| On the Complexity of Neural Computation in Superposition Micah Adler, Nir Shavit Published: 2024-09-05Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, mechanistic-interp, theoretical | 2024-09-05 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | 11 |
| Mathematical Models of Computation in Superposition Dmitry Vaintrob, Jake Mendel, Kaarel H盲nni, Lawrence Chan Published: 2024-08-10Area: Mechanistic Interp.Citations: 23 Tags: ai-safety, mechanistic-interp, theoretical | 2024-08-10 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | 23 |
| Can DPO Learn Diverse Human Values? A Theoretical Scaling Law Shawn Im, Yixuan Li Published: 2024-08-06Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2024-08-06 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R2 (94%) | 7 |
| Mission Impossible: A Statistical Perspective on Jailbreaking LLMs Jingtong Su, Julia Kempe, Karen Ullrich Published: 2024-08-02Area: Adversarial RobustnessCitations: 26 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-08-02 | Adversarial Robustness | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 26 |
| Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization Akshay Krishnamurthy, Audrey Huang, Dylan J. Foster, Jason D. Lee Published: 2024-07-18Area: Alignment TrainingCitations: 54 Tags: ai-safety, alignment-training, theoretical | 2024-07-18 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (96%) | 54 |
| Mechanistically Interpreting a Transformer-based 2-SAT Solver: An Axiomatic Approach Corina P膬s膬reanu, Nils Palumbo, Ravi Mangal, Saranya Vijayakumar Published: 2024-07-18Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2024-07-18 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 2 |
| Learning Dynamics of LLM Finetuning Danica J. Sutherland, Yi Ren Published: 2024-07-15Area: Training DynamicsCitations: 67 Tags: ai-safety, alignment-training, theoretical, training-dynamics | 2024-07-15 | Training Dynamics | ai-safety, alignment-training, theoretical, training-dynamics | E5 / R3 (94%) | 67 |
| Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Aaron Mueller Published: 2024-07-05Area: Mechanistic Interp.Citations: 18 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-07-05 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (92%) | 18 |
| A False Sense of Safety: Unsafe Information Leakage in 'Safe' AI Responses David Glukhov, Ilia Shumailov, Nicolas Papernot, Vardan Papyan Published: 2024-07-02Area: Adversarial RobustnessCitations: 11 Tags: adversarial-robustness, ai-safety, theoretical | 2024-07-02 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (94%) | 11 |
| Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs? Elias Stengel-Eskin, Mohit Bansal, Peter Hase, Thomas Hofweber Published: 2024-06-27Area: Model EditingCitations: 24 Tags: ai-safety, model-editing, theoretical | 2024-06-27 | Model Editing | ai-safety, model-editing, theoretical | E5 / R3 (96%) | 24 |
| Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference Anton Xue, Avishree Khare, Eric Wong, Rajeev Alur Published: 2024-06-21Area: Adversarial RobustnessCitations: 4 Tags: adversarial-robustness, ai-safety, theoretical | 2024-06-21 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (96%) | 4 |
| Data Shapley in One Training Run Dawn Song, Jiachen T. Wang, Prateek Mittal, Ruoxi Jia Published: 2024-06-16Area: Training DynamicsCitations: 48 Tags: ai-safety, theoretical, training-dynamics | 2024-06-16 | Training Dynamics | ai-safety, theoretical, training-dynamics | E5 / R3 (95%) | 48 |
| Information Theoretic Guarantees For Policy Alignment In Large Language Models Youssef Mroueh Published: 2024-06-09Area: Formal/TheoreticalCitations: 19 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-06-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 19 |
| Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models Alkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Grigoris Velegkas Published: 2024-06-09Area: Deception & FailureCitations: 2 Tags: ai-safety, deception-failure, theoretical | 2024-06-09 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (94%) | 2 |
| The Geometry of Categorical and Hierarchical Concepts in Large Language Models Kiho Park, Victor Veitch, Yibo Jiang, Yo Joong Choe Published: 2024-06-03Area: Representation AnalysisCitations: 76 Tags: ai-safety, representation-analysis, theoretical | 2024-06-03 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (95%) | 76 |
| One-Shot Safety Alignment for Large Language Models via Optimal Dualization Dongsheng Ding, Edgar Dobriban, Hamed Hassani, Osbert Bastani Published: 2024-05-29Area: Alignment TrainingCitations: 18 Tags: ai-safety, alignment-training, theoretical | 2024-05-29 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (94%) | 18 |
| A Theoretical Understanding of Self-Correction through In-context Alignment Stefanie Jegelka, Yifei Wang, Yisen Wang, Yuyang Wu Published: 2024-05-28Area: Alignment TrainingCitations: 56 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-05-28 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E6 / R4 (93%) | 56 |
| Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer Boyi Liu, Hongyi Guo, Jose Blanchet, Miao Lu Published: 2024-05-26Area: Alignment TrainingCitations: 90 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-05-26 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (97%) | 90 |
| A statistical framework for weak-to-strong generalization Felipe Maia Polo, Mikhail Yurochkin, Moulinath Banerjee, Seamus Somerstep Published: 2024-05-25Area: Scalable OversightCitations: 8 Tags: ai-safety, scalable-oversight, theoretical | 2024-05-25 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (95%) | 8 |
| Using Degeneracy in the Loss Landscape for Mechanistic Interpretability Cindy Wu, Dan Braun, Jake Mendel, Kaarel H盲nni Published: 2024-05-17Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2024-05-17 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (95%) | 11 |
| Human-AI Safety: A Descendant of Generative AI and Control Systems Safety Andrea Bajcsy, Jaime F. Fisac Published: 2024-05-16Area: Formal/TheoreticalCitations: 9 Tags: ai-safety, formaltheoretical, theoretical | 2024-05-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R4 (93%) | 9 |
| Understanding the Learning Dynamics of Alignment with Human Feedback Shawn Im, Yixuan Li Published: 2024-03-27Area: Alignment TrainingCitations: 18 Tags: ai-safety, alignment-training, theoretical | 2024-03-27 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 18 |
| The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists Elliott Thornley Published: 2024-03-07Area: Formal/TheoreticalCitations: 18 Tags: ai-safety, formaltheoretical, theoretical | 2024-03-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | 18 |