Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Mechanistically Interpreting a Transformer-based 2-SAT Solver: An Axiomatic Approach Corina P膬s膬reanu, Nils Palumbo, Ravi Mangal, Saranya Vijayakumar Published: 2024-07-18Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2024-07-18 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 2 |
| Mission Impossible: A Statistical Perspective on Jailbreaking LLMs Jingtong Su, Julia Kempe, Karen Ullrich Published: 2024-08-02Area: Adversarial RobustnessCitations: 26 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2024-08-02 | Adversarial Robustness | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R3 (95%) | 26 |
| Can DPO Learn Diverse Human Values? A Theoretical Scaling Law Shawn Im, Yixuan Li Published: 2024-08-06Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2024-08-06 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R2 (94%) | 7 |
| Mathematical Models of Computation in Superposition Dmitry Vaintrob, Jake Mendel, Kaarel H盲nni, Lawrence Chan Published: 2024-08-10Area: Mechanistic Interp.Citations: 23 Tags: ai-safety, mechanistic-interp, theoretical | 2024-08-10 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | 23 |
| On the Complexity of Neural Computation in Superposition Micah Adler, Nir Shavit Published: 2024-09-05Area: Mechanistic Interp.Citations: 11 Tags: ai-safety, mechanistic-interp, theoretical | 2024-09-05 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | 11 |
| Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols Alessandro Abate, Buck Shlegeris, Charlie Griffin, Louis Thomson Published: 2024-09-12Area: Scalable OversightCitations: 17 Tags: ai-safety, safety-evaluation, scalable-oversight, theoretical | 2024-09-12 | Scalable Oversight | ai-safety, safety-evaluation, scalable-oversight, theoretical | E5 / R3 (94%) | 17 |
| Provable Weak-to-Strong Generalization via Benign Overfitting Anant Sahai, David X. Wu Published: 2024-10-06Area: Scalable OversightCitations: 15 Tags: ai-safety, scalable-oversight, theoretical | 2024-10-06 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (95%) | 15 |
| Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment Nicholas A. Caputo Published: 2024-10-07Area: Alignment TrainingCitations: 2 Tags: ai-safety, alignment-training, theoretical | 2024-10-07 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R4 (94%) | 2 |
| RL, but don't do anything I wouldn't do Marcus Hutter, Michael K. Cohen, Stuart Russell, Yoshua Bengio Published: 2024-10-08Area: Alignment TrainingCitations: 2 Tags: ai-safety, alignment-training, theoretical | 2024-10-08 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (94%) | 2 |
| The Computational Complexity of Circuit Discovery for Inner Interpretability Federico Adolfi, Martina G. Vilas, Todd Wareham Published: 2024-10-10Area: Formal/TheoreticalCitations: 14 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2024-10-10 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (98%) | 14 |
| On Goodhart's Law, with an Application to Value Alignment El-Mahdi El-Mhamdi, L锚-Nguy锚n Hoang Published: 2024-10-12Area: Formal/TheoreticalCitations: 6 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-10-12 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 6 |
| The Persian Rug: solving toy models of superposition using large-scale symmetries Aditya Cowsik, Alex Infanger, Kfir Dolev Published: 2024-10-15Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2024-10-15 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders Charles O'Neill, David Klindt Published: 2024-11-20Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, mechanistic-interp, theoretical | 2024-11-20 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (97%) | 7 |
| Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies Fr茅d茅ric Berdoz, Roger Wattenhofer Published: 2024-11-21Area: Formal/TheoreticalCitations: 1 Tags: ai-safety, formaltheoretical, theoretical | 2024-11-21 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 1 |
| Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration Chun Hei Yip, Jason Gross, Lawrence Chan, Rajashree Agrawal Published: 2024-12-04Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, mechanistic-interp, theoretical | 2024-12-04 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 4 |
| Measuring Goal-Directedness Francesco Belardinelli, James Fox, Matt MacDermott, Tom Everitt Published: 2024-12-06Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, formaltheoretical, theoretical | 2024-12-06 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (95%) | 4 |
| Neural Interactive Proofs Lewis Hammond, Sam Adam-Day Published: 2024-12-12Area: Scalable OversightCitations: 5 Tags: ai-safety, scalable-oversight, theoretical | 2024-12-12 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (92%) | 5 |
| Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations Alessandro Di Stefano, Paolo Bova, The Anh Han Published: 2024-12-19Area: Safety EvaluationCitations: 3 Tags: ai-safety, safety-evaluation, theoretical | 2024-12-19 | Safety Evaluation | ai-safety, safety-evaluation, theoretical | E5 / R3 (94%) | 3 |
| When Can Proxies Improve the Sample Complexity of Preference Learning? Alexander D'Amour, Daniel Augusto de Souza, Matt J. Kusner, Mengyue Yang Published: 2024-12-21Area: Alignment TrainingCitations: 1 Tags: ai-safety, alignment-training, theoretical | 2024-12-21 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R3 (95%) | 1 |
| Information-theoretic Distinctions Between Deception and Confusion Robin Young Published: 2025-01-27Area: Deception & FailureCitations: - Tags: ai-safety, alignment-training, deception-failure, theoretical | 2025-01-27 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E5 / R3 (94%) | - |
| Propositional Interpretability in Artificial Intelligence David J. Chalmers Published: 2025-01-27Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-01-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (94%) | 13 |
| A Three-Branch Checks-and-Balances Framework for Context-Aware Ethical Alignment of Large Language Models Edward Y. Chang Published: 2025-01-31Area: Alignment TrainingCitations: 3 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2025-01-31 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (95%) | 3 |
| LLM Safety Alignment is Divergence Estimation in Disguise Guang Lin, Qifan Song, Rajdeep Haldar, Yue Xing Published: 2025-02-02Area: Alignment TrainingCitations: 3 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2025-02-02 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (93%) | 3 |
| Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions Baharan Mirzasoleiman, Jiping Li, Yihao Xue Published: 2025-02-02Area: Scalable OversightCitations: 5 Tags: ai-safety, scalable-oversight, theoretical | 2025-02-02 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (94%) | 5 |
| On Almost Surely Safe Alignment of Large Language Models at Inference-Time Haitham Bou Ammar, Ilija Bogunovic, Jun Wang, Matthieu Zimmer Published: 2025-02-03Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2025-02-03 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R3 (96%) | 7 |
| Robust LLM Alignment via Distributionally Robust Direct Preference Optimization Deepak Ramachandran, Dileep Kalathil, Kishan Panaganti, Rahul Jain Published: 2025-02-04Area: Alignment TrainingCitations: 11 Tags: ai-safety, alignment-training, theoretical | 2025-02-04 | Alignment Training | ai-safety, alignment-training, theoretical | E6 / R3 (95%) | 11 |
| Assessing confidence in frontier AI safety cases Alejandro Tlaie, Joshua Krook, Philip Fox, Simon Mylius Published: 2025-02-09Area: Safety EvaluationCitations: 4 Tags: ai-safety, safety-evaluation, theoretical | 2025-02-09 | Safety Evaluation | ai-safety, safety-evaluation, theoretical | E5 / R3 (94%) | 4 |
| Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis Aran Nayebi Published: 2025-02-09Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-02-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E4 / R3 (93%) | 4 |
| Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation Trung Nguyen, Yan Leng Published: 2025-02-22Area: Representation AnalysisCitations: 1 Tags: ai-safety, representation-analysis, theoretical | 2025-02-22 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (94%) | 1 |
| Shh, don't say that! Domain Certification in LLMs Adel Bibi, Alasdair Paren, Bernard Ghanem, Cornelius Emde Published: 2025-02-26Area: Formal/TheoreticalCitations: 4 Tags: adversarial-robustness, ai-safety, formaltheoretical, theoretical | 2025-02-26 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, theoretical | E5 / R3 (93%) | 4 |