Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 41 papers (page 1 of 2)路 37 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective Cheol Woo Kim, Davin Choo, Milind Tambe, Tzeh Yuan Neoh Published: 2026-02-06Area: Formal/TheoreticalCitations: - Tags: adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | 2026-02-06 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | E5 / R3 (93%) | - |
| Towards Worst-Case Guarantees with Scale-Aware Interpretability Alexander Stapleton, Andrew Mack, Anindita Maiti, Artemy Kolchinsky Published: 2026-02-05Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, position | 2026-02-05 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, position | E5 / R3 (94%) | - |
| Dynamics Reveals Structure: Challenging the Linear Propagation Assumption B谩lint Mucs谩nyi, Hoyeon Chang, Seong Joon Oh Published: 2026-01-29Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2026-01-29 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (92%) | - |
| Towards Provably Secure Generative AI: Reliable Consensus Sampling Baohan Huang, Bo Ran, Cong Zuo, Haibin Zhang Published: 2025-12-31Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2025-12-31 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | - |
| Provably Extracting the Features from a General Superposition Allen Liu Published: 2025-12-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-12-17 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E4 / R2 (96%) | - |
| RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification Anonymous Authors Published: 2025-10-30Area: Formal/TheoreticalCitations: - Tags: ai-safety, empirical, formaltheoretical | 2025-10-30 | Formal/Theoretical | ai-safety, empirical, formaltheoretical | E5 / R4 (93%) | - |
| Corrigibility Transformation: Constructing Goals That Accept Updates Rubi Hudson Published: 2025-10-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2025-10-17 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (94%) | - |
| AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures? Florian Mai, Leonard Dung Published: 2025-10-13Area: Formal/TheoreticalCitations: - Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-10-13 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E8 / R2 (97%) | - |
| On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Haozhe Jiang, Nika Haghtalab Published: 2025-08-26Area: Formal/TheoreticalCitations: 3 Tags: adversarial-robustness, ai-safety, formaltheoretical, theoretical | 2025-08-26 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, theoretical | E6 / R3 (95%) | 3 |
| The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Xingcheng Xu Published: 2025-07-27Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-07-27 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R4 (96%) | 4 |
| On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment Frauke Kreuter, Greg Gluch, Guy N. Rothblum, Omer Reingold Published: 2025-07-09Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-07-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (97%) | 4 |
| Out of Control - Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) Elija Perrier Published: 2025-06-21Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, alignment-training, formaltheoretical, position | 2025-06-21 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, position | E4 / R2 (94%) | 2 |
| The Limits of Predicting Agents from Behaviour Alexis Bellot, Jonathan Richens, Tom Everitt Published: 2025-06-03Area: Formal/TheoreticalCitations: 1 Tags: ai-safety, formaltheoretical, theoretical | 2025-06-03 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (93%) | 1 |
| Will artificial agents pursue power by default? Christian Tarsney Published: 2025-06-02Area: Formal/TheoreticalCitations: 1 Tags: ai-safety, formaltheoretical, theoretical | 2025-06-02 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (92%) | 1 |
| Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models Mohsen Bayati, William Overman Published: 2025-06-01Area: Formal/TheoreticalCitations: 3 Tags: ai-safety, formaltheoretical, theoretical | 2025-06-01 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 3 |
| AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability Alexander Boyd, Fernando Rosas, Manuel Baltieri Published: 2025-04-06Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-04-06 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (93%) | 2 |
| Shh, don't say that! Domain Certification in LLMs Adel Bibi, Alasdair Paren, Bernard Ghanem, Cornelius Emde Published: 2025-02-26Area: Formal/TheoreticalCitations: 4 Tags: adversarial-robustness, ai-safety, formaltheoretical, theoretical | 2025-02-26 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, theoretical | E5 / R3 (93%) | 4 |
| Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis Aran Nayebi Published: 2025-02-09Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-02-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E4 / R3 (93%) | 4 |
| You Are What You Eat: AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Alexander Gietelink Oldenziel, Daniel Murfet, George Wang, Jesse Hoogland Published: 2025-02-08Area: Formal/TheoreticalCitations: 9 Tags: ai-safety, alignment-training, formaltheoretical, position | 2025-02-08 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, position | E5 / R4 (96%) | 9 |
| Measuring Goal-Directedness Francesco Belardinelli, James Fox, Matt MacDermott, Tom Everitt Published: 2024-12-06Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, formaltheoretical, theoretical | 2024-12-06 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (95%) | 4 |
| Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies Fr茅d茅ric Berdoz, Roger Wattenhofer Published: 2024-11-21Area: Formal/TheoreticalCitations: 1 Tags: ai-safety, formaltheoretical, theoretical | 2024-11-21 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 1 |
| On Goodhart's Law, with an Application to Value Alignment El-Mahdi El-Mhamdi, L锚-Nguy锚n Hoang Published: 2024-10-12Area: Formal/TheoreticalCitations: 6 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-10-12 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 6 |
| The Computational Complexity of Circuit Discovery for Inner Interpretability Federico Adolfi, Martina G. Vilas, Todd Wareham Published: 2024-10-10Area: Formal/TheoreticalCitations: 14 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2024-10-10 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (98%) | 14 |
| Compact Proofs of Model Performance via Mechanistic Interpretability Alex Gibson, Chun Hei Yip, Euan Ong, Jason Gross Published: 2024-06-17Area: Formal/TheoreticalCitations: 12 Tags: ai-safety, empirical, formaltheoretical, interpretability | 2024-06-17 | Formal/Theoretical | ai-safety, empirical, formaltheoretical, interpretability | E4 / R3 (92%) | 12 |
| Information Theoretic Guarantees For Policy Alignment In Large Language Models Youssef Mroueh Published: 2024-06-09Area: Formal/TheoreticalCitations: 19 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2024-06-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R3 (95%) | 19 |
| Human-AI Safety: A Descendant of Generative AI and Control Systems Safety Andrea Bajcsy, Jaime F. Fisac Published: 2024-05-16Area: Formal/TheoreticalCitations: 9 Tags: ai-safety, formaltheoretical, theoretical | 2024-05-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R4 (93%) | 9 |
| Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems Alessandro Abate, Ben Goldhaber, Christian Szegedy, Clark Barrett Published: 2024-05-10Area: Formal/TheoreticalCitations: 102 Tags: ai-safety, formaltheoretical, position | 2024-05-10 | Formal/Theoretical | ai-safety, formaltheoretical, position | E5 / R4 (97%) | 102 |
| The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists Elliott Thornley Published: 2024-03-07Area: Formal/TheoreticalCitations: 18 Tags: ai-safety, formaltheoretical, theoretical | 2024-03-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | 18 |
| Robust agents learn causal world models Jonathan Richens, Tom Everitt Published: 2024-02-16Area: Formal/TheoreticalCitations: 67 Tags: ai-safety, formaltheoretical, theoretical | 2024-02-16 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 67 |
| Quantifying stability of non-power-seeking in artificial agents Evan Ryan Gunter, Victoria Krakovna, Yevgeny Liokumovich Published: 2024-01-07Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, theoretical | 2024-01-07 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (94%) | 2 |