Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 31-41 of 41 papers (page 2 of 2)路 19 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Xingcheng Xu Published: 2025-07-27Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-07-27 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E5 / R4 (96%) | 4 |
| On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Haozhe Jiang, Nika Haghtalab Published: 2025-08-26Area: Formal/TheoreticalCitations: 3 Tags: adversarial-robustness, ai-safety, formaltheoretical, theoretical | 2025-08-26 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, theoretical | E6 / R3 (95%) | 3 |
| AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures? Florian Mai, Leonard Dung Published: 2025-10-13Area: Formal/TheoreticalCitations: - Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-10-13 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E8 / R2 (97%) | - |
| Corrigibility Transformation: Constructing Goals That Accept Updates Rubi Hudson Published: 2025-10-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2025-10-17 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (94%) | - |
| RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification Anonymous Authors Published: 2025-10-30Area: Formal/TheoreticalCitations: - Tags: ai-safety, empirical, formaltheoretical | 2025-10-30 | Formal/Theoretical | ai-safety, empirical, formaltheoretical | E5 / R4 (93%) | - |
| Provably Extracting the Features from a General Superposition Allen Liu Published: 2025-12-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-12-17 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E4 / R2 (96%) | - |
| Towards Provably Secure Generative AI: Reliable Consensus Sampling Baohan Huang, Bo Ran, Cong Zuo, Haibin Zhang Published: 2025-12-31Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2025-12-31 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | - |
| Dynamics Reveals Structure: Challenging the Linear Propagation Assumption B谩lint Mucs谩nyi, Hoyeon Chang, Seong Joon Oh Published: 2026-01-29Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2026-01-29 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (92%) | - |
| Towards Worst-Case Guarantees with Scale-Aware Interpretability Alexander Stapleton, Andrew Mack, Anindita Maiti, Artemy Kolchinsky Published: 2026-02-05Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, position | 2026-02-05 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, position | E5 / R3 (94%) | - |
| Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective Cheol Woo Kim, Davin Choo, Milind Tambe, Tzeh Yuan Neoh Published: 2026-02-06Area: Formal/TheoreticalCitations: - Tags: adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | 2026-02-06 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | E5 / R3 (93%) | - |
| Inaugural Workshop on Provably Safe and Beneficial AI (PSBAI) Stuart Russell Published: -Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, survey | - | Formal/Theoretical | ai-safety, formaltheoretical, survey | E6 / R3 (96%) | - |