Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Representation Learning on a Random Lattice Aryeh Brill Published: 2025-04-28Area: Mechanistic Interp.Citations: 1 Tags: ai-safety, mechanistic-interp, theoretical | 2025-04-28 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E6 / R3 (96%) | 1 |
| Scaling Laws For Scalable Oversight David D. Baek, Joshua Engels, Max Tegmark, Subhash Kantamneni Published: 2025-04-25Area: Scalable OversightCitations: 4 Tags: ai-safety, scalable-oversight, theoretical | 2025-04-25 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E7 / R3 (97%) | 4 |
| Guillotine: Hypervisors for Isolating Malicious AIs James Mickens, Ravi Netravali, Sarah Radway Published: 2025-04-22Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-04-22 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R4 (97%) | 1 |
| Towards Combinatorial Interpretability of Neural Computation Dan Alistarh, Micah Adler, Nir Shavit Published: 2025-04-10Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-04-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (93%) | 7 |
| How to evaluate control measures for LLM agents? A trajectory from today to superintelligence Buck Shlegeris, Geoffrey Irving, Mikita Balesni, Tomek Korbak Published: 2025-04-07Area: Agent SafetyCitations: 11 Tags: agent-safety, ai-safety, safety-evaluation, theoretical | 2025-04-07 | Agent Safety | agent-safety, ai-safety, safety-evaluation, theoretical | E5 / R3 (92%) | 11 |
| AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability Alexander Boyd, Fernando Rosas, Manuel Baltieri Published: 2025-04-06Area: Formal/TheoreticalCitations: 2 Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-04-06 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E6 / R3 (93%) | 2 |
| Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality Adam Davies, Julia Hockenmaier, Marc E. Canby, Sewoong Lee Published: 2025-03-31Area: Mechanistic Interp.Citations: 2 Tags: ai-safety, mechanistic-interp, theoretical | 2025-03-31 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (93%) | 2 |
| Policy Teaching via Data Poisoning in Learning from Human Preferences Adish Singla, Andi Nika, Debmalya Mandal, Goran Radanovi膰 Published: 2025-03-13Area: Adversarial RobustnessCitations: 1 Tags: adversarial-robustness, ai-safety, theoretical | 2025-03-13 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (96%) | 1 |
| Mitigating Preference Hacking in Policy Optimization with Pessimism Adam Fisch, Alekh Agarwal, Christoph Dann, Dhawal Gupta Published: 2025-03-10Area: Alignment TrainingCitations: 3 Tags: ai-safety, alignment-training, theoretical | 2025-03-10 | Alignment Training | ai-safety, alignment-training, theoretical | E6 / R3 (95%) | 3 |
| CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models Jiaxu Liu, Jinwei Hu, Wenjie Ruan, Xiangyu Yin Published: 2025-03-08Area: Multimodal SafetyCitations: - Tags: adversarial-robustness, ai-safety, multimodal-safety, theoretical | 2025-03-08 | Multimodal Safety | adversarial-robustness, ai-safety, multimodal-safety, theoretical | E5 / R4 (96%) | - |
| From superposition to sparse codes: interpretable representations in neural networks Charles O'Neill, David Klindt, Harald Maurer, Nina Miolane Published: 2025-03-03Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, mechanistic-interp, safety-evaluation, theoretical | 2025-03-03 | Mechanistic Interp. | ai-safety, mechanistic-interp, safety-evaluation, theoretical | E5 / R3 (93%) | 7 |
| Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? Francois Portet, Maxime Meloux, Maxime Peyrard, Silviu Maniu Published: 2025-02-28Area: Mechanistic Interp.Citations: 14 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-02-28 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (94%) | 14 |
| Shh, don't say that! Domain Certification in LLMs Adel Bibi, Alasdair Paren, Bernard Ghanem, Cornelius Emde Published: 2025-02-26Area: Formal/TheoreticalCitations: 4 Tags: adversarial-robustness, ai-safety, formaltheoretical, theoretical | 2025-02-26 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, theoretical | E5 / R3 (93%) | 4 |
| Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation Trung Nguyen, Yan Leng Published: 2025-02-22Area: Representation AnalysisCitations: 1 Tags: ai-safety, representation-analysis, theoretical | 2025-02-22 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (94%) | 1 |
| Assessing confidence in frontier AI safety cases Alejandro Tlaie, Joshua Krook, Philip Fox, Simon Mylius Published: 2025-02-09Area: Safety EvaluationCitations: 4 Tags: ai-safety, safety-evaluation, theoretical | 2025-02-09 | Safety Evaluation | ai-safety, safety-evaluation, theoretical | E5 / R3 (94%) | 4 |
| Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis Aran Nayebi Published: 2025-02-09Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, alignment-training, formaltheoretical, theoretical | 2025-02-09 | Formal/Theoretical | ai-safety, alignment-training, formaltheoretical, theoretical | E4 / R3 (93%) | 4 |
| Robust LLM Alignment via Distributionally Robust Direct Preference Optimization Deepak Ramachandran, Dileep Kalathil, Kishan Panaganti, Rahul Jain Published: 2025-02-04Area: Alignment TrainingCitations: 11 Tags: ai-safety, alignment-training, theoretical | 2025-02-04 | Alignment Training | ai-safety, alignment-training, theoretical | E6 / R3 (95%) | 11 |
| On Almost Surely Safe Alignment of Large Language Models at Inference-Time Haitham Bou Ammar, Ilija Bogunovic, Jun Wang, Matthieu Zimmer Published: 2025-02-03Area: Alignment TrainingCitations: 7 Tags: ai-safety, alignment-training, theoretical | 2025-02-03 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R3 (96%) | 7 |
| LLM Safety Alignment is Divergence Estimation in Disguise Guang Lin, Qifan Song, Rajdeep Haldar, Yue Xing Published: 2025-02-02Area: Alignment TrainingCitations: 3 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2025-02-02 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (93%) | 3 |
| Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions Baharan Mirzasoleiman, Jiping Li, Yihao Xue Published: 2025-02-02Area: Scalable OversightCitations: 5 Tags: ai-safety, scalable-oversight, theoretical | 2025-02-02 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (94%) | 5 |
| A Three-Branch Checks-and-Balances Framework for Context-Aware Ethical Alignment of Large Language Models Edward Y. Chang Published: 2025-01-31Area: Alignment TrainingCitations: 3 Tags: adversarial-robustness, ai-safety, alignment-training, theoretical | 2025-01-31 | Alignment Training | adversarial-robustness, ai-safety, alignment-training, theoretical | E5 / R4 (95%) | 3 |
| Information-theoretic Distinctions Between Deception and Confusion Robin Young Published: 2025-01-27Area: Deception & FailureCitations: - Tags: ai-safety, alignment-training, deception-failure, theoretical | 2025-01-27 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E5 / R3 (94%) | - |
| Propositional Interpretability in Artificial Intelligence David J. Chalmers Published: 2025-01-27Area: Mechanistic Interp.Citations: 13 Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-01-27 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R3 (94%) | 13 |
| When Can Proxies Improve the Sample Complexity of Preference Learning? Alexander D'Amour, Daniel Augusto de Souza, Matt J. Kusner, Mengyue Yang Published: 2024-12-21Area: Alignment TrainingCitations: 1 Tags: ai-safety, alignment-training, theoretical | 2024-12-21 | Alignment Training | ai-safety, alignment-training, theoretical | E4 / R3 (95%) | 1 |
| Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations Alessandro Di Stefano, Paolo Bova, The Anh Han Published: 2024-12-19Area: Safety EvaluationCitations: 3 Tags: ai-safety, safety-evaluation, theoretical | 2024-12-19 | Safety Evaluation | ai-safety, safety-evaluation, theoretical | E5 / R3 (94%) | 3 |
| Neural Interactive Proofs Lewis Hammond, Sam Adam-Day Published: 2024-12-12Area: Scalable OversightCitations: 5 Tags: ai-safety, scalable-oversight, theoretical | 2024-12-12 | Scalable Oversight | ai-safety, scalable-oversight, theoretical | E5 / R3 (92%) | 5 |
| Measuring Goal-Directedness Francesco Belardinelli, James Fox, Matt MacDermott, Tom Everitt Published: 2024-12-06Area: Formal/TheoreticalCitations: 4 Tags: ai-safety, formaltheoretical, theoretical | 2024-12-06 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (95%) | 4 |
| Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration Chun Hei Yip, Jason Gross, Lawrence Chan, Rajashree Agrawal Published: 2024-12-04Area: Mechanistic Interp.Citations: 4 Tags: ai-safety, mechanistic-interp, theoretical | 2024-12-04 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (95%) | 4 |
| Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies Fr茅d茅ric Berdoz, Roger Wattenhofer Published: 2024-11-21Area: Formal/TheoreticalCitations: 1 Tags: ai-safety, formaltheoretical, theoretical | 2024-11-21 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (96%) | 1 |
| Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders Charles O'Neill, David Klindt Published: 2024-11-20Area: Mechanistic Interp.Citations: 7 Tags: ai-safety, mechanistic-interp, theoretical | 2024-11-20 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (97%) | 7 |