Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 199 papers (page 1 of 7)路 100 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment Elias Malomgr茅, Pieter Simoens Published: 2026-02-16Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, theoretical | 2026-02-16 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (97%) | - |
| Unifying Stable Optimization and Reference Regularization in RLHF Dadong Wang, He Zhao, Li He, Lina Yao Published: 2026-02-12Area: Alignment TrainingCitations: 1 Tags: ai-safety, alignment-training, theoretical | 2026-02-12 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (94%) | 1 |
| Authenticated Workflows: A Systems Approach to Protecting Agentic AI Mohan Rajagopalan, Vinay Rao Published: 2026-02-11Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, theoretical | 2026-02-11 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R4 (96%) | - |
| How Many Features Can a Language Model Store Under the Linear Representation Hypothesis? Jon Kleinberg, Kenny Peng, Nikhil Garg Published: 2026-02-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2026-02-11 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E4 / R2 (93%) | - |
| Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI Mohan Rajagopalan, Vinay Rao Published: 2026-02-11Area: Adversarial RobustnessCitations: 1 Tags: adversarial-robustness, ai-safety, theoretical | 2026-02-11 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (99%) | 1 |
| Towards Poisoning Robustness Certification for Natural Language Generation Matthew Wicker, Mihnea Ghitu Published: 2026-02-10Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, theoretical | 2026-02-10 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (95%) | - |
| Trustworthy Agentic AI Requires Deterministic Architectural Boundaries Manish Bhattarai, Minh Vu Published: 2026-02-10Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, alignment-training, theoretical | 2026-02-10 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E5 / R4 (97%) | - |
| Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints Andres Saurez, Dongsoo Har, Yousung Lee Published: 2026-02-10Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-02-10 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities Majid Ghasemi, Mark Crowley Published: 2026-02-08Area: Deception & FailureCitations: - Tags: ai-safety, alignment-training, deception-failure, theoretical | 2026-02-08 | Deception & Failure | ai-safety, alignment-training, deception-failure, theoretical | E6 / R3 (96%) | - |
| Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective Cheol Woo Kim, Davin Choo, Milind Tambe, Tzeh Yuan Neoh Published: 2026-02-06Area: Formal/TheoreticalCitations: - Tags: adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | 2026-02-06 | Formal/Theoretical | adversarial-robustness, ai-safety, formaltheoretical, safety-evaluation, theoretical | E5 / R3 (93%) | - |
| Alignment Verifiability in Large Language Models: Normative Indistinguishability under Behavioral Evaluation Igor Santos-Grueiro Published: 2026-02-05Area: Deception & FailureCitations: 1 Tags: ai-safety, alignment-training, deception-failure, safety-evaluation, theoretical | 2026-02-05 | Deception & Failure | ai-safety, alignment-training, deception-failure, safety-evaluation, theoretical | E4 / R2 (94%) | 1 |
| Subliminal Effects in Your Data: A General Mechanism via Log-Linearity Abhishek Shetty, Allen Liu, Ankur Moitra, Ishaq Aden-Ali Published: 2026-02-04Area: Deception & FailureCitations: - Tags: ai-safety, deception-failure, theoretical | 2026-02-04 | Deception & Failure | ai-safety, deception-failure, theoretical | E4 / R3 (94%) | - |
| Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO Akifumi Wachi, Kohei Miyaguchi, Rei Higuchi, Shokichi Takakura Published: 2026-02-02Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, theoretical | 2026-02-02 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (93%) | - |
| Spectral Superposition: A Theory of Feature Geometry Amir Abdullah, Georgi Ivanov, Narmeen Oozeer, Shivam Raval Published: 2026-02-02Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2026-02-02 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |
| Towards Understanding Steering Strength Damien Garreau, Magamed Taimeskhanov, Samuel Vaiter Published: 2026-02-02Area: Representation AnalysisCitations: - Tags: ai-safety, representation-analysis, theoretical | 2026-02-02 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (96%) | - |
| Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics Chenyan Wu, Haiwen Hong, Hengyu Sun, Huajun Chen Published: 2026-02-02Area: Model EditingCitations: - Tags: ai-safety, model-editing, theoretical | 2026-02-02 | Model Editing | ai-safety, model-editing, theoretical | E5 / R3 (95%) | - |
| How RLHF Amplifies Sycophancy Ariel D. Procaccia, Gerdus Benade, Itai Shapira Published: 2026-02-01Area: Deception & FailureCitations: - Tags: ai-safety, deception-failure, theoretical | 2026-02-01 | Deception & Failure | ai-safety, deception-failure, theoretical | E5 / R3 (94%) | - |
| Jailbreaking LLMs via Calibration Yongkang Guo, Yuqing Kong, Yuxuan Lu Published: 2026-01-31Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, theoretical | 2026-01-31 | Adversarial Robustness | adversarial-robustness, ai-safety, theoretical | E5 / R3 (94%) | - |
| Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective Ce Li, Haichuan Wang, Hezi Jiang, Lingkai Kong Published: 2026-01-31Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, theoretical | 2026-01-31 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (96%) | - |
| Dynamics Reveals Structure: Challenging the Linear Propagation Assumption B谩lint Mucs谩nyi, Hoyeon Chang, Seong Joon Oh Published: 2026-01-29Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2026-01-29 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E5 / R3 (92%) | - |
| Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming Abhinav Palia, A. Feder Cooper, Alexandra Chouldechova, Dan Vann Published: 2026-01-26Area: Safety EvaluationCitations: 1 Tags: ai-safety, red-teaming, safety-evaluation, theoretical | 2026-01-26 | Safety Evaluation | ai-safety, red-teaming, safety-evaluation, theoretical | E5 / R3 (95%) | 1 |
| Feature-Space Adversarial Robustness Certification for Multimodal Large Language Models Chenqi Kong, Meiwen Ding, Song Xia, Wenhan Yang Published: 2026-01-22Area: Multimodal SafetyCitations: - Tags: adversarial-robustness, ai-safety, multimodal-safety, theoretical | 2026-01-22 | Multimodal Safety | adversarial-robustness, ai-safety, multimodal-safety, theoretical | E4 / R3 (95%) | - |
| Patterning: The Dual of Interpretability Daniel Murfet, George Wang Published: 2026-01-20Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2026-01-20 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E6 / R4 (95%) | - |
| Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment Felix Jahn, Kevin Baum, Lisa Dargasz, Patrick Schramowski Published: 2026-01-15Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, alignment-training, theoretical | 2026-01-15 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E6 / R4 (96%) | - |
| Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling Weiqiang Zheng, Yang Cai Published: 2026-01-13Area: Alignment TrainingCitations: - Tags: ai-safety, alignment-training, theoretical | 2026-01-13 | Alignment Training | ai-safety, alignment-training, theoretical | E5 / R3 (93%) | - |
| Towards Provably Secure Generative AI: Reliable Consensus Sampling Baohan Huang, Bo Ran, Cong Zuo, Haibin Zhang Published: 2025-12-31Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, theoretical | 2025-12-31 | Formal/Theoretical | ai-safety, formaltheoretical, theoretical | E4 / R3 (94%) | - |
| Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability Ge Yan, Tsui-Wei (Lily) Weng, Tuomas Oikarinen Published: 2025-12-19Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, theoretical | 2025-12-19 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, theoretical | E5 / R3 (96%) | - |
| Provably Extracting the Features from a General Superposition Allen Liu Published: 2025-12-17Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, interpretability, theoretical | 2025-12-17 | Formal/Theoretical | ai-safety, formaltheoretical, interpretability, theoretical | E4 / R2 (96%) | - |
| Practical challenges of control monitoring in frontier AI deployments Alan Cooney, Charlie Griffin, David Lindner, Geoffrey Irving Published: 2025-12-15Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-12-15 | Agent Safety | agent-safety, ai-safety, theoretical | E6 / R4 (93%) | 1 |
| Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality Eric P. Xing, Guangyi Chen, Kun Zhang, Lingjing Kong Published: 2025-12-11Area: Mechanistic Interp.Citations: - Tags: ai-safety, mechanistic-interp, theoretical | 2025-12-11 | Mechanistic Interp. | ai-safety, mechanistic-interp, theoretical | E5 / R3 (94%) | - |