Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents Diogo Cruz, Mohammad Ali Jauhar, Nidhi Sakpal, Tsimur Hadeliya Published: 2025-12-02Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-12-02 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 2 |
| Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective Anne Lauscher, Jae Hee Lee, Stefano V. Albrecht Published: 2025-12-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | 2025-12-04 | Agent Safety | agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | E5 / R3 (94%) | 1 |
| Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents Mingjie Tang, Tianze Hu, Zaiye Chen, Zhibo Liang Published: 2025-12-07Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-12-07 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (96%) | - |
| SoK: Trust-Authorization Mismatch in LLM Agent Interactions Guanquan Shi, Haohua Du, Song Bian, Weiwenpei Liu Published: 2025-12-07Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, survey | 2025-12-07 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 1 |
| MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents Gil Vernik, Jinhao Zhu, Kevin Tseng, Raluca Ada Popa Published: 2025-12-11Area: Agent SafetyCitations: 3 Tags: agent-safety, ai-safety, tool | 2025-12-11 | Agent Safety | agent-safety, ai-safety, tool | E4 / R3 (96%) | 3 |
| Async Control: Stress-testing Asynchronous Control Measures for LLM Agents Alan Cooney, Arathi Mani, Asa Cooper Stickland, Charlie Griffin Published: 2025-12-15Area: Agent SafetyCitations: 1 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-12-15 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (95%) | 1 |
| Practical challenges of control monitoring in frontier AI deployments Alan Cooney, Charlie Griffin, David Lindner, Geoffrey Irving Published: 2025-12-15Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-12-15 | Agent Safety | agent-safety, ai-safety, theoretical | E6 / R4 (93%) | 1 |
| Trust in LLM-controlled Robotics: a Survey of Security Threats, Defenses and Challenges Huaming Chen, Ian R. Manchester, Kim-Kwang Raymond Choo, Mitch Bryson Published: 2025-12-17Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-12-17 | Agent Safety | agent-safety, ai-safety, survey | E7 / R4 (93%) | - |
| Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation Dongqin Liu, Hongchang Yang, Songlin Hu, Wei Zhou Published: 2025-12-18Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2025-12-18 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (95%) | - |
| Distributional AGI Safety Julian Jacobs, Matija Franklin, Nenad Toma拧ev, S茅bastien Krier Published: 2025-12-18Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, position | 2025-12-18 | Agent Safety | agent-safety, ai-safety, position | E4 / R3 (92%) | 4 |
| Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection Vignesh Iyer Published: 2026-01-05Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-01-05 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | - |
| Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making Jaeyoon Seo, Jean Oh, Jihie Kim, Jua Han Published: 2026-01-09Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-01-09 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (98%) | - |
| Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis Boqi Chen, Da Song, Foutse Khomh, Lei Ma Published: 2026-01-13Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2026-01-13 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (97%) | - |
| Too Helpful to Be Safe: User-Mediated Attacks on Planning and Web-Use Agents Carsten Rudolph, Fengchao Chen, Tingmin Wu, Van Nguyen Published: 2026-01-14Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2026-01-14 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (93%) | 2 |
| AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior Asaf Shabtai, David Mimran, Denis Klimov, Gerard Levinov Published: 2026-01-15Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2026-01-15 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (96%) | 2 |
| Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment Felix Jahn, Kevin Baum, Lisa Dargasz, Patrick Schramowski Published: 2026-01-15Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, alignment-training, theoretical | 2026-01-15 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E6 / R4 (96%) | - |
| ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback Jing Shao, Lijun Li, Peiyang Liu, Shikun Zhang Published: 2026-01-15Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2026-01-15 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 2 |
| Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs Daniele Nardi, Federico Pierucci, Francesco Giarrusso, Marcantonio Bracale Syrnikov Published: 2026-01-16Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2026-01-16 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (92%) | 1 |
| MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning Correction Changyue Jiang, Geng Hong, Jiarun Dai, Wenqi Zhang Published: 2026-01-19Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-01-19 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | - |
| Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents Dmitry Namiot, Narek Maloyan Published: 2026-01-24Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-01-24 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | - |
| The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents Chen Chen, Kim Young Il, Kwok-Yan Lam, Qian Wang Published: 2026-01-24Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, alignment-training, benchmark | 2026-01-24 | Agent Safety | agent-safety, ai-safety, alignment-training, benchmark | E4 / R3 (96%) | - |
| Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses Hozefa Lakadawala, Mohsen Hatami, Van Tuan Pham, Yu Chen Published: 2026-01-28Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2026-01-28 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | - |
| StepShield: When, Not Whether to Intervene on Rogue Agents Gloria Felicia, Hemant Kumar, Jinfeng He, Michael Eniolade Published: 2026-01-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2026-01-29 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R4 (97%) | - |
| Multi-Agent Systems Should be Treated as Principal-Agent Problems Mihaela van der Schaar, Paulius Rauba, Simonas Cepenas Published: 2026-01-30Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, position | 2026-01-30 | Agent Safety | agent-safety, ai-safety, position | E5 / R4 (94%) | - |
| Evolving Interpretable Constitutions for Multi-Agent Coordination Alice Saito, Hershraj Niranjani, Phan Xuan Tan, Rayan Yessou Published: 2026-01-31Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-01-31 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (93%) | - |
| To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack Ruijie Meng, Terry Yue Zhuo, Wenbo Guo, Yangruibo Ding Published: 2026-02-01Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, position | 2026-02-01 | Agent Safety | agent-safety, ai-safety, position | E5 / R4 (94%) | - |
| LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios Chujia Hu, Dongrui Liu, Ge Gao, Tianyu Chen Published: 2026-02-03Area: Agent SafetyCitations: 1 Tags: adversarial-robustness, agent-safety, ai-safety, benchmark | 2026-02-03 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, benchmark | E4 / R3 (95%) | 1 |
| Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World Davood Soleymanzadeh, Fan Fei, Joonkyung Kim, Minghui Zheng Published: 2026-02-03Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, position | 2026-02-03 | Agent Safety | agent-safety, ai-safety, position | E5 / R4 (93%) | - |
| Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment An Zhang, Bingnan Liu, Chaochao Lu, Chenhang Cui Published: 2026-02-03Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2026-02-03 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (95%) | - |
| MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems Adish Singla, Goran Radanovic, Jonathan N枚ther Published: 2026-02-04Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2026-02-04 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (93%) | - |