Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation Dongqin Liu, Hongchang Yang, Songlin Hu, Wei Zhou Published: 2025-12-18Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2025-12-18 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (95%) | - |
| Distributional AGI Safety Julian Jacobs, Matija Franklin, Nenad Toma拧ev, S茅bastien Krier Published: 2025-12-18Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, position | 2025-12-18 | Agent Safety | agent-safety, ai-safety, position | E4 / R3 (92%) | 4 |
| Trust in LLM-controlled Robotics: a Survey of Security Threats, Defenses and Challenges Huaming Chen, Ian R. Manchester, Kim-Kwang Raymond Choo, Mitch Bryson Published: 2025-12-17Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-12-17 | Agent Safety | agent-safety, ai-safety, survey | E7 / R4 (93%) | - |
| Async Control: Stress-testing Asynchronous Control Measures for LLM Agents Alan Cooney, Arathi Mani, Asa Cooper Stickland, Charlie Griffin Published: 2025-12-15Area: Agent SafetyCitations: 1 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-12-15 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (95%) | 1 |
| Practical challenges of control monitoring in frontier AI deployments Alan Cooney, Charlie Griffin, David Lindner, Geoffrey Irving Published: 2025-12-15Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-12-15 | Agent Safety | agent-safety, ai-safety, theoretical | E6 / R4 (93%) | 1 |
| MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents Gil Vernik, Jinhao Zhu, Kevin Tseng, Raluca Ada Popa Published: 2025-12-11Area: Agent SafetyCitations: 3 Tags: agent-safety, ai-safety, tool | 2025-12-11 | Agent Safety | agent-safety, ai-safety, tool | E4 / R3 (96%) | 3 |
| Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents Mingjie Tang, Tianze Hu, Zaiye Chen, Zhibo Liang Published: 2025-12-07Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-12-07 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (96%) | - |
| SoK: Trust-Authorization Mismatch in LLM Agent Interactions Guanquan Shi, Haohua Du, Song Bian, Weiwenpei Liu Published: 2025-12-07Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, survey | 2025-12-07 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 1 |
| Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective Anne Lauscher, Jae Hee Lee, Stefano V. Albrecht Published: 2025-12-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | 2025-12-04 | Agent Safety | agent-safety, ai-safety, alignment-training, interpretability, position, safety-evaluation | E5 / R3 (94%) | 1 |
| LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems Jie Zhang, Kun Wang, Li Sun, Sen Su Published: 2025-12-02Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2025-12-02 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (94%) | 1 |
| When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents Diogo Cruz, Mohammad Ali Jauhar, Nidhi Sakpal, Tsimur Hadeliya Published: 2025-12-02Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-12-02 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 2 |
| Systems Security Foundations for Agentic Computing Ashish Hooda, Earlence Fernandes, Johann Rehberger, Khawaja Shams Published: 2025-12-01Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, position | 2025-12-01 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (96%) | 2 |
| Password-Activated Shutdown Protocols for Misaligned Frontier Agents Francis Rhys Ward, Kai Williams, Rohan Subramani Published: 2025-11-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-29 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | - |
| Toward a Safe Internet of Agents George C. Polyzos, Juan A. Wibowo Published: 2025-11-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-11-29 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (97%) | - |
| Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains Arun Chowdary Sanna Published: 2025-11-25Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-25 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R2 (97%) | - |
| ASTRA: Agentic Steerability and Risk Assessment Framework Guy Shtar, Itay Hazan, Itsik Mantin, Ron Bitton Published: 2025-11-22Area: Agent SafetyCitations: - Tags: adversarial-robustness, agent-safety, ai-safety, benchmark | 2025-11-22 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, benchmark | E5 / R3 (96%) | - |
| Why Do Language Model Agents Whistleblow? Asa Cooper Stickland, Frank Xiao, Guido Bergman, Kushal Agrawal Published: 2025-11-21Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-11-21 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E6 / R4 (95%) | 1 |
| From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systems Afnan Shaik, Archana Vaidheeswaran, Brendan Gho, James Begin Published: 2025-11-18Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-18 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (95%) | - |
| Echoing: Identity Failures when LLM Agents Talk to Each Other Adam Earle, Romain Cosentino, Sarath Shekkizhar, Silvio Savarese Published: 2025-11-12Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-11-12 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E6 / R3 (97%) | 2 |
| Catching Contamination Before Generation: Spectral Kill Switches for Agents Valentin No毛l Published: 2025-11-08Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-08 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | - |
| ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations Ahmed Salem, Amr Gomaa, Sahar Abdelnabi Published: 2025-11-07Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2025-11-07 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (93%) | - |
| Evaluating Control Protocols for Untrusted AI Agents Buck Shlegeris, Chloe Loughridge, Henry Sleight, Joe Benton Published: 2025-11-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2025-11-04 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 1 |
| Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Chenghao Du, Quanfeng Huang, Tingxuan Tang, Yue Xiao Published: 2025-10-31Area: Agent SafetyCitations: - Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-10-31 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E4 / R3 (96%) | - |
| The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy Mohsen Bayati, William Overman Published: 2025-10-30Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, theoretical | 2025-10-30 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E5 / R3 (96%) | 1 |
| OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows Ben Kao, Fangzhi Xu, Kanzhi Cheng, Lingpeng Kong Published: 2025-10-28Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, benchmark | 2025-10-28 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (97%) | 4 |
| Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Nahin Anshuman Chhabra, Prasant Mohapatra, Shahriar Kabir, Shrestha Datta Published: 2025-10-27Area: Agent SafetyCitations: 7 Tags: agent-safety, ai-safety, safety-evaluation, survey | 2025-10-27 | Agent Safety | agent-safety, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 7 |
| Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems Harold Triedman, Justin Wagle, Rishi Jha, Vitaly Shmatikov Published: 2025-10-20Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (98%) | 2 |
| Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety Benjamin Plaut, Khanh Nguyen, Ponnurangam Kumaraguru, Vamshi Krishna Bonagiri Published: 2025-10-18Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-18 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 2 |
| Toward Understanding Security Issues in the Model Context Protocol Ecosystem Xiaofan Li, Xing Gao Published: 2025-10-18Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-18 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (97%) | 2 |
| From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails Andrea Bajcsy, Changliu Liu, Duy P. Nguyen, Jaime Fern谩ndez Fisac Published: 2025-10-15Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-10-15 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R3 (94%) | 1 |