Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Information Retrieval Induced Safety Degradation in AI Agents Benedikt Stroebl, Cheng Yu, Diyi Yang, Orestis Papakyriakopoulos Published: 2025-05-20Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-05-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (93%) | - |
| Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks Conghao Zhou, Dusit Niyato, Jiani Fan, Jiawen Kang Published: 2025-05-19Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, survey | 2025-05-19 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (93%) | 14 |
| A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron? Ada Chen, Jen-tse Huang, Jingyu Xiao, Junyuan Zhang Published: 2025-05-16Area: Agent SafetyCitations: 15 Tags: agent-safety, ai-safety, survey | 2025-05-16 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | 15 |
| Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction Changyue Jiang, Min Yang, Xudong Pan Published: 2025-05-16Area: Agent SafetyCitations: 6 Tags: agent-safety, ai-safety, empirical | 2025-05-16 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (94%) | 6 |
| Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents Christian Schroeder de Witt Published: 2025-05-04Area: Agent SafetyCitations: 38 Tags: agent-safety, ai-safety, survey | 2025-05-04 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 38 |
| Characterizing AI Agents for Alignment and Governance Atoosa Kasirzadeh, Iason Gabriel Published: 2025-04-30Area: Agent SafetyCitations: 28 Tags: agent-safety, ai-safety, alignment-training, theoretical | 2025-04-30 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E6 / R5 (95%) | 28 |
| ACE: A Security Architecture for LLM-Integrated App Systems Alina Oprea, Cristina Nita-Rotaru, Evan Li, Evan Rose Published: 2025-04-29Area: Agent SafetyCitations: 16 Tags: agent-safety, ai-safety, empirical | 2025-04-29 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 16 |
| MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools Benjamin Van Durme, Jason Eisner, Justin Svegliato, Nishant Subramani Published: 2025-04-28Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, empirical | 2025-04-28 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 4 |
| Guillotine: Hypervisors for Isolating Malicious AIs James Mickens, Ravi Netravali, Sarah Radway Published: 2025-04-22Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-04-22 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R4 (97%) | 1 |
| Bare Minimum Mitigations for Autonomous AI Development Chaochao Lu, Chris Cundy, Conor McGurk, Fynn Heide Published: 2025-04-21Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, position | 2025-04-21 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (94%) | 1 |
| RepliBench: Evaluating Autonomous Replication Capabilities of Language Model Agents Alan Cooney, Alex Remedios, Asa Cooper Stickland, Ben Millwood Published: 2025-04-21Area: Agent SafetyCitations: 8 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-04-21 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E4 / R3 (96%) | 8 |
| DoomArena: A framework for Testing AI Agents Against Evolving Security Threats Abhay Puri, Alexandre Drouin, Alexandre Lacoste, Avinandan Bose Published: 2025-04-18Area: Agent SafetyCitations: 19 Tags: agent-safety, ai-safety, safety-evaluation, tool | 2025-04-18 | Agent Safety | agent-safety, ai-safety, safety-evaluation, tool | E6 / R4 (97%) | 19 |
| Evaluating the Goal-Directedness of Large Language Models Alexis Bellot, Cristina G芒rbacea, Henry Papadatos, Jonathan Richens Published: 2025-04-16Area: Agent SafetyCitations: 5 Tags: agent-safety, ai-safety, empirical | 2025-04-16 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (96%) | 5 |
| Ctrl-Z: Controlling AI Agents via Resampling Adam Kaufman, Akbir Khan, Aryan Bhatt, Buck Shlegeris Published: 2025-04-14Area: Agent SafetyCitations: 14 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-04-14 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (93%) | 14 |
| How to evaluate control measures for LLM agents? A trajectory from today to superintelligence Buck Shlegeris, Geoffrey Irving, Mikita Balesni, Tomek Korbak Published: 2025-04-07Area: Agent SafetyCitations: 11 Tags: agent-safety, ai-safety, safety-evaluation, theoretical | 2025-04-07 | Agent Safety | agent-safety, ai-safety, safety-evaluation, theoretical | E5 / R3 (92%) | 11 |
| Emerging Cyber Attack Risks of Medical AI Agents Hao Wei, Jianing Qiu, Jiankai Sun, Kyle Lam Published: 2025-04-02Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, empirical | 2025-04-02 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 9 |
| Towards Trustworthy GUI Agents: A Survey Ninghao Liu, Wenhao Yu, Wenhu Chen, Wenlin Yao Published: 2025-03-30Area: Agent SafetyCitations: 19 Tags: agent-safety, ai-safety, safety-evaluation, survey | 2025-03-30 | Agent Safety | agent-safety, ai-safety, safety-evaluation, survey | E5 / R3 (93%) | 19 |
| Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents Atharv Singh Patlan, Peiyao Sheng Published: 2025-03-20Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, empirical | 2025-03-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (98%) | 14 |
| Multi-Agent Systems Execute Arbitrary Malicious Code Harold Triedman, Rishi Jha, Vitaly Shmatikov Published: 2025-03-15Area: Agent SafetyCitations: 23 Tags: agent-safety, ai-safety, empirical | 2025-03-15 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (98%) | 23 |
| Large language model-powered AI systems achieve self-replication with no human intervention Changyi Li, Jiarun Dai, Min Yang, Minyuan Luo Published: 2025-03-14Area: Agent SafetyCitations: 3 Tags: agent-safety, ai-safety, empirical | 2025-03-14 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R3 (96%) | 3 |
| AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Kamalika Chaudhuri Published: 2025-03-12Area: Agent SafetyCitations: 31 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-03-12 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E6 / R3 (97%) | 31 |
| A Survey on Trustworthy LLM Agents: Threats and Countermeasures Bo An, Fanci Meng, Junyuan Mao, Kun Wang Published: 2025-03-12Area: Agent SafetyCitations: 55 Tags: agent-safety, ai-safety, survey | 2025-03-12 | Agent Safety | agent-safety, ai-safety, survey | E5 / R4 (95%) | 55 |
| Safety Guardrails for LLM-Enabled Robots Alexander Robey, George J. Pappas, Hamed Hassani, Vijay Kumar Published: 2025-03-10Area: Agent SafetyCitations: 22 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-03-10 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (95%) | 22 |
| SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Safe Reinforcement Learning Borong Zhang, Jiaming Ji, Josef Dai, Yaodong Yang Published: 2025-03-05Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-03-05 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (98%) | 9 |
| AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection Chaowei Xiao, Huan Sun, Muhao Chen, Shenghong Dai Published: 2025-02-17Area: Agent SafetyCitations: 29 Tags: agent-safety, ai-safety, empirical | 2025-02-17 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R4 (95%) | 29 |
| RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage Ben L. Titzer, Heather Miller, McKenna McCall, Peter Yong Zhong Published: 2025-02-13Area: Agent SafetyCitations: 25 Tags: agent-safety, ai-safety, empirical | 2025-02-13 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (95%) | 25 |
| Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks Ang Li, Micah Goldblum, Tom Goldstein, Vethavikashini Chithrra Raghuram Published: 2025-02-12Area: Agent SafetyCitations: 38 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-02-12 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E6 / R3 (96%) | 38 |
| AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement Jakob Foerster, J Rosser Published: 2025-02-02Area: Agent SafetyCitations: 6 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-02-02 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R4 (94%) | 6 |
| A Sketch of an AI Control Safety Case Benjamin Hilton, Buck Shlegeris, Geoffrey Irving, Joshua Clymer Published: 2025-01-28Area: Agent SafetyCitations: 22 Tags: agent-safety, ai-safety, position, safety-evaluation | 2025-01-28 | Agent Safety | agent-safety, ai-safety, position, safety-evaluation | E6 / R4 (93%) | 22 |
| Contextual Agent Security: A Policy for Every Purpose Eugene Bagdasarian, Lillian Tsai Published: 2025-01-28Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, tool | 2025-01-28 | Agent Safety | agent-safety, ai-safety, tool | E5 / R3 (95%) | 14 |