Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions Bill Yuchen Lin, Faeze Brahman, Frank Xu, Hao Zhu Published: 2024-09-24Area: Agent SafetyCitations: 34 Tags: agent-safety, ai-safety, benchmark | 2024-09-24 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (95%) | 34 |
| Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Chenlu Zhan, Hanrong Zhang, Hongwei Wang, Jingyuan Huang Published: 2024-10-03Area: Agent SafetyCitations: 123 Tags: agent-safety, ai-safety, benchmark | 2024-10-03 | Agent Safety | agent-safety, ai-safety, benchmark | E7 / R4 (99%) | 123 |
| Permissive Information-Flow Analysis for Large Language Models Ahmed Salem, Andrew Paverd, Boris K枚pf, David Krueger Published: 2024-10-04Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, empirical | 2024-10-04 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 9 |
| Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents Elaine Chang, Elaine Lau, Matt Fredrikson, Priyanshu Kumar Published: 2024-10-11Area: Agent SafetyCitations: 54 Tags: agent-safety, ai-safety, benchmark | 2024-10-11 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (96%) | 54 |
| NetSafe: Exploring the Topological Safety of Multi-agent Networks Chenlong Yin, Guibin Zhang, Junyuan Mao, Kun Wang Published: 2024-10-21Area: Agent SafetyCitations: 26 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2024-10-21 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (94%) | 26 |
| MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control Dongyoon Hahm, June Suk Choi, Juyong Lee, Kimin Lee Published: 2024-10-23Area: Agent SafetyCitations: 26 Tags: agent-safety, ai-safety, benchmark | 2024-10-23 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (95%) | 26 |
| Agent-SafetyBench: Evaluating the Safety of LLM Agents Hongning Wang, Jingzhuo Zhou, Junxiao Yang, Minlie Huang Published: 2024-12-19Area: Agent SafetyCitations: 107 Tags: agent-safety, ai-safety, benchmark | 2024-12-19 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R2 (96%) | 107 |
| The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents Anna Squicciarini, Feiran Jia, Tong Wu, Xin Qin Published: 2024-12-21Area: Agent SafetyCitations: 26 Tags: agent-safety, ai-safety, alignment-training, empirical | 2024-12-21 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (96%) | 26 |
| Infrastructure for AI Agents Alan Chan, Elija Perrier, Gillian K. Hadfield, Kevin Wei Published: 2025-01-17Area: Agent SafetyCitations: 24 Tags: agent-safety, ai-safety, position | 2025-01-17 | Agent Safety | agent-safety, ai-safety, position | E6 / R4 (96%) | 24 |
| Episodic Memory in AI Agents Poses Risks that Should be Studied and Mitigated Chad DeChant Published: 2025-01-20Area: Agent SafetyCitations: 8 Tags: agent-safety, ai-safety, position | 2025-01-20 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (96%) | 8 |
| A Sketch of an AI Control Safety Case Benjamin Hilton, Buck Shlegeris, Geoffrey Irving, Joshua Clymer Published: 2025-01-28Area: Agent SafetyCitations: 22 Tags: agent-safety, ai-safety, position, safety-evaluation | 2025-01-28 | Agent Safety | agent-safety, ai-safety, position, safety-evaluation | E6 / R4 (93%) | 22 |
| Contextual Agent Security: A Policy for Every Purpose Eugene Bagdasarian, Lillian Tsai Published: 2025-01-28Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, tool | 2025-01-28 | Agent Safety | agent-safety, ai-safety, tool | E5 / R3 (95%) | 14 |
| AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement Jakob Foerster, J Rosser Published: 2025-02-02Area: Agent SafetyCitations: 6 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-02-02 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R4 (94%) | 6 |
| Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks Ang Li, Micah Goldblum, Tom Goldstein, Vethavikashini Chithrra Raghuram Published: 2025-02-12Area: Agent SafetyCitations: 38 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-02-12 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E6 / R3 (96%) | 38 |
| RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage Ben L. Titzer, Heather Miller, McKenna McCall, Peter Yong Zhong Published: 2025-02-13Area: Agent SafetyCitations: 25 Tags: agent-safety, ai-safety, empirical | 2025-02-13 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (95%) | 25 |
| AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection Chaowei Xiao, Huan Sun, Muhao Chen, Shenghong Dai Published: 2025-02-17Area: Agent SafetyCitations: 29 Tags: agent-safety, ai-safety, empirical | 2025-02-17 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R4 (95%) | 29 |
| SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Safe Reinforcement Learning Borong Zhang, Jiaming Ji, Josef Dai, Yaodong Yang Published: 2025-03-05Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-03-05 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (98%) | 9 |
| Safety Guardrails for LLM-Enabled Robots Alexander Robey, George J. Pappas, Hamed Hassani, Vijay Kumar Published: 2025-03-10Area: Agent SafetyCitations: 22 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-03-10 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (95%) | 22 |
| AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Kamalika Chaudhuri Published: 2025-03-12Area: Agent SafetyCitations: 31 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-03-12 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E6 / R3 (97%) | 31 |
| A Survey on Trustworthy LLM Agents: Threats and Countermeasures Bo An, Fanci Meng, Junyuan Mao, Kun Wang Published: 2025-03-12Area: Agent SafetyCitations: 55 Tags: agent-safety, ai-safety, survey | 2025-03-12 | Agent Safety | agent-safety, ai-safety, survey | E5 / R4 (95%) | 55 |
| Large language model-powered AI systems achieve self-replication with no human intervention Changyi Li, Jiarun Dai, Min Yang, Minyuan Luo Published: 2025-03-14Area: Agent SafetyCitations: 3 Tags: agent-safety, ai-safety, empirical | 2025-03-14 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R3 (96%) | 3 |
| Multi-Agent Systems Execute Arbitrary Malicious Code Harold Triedman, Rishi Jha, Vitaly Shmatikov Published: 2025-03-15Area: Agent SafetyCitations: 23 Tags: agent-safety, ai-safety, empirical | 2025-03-15 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (98%) | 23 |
| Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents Atharv Singh Patlan, Peiyao Sheng Published: 2025-03-20Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, empirical | 2025-03-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (98%) | 14 |
| Towards Trustworthy GUI Agents: A Survey Ninghao Liu, Wenhao Yu, Wenhu Chen, Wenlin Yao Published: 2025-03-30Area: Agent SafetyCitations: 19 Tags: agent-safety, ai-safety, safety-evaluation, survey | 2025-03-30 | Agent Safety | agent-safety, ai-safety, safety-evaluation, survey | E5 / R3 (93%) | 19 |
| Emerging Cyber Attack Risks of Medical AI Agents Hao Wei, Jianing Qiu, Jiankai Sun, Kyle Lam Published: 2025-04-02Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, empirical | 2025-04-02 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 9 |
| How to evaluate control measures for LLM agents? A trajectory from today to superintelligence Buck Shlegeris, Geoffrey Irving, Mikita Balesni, Tomek Korbak Published: 2025-04-07Area: Agent SafetyCitations: 11 Tags: agent-safety, ai-safety, safety-evaluation, theoretical | 2025-04-07 | Agent Safety | agent-safety, ai-safety, safety-evaluation, theoretical | E5 / R3 (92%) | 11 |
| Ctrl-Z: Controlling AI Agents via Resampling Adam Kaufman, Akbir Khan, Aryan Bhatt, Buck Shlegeris Published: 2025-04-14Area: Agent SafetyCitations: 14 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-04-14 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (93%) | 14 |
| Evaluating the Goal-Directedness of Large Language Models Alexis Bellot, Cristina G芒rbacea, Henry Papadatos, Jonathan Richens Published: 2025-04-16Area: Agent SafetyCitations: 5 Tags: agent-safety, ai-safety, empirical | 2025-04-16 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (96%) | 5 |
| DoomArena: A framework for Testing AI Agents Against Evolving Security Threats Abhay Puri, Alexandre Drouin, Alexandre Lacoste, Avinandan Bose Published: 2025-04-18Area: Agent SafetyCitations: 19 Tags: agent-safety, ai-safety, safety-evaluation, tool | 2025-04-18 | Agent Safety | agent-safety, ai-safety, safety-evaluation, tool | E6 / R4 (97%) | 19 |
| Bare Minimum Mitigations for Autonomous AI Development Chaochao Lu, Chris Cundy, Conor McGurk, Fynn Heide Published: 2025-04-21Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, position | 2025-04-21 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (94%) | 1 |