Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 194 papers (page 1 of 7)路 67 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Avoiding Tampering Incentives in Deep RL via Decoupled Approval Jonathan Uesato, Ramana Kumar, Richard Ngo, Shane Legg Published: 2020-11-17Area: Agent SafetyCitations: 18 Tags: agent-safety, ai-safety, empirical | 2020-11-17 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (94%) | 18 |
| Open Problems in Cooperative AI Allan Dafoe, Edward Hughes, Joel Z. Leibo, Kate Larson Published: 2020-12-15Area: Agent SafetyCitations: 239 Tags: agent-safety, ai-safety, position | 2020-12-15 | Agent Safety | agent-safety, ai-safety, position | E6 / R5 (93%) | 239 |
| What Would Jiminy Cricket Do? Towards Agents That Behave Morally Andy Zou, Bo Li, Christine Zhu, Dan Hendrycks Published: 2021-10-25Area: Agent SafetyCitations: 72 Tags: agent-safety, ai-safety, benchmark | 2021-10-25 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (93%) | 72 |
| Natural Selection Favors AIs over Humans Dan Hendrycks Published: 2023-03-28Area: Agent SafetyCitations: 40 Tags: agent-safety, ai-safety, position | 2023-03-28 | Agent Safety | agent-safety, ai-safety, position | E4 / R3 (93%) | 40 |
| Identifying the Risks of LM Agents with an LM-Emulated Sandbox Andrew Wang, Chris J. Maddison, Honghua Dong, Jimmy Ba Published: 2023-09-25Area: Agent SafetyCitations: 217 Tags: agent-safety, ai-safety, tool | 2023-09-25 | Agent Safety | agent-safety, ai-safety, tool | E5 / R3 (95%) | 217 |
| AI Systems of Concern Fazl Barez, Kayla Matteucci, Se谩n 脫 h脡igeartaigh, Shahar Avin Published: 2023-10-09Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, position | 2023-10-09 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (95%) | 1 |
| Testing Language Model Agents Safely in the Wild Adam Tauman Kalai, Craig Swift, David Atkinson, David Bau Published: 2023-11-17Area: Agent SafetyCitations: 40 Tags: agent-safety, ai-safety, empirical | 2023-11-17 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R2 (95%) | 40 |
| Evil Geniuses: Delving into the Safety of LLM-based Agents Hang Su, Jingyuan Zhang, Xiao Yang, Yinpeng Dong Published: 2023-11-20Area: Agent SafetyCitations: 101 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2023-11-20 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E7 / R4 (95%) | 101 |
| Agent Alignment in Evolving Social Norms Qinyuan Cheng, Shimin Li, Tianxiang Sun, Xipeng Qiu Published: 2024-01-09Area: Agent SafetyCitations: 12 Tags: agent-safety, ai-safety, alignment-training, empirical | 2024-01-09 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (96%) | 12 |
| R-Judge: Benchmarking Safety Risk Awareness for LLM Agents Binglin Zhou, Fangqi Li, Gongshen Liu, Lingzhong Dong Published: 2024-01-18Area: Agent SafetyCitations: 156 Tags: agent-safety, ai-safety, benchmark | 2024-01-18 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (96%) | 156 |
| PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety Feng Zhao, Hongzhi Gao, Huchuan Lu, Jing Shao Published: 2024-01-22Area: Agent SafetyCitations: 75 Tags: agent-safety, ai-safety, empirical, safety-evaluation | 2024-01-22 | Agent Safety | agent-safety, ai-safety, empirical, safety-evaluation | E6 / R4 (95%) | 75 |
| Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science Arman Cohan, Jian Tang, Kunlun Zhu, Mark Gerstein Published: 2024-02-06Area: Agent SafetyCitations: 55 Tags: agent-safety, ai-safety, alignment-training, survey | 2024-02-06 | Agent Safety | agent-safety, ai-safety, alignment-training, survey | E5 / R3 (95%) | 55 |
| The Reasons that Agents Act: Intention and Instrumental Goals Francesca Toni, Francesco Belardinelli, Francis Rhys Ward, Matt MacDermott Published: 2024-02-11Area: Agent SafetyCitations: 22 Tags: agent-safety, ai-safety, theoretical | 2024-02-11 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R3 (96%) | 22 |
| Secret Collusion Among Generative AI Agents Christian Schroeder de Witt, Lewis Hammond, Martin Strohmeier, Mikhail Baranchuk Published: 2024-02-12Area: Agent SafetyCitations: 59 Tags: agent-safety, ai-safety, empirical, safety-evaluation | 2024-02-12 | Agent Safety | agent-safety, ai-safety, empirical, safety-evaluation | E5 / R3 (97%) | 59 |
| A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents Boyuan Zheng, Chaowei Xiao, Huan Sun, Lingbo Mo Published: 2024-02-15Area: Agent SafetyCitations: 23 Tags: adversarial-robustness, agent-safety, ai-safety, survey | 2024-02-15 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, survey | E6 / R4 (97%) | 23 |
| InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents Daniel Kang, Qiusi Zhan, Zhixiang Liang, Zifan Ying Published: 2024-03-05Area: Agent SafetyCitations: 251 Tags: agent-safety, ai-safety, benchmark | 2024-03-05 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (96%) | 251 |
| LLM Agents can Autonomously Exploit One-day Vulnerabilities Akul Gupta, Daniel Kang, Richard Fang, Rohan Bindu Published: 2024-04-11Area: Agent SafetyCitations: 125 Tags: agent-safety, ai-safety, empirical | 2024-04-11 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 125 |
| AirGapAgent: Protecting Privacy-Conscious Conversational Agents Borja Balle, Daniel Ramage, Eugene Bagdasaryan, Marco Gruteser Published: 2024-05-08Area: Agent SafetyCitations: 53 Tags: agent-safety, ai-safety, empirical | 2024-05-08 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (93%) | 53 |
| AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways Changzhou Han, Junwu Xiong, Sheng Wen, Wanlun Ma Published: 2024-06-04Area: Agent SafetyCitations: 152 Tags: agent-safety, ai-safety, survey | 2024-06-04 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 152 |
| Security of AI Agents Ethan Wang, Hao Chen, Yifeng He, Yuyang Rong Published: 2024-06-12Area: Agent SafetyCitations: 21 Tags: agent-safety, ai-safety, empirical | 2024-06-12 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 21 |
| GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning Bo Li, Carl Yang, Chulin Xie, Dawn Song Published: 2024-06-13Area: Agent SafetyCitations: 69 Tags: agent-safety, ai-safety, empirical | 2024-06-13 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (97%) | 69 |
| AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents Edoardo Debenedetti, Florian Tramer, Jie Zhang, Luca Beurer-Kellner Published: 2024-06-19Area: Agent SafetyCitations: 94 Tags: agent-safety, ai-safety, benchmark | 2024-06-19 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (98%) | 94 |
| Towards shutdownable agents via stochastic choice Alexander Roman, Christos Ziakas, Elliott Thornley, Leyton Ho Published: 2024-06-30Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2024-06-30 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 1 |
| Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities Gongshen Liu, Haodong Zhao, Jian Xie, Lifeng Liu Published: 2024-07-10Area: Agent SafetyCitations: 63 Tags: agent-safety, ai-safety, empirical | 2024-07-10 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (94%) | 63 |
| Security Matrix for Multimodal Agents on Mobile Devices: A Systematic and Proof of Concept Study Chao Shen, Chenhao Lin, Shuaidong Li, Tianwei Zhang Published: 2024-07-12Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2024-07-12 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (97%) | 2 |
| Preemptive Detection and Correction of Misaligned Actions in LLM Agents Haishuo Fang, Iryna Gurevych, Xiaodan Zhu Published: 2024-07-16Area: Agent SafetyCitations: 6 Tags: agent-safety, ai-safety, empirical | 2024-07-16 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R5 (95%) | 6 |
| Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification Ahmed Salem, Boyang Zhang, Michael Backes, Savvas Zannettou Published: 2024-07-30Area: Agent SafetyCitations: 65 Tags: agent-safety, ai-safety, empirical | 2024-07-30 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (93%) | 65 |
| Caution for the Environment: Multimodal Agents are Susceptible to Environmental Distractions Aston Zhang, Hai Zhao, Tongxin Yuan, Xinbei Ma Published: 2024-08-05Area: Agent SafetyCitations: 47 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2024-08-05 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E6 / R3 (95%) | 47 |
| Operationalizing Contextual Integrity in Privacy-Conscious Assistants Aneesh Pappu, Borja Balle, Chongyang Shi, Eugene Bagdasaryan Published: 2024-08-05Area: Agent SafetyCitations: 28 Tags: agent-safety, ai-safety, empirical | 2024-08-05 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (95%) | 28 |
| RiskAwareBench: Towards Evaluating Physical Risk Awareness for High-level Planning of LLM-based Embodied Agents Baoyuan Wu, Bingzhe Wu, Zhengyou Zhang, Zihao Zhu Published: 2024-08-08Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, benchmark | 2024-08-08 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (95%) | 14 |