Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| RepliBench: Evaluating Autonomous Replication Capabilities of Language Model Agents Alan Cooney, Alex Remedios, Asa Cooper Stickland, Ben Millwood Published: 2025-04-21Area: Agent SafetyCitations: 8 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-04-21 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E4 / R3 (96%) | 8 |
| Guillotine: Hypervisors for Isolating Malicious AIs James Mickens, Ravi Netravali, Sarah Radway Published: 2025-04-22Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-04-22 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R4 (97%) | 1 |
| MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools Benjamin Van Durme, Jason Eisner, Justin Svegliato, Nishant Subramani Published: 2025-04-28Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, empirical | 2025-04-28 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 4 |
| ACE: A Security Architecture for LLM-Integrated App Systems Alina Oprea, Cristina Nita-Rotaru, Evan Li, Evan Rose Published: 2025-04-29Area: Agent SafetyCitations: 16 Tags: agent-safety, ai-safety, empirical | 2025-04-29 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 16 |
| Characterizing AI Agents for Alignment and Governance Atoosa Kasirzadeh, Iason Gabriel Published: 2025-04-30Area: Agent SafetyCitations: 28 Tags: agent-safety, ai-safety, alignment-training, theoretical | 2025-04-30 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E6 / R5 (95%) | 28 |
| Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents Christian Schroeder de Witt Published: 2025-05-04Area: Agent SafetyCitations: 38 Tags: agent-safety, ai-safety, survey | 2025-05-04 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 38 |
| A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron? Ada Chen, Jen-tse Huang, Jingyu Xiao, Junyuan Zhang Published: 2025-05-16Area: Agent SafetyCitations: 15 Tags: agent-safety, ai-safety, survey | 2025-05-16 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | 15 |
| Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction Changyue Jiang, Min Yang, Xudong Pan Published: 2025-05-16Area: Agent SafetyCitations: 6 Tags: agent-safety, ai-safety, empirical | 2025-05-16 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (94%) | 6 |
| Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks Conghao Zhou, Dusit Niyato, Jiani Fan, Jiawen Kang Published: 2025-05-19Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, survey | 2025-05-19 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (93%) | 14 |
| Information Retrieval Induced Safety Degradation in AI Agents Benedikt Stroebl, Cheng Yu, Diyi Yang, Orestis Papakyriakopoulos Published: 2025-05-20Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-05-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (93%) | - |
| Automating Safety Enhancement for LLM-based Agents with Synthetic Risk Scenarios Guiyao Tie, Jiawen Shi, Lichao Sun, Lin Lu Published: 2025-05-23Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, empirical | 2025-05-23 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 4 |
| Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making Chanyoung Park, Dongju Jang, Jian Kim, Minseo Kim Published: 2025-05-26Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, benchmark | 2025-05-26 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (97%) | 4 |
| AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery Haowei Wang, Junjie Wang, Mingyang Li, Qing Wang Published: 2025-05-27Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, empirical | 2025-05-27 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 4 |
| RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments Eric Fosler-Lussier, Huan Sun, Jaylen Jones, Linxi Jiang Published: 2025-05-28Area: Agent SafetyCitations: 12 Tags: adversarial-robustness, agent-safety, ai-safety, benchmark | 2025-05-28 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, benchmark | E6 / R3 (95%) | 12 |
| Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems Chuan Xiao, Jiseong Jeong, Makoto Onizuka, Ronny Ko Published: 2025-05-28Area: Agent SafetyCitations: 10 Tags: agent-safety, ai-safety, position, safety-evaluation | 2025-05-28 | Agent Safety | agent-safety, ai-safety, position, safety-evaluation | E7 / R3 (96%) | 10 |
| AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models Jinchuan Zhang, Lu Yin, Songlin Hu, Yan Zhou Published: 2025-05-29Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-05-29 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (96%) | 4 |
| LLM Agents Should Employ Security Principles Elisa Bertino, Kaiyuan Zhang, Ninghui Li, Pin-Yu Chen Published: 2025-05-29Area: Agent SafetyCitations: 13 Tags: agent-safety, ai-safety, position | 2025-05-29 | Agent Safety | agent-safety, ai-safety, position | E6 / R4 (94%) | 13 |
| MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment John T. Halloran Published: 2025-05-29Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-05-29 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (97%) | 2 |
| SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents Haofei Yu, Jiaxuan You, Jiaxun Zhang, Kunlun Zhu Published: 2025-05-29Area: Agent SafetyCitations: 9 Tags: agent-safety, ai-safety, empirical | 2025-05-29 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (97%) | 9 |
| RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents Dongrui Liu, Jing Shao, Jingyi Yang, Shuai Shao Published: 2025-05-31Area: Agent SafetyCitations: 11 Tags: agent-safety, ai-safety, benchmark | 2025-05-31 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (97%) | 11 |
| Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS Charu C. Aggarwal, Han Xu, Hui Liu, Juanhui Li Published: 2025-06-02Area: Agent SafetyCitations: 5 Tags: agent-safety, ai-safety, survey | 2025-06-02 | Agent Safety | agent-safety, ai-safety, survey | E4 / R3 (95%) | 5 |
| MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Hang Su, Jiawei Chen, Jun Luo, Jun Zhu Published: 2025-06-02Area: Agent SafetyCitations: 17 Tags: agent-safety, ai-safety, benchmark | 2025-06-02 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (96%) | 17 |
| Attention Knows Whom to Trust: Attention-based Trust Management for LLM Multi-Agent Systems Hui Liu, Jiliang Tang, Jingying Zeng, Pengfei He Published: 2025-06-03Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, empirical | 2025-06-03 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 4 |
| Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Ram Potham Published: 2025-06-03Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, benchmark | 2025-06-03 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (94%) | 1 |
| AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents Akshat Naik, Edward James Young, Emma Goun茅, Francisco Javier Campos Zabala Published: 2025-06-04Area: Agent SafetyCitations: 12 Tags: agent-safety, ai-safety, alignment-training, benchmark | 2025-06-04 | Agent Safety | agent-safety, ai-safety, alignment-training, benchmark | E5 / R3 (94%) | 12 |
| Mind the Web: The Security of Web Use Agents Asaf Shabtai, Avishag Shapira, Edan Habler, Parth Atulbhai Gandhi Published: 2025-06-08Area: Agent SafetyCitations: 3 Tags: agent-safety, ai-safety, empirical | 2025-06-08 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (94%) | 3 |
| Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Ahmed Amer, Evan Harris, Nell Watson, Preeti Ravindra Published: 2025-06-08Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-06-08 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (93%) | 1 |
| TAI3: Testing Agent Integrity in Interpreting User Intent Kaiyuan Zhang, Mingwei Zheng, Shiwei Feng, Syed Yusuf Ahmed Published: 2025-06-09Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, benchmark | 2025-06-09 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (94%) | 2 |
| Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety Chanwoo Park, Cynthia Breazeal, Daniel McDuff, Eugene Park Published: 2025-06-14Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-06-14 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R4 (94%) | - |
| We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems Haokai Ma, Junfeng Fang, Ruipeng Wang, Tat-Seng Chua Published: 2025-06-16Area: Agent SafetyCitations: 23 Tags: agent-safety, ai-safety, empirical | 2025-06-16 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 23 |