Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents Chao Shang, Devang Kulshreshtha, Hang Su, Jianfeng He Published: 2025-09-30Area: Agent SafetyCitations: 2 Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-09-30 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E5 / R3 (94%) | 2 |
| Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness Besmira Nushi, Erfan Shayegani, Keegan Hines, Nael Abu-Ghazaleh Published: 2025-10-02Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, benchmark | 2025-10-02 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (98%) | 1 |
| ToolTweak: An Attack on Tool Selection in LLM-based Agents Adel Bibi, Alasdair Paren, Eric Sommerlade, Jialin Yu Published: 2025-10-02Area: Agent SafetyCitations: 5 Tags: agent-safety, ai-safety, empirical | 2025-10-02 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | 5 |
| Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain Abhay Puri, Alexandre Drouin, Alexandre Lacoste, Chandra Kiran Reddy Evuru Published: 2025-10-03Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2025-10-03 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (95%) | 1 |
| Agentic Misalignment: How LLMs Could Be Insider Threats Aengus Lynch, Benjamin Wright, Caleb Larson, Ethan Perez Published: 2025-10-05Area: Agent SafetyCitations: 55 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-10-05 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E5 / R3 (95%) | 55 |
| Adapting Insider Risk mitigations for Agentic Misalignment: an empirical study Francesca Gomez Published: 2025-10-06Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-10-06 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E6 / R3 (96%) | - |
| A Survey on Agentic Security: Applications, Threats and Defenses Asif Shahriar, Farig Sadeque, Md Nafiu Rahman, Md Rizwan Parvez Published: 2025-10-07Area: Agent SafetyCitations: 6 Tags: adversarial-robustness, agent-safety, ai-safety, survey | 2025-10-07 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, survey | E6 / R4 (96%) | 6 |
| MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents Dongsen Zhang, Peipei Li, Wenjun Xu, Xuannan Liu Published: 2025-10-14Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, benchmark | 2025-10-14 | Agent Safety | agent-safety, ai-safety, benchmark | E4 / R3 (98%) | 2 |
| From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails Andrea Bajcsy, Changliu Liu, Duy P. Nguyen, Jaime Fern谩ndez Fisac Published: 2025-10-15Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, theoretical | 2025-10-15 | Agent Safety | agent-safety, ai-safety, theoretical | E5 / R3 (94%) | 1 |
| Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment Boxuan Wang, Guangliang Cheng, Jinwei Hu, Lokesh Singh Published: 2025-10-15Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, alignment-training, position | 2025-10-15 | Agent Safety | agent-safety, ai-safety, alignment-training, position | E4 / R3 (92%) | 4 |
| When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? Beidi Chen, Corina Pasareanu, Haizhong Zheng, James Song Published: 2025-10-15Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-10-15 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (97%) | - |
| Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety Benjamin Plaut, Khanh Nguyen, Ponnurangam Kumaraguru, Vamshi Krishna Bonagiri Published: 2025-10-18Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-18 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (96%) | 2 |
| Toward Understanding Security Issues in the Model Context Protocol Ecosystem Xiaofan Li, Xing Gao Published: 2025-10-18Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-18 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R4 (97%) | 2 |
| Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems Harold Triedman, Justin Wagle, Rishi Jha, Vitaly Shmatikov Published: 2025-10-20Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, empirical | 2025-10-20 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (98%) | 2 |
| Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Nahin Anshuman Chhabra, Prasant Mohapatra, Shahriar Kabir, Shrestha Datta Published: 2025-10-27Area: Agent SafetyCitations: 7 Tags: agent-safety, ai-safety, safety-evaluation, survey | 2025-10-27 | Agent Safety | agent-safety, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 7 |
| OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows Ben Kao, Fangzhi Xu, Kanzhi Cheng, Lingpeng Kong Published: 2025-10-28Area: Agent SafetyCitations: 4 Tags: agent-safety, ai-safety, benchmark | 2025-10-28 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (97%) | 4 |
| The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy Mohsen Bayati, William Overman Published: 2025-10-30Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, alignment-training, theoretical | 2025-10-30 | Agent Safety | agent-safety, ai-safety, alignment-training, theoretical | E5 / R3 (96%) | 1 |
| Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Chenghao Du, Quanfeng Huang, Tingxuan Tang, Yue Xiao Published: 2025-10-31Area: Agent SafetyCitations: - Tags: adversarial-robustness, agent-safety, ai-safety, empirical | 2025-10-31 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, empirical | E4 / R3 (96%) | - |
| Evaluating Control Protocols for Untrusted AI Agents Buck Shlegeris, Chloe Loughridge, Henry Sleight, Joe Benton Published: 2025-11-04Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2025-11-04 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (94%) | 1 |
| ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations Ahmed Salem, Amr Gomaa, Sahar Abdelnabi Published: 2025-11-07Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, benchmark | 2025-11-07 | Agent Safety | agent-safety, ai-safety, benchmark | E5 / R3 (93%) | - |
| Catching Contamination Before Generation: Spectral Kill Switches for Agents Valentin No毛l Published: 2025-11-08Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-08 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | - |
| Echoing: Identity Failures when LLM Agents Talk to Each Other Adam Earle, Romain Cosentino, Sarath Shekkizhar, Silvio Savarese Published: 2025-11-12Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, alignment-training, empirical | 2025-11-12 | Agent Safety | agent-safety, ai-safety, alignment-training, empirical | E6 / R3 (97%) | 2 |
| From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systems Afnan Shaik, Archana Vaidheeswaran, Brendan Gho, James Begin Published: 2025-11-18Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-18 | Agent Safety | agent-safety, ai-safety, empirical | E6 / R3 (95%) | - |
| Why Do Language Model Agents Whistleblow? Asa Cooper Stickland, Frank Xiao, Guido Bergman, Kushal Agrawal Published: 2025-11-21Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, benchmark, safety-evaluation | 2025-11-21 | Agent Safety | agent-safety, ai-safety, benchmark, safety-evaluation | E6 / R4 (95%) | 1 |
| ASTRA: Agentic Steerability and Risk Assessment Framework Guy Shtar, Itay Hazan, Itsik Mantin, Ron Bitton Published: 2025-11-22Area: Agent SafetyCitations: - Tags: adversarial-robustness, agent-safety, ai-safety, benchmark | 2025-11-22 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, benchmark | E5 / R3 (96%) | - |
| Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains Arun Chowdary Sanna Published: 2025-11-25Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-25 | Agent Safety | agent-safety, ai-safety, empirical | E7 / R2 (97%) | - |
| Password-Activated Shutdown Protocols for Misaligned Frontier Agents Francis Rhys Ward, Kai Williams, Rohan Subramani Published: 2025-11-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, empirical | 2025-11-29 | Agent Safety | agent-safety, ai-safety, empirical | E5 / R3 (95%) | - |
| Toward a Safe Internet of Agents George C. Polyzos, Juan A. Wibowo Published: 2025-11-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-11-29 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (97%) | - |
| Systems Security Foundations for Agentic Computing Ashish Hooda, Earlence Fernandes, Johann Rehberger, Khawaja Shams Published: 2025-12-01Area: Agent SafetyCitations: 2 Tags: agent-safety, ai-safety, position | 2025-12-01 | Agent Safety | agent-safety, ai-safety, position | E5 / R3 (96%) | 2 |
| LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems Jie Zhang, Kun Wang, Li Sun, Sen Su Published: 2025-12-02Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, empirical | 2025-12-02 | Agent Safety | agent-safety, ai-safety, empirical | E4 / R3 (94%) | 1 |