Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 186 papers (page 1 of 7)路 37 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Constructing Safety Cases for AI Systems: A Reusable Template Framework Jieshan Chen, Liming Dong, Liming Zhu, Md Shamsujjoha Published: 2026-01-30Area: Safety EvaluationCitations: - Tags: ai-safety, safety-evaluation, survey | 2026-01-30 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (97%) | - |
| How Should AI Safety Benchmarks Benchmark Safety? Cheng Yu, Dalia Ali, Orestis Papakyriakopoulos, Ruoxuan Cao Published: 2026-01-30Area: Safety EvaluationCitations: - Tags: ai-safety, safety-evaluation, survey | 2026-01-30 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (94%) | - |
| Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses Hozefa Lakadawala, Mohsen Hatami, Van Tuan Pham, Yu Chen Published: 2026-01-28Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2026-01-28 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | - |
| Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems Dmitry Namiot, Narek Maloyan Published: 2026-01-24Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, survey | 2026-01-24 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R3 (97%) | - |
| Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions Usman Naseem Published: 2026-01-21Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2026-01-21 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E6 / R4 (96%) | 1 |
| Unlearning in LLMs: Methods, Evaluation, and Open Challenges Larry Heck, Tyler Lizzo Published: 2026-01-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2026-01-19 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E7 / R3 (96%) | - |
| The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism Ben Nassi, Bruce Schneier, Oleg Brodt Published: 2026-01-14Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, survey | 2026-01-14 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R4 (99%) | - |
| Interpreting Transformers Through Attention Head Intervention Mason Kadem, Rong Zheng Published: 2026-01-07Area: Mechanistic Interp.Citations: - Tags: ai-safety, interpretability, mechanistic-interp, survey | 2026-01-07 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E5 / R3 (95%) | - |
| Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defenses Chao Li, Chaozhuo Li, Litian Zhang, Xi Zhang Published: 2026-01-07Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2026-01-07 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 1 |
| Trust in LLM-controlled Robotics: a Survey of Security Threats, Defenses and Challenges Huaming Chen, Ian R. Manchester, Kim-Kwang Raymond Choo, Mitch Bryson Published: 2025-12-17Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-12-17 | Agent Safety | agent-safety, ai-safety, survey | E7 / R4 (93%) | - |
| SoK: Trust-Authorization Mismatch in LLM Agent Interactions Guanquan Shi, Haohua Du, Song Bian, Weiwenpei Liu Published: 2025-12-07Area: Agent SafetyCitations: 1 Tags: agent-safety, ai-safety, survey | 2025-12-07 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 1 |
| LLM Harms: A Taxonomy and Discussion Abhejay Murali, Amit Dhurandhar, David Atkinson, Junfeng Jiao Published: 2025-12-05Area: Surveys & ReviewsCitations: - Tags: ai-safety, survey, surveys-reviews | 2025-12-05 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | - |
| Toward a Safe Internet of Agents George C. Polyzos, Juan A. Wibowo Published: 2025-11-29Area: Agent SafetyCitations: - Tags: agent-safety, ai-safety, survey | 2025-11-29 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (97%) | - |
| Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Bianka Kowalska, Halina Kwa艣nicka Published: 2025-11-24Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-11-24 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 1 |
| Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks Daoyuan Wu, Pingchuan Ma, Shuai Wang, Tian Tian Published: 2025-11-19Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-11-19 | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E6 / R3 (95%) | - |
| A Survey on Unlearning in Large Language Models Fei Sun, Honglin Wang, Jiajun Tan, Jiayue Pu Published: 2025-10-29Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-10-29 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Nahin Anshuman Chhabra, Prasant Mohapatra, Shahriar Kabir, Shrestha Datta Published: 2025-10-27Area: Agent SafetyCitations: 7 Tags: agent-safety, ai-safety, safety-evaluation, survey | 2025-10-27 | Agent Safety | agent-safety, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 7 |
| SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models Ali Arastehfard, Biying Liu, Hanbin Hong, Heqing Huang Published: 2025-10-17Area: Adversarial RobustnessCitations: 1 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-10-17 | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R4 (97%) | 1 |
| Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity Adji Bousso Dieng, Alex John London, Alvaro Velasquez, Amrit Singh Bedi Published: 2025-10-13Area: Safety EvaluationCitations: 2 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-10-13 | Safety Evaluation | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R3 (93%) | 2 |
| Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs Anh Tuan Luu, Jinbo Feng, Linghui Meng, Man Hu Published: 2025-10-09Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, survey | 2025-10-09 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (95%) | - |
| A Survey on Agentic Security: Applications, Threats and Defenses Asif Shahriar, Farig Sadeque, Md Nafiu Rahman, Md Rizwan Parvez Published: 2025-10-07Area: Agent SafetyCitations: 6 Tags: adversarial-robustness, agent-safety, ai-safety, survey | 2025-10-07 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, survey | E6 / R4 (96%) | 6 |
| LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Briland Hitaj, Gabriel Antonio Fontes Rebello, Igor Jochem Sanz, Rodrigo Duarte de Meneses Published: 2025-09-12Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-09-12 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| A Review of Developmental Interpretability in Large Language Models Ihor Kendiukhov Published: 2025-08-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-08-19 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (94%) | - |
| Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Changjia Zhu, Chi Zhang, Junjie Xiong, Lingyao Li Published: 2025-08-07Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, survey, surveys-reviews | 2025-08-07 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (93%) | 5 |
| A Survey on Data Security in Large Language Models Fan Lin, Jinhe Su, Kang Chen, Li Shen Published: 2025-08-04Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-08-04 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Chaochao Chen, Chengye Wang, Fengyuan Yu, Jiaming Zhang Published: 2025-07-26Area: Surveys & ReviewsCitations: 2 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-07-26 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 2 |
| A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Chang Liu, Hang Su, Jun Luo, Jun Zhu Published: 2025-06-30Area: Agent SafetyCitations: 12 Tags: agent-safety, ai-safety, survey | 2025-06-30 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | 12 |
| A Survey on Model Extraction Attacks and Defenses for Large Language Models Kaixiang Zhao, Kaize Ding, Lincan Li, Neil Zhenqiang Gong Published: 2025-06-26Area: Adversarial RobustnessCitations: 11 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-06-26 | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 11 |
| A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures Changting Lin, Chaochao Chen, Dezhang Kong, Hujin Peng Published: 2025-06-24Area: Agent SafetyCitations: 35 Tags: agent-safety, ai-safety, survey | 2025-06-24 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (97%) | 35 |
| Report on NSF Workshop on Science of Safe AI Corina P膬s膬reanu, Greg Durrett, Hadas Kress-Gazit, Rajeev Alur Published: 2025-06-24Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-06-24 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (97%) | 1 |