Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures Fernando Berzal, Francisco Aguilera-Mart铆nez Published: 2025-05-02Area: Adversarial RobustnessCitations: 13 Tags: adversarial-robustness, ai-safety, survey | 2025-05-02 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (95%) | 13 |
| Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents Christian Schroeder de Witt Published: 2025-05-04Area: Agent SafetyCitations: 38 Tags: agent-safety, ai-safety, survey | 2025-05-04 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (94%) | 38 |
| A Survey on Progress in LLM Alignment from the Perspective of Reward Design Jian Yang, Mark Dras, Miaomiao Ji, Shoujin Wang Published: 2025-05-05Area: Surveys & ReviewsCitations: 10 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-05-05 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (95%) | 10 |
| Adversarial Attacks in Multimodal Systems: A Practitioner's Survey Aman Raj, Ankit Shetgaonkar, Dipen Pradhan, Lakshit Arora Published: 2025-05-06Area: Multimodal SafetyCitations: 2 Tags: adversarial-robustness, ai-safety, multimodal-safety, survey | 2025-05-06 | Multimodal Safety | adversarial-robustness, ai-safety, multimodal-safety, survey | E6 / R3 (94%) | 2 |
| A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron? Ada Chen, Jen-tse Huang, Jingyu Xiao, Junyuan Zhang Published: 2025-05-16Area: Agent SafetyCitations: 15 Tags: agent-safety, ai-safety, survey | 2025-05-16 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | 15 |
| A Survey of Attacks on Large Language Models Keshab K. Parhi, Wenrui Xu Published: 2025-05-18Area: Adversarial RobustnessCitations: 10 Tags: adversarial-robustness, ai-safety, survey | 2025-05-18 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R3 (95%) | 10 |
| Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks Conghao Zhou, Dusit Niyato, Jiani Fan, Jiawen Kang Published: 2025-05-19Area: Agent SafetyCitations: 14 Tags: agent-safety, ai-safety, survey | 2025-05-19 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (93%) | 14 |
| Security Concerns for Large Language Models: A Survey Benjamin C. M. Fung, Miles Q. Li Published: 2025-05-24Area: Surveys & ReviewsCitations: 29 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-05-24 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 29 |
| Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression Ping Liu, Yiwei Xie, Zheng Zhang Published: 2025-05-26Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, survey, surveys-reviews | 2025-05-26 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| The Multilingual Divide and Its Impact on Global AI Safety Aakanksha, Ahmet 脺st眉n, Aidan Peppin, Alice Schoenauer Sebag Published: 2025-05-27Area: Safety EvaluationCitations: 4 Tags: ai-safety, safety-evaluation, survey | 2025-05-27 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (94%) | 4 |
| Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies Chenruo Liu, Kenan Tang, Qi Lei, Yao Qin Published: 2025-05-28Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-05-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (95%) | 1 |
| The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Beyza Ermis, Julia Kreutzer, Marzieh Fadaee, Stephen H. Bach Published: 2025-05-30Area: Surveys & ReviewsCitations: 10 Tags: ai-safety, survey, surveys-reviews | 2025-05-30 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | 10 |
| Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS Charu C. Aggarwal, Han Xu, Hui Liu, Juanhui Li Published: 2025-06-02Area: Agent SafetyCitations: 5 Tags: agent-safety, ai-safety, survey | 2025-06-02 | Agent Safety | agent-safety, ai-safety, survey | E4 / R3 (95%) | 5 |
| Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aeree Cho, Duen Horng Chau, Grace C. Kim, Mansi Phute Published: 2025-06-05Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, survey, surveys-reviews | 2025-06-05 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 5 |
| A Systematic Review of Poisoning Attacks Against Large Language Models Edward W. Staley, Joshua Carney, Marie Chau, Nathan Drenkow Published: 2025-06-06Area: Adversarial RobustnessCitations: 6 Tags: adversarial-robustness, ai-safety, survey | 2025-06-06 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E4 / R3 (95%) | 6 |
| The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Chaozhuo Li, Feiran Huang, Jiameng Qiu, Litian Zhang Published: 2025-06-06Area: Safety EvaluationCitations: 23 Tags: ai-safety, safety-evaluation, survey | 2025-06-06 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (94%) | 23 |
| Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions Dacao Zhang, Guangyi Lv, Kui Yu, Kun Zhang Published: 2025-06-08Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2025-06-08 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| SoK: Machine Unlearning for Large Language Models Charu C. Aggarwal, Hui Liu, Jie Ren, Yingqian Cui Published: 2025-06-10Area: Model EditingCitations: 5 Tags: ai-safety, model-editing, safety-evaluation, survey | 2025-06-10 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E6 / R3 (96%) | 5 |
| Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Chuan Qin, Feimin Zhong, Han Wu, Hengshu Zhu Published: 2025-06-11Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-06-11 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (93%) | 5 |
| SoK: Evaluating Jailbreak Guardrails for Large Language Models Daoyuan Wu, Shuai Wang, Wenxuan Wang, Xunguang Wang Published: 2025-06-12Area: Adversarial RobustnessCitations: 6 Tags: adversarial-robustness, ai-safety, survey | 2025-06-12 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E4 / R2 (98%) | 6 |
| From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem Datao You, Hongsong Zhu, Peipei Liu, Tiehan Cui Published: 2025-06-18Area: Adversarial RobustnessCitations: 7 Tags: adversarial-robustness, ai-safety, survey | 2025-06-18 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R4 (98%) | 7 |
| AI Safety vs. AI Security: Demystifying the Distinction and Boundaries Huan Sun, Ness Shroff, Zhiqiang Lin Published: 2025-06-21Area: Surveys & ReviewsCitations: 2 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-06-21 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R4 (95%) | 2 |
| A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures Changting Lin, Chaochao Chen, Dezhang Kong, Hujin Peng Published: 2025-06-24Area: Agent SafetyCitations: 35 Tags: agent-safety, ai-safety, survey | 2025-06-24 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (97%) | 35 |
| Report on NSF Workshop on Science of Safe AI Corina P膬s膬reanu, Greg Durrett, Hadas Kress-Gazit, Rajeev Alur Published: 2025-06-24Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-06-24 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (97%) | 1 |
| A Survey on Model Extraction Attacks and Defenses for Large Language Models Kaixiang Zhao, Kaize Ding, Lincan Li, Neil Zhenqiang Gong Published: 2025-06-26Area: Adversarial RobustnessCitations: 11 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-06-26 | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 11 |
| A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Chang Liu, Hang Su, Jun Luo, Jun Zhu Published: 2025-06-30Area: Agent SafetyCitations: 12 Tags: agent-safety, ai-safety, survey | 2025-06-30 | Agent Safety | agent-safety, ai-safety, survey | E5 / R3 (95%) | 12 |
| A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Chaochao Chen, Chengye Wang, Fengyuan Yu, Jiaming Zhang Published: 2025-07-26Area: Surveys & ReviewsCitations: 2 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-07-26 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 2 |
| A Survey on Data Security in Large Language Models Fan Lin, Jinhe Su, Kang Chen, Li Shen Published: 2025-08-04Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-08-04 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Changjia Zhu, Chi Zhang, Junjie Xiong, Lingyao Li Published: 2025-08-07Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, survey, surveys-reviews | 2025-08-07 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (93%) | 5 |
| A Review of Developmental Interpretability in Large Language Models Ihor Kendiukhov Published: 2025-08-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-08-19 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (94%) | - |