Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 92 papers (page 1 of 4)路 30 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Position: Capability Control Should be a Separate Goal From Alignment Adrian Weller, David Krueger, Eleni Triantafillou, Shoaib Ahmed Siddiqui Published: 2026-02-05Area: Surveys & ReviewsCitations: - Tags: ai-safety, alignment-training, position, surveys-reviews | 2026-02-05 | Surveys & Reviews | ai-safety, alignment-training, position, surveys-reviews | E6 / R4 (94%) | - |
| Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions Usman Naseem Published: 2026-01-21Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2026-01-21 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E6 / R4 (96%) | 1 |
| Unlearning in LLMs: Methods, Evaluation, and Open Challenges Larry Heck, Tyler Lizzo Published: 2026-01-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2026-01-19 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E7 / R3 (96%) | - |
| Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defenses Chao Li, Chaozhuo Li, Litian Zhang, Xi Zhang Published: 2026-01-07Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2026-01-07 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 1 |
| LLM Harms: A Taxonomy and Discussion Abhejay Murali, Amit Dhurandhar, David Atkinson, Junfeng Jiao Published: 2025-12-05Area: Surveys & ReviewsCitations: - Tags: ai-safety, survey, surveys-reviews | 2025-12-05 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | - |
| Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Bianka Kowalska, Halina Kwa艣nicka Published: 2025-11-24Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-11-24 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 1 |
| A Survey on Unlearning in Large Language Models Fei Sun, Honglin Wang, Jiajun Tan, Jiayue Pu Published: 2025-10-29Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-10-29 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Briland Hitaj, Gabriel Antonio Fontes Rebello, Igor Jochem Sanz, Rodrigo Duarte de Meneses Published: 2025-09-12Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-09-12 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| A Review of Developmental Interpretability in Large Language Models Ihor Kendiukhov Published: 2025-08-19Area: Surveys & ReviewsCitations: - Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-08-19 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (94%) | - |
| Towards Integrated Alignment Ben Y. Reis, William La Cava Published: 2025-08-08Area: Surveys & ReviewsCitations: - Tags: ai-safety, alignment-training, position, surveys-reviews | 2025-08-08 | Surveys & Reviews | ai-safety, alignment-training, position, surveys-reviews | E6 / R4 (93%) | - |
| Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Changjia Zhu, Chi Zhang, Junjie Xiong, Lingyao Li Published: 2025-08-07Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, survey, surveys-reviews | 2025-08-07 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (93%) | 5 |
| A Survey on Data Security in Large Language Models Fan Lin, Jinhe Su, Kang Chen, Li Shen Published: 2025-08-04Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-08-04 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Chaochao Chen, Chengye Wang, Fengyuan Yu, Jiaming Zhang Published: 2025-07-26Area: Surveys & ReviewsCitations: 2 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-07-26 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 2 |
| Report on NSF Workshop on Science of Safe AI Corina P膬s膬reanu, Greg Durrett, Hadas Kress-Gazit, Rajeev Alur Published: 2025-06-24Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-06-24 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (97%) | 1 |
| AI Safety vs. AI Security: Demystifying the Distinction and Boundaries Huan Sun, Ness Shroff, Zhiqiang Lin Published: 2025-06-21Area: Surveys & ReviewsCitations: 2 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-06-21 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R4 (95%) | 2 |
| Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Chuan Qin, Feimin Zhong, Han Wu, Hengshu Zhu Published: 2025-06-11Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-06-11 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (93%) | 5 |
| Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions Dacao Zhang, Guangyi Lv, Kui Yu, Kun Zhang Published: 2025-06-08Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2025-06-08 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (95%) | 1 |
| Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aeree Cho, Duen Horng Chau, Grace C. Kim, Mansi Phute Published: 2025-06-05Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, survey, surveys-reviews | 2025-06-05 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 5 |
| The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Beyza Ermis, Julia Kreutzer, Marzieh Fadaee, Stephen H. Bach Published: 2025-05-30Area: Surveys & ReviewsCitations: 10 Tags: ai-safety, survey, surveys-reviews | 2025-05-30 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | 10 |
| Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies Chenruo Liu, Kenan Tang, Qi Lei, Yao Qin Published: 2025-05-28Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, survey, surveys-reviews | 2025-05-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (95%) | 1 |
| Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression Ping Liu, Yiwei Xie, Zheng Zhang Published: 2025-05-26Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, survey, surveys-reviews | 2025-05-26 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| Security Concerns for Large Language Models: A Survey Benjamin C. M. Fung, Miles Q. Li Published: 2025-05-24Area: Surveys & ReviewsCitations: 29 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2025-05-24 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 29 |
| A Survey on Progress in LLM Alignment from the Perspective of Reward Design Jian Yang, Mark Dras, Miaomiao Ji, Shoujin Wang Published: 2025-05-05Area: Surveys & ReviewsCitations: 10 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-05-05 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (95%) | 10 |
| What is AI safety? What do we want it to be? Cameron Domenico Kirk-Giannini, Jacqueline Harding Published: 2025-05-05Area: Surveys & ReviewsCitations: 1 Tags: ai-safety, position, surveys-reviews | 2025-05-05 | Surveys & Reviews | ai-safety, position, surveys-reviews | E5 / R3 (95%) | 1 |
| AI Awareness Haoyuan Shi, Rongwu Xu, Wei Xu, Xiaojian Li Published: 2025-04-25Area: Surveys & ReviewsCitations: 4 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-04-25 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E6 / R5 (97%) | 4 |
| Safety in Large Reasoning Models: A Survey Baolong Bi, Bryan Hooi, Cheng Wang, Duzhen Zhang Published: 2025-04-24Area: Surveys & ReviewsCitations: 53 Tags: ai-safety, survey, surveys-reviews | 2025-04-24 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 53 |
| An Approach to Technical AGI Safety and Security Alexander Matt Turner, Alex Irpan, Allan Dafoe, Anca Dragan Published: 2025-04-02Area: Surveys & ReviewsCitations: 35 Tags: ai-safety, alignment-training, position, surveys-reviews | 2025-04-02 | Surveys & Reviews | ai-safety, alignment-training, position, surveys-reviews | E5 / R3 (96%) | 35 |
| A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of LLMs Daking Rai, Dong Shu, Haiyan Zhao, Mengnan Du Published: 2025-03-07Area: Surveys & ReviewsCitations: 34 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-03-07 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (94%) | 34 |
| Representation Engineering for Large-Language Models: Survey and Research Challenges Bryan Sukidi, Carsten Maple, David Williams-King, Jennifer Yen Published: 2025-02-24Area: Surveys & ReviewsCitations: 8 Tags: ai-safety, survey, surveys-reviews | 2025-02-24 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 8 |
| A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models Arman Zarei, Barry Menglong Yao, Hongxuan Li, Keivan Rezaei Published: 2025-02-22Area: Surveys & ReviewsCitations: 20 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-02-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (96%) | 20 |