Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| AI Safety for Everyone Atoosa Kasirzadeh, Bálint Gyevnar Published: 2025-02-13Area: Surveys & ReviewsCitations: 17 Tags: ai-safety, survey, surveys-reviews | 2025-02-13 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (96%) | 17 |
| A Survey of LLM Alignment: Instruction Understanding, Intention Reasoning, and Reliable Generation Cheng Ji, Feihong Lu, Jianxin Li, Qian Li Published: 2025-02-13Area: Surveys & ReviewsCitations: 2 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-02-13 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (97%) | 2 |
| A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Hieu Minh Nguyen Published: 2025-02-10Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | 2025-02-10 | Surveys & Reviews | ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | E5 / R3 (92%) | 5 |
| Safety at Scale: A Comprehensive Survey of Large Model Safety Baoyuan Wu, Bo Li, Chaowei Xiao, Cihang Xie Published: 2025-02-02Area: Surveys & ReviewsCitations: 18 Tags: ai-safety, survey, surveys-reviews | 2025-02-02 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (99%) | 18 |
| Building Bridges, Not Walls: Advancing Interpretability by Unifying Feature, Data, and Model Component Attribution Hima Lakkaraju, Shichang Zhang, Tessa Han, Usha Bhalla Published: 2025-01-31Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, position, surveys-reviews | 2025-01-31 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (97%) | 3 |
| Open Problems in Mechanistic Interpretability Adria Garriga-Alonso, Alejandro Ortega, Arthur Conmy, Atticus Geiger Published: 2025-01-27Area: Surveys & ReviewsCitations: 107 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-01-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 107 |
| Large Language Model Safety: A Holistic Survey Bojian Jiang, Chuang Liu, Dan Shi, Deyi Xiong Published: 2024-12-23Area: Surveys & ReviewsCitations: 47 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-12-23 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R5 (99%) | 47 |
| The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment HyunJin Kim, Jianxun Lian, Jing Yao, JinYeong Bak Published: 2024-12-21Area: Surveys & ReviewsCitations: 9 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-12-21 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R3 (94%) | 9 |
| Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond Shanshan Han Published: 2024-10-09Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | 2024-10-09 | Surveys & Reviews | adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | E6 / R3 (92%) | 1 |
| Mechanistic? Naomi Saphra, Sarah Wiegreffe Published: 2024-10-07Area: Surveys & ReviewsCitations: 38 Tags: ai-safety, interpretability, position, surveys-reviews | 2024-10-07 | Surveys & Reviews | ai-safety, interpretability, position, surveys-reviews | E5 / R3 (94%) | 38 |
| A Survey on the Honesty of Large Language Models Cheng Yang, Chufan Shi, Deng Cai, Jie Zhou Published: 2024-09-27Area: Surveys & ReviewsCitations: 18 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2024-09-27 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R4 (94%) | 18 |
| Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Fatih Ilhan, Ling Liu, Selim Furkan Tekin, Sihao Hu Published: 2024-09-26Area: Surveys & ReviewsCitations: 82 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-09-26 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (95%) | 82 |
| Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis Oliver Guest, Oscar Delaney, Zoe Williams Published: 2024-09-12Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R3 (94%) | 3 |
| Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks Benji Peng, Junyu Liu, Keyu Chen, Ming Li Published: 2024-09-12Area: Surveys & ReviewsCitations: 31 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (94%) | 31 |
| Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience Anna Ivanova, Chole Li, Danyal Akarca, George Ogden Published: 2024-08-22Area: Surveys & ReviewsCitations: 7 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (95%) | 7 |
| The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms Adam Davies, Ashkan Khakzar Published: 2024-08-11Area: Surveys & ReviewsCitations: 14 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-11 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (95%) | 14 |
| The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability Aaron Mueller, Arnab Sen Sharma, Aruna Sankaranarayanan, Can Rager Published: 2024-08-02Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies Bo Liu, Dayong Ye, Feng He, Philip S. Yu Published: 2024-07-28Area: Surveys & ReviewsCitations: 85 Tags: ai-safety, survey, surveys-reviews | 2024-07-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R4 (94%) | 85 |
| The Art of Refusal: A Survey of Abstention in Large Language Models Bill Howe, Bingbing Wen, Chenjun Xu, Jihan Yao Published: 2024-07-25Area: Surveys & ReviewsCitations: 55 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2024-07-25 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (94%) | 55 |
| AI Safety in Generative AI Large Language Models: A Survey Chen Wang, Jaymari Chua, Lina Yao, Shiyi Yang Published: 2024-07-06Area: Surveys & ReviewsCitations: 38 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-07-06 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R3 (95%) | 38 |
| Jailbreak Attacks and Defenses Against Large Language Models: A Survey Jiaxing Song, Ke Xu, Qi Li, Sibo Yi Published: 2024-07-05Area: Surveys & ReviewsCitations: 220 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-07-05 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E8 / R4 (98%) | 220 |
| A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Abulhair Saparov, Daking Rai, Shi Feng, Yilun Zhou Published: 2024-07-02Area: Surveys & ReviewsCitations: 91 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-07-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 91 |
| AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies Andy Zhou, Bo Li, Dawn Song, Kevin Klyman Published: 2024-06-25Area: Surveys & ReviewsCitations: 48 Tags: ai-safety, survey, surveys-reviews | 2024-06-25 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E8 / R4 (99%) | 48 |
| From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking Siyuan Wang, Zhihao Fan, Zhongyu Wei, Zhuohan Long Published: 2024-06-21Area: Surveys & ReviewsCitations: 22 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2024-06-21 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E7 / R4 (95%) | 22 |
| A Survey on Human Preference Learning for Large Language Models Juntao Li, Kehai Chen, Liqiang Nie, Min Zhang Published: 2024-06-17Area: Surveys & ReviewsCitations: 16 Tags: ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | 2024-06-17 | Surveys & Reviews | ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | E5 / R3 (95%) | 16 |
| Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey Bo Liu, Dayong Ye, Ming Ding, Philip S. Yu Published: 2024-06-12Area: Surveys & ReviewsCitations: 20 Tags: ai-safety, survey, surveys-reviews | 2024-06-12 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (97%) | 20 |
| Safeguarding Large Language Models: A Survey Changshun Wu, Gaojie Jin, Jie Meng, Jinwei Hu Published: 2024-06-03Area: Surveys & ReviewsCitations: 82 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-06-03 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (97%) | 82 |
| AI Risk Management Should Incorporate Both Safety and Security Arvind Narayanan, Bo Li, Boyi Wei, Chaowei Xiao Published: 2024-05-29Area: Surveys & ReviewsCitations: 20 Tags: adversarial-robustness, ai-safety, position, surveys-reviews | 2024-05-29 | Surveys & Reviews | adversarial-robustness, ai-safety, position, surveys-reviews | E5 / R4 (97%) | 20 |
| A Primer on the Inner Workings of Transformer-based Language Models Arianna Bisazza, Gabriele Sarti, Javier Ferrando, Marta R. Costa-jussà Published: 2024-04-30Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-04-30 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (95%) | 80 |
| Mechanistic Interpretability for AI Safety — A Review Efstratios Gavves, Leonard Bereska Published: 2024-04-22Area: Surveys & ReviewsCitations: 335 Tags: ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | 2024-04-22 | Surveys & Reviews | ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 335 |