Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Exploring Vulnerabilities and Protections in Large Language Models: A Survey Chenhui Hu, Frank Weizhen Liu Published: 2024-06-01Area: Adversarial RobustnessCitations: 12 Tags: adversarial-robustness, ai-safety, survey | 2024-06-01 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E8 / R5 (98%) | 12 |
| A Primer on the Inner Workings of Transformer-based Language Models Arianna Bisazza, Gabriele Sarti, Javier Ferrando, Marta R. Costa-jussà Published: 2024-04-30Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-04-30 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (95%) | 80 |
| How to Use and Interpret Activation Patching Neel Nanda, Stefan Heimersheim Published: 2024-04-23Area: Mechanistic Interp.Citations: 109 Tags: ai-safety, interpretability, mechanistic-interp, survey | 2024-04-23 | Mechanistic Interp. | ai-safety, interpretability, mechanistic-interp, survey | E6 / R4 (95%) | 109 |
| Holistic Safety and Responsibility Evaluations of Advanced AI Models Allan Dafoe, Christina Butterfield, Dawn Bloxwich, Iason Gabriel Published: 2024-04-22Area: Safety EvaluationCitations: 17 Tags: ai-safety, safety-evaluation, survey | 2024-04-22 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 17 |
| Mechanistic Interpretability for AI Safety — A Review Efstratios Gavves, Leonard Bereska Published: 2024-04-22Area: Surveys & ReviewsCitations: 335 Tags: ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | 2024-04-22 | Surveys & Reviews | ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 335 |
| Foundational Challenges in Assuring Alignment and Safety of Large Language Models Abulhair Saparov, Alan Chan, Aleksandar Petrov, Alexander Pan Published: 2024-04-15Area: Surveys & ReviewsCitations: 211 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-04-15 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (94%) | 211 |
| RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs Ameet Deshpande, Ashwin Kalyan, Bruno Castro da Silva, Karthik Narasimhan Published: 2024-04-12Area: Alignment TrainingCitations: 105 Tags: ai-safety, alignment-training, survey | 2024-04-12 | Alignment Training | ai-safety, alignment-training, survey | E5 / R4 (95%) | 105 |
| SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety Bertie Vidgen, Dirk Hovy, Fabio Pernisi, Paul Röttger Published: 2024-04-08Area: Surveys & ReviewsCitations: 69 Tags: ai-safety, survey, surveys-reviews | 2024-04-08 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (98%) | 69 |
| Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security Shaofeng Li, Yihe Fan, Yuxin Cao, Ziyao Liu Published: 2024-04-08Area: Multimodal SafetyCitations: 21 Tags: ai-safety, multimodal-safety, survey | 2024-04-08 | Multimodal Safety | ai-safety, multimodal-safety, survey | E6 / R3 (95%) | 21 |
| Digital Forgetting in Large Language Models: A Survey of Unlearning Methods Alberto Blanco-Justicia, Benet Manzanares, David Sánchez, Guillem Collell Published: 2024-04-02Area: Model EditingCitations: 45 Tags: ai-safety, model-editing, safety-evaluation, survey | 2024-04-02 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E5 / R3 (96%) | 45 |
| Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices CJ Barberan, Jia He, Richard Anarfi, Sara Abdali Published: 2024-03-19Area: Surveys & ReviewsCitations: 49 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-03-19 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (98%) | 49 |
| Machine Unlearning: Taxonomy, Metrics, Applications, Challenges, and Prospects Anmin Fu, Chunyi Zhou, Hui Chen, Na Li Published: 2024-03-13Area: Model EditingCitations: 36 Tags: adversarial-robustness, ai-safety, model-editing, safety-evaluation, survey | 2024-03-13 | Model Editing | adversarial-robustness, ai-safety, model-editing, safety-evaluation, survey | E5 / R3 (95%) | 36 |
| Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models Aman Chadha, Arijit Ghosh Chowdhury, Faysal Hossain Shezan, Md Mofijul Islam Published: 2024-03-03Area: Adversarial RobustnessCitations: 47 Tags: adversarial-robustness, ai-safety, survey | 2024-03-03 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (94%) | 47 |
| Exploring Advanced Methodologies in Security Evaluation for LLMs Jiawei Zhang, Jun Huang, Qi Wang, Weihong Han Published: 2024-02-28Area: Safety EvaluationCitations: - Tags: ai-safety, safety-evaluation, survey | 2024-02-28 | Safety Evaluation | ai-safety, safety-evaluation, survey | E6 / R3 (93%) | - |
| Generative AI Security: Challenges and Countermeasures Banghua Zhu, David Wagner, Jiantao Jiao, Norman Mu Published: 2024-02-20Area: Adversarial RobustnessCitations: 14 Tags: adversarial-robustness, ai-safety, survey | 2024-02-20 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (95%) | 14 |
| Towards Uncovering How Large Language Model Works: An Explainability Perspective Fan Yang, Haiyan Zhao, Himabindu Lakkaraju, Mengnan Du Published: 2024-02-16Area: Surveys & ReviewsCitations: 26 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-02-16 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E5 / R3 (93%) | 26 |
| A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents Boyuan Zheng, Chaowei Xiao, Huan Sun, Lingbo Mo Published: 2024-02-15Area: Agent SafetyCitations: 23 Tags: adversarial-robustness, agent-safety, ai-safety, survey | 2024-02-15 | Agent Safety | adversarial-robustness, agent-safety, ai-safety, survey | E6 / R4 (97%) | 23 |
| Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey Chao Yang, Jing Shao, Yu Qiao, Zhanhui Zhou Published: 2024-02-14Area: Surveys & ReviewsCitations: 140 Tags: ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | 2024-02-14 | Surveys & Reviews | ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | E5 / R3 (98%) | 140 |
| Rethinking Machine Unlearning for Large Language Models Chris Yuhao Liu, Hang Li, Jinghan Jia, Kush R. Varshney Published: 2024-02-13Area: Model EditingCitations: 227 Tags: ai-safety, alignment-training, model-editing, safety-evaluation, survey | 2024-02-13 | Model Editing | ai-safety, alignment-training, model-editing, safety-evaluation, survey | E5 / R3 (95%) | 227 |
| A Roadmap to Pluralistic Alignment Andre Ye, Christopher Michael Rytting, Jared Moore, Jillian Fisher Published: 2024-02-07Area: Alignment TrainingCitations: 161 Tags: ai-safety, alignment-training, survey | 2024-02-07 | Alignment Training | ai-safety, alignment-training, survey | E5 / R3 (94%) | 161 |
| Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science Arman Cohan, Jian Tang, Kunlun Zhu, Mark Gerstein Published: 2024-02-06Area: Agent SafetyCitations: 55 Tags: agent-safety, ai-safety, alignment-training, survey | 2024-02-06 | Agent Safety | agent-safety, ai-safety, alignment-training, survey | E5 / R3 (95%) | 55 |
| Safety of Multimodal Large Language Models on Images and Text Chao Yang, Xin Liu, Yichen Zhu, Yunshi Lan Published: 2024-02-01Area: Multimodal SafetyCitations: 65 Tags: ai-safety, multimodal-safety, survey | 2024-02-01 | Multimodal Safety | ai-safety, multimodal-safety, survey | E7 / R3 (95%) | 65 |
| Security and Privacy Challenges of Large Language Models: A Survey Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu Published: 2024-01-30Area: Surveys & ReviewsCitations: 351 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-01-30 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (97%) | 351 |
| Red-Teaming for Generative AI: Silver Bullet or Security Theater? Anusha Sinha, Hoda Heidari, Michael Feffer, Zachary C. Lipton Published: 2024-01-29Area: Safety EvaluationCitations: 126 Tags: ai-safety, safety-evaluation, survey | 2024-01-29 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (94%) | 126 |
| Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems Chuanpu Fu, Junwu Xiong, Ke Xu, Peiyang Li Published: 2024-01-11Area: Surveys & ReviewsCitations: 104 Tags: ai-safety, survey, surveys-reviews | 2024-01-11 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R5 (97%) | 104 |
| A Comprehensive Study of Knowledge Editing for Large Language Models Bozhong Tian, Fei Huang, Huajun Chen, Jia-Chen Gu Published: 2024-01-02Area: Model EditingCitations: 134 Tags: ai-safety, model-editing, survey | 2024-01-02 | Model Editing | ai-safety, model-editing, survey | E5 / R3 (95%) | 134 |
| A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly Jinhao Duan, Kaidi Xu, Yifan Yao, Yuanfang Cai Published: 2023-12-04Area: Surveys & ReviewsCitations: 1022 Tags: ai-safety, survey, surveys-reviews | 2023-12-04 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 1022 |
| Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges Dan Qu, Hao Zhang, Heyu Chang, Nianwen Si Published: 2023-11-27Area: Model EditingCitations: 41 Tags: ai-safety, model-editing, safety-evaluation, survey | 2023-11-27 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E7 / R5 (95%) | 41 |
| AI Alignment: A Comprehensive Survey Aidan O'Gara, Borong Zhang, Boyuan Chen, Brian Tse Published: 2023-10-30Area: Surveys & ReviewsCitations: 320 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2023-10-30 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R4 (97%) | 320 |
| A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking Rose Hadshar Published: 2023-10-27Area: Surveys & ReviewsCitations: 11 Tags: ai-safety, survey, surveys-reviews | 2023-10-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (92%) | 11 |