Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models Chonghan Chen, Haibo Jin, Haohan Wang, Jun Zhuang Published: 2024-06-26Area: Adversarial RobustnessCitations: 61 Tags: adversarial-robustness, ai-safety, survey | 2024-06-26 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (97%) | 61 |
| A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Abulhair Saparov, Daking Rai, Shi Feng, Yilun Zhou Published: 2024-07-02Area: Surveys & ReviewsCitations: 91 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-07-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 91 |
| Jailbreak Attacks and Defenses Against Large Language Models: A Survey Jiaxing Song, Ke Xu, Qi Li, Sibo Yi Published: 2024-07-05Area: Surveys & ReviewsCitations: 220 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-07-05 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E8 / R4 (98%) | 220 |
| AI Safety in Generative AI Large Language Models: A Survey Chen Wang, Jaymari Chua, Lina Yao, Shiyi Yang Published: 2024-07-06Area: Surveys & ReviewsCitations: 38 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-07-06 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R3 (95%) | 38 |
| A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends Daizong Liu, Mingyu Yang, Pan Zhou, Wei Hu Published: 2024-07-10Area: Multimodal SafetyCitations: 81 Tags: adversarial-robustness, ai-safety, multimodal-safety, survey | 2024-07-10 | Multimodal Safety | adversarial-robustness, ai-safety, multimodal-safety, survey | E7 / R5 (96%) | 81 |
| Relational Composition in Neural Networks: A Survey and Call to Action Fernanda B. Vi茅gas, Martin Wattenberg Published: 2024-07-19Area: Mechanistic Interp.Citations: 19 Tags: ai-safety, mechanistic-interp, survey | 2024-07-19 | Mechanistic Interp. | ai-safety, mechanistic-interp, survey | E6 / R3 (91%) | 19 |
| Operationalizing a Threat Model for Red-Teaming Large Language Models Anu Pradhan, Apurv Verma, David Rabinowitz, John Doucette Published: 2024-07-20Area: Safety EvaluationCitations: 42 Tags: ai-safety, safety-evaluation, survey | 2024-07-20 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (97%) | 42 |
| The Art of Refusal: A Survey of Abstention in Large Language Models Bill Howe, Bingbing Wen, Chenjun Xu, Jihan Yao Published: 2024-07-25Area: Surveys & ReviewsCitations: 55 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | 2024-07-25 | Surveys & Reviews | adversarial-robustness, ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (94%) | 55 |
| The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies Bo Liu, Dayong Ye, Feng He, Philip S. Yu Published: 2024-07-28Area: Surveys & ReviewsCitations: 85 Tags: ai-safety, survey, surveys-reviews | 2024-07-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R4 (94%) | 85 |
| Can LLMs be Fooled? Investigating Vulnerabilities in LLMs CJ Barberan, Jia He, Richard Anarfi, Sara Abdali Published: 2024-07-30Area: Adversarial RobustnessCitations: 9 Tags: adversarial-robustness, ai-safety, survey | 2024-07-30 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (95%) | 9 |
| Machine Unlearning in Generative AI: A Survey Guangyao Dou, Meng Jiang, Yijun Tian, Zhaoxuan Tan Published: 2024-07-30Area: Model EditingCitations: 47 Tags: ai-safety, model-editing, safety-evaluation, survey | 2024-07-30 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E7 / R3 (95%) | 47 |
| The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability Aaron Mueller, Arnab Sen Sharma, Aruna Sankaranarayanan, Can Rager Published: 2024-08-02Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-02 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (95%) | 3 |
| Towards AI-Safety-by-Design: A Taxonomy of Runtime Guardrails in Foundation Model based Systems Dehai Zhao, Liming Zhu, Md Shamsujjoha, Qinghua Lu Published: 2024-08-05Area: Adversarial RobustnessCitations: 12 Tags: adversarial-robustness, ai-safety, survey | 2024-08-05 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (97%) | 12 |
| Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey Long Bao Le, Luan Ba Dang, Vu Tuan Truong Published: 2024-08-06Area: Adversarial RobustnessCitations: 48 Tags: adversarial-robustness, ai-safety, survey | 2024-08-06 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E8 / R4 (99%) | 48 |
| The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms Adam Davies, Ashkan Khakzar Published: 2024-08-11Area: Surveys & ReviewsCitations: 14 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-11 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R4 (95%) | 14 |
| Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience Anna Ivanova, Chole Li, Danyal Akarca, George Ogden Published: 2024-08-22Area: Surveys & ReviewsCitations: 7 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-08-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (95%) | 7 |
| Recent Advances in Attack and Defense Approaches of Large Language Models Jianbin Jiao, Jing Cui, Junge Zhang, Shuchang Zhou Published: 2024-09-05Area: Adversarial RobustnessCitations: 9 Tags: adversarial-robustness, ai-safety, survey | 2024-09-05 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R3 (95%) | 9 |
| Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis Oliver Guest, Oscar Delaney, Zoe Williams Published: 2024-09-12Area: Surveys & ReviewsCitations: 3 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E6 / R3 (94%) | 3 |
| Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks Benji Peng, Junyu Liu, Keyu Chen, Ming Li Published: 2024-09-12Area: Surveys & ReviewsCitations: 31 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-09-12 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (94%) | 31 |
| Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI Ambrish Rawat, Beat Buesser, Elizabeth M. Daly, Erik Miehling Published: 2024-09-23Area: Safety EvaluationCitations: 14 Tags: ai-safety, red-teaming, safety-evaluation, survey | 2024-09-23 | Safety Evaluation | ai-safety, red-teaming, safety-evaluation, survey | E6 / R4 (95%) | 14 |
| Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Fatih Ilhan, Ling Liu, Selim Furkan Tekin, Sihao Hu Published: 2024-09-26Area: Surveys & ReviewsCitations: 82 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-09-26 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (95%) | 82 |
| A Survey on the Honesty of Large Language Models Cheng Yang, Chufan Shi, Deng Cai, Jie Zhou Published: 2024-09-27Area: Surveys & ReviewsCitations: 18 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2024-09-27 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R4 (94%) | 18 |
| Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges Chaowei Xiao, Fei Wang, Jiashu Xu, Muhao Chen Published: 2024-09-30Area: Adversarial RobustnessCitations: 13 Tags: adversarial-robustness, ai-safety, survey | 2024-09-30 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R3 (94%) | 13 |
| Challenges and Future Directions of Data-Centric AI Alignment Jeffrey Wang, Leitian Tao, Min-Hsuan Yeh, Seongheon Park Published: 2024-10-02Area: Alignment TrainingCitations: 8 Tags: ai-safety, alignment-training, survey | 2024-10-02 | Alignment Training | ai-safety, alignment-training, survey | E5 / R3 (94%) | 8 |
| Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond Shanshan Han Published: 2024-10-09Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | 2024-10-09 | Surveys & Reviews | adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | E6 / R3 (92%) | 1 |
| Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations Nilay Pochhi, Tarun Raheja Published: 2024-10-09Area: Safety EvaluationCitations: 12 Tags: ai-safety, safety-evaluation, survey | 2024-10-09 | Safety Evaluation | ai-safety, safety-evaluation, survey | E6 / R4 (94%) | 12 |
| SoK: Prompt Hacking of Large Language Models Baha Rababah, Carson Leung, Cuneyt Gurcan Akcora, Matthew Kwiatkowski Published: 2024-10-16Area: Adversarial RobustnessCitations: 8 Tags: adversarial-robustness, ai-safety, survey | 2024-10-16 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (96%) | 8 |
| Jailbreaking and Mitigation of Vulnerabilities in Large Language Models Benji Peng, Caitlyn Heqi Yin, Lawrence K.Q. Yan, Ming Liu Published: 2024-10-20Area: Adversarial RobustnessCitations: 27 Tags: adversarial-robustness, ai-safety, survey | 2024-10-20 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E7 / R3 (96%) | 27 |
| Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis Aishvariya Priya, Amit Giloni, Hisashi Kojima, Inderjeet Singh Published: 2024-10-21Area: Safety EvaluationCitations: 8 Tags: ai-safety, safety-evaluation, survey | 2024-10-21 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (94%) | 8 |
| Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey Chiyu Zhang, Jiafei Wu, Lu Zhou, Xiaogang Xu Published: 2024-10-31Area: Adversarial RobustnessCitations: 31 Tags: adversarial-robustness, ai-safety, survey | 2024-10-31 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (95%) | 31 |