Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of LLMs Daking Rai, Dong Shu, Haiyan Zhao, Mengnan Du Published: 2025-03-07Area: Surveys & ReviewsCitations: 34 Tags: ai-safety, safety-evaluation, survey, surveys-reviews | 2025-03-07 | Surveys & Reviews | ai-safety, safety-evaluation, survey, surveys-reviews | E5 / R3 (94%) | 34 |
| Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models Akshay Gupta, Alberto Purpura, Andy Luo, Jesse Zymet Published: 2025-03-03Area: Safety EvaluationCitations: 7 Tags: ai-safety, red-teaming, safety-evaluation, survey | 2025-03-03 | Safety Evaluation | ai-safety, red-teaming, safety-evaluation, survey | E6 / R4 (96%) | 7 |
| Representation Engineering for Large-Language Models: Survey and Research Challenges Bryan Sukidi, Carsten Maple, David Williams-King, Jennifer Yen Published: 2025-02-24Area: Surveys & ReviewsCitations: 8 Tags: ai-safety, survey, surveys-reviews | 2025-02-24 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 8 |
| A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models Fakhri Karray, Hans-Arno Jacobsen, Herbert Woisetschl盲ger, Jiahui Geng Published: 2025-02-22Area: Model EditingCitations: 23 Tags: ai-safety, model-editing, safety-evaluation, survey | 2025-02-22 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E6 / R4 (97%) | 23 |
| A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models Arman Zarei, Barry Menglong Yao, Hongxuan Li, Keivan Rezaei Published: 2025-02-22Area: Surveys & ReviewsCitations: 20 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-02-22 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (96%) | 20 |
| Computational Safety for Generative AI: A Signal Processing Perspective Pin-Yu Chen Published: 2025-02-18Area: Adversarial RobustnessCitations: 2 Tags: adversarial-robustness, ai-safety, survey | 2025-02-18 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (96%) | 2 |
| Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks Meng Han, Minghao Li, Mohan Li, Wenpeng Xing Published: 2025-02-18Area: Adversarial RobustnessCitations: 29 Tags: adversarial-robustness, ai-safety, survey | 2025-02-18 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R4 (95%) | 29 |
| A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Bo Du, Dacheng Tao, Mang Ye, Nenghai Yu Published: 2025-02-14Area: Multimodal SafetyCitations: 47 Tags: ai-safety, multimodal-safety, safety-evaluation, survey | 2025-02-14 | Multimodal Safety | ai-safety, multimodal-safety, safety-evaluation, survey | E5 / R3 (95%) | 47 |
| AI Safety for Everyone Atoosa Kasirzadeh, B谩lint Gyevnar Published: 2025-02-13Area: Surveys & ReviewsCitations: 17 Tags: ai-safety, survey, surveys-reviews | 2025-02-13 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (96%) | 17 |
| A Survey of LLM Alignment: Instruction Understanding, Intention Reasoning, and Reliable Generation Cheng Ji, Feihong Lu, Jianxin Li, Qian Li Published: 2025-02-13Area: Surveys & ReviewsCitations: 2 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2025-02-13 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (97%) | 2 |
| A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Hieu Minh Nguyen Published: 2025-02-10Area: Surveys & ReviewsCitations: 5 Tags: ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | 2025-02-10 | Surveys & Reviews | ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | E5 / R3 (92%) | 5 |
| A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods Qingchuan Zhao, Tao Ni, Wei-Bin Lee, Yihe Zhou Published: 2025-02-06Area: Adversarial RobustnessCitations: 24 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | 2025-02-06 | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R3 (95%) | 24 |
| Safety at Scale: A Comprehensive Survey of Large Model Safety Baoyuan Wu, Bo Li, Chaowei Xiao, Cihang Xie Published: 2025-02-02Area: Surveys & ReviewsCitations: 18 Tags: ai-safety, survey, surveys-reviews | 2025-02-02 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (99%) | 18 |
| Open Problems in Mechanistic Interpretability Adria Garriga-Alonso, Alejandro Ortega, Arthur Conmy, Atticus Geiger Published: 2025-01-27Area: Surveys & ReviewsCitations: 107 Tags: ai-safety, interpretability, survey, surveys-reviews | 2025-01-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R3 (96%) | 107 |
| Large Language Model Safety: A Holistic Survey Bojian Jiang, Chuang Liu, Dan Shi, Deyi Xiong Published: 2024-12-23Area: Surveys & ReviewsCitations: 47 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-12-23 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R5 (99%) | 47 |
| Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models Benjamin K. Bergen, Cameron R. Jones Published: 2024-12-22Area: Deception & FailureCitations: 16 Tags: ai-safety, deception-failure, survey | 2024-12-22 | Deception & Failure | ai-safety, deception-failure, survey | E6 / R4 (94%) | 16 |
| The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment HyunJin Kim, Jianxun Lian, Jing Yao, JinYeong Bak Published: 2024-12-21Area: Surveys & ReviewsCitations: 9 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-12-21 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R3 (94%) | 9 |
| The superalignment of superhuman intelligence with large language models Jie Tang, Minlie Huang, Pei Ke, Shiyao Cui Published: 2024-12-15Area: Scalable OversightCitations: 1 Tags: ai-safety, alignment-training, scalable-oversight, survey | 2024-12-15 | Scalable Oversight | ai-safety, alignment-training, scalable-oversight, survey | E6 / R3 (93%) | 1 |
| Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective D'Jeff K. Nkashama, Froduald Kabanza, Jean Marie Tshimula, Marc Frappier Published: 2024-11-25Area: Adversarial RobustnessCitations: 4 Tags: adversarial-robustness, ai-safety, survey | 2024-11-25 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R4 (94%) | 4 |
| Sycophancy in Large Language Models: Causes and Mitigations Lars Malmqvist Published: 2024-11-22Area: Deception & FailureCitations: 71 Tags: adversarial-robustness, ai-safety, deception-failure, survey | 2024-11-22 | Deception & Failure | adversarial-robustness, ai-safety, deception-failure, survey | E5 / R3 (96%) | 71 |
| SoK: Unifying Cybersecurity and Cybersafety of Multimodal Foundation Models with an Information Theory Approach Bo Li, Chaowei Xiao, Hammond Pearce, Jiamin Chang Published: 2024-11-17Area: Multimodal SafetyCitations: - Tags: ai-safety, multimodal-safety, survey | 2024-11-17 | Multimodal Safety | ai-safety, multimodal-safety, survey | E4 / R3 (94%) | - |
| Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey Huaibo Huang, Miaoxuan Zhang, Peipei Li, Ran He Published: 2024-11-14Area: Multimodal SafetyCitations: 25 Tags: adversarial-robustness, ai-safety, multimodal-safety, survey | 2024-11-14 | Multimodal Safety | adversarial-robustness, ai-safety, multimodal-safety, survey | E5 / R4 (96%) | 25 |
| Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey Chiyu Zhang, Jiafei Wu, Lu Zhou, Xiaogang Xu Published: 2024-10-31Area: Adversarial RobustnessCitations: 31 Tags: adversarial-robustness, ai-safety, survey | 2024-10-31 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (95%) | 31 |
| Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis Aishvariya Priya, Amit Giloni, Hisashi Kojima, Inderjeet Singh Published: 2024-10-21Area: Safety EvaluationCitations: 8 Tags: ai-safety, safety-evaluation, survey | 2024-10-21 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (94%) | 8 |
| Jailbreaking and Mitigation of Vulnerabilities in Large Language Models Benji Peng, Caitlyn Heqi Yin, Lawrence K.Q. Yan, Ming Liu Published: 2024-10-20Area: Adversarial RobustnessCitations: 27 Tags: adversarial-robustness, ai-safety, survey | 2024-10-20 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E7 / R3 (96%) | 27 |
| SoK: Prompt Hacking of Large Language Models Baha Rababah, Carson Leung, Cuneyt Gurcan Akcora, Matthew Kwiatkowski Published: 2024-10-16Area: Adversarial RobustnessCitations: 8 Tags: adversarial-robustness, ai-safety, survey | 2024-10-16 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R3 (96%) | 8 |
| Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond Shanshan Han Published: 2024-10-09Area: Surveys & ReviewsCitations: 1 Tags: adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | 2024-10-09 | Surveys & Reviews | adversarial-robustness, ai-safety, red-teaming, survey, surveys-reviews | E6 / R3 (92%) | 1 |
| Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations Nilay Pochhi, Tarun Raheja Published: 2024-10-09Area: Safety EvaluationCitations: 12 Tags: ai-safety, safety-evaluation, survey | 2024-10-09 | Safety Evaluation | ai-safety, safety-evaluation, survey | E6 / R4 (94%) | 12 |
| Challenges and Future Directions of Data-Centric AI Alignment Jeffrey Wang, Leitian Tao, Min-Hsuan Yeh, Seongheon Park Published: 2024-10-02Area: Alignment TrainingCitations: 8 Tags: ai-safety, alignment-training, survey | 2024-10-02 | Alignment Training | ai-safety, alignment-training, survey | E5 / R3 (94%) | 8 |
| Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges Chaowei Xiao, Fei Wang, Jiashu Xu, Muhao Chen Published: 2024-09-30Area: Adversarial RobustnessCitations: 13 Tags: adversarial-robustness, ai-safety, survey | 2024-09-30 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E6 / R3 (94%) | 13 |