Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| The History and Risks of Reinforcement Learning and Human Feedback Nathan Lambert, Thomas Krendl Gilbert, Tom Zick Published: 2023-10-20Area: Alignment TrainingCitations: 50 Tags: ai-safety, alignment-training, safety-evaluation, survey | 2023-10-20 | Alignment Training | ai-safety, alignment-training, safety-evaluation, survey | E5 / R3 (96%) | 50 |
| Sociotechnical Safety Evaluation of Generative AI Systems Arianna Manzini, Ben Bariach, Conor Griffin, Iason Gabriel Published: 2023-10-18Area: Safety EvaluationCitations: 190 Tags: ai-safety, safety-evaluation, survey | 2023-10-18 | Safety Evaluation | ai-safety, safety-evaluation, survey | E6 / R4 (96%) | 190 |
| Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks Erfan Shayegani, Md Abdullah Al Mamun, Nael Abu-Ghazaleh, Pedram Zaree Published: 2023-10-16Area: Surveys & ReviewsCitations: 238 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2023-10-16 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (96%) | 238 |
| Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey Adrian Weller, Ali Shahin Shamsabadi, Carolyn Ashurst, Victoria Smith Published: 2023-09-27Area: Surveys & ReviewsCitations: 41 Tags: ai-safety, survey, surveys-reviews | 2023-09-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | 41 |
| Large Language Model Alignment: A Survey Chuang Liu, Deyi Xiong, Renren Jin, Tianhao Shen Published: 2023-09-26Area: Surveys & ReviewsCitations: 292 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | 2023-09-26 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 292 |
| Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review Gongshen Liu, Haodong Zhao, Pengzhou Cheng, Wei Du Published: 2023-09-12Area: Adversarial RobustnessCitations: 51 Tags: adversarial-robustness, ai-safety, survey | 2023-09-12 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E7 / R3 (94%) | 51 |
| AI deception: A survey of examples, risks, and potential solutions Aidan O'Gara, Dan Hendrycks, Michael Chen, Peter S. Park Published: 2023-08-28Area: Deception & FailureCitations: 265 Tags: ai-safety, deception-failure, survey | 2023-08-28 | Deception & Failure | ai-safety, deception-failure, survey | E6 / R4 (99%) | 265 |
| Identifying and Mitigating the Security Risks of Generative AI Amrita Roy Chowdhury, Ankur Taly, Anupam Datta, Brad Boyd Published: 2023-08-28Area: Surveys & ReviewsCitations: 126 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-28 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (95%) | 126 |
| Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities Bennett Kleinberg, Lewis D. Griffin, Maximilian Mozes, Xuanli He Published: 2023-08-24Area: Safety EvaluationCitations: 113 Tags: adversarial-robustness, ai-safety, red-teaming, safety-evaluation, survey | 2023-08-24 | Safety Evaluation | adversarial-robustness, ai-safety, red-teaming, safety-evaluation, survey | E5 / R4 (95%) | 113 |
| From Instructions to Intrinsic Human Values - A Survey of Alignment Goals for Big Models Jindong Wang, Jing Yao, Xiaoyuan Yi, Xing Xie Published: 2023-08-23Area: Surveys & ReviewsCitations: 62 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-23 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (94%) | 62 |
| Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback Anand Siththaranjan, Anca Dragan, Andi Peng, Charbel-Raphael Segerie Published: 2023-07-27Area: Alignment TrainingCitations: 761 Tags: ai-safety, alignment-training, survey | 2023-07-27 | Alignment Training | ai-safety, alignment-training, survey | E5 / R3 (97%) | 761 |
| Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries Jonas Schuett, Leonie Koessler Published: 2023-07-17Area: Surveys & ReviewsCitations: 35 Tags: ai-safety, survey, surveys-reviews | 2023-07-17 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E8 / R4 (98%) | 35 |
| An Overview of Catastrophic AI Risks Dan Hendrycks, Mantas Mazeika, Thomas Woodside Published: 2023-06-21Area: Surveys & ReviewsCitations: 258 Tags: ai-safety, survey, surveys-reviews | 2023-06-21 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (100%) | 258 |
| Training Data Extraction From Pre-trained Language Models: A Survey Shotaro Ishihara Published: 2023-05-25Area: Safety EvaluationCitations: 57 Tags: ai-safety, safety-evaluation, survey | 2023-05-25 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (94%) | 57 |
| Editing Large Language Models: Problems, Methods, and Opportunities Bozhong Tian, Huajun Chen, Ningyu Zhang, Peng Wang Published: 2023-05-22Area: Model EditingCitations: 417 Tags: ai-safety, model-editing, survey | 2023-05-22 | Model Editing | ai-safety, model-editing, survey | E8 / R3 (97%) | 417 |
| AI Safety Subproblems for Software Engineering Researchers David Gros, Prem Devanbu, Zhou Yu Published: 2023-04-28Area: Surveys & ReviewsCitations: 4 Tags: ai-safety, survey, surveys-reviews | 2023-04-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 4 |
| A Survey of Machine Unlearning Alan Wee-Chung Liew, Hongzhi Yin, Phi Le Nguyen, Quoc Viet Hung Nguyen Published: 2022-09-06Area: Surveys & ReviewsCitations: 345 Tags: ai-safety, survey, surveys-reviews | 2022-09-06 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 345 |
| Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Anson Ho, Dylan Hadfield-Menell, Stephen Casper, Tilman R盲uker Published: 2022-07-27Area: Surveys & ReviewsCitations: 174 Tags: ai-safety, interpretability, survey, surveys-reviews | 2022-07-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E8 / R4 (94%) | 174 |
| Taxonomy of Risks posed by Language Models Abeba Birhane, Amelia Glaese, Atoosa Kasirzadeh, Borja Balle Published: 2021-12-08Area: Surveys & ReviewsCitations: 1366 Tags: ai-safety, survey, surveys-reviews | 2021-12-08 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R6 (98%) | 1366 |
| Neuron-level Interpretation of Deep NLP Models: A Survey Fahim Dalvi, Hassan Sajjad, Nadir Durrani Published: 2021-08-30Area: Representation AnalysisCitations: 97 Tags: ai-safety, representation-analysis, safety-evaluation, survey | 2021-08-30 | Representation Analysis | ai-safety, representation-analysis, safety-evaluation, survey | E5 / R3 (94%) | 97 |
| An Overview of 11 Proposals for Building Safe Advanced AI Evan Hubinger Published: 2020-12-04Area: Surveys & ReviewsCitations: 27 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2020-12-04 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E7 / R3 (97%) | 27 |
| A Primer in BERTology: What We Know About How BERT Works Anna Rogers, Anna Rumshisky, Olga Kovaleva Published: 2020-02-27Area: Surveys & ReviewsCitations: 1772 Tags: ai-safety, interpretability, survey, surveys-reviews | 2020-02-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 1772 |
| A CIA Triad-Based Taxonomy of Prompt Attacks on Large Language Models Afnan Alkreisat, Ammar Alazab, Amr Adel, Md. Whaiduzzaman Published: -Area: Adversarial RobustnessCitations: 6 Tags: adversarial-robustness, ai-safety, survey | - | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E7 / R4 (97%) | 6 |
| Adversarial attacks and defenses for large language models (LLMs): methods, frameworks & challenges Pranjal Kumar Published: -Area: Adversarial RobustnessCitations: 30 Tags: adversarial-robustness, ai-safety, safety-evaluation, survey | - | Adversarial Robustness | adversarial-robustness, ai-safety, safety-evaluation, survey | E5 / R3 (94%) | 30 |
| Combating Security and Privacy Issues in the Era of Large Language Models Anima Anandkumar, Chaowei Xiao, Fei Wang, Huan Sun Published: -Area: Surveys & ReviewsCitations: 7 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | - | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E7 / R4 (94%) | 7 |
| Emerging threats in AI: a detailed review of misuses and risks across modern AI technologies 脕ine MacDermott, Farkhund Iqbal, Khalifa Al-Room, Niyat Seghid Published: -Area: Surveys & ReviewsCitations: - Tags: ai-safety, survey, surveys-reviews | - | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | - |
| Exploring Privacy and Security Risks in LLMs: Data Leakage, Prompt Injection, and Membership Inference Giancarlo Sperl矛 Published: -Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, survey | - | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E5 / R4 (93%) | - |
| Foundation Models as Guardrails: LLM-and VLM-Based Approaches to Safety and Alignment Huy H. Nguyen, Koki Wataoka, Pride Kavumba, Tomoya Kurosawa Published: -Area: Adversarial RobustnessCitations: - Tags: adversarial-robustness, ai-safety, alignment-training, red-teaming, safety-evaluation, survey | - | Adversarial Robustness | adversarial-robustness, ai-safety, alignment-training, red-teaming, safety-evaluation, survey | E5 / R3 (92%) | - |
| Inaugural Workshop on Provably Safe and Beneficial AI (PSBAI) Stuart Russell Published: -Area: Formal/TheoreticalCitations: - Tags: ai-safety, formaltheoretical, survey | - | Formal/Theoretical | ai-safety, formaltheoretical, survey | E6 / R3 (96%) | - |
| Red teaming large language models: A comprehensive review and critical analysis Abrar Alotaibi, Moataz Ahmed, Muhammad Shahid Jabbar, Sadam Al-Azani Published: -Area: Safety EvaluationCitations: 2 Tags: ai-safety, red-teaming, safety-evaluation, survey | - | Safety Evaluation | ai-safety, red-teaming, safety-evaluation, survey | E6 / R3 (99%) | 2 |