Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 92 papers (page 1 of 4)· 72 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Primer in BERTology: What We Know About How BERT Works Anna Rogers, Anna Rumshisky, Olga Kovaleva Published: 2020-02-27Area: Surveys & ReviewsCitations: 1772 Tags: ai-safety, interpretability, survey, surveys-reviews | 2020-02-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 1772 |
| AI Research Considerations for Human Existential Safety (ARCHES) Andrew Critch, David Krueger Published: 2020-05-30Area: Surveys & ReviewsCitations: 65 Tags: ai-safety, position, surveys-reviews | 2020-05-30 | Surveys & Reviews | ai-safety, position, surveys-reviews | E5 / R3 (97%) | 65 |
| An Overview of 11 Proposals for Building Safe Advanced AI Evan Hubinger Published: 2020-12-04Area: Surveys & ReviewsCitations: 27 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2020-12-04 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E7 / R3 (97%) | 27 |
| Unsolved Problems in ML Safety Dan Hendrycks, Jacob Steinhardt, John Schulman, Nicholas Carlini Published: 2021-09-28Area: Surveys & ReviewsCitations: 359 Tags: ai-safety, alignment-training, position, surveys-reviews | 2021-09-28 | Surveys & Reviews | ai-safety, alignment-training, position, surveys-reviews | E6 / R3 (95%) | 359 |
| Taxonomy of Risks posed by Language Models Abeba Birhane, Amelia Glaese, Atoosa Kasirzadeh, Borja Balle Published: 2021-12-08Area: Surveys & ReviewsCitations: 1366 Tags: ai-safety, survey, surveys-reviews | 2021-12-08 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R6 (98%) | 1366 |
| X-Risk Analysis for AI Research Dan Hendrycks, Mantas Mazeika Published: 2022-06-13Area: Surveys & ReviewsCitations: 81 Tags: ai-safety, position, surveys-reviews | 2022-06-13 | Surveys & Reviews | ai-safety, position, surveys-reviews | E6 / R4 (96%) | 81 |
| Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Anson Ho, Dylan Hadfield-Menell, Stephen Casper, Tilman Räuker Published: 2022-07-27Area: Surveys & ReviewsCitations: 174 Tags: ai-safety, interpretability, survey, surveys-reviews | 2022-07-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E8 / R4 (94%) | 174 |
| A Survey of Machine Unlearning Alan Wee-Chung Liew, Hongzhi Yin, Phi Le Nguyen, Quoc Viet Hung Nguyen Published: 2022-09-06Area: Surveys & ReviewsCitations: 345 Tags: ai-safety, survey, surveys-reviews | 2022-09-06 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 345 |
| AI Safety Subproblems for Software Engineering Researchers David Gros, Prem Devanbu, Zhou Yu Published: 2023-04-28Area: Surveys & ReviewsCitations: 4 Tags: ai-safety, survey, surveys-reviews | 2023-04-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 4 |
| An Overview of Catastrophic AI Risks Dan Hendrycks, Mantas Mazeika, Thomas Woodside Published: 2023-06-21Area: Surveys & ReviewsCitations: 258 Tags: ai-safety, survey, surveys-reviews | 2023-06-21 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (100%) | 258 |
| Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries Jonas Schuett, Leonie Koessler Published: 2023-07-17Area: Surveys & ReviewsCitations: 35 Tags: ai-safety, survey, surveys-reviews | 2023-07-17 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E8 / R4 (98%) | 35 |
| From Instructions to Intrinsic Human Values - A Survey of Alignment Goals for Big Models Jindong Wang, Jing Yao, Xiaoyuan Yi, Xing Xie Published: 2023-08-23Area: Surveys & ReviewsCitations: 62 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-23 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (94%) | 62 |
| Identifying and Mitigating the Security Risks of Generative AI Amrita Roy Chowdhury, Ankur Taly, Anupam Datta, Brad Boyd Published: 2023-08-28Area: Surveys & ReviewsCitations: 126 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-28 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (95%) | 126 |
| Large Language Model Alignment: A Survey Chuang Liu, Deyi Xiong, Renren Jin, Tianhao Shen Published: 2023-09-26Area: Surveys & ReviewsCitations: 292 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | 2023-09-26 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 292 |
| Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey Adrian Weller, Ali Shahin Shamsabadi, Carolyn Ashurst, Victoria Smith Published: 2023-09-27Area: Surveys & ReviewsCitations: 41 Tags: ai-safety, survey, surveys-reviews | 2023-09-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | 41 |
| Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks Erfan Shayegani, Md Abdullah Al Mamun, Nael Abu-Ghazaleh, Pedram Zaree Published: 2023-10-16Area: Surveys & ReviewsCitations: 238 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2023-10-16 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (96%) | 238 |
| Managing AI Risks in an Era of Rapid Progress Anca Dragan, Andrew Yao, Ashwin Acharya, Atılım Güneş Baydin Published: 2023-10-26Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, position, surveys-reviews | 2023-10-26 | Surveys & Reviews | ai-safety, position, surveys-reviews | E5 / R3 (93%) | 80 |
| A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking Rose Hadshar Published: 2023-10-27Area: Surveys & ReviewsCitations: 11 Tags: ai-safety, survey, surveys-reviews | 2023-10-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (92%) | 11 |
| AI Alignment: A Comprehensive Survey Aidan O'Gara, Borong Zhang, Boyuan Chen, Brian Tse Published: 2023-10-30Area: Surveys & ReviewsCitations: 320 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2023-10-30 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R4 (97%) | 320 |
| A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly Jinhao Duan, Kaidi Xu, Yifan Yao, Yuanfang Cai Published: 2023-12-04Area: Surveys & ReviewsCitations: 1022 Tags: ai-safety, survey, surveys-reviews | 2023-12-04 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 1022 |
| Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems Chuanpu Fu, Junwu Xiong, Ke Xu, Peiyang Li Published: 2024-01-11Area: Surveys & ReviewsCitations: 104 Tags: ai-safety, survey, surveys-reviews | 2024-01-11 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R5 (97%) | 104 |
| Security and Privacy Challenges of Large Language Models: A Survey Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu Published: 2024-01-30Area: Surveys & ReviewsCitations: 351 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-01-30 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (97%) | 351 |
| Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey Chao Yang, Jing Shao, Yu Qiao, Zhanhui Zhou Published: 2024-02-14Area: Surveys & ReviewsCitations: 140 Tags: ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | 2024-02-14 | Surveys & Reviews | ai-safety, alignment-training, safety-evaluation, survey, surveys-reviews | E5 / R3 (98%) | 140 |
| Towards Uncovering How Large Language Model Works: An Explainability Perspective Fan Yang, Haiyan Zhao, Himabindu Lakkaraju, Mengnan Du Published: 2024-02-16Area: Surveys & ReviewsCitations: 26 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2024-02-16 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E5 / R3 (93%) | 26 |
| Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices CJ Barberan, Jia He, Richard Anarfi, Sara Abdali Published: 2024-03-19Area: Surveys & ReviewsCitations: 49 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-03-19 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (98%) | 49 |
| SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety Bertie Vidgen, Dirk Hovy, Fabio Pernisi, Paul Röttger Published: 2024-04-08Area: Surveys & ReviewsCitations: 69 Tags: ai-safety, survey, surveys-reviews | 2024-04-08 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (98%) | 69 |
| Foundational Challenges in Assuring Alignment and Safety of Large Language Models Abulhair Saparov, Alan Chan, Aleksandar Petrov, Alexander Pan Published: 2024-04-15Area: Surveys & ReviewsCitations: 211 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2024-04-15 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (94%) | 211 |
| Mechanistic Interpretability for AI Safety — A Review Efstratios Gavves, Leonard Bereska Published: 2024-04-22Area: Surveys & ReviewsCitations: 335 Tags: ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | 2024-04-22 | Surveys & Reviews | ai-safety, interpretability, safety-evaluation, survey, surveys-reviews | E5 / R3 (93%) | 335 |
| A Primer on the Inner Workings of Transformer-based Language Models Arianna Bisazza, Gabriele Sarti, Javier Ferrando, Marta R. Costa-jussà Published: 2024-04-30Area: Surveys & ReviewsCitations: 80 Tags: ai-safety, interpretability, survey, surveys-reviews | 2024-04-30 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E7 / R4 (95%) | 80 |
| AI Risk Management Should Incorporate Both Safety and Security Arvind Narayanan, Bo Li, Boyi Wei, Chaowei Xiao Published: 2024-05-29Area: Surveys & ReviewsCitations: 20 Tags: adversarial-robustness, ai-safety, position, surveys-reviews | 2024-05-29 | Surveys & Reviews | adversarial-robustness, ai-safety, position, surveys-reviews | E5 / R4 (97%) | 20 |