Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 186 papers (page 1 of 7)路 47 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Primer in BERTology: What We Know About How BERT Works Anna Rogers, Anna Rumshisky, Olga Kovaleva Published: 2020-02-27Area: Surveys & ReviewsCitations: 1772 Tags: ai-safety, interpretability, survey, surveys-reviews | 2020-02-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E5 / R4 (94%) | 1772 |
| An Overview of 11 Proposals for Building Safe Advanced AI Evan Hubinger Published: 2020-12-04Area: Surveys & ReviewsCitations: 27 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2020-12-04 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E7 / R3 (97%) | 27 |
| Neuron-level Interpretation of Deep NLP Models: A Survey Fahim Dalvi, Hassan Sajjad, Nadir Durrani Published: 2021-08-30Area: Representation AnalysisCitations: 97 Tags: ai-safety, representation-analysis, safety-evaluation, survey | 2021-08-30 | Representation Analysis | ai-safety, representation-analysis, safety-evaluation, survey | E5 / R3 (94%) | 97 |
| Taxonomy of Risks posed by Language Models Abeba Birhane, Amelia Glaese, Atoosa Kasirzadeh, Borja Balle Published: 2021-12-08Area: Surveys & ReviewsCitations: 1366 Tags: ai-safety, survey, surveys-reviews | 2021-12-08 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R6 (98%) | 1366 |
| Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Anson Ho, Dylan Hadfield-Menell, Stephen Casper, Tilman R盲uker Published: 2022-07-27Area: Surveys & ReviewsCitations: 174 Tags: ai-safety, interpretability, survey, surveys-reviews | 2022-07-27 | Surveys & Reviews | ai-safety, interpretability, survey, surveys-reviews | E8 / R4 (94%) | 174 |
| A Survey of Machine Unlearning Alan Wee-Chung Liew, Hongzhi Yin, Phi Le Nguyen, Quoc Viet Hung Nguyen Published: 2022-09-06Area: Surveys & ReviewsCitations: 345 Tags: ai-safety, survey, surveys-reviews | 2022-09-06 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 345 |
| AI Safety Subproblems for Software Engineering Researchers David Gros, Prem Devanbu, Zhou Yu Published: 2023-04-28Area: Surveys & ReviewsCitations: 4 Tags: ai-safety, survey, surveys-reviews | 2023-04-28 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (94%) | 4 |
| Editing Large Language Models: Problems, Methods, and Opportunities Bozhong Tian, Huajun Chen, Ningyu Zhang, Peng Wang Published: 2023-05-22Area: Model EditingCitations: 417 Tags: ai-safety, model-editing, survey | 2023-05-22 | Model Editing | ai-safety, model-editing, survey | E8 / R3 (97%) | 417 |
| Training Data Extraction From Pre-trained Language Models: A Survey Shotaro Ishihara Published: 2023-05-25Area: Safety EvaluationCitations: 57 Tags: ai-safety, safety-evaluation, survey | 2023-05-25 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R4 (94%) | 57 |
| An Overview of Catastrophic AI Risks Dan Hendrycks, Mantas Mazeika, Thomas Woodside Published: 2023-06-21Area: Surveys & ReviewsCitations: 258 Tags: ai-safety, survey, surveys-reviews | 2023-06-21 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R3 (100%) | 258 |
| Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries Jonas Schuett, Leonie Koessler Published: 2023-07-17Area: Surveys & ReviewsCitations: 35 Tags: ai-safety, survey, surveys-reviews | 2023-07-17 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E8 / R4 (98%) | 35 |
| Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback Anand Siththaranjan, Anca Dragan, Andi Peng, Charbel-Raphael Segerie Published: 2023-07-27Area: Alignment TrainingCitations: 761 Tags: ai-safety, alignment-training, survey | 2023-07-27 | Alignment Training | ai-safety, alignment-training, survey | E5 / R3 (97%) | 761 |
| From Instructions to Intrinsic Human Values - A Survey of Alignment Goals for Big Models Jindong Wang, Jing Yao, Xiaoyuan Yi, Xing Xie Published: 2023-08-23Area: Surveys & ReviewsCitations: 62 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-23 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E5 / R3 (94%) | 62 |
| Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities Bennett Kleinberg, Lewis D. Griffin, Maximilian Mozes, Xuanli He Published: 2023-08-24Area: Safety EvaluationCitations: 113 Tags: adversarial-robustness, ai-safety, red-teaming, safety-evaluation, survey | 2023-08-24 | Safety Evaluation | adversarial-robustness, ai-safety, red-teaming, safety-evaluation, survey | E5 / R4 (95%) | 113 |
| AI deception: A survey of examples, risks, and potential solutions Aidan O'Gara, Dan Hendrycks, Michael Chen, Peter S. Park Published: 2023-08-28Area: Deception & FailureCitations: 265 Tags: ai-safety, deception-failure, survey | 2023-08-28 | Deception & Failure | ai-safety, deception-failure, survey | E6 / R4 (99%) | 265 |
| Identifying and Mitigating the Security Risks of Generative AI Amrita Roy Chowdhury, Ankur Taly, Anupam Datta, Brad Boyd Published: 2023-08-28Area: Surveys & ReviewsCitations: 126 Tags: ai-safety, alignment-training, survey, surveys-reviews | 2023-08-28 | Surveys & Reviews | ai-safety, alignment-training, survey, surveys-reviews | E6 / R4 (95%) | 126 |
| Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review Gongshen Liu, Haodong Zhao, Pengzhou Cheng, Wei Du Published: 2023-09-12Area: Adversarial RobustnessCitations: 51 Tags: adversarial-robustness, ai-safety, survey | 2023-09-12 | Adversarial Robustness | adversarial-robustness, ai-safety, survey | E7 / R3 (94%) | 51 |
| Large Language Model Alignment: A Survey Chuang Liu, Deyi Xiong, Renren Jin, Tianhao Shen Published: 2023-09-26Area: Surveys & ReviewsCitations: 292 Tags: adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | 2023-09-26 | Surveys & Reviews | adversarial-robustness, ai-safety, alignment-training, interpretability, safety-evaluation, survey, surveys-reviews | E6 / R4 (97%) | 292 |
| Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey Adrian Weller, Ali Shahin Shamsabadi, Carolyn Ashurst, Victoria Smith Published: 2023-09-27Area: Surveys & ReviewsCitations: 41 Tags: ai-safety, survey, surveys-reviews | 2023-09-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E5 / R3 (97%) | 41 |
| Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks Erfan Shayegani, Md Abdullah Al Mamun, Nael Abu-Ghazaleh, Pedram Zaree Published: 2023-10-16Area: Surveys & ReviewsCitations: 238 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2023-10-16 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R3 (96%) | 238 |
| Sociotechnical Safety Evaluation of Generative AI Systems Arianna Manzini, Ben Bariach, Conor Griffin, Iason Gabriel Published: 2023-10-18Area: Safety EvaluationCitations: 190 Tags: ai-safety, safety-evaluation, survey | 2023-10-18 | Safety Evaluation | ai-safety, safety-evaluation, survey | E6 / R4 (96%) | 190 |
| The History and Risks of Reinforcement Learning and Human Feedback Nathan Lambert, Thomas Krendl Gilbert, Tom Zick Published: 2023-10-20Area: Alignment TrainingCitations: 50 Tags: ai-safety, alignment-training, safety-evaluation, survey | 2023-10-20 | Alignment Training | ai-safety, alignment-training, safety-evaluation, survey | E5 / R3 (96%) | 50 |
| A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking Rose Hadshar Published: 2023-10-27Area: Surveys & ReviewsCitations: 11 Tags: ai-safety, survey, surveys-reviews | 2023-10-27 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (92%) | 11 |
| AI Alignment: A Comprehensive Survey Aidan O'Gara, Borong Zhang, Boyuan Chen, Brian Tse Published: 2023-10-30Area: Surveys & ReviewsCitations: 320 Tags: ai-safety, alignment-training, interpretability, survey, surveys-reviews | 2023-10-30 | Surveys & Reviews | ai-safety, alignment-training, interpretability, survey, surveys-reviews | E7 / R4 (97%) | 320 |
| Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges Dan Qu, Hao Zhang, Heyu Chang, Nianwen Si Published: 2023-11-27Area: Model EditingCitations: 41 Tags: ai-safety, model-editing, safety-evaluation, survey | 2023-11-27 | Model Editing | ai-safety, model-editing, safety-evaluation, survey | E7 / R5 (95%) | 41 |
| A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly Jinhao Duan, Kaidi Xu, Yifan Yao, Yuanfang Cai Published: 2023-12-04Area: Surveys & ReviewsCitations: 1022 Tags: ai-safety, survey, surveys-reviews | 2023-12-04 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E6 / R3 (95%) | 1022 |
| A Comprehensive Study of Knowledge Editing for Large Language Models Bozhong Tian, Fei Huang, Huajun Chen, Jia-Chen Gu Published: 2024-01-02Area: Model EditingCitations: 134 Tags: ai-safety, model-editing, survey | 2024-01-02 | Model Editing | ai-safety, model-editing, survey | E5 / R3 (95%) | 134 |
| Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems Chuanpu Fu, Junwu Xiong, Ke Xu, Peiyang Li Published: 2024-01-11Area: Surveys & ReviewsCitations: 104 Tags: ai-safety, survey, surveys-reviews | 2024-01-11 | Surveys & Reviews | ai-safety, survey, surveys-reviews | E7 / R5 (97%) | 104 |
| Red-Teaming for Generative AI: Silver Bullet or Security Theater? Anusha Sinha, Hoda Heidari, Michael Feffer, Zachary C. Lipton Published: 2024-01-29Area: Safety EvaluationCitations: 126 Tags: ai-safety, safety-evaluation, survey | 2024-01-29 | Safety Evaluation | ai-safety, safety-evaluation, survey | E5 / R3 (94%) | 126 |
| Security and Privacy Challenges of Large Language Models: A Survey Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu Published: 2024-01-30Area: Surveys & ReviewsCitations: 351 Tags: adversarial-robustness, ai-safety, survey, surveys-reviews | 2024-01-30 | Surveys & Reviews | adversarial-robustness, ai-safety, survey, surveys-reviews | E6 / R4 (97%) | 351 |