Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 124 papers (page 1 of 5)路 255 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Analyzing Individual Neurons in Pre-trained Language Models Fahim Dalvi, Hassan Sajjad, Nadir Durrani, Yonatan Belinkov Published: 2020-10-06Area: Representation AnalysisCitations: 120 Tags: ai-safety, empirical, representation-analysis | 2020-10-06 | Representation Analysis | ai-safety, empirical, representation-analysis | E7 / R3 (96%) | 120 |
| An Interpretability Illusion for BERT Adam Pearce, Andy Coenen, Ann Yuan, Emily Reif Published: 2021-04-14Area: Representation AnalysisCitations: 96 Tags: ai-safety, empirical, interpretability, representation-analysis | 2021-04-14 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (96%) | 96 |
| Neuron-level Interpretation of Deep NLP Models: A Survey Fahim Dalvi, Hassan Sajjad, Nadir Durrani Published: 2021-08-30Area: Representation AnalysisCitations: 97 Tags: ai-safety, representation-analysis, safety-evaluation, survey | 2021-08-30 | Representation Analysis | ai-safety, representation-analysis, safety-evaluation, survey | E5 / R3 (94%) | 97 |
| On the Pitfalls of Analyzing Individual Neurons in Language Models Omer Antverg, Yonatan Belinkov Published: 2021-10-14Area: Representation AnalysisCitations: 63 Tags: ai-safety, empirical, representation-analysis | 2021-10-14 | Representation Analysis | ai-safety, empirical, representation-analysis | E7 / R3 (95%) | 63 |
| Extracting Latent Steering Vectors from Pretrained Language Models Matthew E. Peters, Nishant Subramani, Nivedita Suresh Published: 2022-05-10Area: Representation AnalysisCitations: 157 Tags: ai-safety, empirical, representation-analysis | 2022-05-10 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (95%) | 157 |
| Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task Aspen K. Hopkins, David Bau, Fernanda Vi脙漏gas, Hanspeter Pfister Published: 2022-10-24Area: Representation AnalysisCitations: 411 Tags: ai-safety, empirical, representation-analysis | 2022-10-24 | Representation Analysis | ai-safety, empirical, representation-analysis | E4 / R3 (96%) | 411 |
| Discovering Latent Knowledge in Language Models Without Supervision Collin Burns, Dan Klein, Haotian Ye, Jacob Steinhardt Published: 2022-12-07Area: Representation AnalysisCitations: 571 Tags: ai-safety, empirical, representation-analysis | 2022-12-07 | Representation Analysis | ai-safety, empirical, representation-analysis | E4 / R2 (96%) | 571 |
| Eliciting Latent Predictions from Transformers with the Tuned Lens Danny Halawi, Igor Ostrovsky, Jacob Steinhardt, Lev McKinney Published: 2023-03-14Area: Representation AnalysisCitations: 344 Tags: ai-safety, empirical, representation-analysis | 2023-03-14 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (95%) | 344 |
| The Internal State of an LLM Knows When It's Lying Amos Azaria, Tom Mitchell Published: 2023-04-26Area: Representation AnalysisCitations: 532 Tags: ai-safety, empirical, representation-analysis | 2023-04-26 | Representation Analysis | ai-safety, empirical, representation-analysis | E6 / R3 (98%) | 532 |
| Finding Neurons in a Haystack: Case Studies with Sparse Probing Dimitris Bertsimas, Dmitrii Troitskii, Katherine Harvey, Matthew Pauly Published: 2023-05-02Area: Representation AnalysisCitations: 302 Tags: ai-safety, empirical, representation-analysis | 2023-05-02 | Representation Analysis | ai-safety, empirical, representation-analysis | E6 / R3 (94%) | 302 |
| Inference-Time Intervention: Eliciting Truthful Answers from a Language Model Fernanda Vi脙漏gas, Hanspeter Pfister, Kenneth Li, Martin Wattenberg Published: 2023-06-06Area: Representation AnalysisCitations: 897 Tags: ai-safety, empirical, representation-analysis | 2023-06-06 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (96%) | 897 |
| LEACE: Perfect Linear Concept Erasure in Closed Form David Schneider-Joseph, Edward Raff, Nora Belrose, Ryan Cotterell Published: 2023-06-06Area: Representation AnalysisCitations: 175 Tags: ai-safety, empirical, representation-analysis | 2023-06-06 | Representation Analysis | ai-safety, empirical, representation-analysis | E4 / R3 (96%) | 175 |
| A Geometric Notion of Causal Probing Alexander Warstadt, Anej Svete, Cl茅ment Guerner, Ryan Cotterell Published: 2023-07-27Area: Representation AnalysisCitations: 25 Tags: ai-safety, representation-analysis, theoretical | 2023-07-27 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (92%) | 25 |
| Linearity of Relation Decoding in Transformer Language Models Arnab Sen Sharma, David Bau, Evan Hernandez, Jacob Andreas Published: 2023-08-17Area: Representation AnalysisCitations: 146 Tags: ai-safety, empirical, representation-analysis | 2023-08-17 | Representation Analysis | ai-safety, empirical, representation-analysis | E4 / R3 (90%) | 146 |
| Activation Addition: Steering Language Models Without Optimization Alexander Matt Turner, David Udell, Gavin Leech, Lisa Thiergart Published: 2023-08-20Area: Representation AnalysisCitations: 362 Tags: ai-safety, empirical, representation-analysis | 2023-08-20 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (97%) | 362 |
| Representation Engineering: A Top-Down Approach to AI Transparency Alexander Pan, Alex Mallen, Andy Zou, Ann-Kathrin Dombrowski Published: 2023-10-02Area: Representation AnalysisCitations: 780 Tags: ai-safety, empirical, representation-analysis | 2023-10-02 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (95%) | 780 |
| Language Models Represent Space and Time Max Tegmark, Wes Gurnee Published: 2023-10-03Area: Representation AnalysisCitations: 259 Tags: ai-safety, empirical, representation-analysis | 2023-10-03 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R4 (93%) | 259 |
| Interpreting CLIP's Image Representation via Text-Based Decomposition Alexei A. Efros, Jacob Steinhardt, Yossi Gandelsman Published: 2023-10-09Area: Representation AnalysisCitations: 159 Tags: ai-safety, empirical, representation-analysis | 2023-10-09 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (95%) | 159 |
| The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets Max Tegmark, Samuel Marks Published: 2023-10-10Area: Representation AnalysisCitations: 402 Tags: ai-safety, empirical, representation-analysis | 2023-10-10 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (97%) | 402 |
| The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models Avi Caciularu, Aviv Slobodkin, Ido Dagan, Omer Goldman Published: 2023-10-18Area: Representation AnalysisCitations: 52 Tags: ai-safety, empirical, representation-analysis | 2023-10-18 | Representation Analysis | ai-safety, empirical, representation-analysis | E7 / R3 (95%) | 52 |
| Linear Representations of Sentiment in Large Language Models Atticus Geiger, Curt Tigges, Neel Nanda, Oskar John Hollinsworth Published: 2023-10-23Area: Representation AnalysisCitations: 131 Tags: ai-safety, empirical, representation-analysis | 2023-10-23 | Representation Analysis | ai-safety, empirical, representation-analysis | E6 / R3 (95%) | 131 |
| Comparing Optimization Targets for Contrast-Consistent Search Hugo Fry, Ian Fan, Jamie Wright, Nandi Schoots Published: 2023-11-01Area: Representation AnalysisCitations: 4 Tags: ai-safety, empirical, representation-analysis | 2023-11-01 | Representation Analysis | ai-safety, empirical, representation-analysis | E4 / R2 (94%) | 4 |
| The Linear Representation Hypothesis and the Geometry of Large Language Models Kiho Park, Victor Veitch, Yo Joong Choe Published: 2023-11-07Area: Representation AnalysisCitations: 371 Tags: ai-safety, representation-analysis, theoretical | 2023-11-07 | Representation Analysis | ai-safety, representation-analysis, theoretical | E5 / R3 (95%) | 371 |
| Identifying Linear Relational Concepts in Large Language Models Anthony Hunter, David Chanin, Oana-Maria Camburu Published: 2023-11-15Area: Representation AnalysisCitations: 8 Tags: ai-safety, empirical, representation-analysis | 2023-11-15 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (94%) | 8 |
| Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness? Dylan Hadfield-Menell, Jacob Andreas, Kevin Liu, Stephen Casper Published: 2023-11-27Area: Representation AnalysisCitations: 55 Tags: ai-safety, empirical, representation-analysis | 2023-11-27 | Representation Analysis | ai-safety, empirical, representation-analysis | E6 / R3 (95%) | 55 |
| Eliciting Latent Knowledge from Quirky Language Models Alex Mallen, Nora Belrose Published: 2023-12-02Area: Representation AnalysisCitations: 46 Tags: ai-safety, empirical, representation-analysis | 2023-12-02 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (93%) | 46 |
| Improving Activation Steering in Language Models with Mean-Centring Dylan Cope, Murray Shanahan, Nandi Schoots, Ole Jorgensen Published: 2023-12-06Area: Representation AnalysisCitations: 57 Tags: ai-safety, empirical, representation-analysis | 2023-12-06 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (94%) | 57 |
| Steering Llama 2 via Contrastive Activation Addition Alexander Matt Turner, Evan Hubinger, Julian Schulz, Meg Tong Published: 2023-12-09Area: Representation AnalysisCitations: 527 Tags: ai-safety, empirical, representation-analysis | 2023-12-09 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (96%) | 527 |
| Challenges with Unsupervised LLM Knowledge Discovery Johannes Gasteiger, Rohin Shah, Sebastian Farquhar, Vikrant Varma Published: 2023-12-15Area: Representation AnalysisCitations: 35 Tags: ai-safety, empirical, representation-analysis | 2023-12-15 | Representation Analysis | ai-safety, empirical, representation-analysis | E5 / R3 (98%) | 35 |
| Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models Adam Pearce, Asma Ghandeharioun, Avi Caciularu, Lucas Dixon Published: 2024-01-11Area: Representation AnalysisCitations: 174 Tags: ai-safety, empirical, interpretability, representation-analysis | 2024-01-11 | Representation Analysis | ai-safety, empirical, interpretability, representation-analysis | E5 / R3 (95%) | 174 |