Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks Chak-Wing Mak, Chiafeng Chu, Chiao-Wei Hsu, Chi Ruan Published: 2026-03-29Area: cs.GRCitations: - Tags: ai-safety, csgr, preprint, safety-evaluation | 2026-03-29 | cs.GR | ai-safety, csgr, preprint, safety-evaluation | E5 / R3 (92%) | - |
| Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness Phat Nguyen, Thang Pham Published: 2026-03-29Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-03-29 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E6 / R3 (91%) | - |
| Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild Amal Raj, Deepak Akkil, Mowafak Allaham, Ravi Kokku Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Reward Hacking as Equilibrium under Finite Evaluation Jiacheng Wang, Jinbin Huang Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation Doan Nam Long Vu, Simone Balloccu Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Multi-Layered Memory Architectures for LLM Agents: An Experimental Evaluation of Long-Term Context Retention Payal Fofadiya, Sunil Tiwari Published: 2026-03-31Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-31 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R2 (95%) | - |
| Reasoning-Driven Synthetic Data Generation and Evaluation Benoit Seguin, Cesar Ilharco, Enrico Bacis, Hamza Harkous Published: 2026-03-31Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-31 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| Semantic Interaction for Narrative Map Sensemaking: An Insight-based Evaluation Brian Felipe Keith-Norambuena, Chris North, Eric Krokos, Fausto German Published: 2026-03-31Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-31 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E5 / R3 (95%) | - |
| SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation Haruto Yoshida, Jun Suzuki, Keito Kudo, Nobuyuki Shimizu Published: 2026-03-31Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-31 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E4 / R3 (97%) | - |
| Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks Adam Kuzee, Brittany S. Harris, Harry Lyu, Jonathan Rosenfeld Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R2 (95%) | - |
| Logarithmic Scores, Power-Law Discoveries: Disentangling Measurement from Coverage in Agent-Based Evaluation HyunJoon Jung, William Na Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (94%) | - |
| MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy Anna Jung, Eunseob Choi, Hyeonseok Jung, Hyuk-Jae Lee Published: 2026-04-01Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-01 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E6 / R4 (98%) | - |
| Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers Atsuyuki Miyai, Kenta Watanabe, Kiyoharu Aizawa, Mashiro Toyooka Published: 2026-04-01Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-01 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R3 (95%) | - |
| Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering Andr茅 Storhaug, Jingyue Li Published: 2026-04-01Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-01 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R2 (98%) | - |
| UK AISI Alignment Evaluation Case-Study Abby D'Cruz, Alexandra Souly, Jacob Merizian, Robert Kirk Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | E6 / R3 (96%) | - |
| Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study Atsushi Takamatsu, Osamu Abe, Shouhei Hanaoka, Takeharu Yoshikawa Published: 2026-04-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy Peiyao Fu, Pinghong Zhou, Quanlin Li, Ruijie Yang Published: 2026-04-02Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-02 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R2 (96%) | - |
| Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection Kashinath Kadaba Shrish, Siddharth Saxena, Subho Halder, Thiyagarajan M Published: 2026-04-02Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-02 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E5 / R3 (96%) | - |
| SHOE: Semantic HOI Open-Vocabulary Evaluation Metric Bihan Dong, Bo Wang, John Young, Maja Noack Published: 2026-04-02Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-02 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (96%) | - |
| An Independent Safety Evaluation of Kimi K2.5 Aengus Lynch, Andy Wang, Dennis Murphy, Elle Najt Published: 2026-04-03Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-03 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R3 (95%) | - |
| A Systematic Security Evaluation of OpenClaw and Its Variants Haichang Gao, Shiguo Lian, Wenjing Zhang, Xiang Wang Published: 2026-04-03Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-03 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E7 / R5 (98%) | - |
| VERT: Reliable LLM Judges for Radiology Report Evaluation Asma Ben Abacha, Federica Bologna, Jean-Philippe Corbeil, Matthew Wilkens Published: 2026-04-03Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-03 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |
| Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode Shuai Wang, Wenyuan Jiang, Yudong Gao, Zimo Ji Published: 2026-04-04Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-04 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge Dmitry Fedrushkov, Ilya Revin, Ivan Smirnov, Sergey Kovalchuk Published: 2026-04-04Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-04 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (98%) | - |
| Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges Hao Liu, Rong Lu, Song Hou Published: 2026-04-05Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-04-05 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar, Shammur Absar Chowdhury Published: 2026-04-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (97%) | - |
| GUIDE: Interpretable GUI Agent Evaluation via Hierarchical Diagnosis Benlei Cui, Bo Xu, Liang Wang, Liwu Xu Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (97%) | - |
| IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents Rongqian Chen, Sizhe Tang, Tian Lan, Weidong Cao Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation Alhasan Mahmood, Hasan Kurban, Samir Abdaljalil Published: 2026-04-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R4 (96%) | - |
| RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets Andrew Borthwick, Anthony Galczak, Stephen Ash Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |