Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| ATANT: An Evaluation Framework for AI Continuity Samuel Sameer Tanguturi Published: 2026-04-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (97%) | - |
| Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation Xiang Zheng, Xingjun Ma, Yutao Wu, Zonghuan Xu Published: 2026-04-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (94%) | - |
| FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios Alex Xue, Chao Zhang, Chengyu Tao, Dacheng Tao Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R4 (97%) | - |
| FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios Alex Xue, Chao Zhang, Chengyu Tao, Dacheng Tao Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R4 (97%) | - |
| GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Hwee Tou Ng, Kevin Qinghong Lin, Mike Zheng Shou, Mingyu Ouyang Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (98%) | - |
| Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Andr茅 F. T. Martins, Jos茅 Pombal, Ricardo Rei Published: 2026-04-08Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (97%) | - |
| ACE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Chaoda Song, Chuang Ma, Debargha Ganguly, Shouren Wang Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (95%) | - |
| AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Chaoda Song, Chuang Ma, Debargha Ganguly, Shouren Wang Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Beyond Behavior: Why AI Evaluation Needs a Cognitive Revolution Amir Konigsberg Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (92%) | - |
| CAKE: Cloud Architecture Knowledge Evaluation of Large Language Models Florian Girardo Lukas, Krzysztof Sierszecki, Phongsakon Mark Konrad, Rahime Yilmaz Published: 2026-04-07Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-07 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E6 / R4 (96%) | - |
| Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents Bowen Ye, Chenxin An, Hanglong Lv, Lei Li Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization Alperen Kantarci, Dustin Eisenhardt, Gemma Roig, Timothy Schauml枚ffel Published: 2026-04-07Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-07 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E7 / R3 (97%) | - |
| LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency David Simchi-Levi, Jiachun Li, Will Wei Sun Published: 2026-04-07Area: stat.MECitations: - Tags: ai-safety, preprint, safety-evaluation, statme | 2026-04-07 | stat.ME | ai-safety, preprint, safety-evaluation, statme | E5 / R3 (94%) | - |
| Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles Ben Wigler, Maria Tsfasman, Tiffany Matej Hrkalovic Published: 2026-04-07Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-07 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (96%) | - |
| The Deployment Gap in AI Media Detection: Platform-Aware and Visually Constrained Adversarial Evaluation Aishwarya Budhkar, Siddhesh Sheth, Trishita Dhara Published: 2026-04-07Area: cs.CVCitations: - Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-04-07 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (93%) | - |
| Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar, Shammur Absar Chowdhury Published: 2026-04-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (97%) | - |
| GUIDE: Interpretable GUI Agent Evaluation via Hierarchical Diagnosis Benlei Cui, Bo Xu, Liang Wang, Liwu Xu Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (97%) | - |
| IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents Rongqian Chen, Sizhe Tang, Tian Lan, Weidong Cao Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation Alhasan Mahmood, Hasan Kurban, Samir Abdaljalil Published: 2026-04-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R4 (96%) | - |
| RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets Andrew Borthwick, Anthony Galczak, Stephen Ash Published: 2026-04-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |
| RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains Carter Ung, Christopher Tan, Dieter Fox, Evan Gubarev Published: 2026-04-06Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-04-06 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges Hao Liu, Rong Lu, Song Hou Published: 2026-04-05Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-04-05 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode Shuai Wang, Wenyuan Jiang, Yudong Gao, Zimo Ji Published: 2026-04-04Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-04 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge Dmitry Fedrushkov, Ilya Revin, Ivan Smirnov, Sergey Kovalchuk Published: 2026-04-04Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-04 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (98%) | - |
| An Independent Safety Evaluation of Kimi K2.5 Aengus Lynch, Andy Wang, Dennis Murphy, Elle Najt Published: 2026-04-03Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-03 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R3 (95%) | - |
| A Systematic Security Evaluation of OpenClaw and Its Variants Haichang Gao, Shiguo Lian, Wenjing Zhang, Xiang Wang Published: 2026-04-03Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-03 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E7 / R5 (98%) | - |
| VERT: Reliable LLM Judges for Radiology Report Evaluation Asma Ben Abacha, Federica Bologna, Jean-Philippe Corbeil, Matthew Wilkens Published: 2026-04-03Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-03 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |
| Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study Atsushi Takamatsu, Osamu Abe, Shouhei Hanaoka, Takeharu Yoshikawa Published: 2026-04-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy Peiyao Fu, Pinghong Zhou, Quanlin Li, Ruijie Yang Published: 2026-04-02Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-02 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R2 (96%) | - |
| Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection Kashinath Kadaba Shrish, Siddharth Saxena, Subho Halder, Thiyagarajan M Published: 2026-04-02Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-02 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E5 / R3 (96%) | - |