Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| PACE: A Proxy for Agentic Capability Evaluation Aditya Bharat Soni, Daniel Lee, Graham Neubig, Jiarui Liu Published: 2026-07-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E38 / R32 (91%) | - |
| PACE: A Proxy for Agentic Capability Evaluation Aditya Bharat Soni, Daniel Lee, Graham Neubig, Jiarui Liu Published: 2026-07-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions Dipo Dunsin, Ed de Quincey, Eduardo Almeida Palmieri, Kim-Kwang Raymond Choo Published: 2026-07-03Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-03 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E8 / R8 (94%) | - |
| CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI Roopam W. Sure Published: 2026-07-03Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-03 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E9 / R8 (93%) | - |
| The Foreign Policy AI Evaluation Gap Charles Pozniak, Jeba Sania Published: 2026-07-03Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-07-03 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E12 / R16 (95%) | - |
| TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation Axel TahmasebiMoradi, Lucas Schott, Martin Royer Published: 2026-07-03Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-07-03 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E10 / R8 (93%) | - |
| A Unified Algebraic Framework for Classification Performance Evaluation Ronaldo C. Prati Published: 2026-07-04Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-04 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R8 (90%) | - |
| ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation Enshuo Hsu, Jin Zhou, Kirk Roberts Published: 2026-07-04Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-04 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R11 (94%) | - |
| How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation Ehsan Javanmardi, Fardin Ayar, Mahdi Javanmardi, Manabu Tsukada Published: 2026-07-04Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-04 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E10 / R9 (95%) | - |
| evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations Shreyas K Chandrahas Published: 2026-07-05Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-05 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R12 (92%) | - |
| Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese) Ken Ito, Shigekazu Ishihara, Yusei Tamura Published: 2026-07-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-07-05 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E8 / R8 (93%) | - |
| RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Baijun Chen, Dandan Zhang, Enze Xie, Guanyu Lin Published: 2026-07-05Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-07-05 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E22 / R27 (94%) | - |
| RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Baijun Chen, Dandan Zhang, Enze Xie, Guanyu Lin Published: 2026-07-05Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-07-05 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E21 / R21 (91%) | - |
| EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Danti Chen, Josh Fleischer, Kenneth Benavides Published: 2026-07-06Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-06 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E20 / R19 (94%) | - |
| Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation Anthony Marchiafava, Arefa Patwary, Atriya Sen, Sadia Kamal Published: 2026-07-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E15 / R8 (94%) | - |
| SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models Ashish Hallur, Georgi Tinchev, Hao Zhang, Laureano Moro-Velazquez Published: 2026-07-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R9 (94%) | - |
| SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Carsten Maple, Linxi Li, Liwei Jin, Qianwei Guo Published: 2026-07-06Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-07-06 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E13 / R13 (94%) | - |
| Three-Phase Evaluation of AI-Assisted Software Development Life Cycle Carson Crockett, Jacob Viehe, Jason Ferraro, Joshua Strubel Published: 2026-07-06Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-06 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R15 (91%) | - |
| AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation Andrey Podivilov, Maksim Parshin, Matvei Startsev, Roman Pozharskiy Published: 2026-07-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R14 (93%) | - |
| Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation Niels Potters, Theo Hofman Published: 2026-07-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R6 (94%) | - |
| Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning Akshay Arora, Ashutosh Aggarwal, Ishan Nigam, Krishna Singh Published: 2026-07-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (90%) | - |
| Data-dependent Evaluations for Budgeted Submodular Maximization Jing Tang, Lejian Zhang, Xueyan Tang Published: 2026-07-07Area: cs.DSCitations: - Tags: ai-safety, csds, preprint, safety-evaluation | 2026-07-07 | cs.DS | ai-safety, csds, preprint, safety-evaluation | E10 / R12 (94%) | - |
| Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation Bin Li, Kan Li, Mingyang Zhao, Shenxi Liu Published: 2026-07-07Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-07 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E17 / R18 (93%) | - |
| PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages Abrorkhon Inomkhujaev, Alexander Fraser, Antonia Karamolegkou, Daryna Dementieva Published: 2026-07-07Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-07 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R7 (91%) | - |
| Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development Adam P. Burden, Kapil Singi, Majd Sakr, Rohit Mehra Published: 2026-07-07Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-07 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E12 / R11 (92%) | - |
| Reliable and Developer-Aligned Evaluation of Agents for Software Engineering Razvan Mihai Popescu Published: 2026-07-07Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-07 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E7 / R6 (89%) | - |
| Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System Junde Lu, Xuefei Huang, Yiming Gai Published: 2026-07-08Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R7 (90%) | - |
| Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Ezgi Korkmaz Published: 2026-07-08Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-08 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E12 / R8 (94%) | - |
| Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations Silvia Santano Published: 2026-07-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E17 / R15 (92%) | - |
| Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation Alejandro Vergara-Richart, Almudena Fuster-Matanzo, Ana Jim茅nez-Pastor, 脕ngel Alberich-Bayarri Published: 2026-07-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E10 / R9 (93%) | - |