Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| SHOE: Semantic HOI Open-Vocabulary Evaluation Metric Bihan Dong, Bo Wang, John Young, Maja Noack Published: 2026-04-02Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-02 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks Adam Kuzee, Brittany S. Harris, Harry Lyu, Jonathan Rosenfeld Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R2 (95%) | - |
| Logarithmic Scores, Power-Law Discoveries: Disentangling Measurement from Coverage in Agent-Based Evaluation HyunJoon Jung, William Na Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (94%) | - |
| MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy Anna Jung, Eunseob Choi, Hyeonseok Jung, Hyuk-Jae Lee Published: 2026-04-01Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-01 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E6 / R4 (98%) | - |
| Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers Atsuyuki Miyai, Kenta Watanabe, Kiyoharu Aizawa, Mashiro Toyooka Published: 2026-04-01Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-01 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R3 (95%) | - |
| Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering Andr茅 Storhaug, Jingyue Li Published: 2026-04-01Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-01 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R2 (98%) | - |
| UK AISI Alignment Evaluation Case-Study Abby D'Cruz, Alexandra Souly, Jacob Merizian, Robert Kirk Published: 2026-04-01Area: cs.AICitations: - Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-04-01 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | E6 / R3 (96%) | - |
| Multi-Layered Memory Architectures for LLM Agents: An Experimental Evaluation of Long-Term Context Retention Payal Fofadiya, Sunil Tiwari Published: 2026-03-31Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-31 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R2 (95%) | - |
| Reasoning-Driven Synthetic Data Generation and Evaluation Benoit Seguin, Cesar Ilharco, Enrico Bacis, Hamza Harkous Published: 2026-03-31Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-31 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| Semantic Interaction for Narrative Map Sensemaking: An Insight-based Evaluation Brian Felipe Keith-Norambuena, Chris North, Eric Krokos, Fausto German Published: 2026-03-31Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-31 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E5 / R3 (95%) | - |
| SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation Haruto Yoshida, Jun Suzuki, Keito Kudo, Nobuyuki Shimizu Published: 2026-03-31Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-31 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E4 / R3 (97%) | - |
| Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild Amal Raj, Deepak Akkil, Mowafak Allaham, Ravi Kokku Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Reward Hacking as Equilibrium under Finite Evaluation Jiacheng Wang, Jinbin Huang Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation Doan Nam Long Vu, Simone Balloccu Published: 2026-03-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks Chak-Wing Mak, Chiafeng Chu, Chiao-Wei Hsu, Chi Ruan Published: 2026-03-29Area: cs.GRCitations: - Tags: ai-safety, csgr, preprint, safety-evaluation | 2026-03-29 | cs.GR | ai-safety, csgr, preprint, safety-evaluation | E5 / R3 (92%) | - |
| Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness Phat Nguyen, Thang Pham Published: 2026-03-29Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-03-29 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E6 / R3 (91%) | - |
| LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications Alexandre Cristov茫o Maiorano Published: 2026-03-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (98%) | - |
| Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis Axel Marmoret Published: 2026-03-28Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-28 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E6 / R3 (95%) | - |
| ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation Chen Tian, Haonan Sun, Lifei Rao, Qincheng Zhang Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-27 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E4 / R2 (95%) | - |
| Generative Modeling in Protein Design: Neural Representations, Conditional Generation, and Evaluation Standards Ken-Tye Yong, Minh-Duong Nguyen, Nguyen H. Tran, Senura Hansaja Wanasekara Published: 2026-03-27Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-27 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles Dan Marshall, Denae Ford, Edward Cutrell, Ruoxi Shang Published: 2026-03-27Area: cs.HCCitations: - Tags: ai-safety, alignment-training, cshc, preprint, safety-evaluation | 2026-03-27 | cs.HC | ai-safety, alignment-training, cshc, preprint, safety-evaluation | E4 / R3 (96%) | - |
| Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI Anna Kozlova, Hanna Plotnitskaya, Pavel Satalkin, Sergey Parfenyuk Published: 2026-03-26Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R3 (96%) | - |
| Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding Chao Zhang, Gregor Baer, Isel Grau, Pieter Van Gorp Published: 2026-03-26Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-26 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E5 / R3 (94%) | - |
| MolQuest: A Benchmark for Agentic Evaluation of Abductive Reasoning in Chemical Structure Elucidation Bing Zhao, Jinghang Wang, Renquan Lv, Shuang Wu Published: 2026-03-26Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation Chanyoung Park, Jayakumar Subramanian, Mehrab Tanjim, Sangwu Park Published: 2026-03-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (96%) | - |
| RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following Bo Xu, Hongwei Feng, Licai Qi, Qianyu He Published: 2026-03-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (94%) | - |
| Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search Ariel Felner, Carlos Hernandez, Hadar Peer, Oren Salzman Published: 2026-03-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R5 (98%) | - |
| Gaze patterns predict preference and confidence in pairwise AI image evaluation Ankur Samanta, Nikolas Papadopoulos, Paul Sajda, Sheng Bai Published: 2026-03-25Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-25 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E6 / R3 (97%) | - |
| Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation Andrea Generosi, Daniele Agostinelli, Maura Mengoni, Thomas Agostinelli Published: 2026-03-25Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-25 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E7 / R3 (97%) | - |
| NeuroVLM-Bench: Evaluation of Vision-Enabled Large Language Models for Clinical Reasoning in Neurological Disorders Ilinka Ivanoska, Ivan Kitanovski, Katarina Trojachanec Dineva, Kostadin Mishev Published: 2026-03-25Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-25 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (96%) | - |