Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains Carter Ung, Christopher Tan, Dieter Fox, Evan Gubarev Published: 2026-04-06Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-04-06 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E5 / R3 (95%) | - |
| ACE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Chaoda Song, Chuang Ma, Debargha Ganguly, Shouren Wang Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (95%) | - |
| AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Chaoda Song, Chuang Ma, Debargha Ganguly, Shouren Wang Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Beyond Behavior: Why AI Evaluation Needs a Cognitive Revolution Amir Konigsberg Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (92%) | - |
| CAKE: Cloud Architecture Knowledge Evaluation of Large Language Models Florian Girardo Lukas, Krzysztof Sierszecki, Phongsakon Mark Konrad, Rahime Yilmaz Published: 2026-04-07Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-07 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E6 / R4 (96%) | - |
| Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents Bowen Ye, Chenxin An, Hanglong Lv, Lei Li Published: 2026-04-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization Alperen Kantarci, Dustin Eisenhardt, Gemma Roig, Timothy Schauml枚ffel Published: 2026-04-07Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-07 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E7 / R3 (97%) | - |
| LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency David Simchi-Levi, Jiachun Li, Will Wei Sun Published: 2026-04-07Area: stat.MECitations: - Tags: ai-safety, preprint, safety-evaluation, statme | 2026-04-07 | stat.ME | ai-safety, preprint, safety-evaluation, statme | E5 / R3 (94%) | - |
| Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles Ben Wigler, Maria Tsfasman, Tiffany Matej Hrkalovic Published: 2026-04-07Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-07 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (96%) | - |
| The Deployment Gap in AI Media Detection: Platform-Aware and Visually Constrained Adversarial Evaluation Aishwarya Budhkar, Siddhesh Sheth, Trishita Dhara Published: 2026-04-07Area: cs.CVCitations: - Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-04-07 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (93%) | - |
| ATANT: An Evaluation Framework for AI Continuity Samuel Sameer Tanguturi Published: 2026-04-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (97%) | - |
| Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation Xiang Zheng, Xingjun Ma, Yutao Wu, Zonghuan Xu Published: 2026-04-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (94%) | - |
| FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios Alex Xue, Chao Zhang, Chengyu Tao, Dacheng Tao Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R4 (97%) | - |
| FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios Alex Xue, Chao Zhang, Chengyu Tao, Dacheng Tao Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R4 (97%) | - |
| GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Hwee Tou Ng, Kevin Qinghong Lin, Mike Zheng Shou, Mingyu Ouyang Published: 2026-04-08Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-08 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (98%) | - |
| Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Andr茅 F. T. Martins, Jos茅 Pombal, Ricardo Rei Published: 2026-04-08Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (97%) | - |
| ACIArena: Toward Unified Evaluation for Agent Cascading Injection Changjiang Li, Chunyi Zhou, Hengyu An, Jinghuai Zhang Published: 2026-04-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (97%) | - |
| An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks Adrian-Marius Dumitran, Gabriel Stefan Published: 2026-04-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |
| AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan Guangtao Zhai, Haonan Cheng, Hengyan Huang, Jian Liu Published: 2026-04-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-04-09 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (96%) | - |
| AtomEval: Atomic Evaluation of Adversarial Claims in Fact Verification Hanze Jia, Hongyi Cen, Jingyi Zheng, Mingxin Wang Published: 2026-04-09Area: cs.CLCitations: - Tags: adversarial-robustness, ai-safety, cscl, preprint, safety-evaluation | 2026-04-09 | cs.CL | adversarial-robustness, ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (97%) | - |
| AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Chong Luo, Lili Qiu, Qi Dai, Rui Wang Published: 2026-04-09Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-09 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E6 / R3 (94%) | - |
| Can Vision Language Models Judge Action Quality? An Empirical Evaluation Miguel Monte e Freitas, Pedro Henrique Martins, Ricardo Rei, Rui Henriques Published: 2026-04-09Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-09 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E8 / R3 (98%) | - |
| CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V Can Gurkan, John Chen, Mingyi Lin, Sihan Cheng Published: 2026-04-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (93%) | - |
| KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation Fei Tang, Guocheng Shao, Jun Xiao, Kaitao Song Published: 2026-04-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (97%) | - |
| On Semiotic-Grounded Interpretive Evaluation of Generative Art Changwen Chen, Ruixiang Jiang Published: 2026-04-09Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-09 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (95%) | - |
| PIArena: A Platform for Prompt Injection Evaluation Chenlong Yin, Jinyuan Jia, Runpeng Geng, Yanting Wang Published: 2026-04-09Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-09 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E4 / R3 (95%) | - |
| BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation C茅line Hudelot, Emmanuel Malherbe, Hippolyte Gisserot-Boukhlef, Nicolas Boizard Published: 2026-04-10Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-10 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (93%) | - |
| Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Kai Yu, Peng Wang, Qinyuan Chen, Wupeng Wang Published: 2026-04-10Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-10 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R3 (97%) | - |
| Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Kai Yu, Peng Wang, Qinyuan Chen, Wupeng Wang Published: 2026-04-10Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-10 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models Avni Mittal, Monojit Choudhury, Sandipan Dandapat, Shanu Kumar Published: 2026-04-10Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-10 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (93%) | - |