Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation Beidi Luan, Dezhi Chen, Jing Li, Mengting Chen Published: 2026-07-31Area: cs.CLCitations: 1 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-31 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R4 (92%) | 1 |
| MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft Beini Hu, Jianxin Gao, Jinyuan Zhang, Linna Deng Published: 2026-07-31Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-31 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R7 (91%) | 45 |
| ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models Jungang Xu, Penglin Zhu Published: 2026-07-31Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-31 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E14 / R11 (92%) | - |
| AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge Muhammad Sajjad Akbar Published: 2026-07-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R8 (92%) | - |
| Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs Minwoo Yu, Young-guk Ha Published: 2026-07-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (93%) | - |
| Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation Jordan Sassoon, Philipp D. Siedler Published: 2026-07-30Area: cs.CLCitations: 1 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E16 / R15 (94%) | 1 |
| Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Benfeng Xu, Hongtao Xie, Jie Gao, Lingyun Yu Published: 2026-07-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R7 (90%) | - |
| Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Benfeng Xu, Hongtao Xie, Jie Gao, Lingyun Yu Published: 2026-07-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (90%) | - |
| Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories Jun Wang, Wanyu Si, Zhaoji Wang Published: 2026-07-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R5 (92%) | - |
| From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis Ademar Fran莽a, Angelo Perkusich, Danyllo Albuquerque, Emanuel Dantas Published: 2026-07-30Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-30 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E9 / R7 (91%) | - |
| OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Ben Kao, Bowen Yang, Fangzhi Xu, Hang Yan Published: 2026-07-30Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (92%) | - |
| Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation Fazhong Liu, Guoxing Chen, Haojin Zhu, Haozhen Tan Published: 2026-07-30Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-30 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E40 / R27 (88%) | - |
| Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation Weining Zhang Published: 2026-07-30Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-30 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (87%) | 15 |
| SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models Bera Koca, Dania Daye, Ebubechukwu D Enwerem, Emine Meltem Published: 2026-07-30Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-30 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E14 / R14 (95%) | - |
| BayesAME: Bayesian Active Model Evaluation Arnaud Doucet, Paula Cordero Encinar, Silvia Chiappa, Taylan Cemgil Published: 2026-07-29Area: cs.LGCitations: 15 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-29 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (91%) | 15 |
| Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance Chao Peng, Hande Dong, Peijie Dong, Qiang Lin Published: 2026-07-29Area: cs.LGCitations: 25 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-29 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R6 (90%) | 25 |
| LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving Faaiq Waqar, Hanchen Yang, Harsono Simka, Ming-Yen Lee Published: 2026-07-29Area: cs.ARCitations: - Tags: ai-safety, csar, preprint, safety-evaluation | 2026-07-29 | cs.AR | ai-safety, csar, preprint, safety-evaluation | E10 / R8 (89%) | - |
| Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models Cheryl Seals, Gerry Dozier, Parishruthi Ganesh Published: 2026-07-29Area: cs.CLCitations: 35 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-29 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E21 / R8 (89%) | 35 |
| Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Abu Bakar Siddik Published: 2026-07-28Area: cs.AICitations: 97 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (92%) | 97 |
| Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models David Chushig-Muzo, Eva Milara, Felipe Grijalva, Luis Bote-Curiel Published: 2026-07-28Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-28 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | - | - |
| Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification Chenrui Shi, Che Sun, Lifeng Fan, Ruining Feng Published: 2026-07-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification Chenrui Shi, Che Sun, Lifeng Fan, Ruining Feng Published: 2026-07-28Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R8 (91%) | 15 |
| Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Alexandre Sallinen, Charlotte Meyer, Guillaume Allegre, Stefan Krsteski Published: 2026-07-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models Deepanshu Mody, Dipesh Mahato, Samarth Agarwal, Utkarsh Mittal Published: 2026-07-28Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-28 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | - | - |
| Position: Evaluation Scores Are Perishable Knowledge Claims Sankalp Gilda, Shlok Gilda Published: 2026-07-28Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (92%) | 45 |
| When benchmark inferences do not compose: Projectibility in AI evaluation Brett Reynolds Published: 2026-07-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (93%) | - |
| A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility Dariusz Jarosz, Hairong Shang, Mathew J. Cherukara, Michael D. Borland Published: 2026-07-27Area: physics.acc-phCitations: - Tags: ai-safety, physicsacc-ph, preprint, safety-evaluation | 2026-07-27 | physics.acc-ph | ai-safety, physicsacc-ph, preprint, safety-evaluation | - | - |
| CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models Dengzhe Hou, Fangzhou Lin, Kazunori D Yamada, Lingyu Jiang Published: 2026-07-27Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E15 / R16 (91%) | - |
| Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim Christian Becker-Asano, Marcel Heisler, Steve Aschenbrenner, Thomas Sievers Published: 2026-07-27Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-07-27 | cs.RO | ai-safety, csro, preprint, safety-evaluation | - | - |
| Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation Da-Tian Peng, Jingkun Luo Published: 2026-07-27Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-27 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |