Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation Carlos S谩ez, David Fern谩ndez-Narro, Marc P茅rez-Roig Published: 2026-08-17Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R7 (93%) | 15 |
| Position: Fairness Failure in Generative Models is an Evaluation Problem Jean-Yves Franceschi, Mariia Vladimirova, Thibaut Issenhuth Published: 2026-08-17Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-17 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E2 / R6 (91%) | - |
| Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback Ashmita Kapoor, Duncan Cambridge, Michael Zimmermann, Pouya Ghiasnezhad Omran Published: 2026-08-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (93%) | - |
| TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation Adrita Anika, Md Messal Monem Miah, Ruihong Huang, Zhiyuan Yu Published: 2026-08-16Area: cs.AICitations: - Tags: adversarial-robustness, ai-safety, csai, preprint, safety-evaluation | 2026-08-16 | cs.AI | adversarial-robustness, ai-safety, csai, preprint, safety-evaluation | E15 / R14 (93%) | - |
| An Evaluation Framework for National AI Regulation Amal Dhivyan Gregory, Avyay M Casheekar, Kaushik Sanjay Prabhakar, Sreeparvathy Sajeev Published: 2026-08-15Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-08-15 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E13 / R8 (93%) | - |
| Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation M P V S Gopinadh Published: 2026-08-15Area: cs.CLCitations: 12 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-15 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R6 (94%) | 12 |
| The Benchmark Trap: Structures of Power and Injustice in AI Evaluations Angelie Kraft, Jason Branford Published: 2026-08-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E19 / R11 (93%) | - |
| Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark Vasiliy Seibert Published: 2026-08-15Area: cs.AICitations: 35 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R10 (95%) | 35 |
| Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Bin Xu, Bolong Feng, Chunhua Shen, Feihan Chen Published: 2026-08-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E11 / R10 (90%) | - |
| Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Bin Xu, Bolong Feng, Chunhua Shen, Feihan Chen Published: 2026-08-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R9 (90%) | - |
| Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations Toby D. Pilditch Published: 2026-08-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation Avyay M. Casheekar Published: 2026-08-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (92%) | - |
| Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Borun Chen, Fei Sun, Hao Tian, Hexiang Tan Published: 2026-08-13Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E12 / R12 (91%) | - |
| Coverage Aware Active Evaluation for Failure Discovery with Paired Systems Anjali Parashar, Apoorva Sharma, Carson Sobolewski, Chuchu Fan Published: 2026-08-13Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R8 (91%) | - |
| Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes Aimilios Hadjiliasi, Louis Nisiotis Published: 2026-08-13Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (95%) | - |
| How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures Ananya Mukherjee, Christian Greisinger, Owusu-Banahene Osei, Paul Osemudiame Oamen Published: 2026-08-13Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-13 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (95%) | 15 |
| Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation Rana Muhammad Ahmed, Sabahat Abbas Published: 2026-08-13Area: cs.CRCitations: 45 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-08-13 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E15 / R8 (91%) | 45 |
| UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations Peng Li, Qianqian Xu, Qingming Huang, Shilong Bao Published: 2026-08-13Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-13 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E9 / R9 (90%) | - |
| Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework Avinash Agarwal, Vridhi Jain Published: 2026-08-12Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-08-12 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E10 / R8 (94%) | - |
| Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework Avinash Agarwal, Vridhi Jain Published: 2026-08-12Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-08-12 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | - | - |
| Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Hao Zhang, Jianqiang Huang, Jiaxin Qi, Zhijiang Tang Published: 2026-08-12Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-12 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | - | - |
| Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges Jie Mu, Mo Xuan, Qun Shao, Xi Chen Published: 2026-08-12Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-12 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation Oded Vainas, Ofir Ben Shoham, Shravan Mohan, Shrutendra Harsola Published: 2026-08-12Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-12 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | - | - |
| Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation Alison R. Panisson, Rodrigo Guedes de Souza Published: 2026-08-12Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-12 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen, Amina Ibrahim Khaleel, Fatima Isa Jibrin Published: 2026-08-11Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-11 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R9 (95%) | - |
| Benchmarking LLM Judges for Mobile Agent Evaluation Li Gu, Seyed Mehdi Ayyoubzadeh, Yang Wang, Yuanhao Yu Published: 2026-08-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Vasundra Srinivasan Published: 2026-08-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes Bowen Jiang, Guorong Li, Jianbin Jiao, Jiashu Li Published: 2026-08-11Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-11 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E14 / R10 (91%) | - |
| From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation Alireza S. Ziabari, Colleen Chan, Ding Tong, Kat Ellis Published: 2026-08-11Area: cs.AICitations: - Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-08-11 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | - | - |
| Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation Gianluca Demartini, Joel Mackenzie, Pietro Bernardelle, Samaneh Mohtadi Published: 2026-08-11Area: cs.IRCitations: 59 Tags: ai-safety, csir, preprint, safety-evaluation | 2026-08-11 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E20 / R16 (96%) | 59 |