Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents Bo Qu, Mingguang Chen Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R4 (98%) | - |
| EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots Camilo Chac贸n Sartori Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R6 (97%) | - |
| EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures Bu臒ra Alperen Ulu谋rmak, Rifat Kurban Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R4 (95%) | - |
| HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Jingshen Wang, Sui Huang, Waverly Wei, Xinrui Ruan Published: 2026-06-29Area: stat.MECitations: - Tags: ai-safety, preprint, safety-evaluation, statme | 2026-06-29 | stat.ME | ai-safety, preprint, safety-evaluation, statme | E7 / R5 (97%) | - |
| SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics Dimitar I. Dimitrov, Ivo Petrov, Kseniia Ibragimova, Maria Drencheva Published: 2026-06-29Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-06-29 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E9 / R5 (98%) | - |
| Cognitive World Models for Process-Level Social Influence Evaluation Bin Guo, Han Wang, Jingqi Liu, Mengqi Chen Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (98%) | - |
| Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R5 (98%) | - |
| SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models Francesco Daghero, Mayhar Tourchi Moghaddam, Tiziano Santilli Published: 2026-06-28Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-06-28 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E13 / R5 (99%) | - |
| The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling Dhruv Kumar, Murari Mandal, Shubh Chapra, Yash Sinha Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R5 (98%) | - |
| Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation Ananto Nayan Bala, Faisal Muhammad Shah Published: 2026-06-27Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-27 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R9 (94%) | - |
| Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Adeeb Zaman, Erik Nordby, Fred Heiding Published: 2026-06-24Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-06-24 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E14 / R17 (94%) | - |
| A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models Jiayi Wang, Xu-Yao Zhang Published: 2026-06-18Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E12 / R5 (99%) | - |
| Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Aaron Fan, Akshat Bhandari, Alimurtaza Mustafa Merchant, Alisha Vinod Published: 2026-06-18Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R4 (94%) | - |
| Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU Andrea Cossettini, Elisa Vasta, Emanuela Keller, Luca Benini Published: 2026-06-18Area: eess.SPCitations: - Tags: ai-safety, eesssp, preprint, safety-evaluation | 2026-06-18 | eess.SP | ai-safety, eesssp, preprint, safety-evaluation | E9 / R4 (96%) | - |
| FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming Chaeyun Kim, Daeyoung Park, Eunji Song, Jinyoung Jeong Published: 2026-06-18Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-06-18 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E12 / R5 (96%) | - |
| Learner-based Concept Drift Detection: Analysis and Evaluation Md Moman Ul Haque Khan, Samira Sadaoui Published: 2026-06-18Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R5 (96%) | - |
| A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies Fangyijie Wang, Gu茅nol茅 Silvestre, Haixia Huang, Jianjun Yu Published: 2026-06-17Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-06-17 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E9 / R5 (97%) | - |
| Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep Amama Mahmood, Bokyung Kim, Chien-Ming Huang, Honghao Zhao Published: 2026-06-17Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-06-17 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Gender Bias in LLM Hiring Decisions: Evidence from a Japanese Context and Evaluation of Mitigation Strategies Akshara Nadayanur Sathis Kanna, Gabriele Trovato, Machiko Hirota, Phan Xuan Tan Published: 2026-06-17Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-06-17 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E8 / R5 (97%) | - |
| AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows Bryan YP Yan, Guangxin Dai, Huizi Yu, Jiahui Niu Published: 2026-06-16Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-16 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R5 (97%) | - |
| DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack Chunyu Sun, Haitao Cui, Jiahao Zhang, Jie Chen Published: 2026-06-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (100%) | - |
| Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines Mostafa Darvishi Published: 2026-06-16Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-16 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R5 (96%) | - |
| How Inference Compute Shapes Frontier LLM Evaluation Cozmin Ududec, Harry Coppock, Jessica McFadyen, Kevin Wei Published: 2026-06-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R5 (97%) | - |
| MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation Jiamin Yan, Jiaxin Dai, Xiang Xiang, Zehang Wei Published: 2026-06-16Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-16 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E13 / R6 (99%) | - |
| Offline Preference-Based Trajectory Evaluation Fernando Diaz Published: 2026-06-16Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-16 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E10 / R4 (97%) | - |
| RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills A. Ali Heydari, Ahmed A. Metwally, Ben Graef, Chloe Zhang Published: 2026-06-16Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-16 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R4 (97%) | - |
| SEAGym: An Evaluation Environment for Self-Evolving LLM Agents Bin Liang, Changshui Zhang, Chuanyi Xue, Congjie Zheng Published: 2026-06-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R4 (97%) | - |
| An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Akriti Vij, Ee Wei Seah, Gabriel Waikin Loh Matienzo, Hankyul Baek Published: 2026-06-15Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-06-15 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E9 / R3 (98%) | - |
| Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Yanan Long Published: 2026-06-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R4 (96%) | - |
| DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models Despoina Paschalidou, Jenny Schmalfuss, Jose M. Alvarez, Kashyap Chitta Published: 2026-06-15Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-06-15 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E8 / R5 (97%) | - |