Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models Jiayi Wang, Xu-Yao Zhang Published: 2026-06-18Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E12 / R5 (99%) | - |
| Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Aaron Fan, Akshat Bhandari, Alimurtaza Mustafa Merchant, Alisha Vinod Published: 2026-06-18Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R4 (94%) | - |
| Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU Andrea Cossettini, Elisa Vasta, Emanuela Keller, Luca Benini Published: 2026-06-18Area: eess.SPCitations: - Tags: ai-safety, eesssp, preprint, safety-evaluation | 2026-06-18 | eess.SP | ai-safety, eesssp, preprint, safety-evaluation | E9 / R4 (96%) | - |
| FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming Chaeyun Kim, Daeyoung Park, Eunji Song, Jinyoung Jeong Published: 2026-06-18Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-06-18 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E12 / R5 (96%) | - |
| Learner-based Concept Drift Detection: Analysis and Evaluation Md Moman Ul Haque Khan, Samira Sadaoui Published: 2026-06-18Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R5 (96%) | - |
| Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Adeeb Zaman, Erik Nordby, Fred Heiding Published: 2026-06-24Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-06-24 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E14 / R17 (94%) | - |
| Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation Ananto Nayan Bala, Faisal Muhammad Shah Published: 2026-06-27Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-06-27 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R9 (94%) | - |
| Cognitive World Models for Process-Level Social Influence Evaluation Bin Guo, Han Wang, Jingqi Liu, Mengqi Chen Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (98%) | - |
| Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R5 (98%) | - |
| SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models Francesco Daghero, Mayhar Tourchi Moghaddam, Tiziano Santilli Published: 2026-06-28Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-06-28 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E13 / R5 (99%) | - |
| The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling Dhruv Kumar, Murari Mandal, Shubh Chapra, Yash Sinha Published: 2026-06-28Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-28 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R5 (98%) | - |
| Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency Nisarg A. Patel Published: 2026-06-29Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-29 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R4 (95%) | - |
| CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents Bo Qu, Mingguang Chen Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R4 (98%) | - |
| EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots Camilo Chac贸n Sartori Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R6 (97%) | - |
| EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures Bu臒ra Alperen Ulu谋rmak, Rifat Kurban Published: 2026-06-29Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-06-29 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R4 (95%) | - |
| HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Jingshen Wang, Sui Huang, Waverly Wei, Xinrui Ruan Published: 2026-06-29Area: stat.MECitations: - Tags: ai-safety, preprint, safety-evaluation, statme | 2026-06-29 | stat.ME | ai-safety, preprint, safety-evaluation, statme | E7 / R5 (97%) | - |
| SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics Dimitar I. Dimitrov, Ivo Petrov, Kseniia Ibragimova, Maria Drencheva Published: 2026-06-29Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-06-29 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E9 / R5 (98%) | - |
| A Large-Scale Empirical Evaluation of MMAO Under Fair-Budget Continuous and Discrete Benchmarks Jinliang Xu, Liping Ma Published: 2026-06-30Area: cs.NECitations: - Tags: ai-safety, csne, preprint, safety-evaluation | 2026-06-30 | cs.NE | ai-safety, csne, preprint, safety-evaluation | E8 / R5 (99%) | - |
| A Large-Scale Empirical Evaluation of MMAO Under Fair-Budget Continuous and Discrete Benchmarks Jinliang Xu, Liping Ma Published: 2026-06-30Area: cs.NECitations: - Tags: ai-safety, csne, preprint, safety-evaluation | 2026-06-30 | cs.NE | ai-safety, csne, preprint, safety-evaluation | E10 / R9 (93%) | - |
| Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian Adrian Paschke, Ciprian-Octavian Truica, Dragos-Mitrut Vasile, Elena-Simona Apostol Published: 2026-06-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R4 (99%) | - |
| Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law Chang Sun, Noah Scharrenberg Published: 2026-06-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R4 (96%) | - |
| SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework Carolina Gustafson, Diane Litman, Gayle Rogers, Norah Almousa Published: 2026-06-30Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-06-30 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R4 (99%) | - |
| Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity Brett Reynolds Published: 2026-07-01Area: cs.CLCitations: - Tags: adversarial-robustness, ai-safety, cscl, preprint, safety-evaluation | 2026-07-01 | cs.CL | adversarial-robustness, ai-safety, cscl, preprint, safety-evaluation | E9 / R4 (95%) | - |
| Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions Zewen Liu Published: 2026-07-01Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-07-01 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R3 (97%) | - |
| TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue Georgi Tinchev, Hao Zhang, Laureano Moro-Velazquez, Thomas Thebaud Published: 2026-07-01Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-01 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R9 (88%) | - |
| TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue Georgi Tinchev, Hao Zhang, Laureano Moro-Velazquez, Thomas Thebaud Published: 2026-07-01Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-01 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R5 (98%) | - |
| Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias Kevin Vorwalder, Nico Pfeifer, Sofiane Ouaari Published: 2026-07-02Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-02 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E7 / R4 (96%) | - |
| CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training Chenyang Zhao, Fangfei Li, Feng Tian, Long Wang Published: 2026-07-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (96%) | - |
| JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode Gabor Antal, Norbert Sandor Szolnoki Published: 2026-07-02Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-02 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E35 / R24 (94%) | - |
| Meta-Benchmarks for Financial-Services LLM Evaluation Blair Hudson Published: 2026-07-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (95%) | - |