Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| CLAIRE: Compressed Latent Autoencoder for Industrial Representation and Evaluation -- A Deep Learning Framework for Smart Manufacturing Mengchu Zhou, Mohammadhossein Ghahramani Published: 2026-03-06Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-06 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (94%) | - |
| CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation Hassan AlOmaish, Mahmoud Alabbad, Mohammed Baharoon, Mona Alhammad Published: 2026-03-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Making AI Evaluation Deployment Relevant Through Context Specification Matthew Holmes, Reva Schwartz, Thiago Lacerda Published: 2026-03-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers Anastasiia Belova, David Heye, Jan Pennekamp, Johannes Lohm枚ller Published: 2026-03-06Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-06 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R4 (93%) | - |
| CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs Ge Zhang, Hamid Alinejad-Rokny, Jiaheng Liu, Jiajun Shi Published: 2026-03-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions Dilip Thakur, Saroj Mishra, Shiva Gaire, Srijan Gyawali Published: 2026-03-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (94%) | - |
| AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation Changyi Li, Fazl Barez, Min Yang, Pengfei Lu Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context Ashish Pandey, Tek Raj Chhetri Published: 2026-03-08Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning Canchen Lyu, Guilin Qi, Ran Gu, Tianhao Qian Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (95%) | - |
| Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers Pengfei Du Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (99%) | - |
| Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety David Gringras Published: 2026-03-08Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-08 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Efficient Policy Learning with Hybrid Evaluation-Based Genetic Programming for Uncertain Agile Earth Observation Satellite Scheduling Junhua Xue, Yuning Chen Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R4 (97%) | - |
| Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision-Language Models Baek Duhyeong, Eungyeol Han, Gukin han, Jaehyun Jeon Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R3 (94%) | - |
| MASEval: Extending Multi-Agent Evaluation from Models to Systems Ahmed Heakl, Alexander Rubinstein, Anmol Goel, Cornelius Emde Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R4 (99%) | - |
| SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation Carsten Maple, Kayvan Atefi, Omar Santos, Petar Radanliev Published: 2026-03-09Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-09 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R3 (96%) | - |
| AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems Athanasios Davvetas, Michael Papademas, Vangelis Karkaletsis, Xenia Ziouvelou Published: 2026-03-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| GNNs for Time Series Anomaly Detection: An Open-Source Framework and a Critical Evaluation Federico Bello, Federico Larroca, Gast贸n Garc铆a Gonz谩lez, Gonzalo Chiarlone Published: 2026-03-10Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-10 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R3 (96%) | - |
| Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges Rongxiang Zeng, Yongqi Dong Published: 2026-03-10Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-03-10 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E5 / R3 (94%) | - |
| MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings Aditya Choudhary, Anupam Purwar Published: 2026-03-10Area: cs.ETCitations: - Tags: ai-safety, cset, preprint, safety-evaluation | 2026-03-10 | cs.ET | ai-safety, cset, preprint, safety-evaluation | E5 / R3 (94%) | - |
| CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents Marta Sumyk, Oleksandr Kosovan Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R3 (98%) | - |
| LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms Denglin Jiang, Haoting Zhang, Jinghai He, Shen Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R5 (97%) | - |
| Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation Aditya Sood, Christine Allen-Blanchette, Dongzhe Zheng, Tao Zhong Published: 2026-03-11Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-11 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (94%) | - |
| RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation Carson Ezell, Dan Bateyko, Ella Guest, Gailius Praninskas Published: 2026-03-11Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-03-11 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E4 / R3 (94%) | - |
| RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents Robin Cohen, Yonas Atinafu Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (94%) | - |
| Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation Jesus Villalba-Lopez, Laureano Moro-Velazquez, Najim Dehak, Thomas Thebaud Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-11 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI David Fraile Navarro, Enrico Coiera, Farah Magrabi Published: 2026-03-12Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-12 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E7 / R2 (95%) | - |
| Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese Anna Matsuoka, Atsushi Ohara, Genichiro Ishii, Hirohiko Miyake Published: 2026-03-12Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-12 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R3 (96%) | - |
| SemBench: A Universal Semantic Framework for LLM Evaluation German Rigau, Mikel Zubillaga, Naiara Perez, Oscar Sainz Published: 2026-03-12Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-12 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R3 (97%) | - |
| A Systematic Evaluation Protocol of Graph-Derived Signals for Tabular Machine Learning Gonzalo Wandosell Fern谩ndez de Bobadilla, Jeffrey Heidemann, Mario Heidrich, R眉diger Buchkremer Published: 2026-03-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models Chris Dongjoo Kim, Dieter Fox, Ranjay Krishna, Suhwan Choi Published: 2026-03-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (99%) | - |