Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails Gregory N. Frank Published: 2026-03-18Area: cs.LGCitations: - Tags: ai-safety, alignment-training, cslg, preprint, safety-evaluation | 2026-03-18 | cs.LG | ai-safety, alignment-training, cslg, preprint, safety-evaluation | E8 / R4 (98%) | - |
| Machine Learning for Network Attacks Classification and Statistical Evaluation of Machine Learning for Network Attacks Classification and Adversarial Learning Methodologies for Synthetic Data Generation Christos Douligeris, Iakovos-Christos Zarkadis Published: 2026-03-18Area: cs.CRCitations: - Tags: adversarial-robustness, ai-safety, cscr, preprint, safety-evaluation | 2026-03-18 | cs.CR | adversarial-robustness, ai-safety, cscr, preprint, safety-evaluation | E5 / R3 (97%) | - |
| MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models Ismail Ben Ayed, Maxime Darrin, Pablo Piantanida, Philippe Formont Published: 2026-03-18Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (93%) | - |
| The Validity Gap in Health AI Evaluation: A Cross-Sectional Analysis of Benchmark Composition Alvin Rajkomar, Angela Lai, Lily Peng, Pavan Sudarshan Published: 2026-03-18Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models Boryeong Cho, Hojung Jung, Jaehyun Kwak, Juhyeong Kim Published: 2026-03-18Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-18 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R3 (97%) | - |
| CoMAI: A Collaborative Multi-Agent Framework for Robust and Equitable Interview Evaluation Bin Zhang, Gengxin Sun, Liangyi Yin, Ruihao Yu Published: 2026-03-17Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-03-17 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E6 / R4 (94%) | - |
| DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models Jiaqi Xiong, Qi Cao, Ruofan Liao, Sichen Liu Published: 2026-03-17Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (98%) | - |
| DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models Jiaqi Xiong, Qi Cao, Ruofan Liao, Sichen Liu Published: 2026-03-17Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models Aosong Feng, Chanjun Park, Irene Li, Jinghui Lu Published: 2026-03-17Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-17 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R4 (96%) | - |
| Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic Bjarki 脕rmannsson, Finnur 脕g煤st Ingimundarson, Iris Edda Nowenstein, Steinunn Rut Fri冒riksd贸ttir Published: 2026-03-17Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-17 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R4 (96%) | - |
| An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc Barry Smith, Hong Zhang, Junchao Zhang, Le Chen Published: 2026-03-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (99%) | - |
| Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains Blanka Horvath, Christoph Bergmeir, Daniel Schmidt, Frank Rudzicz Published: 2026-03-16Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-16 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (93%) | - |
| SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration Bin Wu, Guangquan Xu, Qiannan Si, Tiejun Wu Published: 2026-03-16Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-16 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E4 / R3 (94%) | - |
| SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction David 艩teva艌谩k, Marek 艩uppa Published: 2026-03-16Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-16 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R3 (96%) | - |
| Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis Atin Ghosh, Daniel Dahlmeier, Harshavardhan Abichandani, Jiyuan Shen Published: 2026-03-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R4 (95%) | - |
| Anterior's Approach to Fairness Evaluation of Automated Prior Authorization System Anuj Iravane, Khadija Mahmoud, Sai P. Selvaraj Published: 2026-03-15Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-15 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (95%) | - |
| A Systematic Evaluation Protocol of Graph-Derived Signals for Tabular Machine Learning Gonzalo Wandosell Fern谩ndez de Bobadilla, Jeffrey Heidemann, Mario Heidrich, R眉diger Buchkremer Published: 2026-03-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (96%) | - |
| vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models Chris Dongjoo Kim, Dieter Fox, Ranjay Krishna, Suhwan Choi Published: 2026-03-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (99%) | - |
| Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI David Fraile Navarro, Enrico Coiera, Farah Magrabi Published: 2026-03-12Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-12 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E7 / R2 (95%) | - |
| Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese Anna Matsuoka, Atsushi Ohara, Genichiro Ishii, Hirohiko Miyake Published: 2026-03-12Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-12 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R3 (96%) | - |
| SemBench: A Universal Semantic Framework for LLM Evaluation German Rigau, Mikel Zubillaga, Naiara Perez, Oscar Sainz Published: 2026-03-12Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-12 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R3 (97%) | - |
| CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents Marta Sumyk, Oleksandr Kosovan Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R3 (98%) | - |
| LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms Denglin Jiang, Haoting Zhang, Jinghai He, Shen Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R5 (97%) | - |
| Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation Aditya Sood, Christine Allen-Blanchette, Dongzhe Zheng, Tao Zhong Published: 2026-03-11Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-11 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (94%) | - |
| RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation Carson Ezell, Dan Bateyko, Ella Guest, Gailius Praninskas Published: 2026-03-11Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-03-11 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E4 / R3 (94%) | - |
| RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents Robin Cohen, Yonas Atinafu Published: 2026-03-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (94%) | - |
| Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation Jesus Villalba-Lopez, Laureano Moro-Velazquez, Najim Dehak, Thomas Thebaud Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-11 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (96%) | - |
| AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems Athanasios Davvetas, Michael Papademas, Vangelis Karkaletsis, Xenia Ziouvelou Published: 2026-03-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| GNNs for Time Series Anomaly Detection: An Open-Source Framework and a Critical Evaluation Federico Bello, Federico Larroca, Gast贸n Garc铆a Gonz谩lez, Gonzalo Chiarlone Published: 2026-03-10Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-10 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R3 (96%) | - |
| Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges Rongxiang Zeng, Yongqi Dong Published: 2026-03-10Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-03-10 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E5 / R3 (94%) | - |