Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| No Single Metric Tells the Whole Story: A Multi-Dimensional Evaluation Framework for Uncertainty Attributions Emily Schiller, Luca Longo, Marco Zullich, Teodor Chiaburu Published: 2026-03-25Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-25 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Beyond Binary Correctness: Scaling Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Abhishek Chandwani, Ishan Gupta Published: 2026-03-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Human-in-the-Loop Pareto Optimization: Trade-off Characterization for Assist-as-Needed Training and Performance Evaluation Harun Tolasa, Volkan Patoglu Published: 2026-03-24Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-03-24 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E6 / R3 (94%) | - |
| LLM Olympiad: Why Model Evaluation Needs a Sealed Exam Alham Fikri Aji, Jan Christian Blaise Cruz Published: 2026-03-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (96%) | - |
| MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation Di Zhu, Zixuan Li Published: 2026-03-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Off-Policy Evaluation and Learning for Survival Outcomes under Censoring Kohsuke Kubota, Mitsuhiro Takahashi, Yuta Saito Published: 2026-03-24Area: stat.MECitations: - Tags: ai-safety, preprint, safety-evaluation, statme | 2026-03-24 | stat.ME | ai-safety, preprint, safety-evaluation, statme | E5 / R3 (96%) | - |
| PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset Juhyeon Park, Sumin Yu, Taesup Moon Published: 2026-03-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-03-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Ran Score: a LLM-based Evaluation Score for Radiology Report Generation Bowen Liu, Danni Ai, Deqiang Xiao, Hongliang Sun Published: 2026-03-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| When AI Shows Its Work, Is It Actually Working? Step-Level Evaluation Reveals Frontier Language Models Frequently Bypass Their Own Reasoning Abhinaba Basu, Pavan Chakraborty Published: 2026-03-24Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-24 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R3 (95%) | - |
| BadminSense: Enabling Fine-Grained Badminton Stroke Evaluation on a Single Smartwatch Kai Chen, Pingchuan Ke, Taizhou Chen, Xingyu Liu Published: 2026-03-23Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-23 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E4 / R3 (98%) | - |
| BadminSense: Enabling Fine-Grained Badminton Stroke Evaluation on a Single Smartwatch Kai Chen, Pingchuan Ke, Taizhou Chen, Xingyu Liu Published: 2026-03-23Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-23 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E5 / R3 (97%) | - |
| Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation Hao Wang, Jie Ma, Jing Tao, Pinghui Wang Published: 2026-03-23Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-03-23 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E6 / R3 (96%) | - |
| Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation Hao Wang, Jie Ma, Jing Tao, Pinghui Wang Published: 2026-03-23Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-03-23 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E41 / R32 (97%) | - |
| SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation Cong Cong, Leyi Wei, Linkuan Zhou, Qiangguo Jin Published: 2026-03-23Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-23 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E5 / R4 (97%) | - |
| AdaRubric: Task-Adaptive Rubrics for LLM Agent Evaluation Liang Ding Published: 2026-03-22Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-22 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R3 (98%) | - |
| Efficient Fine-Tuning Methods for Portuguese Question Answering: A Comparative Study of PEFT on BERTimbau and Exploratory Evaluation of Generative LLMs Caio Veloso Costa, Didier A. Vega-Oliveros, Lilian Berton, Mariela M. Nina Published: 2026-03-22Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-22 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R3 (97%) | - |
| RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation Eugene Agichtein, Kaustubh D. Dhole Published: 2026-03-21Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-03-21 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E5 / R3 (94%) | - |
| ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability Chun-Yi Lee, Jay Chiehen Liao, Shang-Tse Chen, Shao-Yuan Lo Published: 2026-03-20Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-20 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E34 / R31 (97%) | - |
| An Industrial-Scale Retrieval-Augmented Generation Framework for Requirements Engineering: Empirical Evaluation with Automotive Manufacturing Data Muhammad Khalid, Yilmaz Uygun Published: 2026-03-20Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-20 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E6 / R3 (96%) | - |
| An Industrial-Scale Retrieval-Augmented Generation Framework for Requirements Engineering: Empirical Evaluation with Automotive Manufacturing Data Muhammad Khalid, Yilmaz Uygun Published: 2026-03-20Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-20 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E43 / R36 (96%) | - |
| CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation Du-Seong Chang, Haejun Yoo, Insung Lee, Myoung-Wan Koo Published: 2026-03-20Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-20 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (97%) | - |
| Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation Richard J. Young Published: 2026-03-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R4 (98%) | - |
| Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation Richard J. Young Published: 2026-03-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R4 (99%) | - |
| Permutation-Consensus Listwise Judging for Robust Factuality Evaluation Elsa Fan, Justin Tang, Nathan Huang, Tianyi Huang Published: 2026-03-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E33 / R25 (94%) | - |
| Span-Level Machine Translation Meta-Evaluation Eric Morales Agostinho, Hugo Zaragoza, Stefano Perrella Published: 2026-03-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R3 (94%) | - |
| Benchmarking PDF Parsers on Table Extraction with LLM-based Semantic Evaluation Janis Keuper, Pius Horn Published: 2026-03-19Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-03-19 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E6 / R3 (94%) | - |
| ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation Haochen Zhao, Shaoyang Cui Published: 2026-03-19Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-19 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E5 / R4 (97%) | - |
| ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs Abhinaba Basu, Pavan Chakraborty Published: 2026-03-19Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-19 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (98%) | - |
| Is Evaluation Awareness Just Format Sensitivity? Limitations of Probe-Based Evidence under Controlled Prompt Structure Viliana Devbunova Published: 2026-03-19Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-19 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Deployment and Evaluation of an EHR-integrated, Large Language Model-Powered Tool to Triage Surgical Patients Abby Pandya, April S Liang, Jane Wang, Jason Hom Published: 2026-03-18Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-03-18 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E5 / R3 (97%) | - |