Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings Aditya Choudhary, Anupam Purwar Published: 2026-03-10Area: cs.ETCitations: - Tags: ai-safety, cset, preprint, safety-evaluation | 2026-03-10 | cs.ET | ai-safety, cset, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Efficient Policy Learning with Hybrid Evaluation-Based Genetic Programming for Uncertain Agile Earth Observation Satellite Scheduling Junhua Xue, Yuning Chen Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R4 (97%) | - |
| Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision-Language Models Baek Duhyeong, Eungyeol Han, Gukin han, Jaehyun Jeon Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R3 (94%) | - |
| MASEval: Extending Multi-Agent Evaluation from Models to Systems Ahmed Heakl, Alexander Rubinstein, Anmol Goel, Cornelius Emde Published: 2026-03-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R4 (99%) | - |
| SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation Carsten Maple, Kayvan Atefi, Omar Santos, Petar Radanliev Published: 2026-03-09Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-09 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R3 (96%) | - |
| AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation Changyi Li, Fazl Barez, Min Yang, Pengfei Lu Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context Ashish Pandey, Tek Raj Chhetri Published: 2026-03-08Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E4 / R2 (97%) | - |
| Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning Canchen Lyu, Guilin Qi, Ran Gu, Tianhao Qian Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R3 (95%) | - |
| Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers Pengfei Du Published: 2026-03-08Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (99%) | - |
| Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety David Gringras Published: 2026-03-08Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-08 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E5 / R3 (94%) | - |
| CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs Ge Zhang, Hamid Alinejad-Rokny, Jiaheng Liu, Jiajun Shi Published: 2026-03-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E4 / R3 (95%) | - |
| SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions Dilip Thakur, Saroj Mishra, Shiva Gaire, Srijan Gyawali Published: 2026-03-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (94%) | - |
| An Interactive Multi-Agent System for Evaluation of New Product Concepts Bin Xuan, Hakyeon Lee, Ruo Ai Published: 2026-03-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (96%) | - |
| CLAIRE: Compressed Latent Autoencoder for Industrial Representation and Evaluation -- A Deep Learning Framework for Smart Manufacturing Mengchu Zhou, Mohammadhossein Ghahramani Published: 2026-03-06Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-06 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E5 / R3 (94%) | - |
| CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation Hassan AlOmaish, Mahmoud Alabbad, Mohammed Baharoon, Mona Alhammad Published: 2026-03-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-03-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E5 / R3 (94%) | - |
| Making AI Evaluation Deployment Relevant Through Context Specification Matthew Holmes, Reva Schwartz, Thiago Lacerda Published: 2026-03-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E5 / R3 (95%) | - |
| Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers Anastasiia Belova, David Heye, Jan Pennekamp, Johannes Lohm枚ller Published: 2026-03-06Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-03-06 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E6 / R4 (93%) | - |
| MOSAIC: A Unified Platform for Cross-Paradigm Comparison and Evaluation of Homogeneous and Heterogeneous Multi-Agent RL, LLM, VLM, and Human Decision-Makers Abdulhamid M. Mousa, Abdulkarim M. Mousa, Jalaledin M. Azzabi, Ming Liu Published: 2026-03-01Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-01 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E16 / R15 (90%) | - |
| SimAB: Simulating A/B Tests with Persona-Conditioned AI Agents for Rapid Design Evaluation Alina Rublea, Francisco Chicharro Sanz, Marian Schneider, Mario Truss Published: 2026-03-01Area: cs.HCCitations: 45 Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-01 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E9 / R9 (93%) | 45 |
| A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction Ruihao Pan, Suhang Wang Published: 2026-02-28Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (93%) | - |
| CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation Linjie Mu, Shaoting Zhang, Xiaofan Zhang, Xizhuo Zhang Published: 2026-02-28Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R9 (91%) | - |
| Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking Jiayi Liu, Shicong Feng, Tianpeng Zheng, Zhehan Jiang Published: 2026-02-28Area: cs.CLCitations: 33 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E15 / R14 (93%) | 33 |
| Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma Gabriella Waters, Reva Schwartz Published: 2026-02-28Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-28 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E10 / R9 (91%) | - |
| AI Evaluation Should Require Standardized Item-Level Data Releases Dongyao Zhu, Han Jiang, Sang T. Truong, Sanmi Koyejo Published: 2026-02-27Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-27 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R11 (89%) | - |
| AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech Akshara Prabhakar, Caiming, Haolin Chen, Huan Wang Published: 2026-02-27Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-02-27 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E14 / R15 (91%) | - |
| A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science Han Li, QingGuo Zhou, Runze Mao, Tianhao Wu Published: 2026-02-27Area: cs.CLCitations: 16 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (93%) | 16 |
| Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection Brianna D'Urso, Syed Rafay Hasan, Tahmid Hasan Sakib, Terry N. Guo Published: 2026-02-27Area: cs.CVCitations: 17 Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-02-27 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E11 / R7 (94%) | 17 |
| Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment Charles L. A. Clarke, Dake Zhang, Mark D. Smucker Published: 2026-02-27Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-02-27 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E7 / R7 (91%) | - |
| When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation Bian Sun, Orvill de la Torre, Zhenjian Wang, Zirui Wang Published: 2026-02-27Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (94%) | 15 |
| Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds Vadim Indelman, Yaacov Pariente Published: 2026-02-26Area: math.STCitations: 15 Tags: ai-safety, mathst, preprint, safety-evaluation | 2026-02-26 | math.ST | ai-safety, mathst, preprint, safety-evaluation | E8 / R6 (95%) | 15 |