Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 1-30 of 1000+ papers (page 1 of 34)路 72 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference Mengfan Li, Xuanhua Shi, Yang Deng, Zesheng Wei Published: 2026-08-27Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R9 (93%) | - |
| From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities Di Yin, Jiayi Kuang, Kai Jin, Keyu Chen Published: 2026-08-27Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-27 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E16 / R13 (91%) | - |
| A Hybrid Usability Approach for Rating Evaluation of M-Commerce Applications Ahmad Ibtisam, Arshad Ali, Bilal Khan Published: 2026-08-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-26 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E21 / R20 (97%) | - |
| FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence Qiming Dai, Yishan Wu, Zaiwen Wen, Ziyu Wang Published: 2026-08-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (96%) | - |
| FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review Jingzhe Zhu, Liyao Sun, Qingqing Sun, Qi Xu Published: 2026-08-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E19 / R18 (93%) | - |
| How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation Farshad Khorrami, Haoran Xi, Kimberly Milner, Meet Udeshi Published: 2026-08-26Area: cs.CRCitations: 15 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-08-26 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E18 / R17 (94%) | 15 |
| How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation Aida Usmanova, Markus Leippold, Ricardo Usbeck, Zangir Iklassov Published: 2026-08-26Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R11 (96%) | 15 |
| NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation Joseph Gardiner, Lichao Wu, Muhammad Firhard Roslan, Sana Belguith Published: 2026-08-26Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-26 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R6 (94%) | - |
| PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction Mary Raymond Published: 2026-08-26Area: cs.CVCitations: 13 Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-26 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E10 / R8 (93%) | 13 |
| Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation Haiyan Hao Published: 2026-08-26Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-08-26 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E6 / R5 (94%) | - |
| Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening David Atienza, J茅r么me Thevenot, Tomas Teijeiro, Wensi Zhang Published: 2026-08-26Area: eess.ASCitations: 45 Tags: ai-safety, eessas, preprint, safety-evaluation | 2026-08-26 | eess.AS | ai-safety, eessas, preprint, safety-evaluation | E12 / R6 (94%) | 45 |
| Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries Alden Do Rosario, Felipe Pires, Hussein Younes Published: 2026-08-26Area: cs.CLCitations: 12 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R9 (93%) | 12 |
| AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval Arun Menon, Arup Kumar Das, Gunja Agarwal, Jitesh Chandra Mishra Published: 2026-08-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E18 / R11 (94%) | - |
| AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval Arun Menon, Arup Kumar Das, Gunja Agarwal, Jitesh Chandra Mishra Published: 2026-08-25Area: cs.AICitations: 28 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R7 (91%) | 28 |
| A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation Chi Man Vong, Jianlin Chen, Wenhui Chen, Ziyao Lin Published: 2026-08-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (93%) | - |
| Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight Anupam Purwar, Kritika Srivastava, Shashank Singh Published: 2026-08-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R9 (92%) | - |
| Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks Ha Thanh Nguyen, Ken Satoh, May Myo Zin, Su Myat Noe Published: 2026-08-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (90%) | - |
| LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes Dakuo He, Runda Jia, Youcheng Zong Published: 2026-08-25Area: eess.SYCitations: - Tags: ai-safety, eesssy, preprint, safety-evaluation | 2026-08-25 | eess.SY | ai-safety, eesssy, preprint, safety-evaluation | E6 / R5 (92%) | - |
| Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing Bikun Yang, Chao Tan, Fang Zhao, Fuyuan Shi Published: 2026-08-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (90%) | - |
| The RAT: A Unified Bayesian Model for RAG Evaluation Felix Matthias Saaro, Jan Deriu, Mark Cieliebak, Pius von D盲niken Published: 2026-08-25Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-25 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | - | - |
| TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models Ali Almadan, Aminullah Tora, Changchun Yang, Faisal Wahbo Published: 2026-08-25Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R7 (92%) | 15 |
| Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B Chi Man Vong, Defu Lin, Jianlin Chen, Peiji Long Published: 2026-08-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models Gaofeng Liu, Xuetong Li Published: 2026-08-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies Bo Gao, Chunli Qiu, Dong Ni, Guangli Zhou Published: 2026-08-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications Jiafei Wu, Jianmin Chen, Jianwei Yin, Jiaqi Tang Published: 2026-08-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization Bo Ding, Dawei Feng, Huaimin Wang, Lin Wang Published: 2026-08-24Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-08-24 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E10 / R9 (93%) | - |
| The Limits of Automatic Evaluation of Creativity in Large Language Models Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi Published: 2026-08-24Area: cs.CLCitations: 29 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-24 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R6 (93%) | 29 |
| The Limits of Automatic Evaluation of Creativity in Large Language Models Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi Published: 2026-08-24Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-24 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E18 / R8 (93%) | - |
| What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels Dong Sun, Jiawei He, Jie jia, Mengyu Shi Published: 2026-08-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts YuanHang Xiao Published: 2026-08-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |