Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems Henry Fordjour Ansah, Pranish Ghimire, Shreya Banerjee Published: 2026-08-23Area: cs.AICitations: 12 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R6 (92%) | 12 |
| ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation Charles L. A. Clarke, Eugene Y. Agichtein, Kaustubh D. Dhole Published: 2026-08-23Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R10 (92%) | 45 |
| ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation Charles L. A. Clarke, Eugene Y. Agichtein, Kaustubh D. Dhole Published: 2026-08-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model Ergan Shang, Weijing Tang, Yinqiu He Published: 2026-08-23Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-23 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | - | - |
| ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents Yi Chang, YiShan Zheng, Yuan Wu Published: 2026-08-23Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-23 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E15 / R12 (92%) | - |
| Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings Aarzoo Dhiman, Farzana Haque, Kartikae Grover, Lydia Brian Smith Published: 2026-08-22Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-22 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (95%) | - |
| Evaluation Awareness in Language Models: Representation, Verbalization, and Control Amin Memarian, Farzaneh Heidari, Guillaume Rabusseau Published: 2026-08-22Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-22 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E18 / R11 (92%) | - |
| Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation Benhui Zhuang, Bo Yuan, Junlan Feng, Pengshuai Yang Published: 2026-08-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (95%) | - |
| Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation Adam Jatowt, Lorenz Brehme Published: 2026-08-21Area: cs.CLCitations: 12 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E12 / R9 (94%) | 12 |
| Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning Arman Cohan, Manasi Patwardhan, Snigdha Paul Published: 2026-08-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R10 (91%) | - |
| Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric Mariette Awad, Sami Shames El Deen Published: 2026-08-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R7 (92%) | - |
| Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge Mohammad Shojafar, Rishiraj Sengupta, Sotiris Chatzimiltis, Xiatian Zhu Published: 2026-08-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E12 / R8 (99%) | - |
| From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry Adriana Watson, Grant Richards, Marco B眉cheler Published: 2026-08-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation David Hartmann, Dimitri Staufer, Ibrahim Baroud Published: 2026-08-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | - | - |
| Source-Free MT Evaluation Is Not MT Evaluation Asif Ekbal, Baban Gain, Ramakrishna Appicharla Published: 2026-08-21Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R6 (93%) | 15 |
| Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems Balkrishna Giri, Jussi Rasku, Md Toufique Hasan, Muhammad Waseem Published: 2026-08-21Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-21 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E12 / R7 (94%) | - |
| Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills Christopher Kevin, Jean-Francois Puget, Meghana Puvvadi, Mohit Gupta Published: 2026-08-20Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-20 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (90%) | - |
| ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib Shane Caldwell Published: 2026-08-20Area: cs.LOCitations: 7 Tags: ai-safety, cslo, preprint, safety-evaluation | 2026-08-20 | cs.LO | ai-safety, cslo, preprint, safety-evaluation | E12 / R8 (92%) | 7 |
| Rethinking the Evaluation and Optimization of LLM-Based Social Simulation Ji-Rong Wen, Pei Wang, Xu Chen Published: 2026-08-20Area: cs.AICitations: 33 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-20 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (94%) | 33 |
| Testing and Evaluation of Agentic AI Systems In Military Command and Control Adrianna Tan, Di Cooke, Heather Frase, Sarah Cao Published: 2026-08-20Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-20 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R10 (92%) | - |
| Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation Alexandre Bayen, Alex Zongo, Jordan Kam, Mahyar Ghazanfari Published: 2026-08-19Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-19 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (92%) | - |
| Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation Dongjoon Yoo, Gyunho Rho, Minje Park, Sunghoon Joo Published: 2026-08-19Area: eess.SPCitations: - Tags: ai-safety, eesssp, preprint, safety-evaluation | 2026-08-19 | eess.SP | ai-safety, eesssp, preprint, safety-evaluation | E8 / R6 (95%) | - |
| SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation Chen Tang, Dijie Zhu, Jiaqi Ma, Ruopeng Huang Published: 2026-08-19Area: cs.ROCitations: 31 Tags: ai-safety, csro, preprint, safety-evaluation | 2026-08-19 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E8 / R7 (90%) | 31 |
| Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting Changjian Chen, Haixin Xie, Rongwen Li, Xiao Wang Published: 2026-08-18Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R6 (92%) | - |
| Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models Christophe Mues, Cristi谩n Bravo, Mar铆a 脫skarsd贸ttir, Noah Kostesku Published: 2026-08-18Area: q-fin.RMCitations: - Tags: ai-safety, preprint, q-finrm, safety-evaluation | 2026-08-18 | q-fin.RM | ai-safety, preprint, q-finrm, safety-evaluation | E9 / R7 (97%) | - |
| Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges Firoj Alam, Hunzalah Hassan Bhatti, Shammur Absar Chowdhury, Syeda Faiza Ahmed Published: 2026-08-18Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R8 (92%) | - |
| Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility Comfort Adesina, Marvellous O. Ajala, Zainab Ashimiyu-Abdusalam Published: 2026-08-18Area: q-bio.QMCitations: - Tags: ai-safety, preprint, q-bioqm, safety-evaluation | 2026-08-18 | q-bio.QM | ai-safety, preprint, q-bioqm, safety-evaluation | E16 / R10 (94%) | - |
| SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition Adel Youssef, Dae Lee, Mihai Delgeanu Published: 2026-08-18Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (91%) | 45 |
| The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation Aniket Wattamwar, Manav Anandani, Mrunal Kakirwar Published: 2026-08-18Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-18 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E16 / R11 (91%) | - |
| Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions Dieter Jaksch, Pia Siegl, Tomohiro Hashizume, Xiao Wang Published: 2026-08-17Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-17 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (94%) | - |