Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation Chengju Liu, Haojie Dai, Jingwei Yang, Jinlong Li Published: 2026-07-09Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-07-09 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E7 / R7 (89%) | - |
| Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment Seongjae Kang, Taehyung Yu Published: 2026-07-09Area: cs.CLCitations: - Tags: ai-safety, alignment-training, cscl, preprint, safety-evaluation | 2026-07-09 | cs.CL | ai-safety, alignment-training, cscl, preprint, safety-evaluation | E9 / R6 (94%) | - |
| Psychological Competence as a Missing Dimension in AI Evaluation Alexis Michelle Abellar, Antoine Ferr猫re, Fendi Tsim, Marcos Economides Published: 2026-07-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E16 / R7 (92%) | - |
| VEGAS: Human-Aligned Video Caption Evaluation via Gaze Emad Barsoum, Po-han Li, Sandeep Chinchali, Shenghui Chen Published: 2026-07-09Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-09 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E8 / R8 (95%) | - |
| A Production-Oriented Framework for Evaluation of SFX Generation Eric Granger, M茅lodie Desbos, Mohammadhadi Shateri, Yara Bahram Published: 2026-07-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-07-10 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E11 / R15 (94%) | - |
| L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning Dinh-Truong Do, Hoang-Trung Nguyen, Huu-Dong Nguyen, Le-Minh Nguyen Published: 2026-07-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (92%) | - |
| Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation Aoyu Zhang, Dongping Liu, Luyao Zhang Published: 2026-07-11Area: quant-phCitations: - Tags: ai-safety, preprint, quant-ph, safety-evaluation | 2026-07-11 | quant-ph | ai-safety, preprint, quant-ph, safety-evaluation | E10 / R8 (96%) | - |
| When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Daming Luo Published: 2026-07-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R7 (95%) | - |
| 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects Alvin Chan, Jihyeon Je, Jingshen Wang, Michael Spedden Published: 2026-07-12Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-12 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R10 (90%) | - |
| AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation Alex Leung, Kentaroh Toyoda, Yi Ting Shen Published: 2026-07-13Area: cs.CRCitations: 45 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-13 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E16 / R8 (89%) | 45 |
| A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models Abhinav Mohanty, Brandon Behlendorf, Connor Harris, Gary Anthony Ackerman Published: 2026-07-13Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E16 / R15 (89%) | - |
| A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation Ali Mokhtari, Anselm Stark, Christoph Grani, Christoph Ryffel Published: 2026-07-13Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-13 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E14 / R9 (89%) | - |
| Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation Md. Mohayminul Mukit, Md. Monir Hossain Shimul, Md. Sadibul Hasan Sadib, Rahmatul Kabir Rasel Sarker Published: 2026-07-13Area: eess.IVCitations: 28 Tags: ai-safety, eessiv, preprint, safety-evaluation | 2026-07-13 | eess.IV | ai-safety, eessiv, preprint, safety-evaluation | E8 / R8 (94%) | 28 |
| From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth Aleh Manchuliantsau Published: 2026-07-13Area: cs.AICitations: 22 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (92%) | 22 |
| Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking Balaji Nagarajan, Faysal Satter, Karthik Nair, Niranjan Kumar M Published: 2026-07-13Area: cs.AICitations: 20 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E23 / R15 (91%) | 20 |
| The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation Chenglin Yu, Hongquan Gui, Hongxia Yang, Ming Li Published: 2026-07-13Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E19 / R15 (92%) | 15 |
| The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation Chenglin Yu, Hongquan Gui, Hongxia Yang, Ming Li Published: 2026-07-13Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-13 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R14 (95%) | - |
| Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric Oleg Solozobov Published: 2026-07-14Area: cs.SECitations: 35 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-14 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E19 / R19 (90%) | 35 |
| Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System Dylan Uys, Ken Jon Miyachi Published: 2026-07-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E14 / R12 (91%) | - |
| Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models Mehmet Iscan Published: 2026-07-14Area: cs.SECitations: 15 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-14 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E8 / R6 (90%) | 15 |
| Rethinking the Evaluation of Harness Evolution for Agents Hannaneh Hajishirzi, Huaisheng Zhu, Pradeep Dasigi, Shakti Senthil Published: 2026-07-14Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (92%) | 15 |
| Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Guoyi Xu, Hongpeng Zhou, Jiachen Tu, Jingyuan Sun Published: 2026-07-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E15 / R7 (91%) | - |
| Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents Bing Zhu, Guanghui Wang, Peiyang He, Wei Qiu Published: 2026-07-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (91%) | - |
| Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation Aleh Manchuliantsau Published: 2026-07-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (87%) | - |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E18 / R14 (92%) | - |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: 51 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R12 (92%) | 51 |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E14 / R10 (94%) | - |
| Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents Honglei Lv, Jiao Dai, Jiwen Zhou, Jizhong Han Published: 2026-07-15Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-07-15 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E11 / R10 (94%) | - |
| Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents Honglei Lv, Jiao Dai, Jiwen Zhou, Jizhong Han Published: 2026-07-15Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-07-15 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E15 / R13 (94%) | - |
| Copy-on-Write Scoring: Application-Specific Agent Evaluations Joanna Roy, Sven Hoelzel Published: 2026-07-15Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-15 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E12 / R11 (95%) | - |