Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries Amber E. Osterholt, Bailee Rue, Bethany C. Bray, Krishna R. Patel Published: 2026-07-18Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (89%) | 15 |
| Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries Amber E. Osterholt, Bailee Rue, Bethany C. Bray, Krishna Riteshkumar Patel Published: 2026-07-18Area: cs.CLCitations: 16 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E14 / R16 (90%) | 16 |
| AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets Ming Chen, Pranav Pai Published: 2026-07-17Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (89%) | - |
| Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling Bo-An Chang, Yu-Chih Chen Published: 2026-07-17Area: cs.CVCitations: 34 Tags: ai-safety, alignment-training, cscv, preprint, safety-evaluation | 2026-07-17 | cs.CV | ai-safety, alignment-training, cscv, preprint, safety-evaluation | E12 / R12 (91%) | 34 |
| Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection Alejandro Seif, Anshu Singh Published: 2026-07-17Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-17 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E13 / R12 (89%) | - |
| Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Blaine Nelson, Paul Kassianik, Yaron Singer Published: 2026-07-16Area: cs.CRCitations: 45 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-16 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E16 / R14 (92%) | 45 |
| Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Blaine Nelson, Paul Kassianik, Yaron Singer Published: 2026-07-16Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-16 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E10 / R7 (92%) | - |
| Can We Trust Item Response Theory for AI Evaluation? Han Jiang, Jinwen Luo, Sunbeom Kwon, Susu Zhang Published: 2026-07-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E24 / R19 (92%) | - |
| Can We Trust Item Response Theory for AI Evaluation? Han Jiang, Jinwen Luo, Sunbeom Kwon, Susu Zhang Published: 2026-07-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R6 (92%) | - |
| Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions Bingyang Wang, Huiqi Zou, Yijiang Li, Ziang Xiao Published: 2026-07-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R14 (91%) | - |
| Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications Leanne Tan, Rohan Jaggi, Roy Ka-Wei Lee, Shaun Khoo Published: 2026-07-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R10 (90%) | - |
| WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays Adnan Labib, Jiahui Wu, John Maurice Gayed, Qiao Wang Published: 2026-07-16Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-16 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E12 / R11 (94%) | - |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E14 / R10 (94%) | - |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E18 / R14 (92%) | - |
| AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Bowen Yang, Dingbo Yuan, Dongsheng Zhu, Jiaye Ge Published: 2026-07-15Area: cs.AICitations: 51 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R12 (92%) | 51 |
| Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents Honglei Lv, Jiao Dai, Jiwen Zhou, Jizhong Han Published: 2026-07-15Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-07-15 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E15 / R13 (94%) | - |
| Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents Honglei Lv, Jiao Dai, Jiwen Zhou, Jizhong Han Published: 2026-07-15Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-07-15 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E11 / R10 (94%) | - |
| Copy-on-Write Scoring: Application-Specific Agent Evaluations Joanna Roy, Sven Hoelzel Published: 2026-07-15Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-15 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E12 / R11 (95%) | - |
| Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0 Priyatham Kattakinda, Soheil Feizi, Wenxiao Wang Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (92%) | - |
| Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition Donghwan Kim Published: 2026-07-15Area: cs.CLCitations: 25 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-15 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R7 (96%) | 25 |
| Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration Justin Bronder Published: 2026-07-15Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-15 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (91%) | - |
| Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System Javad Zarrin, P. George Lovell, Ruth Falconer, Teri Rumble Published: 2026-07-15Area: cs.CYCitations: 28 Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-07-15 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E13 / R10 (92%) | 28 |
| Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric Oleg Solozobov Published: 2026-07-14Area: cs.SECitations: 35 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-14 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E19 / R19 (90%) | 35 |
| Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System Dylan Uys, Ken Jon Miyachi Published: 2026-07-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E14 / R12 (91%) | - |
| Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models Mehmet Iscan Published: 2026-07-14Area: cs.SECitations: 15 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-07-14 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E8 / R6 (90%) | 15 |
| Rethinking the Evaluation of Harness Evolution for Agents Hannaneh Hajishirzi, Huaisheng Zhu, Pradeep Dasigi, Shakti Senthil Published: 2026-07-14Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (92%) | 15 |
| Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Guoyi Xu, Hongpeng Zhou, Jiachen Tu, Jingyuan Sun Published: 2026-07-14Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-07-14 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E15 / R7 (91%) | - |
| Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents Bing Zhu, Guanghui Wang, Peiyang He, Wei Qiu Published: 2026-07-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (91%) | - |
| Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation Aleh Manchuliantsau Published: 2026-07-14Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-14 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (87%) | - |
| AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation Alex Leung, Kentaroh Toyoda, Yi Ting Shen Published: 2026-07-13Area: cs.CRCitations: 45 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-07-13 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E16 / R8 (89%) | 45 |