Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems Isaiah Andrews Published: 2026-08-05Area: econ.THCitations: - Tags: ai-safety, econth, preprint, safety-evaluation | 2026-08-05 | econ.TH | ai-safety, econth, preprint, safety-evaluation | - | - |
| TriQua: Reconciling Granularity and Context in Factuality Evaluation Achim Rettinger, Jin Liu, Steffen Thoma Published: 2026-08-05Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-05 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (93%) | 45 |
| When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit Mingjie Pang, Minjun Yu, Wei Li, Zheyuan Lai Published: 2026-08-05Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-05 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Abdullah Nazly, Januki Wanniarachchi, Ravisha De Alwis, Saqib Shouqi Published: 2026-08-04Area: cs.AICitations: - Tags: adversarial-robustness, ai-safety, csai, preprint, safety-evaluation | 2026-08-04 | cs.AI | adversarial-robustness, ai-safety, csai, preprint, safety-evaluation | - | - |
| Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation Chunlin Liu, Haitong Jiang, Jiabiao He, Jianyu Zhao Published: 2026-08-04Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-04 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary Peiyao Sheng, Pramod Viswanath, S. Ashwin Hebbar, Sewoong Oh Published: 2026-08-04Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-04 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E6 / R5 (90%) | - |
| KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation Hengyang Lu, Huining Li, Jiyang Tan, Kailin Jiang Published: 2026-08-04Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-04 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| LiveEvalBench: Toward Open-World Evaluation for Web Generation Jun Zhou, Lin Yuan, Wei Chen, Xiaolau Zhang Published: 2026-08-04Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-04 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces Andres Algaba, Francesca Carlon, Vincent Ginis Published: 2026-08-04Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-04 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | - | - |
| Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems I. de Zarz脿, J. de Curt貌 Published: 2026-08-04Area: cs.MACitations: - Tags: ai-safety, csma, preprint, safety-evaluation | 2026-08-04 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E13 / R6 (89%) | - |
| Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Amirhossein Samandar, Biyao Zhang, Debargha Ganguly, Jerry Peng Published: 2026-08-04Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-04 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | - | - |
| TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Bhavin Jawade, Cameron R. Wolfe Published: 2026-08-04Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-04 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R6 (89%) | - |
| MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models Dana茅 Metaxa, Ro Encarnaci贸n, Suresh Venkatasubramanian, Victor Ojewale Published: 2026-08-03Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-03 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Bonan Shen, Wei-Jung Huang Published: 2026-08-03Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-03 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks Bing Zhao, Hu Wei, Tianle Pu, Weiqi Zhai Published: 2026-08-03Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-03 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories Boyun Xu, Kaishen Yuan, Shaofeng Liang, Songning Lai Published: 2026-08-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| PATH-Bench: Path-Dependent Evaluation of Lifelong Agents Chuyun Shen, Junjie Sheng, Tao Fang, Wei Yin Published: 2026-08-02Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-02 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments Dotan Davidovich, Hai Rozencwajg, Or Hiltch, Yair Amar Published: 2026-08-02Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-08-02 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E10 / R9 (91%) | - |
| Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics Grant Schoenebeck, Jason Hartline, Shengwei Xu, Yifan Wu Published: 2026-08-02Area: cs.AICitations: - Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-08-02 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | - | - |
| Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation Diego Kreutz, Francis Luis Santos Vargas, Rodrigo Brand茫o Mansilha Published: 2026-08-02Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-08-02 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E12 / R6 (94%) | - |
| AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment Baskhad Idrisov, Garv Vikram Gursahaney, Thorsten Fr枚hlich, Tim Schlippe Published: 2026-08-01Area: cs.AICitations: 11 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E17 / R11 (89%) | 11 |
| An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models Enzo Nicolas Spotorno, Josafat Leal Filho Published: 2026-08-01Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-01 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E10 / R6 (91%) | - |
| Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation William Caban Published: 2026-08-01Area: cs.AICitations: 77 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R9 (90%) | 77 |
| Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation William Caban Published: 2026-08-01Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R10 (91%) | 15 |
| Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation William Caban Published: 2026-08-01Area: cs.AICitations: 77 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (91%) | 77 |
| Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems Abdalla Doleh, Ratna Babu Chinnam Published: 2026-08-01Area: cs.AICitations: 108 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E17 / R16 (93%) | 108 |
| Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations Jing Shao, Lijun Li, Xinshun Feng, Ziqi Miao Published: 2026-08-01Area: cs.AICitations: 20 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (90%) | 20 |
| When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty Jana Gonnermann-M眉ller, Nicolas Leins, Nico Pelleriti, Sebastian Pokutta Published: 2026-08-01Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-01 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E14 / R9 (94%) | - |
| A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation Keita Kinjo Published: 2026-07-31Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-07-31 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (93%) | - |
| ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation Gaetano Perrone, Simon Pietro Romano Published: 2026-07-31Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-07-31 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E20 / R18 (95%) | - |