Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Chao Liu, Hongqiang Lin, Xiaofan Bai, Xipeng Cao Published: 2026-08-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (92%) | - |
| SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Chao Liu, Hongqiang Lin, Xiaofan Bai, Xipeng Cao Published: 2026-08-11Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-11 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R9 (91%) | - |
| Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Alec Harris, Archie Chaudhury, Kasey Corra, Yixiong Hao Published: 2026-08-10Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (89%) | 45 |
| Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation Chenzhi Nie, Hao Zhang, Tie ji, Yifei Xue Published: 2026-08-10Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-08-10 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | - | - |
| Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures Yifei Cheng, Yuxu Ge Published: 2026-08-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R11 (92%) | - |
| Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Fan Zhang, Hongyuan Zhu, Shulin Tian, Yu Qiao Published: 2026-08-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement Jiaying Qian, Lancheng Gao, Xiaorong Zhu, Xiongkuo Min Published: 2026-08-10Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-10 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R7 (91%) | - |
| AI Evaluation Should Measure Verification Cost, Not Correctness Alone Fabio Persia, Generoso Immediato, Stefania Costantini, Viviana Crescitelli Published: 2026-08-09Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-09 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R5 (91%) | - |
| Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol Christoph Trattner Published: 2026-08-09Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-08-09 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E9 / R7 (93%) | - |
| Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol Christoph Trattner Published: 2026-08-09Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-08-09 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E10 / R8 (91%) | - |
| A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization Kali Prasad Vittala, Naman Kabadi, Praveen Kumar Katwe, Rakesh Chandra Balabantaray Published: 2026-08-08Area: cs.CLCitations: 22 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E15 / R11 (95%) | 22 |
| Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? Andrew Schonebaum, Eric Bennett, Marine Carpuat, Osvaldo Quinjica Published: 2026-08-08Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R8 (91%) | 15 |
| Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? Andrew Schonebaum, Eric Bennett, Marine Carpuat, Osvaldo Quinjica Published: 2026-08-08Area: cs.CLCitations: 45 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-08 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R6 (91%) | 45 |
| Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets Annan Li, Dawei Yin, Dou Shen, Jianmin Wu Published: 2026-08-08Area: cs.AICitations: 42 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-08 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R10 (90%) | 42 |
| Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering David Lo, Guoqiang Chen, Jieke Shi, Junda He Published: 2026-08-07Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-08-07 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E10 / R8 (94%) | - |
| Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models Amirhossein Zohrehvand, Rens Anderson, Tessa Verhoef Published: 2026-08-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | - | - |
| Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery Bing Zhao, Chen Zhao, Hu Wei, Jiajia Li Published: 2026-08-07Area: cs.AICitations: 63 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (90%) | 63 |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Madhusudan Singh, Sahil Pardasani Published: 2026-08-07Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-07 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E13 / R6 (93%) | - |
| Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Ruijie Hou, Yingming Li, Yueyang Jiao, Zhao Wang Published: 2026-08-07Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-07 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | - | - |
| Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques Hotaka Maeda, Yikai Lu Published: 2026-08-06Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (95%) | 45 |
| Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques Hotaka Maeda, Yikai Lu Published: 2026-08-06Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (94%) | 15 |
| AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games Boning Li, Longbo Huang, Yu Chen Published: 2026-08-06Area: cs.GTCitations: - Tags: ai-safety, csgt, preprint, safety-evaluation | 2026-08-06 | cs.GT | ai-safety, csgt, preprint, safety-evaluation | E8 / R6 (91%) | - |
| Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping Biruk Ambaw, Dilith Jayakody, Gabriel Spadon, Jaswanth Kumar Published: 2026-08-06Area: cs.LGCitations: 15 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-08-06 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (93%) | 15 |
| ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation Ahmed Mahir Sultan Rumi, Akanta Das, Md Mahbubur Rahman, Tanzima Hashem Published: 2026-08-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (92%) | - |
| Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks Catia Pesquita, David Herron, Ernesto Jim茅nez-Ruiz, Jonathon Dilworth Published: 2026-08-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R7 (89%) | - |
| MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Noah A. Smith, Omer Goldman, Reut Tsarfaty, Tomasz Limisiewicz Published: 2026-08-06Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-08-06 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (91%) | - |
| Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI Jan-Willem Versteeg, Lourens T. Bloem, Maarten D. Schermer, Marie L. De Bruin Published: 2026-08-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-08-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R6 (93%) | - |
| What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Dana茅 Metaxa, Emma Lurie, Ro Encarnaci贸n, Tina Behzad Published: 2026-08-06Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-08-06 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E8 / R6 (91%) | - |
| ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study Arif Hassan Zidan, Bowen Guo, Churan Yu, Haixing Dai Published: 2026-08-05Area: cs.CRCitations: 39 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-08-05 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E33 / R34 (91%) | 39 |
| AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation Jinting Wang, Li Liu, Shan Yang, Shengyu Li Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-08-05 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | - | - |