Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text Van-Truong Le Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-17 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R5 (99%) | - |
| From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation Chen Xie, Feng Jiang, Jade Yang, Jasper Lu Published: 2026-04-17Area: cs.ROCitations: - Tags: ai-safety, csro, preprint, safety-evaluation | 2026-04-17 | cs.RO | ai-safety, csro, preprint, safety-evaluation | E14 / R9 (97%) | - |
| MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Farhad Abtahi, Fernando Seoane Published: 2026-04-17Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E17 / R10 (95%) | - |
| Spotlights and Blindspots: Evaluation Machine-Generated Text Detection Kailash Patil, Kevin Stowe Published: 2026-04-17Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-17 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R4 (95%) | - |
| Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation Changgeon Ko, Hoyun Song, Huije Lee, Jisu Shin Published: 2026-04-18Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R4 (98%) | - |
| Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text Refael Shaked Greenfeld, Reut Tsarfaty Published: 2026-04-18Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R4 (97%) | - |
| Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier Jazmia Henry Published: 2026-04-19Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-19 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E12 / R7 (97%) | - |
| Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation Bing Zhao, Jianqiang Huang, Jiaxin Qi, Zhijiang Tang Published: 2026-04-19Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-04-19 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R9 (95%) | - |
| AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation Fuli Feng, Hui Su, Qi Gu, Wentao Shi Published: 2026-04-20Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-20 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R5 (99%) | - |
| Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies Lorenz Brehme, Ruth Breu, Thomas Str枚hle Published: 2026-04-20Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-04-20 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E10 / R4 (98%) | - |
| MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models Ee-Peng Lim, Neemesh Yadav, Palakorn Achananuparp, Suhyun Lee Published: 2026-04-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E12 / R6 (95%) | - |
| Multilingual Training and Evaluation Resources for Vision-Language Models Andrea Zugarini, Daniela Baiamonte, Elena Fano, Leonardo Rigutini Published: 2026-04-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E13 / R5 (98%) | - |
| On the Importance and Evaluation of Narrativity in Natural Language AI Explanations David Martens, Mateusz Cedro Published: 2026-04-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R4 (98%) | - |
| TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Chuhan Qiao, Haiming Huang, Jinglai Zheng Published: 2026-04-20Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-20 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R5 (98%) | - |
| WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models Chenchen Zhang, Chenyu Zhou, Dailin Li, Han Li Published: 2026-04-20Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-20 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E11 / R5 (98%) | - |
| Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications Abu Noman Md Sakib, Md. Main Oddin Chisty, Zijie Zhang Published: 2026-04-21Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-04-21 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E10 / R4 (97%) | - |
| Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona, Ambuj Kumar, Igor Kozlov Published: 2026-04-21Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-21 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E9 / R4 (98%) | - |
| Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona, Ambuj Kumar, Igor Kozlov Published: 2026-04-21Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-21 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E11 / R6 (97%) | - |
| Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges Ali Al-Kaswan, Arie van Deursen, Maksim Plotnikov, Maliheh Izadi Published: 2026-04-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R5 (98%) | - |
| Evaluation-driven Scaling for Scientific Discovery Caiyin Yang, Chang Su, Chong Gao, Dachao Ding Published: 2026-04-21Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-21 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R5 (98%) | - |
| IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text Rajveer Singh Pall Published: 2026-04-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R4 (98%) | - |
| RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora Hanjun Cho, Jay-Yoon Lee Published: 2026-04-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R5 (98%) | - |
| When Graph Structure Becomes a Liability: A Critical Re-Evaluation of Graph Neural Networks for Bitcoin Fraud Detection under Temporal Distribution Shift Saket Maganti Published: 2026-04-21Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-21 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E7 / R4 (96%) | - |
| ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Jan-Philipp Schmidt Published: 2026-04-22Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-22 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R7 (98%) | - |
| Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation Andrew Klearman, Radu Revutchi, Rishav Chakravarti, Rohin Garg Published: 2026-04-22Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-04-22 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E8 / R4 (95%) | - |
| Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms Ari Azarafrooz Published: 2026-04-22Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-22 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E8 / R4 (95%) | - |
| LAF-Based Evaluation and UTTL-Based Learning Strategies with MIATTs Yongquan Yang Published: 2026-04-22Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-22 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework Christo Zietsman Published: 2026-04-22Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-22 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E11 / R7 (98%) | - |
| TRAVELFRAUDBENCH: A Configurable Evaluation Framework for GNN Fraud Ring Detection in Travel Networks Bhavana Sajja Published: 2026-04-22Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-22 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation Ameen Abu-Hanna, Iacer Calixto, Michele Miranda, Nishant Mishra Published: 2026-04-23Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-23 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E10 / R4 (98%) | - |