Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Understanding the Limits of Automated Evaluation for Code Review Bots in Practice Baykal Mehmet U莽ar, Eray T眉z眉n, Utku Boran Torun, Veli Karakaya Published: 2026-04-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-27 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E9 / R5 (96%) | - |
| Do Transaction-Level and Actor-Level AML Queues Agree? An Empirical Evaluation of Granularity Effects on the Elliptic++ Graph Ankur Malik Published: 2026-04-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (97%) | - |
| Evaluation of Prompt Injection Defenses in Large Language Models Amy Fox, Kelley McAllister, Krisztian Flautner, Kyle Bacon Published: 2026-04-26Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-26 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E8 / R4 (98%) | - |
| Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task Hiroaki Yamada, Jungmin Choi, Keisuke Sakaguchi Published: 2026-04-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (97%) | - |
| K-SENSE: A Knowledge-Guided Self-Augmented Encoder for Neuro-Semantic Evaluation of Mental Health Conditions on Social Media Vijay Yadav Published: 2026-04-26Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R6 (98%) | - |
| An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code Jelena Ili膰 Vuli膰evi膰 Published: 2026-04-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-25 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E10 / R5 (99%) | - |
| Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Sadman Kabir Soumik Published: 2026-04-25Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R4 (98%) | - |
| ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Aditi Kumaresan, Wenjun Zeng, Yizheng Huang, Zi Wang Published: 2026-04-25Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-25 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E6 / R4 (96%) | - |
| Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline Juergen Dietrich Published: 2026-04-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-04-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E9 / R5 (97%) | - |
| Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Adam Botach, Asaf Gendler, Erez Yosef, Igor Kviatkovsky Published: 2026-04-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R4 (98%) | - |
| Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings Carlos Soares, Hugo Ferreira, Iker Perez, In锚s Oliveira e Silva Published: 2026-04-24Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-24 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (97%) | - |
| Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation Ameen Abu-Hanna, Iacer Calixto, Michele Miranda, Nishant Mishra Published: 2026-04-23Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-23 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E10 / R4 (98%) | - |
| Efficient Agent Evaluation via Diversity-Guided User Simulation Ateret Anaby-Tavor, George Kour, Itay Nakash Published: 2026-04-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework Anna Rumshisky, Aram Galstyan, Charith Peris, Dan Rosen Published: 2026-04-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R4 (98%) | - |
| SQLyzr: A Comprehensive Benchmark and Evaluation Platform for Text-to-SQL M. Tamer 脰zsu, Sepideh Abedini Published: 2026-04-23Area: cs.DBCitations: - Tags: ai-safety, csdb, preprint, safety-evaluation | 2026-04-23 | cs.DB | ai-safety, csdb, preprint, safety-evaluation | E11 / R4 (98%) | - |
| SQLyzr: A Comprehensive Benchmark and Evaluation Platform for Text-to-SQL M. Tamer 脰zsu, Sepideh Abedini Published: 2026-04-23Area: cs.DBCitations: - Tags: ai-safety, csdb, preprint, safety-evaluation | 2026-04-23 | cs.DB | ai-safety, csdb, preprint, safety-evaluation | E11 / R5 (98%) | - |
| Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Minjae Lee, Minji Jung, Minsuk Kahng, Sarang Choi Published: 2026-04-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (97%) | - |
| ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Jan-Philipp Schmidt Published: 2026-04-22Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-22 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R7 (98%) | - |
| Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation Andrew Klearman, Radu Revutchi, Rishav Chakravarti, Rohin Garg Published: 2026-04-22Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-04-22 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E8 / R4 (95%) | - |
| Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms Ari Azarafrooz Published: 2026-04-22Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-22 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E8 / R4 (95%) | - |
| LAF-Based Evaluation and UTTL-Based Learning Strategies with MIATTs Yongquan Yang Published: 2026-04-22Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-22 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework Christo Zietsman Published: 2026-04-22Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-04-22 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E11 / R7 (98%) | - |
| TRAVELFRAUDBENCH: A Configurable Evaluation Framework for GNN Fraud Ring Detection in Travel Networks Bhavana Sajja Published: 2026-04-22Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-22 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R5 (98%) | - |
| Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications Abu Noman Md Sakib, Md. Main Oddin Chisty, Zijie Zhang Published: 2026-04-21Area: cs.HCCitations: - Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-04-21 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E10 / R4 (97%) | - |
| Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona, Ambuj Kumar, Igor Kozlov Published: 2026-04-21Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-21 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E9 / R4 (98%) | - |
| Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona, Ambuj Kumar, Igor Kozlov Published: 2026-04-21Area: cs.CRCitations: - Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-04-21 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E11 / R6 (97%) | - |
| Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges Ali Al-Kaswan, Arie van Deursen, Maksim Plotnikov, Maliheh Izadi Published: 2026-04-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-04-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R5 (98%) | - |
| Evaluation-driven Scaling for Scientific Discovery Caiyin Yang, Chang Su, Chong Gao, Dachao Ding Published: 2026-04-21Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-04-21 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E9 / R5 (98%) | - |
| IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text Rajveer Singh Pall Published: 2026-04-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R4 (98%) | - |
| RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora Hanjun Cho, Jay-Yoon Lee Published: 2026-04-21Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-04-21 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R5 (98%) | - |