Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach Jodi M. Casabianca, Maggie Beiting-Parrish Published: 2026-02-26Area: cs.AICitations: 25 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (89%) | 25 |
| Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation Bohan Liu, Fengfan Zhou, Ruixuan Zhang, Shaokang Wang Published: 2026-02-26Area: cs.CVCitations: - Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-02-26 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E11 / R9 (92%) | - |
| General Agent Evaluation Asaf Yehudai, Elad Venezian, Elron Bandel, Leshem Choshen Published: 2026-02-26Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E19 / R14 (95%) | 45 |
| Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation Aashish Anantha Ramakrishnan, Ardavan Saeedi, Dongwon Lee, Fazlolah Mohaghegh Published: 2026-02-26Area: cs.CLCitations: 12 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R10 (93%) | 12 |
| Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation Bojun Cheng, Dian Xie, Jun Wu, Lichen Bai Published: 2026-02-26Area: cs.CVCitations: 15 Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-02-26 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E10 / R5 (93%) | 15 |
| SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation Feng Jiang, Guibing Guo, Hamid Alinejad-Rokny, Jiahao Zhao Published: 2026-02-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (92%) | - |
| Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions Dongrui Liu, Li Xiong, Qian Chen, Wenjie Wang Published: 2026-02-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (95%) | - |
| An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention Madhusudan Ghosh, Rishabh Gupta Published: 2026-02-25Area: cs.SECitations: 45 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-02-25 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R8 (94%) | 45 |
| Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding Aiden Yiliu Li, Karim Obegi, Shashank Durgad, Wee Joe Tan Published: 2026-02-25Area: cs.AICitations: 20 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (92%) | 20 |
| Evaluation of Audio Language Models for Fairness, Safety, and Security Battista Biggio, Lea Sch枚nherr, Ranya Aloufi, Soumya Shaw Published: 2026-02-25Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-02-25 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E14 / R13 (90%) | 15 |
| Explainability-Aware Evaluation of Transfer Learning Models for IoT DDoS Detection Under Resource Constraints Nelly Elsayed Published: 2026-02-25Area: cs.CRCitations: 15 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-02-25 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E7 / R6 (95%) | 15 |
| FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation Huihang Wu, Jiansong Wan, Jian Xie, Jiayu Guo Published: 2026-02-25Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E32 / R24 (91%) | 15 |
| A Governance and Evaluation Framework for Deterministic, Rule-Based Clinical Decision Support in Empiric Antibiotic Prescribing Diego Moreno, Enrique Javier G贸mez, Francisco Jos茅 G谩rate, Judit L贸pez Luque Published: 2026-02-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E15 / R12 (95%) | - |
| Benchmarking Federated Learning in Edge Computing Environments: A Systematic Review and Performance Evaluation Gil Nicholas Cagande, Sales Aribe Published: 2026-02-24Area: cs.DCCitations: 20 Tags: ai-safety, csdc, preprint, safety-evaluation | 2026-02-24 | cs.DC | ai-safety, csdc, preprint, safety-evaluation | E19 / R17 (92%) | 20 |
| CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation Ayush Sawarni, Jiyuan Tan, Vasilis Syrgkanis Published: 2026-02-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R9 (90%) | - |
| Equitable Evaluation via Elicitation Cynthia Dwork, Elbert Du, Han Shao, Linjun Zhang Published: 2026-02-24Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-24 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (88%) | - |
| Pressure Reveals Character: Behavioural Alignment Evaluation at Depth John Burden, Nora Petrova Published: 2026-02-24Area: cs.AICitations: 15 Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-02-24 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | E18 / R12 (94%) | 15 |
| The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations Mariana Lins Costa Published: 2026-02-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E8 / R0 (93%) | - |
| VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation Changdae Oh, Hyeong Kyu Choi, Sean Du, Seongheon Park Published: 2026-02-24Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-02-24 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R11 (92%) | - |
| Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems Arush Verma, Luigi Medrano, Mukul Chhabra Published: 2026-02-23Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-23 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (93%) | 15 |
| MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems Jin Jia, Yingqi Wang, Zhiling Deng, Zhuangbin Chen Published: 2026-02-23Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-02-23 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R8 (94%) | - |
| Multilevel Determinants of Overweight and Obesity Among U.S. Children Aged 10-17: Comparative Evaluation of Statistical and Machine Learning Approaches Using the 2021 National Survey of Children's Health Joyanta Jyoti Mondal Published: 2026-02-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R9 (94%) | - |
| Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations Alysa Zhao, Ayushi Kishore, Bingzhe Li, Dingyi Kang Published: 2026-02-22Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-22 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R12 (94%) | - |
| DREAM: Deep Research Evaluation with Agentic Metrics Adi Kalyanpur, Amir Dudai, Aviad Aberdam, Changhao Li Published: 2026-02-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (89%) | - |
| MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment Advait Swaminathan, Kevin Zhu, Nguyen Dao Minh Anh, Sagarika Banerjee Published: 2026-02-21Area: cs.CVCitations: 45 Tags: ai-safety, alignment-training, cscv, preprint, safety-evaluation | 2026-02-21 | cs.CV | ai-safety, alignment-training, cscv, preprint, safety-evaluation | E15 / R9 (90%) | 45 |
| Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation Yuan An Published: 2026-02-21Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-21 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E8 / R8 (92%) | - |
| Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm Andras Lukacs, Anna Babarczy, Peter Vedres, Zeteny Bujka Published: 2026-02-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R5 (92%) | - |
| Luna-2: Scalable Single-Token Evaluation with Small Language Models Amey Ramesh Rambatla, Nikhil Ega, Rishon Dsouza, Rob Friel Published: 2026-02-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R9 (89%) | - |
| Towards More Standardized AI Evaluation: From Models to Agents Ali El Filali, In猫s Bedar Published: 2026-02-20Area: cs.CLCitations: 20 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (92%) | 20 |
| A Hybrid Tsallis-Polarization Impurity Measure for Decision Trees: Theoretical Foundations and Empirical Evaluation Edouard Lansiaux, Hayfa Zgaya-Biau, Idriss Jairi Published: 2026-02-19Area: stat.MLCitations: - Tags: ai-safety, preprint, safety-evaluation, statml | 2026-02-19 | stat.ML | ai-safety, preprint, safety-evaluation, statml | E10 / R9 (92%) | - |