Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Equitable Evaluation via Elicitation Cynthia Dwork, Elbert Du, Han Shao, Linjun Zhang Published: 2026-02-24Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-24 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (88%) | - |
| Pressure Reveals Character: Behavioural Alignment Evaluation at Depth John Burden, Nora Petrova Published: 2026-02-24Area: cs.AICitations: 15 Tags: ai-safety, alignment-training, csai, preprint, safety-evaluation | 2026-02-24 | cs.AI | ai-safety, alignment-training, csai, preprint, safety-evaluation | E18 / R12 (94%) | 15 |
| The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations Mariana Lins Costa Published: 2026-02-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E8 / R0 (93%) | - |
| VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation Changdae Oh, Hyeong Kyu Choi, Sean Du, Seongheon Park Published: 2026-02-24Area: cs.CVCitations: - Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-02-24 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E12 / R11 (92%) | - |
| An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention Madhusudan Ghosh, Rishabh Gupta Published: 2026-02-25Area: cs.SECitations: 45 Tags: ai-safety, csse, preprint, safety-evaluation | 2026-02-25 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R8 (94%) | 45 |
| Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding Aiden Yiliu Li, Karim Obegi, Shashank Durgad, Wee Joe Tan Published: 2026-02-25Area: cs.AICitations: 20 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R7 (92%) | 20 |
| Evaluation of Audio Language Models for Fairness, Safety, and Security Battista Biggio, Lea Sch枚nherr, Ranya Aloufi, Soumya Shaw Published: 2026-02-25Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-02-25 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E14 / R13 (90%) | 15 |
| Explainability-Aware Evaluation of Transfer Learning Models for IoT DDoS Detection Under Resource Constraints Nelly Elsayed Published: 2026-02-25Area: cs.CRCitations: 15 Tags: ai-safety, cscr, preprint, safety-evaluation | 2026-02-25 | cs.CR | ai-safety, cscr, preprint, safety-evaluation | E7 / R6 (95%) | 15 |
| FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation Huihang Wu, Jiansong Wan, Jian Xie, Jiayu Guo Published: 2026-02-25Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-25 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E32 / R24 (91%) | 15 |
| Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds Vadim Indelman, Yaacov Pariente Published: 2026-02-26Area: math.STCitations: 15 Tags: ai-safety, mathst, preprint, safety-evaluation | 2026-02-26 | math.ST | ai-safety, mathst, preprint, safety-evaluation | E8 / R6 (95%) | 15 |
| Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach Jodi M. Casabianca, Maggie Beiting-Parrish Published: 2026-02-26Area: cs.AICitations: 25 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (89%) | 25 |
| Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation Bohan Liu, Fengfan Zhou, Ruixuan Zhang, Shaokang Wang Published: 2026-02-26Area: cs.CVCitations: - Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-02-26 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E11 / R9 (92%) | - |
| General Agent Evaluation Asaf Yehudai, Elad Venezian, Elron Bandel, Leshem Choshen Published: 2026-02-26Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E19 / R14 (95%) | 45 |
| Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation Aashish Anantha Ramakrishnan, Ardavan Saeedi, Dongwon Lee, Fazlolah Mohaghegh Published: 2026-02-26Area: cs.CLCitations: 12 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-26 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R10 (93%) | 12 |
| Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation Bojun Cheng, Dian Xie, Jun Wu, Lichen Bai Published: 2026-02-26Area: cs.CVCitations: 15 Tags: ai-safety, cscv, preprint, safety-evaluation | 2026-02-26 | cs.CV | ai-safety, cscv, preprint, safety-evaluation | E10 / R5 (93%) | 15 |
| SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation Feng Jiang, Guibing Guo, Hamid Alinejad-Rokny, Jiahao Zhao Published: 2026-02-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R8 (92%) | - |
| Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions Dongrui Liu, Li Xiong, Qian Chen, Wenjie Wang Published: 2026-02-26Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-26 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (95%) | - |
| AI Evaluation Should Require Standardized Item-Level Data Releases Dongyao Zhu, Han Jiang, Sang T. Truong, Sanmi Koyejo Published: 2026-02-27Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-27 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E11 / R11 (89%) | - |
| AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech Akshara Prabhakar, Caiming, Haolin Chen, Huan Wang Published: 2026-02-27Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-02-27 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E14 / R15 (91%) | - |
| A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science Han Li, QingGuo Zhou, Runze Mao, Tianhao Wu Published: 2026-02-27Area: cs.CLCitations: 16 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (93%) | 16 |
| Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection Brianna D'Urso, Syed Rafay Hasan, Tahmid Hasan Sakib, Terry N. Guo Published: 2026-02-27Area: cs.CVCitations: 17 Tags: adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | 2026-02-27 | cs.CV | adversarial-robustness, ai-safety, cscv, preprint, safety-evaluation | E11 / R7 (94%) | 17 |
| Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment Charles L. A. Clarke, Dake Zhang, Mark D. Smucker Published: 2026-02-27Area: cs.IRCitations: - Tags: ai-safety, csir, preprint, safety-evaluation | 2026-02-27 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E7 / R7 (91%) | - |
| When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation Bian Sun, Orvill de la Torre, Zhenjian Wang, Zirui Wang Published: 2026-02-27Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-27 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (94%) | 15 |
| A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction Ruihao Pan, Suhang Wang Published: 2026-02-28Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (93%) | - |
| CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation Linjie Mu, Shaoting Zhang, Xiaofan Zhang, Xizhuo Zhang Published: 2026-02-28Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E8 / R9 (91%) | - |
| Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking Jiayi Liu, Shicong Feng, Tianpeng Zheng, Zhehan Jiang Published: 2026-02-28Area: cs.CLCitations: 33 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-28 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E15 / R14 (93%) | 33 |
| Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma Gabriella Waters, Reva Schwartz Published: 2026-02-28Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-28 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E10 / R9 (91%) | - |
| MOSAIC: A Unified Platform for Cross-Paradigm Comparison and Evaluation of Homogeneous and Heterogeneous Multi-Agent RL, LLM, VLM, and Human Decision-Makers Abdulhamid M. Mousa, Abdulkarim M. Mousa, Jalaledin M. Azzabi, Ming Liu Published: 2026-03-01Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-03-01 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E16 / R15 (90%) | - |
| SimAB: Simulating A/B Tests with Persona-Conditioned AI Agents for Rapid Design Evaluation Alina Rublea, Francisco Chicharro Sanz, Marian Schneider, Mario Truss Published: 2026-03-01Area: cs.HCCitations: 45 Tags: ai-safety, cshc, preprint, safety-evaluation | 2026-03-01 | cs.HC | ai-safety, cshc, preprint, safety-evaluation | E9 / R9 (93%) | 45 |
| An Interactive Multi-Agent System for Evaluation of New Product Concepts Bin Xuan, Hakyeon Lee, Ruo Ai Published: 2026-03-06Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-03-06 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E6 / R4 (96%) | - |