Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| EduEVAL-DB: A Role-Based Dataset for Pedagogical Risk Evaluation in Educational Explanations Alvaro Ortigosa, Aythami Morales, Francisco Jurado, Javier Irigoyen Published: 2026-02-17Area: cs.AICitations: 32 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R14 (92%) | 32 |
| Quantifying construct validity in large language model evaluations Ryan Othniel Kearns Published: 2026-02-17Area: cs.AICitations: 1 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-17 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R9 (94%) | 1 |
| SEval-NAS: A Search-Agnostic Evaluation for Neural Architecture Search Atah Nuh Mih, Hung Cao, Jianzhou Wang, Truong Thanh Hung Nguyen Published: 2026-02-17Area: cs.LGCitations: 35 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-17 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E12 / R10 (91%) | 35 |
| A Systematic Evaluation of Sample-Level Tokenization Strategies for MEG Foundation Models Chetan Gohil, Mark W. Woolrich, Oiwi Parker Jones, Rukuang Huang Published: 2026-02-18Area: cs.LGCitations: 42 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (89%) | 42 |
| GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation Bertrand Thirion, Hugues Talbot, Nicolas Salvy Published: 2026-02-18Area: cs.LGCitations: 45 Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-18 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E8 / R6 (90%) | 45 |
| IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models Abhinaw Jagtap, Gaurav Azad, Nachiket Tapas, Saurabh Bharti Published: 2026-02-18Area: cs.CLCitations: 9 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-18 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E19 / R20 (94%) | 9 |
| Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents Alec Chiu, Avinash Thangali, Chaitanya Kulkarni, Linsey Pang Published: 2026-02-18Area: cs.AICitations: 15 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-18 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R8 (91%) | 15 |
| A Hybrid Tsallis-Polarization Impurity Measure for Decision Trees: Theoretical Foundations and Empirical Evaluation Edouard Lansiaux, Hayfa Zgaya-Biau, Idriss Jairi Published: 2026-02-19Area: stat.MLCitations: - Tags: ai-safety, preprint, safety-evaluation, statml | 2026-02-19 | stat.ML | ai-safety, preprint, safety-evaluation, statml | E10 / R9 (92%) | - |
| AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games Jos茅 Hern谩ndez-Orallo, Joshua B. Tenenbaum, Kaiya Ivy Zhao, Katherine M. Collins Published: 2026-02-19Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-19 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E7 / R8 (90%) | - |
| Algorithmic Collusion at Test Time: A Meta-game Design and Evaluation Daniel Schoepflin, Xintong Wang, Yuhong Luo Published: 2026-02-19Area: cs.MACitations: 59 Tags: ai-safety, csma, preprint, safety-evaluation | 2026-02-19 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E10 / R8 (92%) | 59 |
| DIALECTIC: A Multi-Agent System for Startup Evaluation Andre Retterath, Georg Groh, Jae Yoon Bae, Joyce Galang Published: 2026-02-19Area: cs.MACitations: 45 Tags: ai-safety, csma, preprint, safety-evaluation | 2026-02-19 | cs.MA | ai-safety, csma, preprint, safety-evaluation | E12 / R10 (88%) | 45 |
| Enhancing Scientific Literature Chatbots with Retrieval-Augmented Generation: A Performance Evaluation of Vector and Graph-Based Systems Amin Kamali, Hamideh Ghanadian, Mohammad Hossein Tekieh Published: 2026-02-19Area: cs.IRCitations: 15 Tags: ai-safety, csir, preprint, safety-evaluation | 2026-02-19 | cs.IR | ai-safety, csir, preprint, safety-evaluation | E14 / R12 (96%) | 15 |
| Fundamental Limits of Black-Box Safety Evaluation: Information-Theoretic and Computational Barriers from Latent Context Conditioning Vishal Srivastava Published: 2026-02-19Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-19 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E10 / R6 (87%) | - |
| Position: Evaluation of ECG Representations Must Be Fixed Collin M. Stultz, Daniel Prakah-Asante, John Guttag, Zachary Berger Published: 2026-02-19Area: cs.LGCitations: - Tags: ai-safety, cslg, preprint, safety-evaluation | 2026-02-19 | cs.LG | ai-safety, cslg, preprint, safety-evaluation | E16 / R10 (94%) | - |
| Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation Alexander L枚ser, Bogdan Kosti膰, Conor Fallon, Julian Risch Published: 2026-02-19Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-19 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R9 (94%) | 15 |
| Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal Ihor Kendiukhov Published: 2026-02-19Area: q-bio.GNCitations: 153 Tags: ai-safety, interpretability, preprint, q-biogn, safety-evaluation | 2026-02-19 | q-bio.GN | ai-safety, interpretability, preprint, q-biogn, safety-evaluation | E10 / R8 (91%) | 153 |
| Toward Trustworthy Evaluation of Sustainability Rating Methodologies: A Human-AI Collaborative Framework for Benchmark Dataset Construction Chekun Law, Peng Qi, Rohit Sharma, Wang Yang Published: 2026-02-19Area: cs.AICitations: 45 Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-19 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E8 / R6 (91%) | 45 |
| Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm Andras Lukacs, Anna Babarczy, Peter Vedres, Zeteny Bujka Published: 2026-02-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E7 / R5 (92%) | - |
| Luna-2: Scalable Single-Token Evaluation with Small Language Models Amey Ramesh Rambatla, Nikhil Ega, Rishon Dsouza, Rob Friel Published: 2026-02-20Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R9 (89%) | - |
| Towards More Standardized AI Evaluation: From Models to Agents Ali El Filali, In猫s Bedar Published: 2026-02-20Area: cs.CLCitations: 20 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-20 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E10 / R8 (92%) | 20 |
| DREAM: Deep Research Evaluation with Agentic Metrics Adi Kalyanpur, Amir Dudai, Aviad Aberdam, Changhao Li Published: 2026-02-21Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-21 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R7 (89%) | - |
| MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment Advait Swaminathan, Kevin Zhu, Nguyen Dao Minh Anh, Sagarika Banerjee Published: 2026-02-21Area: cs.CVCitations: 45 Tags: ai-safety, alignment-training, cscv, preprint, safety-evaluation | 2026-02-21 | cs.CV | ai-safety, alignment-training, cscv, preprint, safety-evaluation | E15 / R9 (90%) | 45 |
| Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation Yuan An Published: 2026-02-21Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-21 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E8 / R8 (92%) | - |
| Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations Alysa Zhao, Ayushi Kishore, Bingzhe Li, Dingyi Kang Published: 2026-02-22Area: cs.CLCitations: - Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-22 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E11 / R12 (94%) | - |
| Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems Arush Verma, Luigi Medrano, Mukul Chhabra Published: 2026-02-23Area: cs.CLCitations: 15 Tags: ai-safety, cscl, preprint, safety-evaluation | 2026-02-23 | cs.CL | ai-safety, cscl, preprint, safety-evaluation | E9 / R7 (93%) | 15 |
| MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems Jin Jia, Yingqi Wang, Zhiling Deng, Zhuangbin Chen Published: 2026-02-23Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-02-23 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E14 / R8 (94%) | - |
| Multilevel Determinants of Overweight and Obesity Among U.S. Children Aged 10-17: Comparative Evaluation of Statistical and Machine Learning Approaches Using the 2021 National Survey of Children's Health Joyanta Jyoti Mondal Published: 2026-02-23Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-23 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E15 / R9 (94%) | - |
| A Governance and Evaluation Framework for Deterministic, Rule-Based Clinical Decision Support in Empiric Antibiotic Prescribing Diego Moreno, Enrique Javier G贸mez, Francisco Jos茅 G谩rate, Judit L贸pez Luque Published: 2026-02-24Area: cs.CYCitations: - Tags: ai-safety, cscy, preprint, safety-evaluation | 2026-02-24 | cs.CY | ai-safety, cscy, preprint, safety-evaluation | E15 / R12 (95%) | - |
| Benchmarking Federated Learning in Edge Computing Environments: A Systematic Review and Performance Evaluation Gil Nicholas Cagande, Sales Aribe Published: 2026-02-24Area: cs.DCCitations: 20 Tags: ai-safety, csdc, preprint, safety-evaluation | 2026-02-24 | cs.DC | ai-safety, csdc, preprint, safety-evaluation | E19 / R17 (92%) | 20 |
| CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation Ayush Sawarni, Jiyuan Tan, Vasilis Syrgkanis Published: 2026-02-24Area: cs.AICitations: - Tags: ai-safety, csai, preprint, safety-evaluation | 2026-02-24 | cs.AI | ai-safety, csai, preprint, safety-evaluation | E9 / R9 (90%) | - |