Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Can AI Models Direct Each Other? Organizational Structure as a Probe into Training Limitations Rui Liu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E4 / R3 (94%) | - |
| EZASP -- Facilitating the usage of ASP Matthias Knorr, Rafael Martins, Ricardo Gon莽alves Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (98%) | - |
| GISclaw: An Open-Source LLM-Powered Agent System for Full-Stack Geospatial Analysis Jae-Joon Lee, JinByeong Lee, Jinzhen Han, Jisung Kim Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (97%) | - |
| Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence Eziyo Ehsani, Ivano Malavolta, Luca Giamattei, Roberto Pietrantuono Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback Deepak Kumar Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (98%) | - |
| Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Jie Tang, Mingdao Liu, Wenyi Hong, Xiaotao Gu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Jie Tang, Mingdao Liu, Wenyi Hong, Xiaotao Gu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | - |
| Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy Aman Mehta Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |
| Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence Eray T眉z眉n, Vehid Geruslu, Zulfiyya Aliyeva Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |
| IncreRTL: Traceability-Guided Incremental RTL Generation under Requirement Evolution Lei Wang, Luanrong Chen, Renzhi Chen, Rui Gong Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation Benjamin G. Ascoli, Jinho D. Choi, Jiseung Hong Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase Yannick Roy Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| UCAgent: An End-to-End Agent for Block-Level Functional Verification Fangyuan Song, Jinru Wang, Junyue Wang, Sa Wang Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R4 (99%) | - |
| WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing Chenxi Sun, Daling Wang, Fanheng Kong, Han Li Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | - |
| From Untestable to Testable: Metamorphic Testing in the Age of LLMs Valerio Terragni Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E6 / R3 (97%) | - |
| Learning From Developers: Towards Reliable Patch Validation at Scale for Linux Ajay Rawat, Attreyee Mukherjee, Chih-En Lin, Pedro Fonseca Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| Sketch2Simulation: Automating Flowsheet Generation via Multi Agent Large Language Models Abdullah Bahamdan, Antonio del Rio Chanona, Emma Pajak, John D. Hedengren Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R2 (97%) | - |
| SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks Albert Ge, Alexander Yun, Alex Gu, Aws Albarghouthi Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (93%) | - |
| The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents Camilo Chac贸n Sartori Published: 2026-03-25Area: cs.SECitations: 15 Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | 15 |
| TRAJEVAL: Decomposing Code Agent Trajectories for Fine-Grained Diagnosis Baishakhi Ray, Dingmin Wang, Farima Farmahinifarahani, Myeongsoo Kim Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| Willful Disobedience: Automatically Detecting Failures in Agentic Traces Benjamin Zorn, Reshabh K Sharma, Shraddha Barke Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| Can an LLM Detect Instances of Microservice Infrastructure Patterns? Ademar Aguiar, Carlos Eduardo Duarte, Filipe Figueiredo Correia, Neil B. Harrison Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |
| Code Review Agent Benchmark Abhik Roychoudhury, Haifeng Ruan, Imam Nur Bani Yusuf, Ridwan Shariffdeen Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R3 (95%) | - |
| Code Review Agent Benchmark Abhik Roychoudhury, Haifeng Ruan, Imam Nur Bani Yusuf, Ridwan Shariffdeen Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E5 / R4 (98%) | - |
| Evaluating LLM-Based Test Generation Under Software Evolution Mohammad Taha Khan, Muhammad Ali Gulzar, Sabaat Haroon Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops Dylan Bradshaw, Gunel Jahangirova, Ravin Ravi, Stefano Ruberto Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| LLMORPH: Automated Metamorphic Testing of Large Language Models Stefano Ruberto, Steven Cho, Valerio Terragni Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| Machine Learning Models for the Early Detection of Burnout in Software Engineering: a Systematic Literature Review Andrea Capiluppi, Ayushi Rastogi, Tien Rahayu Tulili Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| ReqFusion: A Multi-Provider Framework for Automated PEGS Analysis Across Software Domains Manuel Oriol, Muhammad Khalid, Yilmaz Uygun Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| Early Discoveries of Algorithmist I: Promise of Provable Algorithm Synthesis at Scale Janardhan Kulkarni Published: 2026-03-23Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-23 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |