Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops Dylan Bradshaw, Gunel Jahangirova, Ravin Ravi, Stefano Ruberto Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| LLMORPH: Automated Metamorphic Testing of Large Language Models Stefano Ruberto, Steven Cho, Valerio Terragni Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| Machine Learning Models for the Early Detection of Burnout in Software Engineering: a Systematic Literature Review Andrea Capiluppi, Ayushi Rastogi, Tien Rahayu Tulili Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| ReqFusion: A Multi-Provider Framework for Automated PEGS Analysis Across Software Domains Manuel Oriol, Muhammad Khalid, Yilmaz Uygun Published: 2026-03-24Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-24 | cs.SE | ai-safety, csse, preprint | E6 / R4 (98%) | - |
| From Untestable to Testable: Metamorphic Testing in the Age of LLMs Valerio Terragni Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E6 / R3 (97%) | - |
| Learning From Developers: Towards Reliable Patch Validation at Scale for Linux Ajay Rawat, Attreyee Mukherjee, Chih-En Lin, Pedro Fonseca Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| Sketch2Simulation: Automating Flowsheet Generation via Multi Agent Large Language Models Abdullah Bahamdan, Antonio del Rio Chanona, Emma Pajak, John D. Hedengren Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R2 (97%) | - |
| SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks Albert Ge, Alexander Yun, Alex Gu, Aws Albarghouthi Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (93%) | - |
| The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents Camilo Chac贸n Sartori Published: 2026-03-25Area: cs.SECitations: 15 Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | 15 |
| TRAJEVAL: Decomposing Code Agent Trajectories for Fine-Grained Diagnosis Baishakhi Ray, Dingmin Wang, Farima Farmahinifarahani, Myeongsoo Kim Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| Willful Disobedience: Automatically Detecting Failures in Agentic Traces Benjamin Zorn, Reshabh K Sharma, Shraddha Barke Published: 2026-03-25Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-25 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy Aman Mehta Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |
| Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence Eray T眉z眉n, Vehid Geruslu, Zulfiyya Aliyeva Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (95%) | - |
| IncreRTL: Traceability-Guided Incremental RTL Generation under Requirement Evolution Lei Wang, Luanrong Chen, Renzhi Chen, Rui Gong Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E4 / R3 (96%) | - |
| ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation Benjamin G. Ascoli, Jinho D. Choi, Jiseung Hong Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase Yannick Roy Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| UCAgent: An End-to-End Agent for Block-Level Functional Verification Fangyuan Song, Jinru Wang, Junyue Wang, Sa Wang Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R4 (99%) | - |
| WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing Chenxi Sun, Daling Wang, Fanheng Kong, Han Li Published: 2026-03-26Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-26 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | - |
| An Object Web Seminar: A Retrospective on a Technical Dialogue Still Reverbarating James J. Cusick Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E7 / R4 (95%) | - |
| An Object Web Seminar: A Retrospective on a Technical Dialogue Still Reverberating James J. Cusick Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E6 / R3 (94%) | - |
| ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation Chen Tian, Haonan Sun, Lifei Rao, Qincheng Zhang Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint, safety-evaluation | 2026-03-27 | cs.SE | ai-safety, csse, preprint, safety-evaluation | E4 / R2 (95%) | - |
| Automating Domain-Driven Design: Experience with a Prompting Framework Husein Jusic, Stefan Wagner, Tobias Eisenreich Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (96%) | - |
| Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering Chris Brown, Hunter Leary, Swanand Vaishampayan, Yoseph Berhanu Alebachew Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (97%) | - |
| Can AI Models Direct Each Other? Organizational Structure as a Probe into Training Limitations Rui Liu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E4 / R3 (94%) | - |
| EZASP -- Facilitating the usage of ASP Matthias Knorr, Rafael Martins, Ricardo Gon莽alves Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (98%) | - |
| GISclaw: An Open-Source LLM-Powered Agent System for Full-Stack Geospatial Analysis Jae-Joon Lee, JinByeong Lee, Jinzhen Han, Jisung Kim Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (97%) | - |
| Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence Eziyo Ehsani, Ivano Malavolta, Luca Giamattei, Roberto Pietrantuono Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |
| SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback Deepak Kumar Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (98%) | - |
| Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Jie Tang, Mingdao Liu, Wenyi Hong, Xiaotao Gu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (93%) | - |
| Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Jie Tang, Mingdao Liu, Wenyi Hong, Xiaotao Gu Published: 2026-03-27Area: cs.SECitations: - Tags: ai-safety, csse, preprint | 2026-03-27 | cs.SE | ai-safety, csse, preprint | E5 / R3 (94%) | - |