Paper deep dive
Repo0: Design-Driven Zero-to-All Code Generation
Silin Chen, Haoyi Teng, Xiaodong Gu, Yuling Shi, Jiale Huang, Yongpan Wang, Hongyu Zhang, Haibing Guan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/21/2026, 3:45:59 AM
Summary
The paper introduces Repo0, a framework for zero-to-all code generation that treats repository architecture as a continuous structural evolution rather than a static blueprint. Repo0 utilizes a Dual-Directed-Acyclic-Graph (Dual-DAG) to maintain alignment between requirement-level functional relationships and component-level implementation dependencies. The system iteratively evolves component boundaries using modularity metrics (cohesion, coupling) to achieve structural convergence before generating code via test-driven development. Evaluated on the RepoCraft benchmark, Repo0 outperforms baselines like RPG in functionality coverage and pass rate.
Entities (9)
Relation Signals (6)
Repo0 â evaluatedon â RepoCraft
confidence 95% ¡ We evaluate Repo0 on six real-world repositories from RepoCraft
Repo0 â uses â Dual-DAG
confidence 95% ¡ Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG)
Repo0 â employs â Cohesion
confidence 90% ¡ cohesion-guided evolution improves repository quality
Repo0 â employs â coupling
confidence 90% ¡ coupling-guided evolution improves repository quality
Repo0 â generatescodeusing â Test-Driven Development
confidence 90% ¡ after which the converged architecture guides test-driven development code generation
Repo0 â outperforms â RPG
confidence 90% ¡ Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
Tags
Links
- Source: https://arxiv.org/abs/2608.19854v1
- Canonical: https://arxiv.org/abs/2608.19854v1
Trouble viewing inline? Open PDF directly â
Full Text
73,951 characters extracted from source content.
Expand or collapse full text
Repo0: Design-Driven Zero-to-All Code Generation Silin Chen1,*, Haoyi Teng1,*, Xiaodong Gu1,â , Yuling Shi1, Jiale Huang1, Yongpan Wang1, Hongyu Zhang2, Haibing Guan1 Affiliation: 1Shanghai Jiao Tong University 2Chongqing University Abstract Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence11 1 Our code and data are available at https://github.com/cslsolow/Repo0. Index Terms: Software engineering agents, Repository Generation, large language models 11footnotetext: Silin Chen and Haoyi Teng contributed equally.22footnotetext: Xiaodong Gu is the corresponding author.â footnotetext: Emails: cslsolow@gmail.com, chunemeng@outlook.com, xiaodong.gu@sjtu.edu.cn, yuling.shi@sjtu.edu.cn, unicornshjl@gmail.com, frankile@sjtu.edu.cn, hbguan@sjtu.edu.cn, and hongyujohn@gmail.com. I Introduction Code generation has seen remarkable progress with the advent of Large Language Models (LLMs), leading to highly capable agents that can synthesize complex code snippets and resolve repository-level issues [21, 56, 2, 53, 36, 7, 32, 23, 45]. However, current methods predominantly operate under a critical assumption: they assume the code has already been architecturally designed. In these structured repository completion settings, agents operate on top of predefined or partially specified repository architectures, where key design decisions, such as component boundaries, package organization, and dependency graphs, are largely given [65]. By focusing on the coding phase, these approaches overlook the crucial role of software architectural design. Fig. 1: Previous methods treat the graph as a fixed planning artifact, whereas Repo0 continuously evolves repository architecture during generation. In this paper, we focus on a more challenging and realistic problem: zero-to-all code generation. We seek to construct software projects without pre-existing repository architecture [31, 37]. This problem is difficult because the agent must jointly infer both the softwareâs functionality and its architecture. A well-designed software project must adhere to core software engineering principles such as high cohesion and low coupling [57]. When a coding agent attempts zero-to-all code generation without explicit architectural guidance, it could violate these modularity goals, resulting in inconsistent component boundaries, fragile dependency structures, and poor cross-file coordination. This suggests that the core bottleneck in zero-to-all code generation is not merely synthesizing code, but establishing and maintaining robust repository modularity throughout the development process. Recent approaches have begun to explore this problem [65, 10, 31, 37, 29]. Approaches like NL2Repo-Bench [10] bypass the design stage by providing a golden repository architecture as part of the input. Multi-agent systems attempt to coordinate planning and implementation through specialized roles [38, 14, 19]. Recent graph-based approaches introduce explicit repository planning [31]. Despite these advances, these approaches suffer from a critical bottleneck: they treat software design as a static artifact. They assume that a perfect architectural blueprint can be generated during a single, initial planning stage and then rigidly executed. In real-world software development, however, project architecture is rarely ready during initial planning. As development progresses, the generated code often reveals whether planned components are truly cohesive or if dependency edges are overly entangled. Emerging complexities, duplicated functionalities, and validation feedback frequently necessitate splitting low-cohesion components or merging redundant ones. Consequently, zero-to-all code generation cannot be treated as a one-shot planning problem; it must be a continuous structural evolution process (as illustrated in Figure 1). Fig. 2: Overview of the continuous decision-driven structural evolution framework. To address this limitation, we propose Repo0, a design-driven structural evolution framework for zero-to-all code generation. Rather than committing to a rigid initial blueprint, Repo0 explicitly models repository generation as a continuous structural evolution problem. Specifically, we maintain an architectural state represented as a Dual-Directed-Acyclic-Graph (Dual-DAG). The state separates requirement-level functional relationships from implementation-level components and dependencies, allowing functional coordination and implementation dependencies to evolve without being conflated in a single static planning graph. This Dual-DAG preserves traceability from the initial user prompt down to the code, while keeping the project architecture open to subsequent structural updates. Starting from a blank state with only natural language requirements, Repo0 executes a full-lifecycle software development process. It first extracts, normalizes, and decomposes the requirements to construct a requirement-level DAG. Then, it incrementally evolves the component-level DAG before code generation and uses the converged architecture to guide repository generation. Through explicit structural actions (i.e., add, split, merge, revise, and save), Repo0 dynamically adjusts component boundaries based on modularity metrics. After structural convergence, validation feedback drives localized repair during code generation. This evolution process is guided by modularity metrics and continues until structural convergence is reached. By treating design and implementation as deeply intertwined, Repo0 represents a paradigm shift in how autonomous agents construct complex software systems from scratch. We evaluate Repo0 on RepoCraft [31, 30], a repository-generation benchmark consisting of six real-world Python repositories, under both GPT-5 mini and DeepSeek V3.2. Results show that Repo0 consistently improves repository-generation quality over baselines. Across all six repositories and both backbone models, Repo0 achieves the highest Functionality Coverage and Pass Rate in all settings. Compared with RPG [31], the strongest repository-planning baseline, Repo0 improves Functionality Coverage by 4.55â20.08 percentage points and Pass Rate by 7.61â29.74 percentage points. Ablation results show that each major design component contributes to end-to-end performance, while the structural-evolution analysis further indicates that metrics-guided convergence is crucial: cohesion- and coupling-guided evolution improves repository quality, whereas unconstrained LLM-decided structural actions tend to over-decompose the repository architecture and degrade downstream correctness. Overall, this paper makes three contributions: 1. We identify a fundamental limitation of existing repository generation methods: they treat repository architecture as a static artifact determined before implementation. We argue that repository modularity is only partially observable from requirements and emerges throughout coding. We formulate repository generation inherently as a continuous structural evolution problem rather than a one-shot planning problem. 2. We reformulate zero-to-all repository generation as a modularity-driven structural evolution process. Instead of committing to a fixed repository plan, repository architecture is continuously evolved through explicit structural actions guided by modularity metrics and requirement-coverage checks until structural convergence is reached. After convergence, validation feedback drives localized repair during code generation. 3. We instantiate this formulation in Repo0, a design-driven repository-generation framework built upon a Dual-DAG representation that separates requirement-level functionalities from implementation-level components and dependencies while maintaining end-to-end traceability across requirements, architecture, and code. I Methodology I-A Problem Formulation We frame zero-to-all repository generation as a problem of continuous structural evolution. Unlike prior methods that first design a repository architecture and then treat it as a fixed blueprint for code generation, Repo0 treats repository design as an adaptive process whose structure is continuously revised as new architectural evidence emerges during requirement decomposition, component construction, code generation, and validation. The central objective is not only to capture complex long-context relationships across a repository for correct code generation, but also to establish and improve repository modularity throughout the development lifecycle. Figure 2 illustrates the overall framework. To support continuous structural evolution, Repo0 maintains a persistent architectural state St=(GtR,GtC,t)S_t=(G_t^R,G_t^C,A_t) for each time step t. =(,)G_t^R=(V_t^R,E_t^R) is a requirement-level DAG where VtRV_t^R represents a requirement unit, including both high-level requirements and their sub-requirements. Edges in this DAG encode requirement-level functional coordination relations instead of implementation dependencies: an edge (u,v)âEtR(u,v)â E_t^R indicates that two requirements are logically used together and should therefore be considered jointly when aligning their behavior, inputs, and outputs. =(,)G_t^C=(V_t^C,E_t^C) is the component-level DAG, where each vâVtCvâ V_t^C represents a component, such as a module, object, service layer, parser, or adapter, that can later be materialized as concrete source code. A component is the basic design unit that represents a bounded implementation responsibility (e.g., a reusable âUnified Modeling APIâ) that will later be materialized into concrete repository artifacts. EtCE_t^C encodes implementation dependencies such as inheritance, reuse, or containment. âĂ A_t V_t^RĂ V_t^C denotes the relations between requirements and components. A pair (q,c)ât(q,c) _t indicates that component c realizes all or part of requirement q. The relation is many-to-many: a requirement may require several components, and a reusable component may support multiple requirements. The alignment relation preserves traceability from requirements to components throughout structural evolution. StS_t is dynamic over time t: while the requirement-level DAG GtRG_t^R generally stabilizes early, the component-level DAG GtCG_t^C and the alignment relation tA_t are continuously updated until modularity is optimized. I-B Phase I: Requirement Decomposition and Initial Architecture The evolution process begins by translating a natural-language requirement document D into the initial architectural state S0S_0. The requirement documents are usually uneven in granularity: some phrases describe broad project goals, while others describe constraints, behaviors, or usage scenarios. Directly mapping such descriptions to design modules often leaves important requirements under-specified and encourages brittle repository architectures. Fig. 3: Illustrative construction of the initial architectural state. To resolve this mismatch, Repo0 first extracts candidate requirement items from D. Specifically, an LLM is prompted to identify capability-level requirements under the following regularizations: each item should describe a distinct repository-level functionality or system constraint, include the relevant operations and sub-features in its description, and avoid splitting low-level operations into separate top-level items. Repo0 then applies a requirement-merge step that consolidates redundant or subsumed items while keeping merely related requirements separate. The remaining merged items are treated as high-level requirements. They serve as top-level decomposition anchors rather than direct file or module specifications. Next, each high-level requirement is decomposed into sub-requirements. A sub-requirement refers to a requirement-side functional unit that clarifies what functionality must be provided without determining packages, files, classes, or component dependencies. Inspired by Atom of Thoughts [49], Repo0 decomposes requirements through a three-stage reasoning-then-labeling process: First, an LLM is prompted to produce as comprehensive a requirement description as possible, elaborating the intended behavior, inputs and outputs, functional constraints, interface expectations, error cases, and ambiguous scope. Based on this enriched requirement description, Repo0 then prompts the LLM to identify sub-requirements and finally label their logical dependency relations. A directed edge from sub-requirement u to sub-requirement v is added when v logically follows u, such as when v relies on the behavior, output, data contract, or interface assumption established by u. The resulting nodes and labeled edges form the local requirement-level DAG for that high-level requirement. Together, the high-level requirements and their decomposed sub-requirements form the initial requirement-level DAG, G0RG_0^R. Once G0RG_0^R is constructed, the system derives the initial components to populate G0CG_0^C and establishes the initial alignment relation 0A_0. Concretely, for each high-level requirement, Repo0 collects its sub-requirements from G0RG_0^R and prompts the LLM to translate them into a set of bounded components. Each proposed component contains a name, a responsibility description, and an explicit list of served sub-requirements. The proposed components form the initial component nodes V0CV_0^C, and each served-sub-requirement entry creates an alignment pair (q,c)â0(q,c) _0. Finally, Repo0 prompts the LLM to infer the implementation dependencies among the proposed components. Requirement-level coordination edges in G0RG_0^R are provided only as soft evidence rather than being directly copied into G0CG_0^C. The inferred dependencies initialize E0CE_0^C. This initialization separates requirement-side artifacts from implementation artifacts. Figure 3 illustrates how Phase I builds the initial architecture for a zero-to-all HttpEasy repository. Starting from D, Repo0 first extracts candidate requirement items and merges overlapping or subsumed items into high-level requirements. In this example, request construction, session persistence, response handling, and authentication are consolidated into two high-level requirements: âHTTP request and response coreâ and âsession and authentication managementâ. For a selected high-level requirement, the LLM then produces an enriched semantic specification that elaborates parameters, headers, body encoding, timeout behavior, error handling, and response objects. Based on this specification, it identifies sub-requirements including constructing a request object, sending the request through a transport adapter, and parsing the returned response, and organizes them into a functional coordination chain from request construction to transport execution and response parsing, thereby instantiating G0RG_0^R. These sub-requirements are then grounded into implementation components such as RequestBuilder, TransportAdapter, and ResponseParser. For instance, RequestBuilder serves the sub-requirement of constructing a request object, which establishes a corresponding entry in the alignment relation 0A_0. Finally, Repo0 infers must-have implementation dependencies among componentsâfor example, SessionClient depends on RequestBuilder and TransportAdapterâto construct G0CG_0^C and yield the initial architectural state S0=(G0R,G0C,0)S_0=(G_0^R,G_0^C,A_0). I-C Phase I: Modularity-Guided Structural Evolution Loop The initial architectural state provides an initial assignment of requirement nodes to components, but may still be overly broad, fragmented, or redundant. Repo0 therefore evolves GtCG_t^C and tA_t through an evolution loop. During this phase, GtRG_t^R is fixed to preserve the functional scope, while GtCG_t^C and tA_t are updated to improve repository modularity. The evolution is realized through four component-level structural actions: ⢠split: For a diffuse component c, Repo0 extracts the induced subgraph of GtRG_t^R over its served sub-requirements SâĄ(c)S(c) and uses graph partitioning [4], viewed as a minimum-cut objective [50], to identify candidate sub-requirement groups. The partition evidence is provided to the LLM, which rewrites c into a set of narrower components, redistributes the corresponding alignment pairs in tA_t, and reconnects incident dependencies in GtCG_t^C. ⢠merge: Consolidates two components into a single component by combining their responsibility descriptions, merging their alignment pairs in tA_t, and redirecting incident dependencies in GtCG_t^C to the merged component. ⢠revise: A boundary-preserving action that rewrites a componentâs responsibility description, interface assumptions, implementation notes, or alignment entries. ⢠save: Marks a component as structurally stable and carries it forward unchanged in the current round unless neighboring structural updates make it eligible again. These structural evolution actions are triggered by two modularity metrics and two auxiliary criteria: Responsibility Responsibility identifies the sub-requirements assigned to a component. For a component c, we define RâSâ(c)=qâVtRâŁ(q,c)ât.RS(c)=\qâ V_t^R (q,c) _t\. (1) The size |RâSâ(c)||RS(c)| measures how many requirement-side units that c is responsible for. Cohesion Cohesion identifies components that group diffuse, unrelated responsibilities [47]. Let Einâ(c)E_in(c) denote the number of edges in GtRG_t^R whose endpoints both lie strictly within RâSâ(c)RS(c). We define cohesion as the density of realized requirement-level functional coordination relations among the requirements grouped into the same component: cohesionâĄ(c)=1,|RâSâ(c)|â¤1Einâ(c)|RâSâ(c)|â(|RâSâ(c)|â1)/2,|RâSâ(c)|>1cohesion(c)= cases1,&|RS(c)|⤠1\\ E_in(c)|RS(c)|(|RS(c)|-1)/2,&|RS(c)|>1 cases (2) Low cohesion indicates a structurally diffuse component. A split action is triggered when a componentâs cohesion falls below a threshold Îłsplit _split and the size of RâSâ(c)RS(c) exceeds a threshold Ďsplit(t) _split^(t). Here, Ďsplit(t) _split^(t) is used to prevent the loop from over-splitting. Coupling Coupling identifies pairs of components that realize highly overlapping sets of sub-requirements [47]. Let RâSARS_A and RâSBRS_B be the sub-requirement sets for which components A and B are responsible, respectively. Coupling is defined as the Jaccard Similarity [41] between these sets: couplingâĄ(A,B)=|RâSAâŠRâSB||RâSAâŞRâSB|coupling(A,B)= |RS_A⊠RS_B||RS_A⪠RS_B| (3) Connectivity Connectivity measures the edges in GtRG_t^R bridging RâSARS_A and RâSBRS_B. A component pair is considered a merge candidate when their coupling exceeds a threshold θmerge _merge and connectivity is greater than 1. These metrics alone may neglect duplicated component responsibilities from valid architectural coupling, such as layered, adapter, or upstream-downstream relationships. Therefore, we ask an LLM to validate the candidates and determine the final merge action. The structural evolution loop proceeds iteratively. In each round, Repo0 recomputes the modularity metrics and auxiliary criteria, then applies eligible structural actions to GtCG_t^C. The iteration terminates when a complete evolution round yields no eligible split or merge actions. At this point, structural convergence is reached with respect to component boundaries. The system then performs a final semantic alignment pass, where the LLM examines each component together with its aligned sub-requirements and neighboring components in the component-level DAG. If inconsistencies are identified between component responsibilities, requirement coverage, or interface assumptions, revise actions are applied. I-D Phase I: Code Generation Once structural convergence is reached, the repository architecture is fixed, and Repo0 enters a code generation phase. The system first converts the converged GtCG_t^C into a concrete generation plan: it derives package assignments, planned file paths, exported symbols, and dependency-aware generation order from GtCG_t^C. For each component, Repo0 constructs a generation context from three sources: the componentâs responsibility description, its aligned requirement nodes recorded in tA_t, and upstream components that must be available before the current component is implemented. This context specifies the componentâs target behavior, placement, public symbols, and reusable dependencies. Code is generated incrementally under a Test-driven development (TDD) workflow [5, 3]. For each component, Repo0 first generates an importable skeleton that fixes the expected public API, then synthesizes tests from the aligned requirement nodes and interface assumptions, and finally fills in the implementation to satisfy those tests. After each generation step, the system runs validation checks, including import checks, interface checks, and pytest-based execution. Failures such as missing symbols, incompatible call signatures, or unmet behavioral expectations trigger localized repair patches to the affected implementation, tests, or package initialization files. When validation feedback reveals that the component description or interface assumptions are inaccurate, Repo0 can apply revise before the next TDD workflow. I Experimental Setup I-A Research Questions We study three research questions: RQ1 (The Overall Effectiveness). How effective is Repo0 at zero-to-all code generation? RQ2 (Ablation Study). How does each core design component of Repo0 contribute to zero-to-all code generation? RQ3 (Structural Evolution Analysis). Do the proposed modularity metrics improve structural convergence compared with LLM-decided structural actions? I-B Datasets TABLE I: Overview of the six repositories and their paraphrased counterparts (Para. Name) in RepoCraft. #Files denotes the total source files, LOC the effective lines of code, and Task Counts the evaluation tasks. Real Repo Para. Name #Files LOC Task Counts scikit-learn MLKit-Py 185 65,972 236 pandas TableKit 217 106,447 175 sympy SymbolicMath 699 218,924 192 statsmodels StatModeler 271 83,325 234 requests HttpEasy 17 2,793 50 django PyWebEngine 681 109,457 165 We evaluate Repo0 on RepoCraft [31], a repository-generation benchmark consisting of six real-world Python repositories. Following RPG [31], repositories are exposed through paraphrased counterpart names to reduce potential pretraining leakage. To better align the benchmark with the specific zero-to-all code generation task, namely, generating a repository from high-level requirements, we further ask two software engineers to revise the original task descriptions [30]. The revised descriptions remove repository-structure cues, implementation-specific architectural hints, and function-level interface details while preserving the original functional requirements. The benchmark includes scikit-learn, pandas, sympy, statsmodels, requests, and django, exposed as MLKit-Py, TableKit, SymbolicMath, StatModeler, HttpEasy, and PyWebEngine, respectively. As summarized in Table I, these repositories span diverse software domains and scales, providing a challenging benchmark for evaluating repository generation across both compact and dependency-intensive systems. I-C Baseline Methods We compare Repo0 with three representative baselines covering direct coding agents, staged repository generation, and graph-based repository planning. We additionally report the original ground-truth repository for each benchmark as a reference. ⢠mini-SWE-agent is a lightweight software engineering agent that solves coding tasks through iterative code editing and command execution [56, 33]. ⢠Paper2Code is a staged multi-agent repository generation framework that decomposes generation into planning, analysis, and implementation phases [42]. ⢠RPG is the state-of-the-art graph-based repository generation method, which constructs a static Repository Planning Graph to guide code generation [31]. ⢠Gold Project is the original repository corresponding to each benchmark task and serves as a reference for validating the evaluation pipeline. I-D Metrics Following RPG [31], we adopt its evaluation pipeline and metric definitions. Functionality Coverage measures the fraction of reference functional categories matched by at least one generated functionality description. Functionality Novelty measures the fraction of generated functionalities that are not matched to any reference category under the same matching procedure. Functionality Accuracy is evaluated using two task-level metrics: Pass Rate, the fraction of benchmark tasks whose adapted ground-truth tests pass on the generated repository, and Voting Rate, the fraction of tasks for which majority-vote semantic evaluation identifies a matched functional interface. To mitigate evaluator bias, we adopt cross-model evaluation for both semantic voting and test-case rewriting. Specifically, DeepSeek V3.2 evaluates the repositories and rewritten test cases generated by GPT-5 mini, while GPT-5 mini evaluates those generated by DeepSeek V3.2. TABLE I: RQ1 main results on three RepoCraft repositories. For each repository, we report Functionality Coverage (Cov.), Functionality Novelty (Nov.), and Pass./Vot., which combines Pass Rate and Voting Rate. GPT-5 mini marks the highest value among methods under GPT-5 mini, and DeepSeek V3.2 marks the highest value among methods under DeepSeek V3.2. Model Method requests statsmodels django Cov. (%) Nov. (%) Pass./Vot. (%) Cov. (%) Nov. (%) Pass./Vot. (%) Cov. (%) Nov. (%) Pass./Vot. (%) GPT-5 mini mini-SWE-agent 68.18 2.63 4.11 / 27.40 18.18 9.09 0.00 / 31.86 47.92 6.96 37.04 / 44.44 Paper2Code 95.50 7.20 24.66 / 24.66 44.32 24.13 4.42 / 30.09 66.67 15.30 30.04 / 78.60 RPG 90.91 13.70 31.51 / 95.89 70.40 13.80 77.90 / 92.00 60.42 11.58 47.33 / 74.07 Repo0 100.00 18.20 50.98 / 100.00 80.68 11.48 85.51 / 98.65 80.50 13.59 74.36 / 97.12 DeepSeek V3.2 mini-SWE-agent 86.36 15.34 21.92 / 47.95 59.09 23.39 2.65 / 53.98 33.33 9.38 10.70 / 47.33 Paper2Code 90.91 11.80 4.11 / 56.16 14.77 5.00 49.56 / 61.95 62.50 43.46 7.82 / 53.50 RPG 95.45 9.23 61.64 / 90.41 64.70 13.70 39.29 / 73.57 68.75 26.50 46.50 / 69.55 Repo0 100.00 24.77 78.08 / 100.00 78.41 14.10 69.03 / 86.46 79.17 14.29 74.07 / 93.83 Human Developer Gold Project 100.00 â 94.12 / 100.00 100 â 94.15 / 100.00 100.00 â 96.34 / 100.00 I-E Implementation Details We use two backbone models in our experiments: the open-source DeepSeek V3.2 [8] and the close-source GPT-5 mini [46]. All generation runs use deterministic decoding with temperature set to 0. For Repo0, we repeat each experimental setting with three independent runs and report the averaged results. For the structural evolution thresholds introduced in Section I, we set the split cohesion threshold to Îłsplit=2/3 _split=2/3 and the merge coupling threshold to θmerge=0.7 _merge=0.7. These values are empirical settings selected on Commit0 Lite [65] using two held-out repositories, wcwidth (small-scale) and sphinx (large-scale), to cover repositories of different complexity. We compared the evolved component-level DAGs against the corresponding golden repository architectures and chose the setting that produced component boundaries closest to the golden architecture by manually comparing the evolved architectures with the corresponding ground-truth repository architectures while avoiding excessive splitting or merging. The held-out repositories are not used in the RepoCraft evaluation. To isolate the effect of repository-structure planning, each method uses its own method-specific procedure to construct the repository structure. Once the repository architecture is produced, all methods share the same downstream code-generation, validation, and repair scaffold under an identical test-driven development (TDD) execution protocol [5, 3]. All evaluation-pipeline hyperparameters follow the default RPG configuration [31]. Due to space constraints, the main experiment reports three representative RepoCraft repositories: requests, statsmodels, and django, which cover lightweight, medium-scale, and large-scale software projects respectively. All prompts and the results on the remaining three repositories are provided in the supplementary material22 2 https://github.com/cslsolow/Repo0/blob/main/supplementary.pdf. IV Results IV-A RQ1: The effectiveness of Repo0 Table I presents the main RQ1 results on three representative RepoCraft repositories. Across both backbone models and repositories of different scales, Repo0 consistently achieves the best overall repository-generation performance. Under GPT-5 mini, Repo0 attains the highest Functionality Coverage on all three repositories, reaching 100.00% on requests, 80.68% on statsmodels, and 80.50% on django. The same pattern appears under DeepSeek V3.2, where Repo0 again achieves the best coverage across all evaluated repositories. This indicates that continuous structural evolution helps realize a broader portion of the required functionality. Beyond covering more required functionality, Repo0 also translates these improvements into stronger implementation correctness. Repo0 also achieves the strongest implementation accuracy in most settings. Under GPT-5 mini, it obtains the highest Pass Rate on all three repositories, outperforming RPG by +19.47 points on requests, +7.61 points on statsmodels, and +27.03 points on django. For Voting Rate, Repo0 achieves the highest results in all six settings. The baselines exhibit complementary weaknesses. mini-SWE-agent struggles to maintain repository-wide consistency on larger repositories, especially in Pass Rate. Paper2Code often produces relatively high Functionality Novelty, but these gains do not consistently translate into stronger correctness. RPG remains the strongest baseline overall, confirming the value of explicit repository-level planning, but it still degrades noticeably on more complex repositories. Since all methods share the same downstream code-generation, validation, and repair pipeline, the primary difference lies in repository architecture. The consistent improvements therefore indicate that continuously refining repository architecture, rather than relying on a one-shot architectural plan, leads to both broader functionality realization and higher implementation correctness. Finding for RQ1 Repo0 consistently achieves the strongest overall repository-generation performance across all baselines by a significant margin. IV-B RQ2: Ablation Study Table I reports the ablation results of Repo0. Removing any individual component degrades performance on at least one repository or evaluation metric, indicating that all four design choices contribute to the final repository-generation quality. Among all ablations, removing Structural Evolution results in the largest overall performance degradation. Across all three repositories, it consistently reduces both functional coverage and implementation correctness, with particularly large drops on requests (Coverage: â5.68-5.68, Pass Rate: â8.47-8.47, Voting Rate: â17.86-17.86) and django (Coverage: â5.92-5.92, Pass Rate: â13.33-13.33, Voting Rate: â8.34-8.34). This suggests that repository generation benefits substantially from continuously evolving the architectural state, rather than treating the initially constructed architecture as fixed. Removing Requirement Context primarily affects implementation correctness. During code generation, the full system conditions each component not only on its directly aligned sub-requirements, but also on the requirement nodes that are functionally coordinated with them through edges in the requirement-level DAG. This additional context helps the model preserve behavior consistency across related functionalities. Removing this contextual information consistently reduces Pass Rate, particularly on django (â10.00-10.00) and requests (â5.59-5.59), indicating that neighboring requirement-level context is important for producing implementations that correctly satisfy interacting functional behaviors. The Component-Graph Ordering mainly influences implementation correctness rather than functional completeness. While Functionality Coverage remains unchanged on both requests and statsmodels, the Pass Rate on statsmodels decreases by 30.0030.00 points without dependency-aware generation order, demonstrating that generation order becomes increasingly important as repository complexity grows. The Dual-DAG representation provides consistent improvements across repositories by separating requirement-level functional coordination from component-level implementation dependencies. Replacing the two-graph representation with a unified graph consistently reduces either functional coverage or implementation accuracy, suggesting that disentangling functional and implementation structures leads to more effective repository planning. A secondary observation is that several ablated variants achieve slightly higher Functionality Novelty on statsmodels. Since this metric measures additional generated functionalities rather than successful realization of the benchmark functionality, the increased novelty together with reduced coverage suggests a tendency to generate extra functionalities at the expense of faithfully implementing the required repository behavior. TABLE I: RQ2 ablation results on requests, statsmodels, and django. Repository Setting Cov. (%) Nov. (%) Pass./Vot. (%) requests Repo0 100.00 18.20 50.98 / 100.00 w/o Requirement Context 95.45 (-4.55) 9.05 (-9.15) 45.39 (-5.59) / 87.86 (-12.14) w/o Component-Graph Ordering 100.00 (0.00) 12.85 (-5.35) 50.98 (0.00) / 100.00 (0.00) w/o Dual-DAG 95.45 (-4.55) 9.14 (-9.06) 48.72 (-2.26) / 87.86 (-12.14) w/o Structural Evolution 94.32 (-5.68) 8.94 (-9.26) 42.51 (-8.47) / 82.14 (-17.86) statsmodels Repo0 80.68 11.48 85.51 / 98.65 w/o Requirement Context 67.92 (-12.76) 11.69 (+0.21) 85.51 (0.00) / 98.65 (0.00) w/o Component-Graph Ordering 80.68 (0.00) 15.57 (+4.09) 55.51 (-30.00) / 88.65 (-10.00) w/o Dual-DAG 78.55 (-2.13) 14.66 (+3.18) 85.51 (0.00) / 95.32 (-3.33) w/o Structural Evolution 75.35 (-5.33) 13.50 (+2.02) 73.51 (-12.00) / 93.65 (-5.00) django Repo0 87.50 13.59 74.36 / 100.00 w/o Requirement Context 78.29 (-9.21) 12.10 (-1.49) 64.36 (-10.00) / 100.00 (0.00) w/o Component-Graph Ordering 83.56 (-3.94) 11.58 (-2.01) 67.70 (-6.66) / 100.00 (0.00) w/o Dual-DAG 82.24 (-5.26) 12.40 (-1.19) 64.36 (-10.00) / 93.33 (-6.67) w/o Structural Evolution 81.58 (-5.92) 10.94 (-2.65) 61.03 (-13.33) / 91.66 (-8.34) Finding for RQ2 All four design components positively contribute to the performance of Repo0, while Structural Evolution yields the largest overall improvements. IV-C RQ3: Structural Convergence Analysis Metrics-Guided Convergence Figure 4 compares Repo0 with two alternatives on statsmodels: (i) w/o evolution, which directly uses the initial architectural state, and (i) LLM-decided structural evolution with fixed budgets of 1, 3, and 5 rounds. Unlike Repo0, these budgeted variants allow the LLM to repeatedly perform structural actions without using the proposed cohesion and coupling metrics to determine which components should evolve or when structural convergence has been reached. In contrast, Repo0 guides split and merge using these metrics and terminates refinement once no metric-triggered structural update remains. Overall, the full Repo0 setting achieves the best performance. Compared with w/o evolution, Repo0 improves all four metrics, increasing Functionality Coverage from 75.90% to 80.68%, Functionality Novelty from 11.05% to 11.48%, Pass Rate from 81.90% to 85.51%, and Voting Rate from 93.00% to 98.65%. More importantly, Repo0 also outperforms all budgeted LLM-decided variants, even though those variants are allowed additional structural action rounds. This suggests that the benefit does not simply come from applying more refinement steps; rather, the cohesion and coupling metrics help the component-level DAG reach structural convergence before code generation. Fig. 4: RQ3 structural-convergence analysis on statsmodels with GPT-5 mini. This behavior is consistent with the design of Repo0. The two modularity metrics, together with the split and merge decision rules defined in Section I, provide an explicit convergence criterion for the evolving component-level DAG. As refinement proceeds, components with diffuse requirement responsibility are split, components with overlapping responsibility are merged, and the loop stops once no further metric-triggered restructuring is needed. This enables subsequent code-generation decisions to operate on a progressively more stable architectural representation. In contrast, LLM-decided structural evolution lacks this explicit convergence criterion. Repeated structural actions can continue beyond the point of structural convergence, leading to unnecessary fragmentation and compensatory restructuring. Although additional LLM-decided refinement may slightly increase novelty, it degrades coverage, pass rate, and voting rate after one round, indicating that unconstrained structural actions can move the repository away from coherent modular boundaries. Action Distribution Figure 5 summarizes the aggregated structural-action distribution during metrics-guided structural evolution across both backbone models. In both settings, split is the dominant action, followed by save, while merge, revise, and add occur less frequently. This indicates that structural evolution is primarily driven by localized updates to component boundaries rather than global changes to the component-level DAG. The relatively high proportion of save further suggests that many initial architectural states already contain stable components and require only limited modification. We further observe systematic differences between backbone models, particularly in the frequency of revise and add. GPT-5 mini triggers these updates more often, and manual inspection confirms that they correspond to valid updates to the architectural state: revise adjusts underspecified component responsibilities and interface assumptions, while add recovers missing requirements or components to improve requirement coverage. In contrast, DeepSeek V3.2 produces fewer such updates, which is consistent with a more complete initial architectural state before structural evolution. To verify whether this difference is driven by the initial architectural state or the backbone model used during structural evolution, we fix the same unoptimized first-round architectural state generated by GPT-5 mini and apply structural evolution using both backbone models. Under this controlled setting, DeepSeek V3.2 produces nearly the same number of revise actions as GPT-5 mini does when evolving the GPT-5-mini-generated initial architectural state (see the supplementary material33 3 https://github.com/cslsolow/Repo0/blob/main/supplementary.pdf). This suggests that the action distribution is largely determined by the initial architectural state generated by the LLM. revise actions are therefore especially important for LLMs that produce vague or underspecified initial component responsibilities and interface assumptions. Fig. 5: Action distributions of different models across the six RepoCraft repositories during structural evolution. Finding for RQ3 Metrics-guided convergence improves repository quality beyond unconstrained LLM-decided structural actions, indicating that explicit modularity metrics help the architecture reach structural convergence. Fig. 6: Case study of Repo0 on StatModeler. The figure shows how requirements are decomposed, aligned with components, updated through structural actions, and materialized into files. IV-D Case Study Figure 6 shows how Repo0 instantiates its Dual-DAG on StatModeler, the RepoCraft counterpart of statsmodels. Starting from a README that specifies capabilities such as regression modeling, generalized linear models, time-series analysis, statistical testing, data backends, result serialization, and model diagnostics, the system first extracts 39 high-level requirements as top-level decomposition anchors. These anchors are decomposed and organized into a requirement-level DAG with 92 sub-requirements, such as Unified API::result-schema, Unified API::model-api, and Data Backend::data-adapters. On the implementation side, these sub-requirements are grounded into 86 implementation components, thereby forming the component-level DAG of the Dual-DAG. Throughout this process, the alignment relation explicitly connects sub-requirements to the components that realize them. Components such as ResultCore, Modeling Core & Result API, and Unified Data Adapters are ultimately materialized as concrete files including StatModeler/unified_api/result_core.py, StatModeler/regression_modeling/modeling_core_result_api.py, and StatModeler/data_backend/unified_data_adapters.py, illustrating how requirement-level structure is carried through to package and file realization. The figure further highlights how the Dual-DAG remains revisable through structural actions during repository generation. Three representative structural-action cases are shown. A split action refines Experimental Namespace Manager into Namespace & Registry Manager and Access Control & Compatibility Manager, making the component boundary more explicit. A merge action consolidates Numerical Core & Sparse Backend and Parallel Execution & External Connectors into Numerical Compute & Execution Runtime, simplifying an over-segmented local structure in the component-level DAG. A revise action keeps the boundary of Covariance Estimation Engine & Integration API fixed while reorganizing its internals into a cleaner algorithm interface, a versioned CovarianceResult schema, and more stable user-facing API contracts. Together, these cases show that Repo0 does not treat the Dual-DAG as a fixed blueprint: both component boundaries and the alignment relation can be updated through structural evolution before final code realization. V Cost Analysis We analyze the monetary cost of repository generation on the three repositories reported in RQ1, namely requests, statsmodels, and django. Across both DeepSeek V3.2 and GPT-5 mini settings, we decompose costs into generation and evaluation phases to isolate where computational overhead arises. Under DeepSeek V3.2, Repo0 incurs generation costs of $11.95, $28.19, and $27.24 on requests, statsmodels, and django, respectively, while evaluation costs dominate for larger repositories, particularly django ($72.82), yielding total costs of $21.28, $41.12, and $100.06. In comparison, RPG exhibits substantially higher generation overhead on requests and statsmodels (+$9.32 and +$35.59, respectively), but lower cost on django (-$8.83), indicating that Repo0 maintains lower end-to-end cost in most reported settings while improving repository-generation quality. Under GPT-5 mini, Repo0 achieves consistently lower generation cost across all repositories, while evaluation costs remain comparable across methods. The complete cost details for all six repositories are provided in the supplementary material44 4 https://github.com/cslsolow/Repo0/blob/main/supplementary.pdf. Further analysis shows that improving cohesion and reducing coupling lowers cost more directly by reducing the amount of redundant or conflicting code that needs to be generated and repaired. As a result, Repo0 requires fewer costly TDD repair iterations caused by ambiguous or overlapping component responsibilities across both backbone models. VI Threats to Validity Internal Validity. A potential threat is that the quality of structural updates depends on the architectural reasoning capability of the backbone LLM. Although the modularity metrics select candidate split and merge actions, the LLM still instantiates these actions over the current architectural state by rewriting component boundaries, responsibility descriptions, alignment entries, and interface assumptions. This threat is partially mitigated by the iterative structural evolution loop: later evolution rounds can re-evaluate the component-level DAG and the alignment relation under the same metrics, while boundary-preserving revise actions can update component descriptions or interface assumptions without changing component boundaries. During TDD-based code generation, validation feedback can further expose inaccurate component descriptions or interface assumptions, triggering localized repair within the fixed repository architecture. External Validity. Our experiments are conducted on six Python repositories from RepoCraft, covering diverse domains and repository scales, but they remain limited to a single programming language and benchmark. The effectiveness of structural evolution on other language ecosystems remains to be investigated. Nevertheless, the proposed architectural state, component-level structural actions, and modularity metrics operate at the level of repository architecture rather than language-specific syntax, suggesting that the framework is not inherently restricted to Python. Evaluating this generality on additional languages and benchmarks is left to future work. VII Related Work VII-A Benchmarks for Repository Generation Repository generation [16, 34, 52, 9, 1, 20, 51, 48] has recently emerged as a distinct benchmark setting beyond function-level code synthesis. DevBench evaluates broader software-engineering workflows, including implementation, testing, and debugging, but does not primarily target zero-to-repository generation [5, 22, 54]. Commit0 and NL2Repo-Bench move closer to repository-scale tasks, yet both still provide clear structural priors: Commit0 asks agents to complete missing code within a predefined repository architecture, file layout, and function interfaces [65], while NL2Repo-Bench includes the golden repository architecture and detailed module organization in its input specifications [10]. RepoCraft, introduced together with RPG, tightens the setting toward complete repository construction from high-level natural language requirements while reducing direct exposure to repository-specific implementation structure [31, 64]. RepoGenesis extends this line to multilingual microservice systems [37], and ProjDevBench studies end-to-end project development by modern coding agents [29]. These benchmarks progressively reduce structural priors available to repository-generation agents, highlighting the need for methods that can organize and refine repository architectures directly from high-level requirements. Our work contributes to this line by using modularity principles to drive iterative optimization of repository architecture during generation. VII-B Agents for Repository Generation Repository-generation agents [13, 63, 11, 55, 62, 58, 15, 60, 59, 61, 12, 18, 25, 6, 43, 44, 28, 40, 27, 24, 39] have evolved from role-based workflows toward increasingly explicit structural reasoning. Early systems such as ChatDev, MetaGPT, and SoA coordinate specialized agents across staged development phases, but still rely primarily on natural language or predefined workflows as the carrier of intermediate planning [7, 26, 38, 14, 19, 35, 32]. Later work introduces more structured planning representations. CodeS decomposes repository generation into repository-, file-, and function-level sketches, enabling hierarchical repository construction through progressively refined planning artifacts [58]. Similarly, Paper2Code employs a multi-stage planning process that derives architectural designs and dependency structures before synthesizing runnable repositories from research papers [42]. EvoMAC further improves adaptability by dynamically evolving the multi-agent collaboration topology according to environmental feedback [17]. Despite these advances, repository architecture is still treated largely as a planning artifact that is generated once and subsequently executed. RPG is the closest prior work to our setting. It introduces a Repository Planning Graph that explicitly represents capabilities, file structures, data flows, and functions, moving repository generation from purely natural-language planning toward graph-guided planning [31]. The key distinction is that RPG treats the planning graph as a blueprint for subsequent generation, whereas Repo0 treats repository architecture as a persistent and evolving state. Guided by modularity metrics, Repo0 continuously updates component boundaries and alignments toward higher cohesion and lower coupling. VIII Conclusion This paper presents Repo0, a continuous structural evolution framework for repository generation from high-level natural-language requirements. Unlike prior methods that rely on a fixed repository plan, Repo0 maintains an explicit architectural state and iteratively refines repository structure to improve modularity throughout generation. Experiments on RepoCraft demonstrate that Repo0 consistently outperforms representative baselines across both open-source and closed-source backbone models. Ablation studies further show that each major component contributes to overall performance, while metrics-guided structural convergence is critical for effective repository evolution. Overall, our findings suggest that successful repository generation requires not only code synthesis, but also explicit architectural reasoning and continuous structural refinement. We hope this work encourages future research on architecture-aware agents for long-horizon software engineering tasks. References [1] S. Abedu, L. Menneron, S. Khatoonabadi, and E. Shihab (2025) RepoChat: an llm-powered chatbot for github repository question-answering. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR), p. 255â259. Cited by: §VII-A. [2] R. Bairi, A. Sonwane, A. Kanade, V. D. C, A. Iyer, S. Parthasarathy, S. Rajamani, B. Ashok, and S. Shet (2023) CodePlan: repository-level coding using llms and planning. External Links: 2309.12499, Link Cited by: §I. [3] K. Beck (2003) Test-driven development: by example. Addison-Wesley Professional. Cited by: §I-D, §I-E. [4] C. Bichot and P. Siarry (2013) Graph partitioning. John Wiley & Sons. Cited by: 1st item. [5] P. Chang, Y. Fang, S. Chen, Y. Shi, B. Shen, and X. Gu (2026) Test vs mutant: adversarial llm agents for robust unit test generation. arXiv preprint arXiv:2602.08146. Cited by: §I-D, §I-E, §VII-A. [6] S. Chen, H. Li, X. Gu, Y. Shi, and H. Guan (2026) SkillForge: self-distilling agents for project-specific issue resolution. External Links: 2608.18933, Link Cited by: §VII-B. [7] S. Chen, S. Lin, Y. Shi, H. Lian, X. Gu, L. Yun, D. Chen, L. Cao, J. Liu, N. Xia, and Q. Wang (2026) SWE-exp: experience-driven software issue resolution. External Links: 2507.23361, Link Cited by: §I, §VII-B. [8] DeepSeek-AI, A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Lu, C. Zhao, C. Deng, C. Xu, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, E. Li, F. Zhou, F. Lin, F. Dai, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Li, H. Liang, H. Wei, H. Zhang, H. Luo, H. Ji, H. Ding, H. Tang, H. Cao, H. Gao, H. Qu, H. Zeng, J. Huang, J. Li, J. Xu, J. Hu, J. Chen, J. Xiang, J. Yuan, J. Cheng, J. Zhu, J. Ran, J. Jiang, J. Qiu, J. Li, J. Song, K. Dong, K. Gao, K. Guan, K. Huang, K. Zhou, K. Huang, K. Yu, L. Wang, L. Zhang, L. Wang, L. Zhao, L. Yin, L. Guo, L. Luo, L. Ma, L. Wang, L. Zhang, M. S. Di, M. Y. Xu, M. Zhang, M. Zhang, M. Tang, M. Zhou, P. Huang, P. Cong, P. Wang, Q. Wang, Q. Zhu, Q. Li, Q. Chen, Q. Du, R. Xu, R. Ge, R. Zhang, R. Pan, R. Wang, R. Yin, R. Xu, R. Shen, R. Zhang, S. H. Liu, S. Lu, S. Zhou, S. Chen, S. Cai, S. Chen, S. Hu, S. Liu, S. Hu, S. Ma, S. Wang, S. Yu, S. Zhou, S. Pan, S. Zhou, T. Ni, T. Yun, T. Pei, T. Ye, T. Yue, W. Zeng, W. Liu, W. Liang, W. Pang, W. Luo, W. Gao, W. Zhang, X. Gao, X. Wang, X. Bi, X. Liu, X. Wang, X. Chen, X. Zhang, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Li, X. Yang, X. Li, X. Chen, X. Su, X. Pan, X. Lin, X. Fu, Y. Q. Wang, Y. Zhang, Y. Xu, Y. Ma, Y. Li, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Qian, Y. Yu, Y. Zhang, Y. Ding, Y. Shi, Y. Xiong, Y. He, Y. Zhou, Y. Zhong, Y. Piao, Y. Wang, Y. Chen, Y. Tan, Y. Wei, Y. Ma, Y. Liu, Y. Yang, Y. Guo, Y. Wu, Y. Wu, Y. Cheng, Y. Ou, Y. Xu, Y. Wang, Y. Gong, Y. Wu, Y. Zou, Y. Li, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Zhao, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Huang, Z. Wu, Z. Li, Z. Zhang, Z. Xu, Z. Wang, Z. Gu, Z. Zhu, Z. Li, Z. Zhang, Z. Xie, Z. Gao, Z. Pan, Z. Yao, B. Feng, H. Li, J. L. Cai, J. Ni, L. Xu, M. Li, N. Tian, R. J. Chen, R. L. Jin, S. S. Li, S. Zhou, T. Sun, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Song, X. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Z. Huang, Z. Xu, Z. Zhang, D. Ji, J. Liang, J. Guo, J. Chen, L. Xia, M. Wang, M. Li, P. Zhang, R. Chen, S. Sun, S. Wu, S. Ye, T. Wang, W. L. Xiao, W. An, X. Wang, X. Sun, X. Wang, Y. Tang, Y. Zha, Z. Zhang, Z. Ju, Z. Zhang, and Z. Qu (2025) DeepSeek-v3.2: pushing the frontier of open large language models. External Links: 2512.02556, Link Cited by: §I-E. [9] C. Dilgren, P. Chiniya, L. Griffith, Y. Ding, and Y. Chen (2025) Secrepobench: benchmarking llms for secure code generation in real-world repositories. arXiv e-prints, p. arXivâ2504. Cited by: §VII-A. [10] J. Ding, S. Long, C. Pu, H. Zhou, H. Gao, X. Gao, C. He, Y. Hou, F. Hu, Z. Li, W. Shi, Z. Wang, D. Zan, C. Zhang, X. Zhang, Q. Chen, X. Cheng, B. Deng, Q. Gu, K. Hua, J. Lin, P. Liu, M. Li, X. Pan, Z. Peng, Y. Qin, Y. Shan, Z. Tan, W. Xie, Z. Wang, Y. Yuan, J. Zhang, E. Zhao, Y. Zhao, H. Zhu, L. Zhu, C. Zou, M. Ding, J. Jiao, J. Liu, M. Liu, Q. Liu, C. Tao, J. Yang, T. Yang, Z. Zhang, X. Chen, W. Huang, and G. Zhang (2026) NL2Repo-bench: towards long-horizon repository generation evaluation of coding agents. External Links: 2512.12730, Link Cited by: §I, §VII-A. [11] L. Fan, J. Liu, Z. Liu, D. Lo, X. Xia, and S. Li (2025) Exploring the capabilities of llms for code-change-related tasks. ACM Transactions on Software Engineering and Methodology 34 (6), p. 1â36. Cited by: §VII-B. [12] S. Gao, W. Zeng, Z. Yu, J. Wangni, C. Wang, K. Cai, S. He, and M. R. Lyu (2026) SWE-mem: learning adaptive memory management for long-horizon coding agents. arXiv preprint arXiv:2606.28434. Cited by: §VII-B. [13] X. Gu, M. Chen, Y. Lin, Y. Hu, H. Zhang, C. Wan, Z. Wei, Y. Xu, and J. Wang (2025) On the effectiveness of large language models in domain-specific code generation. ACM Transactions on Software Engineering and Methodology 34 (3), p. 1â22. Cited by: §VII-B. [14] S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber (2024) MetaGPT: meta programming for a multi-agent collaborative framework. External Links: 2308.00352, Link Cited by: §I, §VII-B. [15] C. Hu, W. Zeng, Y. Shi, B. Shen, and X. Gu (2026) In line with context: repository-level code generation via context inlining. arXiv preprint arXiv:2601.00376. Cited by: §VII-B. [16] R. Hu, C. Peng, J. Xu, and C. Gao (2026) Repo2run: automated building executable environment for code repository at scale. Advances in Neural Information Processing Systems 38, p. 32679â32718. Cited by: §VII-A. [17] Y. Hu, Y. Cai, Y. Du, X. Zhu, X. Liu, Z. Yu, Y. Hou, S. Tang, and S. Chen (2025) Self-evolving multi-agent collaboration networks for software development. In International Conference on Learning Representations, Vol. 2025, p. 23007â23039. Cited by: §VII-B. [18] J. Huang, S. Yun, S. Chen, X. Gu, and B. Shen (2027) Planning over actions: agentic reasoning for semi-structured table question answering. Information Processing & Management 64 (1), p. 105092. Cited by: §VII-B. [19] Y. Ishibashi and Y. Nishimura (2024) Self-organized agents: a llm multi-agent framework toward ultra large-scale code generation and optimization. arXiv preprint arXiv:2404.02183. External Links: Link Cited by: §I, §VII-B. [20] Z. Jiang, L. Deng, J. Cao, M. Pradel, and Z. Liu (2026) Doc2Feat-bench: evaluating documentation-driven feature addition. External Links: 2507.18130, Link Cited by: §VII-A. [21] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan (2024) Swe-bench: can language models resolve real-world github issues?. In International Conference on Learning Representations, Vol. 2024, p. 54107â54157. Cited by: §I. [22] B. Li, W. Wu, Z. Tang, L. Shi, J. Yang, J. Li, S. Yao, C. Qian, B. Hui, Q. Zhang, Z. Yu, H. Du, P. Yang, D. Lin, C. Peng, and K. Chen (2024) Prompting large language models to tackle the full software development lifecycle: a case study. External Links: 2403.08604, Link Cited by: §VII-A. [23] J. Li, H. Zhu, H. Liu, X. Shi, H. Zong, Y. Dong, K. Zhang, S. Jiang, Z. Jin, and G. Li (2025) Aligning llms to fully utilize the cross-file context in repository-level code completion. In 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), p. 1477â1489. Cited by: §I. [24] Y. Li, S. Liu, K. Chen, T. Zhang, and Y. Liu (2025) Impact-driven context filtering for cross-file code completion. External Links: 2508.05970, Link Cited by: §VII-B. [25] H. Lin, S. Chen, X. Gu, Y. Shi, C. Pan, J. Ge, M. Li, J. Huang, M. Chuang, B. Shen, and H. Guan (2026) Know before fix: qa-driven repository knowledge acquisition for software issue resolution. External Links: 2607.11111, Link Cited by: §VII-B. [26] Y. Lin, Y. Ma, R. Cao, B. Li, F. Huang, X. Gu, and Y. Li (2024) Llms as continuous learners: improving the reproduction of defective code in software issues. arXiv preprint arXiv:2411.13941. Cited by: §VII-B. [27] Z. Lin, M. Zhou, Z. Sun, Y. Yang, R. Yang, D. Lo, and L. Li (2026) RepoRescue: an empirical study of llm agents on whole-repository compatibility rescue. External Links: 2607.01213, Link Cited by: §VII-B. [28] Z. Liu, Z. Jiang, Z. Ye, H. Wang, J. Liu, and X. Ren (2026) Effective and efficient context retrieval via partial dependency graph for repository-level code generation. External Links: 2608.01927, Link Cited by: §VII-B. [29] P. Lu, S. Zhang, Y. Hou, L. Ye, C. Huang, Z. Chen, J. Zeng, H. Jiang, P. Liu, Y. Wang, and M. Yang (2026) ProjDevBench: benchmarking ai coding agents on end-to-end project development. External Links: 2602.01655, Link Cited by: §I, §VII-A. [30] J. Luo, C. Yin, X. Zhang, Q. Li, S. Liu, Y. Huang, J. Wu, H. Liu, Y. Huang, Y. Kang, F. Yang, Y. Xin, and S. Li (2026) Closing the loop: universal repository representation with rpg-encoder. External Links: 2602.02084, Link Cited by: §I, §I-B. [31] J. Luo, X. Zhang, S. Liu, J. Wu, J. Liu, Y. Huang, Y. Huang, C. Yin, Y. Xin, Y. Zhan, H. Sun, Q. Chen, S. Li, and M. Yang (2026) RPG: a repository planning graph for unified and scalable codebase generation. External Links: 2509.16198, Link Cited by: §I, §I, §I, 3rd item, §I-B, §I-D, §I-E, §VII-A, §VII-B. [32] D. Ma, S. Chen, Y. Yang, Y. Shi, Y. Yan, and X. Gu (2026) LLM agents can see code repositories. External Links: 2606.14061, Link Cited by: §I, §VII-B. [33] Mini SWE Agent (2026) Mini-swe-agent. Note: https://mini-agent.ai/Accessed: 2026-05-07 Cited by: 1st item. [34] G. Orlanski, D. Roy, A. Yun, C. Shin, A. Gu, A. Ge, D. Adila, N. Roberts, F. Sala, and A. Albarghouthi (2026) SlopCodeBench: benchmarking how coding agents degrade over long-horizon iterative tasks. arXiv preprint arXiv:2603.24755. Cited by: §VII-A. [35] R. Pan, J. Wang, Q. Zhang, Y. Zhu, L. Wu, Z. Yang, Y. Zhang, L. Zhang, and H. Zhang (2026) Persistent cross-attempt state optimization for repository-level code generation. External Links: 2604.03632, Link Cited by: §VII-B. [36] W. Peng, Y. Shi, Y. Wang, X. Zhang, B. Shen, and X. Gu (2025) Swe-qa: can language models answer repository-level code questions?. arXiv preprint arXiv:2509.14635. Cited by: §I. [37] Z. Peng, X. Yin, P. Zhao, F. Yang, L. Wang, R. Jia, X. Chen, Q. Lin, S. Rajmohan, and D. Zhang (2026) RepoGenesis: benchmarking end-to-end microservice generation from readme to repository. External Links: 2601.13943, Link Cited by: §I, §I, §VII-A. [38] C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun (2024) ChatDev: communicative agents for software development. External Links: 2307.07924, Link Cited by: §I, §VII-B. [39] Y. Qin, H. Wang, C. Xu, X. Ma, and J. Lu (2018) Syneva: evaluating ml programs by mirror program synthesis. In 2018 IEEE International Conference on Software Quality, Reliability and Security (QRS), p. 171â182. Cited by: §VII-B. [40] F. Rabbi, Z. Ding, and J. Yang (2026) A multi-language perspective on the robustness of llm code generation. External Links: 2504.19108, Link Cited by: §VII-B. [41] R. Real and J. M. Vargas (1996) The probabilistic basis of jaccardâs index of similarity. Systematic biology 45 (3), p. 380â385. Cited by: §I-C. [42] M. Seo, J. Baek, S. Lee, and S. J. Hwang (2025) Paper2Code: automating code generation from scientific papers in machine learning. arXiv preprint arXiv:2504.17192. External Links: Link Cited by: 2nd item, §VII-B. [43] Y. Shi, Y. Qian, H. Zhang, B. Shen, and X. Gu (2025) LongCodeZip: compress long context for code language models. In 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), p. 141â153. External Links: Link, Document Cited by: §VII-B. [44] Y. Shi, S. Wang, C. Wan, M. Wang, and X. Gu (2025) From code to correctness: closing the last mile of code generation with hierarchical debugging. External Links: 2410.01215, Link Cited by: §VII-B. [45] Y. Shi, J. Xu, K. Fu, W. Zeng, S. He, L. Zhang, Y. Liu, Z. Zhao, T. Y. Zhuo, J. Cao, S. Ye, T. Liu, K. Cai, S. Cheung, and X. Gu (2026) SWE-bench promax: benchmarking agents on large-scale multilingual code refactoring. External Links: 2608.09802, Link Cited by: §I. [46] A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram, A. Nathan, A. Luo, A. Helyar, A. Madry, A. Efremov, A. Spyra, A. Baker-Whitcomb, A. Beutel, A. Karpenko, A. Makelov, A. Neitz, A. Wei, A. Barr, A. Kirchmeyer, A. Ivanov, A. Christakis, A. Gillespie, A. Tam, A. Bennett, A. Wan, A. Huang, A. M. Sandjideh, A. Yang, A. Kumar, A. Saraiva, A. Vallone, A. Gheorghe, A. G. Garcia, A. Braunstein, A. Liu, A. Schmidt, A. Mereskin, A. Mishchenko, A. Applebaum, A. Rogerson, A. Rajan, A. Wei, A. Kotha, A. Srivastava, A. Agrawal, A. Vijayvergiya, A. Tyra, A. Nair, A. Nayak, B. Eggers, B. Ji, B. Hoover, B. Chen, B. Chen, B. Barak, B. Minaiev, B. Hao, B. Baker, B. Lightcap, B. McKinzie, B. Wang, B. Quinn, B. Fioca, B. Hsu, B. Yang, B. Yu, B. Zhang, B. Brenner, C. R. Zetino, C. Raymond, C. Lugaresi, C. Paz, C. Hudson, C. Whitney, C. Li, C. Chen, C. Cole, C. Voss, C. Ding, C. Shen, C. Huang, C. Colby, C. Hallacy, C. Koch, C. Lu, C. Kaplan, C. Kim, C. Minott-Henriques, C. Frey, C. Yu, C. Czarnecki, C. Reid, C. Wei, C. Decareaux, C. Scheau, C. Zhang, C. Forbes, D. Tang, D. Goldberg, D. Roberts, D. Palmie, D. Kappler, D. Levine, D. Wright, D. Leo, D. Lin, D. Robinson, D. Grabb, D. Chen, D. Lim, D. Salama, D. Bhattacharjee, D. Tsipras, D. Li, D. Yu, D. Strouse, D. Williams, D. Hunn, E. Bayes, E. Arbus, E. Akyurek, E. Y. Le, E. Widmann, E. Yani, E. Proehl, E. Sert, E. Cheung, E. Schwartz, E. Han, E. Jiang, E. Mitchell, E. Sigler, E. Wallace, E. Ritter, E. Kavanaugh, E. Mays, E. Nikishin, F. Li, F. P. Such, F. de Avila Belbute Peres, F. Raso, F. Bekerman, F. Tsimpourlas, F. Chantzis, F. Song, F. Zhang, G. Raila, G. McGrath, G. Briggs, G. Yang, G. Parascandolo, G. Chabot, G. Kim, G. Zhao, G. Valiant, G. Leclerc, H. Salman, H. Wang, H. Sheng, H. Jiang, H. Wang, H. Jin, H. Sikchi, H. Schmidt, H. Aspegren, H. Chen, H. Qiu, H. Lightman, I. Covert, I. Kivlichan, I. Silber, I. Sohl, I. Hammoud, I. Clavera, I. Lan, I. Akkaya, I. Kostrikov, I. Kofman, I. Etinger, I. Singal, J. Hehir, J. Huh, J. Pan, J. Wilczynski, J. Pachocki, J. Lee, J. Quinn, J. Kiros, J. Kalra, J. Samaroo, J. Wang, J. Wolfe, J. Chen, J. Wang, J. Harb, J. Han, J. Wang, J. Zhao, J. Chen, J. Yang, J. Tworek, J. Chand, J. Landon, J. Liang, J. Lin, J. Liu, J. Wang, J. Tang, J. Yin, J. Jang, J. Morris, J. Flynn, J. Ferstad, J. Heidecke, J. Fishbein, J. Hallman, J. Grant, J. Chien, J. Gordon, J. Park, J. Liss, J. Kraaijeveld, J. Guay, J. Mo, J. Lawson, J. McGrath, J. Vendrow, J. Jiao, J. Lee, J. Steele, J. Wang, J. Mao, K. Chen, K. Hayashi, K. Xiao, K. Salahi, K. Wu, K. Sekhri, K. Sharma, K. Singhal, K. Li, K. Nguyen, K. Gu-Lemberg, K. King, K. Liu, K. Stone, K. Yu, K. Ying, K. Georgiev, K. Lim, K. Tirumala, K. Miller, L. Ahmad, L. Lv, L. Clare, L. Fauconnet, L. Itow, L. Yang, L. Romaniuk, L. Anise, L. Byron, L. Pathak, L. Maksin, L. Lo, L. Ho, L. Jing, L. Wu, L. Xiong, L. Mamitsuka, L. Yang, L. McCallum, L. Held, L. Bourgeois, L. Engstrom, L. Kuhn, L. Feuvrier, L. Zhang, L. Switzer, L. Kondraciuk, L. Kaiser, M. Joglekar, M. Singh, M. Shah, M. Stratta, M. Williams, M. Chen, M. Sun, M. Cayton, M. Li, M. Zhang, M. Aljubeh, M. Nichols, M. Haines, M. Schwarzer, M. Gupta, M. Shah, M. Y. Guan, M. Huang, M. Dong, M. Wang, M. Glaese, M. Carroll, M. Lampe, M. Malek, M. Sharman, M. Zhang, M. Wang, M. Pokrass, M. Florian, M. Pavlov, M. Wang, M. Chen, M. Wang, M. Feng, M. Bavarian, M. Lin, M. Abdool, M. Rohaninejad, N. Soto, N. Staudacher, N. LaFontaine, N. Marwell, N. Liu, N. Preston, N. Turley, N. Ansman, N. Blades, N. Pancha, N. Mikhaylin, N. Felix, N. Handa, N. Rai, N. Keskar, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, O. Gleeson, P. Mishkin, P. Lesiewicz, P. Baltescu, P. Belov, P. Zhokhov, P. Pronin, P. Guo, P. Thacker, Q. Liu, Q. Yuan, Q. Liu, R. Dias, R. Puckett, R. Arora, R. T. Mullapudi, R. Gaon, R. Miyara, R. Song, R. Aggarwal, R. Marsan, R. Yemiru, R. Xiong, R. Kshirsagar, R. Nuttall, R. Tsiupa, R. Eldan, R. Wang, R. James, R. Ziv, R. Shu, R. Nigmatullin, S. Jain, S. Talaie, S. Altman, S. Arnesen, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Yoo, S. Heon, S. Ethersmith, S. Grove, S. Taylor, S. Bubeck, S. Banesiu, S. Amdo, S. Zhao, S. Wu, S. Santurkar, S. Zhao, S. R. Chaudhuri, S. Krishnaswamy, Shuaiqi, Xia, S. Cheng, S. Anadkat, S. P. Fishman, S. Tobin, S. Fu, S. Jain, S. Mei, S. Egoian, S. Kim, S. Golden, S. Mah, S. Lin, S. Imm, S. Sharpe, S. Yadlowsky, S. Choudhry, S. Eum, S. Sanjeev, T. Khan, T. Stramer, T. Wang, T. Xin, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Degry, T. Shadwell, T. Fu, T. Gao, T. Garipov, T. Sriskandarajah, T. Sherbakov, T. Korbak, T. Kaftan, T. Hiratsuka, T. Wang, T. Song, T. Zhao, T. Peterson, V. Kharitonov, V. Chernova, V. Kosaraju, V. Kuo, V. Pong, V. Verma, V. Petrov, W. Jiang, W. Zhang, W. Zhou, W. Xie, W. Zhan, W. McCabe, W. DePue, W. Ellsworth, W. Bain, W. Thompson, X. Chen, X. Qi, X. Xiang, X. Shi, Y. Dubois, Y. Yu, Y. Khakbaz, Y. Wu, Y. Qian, Y. T. Lee, Y. Chen, Y. Zhang, Y. Xiong, Y. Tian, Y. Cha, Y. Bai, Y. Yang, Y. Yuan, Y. Li, Y. Zhang, Y. Yang, Y. Jin, Y. Jiang, Y. Wang, Y. Wang, Y. Liu, Z. Stubenvoll, Z. Dou, Z. Wu, and Z. Wang (2026) OpenAI gpt-5 system card. External Links: 2601.03267, Link Cited by: §I-E. [47] W. P. Stevens, G. J. Myers, and L. L. Constantine (1974) Structured design. IBM systems journal 13 (2), p. 115â139. Cited by: §I-C, §I-C. [48] Z. Sun, X. Du, Z. Yang, L. Li, and D. Lo (2024) AI coders are among us: rethinking programming language grammar towards efficient code generation. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024, New York, NY, USA, p. 1124â1136. External Links: ISBN 9798400706127, Link, Document Cited by: §VII-A. [49] F. Teng, Q. Shi, Z. Yu, J. Zhang, Y. Luo, C. Wu, and Z. Guo (2026) Atom of thoughts for markov llm test-time scaling. Advances in Neural Information Processing Systems 38, p. 74010â74040. Cited by: §I-B. [50] D. Wagner and F. Wagner (1993) Between min cut and graph bisection. In International Symposium on Mathematical Foundations of Computer Science, p. 744â750. Cited by: 1st item. [51] K. Wang, P. Lan, J. Liu, S. Ren, L. Bao, J. Han, D. Lo, and Z. Liao (2026) Code refinement with repository context: how far are we?. ACM Trans. Softw. Eng. Methodol.. Note: Just Accepted External Links: ISSN 1049-331X, Link, Document Cited by: §VII-A. [52] S. Wang, Z. Wang, D. Ma, Y. Yu, R. Ling, Z. Li, F. Xiong, and W. Zhang (2025) Codeflowbench: a multi-turn, iterative benchmark for complex code generation. arXiv preprint arXiv:2504.21751. Cited by: §VII-A. [53] X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig (2025) OpenHands: an open platform for ai software developers as generalist agents. External Links: 2407.16741, Link Cited by: §I. [54] Y. Wang, Z. Wang, Y. Shi, S. Chen, X. Wang, Y. Wang, B. Shen, L. Li, X. Gu, J. McAuley, and D. D. Zeng (2026) Context compression for llm agents: a survey of methods, failure modes, and evaluation. Preprints. External Links: Document, Link Cited by: §VII-A. [55] M. Wen, J. Chen, R. Wu, D. Hao, and S. Cheung (2018) Context-aware patch generation for better automated program repair. In Proceedings of the 40th international conference on software engineering, p. 1â11. Cited by: §VII-B. [56] J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press (2024) SWE-agent: agent-computer interfaces enable automated software engineering. External Links: 2405.15793, Link Cited by: §I, 1st item. [57] E. Yourdon and L. L. Constantine (1979) Structured design: fundamentals of a discipline of computer program and systems design. Prentice-Hall, Inc.. Cited by: §I. [58] D. Zan, A. Yu, W. Liu, D. Chen, B. Shen, Y. Yao, W. Li, X. Chen, Y. Gong, B. Guan, Z. Yang, Y. Wang, L. Cui, and Q. Wang (2026) CodeS: natural language to code repository via multi-layer sketch. ACM Trans. Softw. Eng. Methodol. 35 (7). External Links: ISSN 1049-331X, Link, Document Cited by: §VII-B. [59] W. Zeng, Y. Shi, X. Gu, C. Hu, C. Wang, Y. Cui, H. Zhou, M. Qi, J. Wangni, Z. Yu, S. Gao, K. Cai, and S. He (2026) Dockerless: environment-free program verifier for coding agents. External Links: 2606.28436, Link Cited by: §VII-B. [60] W. Zeng, Y. Wang, C. Hu, Y. Shi, C. Wan, H. Zhang, and X. Gu (2025) Pruning the unsurprising: efficient code reasoning via first-token surprisal. arXiv preprint arXiv:2508.05988. Cited by: §VII-B. [61] W. Zeng, X. Zhang, Y. Shi, C. Hu, Y. Chen, B. Shen, and X. Gu (2026) Glimprouter: efficient collaborative inference by glimpsing one token of thoughts. arXiv preprint arXiv:2601.05110. Cited by: §VII-B. [62] F. Zhang, B. Chen, Y. Zhang, J. Keung, J. Liu, D. Zan, Y. Mao, J. Lou, and W. Chen (2023) Repocoder: repository-level code completion through iterative retrieval and generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 2471â2484. Cited by: §VII-B. [63] L. Zhang, S. He, C. Zhang, Y. Kang, B. Li, C. Xie, J. Wang, M. Wang, Y. Huang, S. Fu, E. Nallipogu, Q. Lin, Y. Dang, S. Rajmohan, and D. Zhang (2025) SWE-bench goes live!. External Links: 2505.23419, Link Cited by: §VII-B. [64] Y. Zhang, C. Wan, and B. Jin (2016) An empirical study on recovering requirement-to-code links. In 2016 17th IEEE/ACIS International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing (SNPD), p. 121â126. Cited by: §VII-A. [65] W. Zhao, N. Jiang, C. Lee, J. T. Chiu, C. Cardie, M. GallĂŠ, and A. M. Rush (2025) Commit0: library generation from scratch. In International Conference on Learning Representations, External Links: Link Cited by: §I, §I, §I-E, §VII-A.