Paper deep dive
SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents
Qiao Zhao, JianYing Qu, Jun Zhang, Yehua Yang, Hanwen Du, Zhongkai Sun
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 6/21/2026, 2:59:33 AM
Summary
SWE-Future is a novel forecast-conditioned data synthesis method designed to create realistic, future-oriented coding-agent benchmarks while avoiding the data contamination risks associated with replaying historical GitHub issues and pull requests. The method uses a temporal boundary approach: it first uses repository evidence from a pre-snapshot (T0) to forecast 'task families' (feature implementation, bugfix, or refactor). These forecasts are retrospectively validated against later pull requests to ensure they capture real repository evolution. Validated families then serve as conditioning signals to synthesize a 200-task dataset across 61 repositories using a later task-generation snapshot (Tgen), ensuring the tasks are grounded in the repository's evolution without directly replaying historical artifacts.
Entities (6)
Relation Signals (4)
SWE-Future → uses → forecast-conditioned data synthesis
confidence 100% · We propose SWE-Future, a forecast-conditioned data synthesis method for future-oriented coding tasks.
SWE-Future → validatesagainst → post-T0 PR metadata
confidence 100% · We first validate this forecasting step retrospectively: after forecasts are fixed, later pull requests are used only to measure whether the predicted task families match future repository work.
SWE-Future → addresses → Benchmark Contamination
confidence 90% · SWE-Future targets the middle ground... reducing direct dependence on historical pull-request replay.
Task Family → isconditionedby → T0 evidence
confidence 90% · The method uses only pre-T0 repository evidence to forecast future feature implementation/enhancement, bugfix, and refactor task families.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generation, or benchmark-driven model selection. Fully synthetic tasks avoid direct historical replay, but can drift away from real repository needs. We propose SWE-Future, a forecast-conditioned data synthesis method for future-oriented coding tasks. Given a forecast snapshot at time $T_0$, the method uses only pre-$T_0$ repository evidence to forecast future feature implementation/enhancement, bugfix, and refactor task families. We first validate this forecasting step retrospectively: after forecasts are fixed, later pull requests are used only to measure whether the predicted task families match future repository work. In an 80-repository study, the forecaster achieves 58.1\% future-work relevance under the main semantic matching metric. We then use validated forecast families as conditioning signals to synthesize a 200-task coding-agent dataset across 61 repositories from a task-generation snapshot, rather than replaying the later pull requests used for validation. SWE-Future shows that repository-evolution forecasts can guide realistic, future-oriented coding-task synthesis while reducing direct dependence on historical pull-request replay.
Tags
Links
- Source: https://arxiv.org/abs/2606.18733v1
- Canonical: https://arxiv.org/abs/2606.18733v1
Trouble viewing inline? Open PDF directly →
Full Text
50,896 characters extracted from source content.
Expand or collapse full text
SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents Qiao Zhao JianYing Qu Jun Zhang Yehua Yang Hanwen Du Zhongkai Sun Baidu Inc (June 9, 2026) Abstract Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generation, or benchmark-driven model selection. Fully synthetic tasks avoid direct historical replay, but can drift away from real repository needs. We propose SWE-Future, a forecast-conditioned data synthesis method for future-oriented coding tasks. Given a forecast snapshot at time T0T_0, the method uses only pre-T0T_0 repository evidence to forecast future feature implementation/enhancement, bugfix, and refactor task families. We first validate this forecasting step retrospectively: after forecasts are fixed, later pull requests are used only to measure whether the predicted task families match future repository work. In an 80-repository study, the forecaster achieves 58.1% future-work relevance under the main semantic matching metric. We then use validated forecast families as conditioning signals to synthesize a 200-task coding-agent dataset across 61 repositories from a task-generation snapshot, rather than replaying the later pull requests used for validation. SWE-Future shows that repository-evolution forecasts can guide realistic, future-oriented coding-task synthesis while reducing direct dependence on historical pull-request replay. 1 Introduction Software engineering agents need training and evaluation tasks that are realistic, repository-specific, and aligned with how active projects evolve. Many strong coding-agent datasets and benchmarks get this realism by replaying public GitHub issues and pull requests [6, 13, 8, 4]. That strategy is valuable, but it creates a growing validity problem: public repositories, benchmark releases, and evaluation traces may appear in pretraining, fine-tuning, synthetic-data generation, or model-selection loops. For widely tracked coding-agent benchmarks, direct historical replay therefore carries data-contamination risk for rapidly evolving coding models [11, 9, 1]. Pure synthetic generation can reduce direct replay, but it often loses repository pressure. A synthetic task may be plausible in isolation while ignoring project conventions, maintainer priorities, realistic dependency constraints, or the kinds of tests that the repository would actually accept. SWE-Future targets the middle ground. Instead of replaying historical pull requests or asking an agent to invent arbitrary tasks, it first predicts likely future repository evolution from pre-snapshot evidence. The validated prediction unit is a task family, not an exact PR title. A family can describe, for example, “Zarr/backend capability growth in xarray”, “false-positive fixes in pylint”, or “callback API extension in Dash”. The final synthesis step then uses a task-generation snapshot, denoted TgenT_gen, and the validated family as a conditioning signal to construct future-oriented coding tasks. The method has four stages. First, for each repository snapshot at time T0T_0, SWE-Future builds an evidence bundle from pre-T0T_0 issues, pull requests, labels, and text. Second, it forecasts feature implementation/enhancement, bugfix, and refactor task families from that evidence. Third, in a retrospective study only, it freezes the forecasts and compares them with later pull-request metadata to measure whether the predicted families match future repository work. Fourth, at TgenT_gen, it uses validated families as conditioning signals to synthesize coding tasks from the repository state visible to the task constructor, without directly replaying historical pull requests. Figure 1 in the method section summarizes this information boundary. This paper makes three contributions. First, it introduces forecast-conditioned data synthesis as a construction method for realistic, future-oriented coding-agent tasks. Second, it gives a retrospective validation procedure for testing whether forecast families capture future repository work; in an 80-repository study, the forecaster reaches 58.1% future-work relevance under the main semantic matching metric. Third, it demonstrates forecast-to-task conversion by grounding validated families in TgenT_gen repository snapshots and producing a 200-task coding-agent dataset across 61 repositories with executable validation. 2 Related Work Historical GitHub task construction. SWE-Bench established repository-level issue resolution as a standard coding agent evaluation setting by drawing tasks from real GitHub issues and their corresponding pull requests [6]. SWE-Bench++ scales this family of GitHub-derived evaluation across more repositories and languages, and FEA-Bench narrows the target to feature implementation by collecting feature-focused pull requests and relevant tests [13, 8]. SWE-Factory further automates GitHub issue-resolution data construction through test recovery, environment building, and fail-to-pass validation [4]. These systems provide realism because the target work has already happened. SWE-Future instead uses post-T0T_0 pull requests only as hidden retrospective evidence for frozen forecasts; public tasks are generated from a task-generation repository snapshot and a validated forecast family. Benchmark contamination and leaderboard pressure. Public coding benchmarks are useful precisely because they create shared targets for model developers. That same visibility creates contamination risk: benchmark repositories, issue text, pull-request metadata, reference solutions, or derivative evaluation traces can enter pretraining, fine-tuning, synthetic-data, or model-selection pipelines [11, 9, 1]. This risk is especially salient for widely used GitHub-derived benchmarks such as SWE-Bench: model developers can track public leaderboards, optimize prompts or data against benchmark behavior, and repeatedly inspect public source projects and task texts. Our goal is not to prove that model parameters have never seen a repository; it is to avoid direct replay of observed historical pull requests as task material. Task factories and synthetic environments. SWE-Hub and daVinci-Env focus on large-scale environment construction and executable task generation, while SWE-Playground synthesizes projects, tasks, tests, and implementations from scratch to improve scale and control [14, 2, 15]. These systems reduce the engineering cost of runnable tasks; SWE-Future targets the upstream question of which repository-specific task families should be synthesized before task construction, using forecast validity as the conditioning signal. Software evolution and temporal validity. Software engineering has long used repository histories to predict risky modules, defect-prone changes, and likely co-changes [3, 10, 5, 7, 16]. Recent time-consistent benchmark work also emphasizes temporal discipline by snapshotting a repository at T0T_0, constructing repository knowledge from pre-T0T_0 artifacts, and evaluating on tasks derived from later pull requests [12]. SWE-Future borrows this temporal boundary but changes the target: it forecasts coarse future work families and uses those families to guide synthesis at TgenT_gen. Together, these lines of work leave a gap. Historical GitHub replay supplies realism but increases direct contamination pressure; fully synthetic task factories improve scale and control but can lose repository-specific demand; and time-consistent evaluation still often relies on later realized work as the task itself. SWE-Future addresses this gap by using later work only to validate forecast families, then synthesizing tasks from a task-generation snapshot conditioned on those families. 3 Method Figure 1: SWE-Future overview. The top lane validates frozen forecast families: pre-T0T_0 repository evidence produces task-family forecasts, and post-T0T_0 PR metadata is used only to judge semantic matches. The bottom lane synthesizes tasks from the TgenT_gen repository snapshot and validated directions, then releases only tasks with target-test and gold-patch executable evidence. SWE-Future is a forecast-conditioned data synthesis method for constructing repository-specific coding-agent tasks without directly replaying realized pull requests. For each repository R, the method separates three time points: a forecast snapshot T0T_0, a retrospective validation horizon (T0,T1](T_0,T_1], and a task-generation snapshot TgenT_gen. Evidence at or before T0T_0 is used to predict future task families; post-T0T_0 PR metadata is used only to validate whether those frozen forecasts correspond to later repository work; and task generation is performed from the repository state visible at TgenT_gen. Figure 2 makes this temporal boundary explicit. Figure 2: Temporal information boundary in SWE-Future. Forecasting uses only repository evidence visible at or before T0T_0; retrospective validation uses post-T0T_0 PR metadata only to judge frozen forecast families; task synthesis is grounded in the TgenT_gen repository snapshot, while task-specific tests and gold patches remain hidden evaluation assets. The method has four stages. It first builds task-family forecasts from pre-T0T_0 repository evidence. It then validates the frozen families against post-T0T_0 PR metadata using semantic matching. Next, it uses validated families as conditioning signals to synthesize issue-style public task specifications from the TgenT_gen repository snapshot. Finally, it admits only tasks that pass target-test construction, leakage checks, FAIL_TO_PASS, and selected PASS_TO_PASS validation. The forecast unit is a family rather than a realized pull request: fi=(category,family anchor,expected behavior,evidence refs,target hints,acceptance criteria).f_i=(category,family anchor,expected behavior,evidence refs,target hints,acceptance criteria). A family describes a repository demand direction, such as a recurring bug class, an expected feature implementation/enhancement, or a refactoring need. In this paper we keep feature implementation/enhancement, bugfix, and refactor families, and exclude dependency/platform updates, CI/release work, documentation-only changes, formatting-only edits, and test-only cleanup. This scope keeps the method focused on coding-agent data synthesis rather than general maintenance activity prediction. 3.1 Forecasting and Retrospective Validation Repository pool. The retrospective validation pool contains 80 repositories. We begin from Python repositories covered by or adjacent to SWE-bench and the SWE-bench harness, then expand to active Python and software-engineering ecosystem projects across testing, linting, packaging, web and networking, data and scientific computing, AI and agent frameworks, and infrastructure. Candidate repositories are first filtered by merged-PR activity in the common retrospective window, because forecast validation requires enough post-T0T_0 repository work to compare against frozen forecasts. We then require local GitHub history payloads containing pre-T0T_0 issue and PR signals for forecasting and post-T0T_0 PR metadata for retrospective validation. The remaining candidates are ranked by the density of post-T0T_0 work in the retained categories: feature implementation/enhancement, bugfix, and refactor, together with the amount of usable pre-T0T_0 evidence. Dependency/platform, CI/release, documentation-only, formatting-only, and test-only activity is not treated as positive evidence unless it also contains direct bug or regression evidence. Repositories with very few post-T0T_0 PRs or very few in-scope post-T0T_0 PRs are down-weighted. The final pool is selected to provide enough post-T0T_0 work for retrospective validation and to maintain domain diversity. It is not meant to be an unbiased sample of GitHub repositories. All repositories use the same retrospective window: T0=2025-11-28,T1=2026-05-28.T_0=2025-11-28, T_1=2026-05-28. Forecast construction. Forecast construction uses four explicit steps, all restricted to information visible at or before T0T_0. The forecasting assumption is persistence-based: repeated pre-T0T_0 demand around the same repository-specific anchor is treated as evidence of a continuing work direction. The method extrapolates that direction into a future-oriented family specification rather than a concrete PR, patch, or task. Step 1: Collect pre-T0T_0 signals. For each repository, the forecaster reads issue and merged-PR records visible before T0T_0, using labels, titles, and short body excerpts as evidence. It applies rule-based exclusion filters to remove dependency bumps, platform updates, CI/release work, documentation-only edits, formatting-only edits, and test-only cleanup. If a record contains a maintenance keyword but also direct bug evidence, such as regression, crash, incorrect behavior, or exception, it is kept for the bugfix channel. The remaining records are assigned to feature implementation/enhancement, bugfix, or refactor using category keyword patterns over the title, labels, and body excerpt. Step 2: Cluster signals into candidate directions. The forecaster extracts project-specific anchor terms from labels, titles, module mentions, and short body excerpts. Generic project-management words, repository boilerplate, URLs, overly long path fragments, and noisy anchors are removed. Each retained record contributes its strongest anchors, and the first three anchors are used to form clusters keyed by (category,anchor)(category,anchor). A candidate cluster is considered only when it contains at least two pre-T0T_0 signals. This requirement prevents an isolated mention from becoming a forecast family by itself. Step 3: Score and filter clusters. Candidate clusters are ranked with a fixed heuristic score: s=n+0.45nissue+0.25nprePR+0.05nlabel,s=n+0.45n_issue+0.25n_prePR+0.05n_label, where n is the number of clustered signals, nissuen_issue is the number of issue signals, nprePRn_prePR is the number of pre-T0T_0 merged-PR signals, and nlabeln_label is the number of distinct labels in the cluster. These weights are not learned parameters; they are a deterministic ranking rule used in this study. Term Weight Rationale n 1.0 Base evidence volume from repeated signals. nissuen_issue +0.45 Issues state user-visible demand. nprePRn_prePR +0.25 Pre-T0T_0 merged PRs show project precedent. nlabeln_label +0.05 Label diversity is used only as a weak tie breaker. Clusters below 2.8 are discarded. The remaining clusters are sorted by score, capped at five families per repository, and capped at two families per category. The confidence shown for a family is a deterministic display score derived from the cluster score, min(0.91,0.54+0.045s) (0.91,0.54+0.045s), rather than a learned probability. Step 4: Write each cluster as a forecast family. For each selected cluster, the forecaster instantiates a category-specific template using the cluster anchor: feature implementation/enhancement clusters become capability, API, provider, or format-extension families; bugfix clusters become recurring failure, edge-case, regression, or incorrect-behavior families; and refactor clusters become behavior-preserving restructuring families. The emitted family contains the category, family anchor, expected behavior, acceptance checklist, target hints, supporting pre-T0T_0 signal references, and non-goals. Appendix A lists the templates used for this cluster-to-family step, and Appendix B gives concrete forecast examples. The forecaster emitted 260 families across 76 of 80 repositories; the remaining four repositories abstained because no cluster crossed the evidence threshold. Each emitted family is a structured forecast specification rather than a task, and it is the only forecast-side input available to later task synthesis. Category Forecast families Share Bugfix 139 53.5% Feature impl./enh. 93 35.8% Refactor 28 10.8% All 260 100.0% Table 1: Feature implementation/enhancement, bugfix, and refactor family forecast yield on the 80-repository eligibility pool. Retrospective validation. After forecasts are frozen, retrospective validation checks whether they match later repository work in the six-month window (T0,T1](T_0,T_1], ending at T1=2026-05-28T_1=2026-05-28. The endpoint T1T_1 closes the validation window; it is not a point after which new validation evidence is gathered. Validation has two stages. Retrieval. For each forecast family, a matcher retrieves post-T0T_0 PR candidates from the same repository. It applies the same retained-category scope filter: dependency, platform, CI/release, documentation-only, formatting-only, and test-only PRs are removed unless their text also contains direct bug or regression evidence. A post-T0T_0 PR remains a retrieval candidate only if it shares the forecast category or contains the forecast anchor. Candidate pairs are ranked by m(f,p)=0.45J(f,p)+0.12Stitle(f,p)+Bcat(f,p)+Banchor(f,p),m(f,p)=0.45J(f,p)+0.12S_title(f,p)+B_cat(f,p)+B_anchor(f,p), where J is token Jaccard overlap between the forecast-family text and PR metadata, StitleS_title is title similarity, BcatB_cat is 0.18 for category agreement and 0.04 for a classifiable category mismatch, and Banchor=min(0.34,0.12h)B_anchor= (0.34,0.12h) for h anchor or target-module hits in the PR text. Retrieval forms a top-3 shortlist; semantic validation uses the top-1 candidate. Post-T0T_0 patches are never shown to the judge. Judging. The primary semantic judge labels each frozen forecast and top-1 post-T0T_0 PR metadata on a single relevance dimension with three ordered tiers: strong, related, and rejected. Strong means the forecast directly matches later work, including exact, partial, and event-family hits. Related means the forecast is not the same later task but still matches the category or module direction. Rejected covers no-match and unclear cases. The main metric is forecast-family relevance: #strong+#related#forecast families, \#strong+\#related\#forecast families, which measures synthesis eligibility rather than exact post-T0T_0 PR prediction. This metric counts related families because the downstream goal is not to predict the exact post-T0T_0 PR, but to identify repository-natural directions that can condition task synthesis. A related family is therefore useful when it points to the same category or module pressure even if it does not match a single later PR exactly. The resulting forecast-family relevance is 151/260 (58.1%). The stricter strong hit rate is 111/260 (42.7%). There are 61 repositories with at least one strong+related family. Bugfix families are currently the strongest source of eligible synthesis directions, with 89/139 strong+related; feature implementation/enhancement families are harder, with 45/93 strong+related. An independent semantic audit over all 260 forecasts agrees with the primary relevance-vs-rejected decision in 216/260 cases (83.1%). We use this agreement as a consistency check, while keeping 151/260 as the main future-work relevance metric. Figure 3: Retrospective validation outcome for 260 forecast families. Bars show strong, related, and rejected top-1 semantic matches against post-T0T_0 merged PR metadata in (T0,T1](T_0,T_1]; the main relevance metric counts strong+related families. 3.2 Task Generation and Data Construction Forecast families that pass the retrospective relevance check, either as strong or related matches, become conditioning signals for task synthesis at TgenT_gen. The forecast family supplies the repository-specific direction; inspection of the TgenT_gen codebase turns that direction into a concrete change request. This separation is central to the construction. Post-T0T_0 pull requests are used only at the forecast-family validation stage, after forecasts are frozen, to test whether a predicted direction matches later repository work. They are a family-level validation signal, not a task source. The 200 released tasks do not have target PRs: each task is synthesized from the TgenT_gen repository snapshot and repository evidence visible to the constructor. The generator therefore uses retrospective validation only to select reliable directions, not to replay or transform post-T0T_0 PR artifacts into prompts, tests, or gold patches. 3.2.1 From Forecast Families to Tasks The synthesis stage materializes a 200-task coding-agent dataset across 61 repositories. The 151 strong-or-related forecast families do not map one-to-one onto tasks. A family can contribute multiple task rows when TgenT_gen code inspection exposes separable behavior gaps, capability extensions, or refactor targets; it can also contribute no released row if no executable oracle survives construction. The 200-task mix is used as a candidate-construction target, not a release shortcut: the queue is built with caps per repository and per forecast family, with the related tier capped so that strong matches dominate the pool. Rows that fail target audit, leakage review, or executable validation are revised or replaced rather than admitted to satisfy a quota. We allocate 120 bugfix, 60 feature implementation/enhancement, and 20 refactor tasks because bugfix and feature directions more often yield behavior-level oracles, while refactors are harder to validate with focused executable checks. Figure 4: Forecast-to-task selection funnel. The 80-repository pool yields 260 forecast families from 76 repositories; 151 families validate as strong or related under semantic judging and condition the 200 synthesized task set through gate-based admission. Task synthesis follows the same three-stage route for every row: 1. Select a validated direction. Choose a strong or related forecast family, keep its pre-T0T_0 evidence references as conditioning context, and load the corresponding task-generation snapshot, meaning the repository state visible to the constructor rather than validation artifacts outside that snapshot. 2. Ground it in the TgenT_gen codebase. Inspect code, public APIs, tests, and project conventions to identify a concrete behavior gap, capability extension, or behavior-preserving refactor that is natural for that project. 3. Release only executable tasks. Write a normal issue-style request, construct the test patch and gold patch, and admit the task only after leakage review, FAIL_TO_PASS validation, and selected preservation checks. The grounding and executable-release stages are organized as a multi-agent workflow, with separate artifacts for context selection, task writing, oracle design, patch construction, and verification. The loop starts from a validated family and searches the TgenT_gen repository tree for implementation and test candidates using the family anchor, target-module hints, task category, and variant tokens. If no plausible implementation or test target is found, the row is sent back for repository inspection rather than patch generation. Next, a task-writing role produces the model-visible issue-style request from the forecast family and TgenT_gen evidence only. This role does not see post-T0T_0 PR metadata, judge rationales, test-patch drafts, or gold-patch content. An oracle-design role then converts the request into a concrete target behavior, a test location, and a test-patch plan; it can reject tasks that are too broad, already solved in the snapshot, or not testable with a focused oracle. A patch-construction role writes the test patch first and then a gold patch scoped to the smallest implementation path. Finally, a verification role starts from a clean checkout, checks that both patches apply, runs the FAIL_TO_PASS sequence, runs selected preservation checks, and either admits the row or returns it to the earlier roles for repair. Table 2 summarizes the artifacts, and Figure 5 summarizes the construction workflow and its public/hidden artifact boundary. Appendix C details the role strategies, and Appendix D shows a complete synthesized task row. This separation matters because the task writer should not know the realized post-T0T_0 fix or the hidden oracle. Without that barrier, the public task could be phrased around a known solution and become a disguised historical replay. Task-specific tests, gold patches, provenance records, validation labels, and execution logs are therefore kept as hidden construction and evaluation assets. Role Input Output or gate Context selection Forecast family and task-generation snapshot Implementation/test candidates and target-module evidence Task writing Forecast family plus TgenT_gen repository evidence Model-visible issue statement without validation metadata, judge rationales, or test/gold patch material Oracle design Public task, target files, and TgenT_gen code excerpts Behavior under test, test-patch plan, gold-patch scope, or rejection Patch construction Oracle plan and task-generation snapshot Task-specific test patch and gold patch Verification Public task, test patch, and gold patch FAIL_TO_PASS result plus selected PASS_TO_PASS checks Table 2: Multi-agent construction workflow for grounding and executable release. Figure 5: Multi-agent task construction. Validated directions and the TgenT_gen snapshot enter a construction loop with separate context selection, task-writing, oracle-design, patch-construction, and verification roles. Only the public task and repository snapshot are model-visible; target tests, gold patches, and verification logs remain hidden evaluation assets. Dataset view Tasks Notes Bugfix 120 Correctness and edge-case repair tasks Feature impl./enh. 60 Capability, API, format, and option-support tasks Refactor 20 Behavior-preserving internal improvement tasks Strong-family conditioned 160 Conditioned on retrospective strong matches Related-family conditioned 40 Conditioned on retrospective related matches All tasks 200 TgenT_gen task specifications with executable validation assets Table 3: SWE-Future dataset composition. Tasks are generated from TgenT_gen repository evidence and conditioned on validated forecast families. Repository Category Forecast family Public task example agronholm/anyio Bugfix BlockingPortal edge-case failures Forward provider shutdown exceptions through a focused regression and minimal BlockingPortal cleanup-path fix. celery/kombu Feature Redis broker capability growth Honor the Redis global key-prefix option for expiration commands so broker keys remain isolated across shared databases. scrapy/scrapy Refactor Cleanup helper extraction Extract shutdown-sentinel logic from _parallel_asyncio() while preserving behavior under regression tests. Table 4: Examples of forecast-conditioned public task specifications. The examples show the synthetic public task surface only; validation assets and internal provenance are hidden. 3.2.2 Task Surface and Hidden Validation Assets The released task exposes the task-generation repository snapshot together with a model-visible issue-style request. The request states the expected behavior, acceptance criteria, and non-goals, while the repository snapshot supplies the code context needed by the evaluated coding agent. The task does not claim to reproduce a known pull request, and it does not expose the validation or construction artifacts used to build the dataset row. Those internal artifacts are kept on the hidden side: forecast-family provenance, retrospective validation labels, oracle plans, task-specific test patches, gold patches, and execution logs. They are used to audit why a task direction was admitted and to verify that the task has an executable target. 3.2.3 Executable Dataset Validation Dataset admission is gate-based rather than quota-based. A task must pass three checks before release. First, a target audit verifies that the request can be justified from the task-generation snapshot. This audit must identify a concrete implementation target, a test target or test location, and a narrow behavior that can be checked without requiring full-project redesign. Second, a leakage review checks the public task surface against hidden construction and validation artifacts. The public task may name current files, APIs, symptoms, and acceptance criteria, but it must not contain test-patch logic, gold-patch details, retrospective validation labels, judge rationales, or retrospective validation metadata. The hidden provenance file records the forecast family and validation evidence for audit only; it is never shown to the evaluated coding agent. Third, executable validation requires two independently constructed artifacts: a task-specific test patch and a gold patch. The test patch defines the task-specific checks and is withheld from evaluated coding agents during evaluation runs. The gold patch is a maintainer-quality solution patch used to prove that the task is solvable and that the test patch is not ill-posed. It is also withheld from evaluated agents and is not treated as the only possible solution. In SWE-Future, both patches are constructed during data generation rather than extracted from a historical PR. Both patches must apply cleanly. The target oracle is then evaluated as a FAIL_TO_PASS target: after applying only the test patch, the task must fail on the task-generation snapshot; after also applying the gold patch, the same oracle must pass. For behavior-preserving refactors, the target may be a focused structural or characterization oracle, but the gold patch must still preserve the relevant public behavior. We also run selected PASS_TO_PASS checks to guard against patches that satisfy the new target while breaking existing behavior. These preservation checks are chosen from repository-native public tests that cover the touched module or nearby behavior, pass on the task-generation snapshot, and avoid unstable external services or unusually heavy full-project setup. This validation design makes the public dataset useful as coding-agent data rather than just as a list of plausible issue descriptions. The forecast family decides the repository-specific direction; TgenT_gen code inspection grounds the task in the repository; executable validation ensures that the task has a concrete behavioral target. 4 Discussion SWE-Future depends on the quality of repository-evolution forecasts, and this forecasting step remains uncertain. The forecaster can pick up noisy anchors from labels, tracker terms, or repository-specific process language rather than clean code modules. The generated test patches make each task executable, but they may be less precise than tests written by long-term maintainers with deeper project context. This limitation is especially visible for refactor tasks, where the intended change is internal and behavior-preserving rather than a new externally observable behavior. Future work should reduce uncertainty at both the forecast and execution levels. On the forecasting side, stronger retrieval over repository evidence and better noise filtering for labels and tracker language could improve the quality of admitted forecast families. On the dataset-construction side, richer project-specific test synthesis could make generated test patches closer to maintainer-written tests. 5 Conclusion SWE-Future turns repository-evolution forecasting into a conditioning signal for coding-agent data synthesis. The study shows that feature implementation, bugfix, and refactor families can be forecast from pre-T0T_0 evidence with measurable future relevance, and that these validated families can drive the construction of 200 coding-agent tasks across 61 repositories. The central contribution is a contamination-aware synthesis route: freeze repository history, forecast likely project directions, ground those directions in TgenT_gen repository snapshots, and admit only tasks with executable validation. Future improvements should focus on reducing forecast noise and making generated test patches closer to maintainer-written tests. Appendix A Forecast Family Templates The cluster-to-family step uses category-specific templates. These templates are not task prompts; they convert a repeated pre-T0T_0 evidence cluster into a coarse future-work direction that later task construction can ground in the TgenT_gen codebase. Category Forecasted direction template Acceptance emphasis Feature implementation/enhancement Extend user-visible anchor behavior, API coverage, provider support, or format compatibility. Require a concrete capability or compatibility path, focused tests, and no dependency/platform or documentation-only change. Bugfix Fix recurring anchor failures, edge cases, regressions, or incorrect behavior reported by users. Require a reproducible failure or edge case, a minimal implementation patch, and a regression test that fails before and passes after the fix. Refactor Restructure anchor internals, typing, deprecated paths, or helper boundaries while preserving externally visible behavior. Require behavior preservation, local maintainability rationale, and tests or checks that guard the preserved public behavior. Template summary used to turn high-scoring pre-T0T_0 clusters into structured future-work directions. Appendix B Forecast Examples The following examples show how repeated pre-T0T_0 signals become forecast families and how those families are later judged against post-T0T_0 PR metadata. These examples illustrate the method boundary: the post-T0T_0 PR is used only for retrospective validation, not to construct the released public task. Repository Pre-T0T_0 cluster Forecast family Retrospective validation PrefectHQ/prefect Feature implementation/enhancement anchor deployment; source signals include issues #19037, #18224, and #19415. Extend deployment capability/API support. Strong: PR #21954 added resizable columns on a deployments table, a user-visible deployment enhancement. plotly/dash Feature implementation/enhancement anchor callback; source signals include pre-T0T_0 PRs #3407, #3397, and issue #3493. Extend callback capability/API support. Strong: PR #3563 added hidden as a configurable parameter for clientside callbacks. celery/kombu Feature implementation/enhancement anchor broker; source signals include pre-T0T_0 PRs #2408, #2367, and issue #2334. Extend broker capability/API support. Strong: PR #2251 added Redis queue expiration support. Kludex/uvicorn Bugfix anchor hangs; source signals include issues #2735 and #2754. Fix recurring hangs correctness and edge-case failures. Strong: PR #2815 replaced fixed sleeps with polling in multiprocess tests and run_server, matching a hang/timing edge-case family. astropy/astropy Bugfix anchor fits; source signals include pre-T0T_0 PRs #18867 and #18818 and issue #18815. Fix recurring fits correctness and edge-case failures. Strong: PR #19438 fixed verification of CompImageHDU headers, matching a FITS correctness family. python/mypy Refactor anchor librt; source signals include pre-T0T_0 PRs #20233, #20010, and #19989. Refactor librt internals while preserving behavior. Strong: PR #20623 made several librt.internal functions positional-only, matching an internal cleanup direction. scrapy/scrapy Refactor anchor cleanup; source signals include pre-T0T_0 PRs #6802, #7043, and issue #6924. Refactor cleanup internals while preserving behavior. Strong: PR #7177 refactored the media-pipeline item handler to async def, matching behavior-preserving internal modernization. Examples of forecast construction and retrospective validation. Examples report validation evidence only; post-T0T_0 PR artifacts are not task sources. Appendix C Task Construction Details Task construction uses a multi-agent workflow so that task writing, evaluation-asset construction, and executable verification remain distinct. The role prompts below are simplified versions intended to show the construction strategy for each role. Context-selection role. The input is a validated forecast family, including its category, anchor, expected behavior, target hints, and supporting pre-T0T_0 signal references, together with the TgenT_gen repository snapshot. The role searches the repository using the anchor, spelling variants, target-module hints, and category-specific words such as option, parser, cache, serializer, exception, typing, or cleanup. It ranks candidate implementation files by anchor hits, proximity to public APIs, import paths, and prior local test coverage, then ranks candidate test files by whether they already exercise the same module, public behavior, or CLI/API surface. It rejects rows whose only plausible targets are documentation, packaging metadata, CI files, release scripts, or broad project-wide rewrites. Context-selection prompt skeleton. 1. Inputs: forecast family, category, anchor, target hints, TgenT_gen file tree, and retrieved code/test excerpts. 2. Select 3–8 implementation candidates and 2–6 test candidates. 3. For each selected path, give a one-sentence repository-evidence rationale. 4. Choose one primary implementation target and one primary test target. 5. Reject instead of selecting when no focused code-and-test path exists. Task-writing role. The input is the forecast family and the selected TgenT_gen repository evidence. This role writes the model-visible issue-style request only. It turns the family into a concrete request by naming the user-visible behavior, the expected change, and the boundaries of the task. It must not see or mention post-T0T_0 PR metadata, semantic-judge rationales, task-specific test-patch content, gold-patch content, or execution logs. It avoids vague requests such as “improve the module” and broad rewrites such as “redesign the parser”; the request must be solvable from the repository snapshot and must include concrete acceptance criteria. Task-writing prompt skeleton. 1. Inputs: forecast family, selected repository evidence, category, and non-goals. 2. Write: title, repository, category, issue, requested change, acceptance criteria, and forecast-conditioned context. 3. Make the request natural as a project issue and specific enough for a coding agent. 4. Do not reveal test-patch logic, gold-patch logic, validation commands, or post-T0T_0 PR evidence. 5. Reject if the request cannot be made focused, testable, and local. Oracle-design role. The input is the public task, selected implementation and test targets, and repository test conventions. This role turns the public request into a concrete oracle plan. It checks whether the behavior is already implemented in the snapshot, identifies the smallest observable behavior that should fail before the fix, chooses a test entry point, and specifies the expected failure and pass-after condition. It rejects candidates when the intended behavior requires external services, nondeterministic timing, large integration environments, whole-repository migrations, or subjective maintainability judgments that cannot be captured by a focused oracle. Oracle-design prompt skeleton. 1. Inputs: public task, selected files, relevant code excerpts, and test style examples. 2. Define the target behavior in one sentence. 3. Specify the fail-before observation and pass-after observation. 4. Specify the test-patch location, test name, and setup assumptions. 5. Specify the smallest gold-patch scope, or reject with a reason. Patch-construction role. The input is the oracle plan and a clean TgenT_gen checkout. The role first writes the task-specific test patch, because the test patch defines the target behavior. It then writes the gold patch, scoped to the smallest implementation path that satisfies the oracle. The test patch may add a focused regression test, structural check, or public-behavior check depending on the task category. The gold patch must not include unrelated cleanup, formatting churn, dependency updates, or historical PR content. If the task requires a larger design change than the public request implies, the row is returned to task writing or replaced. Patch-construction prompt skeleton. 1. Apply no historical PR patch and start from the clean TgenT_gen tree. 2. Write the task-specific test patch first; keep it focused on the oracle. 3. Run or reason about the expected fail-before behavior. 4. Write the gold patch with the smallest behavior-changing implementation. 5. Return both patches plus touched paths and any repair notes. Verification role. The input is the public task, task-specific test patch, gold patch, and a clean checkout of the task-generation snapshot. Verification first checks that both patches apply. It then applies only the test patch and expects the target oracle to fail. Next, it applies the gold patch on top of the test patch and expects the same oracle to pass. This is the FAIL_TO_PASS gate. Selected PASS_TO_PASS checks are chosen from repository-native public tests covering touched or nearby behavior. A row that fails target audit, leakage review, patch application, FAIL_TO_PASS, or selected preservation checks is returned for repair or replaced rather than admitted to satisfy the 200-task construction target. Appendix D Synthesized Task Example This appendix shows one complete synthesized task row. The example is a bugfix task from pallets/click. It is conditioned on a validated click bugfix family, but the public task, task-specific test patch, and gold patch are constructed from the task-generation snapshot rather than from a post-T0T_0 PR patch. Field Value Task id swe-future-v301-0001-click-choice-casefold-completion Repository pallets/click Category Bugfix Forecast family Fix recurring click correctness and edge-case failures. Public behavior Case-insensitive Choice conversion and shell completion should use the same Unicode normalization. Task-specific test Add a focused shell-completion regression test for straße/STRAS. Gold patch Replace lower() with casefold() in Choice.shell_complete. Model-visible public task. The following task text is visible to the evaluated coding agent. ⬇ # Fix case-insensitive Choice completion for Unicode casefolding Repository: ‘pallets/click‘ Category: ‘bugfix‘ ## Issue ‘click.Choice(..., case_sensitive=False)‘ uses Unicode ‘casefold()‘ when converting command-line values, but its shell-completion path still compares choices with ‘lower()‘. This makes completion disagree with conversion for valid non-ASCII choices such as ‘stra 00dfe‘, where ‘STRASSE‘ converts successfully but the same prefix is not completed. ## Requested Change Update ‘click.Choice.shell_complete‘ so case-insensitive completion uses the same Unicode casefold normalization as ‘Choice.convert‘. Keep the change local to choice completion and preserve existing case-sensitive behavior. ## Acceptance Criteria - ‘click.Choice(["stra 00dfe"], case_sensitive=False).convert("STRASSE", None, None)‘ continues to return ‘"stra 00dfe"‘. - The same choice returns ‘"stra 00dfe"‘ from ‘shell_complete‘ for an incomplete value such as ‘"STRAS"‘. - Case-sensitive completions remain unchanged. - Do not add broad rewrites, new dependencies, network calls, or unrelated cleanup. ## Forecast-Conditioned Context This task instantiates the ‘click‘ bugfix forecast family by targeting a small user-visible edge case in command-line completion behavior. The task is generated from the repository snapshot and does not use any post-$T_0$ PR patch. Task-specific test patch. The test patch defines the target oracle. With only this patch applied, the task-generation snapshot should fail the new casefold-completion test. ⬇ diff --git a/swe_future_tests/test_click_choice_casefold_completion.py b/swe_future_tests/test_click_choice_casefold_completion.py new file mode 100644 --- /dev/null +++ b/swe_future_tests/test_click_choice_casefold_completion.py @@ -0,0 +1,22 @@ +from pathlib import Path +import sys + +sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src")) + +import click + +def test_case_insensitive_choice_completion_uses_casefold(): + choice = click.Choice(["stra 00dfe"], case_sensitive=False) + assert choice.convert("STRASSE", None, None) == "stra 00dfe" + completions = choice.shell_complete(None, None, "STRAS") + assert [item.value for item in completions] == ["stra 00dfe"] + +def test_case_sensitive_choice_completion_is_unchanged(): + choice = click.Choice(["Alpha"], case_sensitive=True) + assert [item.value for item in choice.shell_complete(None, None, "Al")] == ["Alpha"] + assert choice.shell_complete(None, None, "al") == [] Gold patch. The gold patch demonstrates that the task is solvable with a minimal local implementation change. ⬇ diff --git a/src/click/types.py b/src/click/types.py --- a/src/click/types.py +++ b/src/click/types.py @@ -382,8 +382,8 @@ class Choice(ParamType, t.Generic[ParamTypeValue]): if self.case_sensitive: matched = (c for c in str_choices if c.startswith(incomplete)) else: - incomplete = incomplete.lower() - matched = (c for c in str_choices if c.lower().startswith(incomplete)) + incomplete = incomplete.casefold() + matched = (c for c in str_choices if c.casefold().startswith(incomplete)) return [CompletionItem(c) for c in matched] Executable validation. The task passes the executable release gates used by the constructor. Gate Expected result Return code Passed Public syntax check before patches Pass 0 Yes Apply task-specific test patch Pass 0 Yes Run target oracle before gold patch Fail 1 Yes Apply gold patch Pass 0 Yes Run target oracle after gold patch Pass 0 Yes Public syntax check after patches Pass 0 Yes Selected public preservation gate Pass n/a Yes The selected public preservation gate for this row is a static public check. The row records zero post-T0T_0 PR patch use. The evaluated model receives the public task and repository snapshot; the task-specific test patch, gold patch, and validation logs remain evaluation-side assets. References [1] I. Badertdinov, A. Golubev, M. Nekrashevich, A. Shevtsov, S. Karasik, A. Andriushchenko, M. Trofimova, D. Litvintseva, and B. Yangel (2025) SWE-rebench: an automated pipeline for task collection and decontaminated evaluation of software engineering agents. External Links: 2505.20411, Link, Document Cited by: §1, §2. [2] D. Fu, S. Wu, Y. Wu, Z. Peng, Y. Huang, J. Sun, J. Zeng, M. Jiang, L. Zhang, Y. Li, J. Hu, L. Liu, J. Hou, and P. Liu (2026) daVinci-Env: open SWE environment synthesis at scale. External Links: 2603.13023, Link, Document Cited by: §2. [3] T. L. Graves, A. F. Karr, J. S. Marron, and H. Siy (2000) Predicting fault incidence using software change history. IEEE Transactions on Software Engineering 26 (7), p. 653–661. External Links: Document Cited by: §2. [4] L. Guo, Y. Wang, C. Li, W. Tao, P. Yang, J. Chen, H. Song, D. Tang, and Z. Zheng (2026) SWE-Factory: your automated factory for issue resolution training data and evaluation benchmarks. Note: To appear at FSE 2026 External Links: 2506.10954, Link, Document Cited by: §1, §2. [5] A. E. Hassan (2009) Predicting faults using the complexity of code changes. In Proceedings of the 31st International Conference on Software Engineering, p. 78–88. External Links: Document Cited by: §2. [6] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan (2024) SWE-bench: can language models resolve real-world github issues?. In International Conference on Learning Representations, External Links: Link, Document Cited by: §1, §2. [7] Y. Kamei, E. Shihab, B. Adams, A. E. Hassan, A. Mockus, A. Sinha, and N. Ubayashi (2013) A large-scale empirical study of just-in-time quality assurance. IEEE Transactions on Software Engineering 39 (6), p. 757–773. External Links: Document Cited by: §2. [8] W. Li, X. Zhang, Z. Guo, S. Mao, W. Luo, G. Peng, Y. Huang, H. Wang, and S. Li (2025) FEA-Bench: a benchmark for evaluating repository-level code generation for feature implementation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, External Links: Link, Document Cited by: §1, §2. [9] A. Matton, T. Sherborne, D. Aumiller, E. Tommasone, M. Alizadeh, J. He, R. Ma, M. Voisin, E. Gilsenan-McMahon, and M. Gallé (2024) On leakage of code generation evaluation datasets. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, p. 13215–13223. External Links: Link, Document Cited by: §1, §2. [10] N. Nagappan and T. Ball (2005) Use of relative code churn measures to predict system defect density. In Proceedings of the 27th International Conference on Software Engineering, p. 284–292. External Links: Document Cited by: §2. [11] M. Riddell, A. Ni, and A. Cohan (2024) Quantifying contamination in evaluating code generation capabilities of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, p. 14116–14137. External Links: Link, Document Cited by: §1, §2. [12] X. Sun, H. Sun, T. Yu, S. Ma, Q. Zhang, L. Rao, and C. Tian (2026) A time-consistent benchmark for repository-level software engineering evaluation. External Links: 2603.26137, Link, Document Cited by: §2. [13] L. Wang, L. Ramalho, A. Celestino, P. A. Pham, Y. Liu, U. K. Sinha, A. Portillo, O. Osunwa, and G. Maduekwe (2025) SWE-Bench++: a framework for the scalable generation of software engineering benchmarks from open-source repositories. External Links: 2512.17419, Link, Document Cited by: §1, §2. [14] Y. Zeng, S. Li, D. Dong, R. Xu, Z. Chen, L. Zheng, Y. Li, Z. Zhou, H. Zhao, L. Tian, H. Xiao, T. Zhu, L. Hao, and J. Wu (2026) SWE-Hub: a unified production system for scalable, executable software engineering tasks. External Links: 2603.00575, Link, Document Cited by: §2. [15] Y. Zhu, A. Gandhi, and G. Neubig (2026) Training versatile coding agents in synthetic environments. Note: SWE-Playground project repository External Links: Link Cited by: §2. [16] T. Zimmermann, P. Weissgerber, S. Diehl, and A. Zeller (2004) Mining version histories to guide software changes. In Proceedings of the 26th International Conference on Software Engineering, p. 563–572. External Links: Document Cited by: §2.