Paper deep dive
One Run Is Not an Idea: The Implementation Lottery in Automated Research
Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Chenyan Xiong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/4/2026, 10:47:06 AM
Summary
The paper introduces the 'implementation lottery,' a phenomenon where automated research systems make idea-level decisions based on single-run scores of specific implementations, leading to unreliable conclusions. The authors propose the 'Idea Reliability Audit' to measure idea reliability by freezing candidate ideas, sampling multiple fresh implementations, and assessing stability via Intra-class Correlation Coefficient (ICC) and Leave-One-Implementation-Out (LOO) winner reversal. Experiments across 312 assignments on tabular and materials tasks show that implementation variance significantly exceeds rerun noise, causing winner reversals in 25.6% to 43.6% of decisions, highlighting the need for multiple implementations before guiding research branching or transfer.
Entities (10)
Relation Signals (7)
Idea Reliability Audit → measures → Idea Reliability
confidence 96% · The Idea Reliability Audit measures idea reliability by validating and freezing candidate cards...
Implementation Lottery → causes → Unreliable Idea-Level Decisions
confidence 95% · Crediting that realization-level score as evidence about the parent mechanism creates the implementation lottery, in which an idea-level conclusion depends on which plausible implementation was sampled.
Idea Reliability Audit → uses → Idea ICC
confidence 94% · It reports idea ICC and leave-one-implementation-out (LOO) winner reversal.
Implementation Variance → exceeds → Rerun Noise
confidence 93% · Implementation variance was more than five and ten times same-artifact rerun variance, respectively...
Implementation Lottery → resultsin → LOO Winner Reversal
confidence 92% · LOO winner reversal occurs in 25.6% and 43.6% of decisions.
Tabular Tasks → auditedusing → Idea Reliability Audit
confidence 90% · Across 312 assignments on 13 tabular tasks... we audit 13 public OpenML classification tasks...
DeepSeek V4 Pro → usedin → Tabular Audit
confidence 89% · In the Tabular audit, DeepSeek v4 Pro generates the cards... Both implementation setups also use DeepSeek v4 Pro...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.
Tags
Links
- Source: https://arxiv.org/abs/2607.26587v1
- Canonical: https://arxiv.org/abs/2607.26587v1
Trouble viewing inline? Open PDF directly →
Full Text
49,199 characters extracted from source content.
Expand or collapse full text
One Run Is Not an Idea: The Implementation Lottery in Automated Research Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Chenyan Xiong Abstract Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the implementation lottery, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The Idea Reliability Audit measures idea reliability by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6% and 43.6% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-N artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations. 1 Introduction We use automated research to describe a closed loop that proposes, implements, tests, and selects research directions. The AI Scientist, Dolphin, and Agent Laboratory automate growing parts of this loop (Lu et al. 2024; Yuan et al. 2025; Schmidgall et al. 2025). Tree-search variants explicitly explore candidate artifacts (Yamada et al. 2025), while executable benchmarks ask agents to improve concrete programs (Nathani et al. 2025; Garikaparthi et al. 2026). Exploring several realizations and retaining the best is appropriate when the deliverable is that artifact. The same loop also uses results to decide what research to pursue. Scientific learning then asks whether a selected result is evidence for the named idea, so it can guide branching, transfer, and cumulative knowledge. A selected maximum identifies what one search budget found. It need not rank ideas as they perform across plausible implementations. This distinction matters when systems use their own experiments to allocate future research effort (Wang et al. 2023a). Suppose an agent proposes early stopping for a booster. The mechanism fixes a training-only validation set, a stopping rule, and restoration of the best iteration. Faithful implementations may still choose the validation fraction, patience, monitored metric, and library interface. The score from one run therefore measures both the fixed idea and one realization of its permitted choices. An illustrative case from our audits makes the consequence concrete. Three valid, blind-faithful implementations of one frozen card span +2.31%+2.31\% to −1.85%-1.85\%, while all same-artifact reruns are exact. Against the same competitor, every one-draw winner reverses under the other-two mean (Figure 3). Trying several realizations and keeping the maximum can improve the returned artifact, but it does not resolve attribution. The maximum targets an upper tail and can favor ideas with greater implementation variance or search effort. That is useful for delivery but consequential when the score is credited to the idea and steers the next hypothesis. Source of within-task ITT utility variationWinner changes across attemptsIdeaImplementationRerunTabular Bounded61336ICC .612 [.477,.745]Tabular Agentic51454ICC .511 [.248,.747]Materials (exploratory; Section 5)Materials Bounded2971ICC .290 [.000,.824]Materials Agentic1882ICC .177 [.040,.777]025507510025.6%43.6%33.3%55.6% Figure 1: The implementation lottery in the primary Tabular audit (top) under capped-call Bounded and SDK-session Agentic setups, and in the exploratory Materials diagnostic (shaded). Stacked bars show the share of within-task intention-to-treat (ITT) variance due to frozen candidate, implementation, and same-artifact rerun. The idea ICC equals the first share. Points show LOO reversal with 95% task-level intervals. Tabular uses all frozen candidates. Card-filtered sensitivities appear in Table 1. Materials is never pooled with Tabular, and its rerun variance is zero because the evaluator is deterministic. Written ideas can be rated for novelty and feasibility (Si et al. 2025), and execution can change those ratings (Si et al. 2026a). One idea can admit several plausible implementations. We call dependence on one sampled realization the implementation lottery. Idea reliability is the stability of idea-level evidence across implementations; the Idea Reliability Audit estimates it. Prior repeated-run evaluations treat a task or query as the repeated object. We instead hold a mechanism-level idea fixed, repeat its translation into artifacts, audit semantic adherence, and separately rerun saved artifacts. Automated research often makes idea-level decisions from one realization. This unit mismatch is structural; only its magnitude is empirical. Our audits find 25.6–43.6% selection reversal (Figure 1). We make three contributions. • Phenomenon and scale. Across 312 assignments on 13 tabular tasks, implementation variance exceeds same-artifact rerun variance by more than fivefold for Bounded and tenfold for Agentic. LOO winner reversal occurs in 25.6% and 43.6% of decisions. • Measurement target. We formalize idea reliability as the stability of evidence across independent idea-to-artifact runs. We distinguish average-realization operational idea quality Q from best-of-N artifact utility BNB_N and connect ICC to winner stability. • Instrument. We introduce an audit that checks portfolio validity, freezes cards, samples fresh-session implementations, labels fidelity, and reruns identical artifacts. It reports idea ICC, LOO reversal, and task-level uncertainty. To our knowledge, this is the first study to hold a research idea semantically fixed while repeatedly sampling its translation into artifacts. This design separates implementation variance from same-artifact rerun noise and uses outcome-blind fidelity review. It is also the first to report idea-level selection reversal caused by resampling independent implementations of a fixed idea portfolio. The results compare implementation variation with rerun noise, examine both processes and a fully valid case, repeat the analysis after outcome-blind card filtering, and provide an exploratory Materials check. In a closed loop, a realization-specific winner can redirect later experiments. Search determines what to deploy; idea reliability determines what the research loop should remember. 2 Idea Reliability as Measurement What One Score Contains A task context fixes data, code, split, metric, and compute budget. An idea states a mechanism while leaving reasonable implementation choices open. An agent turns it into code, which an evaluator scores. Repetition separates variation across ideas, across implementations of one idea, and across reruns of one saved implementation. Let s denote a task, i an idea, e a coding-agent setup, j an independent implementation attempt, and k a training or evaluation rerun. For a fixed setup e, we write the observed utility as Ysiejk=μe+Ase+Isie+Msiej+Rsiejk,Y_siejk= _e+A_se+I_sie+M_siej+R_siejk, (1) where A is the task contribution, I is the idea contribution within a task, M is the implementation contribution within an idea, and R is variation observed when the same implementation is rerun. The random terms are zero mean and exchangeable within each level. For a fixed task and implementation agent, operational idea quality is the average ITT utility over independent implementation attempts and evaluator reruns. Q(s,i,e)=j,k[Ysiejk].Q(s,i,e)=E_j,k[Y_siejk]. (2) For decision quantities, let ZsiejZ_siej denote the implementation-level score exposed to the selector. In our experiments this is the frozen primary ITT utility, while separate same-artifact reruns estimate σR2 _R^2. Best-of-N artifact search targets an upper-tail quantity such as BN(s,i,e)=[maxj≤NZsiej]B_N(s,i,e)=E[ _j≤ NZ_siej]. Both are legitimate estimands but answer different questions. BNB_N describes what a budgeted search can return, whereas Q describes the idea across realizations. Crediting a realization-level Z as evidence about Q creates the implementation lottery because the idea-level conclusion then depends on which plausible implementation was sampled. A change in the selected idea across implementation draws is its decision-level manifestation. This definition follows reliability and generalizability theory (Shrout and Fleiss 1979; Brennan 2001). The measured object is a research idea and the repeated measurements are its implementations. Two Measures of Reliability For each coding-agent setup, we estimate the variance components with a balanced nested analysis of variance. Negative sampling estimates are set to zero. Task variance is excluded because an automated research system compares ideas within the same task. The idea ICC is ICCidea=σI2σI2+σM2+σR2.ICC_idea= _I^2 _I^2+ _M^2+ _R^2. (3) The ICC is the share of within-task ITT utility variation that separates ideas. A value near one means that independent implementation attempts preserve the idea signal. Lower values mean that one run is less representative. The second measure asks whether implementation choice changes the winner. We form three one-draw worlds, each containing one implementation per idea, and ask whether each winner survives the other two implementations. The preregistered alignment h∈1,2,3h∈\1,2,3\ groups same-index draws across ideas into worlds; it is a label, not shared provider randomness. For each h, one draw selects i^s,e,h=argmaxiZsieh, i_s,e,h= _iZ_sieh, (4) and we compare it with argmaxiZ¯sie,−h _i Z_sie,-h. This other-two mean is a finite-budget reference, not ground truth, so LOO measures decision stability rather than latent-ranking error. A disagreement is a leave-one-implementation-out (LOO) winner reversal. Exact observed ties select the lexicographically largest frozen idea label; reference means tied within 10−1210^-12 all count as best. Each task contributes three dependent decisions, so uncertainty is clustered by task. These measures apply when a result prunes an idea branch, motivates transfer, enters research memory, or supports a mechanism-level claim. An audit draw may itself be a complete tree-search run; Section 8 discusses promotion policies. 3 The Idea Reliability Audit Figure 2 summarizes the prospective workflow. Each component addresses a distinct alternative explanation for an unstable idea-level decision. Mechanism-Level Idea Cards Mechanism-level cards rule out the alternative that variation reflects an underspecified slogan or a fully specified recipe. A card fixes the scientific intervention while leaving ordinary choices about models, features, preprocessing, and parameters open. It records the mechanism, required properties, open choices, expected result, and counterevidence. Exact constructors, hyperparameters, random seeds, and line-by-line steps are excluded because they define a recipe. Idea reliability is defined at this card granularity. Under our no-search policy, each draw is one translation; with tuning, the search procedure and budget become part of the implementation process being audited. For example, one frozen tabular card fixes early stopping. It reserves training data for validation, monitors it during boosting, stops under a patience rule, and returns the best iteration under a 300-iteration ceiling. The executor still chooses the validation fraction, patience, metric, split seed, and built-in versus manual interface. Leaf regularization or feature subsampling would change the mechanism and violate the card. For every task, cards are generated before implementation and remain blind to hidden outcomes. Frozen baseline development lineage remains visible. The task context fixes the data, code snapshot, split, metric, environment, and compute policy. Portfolio validity must also be checked outcome-blind before an idea-level interpretation is made. Each card should differ from the baseline, encode one coherent and representation-valid mechanism, leave genuine implementation freedom, and remain semantically distinct from the other candidates. Lexical, coarse-family, and anti-recipe screens can assist this gate but do not replace semantic review. Independent Implementations Fresh sessions and outcome-blind fidelity review target the alternative that observed differences come from shared context, leakage, drift, or failure. Each card receives three implementation draws in separate sessions. We save every patch before revealing evaluation outcomes and audit provider response or session identities rather than assuming physical independence. The two coding-agent setups receive the same card portfolio and follow the same task and compute rules. Following the LLM-as-judge paradigm (Zheng et al. 2023), a blinded model-based reviewer checks whether each patch follows the card. A second reviewer labels a frozen 30% subset, and constructed controls probe whether the review distinguishes faithful, drifting, and invalid patches. This step measures whether different implementations still represent the same idea. Separate Implementation Choice from Rerun Noise Same-code reruns target the alternative that observed variation is merely training or replay noise. Each executable implementation is replayed unchanged under frozen rerun conditions; differences among independently produced implementations estimate the additional effect of implementation choice. Failed and drifting runs remain in the primary analysis because failure to produce a working, faithful implementation is part of an idea’s operational cost. Filtering after execution would condition reliability on successful translation. We use frozen penalties for these runs and repeat the analysis across the preregistered penalty settings. This separation is essential. Repeating training seeds for one program measures whether that program is stable, but it cannot show whether another implementation attempt would support the same decision. We also save held-out prediction vectors. This lets the audit distinguish code that looks different but produces the same model behavior from implementations that lead to genuinely different predictions. Prospective outcome-blind portfolio freezing and validation Figure 2: Recommended prospective audit workflow. Portfolios should be frozen and undergo outcome-blind semantic validation before implementation. Our study applies the semantic gate after execution as a sensitivity (Section 5). Three fresh sessions implement each admitted card, fidelity is labeled outcome-blind, executable artifacts receive byte-identical reruns, and all attempts remain in ITT. The audit reports within-task idea ICC and LOO winner reversal. 4 Study Design Tasks Tabular machine learning. We audit 13 public OpenML classification tasks (Vanschoren et al. 2014). Eligibility was frozen before card generation. Tasks could have at most 8,000 rows, 60 features, and ten classes, with no missing values, imbalance at most 10, at least .03 development headroom, and no real agent patch previously exposed to a hidden outcome. Nonrandom Phase 1 contains task IDs 18, 2074, 3903, 3917, 9957, and 10101. A separately frozen extension exhausts the remaining eligible tasks, which are 11, 22, 43, 3902, 3913, 9952, and 146822. They span 522–6,430 examples, 4–57 features, and binary through ten-class prediction. Official OpenML fold 0 is held out. Folds 1–9 are visible only through development evaluation. The common one-thread baseline for every task is scikit-learn’s HistGradientBoostingClassifier (learning rate .08, 150 iterations, 31 leaves, minimum leaf size 20, and ℓ2 _2 penalty 1.0). Runs use Python 3.12.12, NumPy 2.2.6, and scikit-learn 1.9.0. Search, ensembles, external data, and more than 300 boosting iterations are forbidden. Four accepted cards per task, three implementations per card, and two coding setups give 13×4×3×2=31213× 4× 3× 2=312 assignments. Materials prediction. We separately apply the frozen-card design to three Matbench-derived regression workflows (Dunn et al. 2020). They cover steel yield strength and experimental band gap from composition, and phonon frequency from crystal structure, using descriptors from standard materials-informatics tools (Ward et al. 2018). A fourth task, dielectric response, was excluded by an outcome-blind decision when 20 generation attempts yielded only two distinct mechanisms. The retained tasks contribute 72 assignments. Because they were exposed in an earlier proof of concept and the preregistered human portfolio gate was not validly completed, Materials is exploratory. After execution, an outcome-blind model in a fresh context reviewed only the baselines and cards. Section 5 reports its findings. Models and Implementation Processes In the Tabular audit, DeepSeek v4 Pro generates the cards from frozen code, task context, and visible development lineage, without hidden outcomes. Lexical overlap and a mechanism-family classifier reject duplicates, while a recipe guard rejects cards that pin constructors or hyperparameters. Both implementation setups also use DeepSeek v4 Pro and receive identical cards and task rules. Bounded permits at most four model calls, limited code inspections, two repairs, and a fixed token allowance. Agentic uses one fresh Read/Edit SDK session, the SDK max_turns option set to 20, and no external repair pass. Full call, token, and inspection limits are documented in the accompanying supplementary artifact. The max_turns setting is not a verified hard cap. Reliability is estimated separately within each setup. The comparison concerns the full implementation processes rather than base models. Blinding, Fidelity, and Reruns Each Tabular assignment uses a fresh context and a frozen, outcome-blind nonce. Patches, transcripts, rerun seeds, and fidelity judgments are sealed before hidden evaluation. The provider did not support sampling seeds, so independence rests on fresh sessions, nonces, and response/session identities rather than reproducible model sampling. Those identifiers verify physical independence for 311 of 312 assignments. One Bounded draw lacks a provider response identifier and is reported as unverifiable rather than certified. The primary fidelity reviewer is DeepSeek v4 Flash. An independently prompted DeepSeek v4 Pro labels a preregistered 30% subset, rounded up within each run (98/312 overall). Reviewers see only the card and patch diff, not task identity, scores, other labels, or deterministic checks. The primary label remains the ITT label rather than being adjudicated by the second. Constructed controls pair a card with another card’s patch (drift) or an empty patch (invalid). Every executable patch is evaluated with byte-identical code and three frozen seeds, giving three scores without another executor call. All assignments remain in ITT. Utility and Inference Tabular utility is relative held-out accuracy gain, (patched−baseline)/|baseline|(patched-baseline)/|baseline|. Penalties are −.10-.10 for evaluation failure or fidelity invalidity, −.02-.02 for drift or missing fidelity, and zero for faithful or partial patches. We repeat the analysis over the frozen failure (−.05,−.10,−.20-.05,-.10,-.20) and drift (0,−.02,−.050,-.02,-.05) grids, and with deterministic fidelity labels. We estimate Equation 1 by balanced method of moments. The mean within-patch sample variance from the three seeds is the rerun component and is subtracted from the implementation mean square. Negative sampling estimates are retained in the artifacts and bounded at zero for ICC. Confidence intervals resample entire tasks and their rerun records. Phase 1 enumerates all 66=46,6566^6=46,656 task-cluster bootstrap draws, while the extension and pooled analysis use 200,000 frozen-seed draws. Repeatedly sampled tasks receive new cluster identifiers and carry resampled task-level rerun records. We also use leave-one-task-out (LOTO) estimates and exact two-sided, task-level paired sign-flip tests for setup contrasts. Materials utility is task-normalized improvement in mean absolute error and is never pooled with Tabular. 5 Results Implementation Variance Exceeds Rerun Noise Figure 1 gives the pooled frozen-card result. Among the non-idea sources of within-task ITT variation, implementation choice is larger than same-artifact rerun noise in both setups. Implementation accounts for 33% and 45%, versus 6% and 4% for reruns. The corresponding implementation standard deviations are .0202 and .0295 on the relative-held-out-gain scale, versus .00871 and .00917 for reruns. Idea ICC is .612 [95% task-cluster interval .477, .745] for Bounded and .511 [.248, .747] for Agentic. Estimated implementation variance is more than five times rerun variance for Bounded and more than ten times for Agentic. LOO winner reversal occurs in 10/39 decisions (25.6% [7.7%, 46.2%]) and 17/39 (43.6% [20.5%, 66.7%]). Using each idea’s other-two-attempt mean as the internal reference, descriptive mean LOO selection regret is .00342 and .00693 in relative-gain units. Low immediate regret does not make reversal vacuous. Automated loops apply selection pressure precisely when candidates are close, and a flipped label can redirect the next branch. Same-artifact reruns isolate a smaller component and cannot measure the effect of retranslating a card. Because ITT treats failed or drifting translations as operational outcomes, their dispersion enters the implementation component. Even so, implementation variation exceeds rerun variation across the preregistered penalty grid. LOTO ICC ranges from .566 to .627 for Bounded and .459 to .568 for Agentic. Across the grid, ICC ranges from .464 to .707 and .465 to .511, while reversal ranges from 25.6–30.8% and 35.9–43.6%. Deterministic fidelity labels do not reverse the finding. (a) One card, three faithful realizations Frozen I02 Pairwise products + one HGB Draw 0 full quadratic; pre- and post-scale Draw 1 interactions + duplicated originals; pre-scale Draw 2 interactions + originals; post-scale; η=.05,ℓ2=.1η=.05,\ _2=.1 ✓ valid✓ blind-faithful✓ exact 3-seed replay(b) One score changes the selected ideaI02 realizationI00 competitor3/3 reversedDraw 0Draw 1Draw 2-2-1012relative held-out utility (%)+0.46+2.31+0.46−1.85-1.85I02 → I00I02† → I00I00 → I02 Figure 3: Implementation lottery in one fresh-extension case (OpenML task 146822, Bounded), chosen after unblinding to illustrate the mechanism. Pooled inference is in Figure 1. The illustrated I02–I00 pair is semantically distinct. Three valid, blind-faithful I02 realizations yield +2.31%+2.31\%, +0.46%+0.46\%, and −1.85%-1.85\% relative utility, while each saved artifact replays exactly across three seeds. Against I00 (+0.46%+0.46\%), every one-draw selection changes under the other-two mean. Omitted cards score zero. The dagger marks the frozen exact-tie rule. A concrete implementation lottery. Figure 3 makes the estimand tangible without adding a new inferential test. The three allowed realizations of the same interaction-feature card range from +2.31%+2.31\% to −1.85%-1.85\% relative held-out utility, and their prediction vectors disagree on as many as 4.76% of examples. Against a competing card that remains at +0.46%+0.46\%, all three one-shot choices are reversed by the other two implementations. Because every patch is executable and blind-faithful and its same-patch reruns agree exactly, ordinary implementation choices rather than failure, fidelity penalties, or rerun noise change the selected idea in this case. A high-scoring artifact can remain valid even when its score provides limited evidence about the parent idea. The Lottery Persists Across Processes The implementation lottery appears in both setups. Phase 1 ICC is .661 for Bounded and .260 for Agentic, with 22.2% and 66.7% reversal. On seven fresh tasks the ordering flips. ICC is .588 versus .711, while reversal is 28.6% versus 23.8% (paired p=1.00p=1.00). Pooled Agentic-minus-Bounded intervals include zero for ICC (−.101-.101 [−.337,.173-.337,.173]) and reversal (+17.9% [0.0%, 38.5%], p=.1875p=.1875). These audits therefore do not identify a consistent reliability ordering between the two implementation processes. Reliability must be estimated for the process being evaluated. Execution Validity and Audit Trail All 312 assignments remain in ITT. Each setup has 147/156 valid assignments. The remaining nine are evaluation errors, with zero code-invalid patches. Fidelity is faithful or partial for 150/156 Bounded and 152/156 Agentic assignments. On the double-labeled subset, exact agreement is 47/49 and 43/49. The primary reviewer detects 50/52 and 51/52 constructed drifts and all 52/52 constructed invalids in each setup. The single missing response identifier noted above affects the completeness audit, not assignment inclusion. Two outcome-blind repairs remain in the audit trail. A family classifier was corrected to use the mechanism field. Task 18 then has three rather than four families, but frozen cards were not regenerated. The registered ten generation attempts became a floor with a ceiling of 40 when some tasks produced too few distinct candidates. Tasks 3913 and 9952 required 14 and 13 attempts. Both repairs preceded hidden outcomes and left accepted cards unchanged. Portfolio Validity Is a Separate Gate To test whether degenerate cards account for the reversals, we repeated the analysis after applying two outcome-blind card filters. A fresh-context OpenAI Codex instance from outside the DeepSeek pipeline reviewed only the 52 frozen Tabular cards and visible baselines after execution. A conservative rule flags any card that is not baseline-distinct, single-mechanism, representation- and policy-applicable, and idea-like, admitting 38/52 cards (73.1%). Some flags reflect a review instruction stricter than the frozen protocol on fitted preprocessing and idea-versus-recipe boundaries. A construct-focused adjudication using the frozen policy therefore admits 43/52 (82.7%). Substantive flags include monotone transforms that are effectively equivalent for histogram trees, one multi-mechanism umbrella, and one duplicate pair. Table 1: LOO winner reversal (%) before and after outcome-blind card filtering. Every set retains all 13 tasks. Candidate set Cards Bounded Agentic All frozen 52 25.6 43.6 Conservative rule 38 43.6 35.9 Construct rule 43 33.3 38.5 Requiring all four cards in a task to pass is stricter than card-level filtering. Under the construct rule, five of seven excluded tasks contain only one flagged card. We therefore make card-level filtering the primary sensitivity. Keeping admitted cards and every task with at least two candidates retains all 13 tasks. Table 1 reports reversal of 33.3–43.6%, with task-cluster intervals spanning 12.8–66.7%. The unequal two-to-four cards per task preclude the registered balanced method-of-moments ICC. As a maximally conservative check requiring every card to pass, the conservative rule leaves four tasks (ICC .550/.250 and 50.0% reversal in each setup). The construct rule leaves six (ICC .642/.443 and 33.3% reversal in each), albeit with wide intervals. We use the post-execution construct audit only for sensitivity analysis and leave both the cards and primary ITT sample unchanged. The all-13 estimates measure the reliability of the deployed candidate-card workflow. Clean mechanism-level promotion additionally requires the portfolio gate. The Lottery Persists in Materials Materials stress-tests cross-domain transfer and exposes a second failure mode. Its outcome-blind review passes band gap but flags baseline-equivalent, bundled, duplicate, or representation-dependent candidates in steels and phonons. Performance-only evaluation would miss this portfolio failure; prospective semantic admission is designed to catch it. We therefore treat all 72 assignments as exploratory and do not pool them with Tabular. Winner reversal is 33.3% [0.0%, 66.7%] for Bounded and 55.6% [0.0%, 100.0%] for Agentic. Frozen-card ICC is .290 [.000, .824] and .177 [.040, .777], with implementation shares of 71% and 82%. Same-artifact rerun variance is zero by construction because the evaluator is deterministic. Bounded validity is 29/36 (80.6%), below the registered .85 gate, while Agentic validity is 34/36 (94.4%). Intervals are wide because only three workflows are available. Even with rerun variation fixed at zero, however, implementation variation dominates and winner reversals remain. 6 Related Work Automated research. Automated research systems increasingly connect ideation, coding, execution, and selection. The AI Scientist, Dolphin, Agent Laboratory, Co-Scientist, ERA, execution-grounded research, specialist-agent recipe search, and closed-loop molecular workflows automate different parts of this loop (Lu et al. 2024; Yuan et al. 2025; Schmidgall et al. 2025; Gottweis et al. 2026; Aygün et al. 2026; Si et al. 2026b; Ning et al. 2026a, b). Related work studies end-to-end automation of AI research directly (Lu et al. 2026). MLGym, MLRC-Bench, MLR-Bench, ResearchGym, and ResearchArena provide executable environments for research agents (Nathani et al. 2025; Zhang et al. 2025; Chen et al. 2025; Garikaparthi et al. 2026; Zhang et al. 2026). They test whether agents can complete research tasks or improve experimental artifacts. We study when this artifact-level feedback can be attributed to the idea that generated it. Ideas and execution. Research idea evaluation has scored written proposals for novelty and feasibility (Si et al. 2025). A closely related execution study assigns one completed project to each idea and shows that execution can change an idea’s evaluation (Si et al. 2026a). We repeat fresh-session implementation attempts of the same idea to study the next measurement step. ResearchCodeBench, AutoExperiment, and AutoReproduce test whether agents can translate scientific intent into working code (Hua et al. 2025; Kim et al. 2026; Zhao et al. 2026). Complementary analysis identifies implementation capability as a central bottleneck (Zhu et al. 2025). Our audit asks whether independently generated translations preserve the same idea-selection decision, with fidelity labeled separately. Repeated trials. Repeated-run evaluations expose agent inconsistency (Rabanser et al. 2026; Mehta 2026). Mustahsan et al. quantify it with ICC (Mustahsan et al. 2025). Pass@k and self-consistency aggregate samples, while diverse query initialization searches trajectories (Chen et al. 2021; Wang et al. 2023b; Murali et al. 2026). Rollout Cards preserves rollout-level evidence and its reporting rules (Masters et al. 2026). Controlled second-pass studies decompose gains into re-solving, scaffold, and content effects (Ning et al. 2026c). Where these methods repeat a task or query, our audit fixes a semantic idea, samples independent translations, reruns saved artifacts, and tests winner stability. Best-of-k remains complementary. Measurement and experimental variance. Machine learning experiments vary across seeds, splits, hyperparameters, models, and benchmarks (Henderson et al. 2018; Bouthillier et al. 2021; Agarwal et al. 2021; Wortsman et al. 2022; Dehghani et al. 2021; Liang et al. 2023; Pineau et al. 2021). Those studies quantify uncertainty around specified pipelines. We repeat the translation of a semantic mechanism and ask whether legal realizations preserve its idea-level decision. Hyperparameters may be one open choice; fixing all choices yields a recipe, while tuning defines a richer implementation process. Construct-validity work separates a claimed concept from its measurement (Jacobs and Wallach 2021; Bean et al. 2025), and we apply reliability tools to the idea-to-implementation mapping (Shrout and Fleiss 1979; Brennan 2001). Artifact-search systems such as FunSearch and AlphaEvolve make artifact-level claims (Romera-Paredes et al. 2024; Novikov et al. 2025); idea reliability matters when their outputs credit a broader hypothesis or steer idea-level learning. 7 Scope and Limitations The reported rates apply to this task pool, four candidates per task, three implementations, two same-provider implementation processes, and one evaluator. The all-13 estimates describe the deployed workflow. Because semantic review occurred after execution and depends on policy judgments, mechanism-level promotion should instead use prospective gating. ICC varies with candidate spread and our zero-bounding rule. LOO also depends on a frozen cross-card alignment because the provider did not expose sampling seeds. Accordingly, 25.6–43.6% is conditional rather than an error rate for arbitrary projects. Fidelity labels were blinded and stress-tested but were not provided by humans, and ITT retains failure and drift. Materials was historically exposed and remains exploratory. We do not evaluate best-of-N stability, end-to-end tree search, universal thresholds, or downstream audit gains. OpenML is comparatively constrained. Larger deep-learning workflows add choices in data, architecture, optimization, checkpoints, and systems, creating more routes from one realization to a different idea-level conclusion. Their reversal rates require measurement; attribution grows with workflow complexity. 8 Implications for Automated Research Match evidence to the objective. Best-of-N search is appropriate for artifact delivery, whereas mechanism-level claims also require portfolio validity and reliability across implementations. Before a score prunes a branch, justifies transfer, enters research memory, or supports a mechanism-level claim, the candidate must denote a coherent intervention whose ranking survives another realization. Supply evidence for promotion decisions. Consequential checkpoints need not audit every candidate. The three-draw design triples executor assignments, while rerun calibration and fidelity review use only evaluator calls. A practical policy validates portfolios before implementation, requests additional implementations when candidates are close or implementation variation is high, and audits finalists before reusing their scores as evidence. Benchmark and leaderboard maintainers can adopt the schema to distinguish artifact scores from idea-level evidence. Large implementation variation calls for new implementations, rerun variation for a more stable evaluator, and low fidelity for tighter control. The following schema is not a universal threshold. Idea Reliability Audit checklist Unit: mechanism, invariants/open choices, portfolio gate. Process: executor/model, budget, independence, repeats. Adherence: blind validity/fidelity; failures retained in ITT. Noise: byte-identical reruns and evaluator policy. Evidence: idea ICC/uncertainty, LOO reversal, sensitivities. Build cumulative scientific knowledge. Cumulative use requires evidence that a mechanism, not only one artifact, remains reliable across implementations. Idea reliability matters when mechanisms are compared, falsified, transferred, or recombined across machine learning, materials workflows, simulations, data-analysis plans, wet-lab procedures, and multi-agent research trajectories (Wang et al. 2023a; Gottweis et al. 2026). It prevents a lucky realization or malformed candidate from steering the cumulative trajectory. 9 Conclusion One run is not an idea. Automated research systems often ask one realization-level score to do two jobs: deliver an artifact and decide which parent idea to retain, transfer, or remember. Best-of-N search can serve the first objective, but the second requires evidence that survives another legitimate translation of the same mechanism. We call dependence on the sampled translation the implementation lottery; idea reliability is the stability of idea-level evidence across such translations. The Idea Reliability Audit holds mechanism-level cards fixed, translates each card in three fresh sessions, uses outcome-blind fidelity labels, retains failure and drift under ITT, and reruns saved artifacts unchanged. Idea ICC measures how much within-task variation separates ideas, while LOO reversal asks whether a one-draw winner survives the other two implementations. Across 312 assignments on 13 tabular tasks, implementation choice accounts for 33% and 45% of within-task ITT variance, more than fivefold and tenfold the same-artifact rerun variance. One-draw winners reverse in 25.6% and 43.6% of sampled LOO decisions. The lottery persists in Bounded and Agentic, although their ordering reverses in the fresh extension. Outcome-blind card filtering retains all 13 tasks and leaves reversal at 33.3–43.6%, so flagged degeneracies do not explain the result. An exploratory Materials diagnostic shows the same directional pattern under deterministic evaluation and exposes portfolio validity as a separate gate. Together, these findings distinguish average-realization idea quality from best-of-N artifact utility. Search can optimize the delivered artifact; mechanism-level branching, transfer, and research memory additionally require evidence that survives multiple implementations. The measured rates vary by setting, but the attribution problem is structural whenever a realization updates beliefs about an idea. Search delivers artifacts; the audit supports cumulative science. References R. Agarwal, M. Schwarzer, P. S. Castro, A. Courville, and M. G. Bellemare (2021) Deep reinforcement learning at the edge of the statistical precipice. In Advances in Neural Information Processing Systems, Vol. 34. Cited by: §6. E. Aygün, A. Belyaeva, G. Comanici, et al. (2026) An AI system to help scientists write expert-level empirical software. Nature 654 (8120), p. 909–916. External Links: Document Cited by: §6. A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, et al. (2025) Measuring what matters: construct validity in large language model benchmarks. In Advances in Neural Information Processing Systems, Vol. 38. Cited by: §6. X. Bouthillier, P. Delaunay, M. Bronzi, A. Trofimov, B. Nichyporuk, J. Szeto, N. M. Sepahvand, E. Raff, K. Madan, V. Voleti, S. E. Kahou, V. Michalski, D. Serdyuk, T. Arbel, C. Pal, G. Varoquaux, and P. Vincent (2021) Accounting for variance in machine learning benchmarks. In Proceedings of Machine Learning and Systems, Vol. 3, p. 747–769. Cited by: §6. R. L. Brennan (2001) Generalizability theory. Springer, New York. External Links: Document Cited by: §2, §6. H. Chen, M. Xiong, Y. Lu, W. Han, A. Deng, Y. He, J. Wu, Y. Li, Y. Liu, and B. Hooi (2025) MLR-Bench: evaluating AI agents on open-ended machine learning research. In Advances in Neural Information Processing Systems, Vol. 38. External Links: Link Cited by: §6. M. Chen, J. Tworek, H. Jun, et al. (2021) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Cited by: §6. M. Dehghani, Y. Tay, A. A. Gritsenko, Z. Zhao, N. Houlsby, F. Diaz, D. Metzler, and O. Vinyals (2021) The benchmark lottery. External Links: 2107.07002 Cited by: §6. A. Dunn, Q. Wang, A. Ganose, D. Dopp, and A. Jain (2020) Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm. npj Computational Materials 6 (1), p. 138. External Links: Document Cited by: §4. A. Garikaparthi, M. Patwardhan, and A. Cohan (2026) ResearchGym: evaluating language model agents on real-world AI research. External Links: 2602.15112 Cited by: §1, §6. J. Gottweis, W. Weng, A. Daryin, et al. (2026) Accelerating scientific discovery with Co-Scientist. Nature 655 (8122), p. 487–496. External Links: Document Cited by: §6, §8. P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger (2018) Deep reinforcement learning that matters. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32. External Links: Document Cited by: §6. T. Hua, H. Hua, V. Xiang, B. Klieger, S. T. Truong, W. Liang, F. Sun, and N. Haber (2025) ResearchCodeBench: benchmarking LLMs on implementing novel machine learning research code. In Advances in Neural Information Processing Systems, Vol. 38. External Links: 2506.02314 Cited by: §6. A. Z. Jacobs and H. Wallach (2021) Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, p. 375–385. External Links: Document Cited by: §6. G. J. Kim, A. Wilf, L. Morency, and D. Fried (2026) From reproduction to replication: evaluating research agents with progressive code masking. In International Conference on Learning Representations, External Links: 2506.19724 Cited by: §6. P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al. (2023) Holistic evaluation of language models. Transactions on Machine Learning Research. Cited by: §6. C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha (2024) The AI scientist: towards fully automated open-ended scientific discovery. External Links: 2408.06292 Cited by: §1, §6. C. Lu, C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, D. Ha, and J. Clune (2026) Towards end-to-end automation of AI research. Nature 651 (8107), p. 914–919. External Links: Document Cited by: §6. C. Masters, Z. Liu, and S. V. Albrecht (2026) Rollout cards: a reproducibility standard for agent research. External Links: 2605.12131 Cited by: §6. A. Mehta (2026) When agents disagree with themselves: behavioral consistency as an uncertainty signal for LLM agents. External Links: 2602.11619 Cited by: §6. S. Murali, J. Coelho, J. Ning, J. Magalhães, B. Martins, and C. Xiong (2026) Beyond parallel sampling: diverse query initialization for agentic search. External Links: 2606.17209 Cited by: §6. Z. Mustahsan, A. Lim, M. Anand, S. Jain, and B. McCann (2025) Stochasticity in agentic evaluations: quantifying inconsistency with intraclass correlation. External Links: 2512.06710 Cited by: §6. D. Nathani, L. Madaan, N. Roberts, N. Bashlykov, A. Menon, V. Moens, M. Plekhanov, A. Budhiraja, D. Magka, V. Vorotilov, G. Chaurasia, D. Hupkes, R. S. Cabral, T. Shavrina, J. N. Foerster, Y. Bachrach, W. Y. Wang, and R. Raileanu (2025) MLGym: a new framework and benchmark for advancing AI research agents. In Conference on Language Modeling, External Links: Link, 2502.14499 Cited by: §1, §6. J. Ning, X. Li, J. Zeng, H. Kang, and C. Xiong (2026a) Auto research with specialist agents develops effective and non-trivial training recipes. External Links: 2605.05724 Cited by: §6. J. Ning, X. Li, J. Zeng, C. Xiong, and G. Ke (2026b) Closed-loop auto research for molecular property prediction: discovering and certifying generalizable improvements. External Links: 2606.22731 Cited by: §6. J. Ning, X. Li, and C. Yu (2026c) Revision or re-solving? decomposing second-pass gains in multi-LLM pipelines. External Links: 2604.01029 Cited by: §6. A. Novikov, N. Vũ, M. Eisenberger, et al. (2025) AlphaEvolve: a coding agent for scientific and algorithmic discovery. External Links: 2506.13131 Cited by: §6. J. Pineau, P. Vincent-Lamarre, K. Sinha, V. Larivière, A. Beygelzimer, F. d’Alché-Buc, E. Fox, and H. Larochelle (2021) Improving reproducibility in machine learning research: a report from the NeurIPS 2019 reproducibility program. Journal of Machine Learning Research 22 (164), p. 1–20. Cited by: §6. S. Rabanser, S. Kapoor, P. Kirgis, K. Liu, S. Utpala, and A. Narayanan (2026) Towards a science of AI agent reliability. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2602.16666 Cited by: §6. B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. R. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, P. Kohli, and A. Fawzi (2024) Mathematical discoveries from program search with large language models. Nature 625 (7995), p. 468–475. External Links: Document Cited by: §6. S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, M. Moor, Z. Liu, and E. Barsoum (2025) Agent laboratory: using LLM agents as research assistants. In Findings of the Association for Computational Linguistics: EMNLP 2025, p. 5977–6043. External Links: Document, Link Cited by: §1, §6. P. E. Shrout and J. L. Fleiss (1979) Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin 86 (2), p. 420–428. External Links: Document Cited by: §2, §6. C. Si, T. Hashimoto, and D. Yang (2026a) The ideation–execution gap: execution outcomes of LLM-generated versus human research ideas. In International Conference on Learning Representations, Cited by: §1, §6. C. Si, D. Yang, and T. Hashimoto (2025) Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers. In International Conference on Learning Representations, Cited by: §1, §6. C. Si, Z. Yang, Y. Choi, E. Candès, D. Yang, and T. Hashimoto (2026b) Towards execution-grounded automated AI research. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2601.14525 Cited by: §6. J. Vanschoren, J. N. van Rijn, B. Bischl, and L. Torgo (2014) OpenML: networked science in machine learning. SIGKDD Explorations 15 (2), p. 49–60. External Links: Document Cited by: §4. H. Wang, T. Fu, Y. Du, W. Gao, K. Huang, Z. Liu, P. Chandak, S. Liu, P. Van Katwyk, A. Deac, et al. (2023a) Scientific discovery in the age of artificial intelligence. Nature 620 (7972), p. 47–60. External Links: Document Cited by: §1, §8. X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou (2023b) Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations, Cited by: §6. L. Ward, A. Dunn, A. Faghaninia, N. E. R. Zimmermann, S. Bajaj, Q. Wang, J. Montoya, J. Chen, K. Bystrom, M. Dylla, K. Chard, M. Asta, K. A. Persson, G. J. Snyder, I. Foster, and A. Jain (2018) Matminer: an open source toolkit for materials data mining. Computational Materials Science 152, p. 60–69. External Links: Document Cited by: §4. M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt (2022) Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Proceedings of the 39th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 162, p. 23965–23998. External Links: Link Cited by: §6. Y. Yamada, R. T. Lange, C. Lu, S. Hu, C. Lu, J. Foerster, J. Clune, and D. Ha (2025) The AI scientist-v2: workshop-level automated scientific discovery via agentic tree search. External Links: 2504.08066 Cited by: §1. J. Yuan, X. Yan, B. Zhang, T. Chen, B. Shi, W. Ouyang, Y. Qiao, L. Bai, and B. Zhou (2025) Dolphin: moving towards closed-loop auto-research through thinking, practice, and feedback. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 21768–21789. External Links: Document Cited by: §1, §6. Y. Zhang, M. Khalifa, S. Bhushan, G. Murphy, L. Logeswaran, J. Kim, M. Lee, H. Lee, and L. Wang (2025) MLRC-Bench: can language agents solve machine learning research challenges?. In Advances in Neural Information Processing Systems, Vol. 38. External Links: Link Cited by: §6. Z. Zhang, N. Wang, S. Galhotra, and C. Cardie (2026) How far are we from true auto-research?. External Links: 2605.19156 Cited by: §6. X. Zhao, Z. Sang, Y. Li, Q. Shi, W. Zhao, S. Wang, D. Zhang, X. Han, Z. Liu, and M. Sun (2026) AutoReproduce: automatic AI experiment reproduction with paper lineage. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, p. 21920–21942. External Links: Document, Link, 2505.20662 Cited by: §6. L. Zheng, W. Chiang, Y. Sheng, et al. (2023) Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, Vol. 36. Cited by: §3. M. Zhu, Q. Xie, Y. Weng, J. Wu, Z. Lin, L. Yang, and Y. Zhang (2025) AI scientists fail without strong implementation capability. External Links: 2506.01372 Cited by: §6.