Paper deep dive
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
Hui Mao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/23/2026, 2:45:46 AM
Summary
The paper introduces 'Historical Backtesting,' a protocol for evaluating AI-generated scientific research questions by measuring their engagement in future scientific literature. Using an astronomy pilot dataset, the authors demonstrate that evidence-structure-first generation outperforms LLM-only prompting in identifying questions that are answered or have their premises refuted by future science. The study also highlights low inter-rater reliability among humans and LLM judges, suggesting that current evaluation taxonomies are flawed rather than the judges themselves.
Entities (7)
Relation Signals (5)
Historical Backtesting â evaluates â Scientific Questions
confidence 95% · We formalize historical backtesting as an evaluation protocol for scientific question discovery... a temporally isolated future corpus then determines whether each question was subsequently answered...
Evidence-structure-first generation â outperforms â LLM-only prompting
confidence 92% · evidence-structure-first generation outperforms LLM-only prompting at scale... evidence-structure-first generation resolves and refutes far more than direct LLM prompting (39% vs. 15% answered; 13% vs. 0% premise refutation...)
Astronomy v1 â uses â Historical Backtesting
confidence 90% · We release a reproducible astronomy benchmark instance... evaluated through the identical pipeline.
Frontier LLMs â agreeswith â Frontier LLMs
confidence 88% · frontier models agree with one another at Îș=0.60âso the common practice of certifying an LLM judge by modelâmodel agreement would have overstated its reliability threefold here.
Human Annotators â agreeswith â Human Annotators
confidence 88% · two careful humans agree with each other at only Îș=0.17
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus then determines whether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored. We release reproducible astronomy instances with temporally isolated corpora, frozen questions, auditable labels, four reference baselines, and a submission interface. Two findings result. First, evidence-structure-first generation outperforms LLM-only prompting: across a generator decomposition crossed with a four-cutoff stress test (2010-2024, 798 judged questions) whose last window postdates model training, LLM-only generation shows memorized relevance without specific foresight, while a generator using no model weights at all finds questions whose premises the future refutes in every era. Second, a seven-rater agreement study (two blinded human annotators, five judge models, 90 items) indicts the outcome taxonomy rather than the judge: two careful humans agree at kappa = 0.17, every judge model agrees with the professional annotator as well or better (0.17-0.26), and frontier models agree with one another at 0.60 -- certifying an LLM judge by model-model agreement would have overstated its reliability threefold. A prospective instance -- 200 questions frozen 2026-08-17, scored 2027-2030 -- is released so the central claims become contamination-free tests that time itself will grade.
Tags
Links
- Source: https://arxiv.org/abs/2608.16795v1
- Canonical: https://arxiv.org/abs/2608.16795v1
Trouble viewing inline? Open PDF directly â
Full Text
89,039 characters extracted from source content.
Expand or collapse full text
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot Evaluating AI-Generated Scientific Questions Against Future Scientific Progress Hui Mao Affiliation: Independent Researcher Email: hui.mao@alumni.upenn.edu August 17, 2026 Abstract Systems that generate scientific research questions are currently evaluated by expert scores, LLM-as-judge ratings, or curated case studiesâall subjective, none falsifiable. We propose a different standard: future scientific engagement as an observable, falsifiable proxy for one important dimension of a questionâs value. We formalize historical backtesting as an evaluation protocol for scientific question discovery: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus determinesâvia fixed retrieval, a citation-constrained judge, and a declared adjudication tierâwhether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored, and all metrics are defined independently of how questions are produced. We release a reproducible astronomy benchmark instance (cutoff 2020-12-31): a 2,512-paper past corpus, a temporally isolated 1,891-paper future corpus (2021â2026), frozen questions, retrieval records, adjudicated outcome labels, and one-command metric computation, plus a submission interface and four reference baselines evaluated through the identical pipeline. In an initial set of ten questions generated by an evidence-graph system from pre-cutoff literature only, all ten were substantively engaged by later literature: two were answered, seven partially addressed, one independently posed and still openâand one questionâs underlying premise (a strongly subsolar water abundance for HD 209458 b) was subsequently refuted by three independent analyses, the exact convergence the question called for. A scaled second instance (astronomy v1L: 424 frozen baseline questions against a 5,754-paper future corpus) then stress-tests the small-sample conclusions and revises two of them: engagement rates do discriminate at n=125n=125 (random 73% vs. direct-LLM 96%, p<10â4p<10^-4), and premise refutation is rare but not uniqueâchasing highly cited results catches refutations on 3.2% of questions, while random templates stay at zero and acquire a measurable 11%-answered floor. Finally, we turn the benchmarkâs deepest threatâLLM weights that have read the futureâinto its subject: a generator decomposition (LLM-only vs. deterministic evidence-structure vs. structure-plus-LLM verbalization) crossed with a four-cutoff temporal stress test (2010â2024, 798 judged questions) whose last window postdates the modelâs training. LLM-only generation shows memorized relevance without specific foresight: near-ceiling engagement and the closest phrasing to future literature at every cutoff, but an answered rate flat across the training boundary, indistinguishable from random templates, and zero premise refutations outside the deepest-history era. A weight-free structural generator finds engaged questions at every cutoff, and adding the LLM back as a pure verbalizer refutes premises in every era including the post-training oneâlocating the foresight signal in pre-cutoff evidence structure, with the LLM as a separable realization layer. We then validate the measurement instrument itself with a seven-rater agreement study (two independent blinded human annotators, five judge models, 90 items) and report what it shows: two careful humans agree with each other at only Îș=0.17Îș=0.17, every judge model agrees with the professional annotator as well as or better than the humans agree with each other (Îș=0.17Îș=0.17â0.260.26), and frontier models agree with one another at Îș=0.60Îș=0.60âso the common practice of certifying an LLM judge by modelâmodel agreement would have overstated its reliability threefold here. The outcome taxonomy, not the judge, fails validation; absolute rates are therefore rater-relative throughout, while the paperâs comparative claims are checked under three judges and survive with no reversals. Two findings result: evidence-structure-first generation outperforms LLM-only prompting at scale, and outcome taxonomies for scientific-question evaluation need a measured humanâhuman reliability gate before any judge, human or model, is scored against them. A prospective instanceâ200 questions from four generators, frozen 2026-08-17 with a 2027â2030 scoring windowâis released so the central claims become contamination-free tests that time itself will grade. 1 Introduction A growing family of systems claims to generate scientific research questions, hypotheses, or ideas (7; 13; 1). How do we know whether any of them is good at it? Today, essentially every evaluation falls into one of three patterns: an expert score (a panel rates novelty and significance on a Likert scale), an LLM score (a language model rates the same properties), or a case study (a handful of generated ideas is narrated persuasively). All three share the same defect: they are subjective. Expert panels disagree with each other and with themselves (11); LLM judges inherit the biases of their training distribution and can be steered by phrasing; case studies are selected by the authors. None of these evaluations can be wrong in a way that data could demonstrate. We hold our own instrument to that standard too: Section 10 subjects it to a seven-rater reliability study and reports the result, which is not flattering, in full. Science itself offers a harder criterion. Research questions are bets about where inquiry should go next, and the scientific community eventually settles those bets: it invests observing time, funds follow-ups, writes papers that answer some questions, poses others independently, and refutes the premises of a few. This suggests a measurement standardâdeliberately a proxy, not a definition: Future scientific engagement provides an observable, falsifiable proxy for one important dimension of a questionâs value. We do not claim engagement defines value: community attention carries popularity bias (Section 12), and a question can be excellent yet ignored for want of an instrument. The claim is narrower and stronger where it countsâengagement is the one dimension of value that is observable from frozen public data, and therefore the one on which systems can be compared without asking anyoneâs opinion. The standard becomes an evaluation protocol the moment we rewind the clock. Fix a historical cutoff. Give a system only the literature available before the cutoff. Freeze the questions it generates. Then let the literature published after the cutoffâwhich the system never sawâgrade the bet: Was the question answered? Substantially advanced? Independently posed by working scientists? Ignored? Was its underlying premise confirmed, or refuted? We call this procedure historical backtesting, by analogy with the evaluation of trading strategies on held-out past data (2), and with recent forecasting benchmarks that score models against events occurring after training (14). Past Literature (†cutoff)Generate Scientific QuestionsFreeze QuestionsFuture Literature (after cutoff)Retrieve EvidenceHistorical Outcome AssessmentBenchmark Metrics any system: evidence graph, prompted LLM, heuristic, human temporally isolated; never seen at generation time constrained judge; adjudicated tier for released labels Figure 1: The historical backtesting protocol. No question generator, evidence graph, or particular LLM appears in the loop: the protocol evaluates frozen question lists, whatever produced them. Crucially, the protocol contains no question generator (Figure 1). It takes a frozen list of questions as input and returns outcome labels and metrics as output. Any systemâan evidence-graph pipeline, a prompted LLM, a citation heuristic, a human scientistâcan be evaluated under identical conditions. This is what makes it a benchmark rather than a validation appendix for one particular architecture. Contributions. We make five contributions; the first two are claims ordered by strength, the remaining three are measurements and artifacts: 1. Protocol (strong). We formalize historical backtesting as an evaluation protocol for scientific question discovery: temporal isolation rules, a question-freezing requirement, fixed future-evidence retrieval, a two-dimensional outcome taxonomy that separates a questionâs fate (answered, partially_addressed, posed_but_open, not_addressed) from its premiseâs fate (supported, refuted, weakened, still_plausible, not_applicable), and metrics defined over those labels (Sections 3 and 4). 2. Benchmark instance (medium). We release a reproducible astronomy instance with temporally isolated past and future corpora (2,512 and 1,891 papers; cutoff 2020-12-31), frozen questions, released retrieval records, adjudicated labels, one-command metric computation with CI-enforced reproducibility, a submission format, and four reference baselines evaluated through the identical pipeline (Sections 5 and 6). 3. Empirical findings (cautious, and separated by sample size). Our statistically supported claim concerns a class of methods, not a single system: across 125-question submissions, evidence-structure-first generation resolves and refutes far more than direct LLM prompting (39% vs. 15% answered; 13% vs. 0% premise refutation, p=3Ă10â5p=3Ă 10^-5), and this holds at four historical cutoffs including one whose future postdates the modelâs training data (Section 9). A weight-free structural generatorâno LLM anywhereâoutperforms LLM-only prompting on resolution, locating the foresight signal in pre-cutoff evidence structure rather than in model weights. Separately, and as an illustrative case rather than a statistical claim, a ten-question evidence-graph submission had every question engaged by later literature, one of them by refuting the premise it challenged (Section 7). At n=10n=10 that submission cannot be ranked against baselinesâdetecting its apparent advantage would require nâ209nâ 209 per armâand we make no such ranking claim. We do not claim that AI reliably identifies the most valuable future scientific questions; we claim the protocol can tell us, eventually, whether it can, and that it already discriminates between generator families. 4. Measurement validity (adverse, and general). We validate the measurement instrument itself with a seven-rater agreement study: two independent blinded human annotators and five judge models on 90 items. Humans agree with each other at Îș=0.17Îș=0.17; every model matches the professional annotator as well as the humans match each other; frontier models agree with one another at Îș=0.60Îș=0.60. The outcome taxonomy, not the judge, fails validationâand certifying an LLM judge by modelâmodel agreement, the fieldâs common shortcut, would have overstated reliability threefold here (Section 10). 5. Prospective instance (frozen, unscoreable until 2031). Two hundred questions from four generators, frozen at cutoff 2026-08-17 with a pre-registered 2027â2030 scoring window and published corpus manifests and hashesâthe contamination-free test that time itself will grade (Section 14). What we believe is ultimately most useful here is not any single result but the change of category: scientific question quality moves from a matter of taste (âthis question seems interestingâ) to a measured, comparable quantity (âunder historical backtesting, 100% of this systemâs questions were engaged by later literature; 10% led to a premise refutationâ). Every artifact needed to run the protocolâdata, code, labels, and checksâis public, and Astronomy v1 is offered as the first instance of the benchmark, not as its definition. 2 Related Work Automated scientific discovery and question generation. Computational discovery has a long lineage, from rule-based rediscovery of physical laws (6) through closed-loop robot scientists (4) to the Nobel Turing Challengeâs call for AI scientists (5). Recent LLM-based systems generate research ideas, hypotheses, or full papers: literature-based generation (13), agentic idea refinement (1), and end-to-end automated research (7). Literature-based discovery pioneered the underlying intuition that recombining published evidence can anticipate findings later verified empirically (12). Our work is orthogonal to all of these: we do not propose a better generator; we propose the missing evaluation. Evaluating generated ideas. Existing evaluations are dominated by human preference and LLM scoring. 11 ran a large expert study comparing human and LLM research ideas on rated novelty and excitementâthe most rigorous instance of the expert-score paradigm, and still a measurement of opinion at generation time rather than of what the ideas turned out to be worth. LLM-as-judge scoring inherits known biases (position, verbosity, self-preference) and, for questions about the future, cannot be validated against ground truth at all. Historical backtesting replaces both with an outcome variable that exists independently of any rater: the subsequent behavior of the scientific community. Backtesting and forecasting benchmarks. Scoring a strategy on held-out history is standard in quantitative finance, along with well-documented failure modesâoverfitting to the backtest itself (2)âthat motivate our freezing and no-overwrite rules. Forecasting benchmarks score models on events that resolve after training (14); retrodictive evaluation with temporal holdouts is likewise used to test whether models anticipate later discoveries (12; 4). We transplant this design to a harder target: not whether a stated event occurs, but whether an open-ended research question earns the communityâs future investment. Code benchmarks such as SWE-bench (3) demonstrated how a well-specified task format plus frozen data can reorganize a research area around measurable progress; we aim the same mechanism at question discovery. Data contamination. Temporal splits are increasingly used to control LLM memorization in evaluation. Our protocol controls the retrieval channel completely (corpus manifests, isolation rules, CI checks) and treats the weights channelâmodels whose training data postdates the cutoffâas a declared, audited threat rather than a solved problem (Section 12). 3 The Historical Backtesting Protocol The protocol evaluates a set of frozen questions Q=q1,âŠ,qnQ=\q_1,âŠ,q_n\ against a future corpus. It has six steps; each is fully specified so that two groups running the same instance obtain the same measurement. 3.1 Step 1: Choose a cutoff A historical date T (Astronomy v1: 2020-12-31) splits the literature into a past corpus Corpus A (everything available up to T) and a future window realized as an isolated corpus Corpus B (strictly after T). The cutoff must be far enough in the past for the community to have had time to actâwe recommend â„4â„ 4 yearsâand recent enough that the past corpus reflects a modern research frontier. 3.2 Step 2: Generate questions Any method may generate questions: an evidence-graph pipeline, a prompted LLM, a heuristic over citation statistics, a human expert. The only requirements are (i) the generator consumes only Corpus A evidence, and (i) every question records the pre-cutoff evidence it is grounded in (source_evidence_ids). A submission whose source evidence postdates T is invalid, mechanically (scripts/validate_cutoff.py). 3.3 Step 3: Freeze Questions are serializedâidentifier, text, cutoff, generating system, source evidence, system-assigned rankâwith frozen: true before any access to post-cutoff literature, and are never edited afterwards. Freezing is the protocolâs load-bearing rule: without it, question text drifts toward what the evaluator has meanwhile learned the future contains, and the backtest silently becomes a description of the future rather than a prediction of it (2). In the released implementation frozen files are append-only and guarded by CI; editing a released question mints a new versioned instance rather than overwriting the old one. 3.4 Step 4: Define the future window Corpus B is collected under its own frozen manifest (query set, date window, deduplication rules) and stored separately from Corpus A; records from Corpus B must never enter the generation pipeline. The Astronomy v1 window is 2021â2026. Bounding the window matters for comparability: âeventually engagedâ is not a fixed target, but âengaged within k yearsâ is. 3.5 Step 5: Retrieve future evidence For each frozen question, the question text and every Corpus B document (title + abstract) are embedded (text-embedding-3-small); the top-k documents by cosine similarity (k=8k=8) become the candidate evidence. Retrieval is deliberately fixed and deliberately simple: systems are compared on their questions, not their retrievers, and reviewers can inspect exactly which documents the judge saw because retrieval records are part of the release. 3.6 Step 6: Assess outcomes A judge reads the question and its retrieved candidates and assigns two independent labels (Table 1): the fate of the question and the fate of its premise. Table 1: Outcome taxonomy v1.0. The two dimensions are labeled independently. Dimension Label Meaning outcome answered Future evidence substantially answers the question (including by refuting its premise) partially_addressed Important, directly relevant progress; core question unresolved posed_but_open The community independently poses essentially the same question without resolving it not_addressed No meaningful follow-up in the future corpus premise_status supported Future evidence confirms the underlying premise refuted Future evidence falsifies the underlying premise weakened Substantial doubt cast without falsification still_plausible The premise was not directly tested after the cutoff not_applicable The question rests on no contestable premise Two design decisions deserve emphasis. First, the dimensions are separated because the single most informative outcome a backtest can surfaceâthe community answered this question by refuting its premiseâis inexpressible in a flat label set: it is simultaneously a resolution (answered) and a falsification (refuted). Our pilotâs headline case (Section 7.1) is exactly of this type. Second, premise refutation is scored as a success of the question, not a failure: a question that provokes the community into overturning one of its own published conclusions has done the most a question can do. The judge operates under hard constraints enforced outside the model: it may cite only bibcodes from the retrieved candidates (violations are errors, never silently dropped); topical similarity is explicitly insufficientâthe cited paper must bear on the questionâs actual test or premise; unknown labels fall back to the most conservative value and are flagged. Labels then occupy one of two declared tiers. Adjudicated labels have passed human review under written guidelines (citation validity, the engagement bar, the outcome/premise split), with the adjudication log released; the headline labels of a released instance are required to be of this tier (Astronomy v1âs are). Judge-only labels have not, are marked as such wherever reported, and are the tier at which this paperâs large-n comparative studies run (Sections 8â9). The tier is part of every resultâs provenance; conflating them is a protocol violation. 3.7 Submissions A system is evaluated by submitting a directory containing questions.jsonl (the frozen questions) and metadata.json (system description, including a mandatory declaration of any LLM components and their versions, for contamination auditing). The benchmark pipelineâisolation checks, retrieval, judging, adjudication where the tier requires it, metrics, reportâis identical for every submission; the system whose questions we evaluate in Section 7 interacts with the benchmark only through this interface. 4 Benchmark Metrics Let Q be the n frozen questions of a submission, with outcome labels oâĄ(q)o(q) and premise labels pâĄ(q)p(q) as in Table 1, and let EâĄ(q)E(q) be the set of independent post-cutoff papers cited as supporting evidence for qâs label. All rates are over n, so the four outcome rates sum to one. Table 2: Benchmark metrics v1.0. All are computed by benchmark/metrics.py from the released annotation records. Metric Definition Coverage (future attention rate) 1nâ|q:oâĄ(q)â not_addressed| 1n\,|\q:o(q)â not\_addressed\| â did future science engage the question at all? Answer rate 1nâ|q:oâĄ(q)=answered| 1n\,|\q:o(q)= answered\| Partial rate 1nâ|q:oâĄ(q)=partially_addressed| 1n\,|\q:o(q)= partially\_addressed\| Open rate 1nâ|q:oâĄ(q)=posed_but_open| 1n\,|\q:o(q)= posed\_but\_open\| â the community independently recognized the question Premise refutation rate 1nâ|q:pâĄ(q)=refuted| 1n\,|\q:p(q)= refuted\| â questions that led to an established conclusion being overturned Evidence strength mean |EâĄ(q)||E(q)| over engaged questions, and the multi-source rate 1nâ|q:|EâĄ(q)|â„2| 1n\,|\q:|E(q)|â„ 2\| â is the label supported by multiple independent papers? Lead time years between question submission and the first time the community independently poses the same question Community attention volume of future investment engaging the question: papers, citations, review mentions, and major observing programs (e.g. JWST/HST proposals) Coverage vs. answer rate. Coverage asks whether the question pointed anywhere the community went at all; the answer/partial/ open decomposition asks what happened when it got there. A system can maximize coverage with fashionable-topic questions, which is why coverage is never reported alone (Section 12 discusses the popularity confound). Premise refutation rate. This is the metric we most want the field to adopt. Questions that trigger refutations are the rarest and arguably most valuable output of question discoveryâthey mark places where the literatureâs accepted conclusions were wrong and where a well-aimed question preceded the correction. Under the two-dimensional taxonomy the refutation is recorded without erasing the fact that the question was thereby answered. Lead time. If a system poses a question at the cutoff and the community first independently poses it in year T+âT+ , the system led the field by â years; averaged over questions this yields a comparable earliness score. Measuring â requires identifying community first-posed dates, which demands careful review-literature annotation we do not yet have. Astronomy v1 therefore reports mean_lead_time_years: null rather than a number we cannot defend; the released records do include a weaker, well-defined lower boundâfirst-engagement lag, the years from cutoff to the earliest judge-cited supporting paper (pilot mean 2.9, range 1â5)â which should not be confused with lead time. Community attention. Beyond binary engagement, the volume of future investment (paper counts, citations to engaging papers, review mentions, dedicated observing programs) reflects how much the community cared. v1 records the ingredients (supporting bibcodes, their venues and years) and reports evidence strength; a calibrated attention index is future work, and Section 12 explains why raw attention must never be the headline metric. Reporting requirements. A benchmark report must state: the instance and protocol versions, n, all outcome and premise rates, the evidence-strength pair, and either lead time or an explicit null. The released implementation produces exactly this (results/astronomy_v1/metrics.json) with one command, and CI fails if the committed numbers do not reproduce from the raw annotations. 5 The Astronomy v1 Instance Astronomy v1 instantiates the protocol in exoplanet atmospheresâa domain chosen because it uniquely combines a fast-moving literature, structured catalogs, and space-telescope archives, and because the 2021â2026 window contains a natural experiment: JWST began delivering data mid-window, resolving questions that were unanswerable at the cutoff. Table 3 summarizes the instance. Table 3: Astronomy v1 at a glance. Corpus manifests, frozen questions, retrieval records, and adjudicated labels are all released. Historical cutoff 2020-12-31 Domain scope atmospheric composition in transmission spectra; cloud/haze degeneracies; instrument systematics Corpus A (past) 2,512 deduplicated papers, 2015â2020 (NASA ADS; 500-paper full-text core) Corpus B (future) 1,891 unique papers, 2021â2026, temporally isolated Questions 10, frozen, ranked, with pre-cutoff source evidence Retrieval text-embedding-3-small, cosine, top-8; records released Judge gpt-4.1, temperature 0, citation-constrained; human-adjudicated Corpora. Both corpora are defined by frozen manifestsâADS query sets, date windows, deduplication and filtering rules (records without abstracts are dropped)ârather than by bulk data dumps: the manifests are committed, and a script rebuilds either corpus from its manifest via the ADS API. Corpus Bâs manifest adds targeted follow-up queries for the questionsâ objects (HD 189733, HD 209458, WASP-12, WASP-121, TRAPPIST-1) so that engagement is measured against the relevant future literature rather than against whatever a generic query happens to return. Leakage controls. Temporal isolation is enforced mechanically, not editorially: (1) nothing dated after the cutoff may enter Corpus A, including catalog rows updated post-cutoff; (2) every questionâs source evidence must predate the cutoff; (3) every retrieved document must postdate it; (4) the judge may cite only retrieved candidates; (5) question/retrieval/annotation records must align one-to-one; (6) Corpus Bâs window must start strictly after the cutoff. All six checks run in continuous integration on every change to the released data, together with a check that the released metrics.json reproduces bit-identically from the raw annotations. Questions. The ten frozen questions were generated by an evidence-graph system (10) from Corpus A only: claims with provenance are extracted from the full-text core, cross-paper tensions are detected and typed (observational tensions, methodological challenges, single-dataset conclusions, independent qualifications), and surviving signals are refined into ranked, falsifiable questions. For the benchmark, that system is submission evidence_graph_v1âevaluated through the same interface as any future submission. Each released record carries the question text, system rank, signal type, target objects, and pre-cutoff source bibcodes; a curation log documenting human edits made before freezing (including a near-duplicate merge and presupposition fixes, Section 11) is released for audit. What is released. Frozen questions; corpus manifests; per- question retrieval records (model, window, top-k, judge-cited documents, top-1 similarity; full ranked lists with scores are scheduled for v1.1); adjudicated outcome annotations with rationales and supporting bibcodes; the adjudication log; computed metrics and per-question results; four frozen baseline submissions with their full retrieval records (per-document scores included), judge annotations, and reports (Section 6); the scaled v1L instance (manifests, configs, 424 frozen baseline questions, retrieval records with per-document scores, judge annotations, per-system reports, and the statistical comparison script of Section 8); the tension-pair generators, the four-cutoff stress-test corpora manifests, all 798 stress-test judgments, and the deterministic specificity rubric of Section 9; and the complete pipeline code with tests. The data format is deliberately plain (JSONL + JSON manifests) so that other groups can mint new instancesâdifferent domain, different cutoffâby writing two manifests and one config file. 6 Baselines A benchmark that only ever scored one system would be a validation appendix. Astronomy v1 ships four reference baselines, chosen to bracket the interesting comparisons; each consumes only Corpus A records, freezes its output before any future-corpus access, and is evaluated through the identical pipeline. B1: Random claims. Sample random pre-cutoff papers and template their headline result into a robustness question. The floor: any system must beat chance-directed attention. B2: Direct LLM. Give an LLM (gpt-4.1, 60 sampled pre-cutoff abstracts) a request for the most valuable open questions. The âwhy not just ask GPT?â comparison. Because a modern LLMâs weights postdate the cutoff, this baseline is also a contamination probe: performance that vanishes for post-training-cutoff instances indicates memorized hindsight rather than generation ability (Section 12). B3: Review future work. Extract explicitly posed open questions from pre-cutoff review papers. The strongest natural reference: a discovery system is interesting only if it adds value over questions the community had already written down. Note this baseline should score highly on coverage by constructionâthese questions are known community prioritiesâso the discriminating metrics are premise refutation and lead time, where a copied question can never lead the field. B4: Citation leaders. Template follow-up questions from the most-cited pre-cutoff papers. Tests whether chasing prominence matches structured evidence analysis. One caveat is built in: ADS citation counts are fetched at corpus-rebuild time and therefore include post-cutoff citations, so this baseline selects papers with hindsight knowledge of which pre-2021 work the future found importantâa bias in its favor that a cutoff-dated citation snapshot would remove. Evaluation conditions. All four baselines were run through the released pipeline: top-8 retrieval with text-embedding-3-small against the manifest-rebuilt future corpus restricted to the frozen window (2,384 records, a superset of the frozen 1,891 due to retroactive ADS indexing, containing all 22 citations in the released annotations), then gpt-4.1 judging at temperature 0 under the citation constraints of Section 3. Baseline labels are judge-only: they have not received the human adjudication that the released evidence_graph_v1 labels did. To make that comparison honest, Table 4 also reports a judge-only rerun of the evidence-graph submission under exactly the baseline conditions. The rerun doubles as a replication check: it reproduces the released mean top-1 retrieval similarity (0.668 vs. 0.666) and lands within one label of the adjudicated results (coverage 90% vs. 100%, answered 10% vs. 20%, refutation 10% = 10%)âthe deltas are exactly the two labels adjudication had strengthened, so judge-only scoring reads as the conservative floor of the adjudicated score. Table 4: Astronomy v1 leaderboard (n=10n=10 questions per system). Adjud. = human-adjudicated labels; judge-only rows are directly comparable to each other. Cov. = future-attention rate; Ans. = answered; Part. = partially addressed; Open = posed but open; Ref. = premise refuted; s1s_1 = mean top-1 retrieval similarity. Lead time is null for all systems until community first-posed dates are annotated (Section 4). System Eval Cov. Ans. Part. Open Ref. s1s_1 evidence_graph_v1 (10) adjud. 100% 20% 70% 10% 10% 0.666 evidence_graph_v1 (rerun) judge 90% 10% 80% 0% 10% 0.668 B4 citation leaders judge 100% 30% 70% 0% 0% 0.632 B3 review future work judge 100% 10% 70% 20% 0% 0.593 B2 direct LLM judge 90% 0% 90% 0% 0% 0.732 B1 random claims judge 70% 0% 70% 0% 0% 0.611 Reading the leaderboard. Three observations, offered with the n=10n=10 caution of Section 7 applying to every row. Coverage saturates. Every non-random system scores 90â100% on future attention: in a field this active, any fluent, topical question attracts partial engagement within five years. Coverage separates the floor (random claims, 70%, the only system with three not_addressed labels) from everything else, and nothing elseâwhich is why the protocol never reports it alone (Section 12). Answered-rate comparisons need reading, not just ranking. The citation-leader baseline posts the highest judge-only answered rate (30%). Two mechanisms inflate it: its hindsight-biased paper selection (above), and its templateââdoes the conclusion of highly cited paper X hold?ââwhich pattern-matches the replication studies that prominent results reliably attract, so the judge can mark it answered whenever the community re-examined a famous result for any reason. What the template cannot do is risk anything: B4 refuted no premise, and by construction a question of the form âis the famous result right?â poses nothing the community was not already testing. The evidence-graph submissionâs answered questions, by contrast, specified particular tests (Section 7.1) and include the leaderboardâs only refuted premiseâon both adjudicated and judge-only rows. The contamination probe registers a signal. The direct-LLM baseline has by far the highest retrieval similarity to future literature (s1=0.732s_1=0.732 vs. 0.593â0.668 for every other system) and the highest mean support count (4.6 cited papers per question)âits questions are phrased in the way the 2021â2026 literature would come to phrase them, consistent with weights that have read that literature. Yet it resolves nothing: 0% answered, 0% refuted, 90% partial. The pattern suggests phrasing-level contamination without commitment to falsifiable specificsâbroad, well-aimed questions that everything engages and nothing settles. Prospective instances (Section 14) will separate the two channels definitively. Baseline generators, frozen submissions, retrieval records with full per-document scores, judge annotations, and per-system reports are all released; the leaderboard file is designed for external submissions to append to. All three observations above are drawn from ten questions per system; Section 8 re-examines them on a scaled instance with 424 baseline questions, and two of the three require revision there. 7 Pilot Results: Historical Validation We ran the full protocol on the ten frozen evidence_graph_v1 questions. Headline numbers: every question was substantively engaged by the 2021â2026 literature (coverage 100%); two were answered, seven partially addressed, one independently posed and still open; one premise was refuted. Evidence strength: 2.4 supporting papers per question on average, with 60% of labels supported by â„2â„ 2 independent papers. The earliest judge-cited engagement came 1â5 years after the cutoff (mean 2.9). Table 5 gives the per-question picture. Table 5: Per-question outcomes for evidence_graph_v1 on Astronomy v1 (system rank order; |E||E| = independent supporting papers; year = earliest cited engagement). Rank ID Outcome Premise |E||E| Year 1 q_002 posed_but_open still_plausible 1 2022 2 q_001 partially_addressed supported 3 2021 3 q_004 partially_addressed supported 1 2021 4 q_008 answered refuted 3 2025 5 q_007 partially_addressed still_plausible 3 2024 6 q_010 partially_addressed still_plausible 1 2022 7 q_005 answered supported 2 2024 8 q_009 partially_addressed supported 5 2021 9 q_006 partially_addressed still_plausible 1 2025 10 q_011 partially_addressed still_plausible 4 2024 7.1 Case study: a premise refuted (q_008, HD 209458 b) From pre-2021 evidence, the system flagged a single-dataset conclusion: the influential retrieval of a strongly subsolar water abundance at the terminator of HD 209458 b (9). The frozen question asked, in 2020 terms: Is the strongly subsolar terminator water abundance retrieved for HD 209458 b a property of the atmosphere or an artifact of retrieval assumptions, as tested by comparing independent retrieval frameworks on the same and on independent datasets? The 2021â2026 literature then performed exactly the test the question specified (Figure 2): a reanalysis of HST and JWST spectra with improved systematics treatment and Bayesian model averaging, an independent retrieval framework demonstrating that free-vs-equilibrium chemistry assumptions span the subsolar-to-solar range, and ground-based high-resolution spectroscopy constraining the abundance independently of space-based data. All three converge: the terminator water abundance is consistent with solar, and the strongly subsolar value was an artifact of earlier retrieval assumptions and data systematics. Under the taxonomy this is answered + refuted: the question was resolved by the community overturning the premise the question challengedâan outcome that no expert score assigned in 2020 could have certified, and precisely what backtesting exists to detect. year201720202021202320252026 subsolar H2O retrieved for HD 209458 b frozenquestionartifact or atmosphere?cutoff2020-12-31 three independent reanalyses converge: H2O consistent with solar answered + premise refuted future corpus window (2021â2026) Figure 2: Timeline of q_008. The question was frozen from pre-2021 evidence; by 2025 three independent analyses had performed the test it specified and refuted its premiseâthe strongly subsolar water abundance was a retrieval artifact. 7.2 Secondary observations The top-ranked question is independently posed and open. The systemâs rank-1 question (q_002: how terminator heterogeneity biases the WASP-12 b water abundance and C/O ratio) tracks a methodological concern the community has since engaged in general form â inhomogeneous-terminator biases in retrievals â without resolving it for WASP-12 b specifically: posed_but_open. A question the field poses but has not answered is a live research target; that the systemâs top pick lands there is the behavior a ranking is supposed to produce, though n=1n=1 at rank 1 proves nothing by itself. An answered null result. q_005 asked whether HST transmission data independently support NH3 or HCN in HD 209458 b â a molecular-detection claim from the same single-dataset analysis as q_008 (9). By 2024â2025, high-resolution spectroscopy had placed stringent upper limits on both species: answered, premise supported (the data indeed do not independently support the detection). Backtesting counts a cleanly resolved null exactly as it counts a positive. Instrument-gated engagement. q_011 (whether JWST validated pre-launch predictions of TRAPPIST-1 CO2 detectability, 8) could not have been engaged before JWST flew; its first cited engagement is 2024 and stellar contamination has so far prevented a definitive test. Engagement timing is partly an instrument schedule, not purely a question-quality signalâa confound Section 12 treats explicitly. What these ten questions do not show. With ten questions from one system in one domain, rates carry wide intervals (the exact 95% ClopperâPearson interval for 10/10 coverage is [0.69,1.0][0.69,1.0]), the baseline rows are judge-only rather than adjudicated (Section 6), and the generating systemâs LLM components postdate the cutoff (Section 12). Moreover, Section 10 shows all absolute rates are rater-relative: â10/10 engagedâ is this instanceâs adjudicated reading, made by the authors, not a rater-free fact. The pilot demonstrates that the protocol runs end-to-end, yields auditable, reproducible labels, and prices its baselines; it does not establish that any system reliably anticipates future science. 8 Scaling the Baselines: Astronomy v1L Every baseline conclusion in Section 6 rests on ten questions per system. To test which of them survive a larger sample, we minted a second instance, astronomy v1L: same cutoff, same retrieval and judge settings, but corpora rebuilt from broadened manifests (12 past-corpus and 18 future-corpus ADS query sets covering clouds and hazes, atmospheric escape, phase curves, high-resolution spectroscopy, and JWST; 4,040 past and 5,754 frozen-window future recordsâ1.8Ă and 2.4Ă the v1 corpora) and baselines scaled to 125 questions each. The reviewâfuture-work extractor is the exception by necessity: it exhausts the supply of explicitly posed questions in all 4,040 pre-cutoff abstracts at 49âitself a finding; the communityâs already-written-down questions are a finite resource. In total v1L evaluates 424 frozen baseline questions plus the ten evidence-graph questions re-run as a cross-instance anchor, all judge-only. The v1 instance and its released data are untouched. Table 6: Astronomy v1L results (judge-only). Brackets are exact 95% ClopperâPearson intervals. Eng. = future-attention rate; Ans. = answered; Ref. = premise refuted; s1s_1 = mean top-1 retrieval similarity; lag = mean years to earliest cited engagement. The evidence-graph row is the ten v1 questions re-evaluated on the v1L corpus (anchor), not a scaled submission. System n Eng. Ans. Ref. s1s_1 lag evidence_graph (anchor) 10 80% [44,97] 30% [7,65] 10% [0,45] 0.682 3.5 B2 direct LLM 125 96% [91,99] 15% [9,23] 0% [0,3] 0.722 1.9 B4 citation leaders 125 87% [80,93] 26% [18,34] 3.2% [1,8] 0.632 2.5 B3 review future work 49 92% [80,98] 22% [12,37] 0% [0,7] 0.567 2.4 B1 random claims 125 73% [64,80] 11% [6,18] 0% [0,3] 0.622 2.8 Table 6 gives the scaled results. Sample size changes two of Section 6âs three conclusions and sharpens the thirdâwhich is the point of running the experiment. Revised: coverage does discriminate at scale. At n=10n=10 every non-random system sat at 90â100% engagement and we concluded coverage separates only the floor. At n=125n=125 the rates pull apart: random claims 72.8% [64,80], citation leaders 87.2%, direct LLM 96.0% [91,99]; random vs. direct-LLM engagement differs at p<10â4p<10^-4 and random vs. citation leaders at p=0.007p=0.007 (two-sided Fisher). The v1 reading was a small-sample artifact. What survives is the ceiling: fluent, topical LLM questions still approach saturation, so coverage separates the bottom and middle of the range while compressing the topâit remains unusable as a sole metric. Revised: premise refutation is not unique to the evidence-graph system. At n=10n=10 no baseline refuted a premise; at n=125n=125 the citation-leader baseline catches four refutations (3.2% [1,8]): the subsolar-water conclusion for HD 209458 b (the same 9 result behind q_008), the methane-depleted atmosphere of K2-18 b overturned by JWST, TiO in WASP-121 b unconfirmed by later data, and systematic bias found in benchmark ultracool-dwarf retrievals. Asking âis the famous result right?â of enough famous results does eventually catch the ones that fallânote its hindsight-biased selection (Section 6) works in its favor here, since post-cutoff citation counts are inflated by exactly the controversies that produce reversals. Three things remain true. Refutation is the rarest outcome for every system (random templates: 0/125, upper bound 2.9%ârefutations are not free; a question must aim at a contestable claim). The evidence-graph submissionâs nominal rate stays highest (10% vs. 3.2%), but n=10n=10 cannot establish superiority (p=0.32p=0.32; the comparison needs nâ209nâ 209 per arm for 80% power, Section 12). And the two routes to a refutation differ qualitatively: prominence-chasing rediscovers that famous claims attract scrutiny, while the evidence-graph question specified the decisive test from pre-cutoff evidence tensions (Section 7.1). The scaled data cannot yet separate those routes quantitatively; a scaled evidence-graph submission could. Sharpened: there is a nonzero answered floor. Random robustness templates get answered 11.2% [6,18] of the timeâthe field re-examines even arbitrarily chosen results at a measurable base rate. The v1 estimate of that floor (0/10) was too flattering to every other system: an answered rate is meaningful only against ⌠11%, not zero. Citation leaders (25.6%) clear the floor (p=0.005p=0.005); direct LLM (15.2%) does not (p=0.42p=0.42). Persistent: the contamination signature. The direct-LLM baseline keeps the highest similarity to future literature at scale (s1=0.722s_1=0.722 vs. 0.567â0.682 for all others), the broadest engagement (96%, multi-source rate 94%), and the earliest mean engagement (1.9 yearsâits cited evidence concentrates in 2021â2022, the years closest to its training distribution). Scaling revises one part of the v1 reading: it does convert engagement into answers (15.2% vs. 0/10 at n=10n=10), but at a rate statistically indistinguishable from random templates, despite engaging twice as much of the literature. Breadth without resolution remains the signature. Cross-instance anchor: metrics are corpus-relative. Re-judging the ten v1 questions on the 2.4Ă larger corpus flips four labels in both directions: two questions gain answered (richer evidence pools), one drops to not_addressed (its engaging papers pushed out of a top-8 that a larger corpus makes more competitive), and one premise moves weakenedâ still_plausible. q_008âs answered + refuted reproduces. The lesson is structural: rates are functions of the (corpus, retriever, judge) triple, so rows are comparable only within an instanceâv1 and v1L rows must never be ranked against each other, and the frozen-instance design exists precisely to make the triple explicit. What scaling did and did not change. The scaled study strengthens the benchmarkâs discriminative claims (coverage now separates three tiers; answered rates have a measurable floor) and weakens one system-level claim (refutation exclusivity). It does not change the evidence-graph submissionâs standingâits rates are unchanged and its refutation reproducesâbut it narrows what that standing demonstrates: at current sample sizes, the defensible statement is that structured evidence analysis found a refutation by specifying its test in advance, not that it finds refutations at a higher rate than strong heuristics. Settling the rate question requires scaling the submission, not just the baselinesâthe first item on the revised roadmap. 9 Separating Hindsight Memorization from Foresight The deepest objection to any backtest run with a modern LLM anywhere in the loop (Section 12) is that its apparent performance decomposes into three terms: Performance=reasoning over pre-cutoff evidenceâwhat we want+memorized futureâcontamination+topic priorâfashionPerformance\;=\; reasoning over pre-cutoff evidence_what we want\;+\; memorized future_contamination\;+\; topic prior_fashion and a single retrospective instance cannot tell the terms apart. A question that 2024 answered may have been predicted from 2020 evidenceâor remembered from the modelâs training data. This section reports two experiments designed to pry the terms apart: holding the cutoff fixed while varying which component generates the question, and holding the generators fixed while moving the cutoff across the judge-modelâs training boundary. 9.1 Same cutoff, different generators: where does the signal live? Three generation pipelines share the 2020 cutoff and the v1L evaluation conditions but differ in what produces the question: A â LLM only. The direct-LLM baseline: gpt-4.1 reads sampled pre-cutoff abstracts and proposes questions. Weights fully exposed to post-2020 literature. B â structure â LLM. A deterministic, LLM-free detector finds pairs of pre-cutoff abstracts about the same catalogued object with opposing stances on the same species (detection vs. non-detection / upper limit); gpt-4.1âs only job is to verbalize each detected tension as one falsifiable question. The evidence structure is fixed before any LLM sees anything. C â structure only. The same detected pairs rendered by a fixed template. No LLM anywhere: this generator has no weights to contaminate. B and C share identical evidence structures, so their gap isolates the language-realization layer; A and B share the same LLM, so their gap isolates evidence structure. (B/C are an evidence-structure-lite probeâobject co-mention plus stance cuesânot a reimplementation of the evidence-graph system.) Table 7 gives the n=125n=125 results. Table 7: The A/B/C decomposition on astronomy v1L (judge-only, n=125n=125 each; brackets are 95% ClopperâPearson intervals). sobjs_obj = share of questions naming a specific catalogued object (deterministic rubric, Section 9.3). Pipeline Eng. Ans. Ref. s1s_1 sobjs_obj A: LLM only 96% [91,99] 15% [9,23] 0% [0,3] 0.722 5% B: structure â LLM 74% [65,81] 39% [31,48] 13% [8,20] 0.694 95% C: structure only 77% [68,84] 25% [17,33] 2.4% [0,7] 0.721 100% Three readings. First, the contamination-free pipeline C beats the fully exposed pipeline A on answered rate (24.8% vs. 15.2%, p=0.08p=0.08) and on refutations (3 vs. 0)âevidence that a foresight signal exists in pre-cutoff evidence structure alone, extractable with zero model weights. Second, adding the LLM back as a pure verbalizer (B) roughly doubles resolution over the same structure (39.2% vs. 24.8% answered, p=0.02p=0.02âsuggestive only; this contrast does not survive the multiple-comparison correction of Section 12âand 12.8% vs. 2.4% refuted, p=0.003p=0.003, which does): phrasing a tension as a crisp either/or question makes it judgeable, and B ends with five times pipeline Aâs refutation count while citing the same model. Third, Aâs questions are structurally different, not just weaker: only 5% name a specific object (vs. 95â100% for B/C), and its engagement is the highest of any systemâbroad questions that everything touches and little settles. 9.2 Same generators, moving cutoff: the temporal stress test If pipeline Aâs performance were substantially the memorized-future term, it should degrade as the cutoff crosses the modelâs training boundary (June 2024 for gpt-4.1). We ran four generators (A, B, C, and random claims as an era control) at four cutoffsâ2010, 2015, 2020, 2024âwith uniform four-year future windows (n=50n=50 per cell; the 2024 window is censored at 2026-06 and flagged; corpora per era rebuilt from released manifests). Table 8 reports the grid. Table 8: Temporal contamination stress test (judge-only, n=50n=50 per cell; c2010 tension cells have n=48n=48âthe 855-paper 2005â2010 corpus yields only 48 detectable tension pairs). Cutoffs 2010â2020 lie inside the LLMâs training data; the 2024 cutoffâs future window (2025â2026-06, censored) postdates it. Ans. = answered; Ref. = refuted count. c2010 c2015 c2020 c2024 A: LLM only Ans. 22% 12% 10% 16% Ref. 2 0 0 0 B: structure â LLM Ans. 31% 26% 34% 32% Ref. 6 3 6 5 C: structure only Ans. 10% 6% 16% 4% Ref. 0 0 0 0 Random claims Ans. 12% 6% 4% 6% Ref. 1 1 0 0 The naive collapse does not happen. Pipeline Aâs answered rate is statistically flat across the training boundary (14.7% pooled in-training vs. 16.0% post-training, p=0.82p=0.82), as is every other generatorâs. At the outcome level, the memorized-future term is not where Aâs performance comes fromâbecause, as the decomposition shows, Aâs performance never rested on specifics that memorization could supply. Its engagement sits at 92â98% at every cutoff, its specificity at the floor (â 1.1 of 3) at every cutoff: broad questions about each eraâs active topics, engaged everywhere, resolving little, in any era. That is the topic-prior term at work, and a topic prior does not need to remember the futureâthe present is enough. Where a memorization trace does appear. Two places, both in the channels the topic prior cannot supply. Pipeline Aâs only premise refutations in the entire stress test (2 of 200) occur at the deepest-contamination cutoff, 2010âthe one era whose reversals the model has certainly read aboutâand never after (0 of 150; too few for significance, CI [0.5%, 13.7%] at c2010). And Aâs phrasing similarity to future literature is highest inside its training window (0.717â0.724) with a mild post-training dip (0.703)âdirectionally consistent with phrasing-level memorization, though era confounds keep this suggestive rather than conclusive. The structural signal is era-robust. Pipeline B refutes premises at every cutoffâ6, 3, 6, 5âincluding the one whose future the model cannot have seen (10.0% post-training vs. 10.1% in-training). Across all cutoffs B refutes at 10.1% vs. Aâs 1.0% (p=4Ă10â5p=4Ă 10^-5), and Bâs answered margin over C persists post-training (+28+28 points at c2024). One residual channel remains open and is worth stating precisely: c2024 tension pairs are drawn from 2019â2024 abstracts, and a tension resolved in early-2024 literature the model saw could steer Bâs phrasing even though the evaluation window postdates training. C is immune by construction, which is why C finding any future engagement at every cutoff (70â83%) is the cleanest single fact in the grid. 9.3 Specificity-adjusted foresight Coverage can be farmed by asking broad questionsââHow can we better understand exoplanet atmospheres?â will be engaged with probability 1 in any active field. We therefore score every question with a deterministic, released rubric: +1+1 for naming a specific catalogued object, +1+1 for naming a measurable claim (species, quantity with units, abundance comparative), +1+1 for an explicit discriminative construction (ââŠor an artifact ofâŠâ, âas tested byâ); and define specificity-adjusted foresight SAF=âĄ[wâĄ(outcome)â specificity/3]SAF=E[\,w(outcome)·specificity/3\,] with w = 1 / 0.5 / 0.25 / 0 for answered / partial / posed-open / not addressed. (Template pipelines inherit the test-construction point from their templateâthe rubricâs components are reported separately for exactly that reason; a novelty term awaits first-posed dates, Section 14.) Two facts survive every era and instance: pipeline Aâs anchoring is an order of magnitude below the structure pipelinesâ (sobjs_obj 5% vs. 95â100% on v1L), and its SAF never exceeds the random-template floor by more than a few points (0.18â0.23 vs. 0.24â0.26), while structure pipelines reach 0.34â0.43. The evidence-graph submissionâs ten questions score sobj=90%s_obj=90\%, SAF 0.40. 9.4 What this section establishes Memorized relevance is not scientific foresight. The LLM-only pipeline exhibits relevance everywhereânear-ceiling engagement, the closest phrasing to the future literature at every cutoffâand specific foresight nowhere: no refutation outside the era it could have memorized, an answered rate indistinguishable from random templates, specificity at the floor. The structure-first pipelines invert the picture: lower engagement, higher resolution, refutations in every era including the one that postdates the modelâs training. On this evidence, the foresight signal measured by historical backtesting lives chiefly in pre-cutoff evidence structure; the LLM contributes a real but separable serviceâturning a detected tension into a question sharp enough to be judged. Contamination, meanwhile, turns out to be measurable rather than merely confessable: the stress test bounds its outcome-level effect (small for this task family) and localizes its traces (phrasing proximity; refutations only in deep history). All 798 stress-test judgments, the A/B/C submissions, and the per-cell corpora manifests are released; every number in this section is judge-only and carries n=48n=48â125125 intervalsâthe qualitative pattern, not any single rate, is the finding. 10 Judge Validation Replacing expert scores with an LLM judge only helps if the judge is itself accountable. This section reports four validation experiments on a frozen 90-question sample, stratified across systems and outcome labels (results/judge_validation/), plus a seven-rater agreement study built on the released blinded annotation apparatus. The internal checks pass, or fail in ways the data explain; the external one â two independent human annotators against five judge models â does not, and it reframes what an LLM judge can even be validated against (Section 10.4). 10.1 Does the judge discriminate, or merely detect topic? The sharpest failure mode for an engagement metric is that it measures topical similarity and calls it engagement. We test this directly: re-judge every sampled question after swapping in another questionâs retrieved evidence. A discriminating judge must collapse to not_addressed. Engagement falls from 74% (67/90) on true pairings to 20% (18/90) on mismatched ones (p<10â4p<10^-4, two-sided Fisher), and the drop is individually significant for five of six systems. Two further readings matter. The 20% residual is a generosity bound: on evidence that cannot possibly bear on the question, this judge still reports engagement one time in five, so every engagement rate in this paper should be read against that floor rather than against zero. And the drop tracks question specificityâcleanest for the most anchored questions (tension-LLM 12/16 â 1/16, p=0.0002p=0.0002; evidence-graph 8/10 â 1/10) and weakest for the generic citation-leader template (11/16 â 5/16, p=0.076p=0.076, the only non-significant cell). A question vague enough to accept unrelated evidence is vague enough to fool the judge, which is independent support for the specificity rubric of Section 9.3. 10.2 Judge stochasticity and prompt sensitivity Temperature 0 is not determinism. Re-running the identical configuration twice gives outcome agreement 96.7% and 93.3% (Îș=0.95Îș=0.95, 0.910.91) and premise agreement 95.6% and 94.4% (Îș=0.92Îș=0.92, 0.900.90): roughly 3â7 points of label noise, small relative to the effects the paper reports but not zero, and it should be assumed present in every rate. Prompt sensitivity is larger. A semantically equivalent rewrite of the judge prompt agrees at Îș=0.68Îș=0.68 (outcome) and 0.720.72 (premise). A harder variant, replacing the label names with neutral codes (L1âL4, P1âP5) while keeping the definitions verbatim, drops to Îș=0.57Îș=0.57 and 0.550.55, and never once uses the code corresponding to posed_but_open. Part of the judgeâs behaviour therefore rests on the connotations of the label names, not on their stated definitions. We report this as a real limitation: label naming is part of the protocol and must be frozen along with everything else. 10.3 Cross-model agreement, and what low agreement means here Re-judging with gpt-4o and gpt-4-turbo gives low nominal agreement with gpt-4.1: pairwise Cohenâs Îș=0.34Îș=0.34 and 0.200.20 on outcome, 0.190.19 and 0.050.05 on premise status; three-way Fleiss Îș=0.38Îș=0.38 and 0.030.03. Taken alone these numbers say the instrument is unreliable, and we report them unadorned. The label distributions complicate that reading. On 90 items gpt-4-turbo assigns still_plausible 87 times and never once uses refuted or weakened; gpt-4o assigns it 78 times. On the outcome dimension both almost never use answered (2 and 4 times of 90, against 25 for gpt-4.1), collapsing a four-way judgement into a two-way one. Their comparatively high mutual agreement (Îș=0.71Îș=0.71 on outcome) is therefore consistent with a shared conservative default rather than with shared judgement, and Îș is in any case depressed when one raterâs marginals are near-degenerate. We flag plainly that this reading is self-servingââthe judges who disagree with ours are the incompetent onesâ is exactly what a motivated author would sayâand that only human annotation can arbitrate it. Section 10.4 reports that arbitration; it part-vindicates the reading (the human sides with the non-degenerate judge) while overturning the larger assumption that any of the judges tracks human judgement well. What can be settled without humans is whether the paperâs conclusions depend on the judge. Because the stratified sample equalises label mixes across systems and so erases between-system rate differences, this requires a second, unstratified draw (40 random questions from each of four systems); we note the distinction because computing conclusion robustness on a label-stratified sample is a mistake that is easy to make and that we made first. Table 9: Conclusion robustness under three judges, unstratified sample (n=40n=40 per system). â = the stated ordering holds; Ă = it does not. Every Ă is a tie at zero, where the judge assigns the label to no system; no judge reverses any conclusion. Conclusion gpt-4.1 gpt-4o gpt-4-turbo structureâ answers more than LLM-only â â â LLM-only engages more than random templates â â â structureâ refutes more than LLM-only â â Ă structure-only answers more than LLM-only â Ă Ă Table 9 gives the result. The paperâs headline claimâevidence-structure-first generation resolves more than LLM-only promptingâholds under all three judges. So does the engagement ordering. The two conclusions that fail do so in a specific and benign way: gpt-4-turbo reports 0% refutation for every system, and both weaker judges report 0% answered for both compared systems, so the comparison has no resolution rather than the opposite sign. Across all twelve judgeâconclusion cells, no judge ever orders the systems the other way. The practical implication is a requirement, not a reassurance: this benchmark has a judge capability floor. The premise dimension in particular is unmeasurable with models that default to still_plausible, so an instance is only reproducible on a judge that demonstrably uses the full label space. We recommend reporting the judgeâs label distribution alongside any submission, and treating a near-degenerate distribution as a failed run. 10.4 Human annotation, and a seven-rater agreement study The decisive experiment is agreement with human readers. Two annotators labelled all 90 blinded items independently from the abstracts alone: a non-expert (the first author of the submission under test, blinded to system identity and model labels) and a commissioned professional annotation team. We then added two frontier judge models â claude-fable-5 (judged in an agent harness rather than a temperature-0 API call; records are marked accordingly) and gpt-5.6-sol â to the three already run, giving a seven-rater matrix (Table 10). Table 10: Pairwise Cohenâs Îș on outcome, all seven raters, n=90n=90. H1 = non-expert human; H2 = professional annotation team. Read against the humanâhuman cell (0.17): no model reaches agreement with a human that could pass for reliability, and every modelâmodel pair agrees more strongly than any humanâmodel pair. H2 4.1 4o 4-t fable-5 5.6-sol H1 (non-expert) 0.17 0.10 â-0.03 â-0.00 0.02 0.02 H2 (professional) 0.26 0.17 0.20 0.21 0.21 gpt-4.1 0.34 0.20 0.47 0.31 gpt-4o 0.71 0.47 0.58 gpt-4-turbo 0.32 0.51 claude-fable-5 0.60 Four facts, in decreasing order of comfort. 1. Humans do not agree with each other. Humanâhuman agreement is Îș=0.17Îș=0.17 on outcome and 0.170.17 on premise status (41â44% raw). This is the studyâs most consequential number, because it caps everything: no judge, human or model, can be validated against a reference that does not exist. The taxonomy, as specified â even with ordered decision procedures and worked examples â does not produce convergent labels from independent careful readers. The professional teamâs own confidence does not rescue it: on the 31 items they marked highest-confidence, their agreement with the judge is no better (Îș=0.11Îș=0.11). 2. Every model clears the humanâhuman bar with the expert â and none clears it by much. Against the professional team, the five models span Îș=0.17Îș=0.17â0.260.26, with the original gpt-4.1 judge highest (0.26), and the two frontier models at 0.21 despite two additional model generations. In this specific sense the LLM judge is vindicated: it agrees with the expert about as well as another human does, and slightly better. In every other sense it is not: Îș=0.26Îș=0.26 is far below any conventional reliability threshold, and newer, stronger models do not close the gap. One narrower check does lean the deployed judgeâs way: on the 42 items where gpt-4.1 and gpt-4o disagree, both humans side with gpt-4.1 more often (19â9 for the non-expert, uncorrected p=0.015p=0.015 and suggestive only; 18â13 for the professional team, not significant), consistent with the degeneracy reading above. 3. Models agree with each other far more than with any human. Modelâmodel agreement runs 0.20â0.71, with the two frontier models â different vendors, different harnesses â at Îș=0.60Îș=0.60 (67/90 identical labels), triple the humanâhuman figure. Some of the high modelâmodel cells are degeneracy artifacts (gpt-4o/gpt-4-turbo at 0.71 share a two-label collapse), but fable-5 and gpt-5.6-sol both use the full label space and still converge. The models constitute an internal consensus that correlates only weakly with either human reader. For LLM-as-judge practice generally, this is the sharpest caution in the paper: measuring judge reliability by modelâmodel agreement â the cheap and common method â would have reported Îșâ0.6Îșâ 0.6 here, three times what validation against humans supports. 4. The non-expert is the outlier, informatively. The first annotator agrees with nobody (Îșâ€0.17Îș†0.17 with every other rater), labelling far more items answered (33) and far fewer not_addressed (9) than the professional team (33/24) or any model. The expertâs marginal distribution closely tracks the strict judgesâ. This ordering â expert closest to models, non-expert loosest â suggests the disagreement is partly about how much domain scepticism a reader brings to âsubstantially resolved,â which is a calibration norm the codebook failed to pin down, not a fact about either raterâs diligence. What this settles. Of the three readings left open after the first pass, the evidence now favours the third: the taxonomy is underdetermined. The judge is not distinguishably worse than a human rater â it sits at the top of the observed agreement range with the expert â but nothing, human or model, converges on these labels reliably. Three consequences follow for this benchmark and for the genre. Absolute rates (any systemâs âanswered 39%â) are rater-relative and should never be quoted without the rater attached. Comparative claims measured under a fixed judge remain defensible â Section 10 showed the paperâs headline orderings survive three judges with all failures being ties â and they are the only currency this instrument currently supports. And the v2 protocol must redesign the outcome taxonomy itself: fewer labels, hard decision criteria phrased as checkable conditions, and a measured humanâhuman Îș as a release gate before any judge, human or model, is scored against it. We release all seven label sets, the annotation apparatus, and the agreement matrix; the professional teamâs labels were commissioned for a fixed fee with a published undertaking that no entry would be adjusted to improve agreement with any model, and none was. 11 Error Analysis The released curation log records every human intervention made before freezing; the failure modes below are taken from it and from the adjudication log, and each motivates a benchmark rule. Question duplication. The generator produced near-duplicate questions (q_002/q_003) from the same claim pair, differing mainly in emphasis; they were merged during curation into one question with two sub-questions. Left unmerged, duplicates would double-count a single insight in every rate. Rule: submissions are screened for near-duplicates, and instances should report a deduplication note; a mechanical similarity screen is a v1.1 roadmap item. Leading questions (presupposition). Several generated questions presupposed their own answer. q_001 originally asserted that vertical wind shear exists rather than asking what causes the wind-speed discrepancy; q_008 originally presupposed the subsolar water abundance rather than framing artifact-vs-atmosphere attribution. A question that presupposes its answer cannot be cleanly refutedâand q_008âs later refutation was only expressible because curation reframed it as attribution. Rule: the quality gate flags presupposing phrasings before freezing; the reframing is logged. Self-contradictory quantifiers. q_003âs original phrasing asked whether a detection was ârobust to biases exceeding an order of magnitudeââa bias that large is non-robustness. Templated quantifier language can silently produce unanswerable questions; the clarity gate exists for this. Conflated statistical and physical framing. q_010 originally asked what âtemperature and pressure conditionsâ produce a 5.4Ï water detection, conflating atmospheric state with the detection pipeline (significance depends on noise model, priors, null hypothesisânot on the atmosphere). It was reframed as a sensitivity analysis; its score on the generatorâs internal 0â10 clarity gate (6.0) was the lowest of the set, and its outcome (partially_addressed on one supporting paper) remains among the weakest-evidenced labels. Premise bias in generation. The generator inherits the premises of the papers it reads: single-dataset conclusions taken at face value can produce questions that merely restate a claim rather than test it. The tension-typing stage (which explicitly marks single-dataset conclusion as a signal type) partially controls thisâq_008 and q_005 are that control workingâbut the benchmarkâs premise dimension is the systematic check: a healthy portfolio should show a mix of supported and refuted, not uniform support of its sources. Retrieval near-misses. Judged evidence for q_007 (HD 189733 b / HAT-P-11 b patchy clouds) leans partly on a three-retrieval-framework study of HAT-P-18 bâdirectly relevant methodologically, but not the named targets. The adjudication log records the engagement-bar judgment; releasing full top-8 lists (v1.1) will let others re-litigate such calls, which is the point of releasing them. 12 Threats to Validity Historical backtesting removes rater subjectivity; it does not remove every confound. We enumerate the serious ones and what the protocol doesâand cannot doâabout each. Future inattention is not question badness. A question can be ignored because the enabling instrument never flew, the communityâs funding shifted, or the subfield is smallânot because the question was poor. q_011 was unanswerable before JWST delivered TRAPPIST-1 spectra; had JWST slipped five years, an excellent question would have scored not_addressed. Mitigations: bounded windows make the censoring explicit; posed_but_open separates ârecognized but unresolvedâ from âignoredâ; instances should be read as engagement within k years given the eraâs instruments, not as timeless value. Future attention is not question goodness. Symmetrically, a question on a fashionable topic collects engagement for reasons other than merit; coverage alone can be gamed by asking about whatever is popular. Mitigations: coverage is never reported alone; premise refutation and (future) lead time cannot be earned by fashion-chasing; baseline B4 (citation leaders) exists precisely to price in prominence; and the engagement bar requires the cited paper to bear on the questionâs actual test, not its topic. LLM training contamination. The generating system and the judge both use models whose training data postdates the cutoff. The protocol seals the retrieval channel, not the weights channel: a model may âknowâ the 2025 refutation while drafting a 2020-framed question. This is the deepest threat to any backtest run with modern models. Mitigations, none complete: source-evidence audit forces every question to be grounded in cited pre-cutoff evidence; submission metadata must declare all LLM components and versions; baseline B2 doubles as a contamination probeâand registers one: its questions sit measurably closer to the future literatureâs phrasing than any other systemâs (s1=0.732s_1=0.732 at n=10n=10, 0.722 at n=125n=125; Tables 4 and 6), with its cited engagement concentrated in the years nearest its training distribution, while its answered rate stays indistinguishable from random templatesâthe signature of phrasing-level memorization without specific foresight. Section 9 upgrades this confession to a measurement: a generator decomposition plus a four-cutoff stress test bound the contaminationâs outcome-level effect and localize its traces; and the decisive test is prospective instancesâquestions frozen today and scored in 2030 cannot be contaminated. The protocol is explicitly designed so its instances convert from retrospective to prospective by just letting time pass. Judge and adjudicator reliability. A single LLM judge, even citation-constrained, is one reading of the evidence, and abstracts (not full texts) bound what it can see. Section 10 measures rather than asserts what this costs: the judge discriminates real evidence from topical similarity (engagement 74% â 20% under mismatched evidence) but with a 20% generosity floor; it is stable under resampling (Îșâ0.9Îșâ 0.9) and moderately sensitive to prompt wording (Îș=0.57Îș=0.57â0.720.72); and every headline conclusion survives a judge swap, with all failures being ties at zero rather than reversals. The most serious question is now answered, and the answer indicts the taxonomy rather than the judge: two independent human annotators agree with each other at Îș=0.17Îș=0.17, every judge model agrees with the professional annotator at Îș=0.17Îș=0.17â0.260.26 (the deployed judge highest), and frontier models agree with each other at up to Îș=0.60Îș=0.60âfar above their agreement with any human (Section 10.4). Absolute rates are therefore rater-relative throughout this paper; only comparisons under a fixed judge carry weight, and the v2 protocol owes the field a taxonomy with a measured humanâhuman Îș before any judge is scored against it. Blinded adjudication is now built into the released annotation apparatus rather than promised. Retrieval as a bottleneck. Top-8 abstract-embedding retrieval can miss engaging papers (undercounting engagement) or surface topically similar non-engagement (which the judge must reject). Fixed retrieval is a deliberate trade: it makes comparisons across systems fair and auditable at the cost of an engagement floor. Sensitivity of labels to k and to the retriever is measurable within the released data schema and belongs in v1.1. Small n where it matters most, one domain, self-evaluation. The scaled instance lifts the baseline side to n=125n=125 per system, and Section 8 shows how much that matters: two of three small-sample conclusions did not survive. The submission under test, however, still has ten questions, and we can now say exactly what that costs. At the observed effect size (10% vs. 3.2% premise refutation), detecting the difference at 80% power and α=0.05α=0.05 requires nâ209nâ 209 per arm; the ten-question instance has a power of roughly 12%. The comparison is therefore not merely unresolved, it was never resolvable at this sample size, and nâ„209nâ„ 209 is the design specification we adopt for the next instance rather than an aspiration. That specification collides with a structural fact worth reporting, because it constrains anyone building a high-precision question generator. The evidence-graph systemâs strongest signals (confirmed observational tensions, method challenges, qualifications between independent datasets) are gated on human-reviewed claim relations: its released instance rests on 37 annotated claims and 16 reviewed relation edges. Its question supply is bounded by annotation labour, not by compute or API budgetâwhich is precisely why it ships ten questions while the automatic baselines ship 125 each. The benchmark thus measures a real precision/scale trade-off rather than mere effort: automatic generators reach n easily and mostly produce questions no one can settle, while the human-gated generator produces few questions with high anchoring (90% naming a specific object). Closing the gap requires either scaled annotation or an automatic tension detector of comparable precisionâthe latter is what Section 9âs B5/B6 probe begins, at visibly lower precision. One domain, evaluated by the group that built the leading system, judge-only baseline labels: all still true. We report priced reference points, not rankings; the protocolâs value grows with adversarial use by systems we did not build. Multiple comparisons. This paper reports roughly twenty significance tests. Under a conservative Bonferroni correction at that count, the headline contrasts survive comfortably: the structure-vs-LLM refutation gap (p=4Ă10â5p=4Ă 10^-5), the engagement orderings at scale (p<10â4p<10^-4), the mismatched-evidence control (p<10â4p<10^-4), and the B/C refutation gap (p=0.003p=0.003). Contrasts reported at pâ0.01pâ 0.01â0.050.05 (the B/C answered gap, the citation-leader engagement gap, the humansâ arbitration splits) do not, and are labelled suggestive where they appear. No headline claim rests on a contrast that fails correction. Curation hindsight. Humans who edited questions before freezing know the post-2020 literature. The curation log is released so every edit is auditable (e.g. q_008âs reframing strengthened falsifiability without smuggling in the answer), but retrospective instances cannot fully exclude this channel; prospective ones can. 13 Outlook: Discovery as Search, Language as Realization A scope statement first. Nothing in this paper shows that large language models cannot originate scientific questions in principle; a single prompting strategy against a single model family in a single domain cannot support that claim, and we do not make it. What the data do support is narrower and more useful: bare prompting does not reliably perform problem discovery, and the components of the systems that do perform better can be named. Three capabilities the experiments separate. Read together, the decomposition (Section 9.1), the anchoring rubric (Section 9.3), and the temporal stress test (Section 9.2) distinguish three things that âasking good questionsâ conflates. First, a topic prior: knowing what a field is likely to work on next. This is what LLM-only generation exhibitsânear-ceiling engagement in every era (92â98%), questions that name object classes rather than objects (5% anchoring against 95â100% for the structural pipelines), and specificity-adjusted foresight within a few points of the random-template floor. A topic prior is genuinely predictive of where attention flows, and genuinely cheap: it requires no memory of the future, which is why it survives the training boundary unchanged. Second, structural problem discovery: locating specific, contestable configurations of existing evidenceâa claim, a counter-claim, a method dependency, an untested premise. The weight-free pipeline C is the clean witness that this capability does not reside in model weights: with no language model anywhere, it finds questions the future substantively engages at every cutoff, and it beats LLM-only prompting on resolution. Third, articulation: turning a detected configuration into a question a scientist would recognize as askable. This is where the LLM earns its placeâover identical evidence structures, LLM verbalization roughly doubles resolution and quintuples refutations relative to a fixed templateâand it is a capability the stress test shows to be era-robust rather than memorized. The architecture this implies. These results point away from âmake the model smarter and ask it for a hundred ideasâ and toward a division of labour: machine-scale structural search over the evidence space â\;â\; candidate scientific tensions â\;â\; LLM articulation â\;â\; testable questions. The asymmetry that motivates it is quantitative. A literature of thousands of papers yields tens of thousands of claims and a combinatorially larger space of claim pairs, evidence paths, and method dependenciesâfar beyond what any single reader, human or prompted model, holds in attention at once, but squarely within what a machine can sweep in parallel. Under this framing the LLM is never asked to conjure novelty from nothing; it is asked to do what it demonstrably does wellâlocal semantic understanding, claim extraction, relation judgment, and finally phrasingâwhile the search system carries the burden of combinatorial exploration. Scientific question discovery becomes a computable search problem over an explicit representation of what the literature claims, and the interesting engineering question shifts from prompting to representation and search: what to index, which configurations to enumerate, and how to rank them. The next measurable question. Ranking is where this benchmark and that architecture meet. Every question in the released instances carries its generating signalâtension type, object, method-dependency, source claimsâand its measured fate. That pairing makes a new question answerable: which structural patterns most often lead to questions the future answers, advances, or refutes? Our own probe is deliberately crudeâobject co-mention plus stance cues, visibly below the human-gated graph in precisionâso the headroom is real: contradiction typing, dataset-dependency detection, and archival-data availability are all candidate features for a learned prior over the tension space. We flag the honest status of all of this: an interpretation consistent with our data, not an established causal accountâand the frozen prospective instance (Section 14) is the experiment that will test it without any of this paperâs retrospective caveats. Learning that prior from backtested outcomes, and validating it prospectively, is the natural next paper. 14 Roadmap and Conclusion A prospective instance, frozen now. The one experiment no retrospective design can deliver is the one this release starts: 200 questions â 50 each from the four automatic generators of Sections 6 and 9 â generated from a 2015â2026 corpus, frozen at cutoff 2026-08-17, and committed to the public repository with per-file SHA-256 digests (combined digest f1c61a5107e55e7a7e47e2bcab35ae35e484270d892c98ec9c63f18f2c17994d over the per-file list in submissions/prospective_2026/SHA256SUMS). The generation corpus is 9,133 papers, 2015 through freeze day. The scoring window is pre-registered as 2027-01-01 to 2030-12-31, with a 4.5-month buffer between cutoff and window so papers already in flight at freeze time do not contaminate the future corpus, whose ADS query manifest is likewise frozen now and may be executed no earlier than 2031. Evaluation is pre-registered to the released pipeline (frozen retrieval settings, judge protocol sqb-v1, judge label distributions reported alongside results), with one explicitly permitted amendment: if a reliability-gated v2 taxonomy exists before the window closes, results are to be reported under both taxonomies. No model that exists today has seen 2027; whatever these 200 questions score in 2031 is foresight or its absence, untouched by memorization, hindsight, or curation. Roadmap. Beyond waiting: (1) a reliability-gated v2 outcome taxonomyâfewer labels, checkable conditions, and a measured humanâhuman Îș as a release gate before any judge is scored against it (Section 10); (2) scale the submission sideâa 100+ question evidence-graph run on v1Lâto resolve the refutation-rate comparison that n=10n=10 leaves open; (3) release full top-8 retrieval lists for v1 (v1.1; the v1L records already ship with full scores); (4) annotate community first-posed dates to activate lead time, and calibrate a community-attention index against the popularity confound; (5) mint instances in additional domains. Conclusion. We formalized historical backtesting as an evaluation protocol for scientific question discovery, released two retrospective astronomy instances and one prospective one, with temporally isolated corpora and fully auditable labels, and ran a ten-question pilot in which every frozen question was substantively engaged by literature the generating system never sawâincluding one whose premise the community subsequently refuted, the exact convergence test the question had specified. Scaling the baselines to 424 questions then did what a benchmark is supposed to do: it overturned two of our own small-sample conclusions, put a measurable floor under a third, and left the qualitative distinctionâa refutation reached by specifying its test in advanceâstanding but explicitly unresolved at current sample sizes. Finally, the generator decomposition and temporal stress test turned the benchmarkâs deepest limitation into its sharpest result: memorized relevance is not scientific foresight, and the foresight signal that historical backtesting measures survives in a generator with no weights at all. The individual numbers matter less than the category they inhabit: for the first time, âthis system asks good scientific questionsâ is a claim with a denominatorâfalsifiable, comparable across systems, and computable by anyone from frozen public data. Question-asking has been argued to be a core capability on the path to more general scientific intelligence (5; 10); if that is so, the field will need to measure it. This protocol, and this first instance, are offered as the place to startânot as the definition of the benchmark, but as its initial version, built to be superseded by instances with more questions, more systems, more domains, and cutoffs whose futures have not yet happened. References Baek et al. (2024) J. Baek, S. K. Jauhar, S. Cucerzan, and S. J. Hwang ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. Note: arXiv:2404.07738 Cited by: §1, §2. Bailey et al. (2014) D. H. Bailey, J. M. Borwein, M. LĂłpez de Prado, and Q. J. Zhu Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the American Mathematical Society 61 (5), p. 458â471. Cited by: §1, §2, §3.3. Jimenez et al. (2024) C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. In International Conference on Learning Representations, Cited by: §2. King et al. (2009) R. D. King, J. Rowland, S. G. Oliver, M. Young, W. Aubrey, E. Byrne, M. Liakata, M. Markham, P. Pir, L. N. Soldatova, A. Sparkes, K. E. Whelan, and A. Clare The Automation of Science. Science 324 (5923), p. 85â89. Cited by: §2, §2. Kitano (2021) H. Kitano Nobel Turing Challenge: creating the engine for scientific discovery. npj Systems Biology and Applications 7, p. 29. Cited by: §14, §2. Langley et al. (1987) P. Langley, H. A. Simon, G. L. Bradshaw, and J. M. Zytkow Scientific Discovery: Computational Explorations of the Creative Processes. MIT Press. Cited by: §2. Lu et al. (2024) C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. Note: arXiv:2408.06292 Cited by: §1, §2. Lustig-Yaeger et al. (2019) J. Lustig-Yaeger, V. S. Meadows, and A. P. Lincowski The Detectability and Characterization of the TRAPPIST-1 Exoplanet Atmospheres with JWST. The Astronomical Journal. Note: ADS bibcode: 2019AJâŠ.158âŠ27L External Links: Document Cited by: §7.2. MacDonald and Madhusudhan (2017) R. J. MacDonald and N. Madhusudhan HD 209458b in new light: evidence of nitrogen chemistry, patchy clouds and sub-solar water. Monthly Notices of the Royal Astronomical Society. Note: ADS bibcode: 2017MNRAS.469.1979M External Links: Document Cited by: §7.1, §7.2, §8. Scientific Question Discovery project (2026) Scientific Question Discovery project Scientific Question Discovery: Toward Question-Asking as a Core Capability of AGI. Note: Manuscript and code: https://github.com/nonameisready/scientific-question-discovery Cited by: §14, §5, Table 4. Si et al. (2024) C. Si, D. Yang, and T. Hashimoto Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. Note: arXiv:2409.04109 Cited by: §1, §2. Swanson (1986) D. R. Swanson Fish Oil, Raynaudâs Syndrome, and Undiscovered Public Knowledge. Perspectives in Biology and Medicine 30 (1), p. 7â18. Cited by: §2, §2. Wang et al. (2024) Q. Wang, D. Downey, H. Ji, and T. Hope SciMON: Scientific Inspiration Machines Optimized for Novelty. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Cited by: §1, §2. Zou et al. (2022) A. Zou, T. Xiao, R. Jia, J. Kwon, M. Mazeika, R. Li, D. Song, J. Steinhardt, O. Evans, and D. Hendrycks Forecasting Future World Events with Neural Networks. In Advances in Neural Information Processing Systems, Cited by: §1, §2.