Paper deep dive
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
Fabricio F Costa
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/18/2026, 4:50:56 AM
Summary
This paper audits the public measurement record for frontier AI forecasting, identifying three major bottlenecks: sparse joint observability of training compute and capability metrics, benchmark-linking uncertainty due to version changes, and high provenance concentration. The author constructs an event-centric graph of 62 systems and 144 events to demonstrate that defensible forecasting requires explicit measurement system claims rather than simple trend fitting.
Entities (7)
Relation Signals (5)
Fabricio F. Costa â authored â Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
confidence 99% ¡ Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence Fabricio F. Costa
METR Time Horizon 1.0 â linkedto â METR Time Horizon 1.1
confidence 95% ¡ a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206
MMLU â linkedto â MMLU-Pro
confidence 95% ¡ a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links
METR â supplieddatafor â METR Time Horizon 1.1
confidence 90% ¡ 52 of 71 substantive quantitative events... come from one measurement programme
Epoch AI â provideddatafor â Training Compute
confidence 85% ¡ Training-compute values were frozen from Epoch AI
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.
Tags
Links
- Source: https://arxiv.org/abs/2608.14903v1
- Canonical: https://arxiv.org/abs/2608.14903v1
Trouble viewing inline? Open PDF directly â
Full Text
53,555 characters extracted from source content.
Expand or collapse full text
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence Fabricio F. Costa, PhD, MBA, PMP 1,2,3,4â 1 AIx4All, LLC, Sunnyvale, CA, USA 2 HCLTech, Santa Clara, CA, USA 3 Genomic Sciences and Biotechnology Program, UCB, Bras ĚÄąlia, Brazil 4 Cancer Biology and Epigenomics Program, Stanley Manne Childrenâs Research Institute, Ann & Robert H. Lurie Childrenâs Hospital of Chicago, Northwestern University Feinberg School of Medicine, Chicago, IL, USA Evidence cutoff: 12 August 2026. Preprint, comments welcome. Abstract Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50% task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95% CI 1.021â1.390), whereas a six- system MMLUâMMLU-Pro comparison appears approximately shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80% power only for slope departures near 25%. Provenance is concentrated: 52 of 71 sub- stantive quantitative events (73.2%) come from one measurement programme, and 76.1% are laboratory releases. A review of 56 additional methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier-AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measure- ment systemâits joins, protocols, links and source dependenceânot merely a fitted curve or calendar date. Keywords: AI evaluation; measurement; benchmark linking; forecast validation; training compute; provenance; metascience 1 Introduction Forecasts of advanced AI now appear as expert probability distributions, benchmark extrapola- tions, compute trends and backtested models of agent performance [18, 19, 45]. The quantitative form is attractive: select an indicator, estimate its rate of change and solve for a date at which â Correspondence: fcosta@aix4all.com. This independent scholarly work uses only public sources, received no external funding, and represents the authorâs views rather than those of the listed affiliations. The author declares no financial interest in organisations whose systems or forecasts are discussed. 1 arXiv:2608.14903v1 [cs.AI] 14 Aug 2026 a threshold is crossed. Yet the date is meaningful only if the target is operational, the obser- vations are comparable, the progression variable is jointly observed and the uncertainty model includes changes in the measuring instrument. This paper audits those preconditions. It began as an attempt to construct an event-based forecast of frontier-AI capability milestones. The supporting infrastructure could be built; the intended progression analysis could not be identified without strong additional assumptions. The failure was not a software failure. It was a measurement finding: capability and resource records were sparsely joined, benchmark versions did not always preserve scale, evaluation protocols changed, and repeated observations were heavily dependent on a small number of sources. The distinction matters. A criterion can be well defined while its public evidence is not forecastable. Conversely, a record can contain many scores yet fail to measure one stable quantity. Existing work has proposed taxonomies of performance, generality and autonomy [40], holistic evaluation suites [31], ability-oriented measurement [22], explicit evaluation estimands [6] and statistical models that distinguish fixed-benchmark accuracy from generalized accuracy [26]. The contribution here is to connect those ideas to the actual longitudinal record used in frontier-AI forecasting. The paper makes three contributions. First, it provides a reproducible event-centric audit in which releases, results, revisions, corrections, contamination findings, field experiments and source assertions are separately timestamped and versioned. Second, it quantifies three fore- castability bottlenecks in that record: sparse joint observability, benchmark-linking uncertainty and provenance concentration. Third, it translates a deep source review into a measurement portfolio and a set of study-design requirements rather than proposing another universal leader- board score. The semantic graph is therefore infrastructure, not the main scientific claim. Its role is to make questions such as âwhich systems connect compute to capability?â, âwhich result changed after a benchmark revision?â and âwhich conclusions depend on one programme?â answerable from the same evidence ledger. 2 From a Score to a Forecastable Quantity A benchmark result is not an intrinsic property of a model. It is an observation generated by a model, an instrument and a protocol at a time. A compact representation is Y m,b,v,p,r,j,t = h b,v (θ m,t ;p,r,j) + Îľ m,b,v,p,r,j,t ,(1) where m is the model, b the benchmark family, v its version, p the prompting or agent protocol, r the inference-resource budget, j the judge or scoring procedure and t the observation date. The latent quantity θ may itself be multidimensional. Equation 1 is not fitted as a complete latent-variable model here; it is an accounting identity that makes the assumptions behind a trend explicit. A dated forecast adds at least three more objects: a target estimand, a progression axis and a threshold rule. If training compute is the progression axis, compute and the outcome must be co-observed or missingness must be modelled. If benchmark version v is replaced by v + 1, a linking function must place the two instruments on a defensible common scale. If inference budget or scaffolding changes, the score may move even when the underlying system does not. If many observations share a source, their errors need not be independent. This accounting view is increasingly reflected in measurement science. Statistical evaluation can distinguish performance on one fixed item set from a generalized-accuracy estimand over item and trial populations [26]. Psychometric aggregation can place models and items on a latent scale when the response design and anchors support it; CAISI, for example, has combined item-response modelling with fixed scaffolding, weighted-token budgets and cost measurement 2 [10]. Adaptive testing and fixed-parameter calibration can reduce evaluation cost [21, 30], but simulation evidence also shows that small, clustered or non-normal model samples can make item and ranking inferences unstable [25]. The method is therefore not âuse IRTâ or âuse a better benchmark.â It is to state the estimand, design the linking observations and preserve the protocol that makes the scale interpretable. The audit therefore asks five questions: 1. Is the target construct and intended population explicit? 2. Are the proposed predictor and outcome jointly observed on enough systems? 3. Are benchmark versions linked rather than concatenated? 4. Are protocol and resource choices part of the measurement record? 5. Are provenance, replication and source dependence represented in uncertainty? These are deliberately stricter than asking whether a regression line can be drawn. 3 Method and Data Audit 3.1 Evidence freeze and sampling contract The evidence cutoff is 12 August 2026. The analytic set contains 62 selected systems: 27 closed and 35 open-weight. Inclusion was driven by the records needed for the attempted computeâ capability analysis and by historically or contemporaneously important systems with public release metadata. It is not a complete catalogue of releases. A separate cutoff audit records four known pre-cutoff systems absent from the analytic set so that the package cannot be mistaken for a census. Training-compute values were frozen from Epoch AI and the processed Our World in Data series, with release reports used where appropriate [15, 16, 46]. METR Time Horizon 1.1 observations and uncertainty intervals were frozen locally for reproducibility [28, 37]. Benchmark links, lifecycle events, field experiments and provenance records were checked against primary or official sources where available. 3.2 Event-centric representation The normalized corpus contains 12 benchmark-version records, seven criterion records, 27 source records and 144 events. The event table includes 62 model releases, 62 benchmark results and 20 methodology, correction, contamination, efficiency, field, elicitation, substrate, replication, instability or measurement-ceiling events. Every substantive quantitative event has a source identifier and an evidence grade. The graph contains 269 nodes and 408 typed relations. Its vocabulary is informed by event and provenance standards [29, 54]; it does not claim that semantic representation resolves statistical identification. 3.3 Empirical tests Joint observability. For each model I record whether training compute, METR p50, METR p80, any capability result, a non-METR capability result and a peer-reviewed measurement are available. Pairwise overlap counts reveal whether candidate variables can enter the same model-level analysis. Benchmark linking. Two small common-system bridges are used as stress tests. METR Time Horizon 1.0 and 1.1 are linked on a base-2 logarithmic scale. MMLU and MMLU-Pro are 3 Method: from public records to a forecastability audit The event graph is supporting infrastructure; the scientific output is an audit of joins, scale continuity, protocol dependence and provenance. 1 Freeze scope cutoff, inclusion rule, raw snapshots 2 Normalize records models, versions, metrics, configurations 3 Eventize evidence results, changes, failures, provenance 4 Audit the links co-observation, lifecycle, source independence 5 Test and redesign linking, power, portfolio, forecast backtests Event topology used by the audit Model Configuration and resources Event Benchmark version Criterion Source has participates in evaluated on bears on derived from operationalizes Audit outputs joint-observation topology | benchmark-link diagnostics | lifecycle and source dependence | alternative measurement designs Forecast only after the target, observable, protocol, version link and uncertainty budget are explicit. Figure 1: Method used in the audit. Public evidence is frozen under an explicit cutoff and sampling contract, normalized into versioned entities, represented as events with provenance, tested for missing joins and instrument changes, and then evaluated for forecastability. The lower topology shows the minimum graph structure used to connect models, events, benchmark versions, criteria and sources. compared under logit, probit, linear and logarithmic representations. In a simple link g(Y new ) = Îą + β g(Y old ) + Îľ,(2) β = 1 represents a shift-like relationship on the chosen scale; β ̸= 1 indicates shape change on that scale. Ordinary least squares is used as a transparent diagnostic, with leave-one-out sensitivity. It is not a full errors-in-variables or construct-invariance model. Power. The minimum detectable departure|βâ1| is calculated for a two-sided test at Îą = 0.05 and 80% power using the noncentral t distribution. Projected sample sizes hold residual noise and anchor-score dispersion at their observed values. They are design illustrations, not universal rules; anchor placement is part of the power calculation [8, 27]. Source and literature audit. Administrative release rows are excluded from provenance- concentration statistics. âSubstantive quantitative eventsâ comprise benchmark results, field experiments, substrate estimates and efficiency observations. In parallel, 56 external sourcesâ 29 peer-reviewed and 27 preprints, standards reports, official releases or living datasetsâwere coded into 16 measurement directions and normalized into a machine-readable directionâsource map. Nine recurring evidence-design requirements are documented separately from the empir- ical event graph. Scores in Figure 5 range from zero to three and summarize design maturity; they are not empirical capability estimates or a formal systematic-review quality score. Figure 1 summarizes the full workflow. 4 Result 1: The Intended Measurement Join Is Sparse The public record is block-structured rather than merely incomplete (Figure 2). Training com- pute is available for 43 of 62 systems, but the missingness is concentrated: 19 of 27 closed 4 systems lack a compute estimate, while all 35 open-weight systems in the sample have one. METR p50 and p80 measurements exist for 26 systems, all in the closed block. Consequently, only seven systems jointly observe training compute and METR p50. No open-weight system in the sample supplies the intended computeâhorizon join. Compute only 36 Horizon only 19 Both7 Neither0 ComputeHorizon (a) The intended join exists for only 7 of 62 systems Training computeHorizon Other capability Peer-reviewed measurement Closed Open-weight 8/27 30% 26/27 96% 3/27 11% 1/27 4% 35/35 100% 0/35 0% 0/35 0% 0/35 0% (b) Missingness follows a closed/open-weight split 10 21 10 22 10 23 10 24 10 25 10 26 Estimated training compute (FLOP) 10 â2 10 â1 10 0 10 1 10 2 METR p50 task horizon (minutes) GPT-2 GPT-3 GPT-3.5 GPT-4 Claude 3.5 S. Claude 3.7 S. GPT-5 Descriptive only: the sample is not a census, compute is estimated, and no recent open-weight system carries a METR horizon. (c) The apparent relationship is identified by seven selected systems descriptive log-log slope 0.73 Figure 2: Observability of the selected 62-system record. Panel (a) shows the intended computeâhorizon join as four mutually exclusive observation patterns: compute only, horizon only, both and neither. Panel (b) shows that missingness follows the closed/open-weight split and extends to non-METR and peer-reviewed measurements. Panel (c) displays the seven jointly observed systems; the fitted line is descriptive and is not interpreted causally or as a stable scaling law. This pattern is more consequential than a low marginal coverage rate. A compute-indexed capability model requires the intersection, not the union, of the two records. The seven systems are also historically clustered and not a designed sample across developers, architectures or access regimes. A descriptive logâlog slope can be fitted, but its apparent smoothness does not identify whether compute, algorithms, data, inference effort, model family or evaluation protocol caused the change. The missingness is unlikely to be ignorable. Public compute disclosure is related to access regime, developer practice and time. Treating absent compute as random would therefore turn an institutional disclosure pattern into a statistical assumption. Controlled training suites such as Pythia demonstrate the value of dense checkpoints and fixed data order for identifying training dynamics [5]; frontier-release records do not offer the same design. 5 Result 2: Changing Instruments, Concentrated Provenance Benchmark results are versioned events, not timeless labels. The audit records methodology changes, corrections, contamination findings, reward-hacking findings, measurement ceilings and replacements across METR, ARC-AGI, SWE-bench, MMLU and other instruments (Figure 3). Dynamic and contamination-resistant benchmarks are promising responses [57, 62], but refreshes and repairs still need explicit linking if the objective is longitudinal inference. Recent work on 5 saturation and verified benchmark revision likewise shows that item lifecycle is part of the measurement process rather than a footnote [1, 44, 61]. 20192020202120222023202420252026 Calendar year LifeSciBench 1 CHC AGI Score 1 Humanity's Last Exam 1 MMLU 1 SWE-bench Pro SWE-bench Verified ARC-AGI 3 ARC-AGI 2 ARC-AGI 1 METR Time Horizon 80pct 1.1 METR Time Horizon 1.1 METR Time Horizon 1.0 cutoff (a) A benchmark is a versioned instrument with a lifecycle MC ! L 52/71 (73.2%) from METR (b) Measurement provenance is concentrated METR TH1.1: 52 Other lab releases: 2 Preprints: 10 Peer-reviewed: 6 Vendor: 1 M/C/!/L mark lifecycle events. Each square is one of 71 quantitative events; administrative release records are excluded. Figure 3: Benchmark lifecycle and source dependence. Panel (a) places benchmark versions and recorded methodology, correction, contamination and measurement-ceiling events on a common timeline. Panel (b) combines a Pareto view of the 71 substantive quantitative events, their cumulative source share and their venue-class composition. METR supplies 52 events (73.2%); the source Herfindahl index is 0.542, or 1.85 effective equally represented sources. Repeated rows from one programme are not independent replication. The provenance distribution is similarly uneven. The 71 substantive events consist of 62 benchmark results, four field experiments, four substrate estimates and one efficiency observa- tion. Fifty-two events (73.2%) come from the METR Time Horizon 1.1 programme. By venue class, 54 events (76.1%) are laboratory releases, ten (14.1%) preprints, six (8.5%) peer-reviewed publications and one (1.4%) a vendor source. The corresponding source Herfindahl index is 0.542, equivalent to only 1.85 equally represented sources, even though 16 source identifiers appear at least once. Protocol history can matter even without an explicit benchmark revision. Repeated behav- ioral measurements have been shown to change with item order, reasoning mode, persona and conversation history [53]; a fixed candidate-response set can also receive different scores when the LLM judge is replaced [60]. These are not side effects to be averaged away after the fact. They are versioned measurement events that belong in the evidence record. These percentages describe this audit, not the full evaluation literature. They neverthe- less bound what can be claimed from this record. Fifty-two measurements generated by one task collection and adjudication process are valuable observations, but they do not provide 52 independent replications. Shared tasks, protocols and model-fitting choices can induce corre- lated error. The concentration is not a criticism of METR; its documentation of limitations, task changes and failure modes makes the analysis possible [35, 36]. It is evidence that the ecosystem lacks overlapping longitudinal programmes with which to estimate programme-specific bias. 6 Result 3: Benchmark Links Depend on Scale and Design Figure 4 separates a within-family revision from a cross-benchmark stress test. For seven systems measured on METR Time Horizon 1.0 and 1.1, the fitted log 2 slope is 1.206 (95% CI 1.021â1.390). Leave-one-out slopes range from 1.126 to 1.275. In this bridge, the new version 6 is therefore not represented well by a pure additive shift on the log scale. 10 1 10 2 TH1.0 horizon (min) 10 1 10 2 TH1.1 horizon (min) log 2 slope 1.206 95% CI [1.021, 1.390] (a) Within-family revision 60657075808590 MMLU accuracy (%) 30 40 50 60 70 80 90 MMLU-Pro accuracy (%) logit slope 0.976 95% CI [0.793, 1.159] (b) Cross-benchmark stress test 0.250.500.751.001.251.501.752.00 Linking slope (95% CI) METR log 2 MMLU logit MMLU probit MMLU linear MMLU log 2 1.21 0.98 1.04 1.31 1.98 (c) The diagnosis depends on scale 102030405060 Common systems (illustrative) 0.0 0.2 0.4 0.6 0.8 1.0 Power for |slope - 1| = 0.10 n=31 n=23 (d) Anchor count and placement determine detectability METR-like bridge MMLU-like bridge Power curves hold residual noise and transformed-score dispersion at their observed values; they are design illustrations, not universal sample-size rules. Figure 4: Linking diagnostics. Panel (a) links METR Time Horizon 1.0 to 1.1 on a logarithmic scale. Panel (b) shows the six common MMLU/MMLU-Pro systems with paired uncertainty intervals in raw percentage space; MMLU-Pro changes content, choice structure and reasoning demands, so this is a cross-benchmark stress test rather than a claim of construct invariance. Panel (c) shows how the fitted slope changes with score representation; the vertical line marks a shift-like slope of one. Panel (d) shows that detectability depends jointly on anchor count and the observed score dispersion. For six common systems on MMLU and MMLU-Pro, the logit slope is 0.976 (95% CI 0.793â 1.159), the probit slope is 1.043, the linear slope is 1.312 and the logarithmic slope is 1.984. Logit and probit are natural bounded-response links, but boundedness alone does not make either uniquely correct. More importantly, MMLU-Pro is not simply a new form of the same test: it changes item selection, answer options and reasoning demands [55]. The comparison shows why an analyst must state the measurement model; it does not establish strict invariance. Table 1: Power of the observed linking designs. The final column is an illustrative projection assuming the same residual noise and transformed-score dispersion. Deliberately spanning the old score range can materially change the required sample. Bridge and scalen95% CI 80% MDE Power observed n for 10% METR TH1.0 â TH1.1, log 2 7[1.021, 1.390]0.25263%31 MMLU â MMLU-Pro, logit6[0.793, 1.159]0.2486%23 The designs have roughly 80% power only for slope departures near 25% (Table 1). The METR bridge has about 63% power at its observed departure; the MMLU logit bridge has about 6%. Under the strong same-noise and same-spread assumption, detecting a 10% departure would require approximately 23â31 common systems. This is why benchmark revisions need designed anchor panels: common systems or common items chosen across the previous score range and model families, with uncertainty on both axes. Psychometric linking, benchmark- agreement analysis and fixed-parameter calibration provide more appropriate foundations than concatenating versioned leaderboards [21, 24, 25, 27, 43]. 7 7 Beyond One Ruler: A Measurement Portfolio The source audit was expanded to ask what else could be measured. The answer is not another benchmark that replaces all others. Frontier-AI progress has several scientifically distinct layers, and the most defensible record is a portfolio in which each layer has a stated role. Resources and mechanisms. Training FLOP, parameters, tokens and data describe inputs, not capability by themselves. Algorithmic efficiency asks how much resource is needed to reach a fixed performance level [23]. Inference-compute response curves then measure how achieved per- formance changes with tokens, attempts, feedback, parallelism or tool use; fixed-budget scores can understate or reorder systems [34, 51]. Price-performance and energy provide deployment- relevant resource frontiers rather than assuming the cost of eliciting a score is constant [20, 39]. A recent CAISI evaluation demonstrates the combined design: latent-capability aggregation, controlled agent scaffolds and token budgets, held-out tasks and end-to-end cost are reported together rather than as separate leaderboards [10]. Behavior as a distribution, not a point. Psychometric capability profiles and efficient item selection can estimate multiple abilities while retaining item difficulty and uncertainty [26, 30, 32]. Dynamic streams address freshness and contamination, while anchor-based sys- tems preserve longitudinal comparability as instruments grow [11, 21, 24]. Reliability curves replace a single p50 horizon with success probability across task duration, repeated attempts and time budgets. Agentic evaluations such as RE-Bench and PaperBench add realistic long-form work, human baselines and decomposed rubrics [52, 58]. Capability alone is still incomplete: propensity profiles ask whether systems tend toward or away from behaviors that matter for performance and safety [49]. Human-preference systems such as Chatbot Arena measure a dif- ferent endpoint again: perceived comparative usefulness under a changing prompt population [12]. Human ratings are themselves an instrument: multifaceted item-response models can sep- arate rater severity and centrality from output quality, while safety-profile IRT can preserve multidimensional tendencies rather than collapsing them into one score [9, 48]. Validate the construct and the human reference. Two further directions become im- portant whenever a forecast uses phrases such as âreasoningâ, âautonomyâ, âhuman-levelâ or âsuperhumanâ. Measurement-theory work distinguishes the construct from the particular score used to operationalize it, and empirical benchmark audits show that item defects or nuisance phrasing can move results without the intended capability changing [2, 41, 59]. The appropri- ate response is triangulation: pre-specify the construct, seek convergent evidence from more than one instrument, test discriminant and criterion validity, and inspect item- or process-level failure modes. Human thresholds require the same discipline. Human baselines should define the participant population and expertise, use matched tools and time budgets, retain repeated observations and uncertainty, and document exclusions rather than treating a single historical score as a fixed species-level constant [56, 63]. Risk and deployment. Evaluations of dangerous capability and robust refusal are policy- relevant even when they do not track general benchmark averages [33, 50]. Applied-work bench- marks such as LifeSciBench can better represent realistic domain workflows, but still require controls for protocol, judge and provenance [42]. Field experiments measure productivity, qual- ity, heterogeneous effects and task boundaries that laboratory scores cannot supply [4, 7, 13, 14]. Post-deployment monitoring is therefore a measurement layer, not merely operational housekeep- ing [3, 47]. Figure 5 shows why these directions should not be averaged into one âAGI score.â Public availability, joinability, protocol control and external validity are different properties. Human 8 Public longitudinal record Model joinability Protocol control External validity Current readiness Anchored common-scale calibration Repeated open-weight sentinel panel Capability elicitation envelope Resource-performance frontier Task-duration by reliability curve Dynamic contamination-resistant streams Stochastic protocol and judge reliability Multidimensional capability and propensity profile Novel transfer and interactive generalisation Field and workflow outcomes Multi-lab replication and assertion- level provenance Explicit estimands and variance decomposition Post-deployment monitoring and forecast backtesting Human-rater and preference calibration Matched human-reference calibration Construct-validity triangulation longitudinal scale bridge infrastructure elicited capability progression decomposition agentic reliability fresh capability evidence measurement reliability capability-risk profile generalisation evidence external outcome source-dependence control statistical estimand forecast validation human utility signal human-reference threshold claim validity Primary role in a forecast 11212 03311 12322 22222 32222 22212 13322 22222 11221 21232 22221 13322 11231 32232 12332 12332 16 research-backed directions: measurement needs a portfolio Each row addresses a different failure mode; the scores summarize the current public design landscape. score 0123 Scores 0-3 are transparent literature-based design assessments (low to high), not empirical estimates. Public availability, joinability and scientific validity are distinct. Behavior/instrumentSystem/protocolResourcesHuman/deploymentProvenanceEstimand/forecast Figure 5: Sixteen complementary measurement directions identified in the expanded source audit. Circle size and label give a literature-based design assessment from zero (absent) to three (comparatively mature) for public longitudinal record, model-level joinability, protocol control, external validity and current readiness. These are transparent synthesis scores, not empirical estimates. The right column states the role each direction could play in a forecast. preference is longitudinally rich but protocol-sensitive; field outcomes are externally strong but hard to attribute to one model; training compute is joinable for open systems but selectively missing at the closed frontier; safety endpoints can be highly decision-relevant without being monotonic in general capability. Across the portfolio, nine cross-cutting design requirements recur: anchored common-scale calibration; a repeated open-weight sentinel panel; capability-elicitation envelopes; resource- performance frontiers; task-duration by reliability curves; fresh and contamination-resistant item streams; causal field outcomes; novel-transfer and interactive-generalisation designs; and multi-lab replication with assertion-level provenance. Model cards, dataset sheets and formal provenance standards supply useful reporting primitives [17, 29, 38], but they become scientifi- cally consequential only when tied to versioned configurations, reruns and independent source programmes. The package records proposed designs separately from the empirical event graph so that remedies are not confused with measurements already observed. 9 8 Implications for Dated Forecasts The full argument is a chain (Figure 6). Resources are converted by a deployed system config- uration into behavior; behavior interacts with people and institutions to produce outcomes; a forecast maps a versioned endpoint to a date. Four failure modes occur at the links: structured missingness, protocol dependence, benchmark drift and source dependence. A forecast is the end of a measurement chain, not the beginning Every dated claim inherits choices about the target, ruler, version link, data join, source dependence and validation rule. 1 Define the estimand What population, task and quantity? 2 Choose the ruler Score, horizon, cost, profile or outcome? 3 Link versions What stays invariant when instruments change? 4 Join the records Which systems carry both variables? 5 Test independence How many programmes actually replicated? 6 Forecast and backtest Was the dated target frozen and scored? AUDIT FAILURE DESIGN RESPONSE Vague or shifting target Operational criterion and variance model One-point score as intrinsic ability Response curve or multidimensional profile Successive scores concatenated Designed anchors and link uncertainty Sparse or MNAR overlap Sentinel panel and explicit missingness Repeated rows treated as replication Multi-lab reruns and assertion provenance Date published; measurement untested Archived target, monitoring and scoring Measurement portfolio: choose the ruler that matches the research question Linked latent capability Task duration x reliability Resource, cost and energy frontier Construct-valid capability profile Matched human reference Field and deployment outcomes Cross-cutting controls fresh item lifecycles | construct validation | matched human baselines | judge calibration | provenance | uncertainty and missingness The research question should choose the ruler. The ruler should not choose the conclusion. Figure 6: Synthesis of the paper. A dated frontier-AI forecast links resources, the system as used, measured behavior and deployment outcomes. Missingness, protocol dependence, benchmark drift and source dependence can break that chain. The corresponding remedies are disclosure with uncertainty bounds, response curves at matched budgets, versioned anchors and independent replication with back- testing. A defensible forecast packet should therefore report, at minimum: 1. the operational target, population and estimand; 2. the benchmark, version, item lifecycle and linking function; 3. the exact model, scaffold, tools, judge and inference-resource budget; 4. the joint observations supporting the progression relationship and the missingness model; 5. provenance, dependence among sources and independent replications; and 6. a frozen backtest protocol using proper scoring rules or interval coverage after measurement changes. This standard is compatible with provocative forecasting. It does not require waiting for perfect measurement. It requires that the forecast expose where extrapolation begins. A release- date model may be useful even when compute is missing; an agent benchmark may be informa- tive even when it changes; an expert survey may be decision-relevant even when criteria differ. The error is to treat those choices as invisible and then interpret a narrow confidence band as if it covered instrument, protocol and provenance uncertainty. 10 9 Limitations The 62 systems form a curated analytic sample, not a population estimate of all frontier re- leases through 12 August 2026. Coverage percentages must therefore be read as properties of this record. The package explicitly lists known pre-cutoff omissions, but that list is itself not guaranteed exhaustive. The benchmark bridges are small and observational. OLS treats the old transformed score as measured without error; a Bayesian or Deming-style errors-in-variables model would be preferable if comparable uncertainty were available on both versions. The MMLU/MMLU-Pro case changes construct-relevant content and protocol, so it is a stress test of linking assumptions rather than a psychometric equating study. The portfolio review is structured and source-audited but is not a registered systematic review or meta-analysis. Its zero-to-three ratings are reasoned design assessments documented in CSV form. They are intended to make tradeoffs inspectable, not to rank research programmes. Only public evidence is represented. Private evaluations, internal training records, undis- closed inference budgets and proprietary deployment outcomes may be much denser. Their existence would not repair the public record used for independent forecasting, but it limits claims about what developers themselves can infer. Finally, event semantics improve traceability, not causal identification. A graph can reveal that two variables do not join or that one source dominates; it cannot supply missing counter- factuals, invariant constructs or independent replications. 10 Conclusion Frontier AI is being forecast without a common ruler. In the audited public record, the resource and capability tables meet on seven systems; missing compute is structured by access regime and time; benchmark versions can change scale; the diagnosis depends on the chosen link; and most substantive measurements come from one programme. Those are not reasons to abandon forecasting. They are reasons to move the measurement system into the forecast rather than leaving it in the background. The expanded research review also changes the remedy. The field does not need one more scalar presented as the answer. It needs a versioned measurement portfolio: resources and algo- rithmic efficiency; inference-budget response curves; reliability and agentic work; dynamic and psychometric instruments; construct validity and matched human references; safety, preference and field outcomes; and forecast backtesting. Each layer answers a different question and carries a different uncertainty structure. A dated forecast should therefore be read as a composite scientific claim: that the target is operational, the observations are joined, the ruler has been linked across revisions, the protocol is controlled, and the evidence is sufficiently independent. Until those claims are made explicit, the most precise part of many frontier-AI forecasts may be the dateâand the least precise part may be what, exactly, is being measured on the way there. Data and Code Availability The reproducibility package contains 26 CSV tables and an RDF graph in Turtle, together with a 27-source verification log and cutoff audit. It also includes the 56-source literature table, the normalized directionâsource map, the 16-direction measurement portfolio, nine cross-cutting design requirements, analysis and figure scripts, and one reproduction entry point. A clean run rebuilds the normalized data, graph, linking and power analyses, all six figures, bibliography and manuscript PDF without author-specific paths. The arXiv bundle includes the compiled bibliography and the ancillary audit under anc/. 11 References [1] Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammana- manchi, Ruchit Rawal, Vilâem Zouhar, et al. When AI benchmarks plateau: A sys- tematic study of benchmark saturation. arXiv preprint arXiv:2602.16763, 2026. doi: 10.48550/arXiv.2602.16763. [2] Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, In- ioluwa Deborah Raji, and Travis Zack. Position: Medical large language model benchmarks should prioritize construct validity. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 80991â81004, 2025. URL https://proceedings.mlr.press/v267/alaa25a.html. [3] Razvan Amironesei, Afzal Godil, Craig Greenberg, Kristen Greene, Johnston Patrick Hall, Theodore Jensen, Jonathan Fiscus, and Noah Schulman. Assessing risks and impacts of AI (ARIA): Pilot evaluation report. Technical Report NIST AI 700-2, National Institute of Standards and Technology, 2025. [4] Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein.Measuring the impact of early-2025 AI on experienced open-source developer productivity.arXiv preprint arXiv:2507.09089, 2025. [5] Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle OâBrien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 2397â2430, 2023. [6] Olivier Binette and Jerome P. Reiter. Improving the validity and practical usefulness of AI/ML evaluations using an estimands framework. arXiv preprint arXiv:2406.10366, 2024. [7] Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond. Generative AI at work. The Quarterly Journal of Economics, 140(2):889â942, 2025. doi: 10.1093/qje/qjae044. [8] Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. With little power comes great responsibility. In Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing, pages 9263â9274, 2020. doi: 10.18653/v1/2020.emnlp-main.745. [9] Jodi M. Casabianca and Maggie Beiting-Parrish. Correcting human labels for rater effects in AI evaluation: An item response theory approach. arXiv preprint arXiv:2602.22585, 2026. doi: 10.48550/arXiv.2602.22585. [10] Center for AI Standards and Innovation. CAISI evaluation of DeepSeek V4 Pro. Na- tional Institute of Standards and Technology, May 2026. URL https://w.nist.gov/ news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro.Released 1 May 2026; updated 2 May 2026. [11] Simin Chen, Pranav Pusarla, and Baishakhi Ray. Dycodeeval: Dynamic benchmark- ing of reasoning capabilities in code large language models under data contamination. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 8890â8909, 2025.URL https: //proceedings.mlr.press/v267/chen25ba.html. 12 [12] Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 8359â8388, 2024. [13] Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. Management Science, 2026. doi: 10.1287/mnsc.2025.00535. [14] Fabrizio DellâAcqua, Edward McFowland, Ethan Mollick, Hila Lifshitz, Katherine C. Kel- logg, Saran Rajendran, Lisa Krayer, Fran ̧cois Candelon, and Karim R. Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intel- ligence on knowledge worker productivity and quality. Organization Science, 37(2):403â423, 2026. doi: 10.1287/orsc.2025.21838. [15] Epoch AI. Data on AI models. https://epoch.ai/data/ai-models-documentation, 2026. Notable AI Models database; frozen for this audit on 12 August 2026. [16] Epoch AI, with major processing by Our World in Data.Computation used to train notable artificial intelligence systems. https://ourworldindata.org/grapher/ computation-used-to-train-notable-artificial-intelligence-systems, 2026. Re- trieved 12 August 2026. [17] Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daum Ěe I, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86â92, 2021. doi: 10.1145/3458723. [18] Katja Grace, John Salvatier, Allan Dafoe, Baobao Zhang, and Owain Evans. When will AI exceed human performance? evidence from AI experts. Journal of Artificial Intelligence Research, 62:729â754, 2018. doi: 10.1613/jair.1.11222. [19] Katja Grace, Julia Fabienne Sandk Ěuhler, Harlan Stewart, Benjamin Weinstein-Raun, Stephen Thomas, Zach Stein-Perlman, John Salvatier, Jan Brauner, and Richard C. Ko- rzekwa. Thousands of AI authors on the future of AI. Journal of Artificial Intelligence Research, 84, 2025. doi: 10.1613/jair.1.19087. [20] Hans Gundlach, Jayson Lynch, Matthias Mertens, and Neil Thompson. The price of progress: Algorithmic efficiency and the falling cost of AI inference.arXiv preprint arXiv:2511.23455, 2025. [21] Eliya Habba, Itay Itzhak, Asaf Yehudai, Yotam Perlitz, Elron Bandel, Michal Shmueli- Scheuer, Leshem Choshen, and Gabriel Stanovsky. Growing pains: Extensible and effi- cient LLM benchmarking via fixed parameter calibration. arXiv preprint arXiv:2604.12843, 2026. [22] Jos Ěe Hern Ěandez-Orallo. Evaluation in artificial intelligence: From task-oriented to ability- oriented measurement. Artificial Intelligence Review, 48(3):397â447, 2017. doi: 10.1007/ s10462-016-9505-7. [23] Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen, Robi Rahman, Zifan Carl Guo, David Atkinson, Neil Thompson, and Jaime Sevilla. Algorithmic progress in language models. In Advances in Neural Information Processing Systems, volume 37, 2024. arXiv:2403.05812. [24] Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. A Rosetta stone for AI benchmarks. arXiv preprint arXiv:2512.00193, 2025. doi: 10.48550/ arXiv.2512.00193. 13 [25] Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, and Susu Zhang. Can we trust item response theory for AI evaluation?arXiv preprint arXiv:2607.15190, 2026. doi: 10.48550/arXiv.2607.15190. [26] Andrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita Rao, Julia Sharp, and Amanda Bergman. Expanding the AI evaluation toolbox with statistical models. Technical Report NIST AI 800-3, National Institute of Standards and Technology, 2026. [27] Michael J. Kolen and Robert L. Brennan. Test Equating, Scaling, and Linking: Methods and Practices. Springer, 3rd edition, 2014. doi: 10.1007/978-1-4939-0317-7. [28] Thomas Kwa, Ben West, Joel Becker, et al. Measuring AI ability to complete long tasks. arXiv preprint arXiv:2503.14499, 2025. [29] Timothy Lebo, Satya Sahoo, and Deborah McGuinness. PROV-O: The PROV ontology. https://w.w3.org/TR/prov-o/, 2013. W3C Recommendation. [30] Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, and Nitesh V. Chawla. Adaptive testing for LLM evaluation: A psychometric alternative to static bench- marks. arXiv preprint arXiv:2511.04689, 2025. doi: 10.48550/arXiv.2511.04689. [31] Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holis- tic evaluation of language models. Transactions on Machine Learning Research, 2023. arXiv:2211.09110. [32] Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. tinybenchmarks: Evaluating LLMs with fewer examples. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 34303â34326, 2024. [33] Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harm- Bench: A standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 35181â35224, 2024. [34] Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, and Cozmin Ududec. How inference compute shapes frontier LLM evaluation. arXiv preprint arXiv:2606.17930, 2026. [35] METR.Clarifying limitations of time horizon. https://metr.org/notes/ 2026-01-22-time-horizon-limitations/, 2026. [36] METR.Frontier risk report (february to march 2026). https://metr.org/blog/ 2026-05-19-frontier-risk-report/, 2026. [37] METR. Time horizon 1.1. https://metr.org/time-horizons/, 2026. Accessed 12 August 2026. [38] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Trans- parency, pages 220â229, 2019. doi: 10.1145/3287560.3287596. [39] MLCommons. MLPerf Power: A benchmark and measurement framework for power and energy. https://mlcommons.org/2025/03/ml-commons-power-hpca/, 2025. 14 [40] Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. Position: Levels of AGI for oper- ationalizing progress on the path to AGI. In Proceedings of the 41st International Con- ference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 36308â36321, 2024. [41] Seyed Mahed Mousavi, Edoardo Cecchinato, Lucia Hornikova, and Giuseppe Riccardi. Garbage in, reasoning out? why benchmark scores are unreliable and what to do about it. In Findings of the Association for Computational Linguistics: EACL 2026, pages 1747â 1759, 2026. doi: 10.18653/v1/2026.findings-eacl.89. URL https://aclanthology.org/ 2026.findings-eacl.89/. [42] OpenAI.Introducing LifeSciBench:An expert-written, expert-reviewed bench- mark grounded in real-world life science research. https://openai.com/index/ introducing-life-sci-bench/, 2026. Published 17 June 2026. [43] Yotam Perlitz, Ariel Gera, Ofir Arviv, et al. Do these LLM benchmarks agree? fixing benchmark evaluation with BenchBench. arXiv preprint arXiv:2407.13696, 2024. [44] Long Phan et al. A benchmark of expert-level academic questions to assess AI capabilities. Nature, 649:1139â1146, 2026. doi: 10.1038/s41586-025-09962-4. [45] Govind Pimpale, Axel Højmark, J Ěer Ěemy Scheurer, and Marius Hobbhahn. Forecasting frontier language model agent capabilities. arXiv preprint arXiv:2502.15850, 2025. [46] Robi Rahman and David Owen.The training compute of notable AI models has been doubling roughly every six months. https://epoch.ai/data-insights/ compute-trend-post-2010, 2024. [47] Anita Rao, Andrew Keller, Neha Kalra, Ryan Steed, Kweku Kwegyir-Aggrey, Kevin Kly- man, Diane Staheli, and Amanda Bergman. Challenges to the monitoring of deployed AI systems. Technical Report NIST AI 800-4, National Institute of Standards and Technology, 2026. [48] Joshua Fonseca Rivera, Neil Shah, David Demitri Africa, and Konstantinos Voudouris. Item response theory for AI safety. arXiv preprint arXiv:2608.05086, 2026. doi: 10.48550/ arXiv.2608.05086. [49] Daniel Romero-Alvarado, Fernando Mart ĚÄąnez-Plumed, Lorenzo Pacchiardi, Hugo Save, Sid- dhesh Milind Pawar, Behzad Mehrbakhsh, Pablo Antonio Moreno Casares, Ben Slater, Paolo Bova, Peter Romero, Zachary R. Tyler, Jonathan Prunty, Luning Sun, and Jos Ěe Hern Ěandez-Orallo. Capabilities ainât all you need: Measuring propensities in AI. arXiv preprint arXiv:2602.18182, 2026. doi: 10.48550/arXiv.2602.18182. [50] Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023. [51] Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling LLM test-time com- pute optimally can be more effective than scaling model parameters.arXiv preprint arXiv:2408.03314, 2024. [52] Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. Paperbench: Evaluating AIâs ability to replicate AI research. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 56843â56873, 2025. 15 [53] Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos, Mahmood Hegazy, Al- berto Tosato, David John Lemay, Irina Rish, and Guillaume Dumas. Persistent instability in LLMâs personality measurements: Effects of scale, reasoning, and conversation history. Proceedings of the AAAI Conference on Artificial Intelligence, 40(44):37961â37969, 2026. doi: 10.1609/aaai.v40i44.41133. [54] Willem Robert van Hage, V Ěeronique Malais Ěe, Roxane Segers, Laura Hollink, and Guus Schreiber. Design and use of the simple event model (SEM). Journal of Web Semantics, 9 (2):128â136, 2011. doi: 10.1016/j.websem.2011.03.003. [55] Yubo Wang, Xueguang Ma, Ge Zhang, et al. MMLU-Pro: A more robust and challeng- ing multi-task language understanding benchmark. In Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 2024. arXiv:2406.01574. [56] Kevin Wei, Patricia Paskov, Sunishchal Dev, Michael J. Byun, Anka Reuel, Xavier Roberts- Gaal, Rachel Calcott, Evie Coxon, and Chinmay Deshpande. Position: Human baselines in model evaluations need rigor and transparency. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 82265â82325, 2025. URL https://proceedings.mlr.press/v267/wei25s.html. [57] Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum. LiveBench: A challeng- ing, contamination-free LLM benchmark. arXiv preprint arXiv:2406.19314, 2024. [58] Hjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Joshua M. Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Jun Koba Sato, William Saunders, Maksym Taran, Ben West, and Elizabeth Barnes. RE-bench: Evaluating frontier AI r&d capabilities of language model agents against human experts. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 66772â66832, 2025. [59] Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10967â 10982, 2023. doi: 10.18653/v1/2023.emnlp-main.676. URL https://aclanthology.org/ 2023.emnlp-main.676/. [60] Zongyou Yang, Yinghan Hou, and Xiaokun Yang. When the judge changes, so does the measurement: Auditing LLM-as-judge reliability. arXiv preprint arXiv:2607.08535, 2026. doi: 10.48550/arXiv.2607.08535. [61] Yuming Zhai et al. HLE-Verified: A systematic verification and structured revision of humanityâs last exam. arXiv preprint arXiv:2602.13964, 2026. [62] Wenting Zhao et al. MMLU-CF: A contamination-free multi-task language understanding benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computa- tional Linguistics, 2025. doi: 10.18653/v1/2025.acl-long.656. [63] Yan Zhuang, Qi Liu, Zachary Pardos, Patrick C. Kyllonen, Jiyun Zu, Zhenya Huang, Shijin Wang, and Enhong Chen. Position: Ai evaluation should learn from how we test humans. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 82483â82508, 2025. URL https: //proceedings.mlr.press/v267/zhuang25e.html. 16