Paper deep dive
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu, Yuzhuo Bai, Sang T. Truong, Xiaoyuan Yi, Sanmi Koyejo, Xing Xie, Ziang Xiao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/20/2026, 7:53:46 AM
Summary
This position paper argues that standardized item-level benchmark data should be the default infrastructure for AI evaluation to address validity issues like construct misalignment and poor generalization caused by a focus on aggregate scores. The authors introduce OpenEval, an item-level archive containing 10 million responses across 155,000 items from widely-used benchmarks, demonstrating how such data enables transparency, replicability, and rigorous validity evidence.
Entities (11)
Relation Signals (11)
OpenEval → contains → 10M responses
confidence 95% · OpenEval, an item-level archive of 10M responses across 155k items
OpenEval → covers → 155k items
confidence 95% · OpenEval, an item-level archive of 10M responses across 155k items
RealToxicityPrompts → exampleof → Toxicity Benchmark
confidence 90% · REALTOXICITYPROMPTS [26] is one of the most widely used AI toxicity benchmarks
OpenEval → usesschema → Unified Schema
confidence 90% · under a unified schema that the AI evaluation community can develop upon
Item Response Theory → usedfor → Item Analysis
confidence 88% · Classical Test Theory (CTT) and Item Response Theory (IRT) are two foundational frameworks for this purpose
Classical Test Theory → usedfor → Item Analysis
confidence 88% · Classical Test Theory (CTT) and Item Response Theory (IRT) are two foundational frameworks for this purpose
CULTURALBENCH → incorporatedinto → OpenEval
confidence 85% · incorporating interdisciplinary datasets on social aspects of AI, such as CULTURALBENCH
EMOBENCH → incorporatedinto → OpenEval
confidence 85% · and EMOBENCH [61]
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified item selection, construct misalignment, and poor generalization. The root cause of these failures is a misplaced focus on aggregate model scores. Without item-level evidence, validity claims cannot be assessed, resulting in inflated capability claims, misdirected research, and unwarranted trust in deployed systems. Our position is that designing valid evaluations requires empirical evidence from item-level model responses, and the standardized release of such data should be treated as core AI evaluation infrastructure. Such a release, in addition, enables transparency, replicability, and auditability of evaluation results. To show the norm is both feasible and consequential, we construct OpenEval, an item-level archive of 10M responses across 155k items from widely-used benchmarks, under a unified schema that the AI evaluation community can develop upon. We demonstrate how item-level data can identify low-quality items, document construct misalignment, and recover validity evidence about benchmarks' internal structure. We address objections around contamination and author burden, and show each is tractable relative to the cost of decisions made on claims that cannot be trusted.
Tags
Links
- Source: https://arxiv.org/abs/2604.03244v2
- Canonical: https://arxiv.org/abs/2604.03244v2
Trouble viewing inline? Open PDF directly →
Full Text
68,874 characters extracted from source content.
Expand or collapse full text
AI Evaluation Should Require Standardized Item-Level Data Releases Han Jiang 1 Susu Zhang 2 Dongyao Zhu 5 Yuzhuo Bai 6 Sang T. Truong 4 Xiaoyuan Yi 3 Sanmi Koyejo 4 Xing Xie 3 Ziang Xiao 1 1 Johns Hopkins University 2 University of Illinois Urbana-Champaign 3 Microsoft Research Asia 4 Stanford University 5 North Carolina State University 6 Tsinghua University hjiang66@jh.edu, ziang.xiao@jhu.edu Abstract This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified item selection, construct misalignment, and poor generalization. The root cause of these failures is a misplaced focus on aggregate model scores. Without item-level evidence, validity claims cannot be assessed, resulting in inflated capability claims, misdirected research, and unwarranted trust in deployed systems. Our position is that designing valid evaluations requires empirical evidence from item-level model responses, and the standardized release of such data should be treated as core AI evaluation infrastructure. Such a release, in addition, enables transparency, replicability, and auditability of evaluation results. To show the norm is both feasible and consequential, we construct OPENEVAL, an item-level archive of 10M responses across 155k items from widely-used benchmarks, under a unified schema that the AI evaluation community can develop upon. We demonstrate how item-level data can identify low-quality items, document construct misalignment, and recover validity evidence about benchmarks’ internal structure. We address objections around contamination and author burden, and show each is tractable relative to the cost of decisions made on claims that cannot be trusted. 1 Introduction Generative AI is moving rapidly into high-stakes deployments, while AI evaluation, dominated by benchmarking practice [21], has become the primary instrument for understanding model capabilities, informing AI policy, and guiding responsible deployment. However, the empirical foundation of such instruments is thinner than their influence suggests. Benchmarks are released, leaderboards are populated, and deployment and decisions are made, while the underlying response data that would let the community audit whether benchmark measures what they claim to measure and evaluate benchmark reliability are rarely shared and analyzed. This gap entails consequences. Critical design choices, including capability definitions, content curation, and metric selection, often lack transparency or formal justification [49]. This opacity undermines the validity evidence [10,49] needed to support the interpretations of results, making it unclear whether benchmarks genuinely measure their intended constructs, despite explicit metadata or task descriptions [2]. Compounding these design limitations, the speed of AI development manifests as benchmark saturation [56], rapidly outdated content [37], and widespread data contamination [76], rendering aggregate scores uninformative or misleading for deployment decisions [27]. New bench- marks proliferate to address these concerns, but often recycle existing practices without adding Preprint. arXiv:2604.03244v2 [cs.AI] 22 May 2026 evaluative insight [9], widening socio-technical gap between technical solutions and real-world requirements [68, 45]. Science of AI Evaluation Principled Benchmark Design & Assessment Missing Foundation: Item-Level Benchmark Data Validity Current Paradigm: Aggregate Benchmark Scores ReliabilityEfficiency Reliable & Insightful Evaluative Claims Efficient Maintenance, Use, & Cumulative Development Figure 1: Illustrative overview of our position. Item-level benchmark data serves as the miss- ing foundation between the current aggregate paradigm and the science of AI evaluation. Crucially, many validity issues are not diag- nosable from benchmark-level aggregate scores alone.Foundational questions—including whether items effectively differentiate model ca- pabilities, how construct-irrelevant nuisance fac- tors drive performance, or whether gains reflect genuine improvement rather than artifacts—are inherently item-level inquiries. Without item- level benchmark data, our field lacks the em- pirical evidence required to evaluate and curate effective benchmarks. We argue that standard- ized item-level benchmark data should be- come the default infrastructure for AI eval- uation, released alongside aggregate scores whenever legally and ethically feasible: de- signing valid evaluations requires empirical ev- idence from item-level model responses, and the standardized release of such data should be treated as core evaluation infrastructure on par with the benchmarks themselves. Shifting the focus from aggregated model score to item-level model outputs has deep roots in measurement science, where item-level assessment data have long been significant to test development and validation across education, psychology, and other social sciences disciplines. Bringing this granularity to AI evaluation would transform fragmented leaderboard results into cumulative empirical evidence: item content, score statistics, and per-item responses jointly enable rigorous construct validation, reliability analysis through item consistency and discrimination patterns, and efficiency improvements through identification of redundant or uninformative items. Item-level data is the foundation for AI evaluation science. To show the norm is feasible and consequential, we construct OPENEVAL, an item-level archive of 10M responses across 155k items from widely-used benchmarks, organized under a unified schema the community can contribute to. We use OpenEval to demonstrate how item-level data identifies low- quality items, documents construct misalignment, and recovers validity evidence about benchmarks’ internal structure. At the end, we discussed broader opportunities and addressed objections around storage, contamination, and author burden. 2 Validity Challenges in AI Benchmarking Despite their pivotal role, AI benchmarks face validity challenges from two directions, namely internal methodological limitations and external pressures from the rapid evolution of AI systems, both of which are exacerbated by the absence of item-level data. 2.1 Methodological Issues in Benchmark Design The term Validity, specifically Construct Validity, was formalized by Cronbach and Meehl [17] in modern psychometrics and is directly applicable to AI evaluation [74]. A Construct refers to a theoret- ical, unobservable attribute or trait (e.g., a perceived AI capability) that a test is intended to measure. Validity thus reflects how well an assessment measures the intended underlying outcome [64] and is crucial for ensuring benchmark quality. As noted in [71], construct validity is a central concept in psychology, particularly for informing the design of psychological measures. In AI evaluation, however, validity has received limited attention, despite AI capabilities rapidly expanding and benchmark creation involving numerous design decisions, such as target capability identification, construct-to-task operationalization, item adaptation for AI systems, and metric selection. These decisions are often oversimplified due to a dominant focus on benchmark-level results, leaving only standard or default settings (e.g., [44,69]) and insufficient evidence for validity justification [49]. The hidden ambiguities have resulted in a 2 lack of a common language for communicating validity within the AI community. Although there are several studies exploring validity-relevant properties (e.g., [79,73]), research explicitly assessing AI benchmark validity remains limited. Existing analyses suggest that most benchmarks’ validity remains underdeveloped and call for evidence beyond the benchmark level (e.g., [10, 6, 62]). Benchmarks designed with validity issues act as traps: their claimed goals may sound aligned with users’ intended purposes, yet the underlying constructs being measured could diverge substantially, yielding unreliable conclusions, as demonstrated by [6], that propagate through benchmark selection and all subsequent evaluation practices. Such misalignment can be further exacerbated by confounders unrelated to the intended system ability [78], such as erroneous items, spurious correlations, or unintended shortcuts (e.g., [20, 54]). Notably, both construct misalignment and these overlooked confounders are only visible at the item level, where the content of each item can be scrutinized and its empirical behavior systematically examined. The absence of item-level data hinders reasoning about what actually drives benchmark performance — a question left unresolved from the design stage — which not only wastes the effort invested in benchmark selection, but also leaves the validity of any downstream claim unjustified. Yet existing benchmark data remains largely untapped at this level, leaving benchmark design without the empirical feedback needed to become more principled. 2.2 AI Advances Pressures on Benchmarking Practices Beyond the design limitations, the rapid evolution and increasing opacity of AI systems impose external pressures that further erode benchmark validity over time. Evaluation Setting Pre-Nov. 2023 models on MMLU Post-Jun. 2024 models on MMLU-Pro Figure 2: Benchmark-level accuracy distributions for 66 pre–Nov. 2023 models on MMLU and 72 post–Jun. 2024 models on MMLU-PRO. As AI systems and real-world knowledge co- evolve, benchmarks are subject to various forms of validity degradation. [18, 38]. Benchmark saturation. Some benchmarks have gradually become too easy to distinguish be- tween latest models [56,19]. For instance, REALTOXICITYPROMPTS [26] is one of the most widely used AI toxicity benchmarks; al- though by 2023 it could no longer differentiate between multiple versions of GPT [36], it re- mained widely used beyond 2024. Outdated knowledge. Benchmarks requiring fac- tual knowledge are particularly time-sensitive, as outdated references render the corresponding test items unreliable [37]. Data contamination. The opacity and informa- tion asymmetry in AI system development have made the risk increasingly pervasive. As AI training and evaluation scale, many omnibus benchmarks further aggregate multiple previous bench- marks, making data provenance harder to trace. Both intentional and inadvertent test-train overlap can secretly lead to unfair evaluations unless explicitly reported by developers [80]. All these issues are nearly impossible to detect at benchmark level, and Fig. 2 exemplifies this diagnostic gap: the benchmark-level accuracy distribution has shifted markedly rightward as newer models are evaluated, even though MMLU-PRO was designed to be more challenging. Without item- level inspection, it remains unknown whether the observed improvements reflect genuine capability gains, benchmark saturation, or data contamination, leaving no valid claims about model capabilities to be drawn. This progressive degradation at scale creates a fundamental tension between evaluation efficiency and validity. Manual benchmark updates (e.g., [46,72]) are believed to generate higher-quality items but are typically time-consuming and costly. LLM-powered benchmarking (e.g., [39,23]) improves efficiency but at the expense of validity, as the quality of synthesized test items has been questioned in [11]. More recent benchmark generation techniques, such as adversarial filtering (e.g., [41,55]), item generation (e.g., [47,40]), and adaptive difficulty adjustment (e.g., [66]), offer promising 3 directions but require exploring inter-item dynamics and engagement with measurement theory. In this regard, drawing on psychometric methods for item analysis could provide principled guidance for navigating this tension, making it possible to isolate contaminated or outdated items, identify affected constructs, and inform decisions about item retention, replacement, and generation under evolving conditions. Additionally, a common bottleneck underlying all of these efforts is the lack of shared data in- frastructure. Current evaluation results are aggregated locally and not consistently released across leaderboards, making them neither comparable nor easily built upon, and forcing each study to independently assemble and preprocess its data. This duplicates effort across research groups and prevents cumulative progress. Moreover, ensuring the long-term sustainability of a unified item-level data release is itself a real challenge, as the effort required to curate, standardize, and update data across diverse benchmarks is prohibitive for any single research group. That is to say, without a community-driven, consistent release of item-level benchmark data, individual efforts to counter these ecosystem pressures remain isolated and unscalable. 3 Item-Level Data: The Missing Foundation for AI Evaluation Science Given the twofold validity challenges in AI benchmarking and the limits of existing benchmarking practices, item-level benchmark data is a crucial missing piece in the science of AI evaluation. Every benchmark run already generates such data: detailed test conditions, the content of each item, the model’s response, and per-response scores and statistics. However, this evidence is routinely discarded once aggregate scores are computed. What remains is a rich yet largely inaccessible foundation for investigating the validity, reliability, and efficiency of AI benchmarking. We argue that releasing item-level data under a unified schema (e.g., 10) should become a community norm when publishing evaluation results. Reconstituting this evidence into shared infrastructure is a precondition for evaluation to function as a cumulative and open science. Item-level data enables principled benchmark design and assessment. Item-level data is a prerequisite for developing measurement theories for AI systems, including principled notions of construct and validity, analogous to those in psychometrics. These theories, combined with the rich evidence from item-level analysis, make it possible to scrutinize, communicate, and enhance AI benchmark validity. With principled practices, such as quantifying item characteristics, identifying decisive benchmark factors, and modeling relationships between items and intended constructs, future benchmarks can better clarify the mapping between AI competencies and evaluation objectives. Item-level data enables reliable and insightful evaluative claims. Item-level data facilitates a bidirectional attribution that addresses the opacity in AI evaluation caused by the aggregate-focusing paradigm. First, observed model performance can be traced to measurable factors, disentangling intended constructs from confounders such as construct-irrelevant shortcuts and data leakage. Second, item properties such as difficulty and discrimination, as well as the latent constructs they collectively reflect, can be investigated for how they manifest across AI systems, revealing where models systematically succeed or fail. Together, these two perspectives ground evaluative claims in empirical item-level evidence, supporting more defensible conclusions about capability and deployment. Item-level data enables efficient maintenance, use, and cumulative development of AI bench- marks. With increased access to item-level data, score statistics and detailed test cases make it possible to diagnose the validity threats discussed in Sec. 2.2, including benchmark saturation and data contamination, in a timely manner. Item-level analyses can further identify which items have de- graded and which remain informative, providing fine-grained guidance for benchmark re-composition, updates, and targeted item generation, saving effort in maintenance rather than multiplying it. More- over, a consistent, community-wide release of item-level data lowers the costs of data engineering and allows findings to be reproduced, audited, and compared across studies, enabling cumulative progress and prolonging benchmark lifecycles. 4OPENEVAL: An Item-Centered Benchmark Repository To show what data infrastructure in our position may look like, we propose an item-centered reposi- tory, OPENEVAL, designed as a community-driven data infrastructure to standardize benchmark items with model responses, metric scores, and other associated information across different tasks, formats, 4 Model item_adaptation Score value response_content* item_metadata item_content metric namesize model_adaptation Item Response Benchmark Indexed by benchmark_name Indexed by item_id starting with benchmark_name Indexed by response_id starting with item_id 1:N 1:N benchmark_version paper_urldataset_url benchmark_tags input* source contributoringestion_time references* system_instruction generation_parameters tools* extra_artifacts*models name request_input* external_resources* demonstrations 1:N 1:1 Figure 3: Unified data schema of OPENEVAL. The left column shows the hierarchical storage structure with indexing keys; the right side expands each entity into its constituent fields. Diamond- tipped connectors point from parent to subfield; dashed connectors indicate folded siblings. 1:1 and 1:N denote cardinality, e.g., each response contains exactly one model but may contain multiple scores. Asterisks (*) mark fields accepting values in arbitrary formats (implemented asList[Dict]). and evaluation harnesses. OPENEVAL is not the endpoint of our proposal. It is a demonstration that item-level standardization is feasible and a seed for community norms. There have been many large-scale, high-quality AI benchmark repositories (e.g., HELM [44], Chatbot Arena [13], and Open LLM [24]), some of which release unstructured or semi-structured item-level details alongside the benchmark-level results. While these resources provide a basis for recent AI evaluation research, the shared data infrastructure remains underdeveloped, hindering cumulative progress across studies and groups. OPENEVAL therefore addresses the infrastructure challenge through a unified, item-centered schema, as illustrated on the right side of Fig. 3 and detailed in App. C, in which each data entry represents a unique item. Specifically, the schema is designed around four principles: •Unified while remaining flexible. While effectively standardizing item-level data across different sources, the schema poses relatively loose constraints on fields such asresponse_content anditem_contentto enable incorporation of heterogeneous content.The typed aux- iliary arrays (model_adaptation.tools,item_adaptation.external_resources, and metric.extra_artifacts) further accommodate items from various evaluation paradigms, in- cluding agent and tool-augmented settings, allowing researchers to focus on research opportunities rather than data engineering overhead. •Faithful to evaluation context. The schema separates the original item content from the adapted input actually given to the model and records the corresponding model configuration, providing a clean snapshot of the test condition for each response rather than collapsing various experimental conditions into a single record. This self-contained, faithful preservation supports reproducibility and enables meaningful comparisons across different experimental setups. •Informatively rich and versatile. The schema archives diverse information at multiple levels, including benchmark-level metadata, test environments, and evaluation metrics, enabling a wide range of practices from provenance tracking and benchmark auditing to coarse- and fine-grained querying by domain, model, or other properties. This richness reflects the paper’s core premise: item-level benchmark data serves the science of AI evaluation not for any single analytical goal, but for a wide range of research opportunities. • Scalable for long-term maintenance. The unified, flexible, and self-contained nature of each entry naturally supports scalability. As shown on the left side of Fig. 3, the repository organizes data into three levels, each indexed by hierarchical identifiers (e.g.,response_idprefixed byitem_id). New benchmarks, items, or model responses can be easily incorporated by appending new entries at the appropriate level without restructuring existing data; this design further facilitates community-driven contributions. So far, to jump-start the community effort, we have been (1) collecting evaluation results to supplement item-level data coverage across existing benchmark repositories, (2) incorporating interdisciplinary datasets on social aspects of AI, such as CULTURALBENCH [14] and EMOBENCH [61], and (3) trans- 5 forming existing item-level resources from benchmark repositories (e.g., HELM) into OPENEVAL schema. OPENEVAL now covers over 155K items across diverse benchmark datasets, with the number of evaluated models per dataset ranging from 11 to 111 (70.3 on average), resulting in 10M item-level responses, each associated with one or multiple metric scores. To lower the barriers for data release, we offer multiple data converters for popular evaluation harnesses. By shifting the norm from reporting aggregate scores to releasing item-level results as a standard part of research publication, we hope that OPENEVAL will grow into a shared data foundation for the science of AI evaluation and responsible AI deployment, leading us toward more rigorous, empirically grounded validation of benchmark-related claims. 5 Empirical Illustrations of Item-Level Analysis To demonstrate the unique insights enabled by item-level benchmark data, we present illustrative analyses on selected item-level data from OPENEVAL that provide finer-grained understanding of existing AI benchmark datasets, specifically item quality and construct alignment. 5.1 Background and Analytical Approaches Item-level benchmark analysis has drawn increasing attention in AI evaluation research, with studies examining properties such as item difficulty (e.g., [79,42]), item discrimination, namely the ability to separate different capability levels (e.g., [32,43]), item diversity (e.g., [53,79]), inter-benchmark agreement (e.g., [57,48]), and downstream performance predictability (e.g., [8,52,73]). As discussed in Sec. 2.2, these analyses are typically isolated and difficult to reproduce due to the lack of a large- scale, consistent item-level data release, leaving substantial room for further exploration. Notably, these studies, to varying degrees, draw on psychometric methodology, which offers a well-established toolkit for item-level analysis grounded in decades of test development and validation practice. Two types of analysis from this discipline are particularly relevant to AI evaluation. The first is item characteristic analysis, which examines whether individual items function as intended with adequate precision for their intended use [3]; Classical Test Theory (CTT) and Item Response Theory (IRT) are two foundational frameworks for this purpose [15,29]. CTT measures item difficulty and discrimination directly from observed score matrices with minimal assumptions, making it readily applicable to large-scale, heterogeneous AI benchmark data [28,51]; IRT provides a more sophisticated, model-based alternative with stronger statistical properties, but requires parametric assumptions about the relationship between model ability and item-level performance, which requires careful validation that item responses are driven by approximately a single latent ability [50,30]. The second is Item Factor Analysis (IFA), which provides evidence about a test’s internal structure [5,59]. The Standards for Educational and Psychological Testing [3] emphasizes that such structural evidence should be provided when a single aggregate score is interpreted as a meaningful summary of examinee ability. In the context of AI evaluation, IFA examines whether a benchmark behaves as a coherent measure of the intended model capability versus reflecting construct-irrelevant variance, such as formatting artifacts, label distribution bias, and data contamination. 5.2 Examining Item Quality via Classical Test Theory An item’s statistical characteristics such as difficulty and discrimination are routinely examined in psychometric test development for quality assurance. We conduct a CTT analysis of item character- istics from (1) 66 pre-Nov. 2023 models on 567 items in MMLU [33] and (2) 72 post–June 2024 models on 1,000 items in MMLU-PRO [70]. MMLU-PRO [70] is an enhanced variant of MMLU intended to increase difficulty and reduce noise via additional distractors, more careful item curation, and expert item review. CTT item analysis provides an empirical way to evaluate these design claims. For thei-th item in a benchmark, the Item Difficulty Index (p i ) can be estimated as [28] the proportion of the maximum score achieved on the item averaged across models; a largerp i indicates an easier item. The Item Discrimination (r i ) is measured by the Pearson correlation between the item score and the Rest-Total Score (s rest i , sum score on all items excepti) across all measured models [34]. Higher 6 Evaluation Setting Pre-Nov. 2023 models on MMLU Post-Jun. 2024 models on MMLU-Pro Item Difficulty (0.5 - 푝 푖 ) Item Discrimination (a) Item Characteristics Distribution (b) Item Characteristic Curves in MMLU Rest-Total Score (푠 푖 rest ) Item Difficulty Index 푝 푖 Item ID 푖 with 푟 푖 #496 (0.8401) #55 (-0.0663) #374 (-0.4535) Figure 4: (a) Item characteristics distribution for items from MMLU and MMLU-PRO. Higher item difficulty values correspond to harder items. (b) ICCs for three items in MMLU. r i indicates that itemican well-differentiate models with strong vs. weak overall performance on the benchmark, whereas negative or near-zero r i suggests a potentially problematic item. Fig. 4(a) shows the distributions of item difficulty and discrimination under the two evaluation settings. (Note that the CTT item difficulties on MMLU and MMLU-PRO are specific to their respective sample of models and are not comparable.) There are two notable observations: (1) The high density of orange observations on the left indicates that a substantial proportion of MMLU-PRO items have very low difficulty for current models. In other words, many items are no longer challenging for the 72 post-June 2024 models, suggesting fast benchmark saturation and the need to accelerate benchmark updates. (2) Compared to MMLU, item quality substantially improves on MMLU- PRO with much fewer items with low or negative discrimination. This empirical observation is aligned with MMLU-PRO designers’ goal [70] to build a more robust, less noisy benchmark. However, some MMLU-PRO items still show poor discrimination. Although these items remain in MMLU-PRO after expert item review, their poor empirical discrimination merits additional scrutiny (e.g., for ambiguity, miskeying, or construct-irrelevant cues). To further understand how item discrimination manifests, we plot the Item Characteristic Curves (ICCs) of three items in MMLU. For each itemi, all respondent models are sorted into six equally- sized bins based on their rest-total scores (s rest i ), representing six levels of overall performance. The item difficulty index (p i ) within each bin is then plotted in Fig. 4(b). Intuitively, a high-discrimination item’sp i should increase withs rest i , meaning that the item appears easier for higher-performing models, resulting in a monotonically increasing ICC. This is the case for item #496 which had a high item discrimination (r 496 =0.84), whereas for items #55 and #374 with near-zero or negativer i , models that did well on the rest of the benchmark did worse on these items. These findings can help identify low-quality items in the benchmark, thereby improving the alignment between benchmark outcomes and design goals at a finer granularity. In this sense, detailed exami- nation of item characteristics supports more targeted benchmark maintenance and item generation without requiring wholesale replacement. 5.3 Revealing Construct Alignment via Item Factor Analyses As discussed in Sec. 2.1, examination of a benchmark’s internal structure is essential for understanding whether it measures the intended capabilities or irrelevant confounders. Here, we perform variants of conventional IFA for high-dimensional data based on Singular Value Decomposition [SVD;81] and Generalized Low Rank Models [GLRM;67]. We report findings from BABIQA Task 15 (basic deduction) and MMLU-PRO. Additional analysis results on MMLU are presented in App. B. 7 Table 1: Example BABIQA item and item counts by reference answers for each cluster in Fig. 5. Example: Item #1295 Sheep are afraid of mice. Cats are afraid of mice. Jess- ica is a sheep. Wolves are afraid of mice. Mice are afr- aid of wolves. Emily is a wolf. Gertrude is a wolf. Winona is a mouse. Question: What is Emily afraid of? Answer: mouse Number of Items Item AnswerCluster #0Cluster #1Cluster #2 Sheep00240 Mouse22500 Cat22130 Wolf63260 Figure 5: Item clusters on BABIQA based on factor loadings. BABIQA Task 15 [71] aims to assess basic deductive reasoning via inheritance of properties. An example item is shown in Table 1. To aid interpretation, we perform K-means clustering on items’ factor loadings on the top 3 factors obtained from SVD-based IFA, with items within each cluster potentially measuring similar sub-constructs. As shown in Fig. 5, 1,000 items in BABIQA form three distinct clusters. A closer examination raises a construct validity flag: Table 1 shows that item clusters are explained by the reference answer to the item. This finding demonstrates that different models’ performance on BABIQA is partially explained by models’ propensity to select specific animals that one is afraid of (e.g., potentially based on common sense if a model tends to select “wolf”), rather than the intended basic deduction capability, suggesting potential construct misalignment in BABIQA Task 15. GLRM-based IFA yields consistent findings. For MMLU-PRO, we interpret the top 4 factors retained from GLRM-based IFA after varimax rotation. The top 100 items with the largest absolute loadings on each factor are sent to GPT-5 for interpretation. Table 2 presents a representative item and a tentative sub-construct interpretation for each factor. The four primary dimensions that best explain differences in model performance appear to reflect different higher-level reasoning capabilities, rather than subject domain proficiency. This empirical finding supports MMLU-PRO’s stated motivation to increase reasoning demands [70] relative to MMLU. Indeed, items within the same subject domain (e.g., Psychology and Physics in Fig. 9) could differ substantially in loadings on the four factors. #1 #2 #3 #4 GPQA Omni- MATH #1#2#3#4GPQAOmni- MATH Convergent/Discriminant Validity: MMLU-Pro Figure 6: Convergent/discriminant evidence of the four sub-constructs (#1 - #4) on MMLU-PRO. As an external plausibility check using conver- gent and discriminant validity evidence [12], we correlate factor subscores (mean scores on the 100 items with the largest absolute loadings) with scores on two external bench- marks: GPQA [60] (graduate-level biology, physics, and chemistry) and OMNI-MATH [25] (Olympiad-level mathematics).Both target high-level formal reasoning, with the former grounded more in applied scientific contexts. We hypothesize that Factor #1 (formal, quanti- tative, multi-step modeling) aligns with both formal reasoning benchmarks, whereas Fac- tor #4 (applied synthesis and case-based judg- ment) aligns more with GPQA. Results in Fig. 6 are broadly consistent with these hypothe- ses. Further, Factors #2 (domain-specific recall and simple reasoning) and #3 (conceptual under- standing and explanation) show weak correlations with both GPQA and OMNI-MATH, providing discriminant validity evidence. We treat these findings as descriptive rather than definitive evidence supporting these sub-construct interpretations. 8 Table 2: Representative items with large absolute item factor loadings and possible constructs for each GLRM factor in MMLU-PRO. The top 100 representative items are interpreted and summarized into a candidate label by GPT-5, then manually revised. FactorRepresentative ItemPotential Sub-Construct MMLU-Pro #1 A 10kVA, 2400/240V, single-phase transformer has the following resistances and leakage reactance. Find the primary voltageFormal, quantitative, required to produce 240V at the secondary terminals at full load,multi-step modeling when the load power factor is 0.8 power factor lagging/leading. MMLU-Pro #2 What are the principal and selective instruments of control of whichDomain-specific recall the Federal Reserve System makes use?and simple reasoning MMLU-Pro #3What is meant by the term “hypothesis testing”? Conceptual understanding and explanation Which of the following might explain how a price decrease might Applied synthesis and MMLU-Pro #4cause a decrease in quantity demanded and an upward-sloping case-based judgment demand curve? 6 Broader Implications beyond Benchmarking Releasing item-level data can benefit beyond AI evaluation. Shared infrastructure of this kind could benefit AI research, ground policy in evidence, and democratize AI through participatory evaluation. For AI Researchers. Releasing such data at scale could contribute to ML theory, where central questions about generalization and learning dynamics are inherently item-level. Methods built on item-level data have advanced principled training: identifying which examples drive learning [65], informing data mixing [16,75], and grounding curriculum learning [7] through measurable item difficulty. The same per-instance signal supports estimating generalization gaps and attributing predictions to training examples [77]. For Policy Makers. Policy and governance decisions increasingly rely on AI benchmark results to justify claims about model capability, risk, and deployment readiness (e.g., [4,31,22]). Item- level data lets policy makers inspect what a benchmark actually measures, enabling contextualized judgments about how much weight a score should carry for a given policy question. The evaluation methods built on such data further equip regulators to audit benchmark design, verify developer claims, and ground policy decisions in independently checkable evidence. For End Users. The communities affected by AI deployment are largely absent from the evaluations that justify it. Item-level responses, not just items, make user engagement substantive: users can see whether a model actually handles cases like theirs, detect performance gaps that aggregate scores hide, and compare options before deployment [58]. Interfaces built on this infrastructure can lower the technical barrier to participation, making participatory evaluation practical: evaluation as a channel through which users help shape priorities, and through them, model development [45]. 7 Alternative Views Item-level release worsens contamination. A natural objection is that releasing item-level data worsens data contamination [63,35]. We view this trade-off as favoring release. Most widely used benchmarks already appear in pretraining corpora. Withholding item-level data does not prevent contamination. It only makes it harder to detect. For example, a recent analysis [1] showed that having private test sets does not prevent benchmark saturation. Item-level release addresses contam- ination directly by making contaminated items detectable and removable, without sacrificing the reproducibility that competitions abandon. In addition, our position is compatible with strategies such as hold-out designs, verified access, and staged release after a fixed embargo period to protect strategic misuse and licenced content. We believe the asymmetry that decides the trade-off: contaminated benchmarks can be replaced, but AI evaluation built on unverifiable scores cannot. Mandating item-level release creates additional burden. Standardized item-level release does require additional effort from benchmark authors and model evaluators. But the data already exists when running the evaluation. There is no additional compute cost. The remaining work is converting 9 responses to a shared schema and uploading them. As demonstrated by OPENEVAL, a centralized, schema-standardized infrastructure absorbs the costs. Contributors submit to a shared archive rather than building bespoke release pipelines, the schema removes the design work of deciding what to release in what format, and converters for widely-used evaluation harnesses minimize the reformatting burden. The marginal cost of release is closer to uploading a results table than to maintaining a benchmark website. The marginal cost of data releasing is closer to the cost of uploading a results table than to the cost of maintaining a benchmark website. We do not deny that some burden remains. Releasing item-level data creates transparent, auditable, and reproducible AI evaluation and a strong foundation of AI evaluation science. References [1]M. Akhtar, A. Reuel, P. Soni, S. Ahuja, P. S. Ammanamanchi, R. Rawal, V. Zouhar, S. Yadav, C. Whitehouse, D. Ki, et al. When ai benchmarks plateau: A systematic study of benchmark saturation. arXiv preprint arXiv:2602.16763, 2026. [2] A. F. Akyürek, M. Y. Kocyigit, S. Paik, and D. T. Wijaya. Challenges in measuring bias via open-ended language generation. In C. Hardmeier, C. Basta, M. R. Costa-jussà, G. Stanovsky, and H. Gonen, editors, Proceedings of the 4th Workshop on Gender Bias in Natural Lan- guage Processing (GeBNLP), pages 76–76, Seattle, Washington, July 2022. Association for Computational Linguistics. [3]American Educational Research Association, American Psychological Association, and National Council on Measurement in Education. Standards for Educational and Psychological Testing. American Educational Research Association, 2014. [4] M. Anderljung, J. Barnhart, A. Korinek, J. Leung, C. O’Keefe, J. Whittlestone, S. Avin, M. Brundage, J. Bullock, D. Cass-Beggs, B. Chang, T. Collins, T. Fist, G. Hadfield, A. Hayes, L. Ho, S. Hooker, E. Horvitz, N. Kolt, J. Schuett, Y. Shavit, D. Siddarth, R. Trager, and K. Wolf. Frontier ai regulation: Managing emerging risks to public safety, 2023. [5]T. W. Anderson, T. W. Anderson, T. W. Anderson, T. W. Anderson, and E.-U. Mathématicien. An introduction to multivariate statistical analysis, volume 2. Wiley New York, 1958. [6]A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, J. Batzner, N. Foroutan, C. Schmitz, K. Korgul, H. Batra, O. Deb, E. Beharry, C. Emde, T. Foster, A. Gausen, M. Grandury, S. Han, V. Hofmann, L. Ibrahim, H. Kim, H. R. Kirk, F. Lin, G. K.-M. Liu, L. Luettgau, J. Magomere, J. Rystrøm, A. Sotnikova, Y. Yang, Y. Zhao, A. Bibi, A. Bosselut, R. Clark, A. Cohan, J. N. Foerster, Y. Gal, S. A. Hale, I. D. Raji, C. Summerfield, P. Torr, C. Ududec, L. Rocher, and A. Mahdi. Measuring what matters: Construct validity in large language model benchmarks. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. [7]Y. Bengio, J. Louradour, R. Collobert, and J. Weston. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning, ICML ’09, page 41–48, New York, NY, USA, 2009. Association for Computing Machinery. [8]S. Bhojanam and S. Mehta. Prompt genotyping: Quantifying the evaluation gap between synthetic benchmarks and real LLM performance. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [9]S. L. Blodgett, S. Barocas, H. Daumé I, and H. Wallach. Language (technology) is power: A critical survey of “bias” in NLP. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online, July 2020. Association for Computational Linguistics. [10]S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In C. Zong, F. Xia, W. Li, and R. Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguis- tics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online, Aug. 2021. Association for Computational Linguistics. 10 [11]S. R. Bowman and G. Dahl. What will it take to fix benchmarking in natural language un- derstanding? In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4843–4855, Online, June 2021. Association for Computational Linguistics. [12]D. T. Campbell and D. W. Fiske. Convergent and discriminant validation by the multitrait- multimethod matrix. Psychological Bulletin, 56(2):81–105, 1959. [13]W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. I. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: an open platform for evaluating llms by human preference. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024. [14] Y. Y. Chiu, L. Jiang, B. Y. Lin, C. Y. Park, S. S. Li, S. Ravi, M. Bhatia, M. Antoniak, Y. Tsvetkov, V. Shwartz, and Y. Choi. CulturalBench: A robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through human-AI red-teaming. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), pages 25663–25701, Vienna, Austria, July 2025. Association for Computational Linguistics. [15]L. L. Cook and M. J. Pitoniak, editors. Educational Measurement. Oxford University Press, 5 edition, 2025. [16] I. Covert, W. Ji, T. Hashimoto, and J. Zou. Scaling laws for the value of individual data points in machine learning. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024. [17]L. J. Cronbach and P. E. Meehl. Construct validity in psychological tests. Psychological Bulletin, 52(4):281–302, 1955. 60 references. (PsycInfo Database Record (c) 2025 APA, all rights reserved). [18]M. Dehghani, Y. Tay, A. A. Gritsenko, Z. Zhao, N. Houlsby, F. Diaz, D. Metzler, and O. Vinyals. The benchmark lottery, 2021. [19] ̇ I. E. Deveci and D. Ataman. The ouroboros of benchmarking: Reasoning evaluation in an era of saturation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [20]M. Du, V. Manjunatha, R. Jain, R. Deshpande, F. Dernoncourt, J. Gu, T. Sun, and X. Hu. Towards interpreting and mitigating shortcut learning behavior of NLU models. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cot- terell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 915–929, Online, June 2021. Association for Computational Linguistics. [21]M. Eriksson, E. Purificato, A. Noroozian, J. Vinagre, G. Chaslot, E. Gomez, and D. Fernandez- Llorca. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1):850–864, Oct. 2025. [22]European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L series, 2024. [23]FadillAmir. Benchmarking and standardization of evaluation protocols: A feedback-driven framework using LLM judges to gatekeep and iteratively improve synthetic benchmarks. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [24] C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, and T. Wolf. Open llm leaderboard v2.https: //huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2024. 11 [25]B. Gao, F. Song, Z. Yang, Z. Cai, Y. Miao, Q. Dong, L. Li, C. Ma, L. Chen, R. Xu, Z. Tang, B. Wang, D. Zan, S. Quan, G. Zhang, L. Sha, Y. Zhang, X. Ren, T. Liu, and B. Chang. Omni- MATH: A universal olympiad level mathematic benchmark for large language models. In The Thirteenth International Conference on Learning Representations, 2025. [26] S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In T. Cohn, Y. He, and Y. Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online, Nov. 2020. Association for Computational Linguistics. [27] S. Golchin and M. Surdeanu. Time travel in LLMs: Tracing data contamination in large language models. In The Twelfth International Conference on Learning Representations, 2024. [28]H. Gulliksen. Theory of Mental Tests. Wiley Publications in Psychology. John Wiley & Sons, Hoboken, NJ, 1950. [29]R. K. Hambleton and R. W. Jones. Comparison of classical test theory and item response theory and their applications to test development. Educational Measurement: Issues and Practice, 12(3):38–47, 1993. [30]R. K. Hambleton, H. Swaminathan, and H. J. Rogers. Fundamentals of Item Response Theory, volume 2 of Measurement Methods for the Social Sciences. Sage Publications, Thousand Oaks, CA, 1991. [31] A. Hardy, A. Reuel, K. Jafari Meimandi, L. Soder, A. Griffith, D. M. Asmar, S. Koyejo, M. S. Bernstein, and M. J. Kochenderfer. More than marketing? on the information value of ai benchmarks for practitioners. In Proceedings of the 30th International Conference on Intelligent User Interfaces, IUI ’25, page 1032–1047, New York, NY, USA, 2025. Association for Computing Machinery. [32]D. Heineman, V. Hofmann, I. Magnusson, Y. Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge. Signal and noise: A framework for reducing uncertainty in language model evaluation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [33]D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Mea- suring massive multitask language understanding. In International Conference on Learning Representations, 2021. [34]S. Henrysson.Correction of item-total correlations in item analysis.Psychometrika, 28(2):211–218, 1963. [35]A. Jacovi, A. Caciularu, O. Goldman, and Y. Goldberg. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5075–5084, Singapore, Dec. 2023. Association for Computational Linguistics. [36] H. Jiang, X. Yi, Z. Wei, Z. Xiao, S. Wang, and X. Xie. Raising the bar: Investigating the values of large language models via generative evolving testing. In Forty-second International Conference on Machine Learning, 2025. [37] X. Jiang, D. Chang, and X. Xu. Time waits for no benchmark: Exploring the temporal misalignment between static benchmarks, modern LLMs, and the real world. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [38]G. Kamradt. There are 4 stages in a benchmark lifecycle. X (formerly Twitter) post, Nov 2025. Accessed: 2026-01-16. [39]D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stenetorp, R. Jia, M. Bansal, C. Potts, and A. Williams. Dynabench: Rethinking benchmarking in NLP. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and 12 Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110–4124, Online, June 2021. Association for Computational Linguistics. [40]E. Kim, S. Li, S. Khalil, and H. J. Shin. STAIR-AIG: Optimizing the automated item generation process through human-AI collaboration for critical thinking assessment. In E. Kochmar, B. Alhafni, M. Bexte, J. Burstein, A. Horbach, R. Laarmann-Quante, A. Tack, V. Yaneva, and Z. Yuan, editors, Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 920–930, Vienna, Austria, July 2025. Association for Computational Linguistics. [41]R. Le Bras, S. Swayamdipta, C. Bhagavatula, R. Zellers, M. E. Peters, A. Sabharwal, and Y. Choi. Adversarial filters of dataset biases. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org, 2020. [42]M. Li, H. Jiao, T. Zhou, N. Zhang, S. Peters, and R. W. Lissitz. Item difficulty modeling using fine-tuned small and large language models. In J. Wilson, C. Ormerod, and M. Beiting Parrish, editors, Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 48–55, Wyndham Grand Pittsburgh, Down- town, Pittsburgh, Pennsylvania, United States, Oct. 2025. National Council on Measurement in Education (NCME). [43]T. Li, W.-L. Chiang, E. Frick, L. Dunlap, T. Wu, B. Zhu, J. E. Gonzalez, and I. Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning, 2025. [44]P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Re, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. WANG, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. S. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. A. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda. Holistic evaluation of language models. Transactions on Machine Learning Research, 2023. Featured Certification, Expert Certification, Outstanding Certification. [45]Q. V. Liao and Z. Xiao. Rethinking model evaluation as narrowing the socio-technical gap, 2025. [46]B. Y. Lin, Y. Deng, K. Chandu, F. Brahman, A. Ravichander, V. Pyatkin, N. Dziri, R. L. Bras, and Y. Choi. Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024. [47]F. Lin, S. Xie, Y. Dai, W. Yao, T. Lang, and Y. Zhang. IDGen: Item discrimination induced prompt generation for LLM evaluation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [48]J. Liu, Y. Nam, X. Cui, and S. Swayamdipta. Evaluation under imperfect benchmarks and ratings: A case study in text simplification. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [49]Y. L. Liu, S. L. Blodgett, J. Cheung, Q. V. Liao, A. Olteanu, and Z. Xiao. ECBD: Evidence- centered benchmark design for NLP. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16349–16365, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics. [50]F. M. Lord. Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980. [51]F. M. Lord and M. R. Novick. Statistical Theories of Mental Test Scores. Addison-Wesley, Reading, MA, 1968. 13 [52]I. Magnusson, N. Tai, B. Bogin, D. Heineman, J. D. Hwang, L. Soldaini, A. Bhagia, J. Liu, D. Groeneveld, O. Tafjord, N. A. Smith, P. W. Koh, and J. Dodge. Datadecide: How to predict best pretraining data with small experiments. In Forty-second International Conference on Machine Learning, 2025. [53] N. Muennighoff, N. Tazi, L. Magne, and N. Reimers. MTEB: Massive text embedding bench- mark. In A. Vlachos and I. Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2014–2037, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. [54]O. Nahum, N. Calderon, O. Keller, I. Szpektor, and R. Reichart. Are LLMs better than reported? detecting label errors and mitigating their effect on model performance. In C. Christodoulopou- los, T. Chakraborty, C. Rose, and V. Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26782–26809, Suzhou, China, Nov. 2025. Association for Computational Linguistics. [55]Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela. Adversarial NLI: A new benchmark for natural language understanding. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Com- putational Linguistics, pages 4885–4901, Online, July 2020. Association for Computational Linguistics. [56]S. Ott, A. Barbosa-Silva, K. Blagec, J. Brauner, and M. Samwald. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13(1):6793, 2022. [57]Y. Perlitz, A. Gera, O. Arviv, A. Yehudai, E. Bandel, E. Shnarch, M. Shmueli-Scheuer, and L. Choshen. Benchmark agreement testing done right: A guide for LLM benchmark evaluation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [58] I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes. Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Account- ability, and Transparency, FAT* ’20, page 33–44, New York, NY, USA, 2020. Association for Computing Machinery. [59]M. D. Reckase. 18 multidimensional item response theory. Handbook of statistics, 26:607–642, 2006. [60] D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024. [61]S. Sabour, S. Liu, Z. Zhang, J. Liu, J. Zhou, A. Sunaryo, T. Lee, R. Mihalcea, and M. Huang. EmoBench: Evaluating the emotional intelligence of large language models. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5986–6004, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics. [62]O. E. Salaudeen, A. Reuel, A. M. Ahmed, S. Bedi, Z. Robertson, S. Sundar, B. W. Domingue, A. Wang, and S. Koyejo. Measurement to meaning: A validity-centered framework for AI evaluation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [63] D. Sculley, W. Cukierski, P. Culliton, S. Dane, M. M. Demkin, R. Holbrook, A. Howard, P. T. Mooney, W. Reade, M. Risdal, and N. Keating. Position: AI competitions provide the gold standard for empirical rigor in genAI evaluation. In Forty-second International Conference on Machine Learning Position Paper Track, 2025. [64]H. S. Son. Validity evaluation for the data used for artificial intelligence system. In Y. Bi, R. Bhatia, and S. Kapoor, editors, Intelligent Systems and Applications, pages 362–369, Cham, 2020. Springer International Publishing. 14 [65]S. Swayamdipta, R. Schwartz, N. Lourie, Y. Wang, H. Hajishirzi, N. A. Smith, and Y. Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In B. Webber, T. Cohn, Y. He, and Y. Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online, Nov. 2020. Association for Computational Linguistics. [66]S. T. Truong, Y. Tu, P. Liang, B. Li, and S. Koyejo. Reliable and efficient amortized model-based evaluation. In Forty-second International Conference on Machine Learning, 2025. [67] M. Udell, C. Horn, R. Zadeh, S. Boyd, et al. Generalized low rank models. Foundations and Trends® in Machine Learning, 9(1):1–118, 2016. [68]K. L. Wagstaff. Machine learning that matters. In Proceedings of the 29th International Coference on International Conference on Machine Learning, ICML’12, page 1851–1856, Madison, WI, USA, 2012. Omnipress. [69]A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. SuperGLUE: a stickier benchmark for general-purpose language understanding systems. Curran Associates Inc., Red Hook, NY, USA, 2019. [70]Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. MMLU-pro: A more robust and challenging multi-task language understanding benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. [71]D. Westen and R. Rosenthal. Quantifying construct validity: Two simple measures. Journal of Personality and Social Psychology, 84(3):608–618, 2003. [72]C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Sai- fullah, S. Dey, Shubh-Agrawal, S. S. Sandha, S. V. Naidu, C. Hegde, Y. LeCun, T. Goldstein, W. Neiswanger, and M. Goldblum. Livebench: A challenging, contamination-limited LLM benchmark. In The Thirteenth International Conference on Learning Representations, 2025. [73] S. Wu, H. Bao, S. Li, A. Holtzman, and J. A. Evans. Mapping overlaps in benchmarks through perplexity in the wild, 2025. [74]Z. Xiao, S. Zhang, V. Lai, and Q. Liao. Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [75]S. M. Xie, H. Pham, X. Dong, N. Du, H. Liu, Y. Lu, P. Liang, Q. V. Le, T. Ma, and A. W. Yu. Doremi: Optimizing data mixtures speeds up language model pretraining. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. [76] C. Xu, N. Yan, S. Guan, C. Jin, Y. Mei, Y. Guo, and T. Kechadi. DCR: Quantifying data contam- ination in LLMs evaluation. In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 23002–23020, Suzhou, China, Nov. 2025. Association for Computational Linguistics. [77]X. Xu, Z. Wu, R. Qiao, A. Verma, Y. Shu, J. Wang, X. Niu, Z. He, J. Chen, Z. Zhou, G. K. R. Lau, H. Dao, L. Agussurja, R. H. L. Sim, X. Lin, W. Hu, Z. Dai, P. W. Koh, and B. K. H. Low. Position paper: Data-centric AI in the age of large language models. In Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 11895–11913, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. [78]Z. Xu, S. Xie, Q. Lv, S. Xiao, L. Song, S. Wenjuan, and F. Lin. Diagnosing failures in large language models’ answers: Integrating error attribution into evaluation framework. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 21148–21165, Vienna, Austria, July 2025. Association for Computational Linguistics. 15 [79]J. Yao, P. Jin, K. Bao, Q. Yu, K. Bhardwaj, C. Su, J. Wang, Y. ZHU, S. Devare, D. Mosk- Aoyama, Z. Dong, V. K. Srinivasan, Y. Zhang, O. Kuchaiev, J. Jiao, and B. Zhu. The measure of all measures: Quantifying LLM benchmark quality. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. [80] A. K. Zhang, K. Klyman, Y. Mai, Y. Levine, Y. Zhang, R. Bommasani, and P. Liang. Position: Language model developers should report train-test overlap. In Forty-second International Conference on Machine Learning Position Paper Track, 2025. [81]H. Zhang, Y. Chen, and X. Li. A note on exploratory item factor analysis by singular value decomposition. Psychometrika, 85(2):358–372, 2020. 16 A Limitations The limitations discussed here pertain to how we present and substantiate our position in this paper, rather than to the value or potential of item-level benchmark data itself. First, the empirical illustrations in Sec. 5 draw primarily on classical test theory and item factor analysis. While these well-established psychometric methods effectively demonstrate the unique insights afforded by item-level data, they represent only a subset of the analytical possibilities. Techniques such as item response theory modeling, differential item functioning analysis, and diagnostic classification models remain unexplored in our illustrations. Broader methodological coverage would more comprehensively showcase the range of research practices that item-level data can support. Second, although Secs. 2, 4, & 7 discuss the tension between open data access and data contamination risk, our current treatment of risk mitigation strategies for open-sourcing item-level benchmark data remains relatively conceptual. More concrete safeguards, such as access control mechanisms, usage monitoring, and community-agreed norms for responsible data use, would strengthen the practical viability of our proposal. B Additional Analysis Results Here, we present additional analysis results complementing Sec. 5. Table 3 and Fig. 7 report the results of the GLRM analysis conducted in Sec. 5.2 on MMLU. Fig. 8 shows item clusters from four benchmark datasets in HELM, revealed by K-means clustering over item factor loadings derived from GLRM. Fig. 9 illustrates that items within the same subject or dataset, despite sharing the same label, can emphasize different aspects when their maximum item factor loadings differ. Table 3: Representative items with large absolute item factor loadings and possible constructs for each GLRM factor in MMLU. The top 100 representative items are interpreted and summarized into a candidate label by GPT-5, then manually revised. FactorRepresentative ItemPotential Sub-Construct MMLU #1 Which one of the following statements best describes the algebraicDomain-specific canonical representation of the fitted regression line?framework knowledge Based on the paper “SoK: SSL and HTTPS: Revisiting past Applied synthesis and MMLU #2challenges and evaluating certificates trust model enhancements”, case-based judgment which of the following statements are false? MMLU #3 Find the product of the given polynomials in the given polynomialFormal, quantitative, ring. f(x) = 4x− 5, g(x) = 2x 2 − 4x + 2 inZ 8 [x].multi-step modeling MMLU #4 Which of the following statements is true concerning the populationDomain-specific recall regression function (PRF) and sample regression function (SRF)?and simple reasoning The ( ) is categorized as an unknown segment of the Deep Web Conceptual understanding MMLU #5which has been purposely kept hidden and is inaccessible using and explanation standard web browsers. C OpenEval Schema Fig. 10 displays the data schema of OpenEval, including required and optional fields, data type constraints, and field descriptions. Theresponse,model, andscoreobjects folded in Fig. 10 are detailed in Figs. 11, 12, and 13, respectively. 17 Convergent/Discriminant Validity: MMLU #1 #2 #3 #4 #5 Truth- fulQA GSM- 8K #1#2#3#4#5Truth- fulQA GSM- 8K Figure 7: Convergent/discriminant evidence of the four sub-constructs (#1 - # 5) on MMLU. BabiQA (k=3)MMLU (k=5)MMLU-Pro (k=4) Item Clusters in GLRM Factor Space Omni-MATH (k=4) Figure 8: Clusters from four benchmark datasets in HELM revealed by K-means clustering over item factor loadings from GLRM. Item #9546 The mutual induction of electric and magnetic fields can produce: Item #9676 The line width of a He-Ne laser is 10 3 Hz. The operating wavelength is 6328Å and the power is 1 milliwatt. (a) How many photons are emitted per second? (b) If the output beam is 1 m in diameter... Item #2321 Describe the four major types of conduct disorders and the main characteristics of each. Item #2348 You receive a letter from the Ethics Committee asking for information about a former client who has filed a complaint against her current therapist. You stopped seeing the client over seven years ago, you should: Item #2321 Item #2348 Item #9546 Item #9676 GLRM Factor #1 (Formal, quantitative, multi-step modeling) #2 (Domain-specific recall and simple reasoning) #3 (Conceptual understanding and explanation) #4 (Applied synthesis and case-based judgment) Figure 9: Example items with different maximum factor loadings within the same subject (psychology and physics) in MMLU-PRO. 18 "item_id": "[str | auto] UNIQUE IDENTIFIER FOR THE ITEM", "item_metadata": "ingestion_time": "[str | auto] TIMESTAMP OF WHEN THE ITEM WAS INGESTED (ISO 8601 FORMAT)", "contributor": "name": "[str | optional] NAME OF THE CONTRIBUTOR", "email": "[str | optional] EMAIL ADDRESS OF THE CONTRIBUTOR", "affiliation": "[str | optional] AFFILIATION OF THE CONTRIBUTOR" , "source": "benchmark_name": "[str | non-empty] NAME OF THE BENCHMARK OR DATASET", "benchmark_version": "[str | required] VERSION OF THE BENCHMARK OR DATASET", "paper_url": "[str | optional] URL TO THE BENCHMARK PAPER", "dataset_url": "[str | optional] URL TO THE BENCHMARK DATASET", "benchmark_tags": "[list[str] | optional] TAGS OR KEYWORDS ASSOCIATED WITH THE BENCHMARK OR DATASET" , "item_content": "input": "[list[str,dict] | non-empty] CONTENT OF THE ITEM, SUCH AS A QUESTION, A DIALOGUE, OR MULTIPLE CHOICE OPTIONS", "references": "[list[str,dict] | required] REFERENCES FOR THE EVALUATION METRIC" , "responses": [· ], "schema_version": "[str | auto] OPENEVAL SCHEMA VERSION" Figure 10: OPENEVAL data schema (item). The value of each field specifies the data type, presence requirement, and a short explanation of the field. Specifically,automeans the field is auto-generated, optionalmeans the field can be absent from the data entry,requiredmeans the data contributors are required to specify this field, and non-empty means the field cannot be left blank. "response_id": "[str | auto] UNIQUE IDENTIFIER FOR THE RESPONSE", "model": · , "item_adaptation": "request_input": "[list[str,dict] | non-empty] ACTUAL INPUT PROVIDED TO THE MODEL FOR THIS RESPONSE, SUCH AS THE ADAPTED QUESTION OR DIALOGUE", "demonstrations": "[list[str,dict] | optional] DEMONSTRATIONS OR EXAMPLES PROVIDED TO THE MODEL FOR THIS RESPONSE", "external_resources": [ "type": "[str | required] TYPE OF THE EXTERNAL RESOURCE", "content": "[any | required] AVAILABLE EXTERNAL RESOURCES FROM THE TEST ENVIRONMENT" ] , "response_content": "[list[str,dict] | non-empty] CONTENT OF THE RESPONSE, SUCH AS THE MODEL'S ANSWER OR OUTPUT", "scores": [· ] Figure 11: OPENEVAL data schema (response). Each response object is an element of theresponses field in Fig. 10. 19 "name": "[str | non-empty] NAME OF THE RESPONDENT MODEL", "size": "[str | optional] SIZE OF THE RESPONDENT MODEL", "model_adaptation": "system_instruction": "[str | required] SYSTEM INSTRUCTION OR PROMPT USED TO ADAPT THE ITEM FOR THIS RESPONSE", "generation_parameters": "temperature": "[float | required] TEMPERATURE USED FOR GENERATION", "do_sample": "[bool | required] WHETHER OR NOT TO USE SAMPLING, FALSE MEANS GREEDY DECODING", "top_k": "[int | required] TOP-K VALUE USED FOR GENERATION", "top_p": "[float | required] TOP-P VALUE USED FOR GENERATION", "max_tokens": "[int | required] MAXIMUM NUMBER OF TOKENS GENERATED" , "tools": [ "type": "[str | required] TYPE OF THE TOOL, SUCH AS A SEARCH ENGINE", "content": "[any | required] TOOL AVAILABLE TO THE MODEL" ] Figure 12: OPENEVAL data schema (model). The model object here corresponds to themodelfield in Fig. 11. "metric": "name": "[str | non-empty] NAME OF THE EVALUATION METRIC", "models": "[list[str] | required] MODELS OR ALGORITHMS USED TO COMPUTE THE EVALUATION METRIC", "extra_artifacts": [ "type": "[str | required] TYPE OF THE ARTIFACT", "content": "[any | required] ADDITIONAL ARTIFACT FOR SCORING THE RESPONSE" ] , "value": "[int,float,bool | non-empty] NUMERIC VALUE OF THE EVALUATION METRIC" Figure 13: OPENEVAL data schema (score). Each score object is an element of thescoresfield in Fig. 11. 20