Paper deep dive
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
Hang Li, Kaiqi Yang, Xianxuan Long, Fedor Filippov, Yucheng Chu, Yasemin Copur-Gencturk, Peng He, Cory Miller, Namsoo Shin, Joseph Krajcik, Hui Liu, Jiliang Tang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/21/2026, 2:25:58 AM
Summary
This paper benchmarks uncertainty quantification (UQ) methods for Large Language Models (LLMs) used in automatic assessment. It evaluates various metrics—categorical (Numset, MAR, CE, FSD) and relation-based—across multiple datasets, LLM families, and generation settings to characterize uncertainty patterns and provide insights for developing reliable, uncertainty-aware grading systems.
Entities (10)
Relation Signals (7)
Hang Li → affiliatedwith → Michigan State University
confidence 99% · Hang Li ∗ lihang4@msu.edu Michigan State University
Jiliang Tang → affiliatedwith → Michigan State University
confidence 99% · Jiliang Tang tangjili@msu.edu Michigan State University
LLM → usedfor → Automatic Assessment
confidence 95% · LLM-powered automatic assessment has become increasingly popular
First-Second Distance → typeof → Uncertainty Quantification
confidence 90% · First–Second Distance (FSD). It is a recently proposed frequency-based uncertainty measure
Categorical-Entropy → typeof → Uncertainty Quantification
confidence 90% · Categorical-Entropy (CE). It is a classical distribution chaos quantification method
Numset → typeof → Uncertainty Quantification
confidence 90% · Numset. It measures uncertainty as the number of unique answers
Max-Agree-Rate → typeof → Uncertainty Quantification
confidence 90% · Max-Agree-Rate (MAR). It captures uncertainty by incorporating the full distribution
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptability to diverse question types and flexibility in output formats, they also introduce new challenges related to output uncertainty, stemming from the inherently probabilistic nature of LLMs. Output uncertainty is an inescapable challenge in automatic assessment, as assessment results often play a critical role in informing subsequent pedagogical actions, such as providing feedback to students or guiding instructional decisions. Unreliable or poorly calibrated uncertainty estimates can lead to unstable downstream interventions, potentially disrupting students' learning processes and resulting in unintended negative consequences. To systematically understand this challenge and inform future research, we benchmark a broad range of uncertainty quantification methods in the context of LLM-based automatic assessment. Although the effectiveness of these methods has been demonstrated in many tasks across other domains, their applicability and reliability in educational settings, particularly for automatic grading, remain underexplored. Through comprehensive analyses of uncertainty behaviors across multiple assessment datasets, LLM families, and generation control settings, we characterize the uncertainty patterns exhibited by LLMs in grading scenarios. Based on these findings, we evaluate the strengths and limitations of different uncertainty metrics and analyze the influence of key factors, including model families, assessment tasks, and decoding strategies, on uncertainty estimates. Our study provides actionable insights into the characteristics of uncertainty in LLM-based automatic assessment and lays the groundwork for developing more reliable and effective uncertainty-aware grading systems in the future.
Tags
Links
- Source: https://arxiv.org/abs/2602.16039v1
- Canonical: https://arxiv.org/abs/2602.16039v1
Trouble viewing inline? Open PDF directly →
Full Text
74,041 characters extracted from source content.
Expand or collapse full text
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment Hang Li ∗ lihang4@msu.edu Michigan State University East Lansing, USA Kaiqi Yang ∗ kqyang@msu.edu Michigan State University East Lansing, USA Xianxuan Long longxia2@msu.edu Michigan State University East Lansing, USA Fedor Filippov filipo2@msu.edu Michigan State University East Lansing, USA Yucheng Chu chuyuch2@msu.edu Michigan State University East Lansing, USA Yasemin Copur-Gencturk copurgen@usc.edu University of Southern California Los Angeles, USA Peng He peng.he@wsu.edu Washington State University Pullman, USA Cory Miller mill3118@msu.edu Michigan State University East Lansing, USA Namsoo Shin namsoo@msu.edu Michigan State University East Lansing, USA Joseph Krajcik krajcik@msu.edu Michigan State University East Lansing, USA Hui Liu liuhui7@msu.edu Michigan State University East Lansing, USA Jiliang Tang tangjili@msu.edu Michigan State University East Lansing, USA Abstract The rapid rise of large language models (LLMs) is reshaping the land- scape of automatic assessment in education. Benefiting from strong prior knowledge and advanced reasoning capabilities, LLM-based assessment systems have moved beyond laboratory prototypes used by small groups of researchers and are increasingly becoming prac- tical tools for everyday use by teachers. While these systems demon- strate substantial advantages in adaptability to diverse question types and flexibility in output formats, they also introduce new chal- lenges related to output uncertainty, stemming from the inherently probabilistic nature of LLMs. Output uncertainty is an inescapable challenge in automatic assessment, as assessment results often play a critical role in informing subsequent pedagogical actions, such as providing feedback to students or guiding instructional deci- sions. Unreliable or poorly calibrated uncertainty estimates can lead to unstable downstream interventions, potentially disrupting students’ learning processes and resulting in unintended negative consequences. To systematically understand this challenge and in- form future research, we benchmark a broad range of uncertainty quantification methods in the context of LLM-based automatic as- sessment. Although the effectiveness of these methods has been ∗ Both authors contributed equally to this research. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference acronym ’X, Woodstock, NY © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X demonstrated in many tasks across other domains, their applicabil- ity and reliability in educational settings, particularly for automatic grading, remain underexplored. Through comprehensive analyses of uncertainty behaviors across multiple assessment datasets, LLM families, and generation control settings, we characterize the un- certainty patterns exhibited by LLMs in grading scenarios. Based on these findings, we evaluate the strengths and limitations of dif- ferent uncertainty metrics and analyze the influence of key factors, including model families, assessment tasks, and decoding strategies, on uncertainty estimates. Our study provides actionable insights into the characteristics of uncertainty in LLM-based automatic as- sessment and lays the groundwork for developing more reliable and effective uncertainty-aware grading systems in the future. CCS Concepts • Computing methodologies→Natural language processing; • Applied computing→ Computer-assisted instruction. Keywords Uncertainty Quantification, Large Language Models, Automatic Assessment, Automated Grading ACM Reference Format: Hang Li, Kaiqi Yang, Xianxuan Long, Fedor Filippov, Yucheng Chu, Yasemin Copur-Gencturk, Peng He, Cory Miller, Namsoo Shin, Joseph Krajcik, Hui Liu, and Jiliang Tang. 2018. How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’X). ACM, New York, NY, USA, 13 pages. https: //doi.org/X.X arXiv:2602.16039v1 [cs.AI] 17 Feb 2026 Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. 1 Introduction With the surge of Large language models (LLMs), LLM-powered automatic assessment has become increasingly popular in many do- mains [23,40]. Among the scenarios employing LLMs as assistants, automatic assessment is one of the most practical and challeng- ing tasks. Automated assessment, also called automated grading, is a system that takes questions and candidate solutions, option- ally with rubrics and other context information, and generates text to present the correctness of the evaluated solution [3,20]. The inputs of automated assessment include questions, answers, rubrics that describe how to assess the solutions, and optionally, demonstrations of solution-grade examples. From the output side, assessment results are not limited to values or labels of grades, but also cover aspect-level assessment [46,60], reasons for assess- ments [13,42], etc. In existing work, the automated assessments are supported by rule-based [8,12] or supervised machine learning models [2,47], which require intensive labor from domain experts to design the rules and features, to annotate large-scale datasets for training and validation, and to address anomalies or out-of- distribution cases [3,52]. Owing to LLMs’ broad prior knowledge and instruction-following ability, LLM-based grading is able to deal with diverse question domains, understand rubrics, and produce structured evaluations together with interpretable rationales in nat- ural language [25,32,41]. Compared with traditional task-specific grading systems, LLM-based systems offer a unified pipeline ca- pable of handling diverse inputs without training or adaptation. This flexibility has enabled their usage across domains and problem formulations in practical educational workflows. Although automated assessment is efficient, the uncertainty is- sue has been a central challenge, and this is even more severe in LLM-based assessment. Uncertainty is the metric denoting how sure the models are when making judgments, with the causes of randomness, absence of precise boundaries of concepts, inaccurate base knowledge, and ambiguity between different objects [50]. Un- certainty causes problems of decreased performance and misleading outputs, especially for cases without adequate information or model abilities. This issue is more severe in LLM-based frameworks, as LLMs hide the internal mechanisms in large-scale parameters, mak- ing it hard to probe or control. In addition, LLMs rely on token probability and sampling, which inherently brings randomness into the generation. Lastly, being powerful in text generation, LLMs tend to produce seemingly correct responses even without enough information and capabilities, which further increases the risk of misleading results. In conclusion, certainty is highly prioritized by automated assessment tasks, and responses of low certainty are harmful to both the learning procedures and the equality between students. Uncertainty quantification (UQ) has been studied in various tasks, including classification [9,37], regression [1,38], and Bayesian in- ference settings [50], covering data in text and image modalities. However, the existing methods rely heavily on access to model structure and hidden states, including prediction probabilities rep- resented by neural networks [55], the feasible space of possibility distributions [9], etc. Although showing convincing performances in measuring uncertainty, the existing methods are inapplicable for LLM-based auto-assessment systems, because the internal states of LLMs are either computationally expensive and hard to handle, or even inaccessible when using proprietary LLMs for users. There are also works on uncertainty estimation for LLMs, which primarily include white-box methods that require LLMs to provide logits, internal layer outputs, and the model parameters [69]; consistency- based methods that ground uncertainty on the comparison between multiple generations [75]; and reflection-based methods that elicit LLMs to self-report or analyze the certainty in generation [68]. Because the access to models and settings of the methods is differ- ent, these works are not systematically compared, and there are few insights into how well they perform. As a prosperous research field, LLM-based auto-assessment has developed many works on modeling and implementing uncertainty, while a comprehensive investigation to validate and compare them is still ignored. To bridge this gap, in this work, we propose a benchmark on uncertainty quantification for the auto-assessment tasks. The chal- lenges of applying UQ methods to auto-assessment derive from the specificity in the wide range of domain knowledge, the finite set of ordinal output features, and the influences of input prompts. First, auto-assessment deploys general-use LLMs as a judge in a specialized domain, which requires benchmarking work to con- sider heterogeneous subjects, rather than being limited to domains of LLMs’ expertise. Second, the expected output (i.e., the grade) is an ordinal variable whose labels have a relative order, but the distances between categories are not comparable. This prevents arithmetic operations on the grade labels and makes numeric met- rics (e.g., differences and mean squared error) invalid. Similarly, metrics designed for semantic distance may not adequately capture the differences of outputs, because both the correct and wrong gradings are centered on the semantic representations of question contexts. Third, LLMs are sensitive to the input text, making many UQ methods incomparable as they use different templates along with the question and rubrics. Taking these points in mind, we design the settings of the bench- mark with multiple learning subjects, different forms of output grades, and unified templates to build the input prompts from ques- tions and rubrics. We collect datasets of both open-sourced and private, covering essay writing, natural science, chemistry, and math education. As for the output grades, the grade sets include binary, ordered, multi-class, etc. In addition, we elicit the LLMs to judge with the text of reasons for the grades, incorporating the semantic features into the analysis. Finally, we use a set of curated- designed templates to build the prompts, trying to make the input text similar to prior works or the other datasets in this benchmark. To satisfy the need for broad usage and flexibility, we focus on the repetition-based (or ensemble-based) methods in this work, which take several runs of generation to estimate uncertainty, without using internal states of LLMs in any way. 2 Related Work 2.1 LLM-based Automatic Assessment Automatic assessment has long been studied before the emergence of LLMs. Earlier systems were typically built on task-specific ma- chine learning models trained for particular grading formats, such as automated essay scoring, short-answer evaluation, or rubric- based classification [3,28,52]. These approaches generally relied How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY on engineered linguistic features and supervised neural architec- tures for fixed domains or tasks. While effective in narrow settings, they often require annotated datasets for each task, making it costly to scale assessment systems to new contexts. Besides, they require rigorous prompt tuning to align the assessment with fine-grained evaluation criteria, and this limits the transferability across subjects, question types, and scoring criteria. The advent of large language models has shifted this paradigm toward general-purpose evaluators that can assess with instructions and rubrics. A representative framework is LLM-as-a-Judge [23], where a prompted model assigns correctness labels [21,45,81] or reviews written works [7,29]. The LLM-as-a-Judge framework is widely used to handle diverse tasks, including mathematics [48], natural and medical science [34,78], programming [36], and lan- guage [61]. Recent studies demonstrate the growing capability of LLM-based assessment systems. For example, LLM evaluators can produce stable grading judgments across diverse prompts and reasoning settings [53]; rubric-guided prompting enables mod- els to perform fine-grained scoring aligned with evaluation cri- teria [18]; aggregation frameworks such as SURE [33] leverage self-consistency and majority voting to improve grading reliability. Along with prompt engineering work, there are also studies on how to refine rubrics to achieve fine-grained grades [11,59]. How- ever, these studies also reveal remaining limitations: evaluation outcomes may change under minor input perturbations, includ- ing the replacement of synonyms and meaning-invariant rephras- ing [26,53,80], symbols indicating grade labels [74], and the order of provided grade labels [43,56]. These findings suggest that, al- though LLM-based assessment improves coverage and flexibility, existing methods remain vulnerable to undesired variation in gener- ated judgments. This limitation highlights the need for uncertainty estimation to complement assessment decisions and mitigate unre- liable outputs. 2.2 Uncertainty Estimation in LLMs With the surge of LLM applications, inaccurate or non-factual gener- ations have motivated uncertainty estimation as a practical signal of when outputs should (or should not) be trusted. A common starting point is to derive confidence from observable generation behavior, especially via repeated sampling. [14] shows that repetition-based metrics computed from multiple samples provide a reliable criterion for abstention decisions against uncertain responses, outperforming likelihood-based and self-verification alternatives. [49] formalizes sample consistency into confidence measures based on agreement or entropy, and demonstrates their effectiveness for confidence cali- bration across models and reasoning tasks. This work also explores how factors such as sampling size and explanation prompting influ- ence calibration quality. In addition, [73] proposes self-probing, a method that elicits verbalized confidence from the LLMs, and finds that the verbalized confidence score helps with calibration and fail- ure prediction, while also presenting overconfident statements like humans. Complementing the method-focused studies, [30] conducts a large-scale empirical comparison with a broad set of uncertainty estimators, LLMs, and domain tasks, claiming that uncertainty sig- nals can help surface risky generations while also exposing practical limitations across settings. Beyond focusing on the match of surface features, more sophisticated approaches refine the notion of equiv- alence used to aggregate samples. For open-ended generation, [35] introduces semantic entropy, which clusters responses by meaning invariance and quantifies uncertainty over semantic modes rather than token sequences, thereby addressing the mismatch between lexical diversity and true epistemic uncertainty. Recently, [71] syn- thesizes inference-time uncertainty estimation for LLMs by defining key uncertainty sources (e.g., incomplete information and model limitations) and organizing methods into a taxonomy that includes verbalized confidence, latent-information, consistency-based, and semantic clustering approaches. In addition, the ordinal characteristic of grade labels makes the assessment task different from other tasks. The grades are presented with a label or value that belongs to a finite set in most cases; for example, in SemEval 2013 - Task 7 [19] where the grade labels indicate whether and for which degree student answer aligns with the given reference, the scores are:3for fully correct,2for partially correct,1for contradictory with reference, and0for irrelevant. Although the labels are inherently ordered in a sequence where higher scores indicate a stronger tie with the reference, the distance between labels is incomparable or even nonsensical [67,72]. This body of work motivates our benchmark design. When evaluating these uncertainty signals in automated assessment, we focus on grading tasks whose predictions lie in a constrained label space. 3 Method 3.1 Uncertainty Definition As described in Section 1, we focus on uncertainty estimation approaches based on repeated generation, as such methods can be applied uniformly to both open-source and proprietary LLMs without requiring access to internal model states. Following prior work [35,69], uncertainty is defined as the variability in a model’s predictions. In the context of LLM-based automatic assessment, uncertainty can be operationalized as the variability of grading outputsOproduced for a given input triplet푥=(푞,푟,푎), consisting of the question, grading rubric, and student answer. Formally, the uncertainty associated with 푥 is defined as 푈= 푓(O), O=표 푖 푁 푖=1 , 표 푖 =푔 휃 (푥,푐)(1) where푓(·)denotes an uncertainty quantification function,푔 휃 is the grading model parameterized by휃,표 푖 denotes the grading out- put from the i-th repeated generation,푁is the repetition times, and푐represents additional contextual information incorporated by recent inference-enhancement techniques, such as few-shot demonstrations [77] and external knowledge retrieved via retrieval- augmented generation (RAG) [10]. 3.2 Uncertainty Quantification Methods Existing studies on uncertainty quantification have primarily fo- cused on open-ended question answering (Q&A) or reasoning tasks [71], where model responses are free-form. Accordingly, many existing methods characterize uncertainty through measures of se- mantic equivalence or variability across generated responses. In this work, we shift the focus to uncertainty quantification in the context of automatic assessment. Unlike open-ended generation, automatic assessment is typically formulated as a classification or Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. scoring task, in which model outputs are constrained to a fixed label or score space. Consequently, we select uncertainty quantification methods [35,49,71] that are suitable for capturing uncertainty from a categorical perspective. Meanwhile, with the recent adop- tion of chain-of-thought (CoT) prompting [70] and the emergence of reasoning-oriented models (e.g., GPT-o1), model outputs in auto- matic assessment increasingly include not only final grading scores but also intermediate grading rationales. To capture uncertainty manifested in these intermediate representations, we additionally incorporate representation-based uncertainty quantification meth- ods into our study. In the following sections, we present the uncer- tainty quantification methods in each category in detail. 3.2.1 Categorical Based Methods. Owing to the pre-defined and discrete output label space, methods in this category are relatively simple to compute and computationally efficient. In general, they derive uncertainty from the empirical frequency distribution of score labels observed across repeated queries to the same request. Numset [49]. It measures uncertainty as the number of unique answers observed across multiple samples: 푈 Numset =|O 푢 |(2) whereO 푢 denotes the set of unique answers obtained from re- peated queries to the same input푥. Intuitively, a larger|O 푢 |indi- cates greater variability in the model’s outputs and thus higher uncertainty. Max-Agree-Rate (MAR) [73]. It captures uncertainty by incor- porating the full distribution of answer frequencies. Compared to Numset, which only counts the number of distinct answers, the MAR reflects the degree of dominance of the most frequent an- swer among all sampled responses. A larger MAR indicates that the model consistently produces the same answer and is therefore more confident. To express this as an uncertainty measure, we define the MAR as 푈 MAR = 1− max 표∈O 푢 1 푁 푁 ∑︁ 푖=1 1[표 푖 =표],(3) Categorical-Entropy (CE) [49]. It is a classical distribution chaos quantification method over the output class probabilities for classification problems [17]. As the automatic assessment is also commonly treated as the classification task, we adopt the CE as the uncertainty evaluation method for the automatic assessment. Com- pared to the MAR, CE not only captures the value distribution of the majority class but also accounts for all categories, ensuring that shifts in the distribution of minority responses are reflected in the computed value. Following the common definition in classification, we define entropy-based consistency푈 CE as: 푈 CE = ∑︁ 표∈O 푢 −푝 표 log(푝 표 ), 푝 표 = 1 푁 푁 ∑︁ 푖 1[표 푖 =표](4) where푝 표 is the normalized frequency of each unique answer표 ∈ O 푢 . Commonly, a higher푈 entropy indicates the more chaos existing in the distributions of the answers, which reflects a high uncertainty. First–Second Distance (FSD) [49]. It is a recently proposed frequency-based uncertainty measure that aims to address limita- tions of commonly used methods such as MAR and Entropy. FSD quantifies the gap between the proportions of samples supporting the most frequent (majority) answer ̄ 푎and the second most frequent answer ̄ ̄ 푎. It is defined as: 푈 FSD = 1− 1 푁 푁 ∑︁ 푖=1 1[표 푖 = ̄ 표]− 푁 ∑︁ 푖=1 1[표 푖 = ̄ ̄ 표] ! (5) Compared with the two prior methods, FSD mitigates the short- comings of Entropy, which can be overly influenced by low-frequency tail categories, and of MAR, which depends solely on the majority answer and discards information about competing alternatives. By focusing on the relative dominance between the top two categories in the frequency distribution, FSD is particularly informative in cases where the model is uncertain between the two most plausible answers. In such scenarios, an FSD-based measure avoids over- confidence induced by considering only the most-voted answer. Intuitively, a larger gap between the top two frequencies indicates higher confidence in the majority answer. To align this measure with our uncertainty formulation (where larger values correspond to higher uncertainty), we use the complement of this gap, as de- fined above, as the uncertainty score. 3.2.2 Relation Based Methods. Although evaluating uncertainty directly in the categorical label space is computationally conve- nient, it has notable limitations. Due to the discrete nature of label categories in automatic assessment tasks, estimated uncertainty values can exhibit large fluctuations when the number of sam- pled responses푁is small. For example, entropy-based measures (e.g., CE) are particularly sensitive to sampling noise before the response distribution stabilizes [16]. In practical automatic assess- ment settings, repeatedly querying an LLM for the same input푥 is constrained by efficiency and cost considerations, making such instability especially problematic. Fortunately, the recent emer- gence of reasoning-path generation in large language models pro- vides rich intermediate signals that help mitigate the limitations of category-based uncertainty quantification. Specifically, by model- ing semantic similarity [54] or logical entailment [35] among the reasoning traces associated with grading outputs (표 푖 and표 푗 ), and constructing relation graphs in which these quantified relationships serve as edge weights, relation-based uncertainty quantification methods enable more fine-grained and continuous assessments of model uncertainty. Existing relation-based approaches typically follow a common two-step paradigm: (1) constructing a relation graph over the set of model outputs, and (2) computing uncertainty metrics based on properties of the resulting graph. In the following, we describe these two steps in detail and present representative implementations. Relation Graph. To evaluate the uncertainty of a model’s re- sponses, we conceptualize the problem as analyzing the “tightness” of relationships among outputs표within a generated grading re- sponse setO. Intuitively, a more compact and coherent group of answers indicates lower uncertainty, whereas a more dispersed group suggests higher uncertainty. Inspired by graph-based clus- tering analysis, recent studies [30,71] construct a relation graph How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY G= (푉,퐸)over the outputsO, where each vertex푣 ∈ 푉corre- sponds to a grading response in표 ∈ O, and each edge푒 푖푗 ∈ 퐸is weighted by the pairwise distance between responses표 푖 and표 푗 . A key factor influencing the analysis lies in how this relation graph is constructed: different definitions of pairwise relationships yield dif- ferent notions of answer similarity and capture uncertainty at vary- ing semantic depths and computational costs. In general, existing relation-based uncertainty quantification methods primarily differ in their choice of pairwise distance or similarity function. Below, we summarize three mainstream and representative approaches that we adopt for subsequent evaluation. • Jaccard Similarity measures lexical overlap between two an- swers based on the Jaccard distance between their token sets. Specifically, given two responses푎 푖 and푎 푗 , the Jaccard similarity (푠 푗푎푐 푖푗 ) is defined as the ratio between the size of the intersection and the size of the union of their token sets. 푠 jac 푖푗 = |푇(표 푖 )∩푇(표 푗 )| |푇(표 푖 )∪푇(표 푗 )| ,(6) where푇(푎)denotes the set of tokens in answer푎. This approach captures surface-level similarity and is computationally efficient, but it is limited in its ability to reflect deeper semantic equivalence between paraphrased or lexically diverse answers. •Embedding Cosine Similarity is defined based on the cosine similarity푠 emb 푖푗 between vector representations of answers ob- tained from a sentence embedding model: 푠 emb 푖푗 = h ⊤ 푖 h 푗 ∥h 푖 ∥ 2 ∥h 푗 ∥ 2 , h 푖 = 푓 emb (표 푖 ),(7) where h 푖 ∈R 푑 denotes the semantic embedding of answer푎 푖 produced by a pre-trained encoder푓 emb . By operating in a con- tinuous semantic embedding space, embedding-based similarity captures higher-level semantic relatedness beyond exact lexical overlap. Compared with Jaccard Similarity, this approach pro- vides a more robust notion of answer similarity, at the cost of additional computational overhead for embedding extraction. •Entailment Score models pairwise relationships using natural language inference (NLI) by estimating the degree to which one answer entails another. Off-the-shelf NLI models output class probabilities over entailment, neutral, and contradiction [39]. However, the limited input context window makes them unsuit- able for directly processing long, multi-step reasoning traces produced by LLMs. To address this, we adopt a BERTScore- inspired strategy [76]: we split each answer into sentences and compute entailment scores over all sentence pairs, followed by a mean–max aggregation. Following prior work [35], we use the entailment probability to construct the relation graph. Since en- tailment is directional, we symmetrize the pairwise relationship by averaging both directions: 푠 nli 푖→푗 = 1 푀 푀 ∑︁ 푚=1 max 푘 푃(entail | 표 푖,푚 ,표 푗,푘 ), 푠 nli 푖푗 = 1 2 푠 nli 푖→푗 +푠 nli 푗→푖 (8) where푃(·)denotes the entailment probability predicted by the NLI model, M and K are the numbers of sentences after split- ting answers표 푖 and표 푗 , respectively, and표 푖,푚 denotes the m-th sentence of answer표 푖 ,표 푗,푘 denotes the k-th sentence of answer 표 푗 . Entailment scores capture fine-grained semantic and logical consistency between responses and are well-suited for identify- ing subtle disagreements in reasoning or factual content. While this method provides the most semantically grounded notion of relational tightness, it is also the most computationally expensive due to the need to run NLI inference over all sentence pairs. Property-driven Uncertainty. Given the relation graphG, dif- ferent graph-theoretic properties can be leveraged to quantify its tightness. Below, we describe four representative property-driven uncertainty measures that have been widely adopted in prior work on uncertainty quantification for Q&A and reasoning tasks [71]. •Normalized Average Degree (NAD), also known as the varia- tional ratio (VR) [30], measures the connectivity density of the relation graph. Intuitively, if each answer is closely connected to many others, the response set forms a dense and compact cluster, indicating low uncertainty. Formally, we define the NAD-based uncertainty as 푈 NAD = 1− 1 푁(푁 − 1) 푁 ∑︁ 푖=1 ∑︁ 푖≠푗 푠 ∗ 푖푗 (9) where푠 ∗ 푖푗 denotes the pairwise similarity between answers푎 푖 and푎 푗 introduced in the previous section. Larger values of푈 NAD indicate weaker local connectivity and thus higher uncertainty. •Graph Eccentricity (GE) quantifies the maximum shortest- path distance from each node to all other nodes in the graph, and the overall tightness is summarized by aggregating node eccentricities. We define the eccentricity-based uncertainty as 푈 ECC = 1 푁 ∑︁ 푖 max 푗 [푑 sp 푖푗 ],(10) where푑 sp denotes the all-pairs shortest-path distance induced by the complement of similarity, 1−푠 ∗ 푖푗 . Smaller values indicate that all answers are mutually close in the relation graph, forming a compact cluster, whereas larger values imply that some answers are distant from the rest, reflecting higher uncertainty. GE thus characterizes the global spread of the response set. • Graph Eigenvalues (Eigen) leverage spectral properties of the graph Laplacian to provide a global measure of connectivity. In particular, the second smallest eigenvalue,휆 2 (퐿), known as algebraic connectivity, reflects how tightly connected the graph is. We define the eigenvalue-based uncertainty as 푈 Eigen = 1 휆 2 (퐿) , 퐿= 퐷− 퐴, 퐴 푖푗 = 푠 ∗ 푖푗 (11) where퐿is the graph Laplacian constructed from the degree matrix퐷and adjacency matrix퐴. A more tightly connected graph exhibits larger휆 2 (퐿)and thus lower uncertainty, whereas smaller 휆 2 (퐿) corresponds to higher uncertainty. •Discrete Semantic Entropy (DSE) extends entropy-based un- certainty measures to the relation-graph setting by computing entropy over semantically clustered responses [35]. By grouping answers into equivalence classes based on pairwise relations inG and computing the entropy of the induced discrete distribution, DSE captures both the diversity of semantic modes and their relative frequencies. Specifically, we define Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. 푈 DSE =− 푀 ∑︁ 푚=1 푝 푚 log푝 푚 , 푝 푚 = |C 푚 | 푁 (12) whereC 푚 푀 푚=1 denotes the clusters obtained from the relation graphG. As clustering overGincurs additional computational cost, we follow prior work [35] and compute this metric only for the NLI-based relation graph, using bidirectional entailment relations. Here,|퐶 푚 |is the size of clusters. Larger푈 DSE indicates a more fragmented semantic landscape and thus higher uncer- tainty. 4 Evaluation 4.1 Metric 4.1.1 Effectiveness. As introduced in Section 3.1, uncertainty val- ues are designed to indicate the confidence or reliability of responses produced by LLMs. In practice, we expect responses with lower uncertainty to exhibit higher accuracy than those with higher uncer- tainty. In the context of automatic assessment, this property enables uncertainty to serve as a signal for human intervention, preventing low-reliability grading results from being released for downstream use. Based on this premise, we evaluate the effectiveness of each uncertainty quantification method as an accuracy indicator using the following four metrics: •AUROC [71] evaluates the discriminative ability of uncertainty scores to separate correct and incorrect predictions. Specifically, by treating incorrect responses as positive cases and uncertainty scores as ranking signals, AUROC measures the probability that a randomly chosen incorrect response is assigned a higher uncer- tainty than a randomly chosen correct response. Higher AUROC indicates better discrimination between reliable and unreliable predictions. •C-Index [62] is a rank-based metric commonly used in survival analysis to assess the consistency between predicted risk scores and observed outcomes. In our setting, where assessment scores are ordinal rather than binary, we use the C-index to evaluate whether responses with larger true errors tend to receive higher uncertainty scores. This extends AUROC to continuous or ordi- nal error magnitudes by measuring the fraction of concordant pairs between uncertainty and error rankings. Higher C-index indicates better alignment between uncertainty and error sever- ity. •AUARC [71] measures the trade-off between prediction accu- racy and rejection rate when selectively abstaining from low- confidence responses. The accuracy–rejection curve plots the accuracy of retained predictions as a function of the fraction of responses rejected based on uncertainty. AUARC summarizes this curve as a single scalar, where higher values indicate that rejecting a small fraction of high-uncertainty responses yields substantial gains in accuracy. This metric reflects the practical utility of uncertainty for selective prediction and human-in-the- loop workflows. •AUERC [57] is the error-based counterpart of AUARC, where the error–rejection curve plots the mean absolute error of re- tained predictions against the rejection rate. This metric evalu- ates whether high-uncertainty responses indeed correspond to larger errors. Unlike AUARC, lower AUERC values are better, as they indicate that rejecting uncertain responses effectively reduces the remaining prediction error. AUERC is particularly suitable when assessment outputs are ordinal or continuous scores rather than binary correctness labels. 4.1.2Stability. All uncertainty quantification methods in this work are derived from stochastic responses generated by LLMs. As a re- sult, the estimated uncertainty values are inherently noisy due to the randomness of the generation process. Although increasing the number of sampled responses can reduce this variance, repeatedly generating a large number of responses for the same grading re- quest is computationally expensive and impractical in real-world deployment. To study this trade-off, we analyze how uncertainty estimates evolve as the response set size n increases, and measure the rate at which the uncertainty values stabilize. Beyond the ab- solute values, uncertainty is primarily used for ranking responses (e.g., flagging the top fraction of high-uncertainty cases for hu- man review). Therefore, we additionally evaluate the stability of uncertainty-induced rankings using Spearman’s rank correlation coefficient (Spearmanr) [6] between estimates computed from dif- ferent response set sizes. Concretely, given a sequence of푁sampled responses, we compute uncertainty values incrementally using the first푘responses for푘=2,3, . . .,푁, and measure (i) the relative change in uncertainty values between successive steps (from푘to 푘+1), and (i) the Spearman rank correlation between the corre- sponding uncertainty rankings. The final stability score is obtained by averaging these stepwise change ratios and rank correlations across all valid steps. Lower change ratios and higher Spearman correlations indicate more stable uncertainty estimates. 4.1.3 Correlation. Although different uncertainty quantification methods assess response uncertainty from diverse perspectives, cor- relations among these methods often arise due to shared underlying evidence sources (e.g., categorical frequency vs. relational similar- ity). To provide an overview of the relationships among methods and to help future work avoid selecting redundant measures, we analyze pairwise correlations between methods using the Pearson correlation coefficient. Table 1: List of questions evaluated by our benchmark. Question Size ScoreSubjectQuestion Size ScoreSubject ASAP5000-3Essay WritingSemveal5000-3Mixture T1-DCI1240-3ChemistryT1-SEP1240-3Chemistry T2-DCI1410-3ChemistryT2-SEP1410-3Chemistry T1-CK2080-3Math PedagogyT1-PCK1880-3Math Pedagogy T2-CK2020-3Math PedagogyT2-PCK1880-3Math Pedagogy U41Q11590-4Earth ScienceU41Q21840-4Earth Science U42Q11980-4Earth ScienceU42Q22090-4Earth Science 4.2 Evaluation Setting 4.2.1 Dataset. To provide a comprehensive evaluation, we col- lect 14 grading questions from five different sources for our ex- periments, including two open-source datasets and three private datasets. Specifically, the two open-source datasets are the Auto- mated Student Assessment Prize (ASAP) dataset [24], which fo- cuses on short-essay grading (150–550 words) written by students How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY Table 2: List of LLMs evaluated by our benchmark. Proprietary Models ModelModel IDScale Calude Haiku 4.5[5]claude-haiku-4-5-20251001/ Calude Sonnet 4.5[4]claude-sonnet-4-5-20250929/ Gemini 2.5 Pro [63]gemini-2.5-pro/ Gemini 3 Flash [63]gemini-3-flash-preview/ GPT-4o mini [51]gpt-4o-mini/ GPT-5 [58]gpt-5/ GPT-5 nano [58]gpt-5-nano/ Open-source Models ModelModel IDScale Qwen3 4B Instruct [64]Qwen/Qwen3-4B-Instruct-25074B Stable LM 2 1.6B Chat [65]stabilityai/stablelm-2-1-6b-chat6B Llama 3.1 8B Instruct [22]meta-llama/Llama-3.1-8B-Instruct8B Falcon3 10B Instruct [66]tiiuae/Falcon3-10B-Instruct10B Ministral 3 14B Instruct [44]mistralai/Ministral-3-14B-Instruct-251214B DeepSeek R1 Distill - Qwen 32B [? ]deepseek-ai/DeepSeek-R1-Distill-Qwen-32B 33B Mixtral 8x7B Instructmistralai/Mixtral-8x7B-Instruct-v0.147B in Grades 7–10, and SemEval-2013 Task 7 (SemEval-13-T7) [19], which scores student responses to questions collected from in-class exercises, tests, and tutorial dialogues. For the three private datasets, we manually collect two sets of questions from elementary and middle school classes covering chemistry and general science fol- lowing the 3DLP standard [27], and one set from adult learners focusing on pedagogical skills [15]. For the open-source datasets, we use the original ground-truth labels to evaluate the correct- ness of LLM-based automatic grading. For the private datasets, the ground-truth labels are obtained via agreement between two expert human graders. In cases of disagreement, a third grader adjudicates to determine the final label. Based on these labels, we compute the correctness of the automatic assessment results. In Table 1, we summarize the statistics and characteristics of each dataset and the corresponding grading questions. For each question, we randomly sample 60–150 answers covering all score categories for evaluation. 4.2.2 Large Language Models. To ensure the robustness of our conclusions, we benchmark a comprehensive set of mainstream LLMs, including seven closed-source API-based models from three leading vendors, i.e., OpenAI, Anthropic, and Google, and seven open-source models released by research teams worldwide. For the closed-source models, we include multiple model sizes within the same model families to assess scaling effects. For the open-source models, we cover recent representative variants of standard LLMs, including mixture-of-experts models [79] and explicitly reasoning- oriented models [31]. With this broad model coverage, we aim to ensure that our benchmark results generalize across diverse deployment scenarios of automatic assessment. Table 2 summarizes the details of all evaluated models. 4.2.3 Generation Settings. To ensure broad coverage of uncer- tainty quantification methods under realistic automatic assessment settings, we implement three widely used generation strategies adopted in recent automatic assessment studies [11]: •zero-shot is the most straightforward way to leverage LLMs for automatic assessment. Given a question, a student answer, and a grading rubric, the LLM is prompted to directly produce a score without any demonstrations or intermediate reasoning steps. •zero-shot + Chain of Thought (COT) extends zero-shot by eliciting intermediate reasoning steps from the LLM. By decom- posing the grading process into a step-by-step reasoning trajec- tory, the model is better able to align with the grading criteria specified in the rubric and to analyze student responses more comprehensively, often resulting in improved grading perfor- mance. •few-shot + Chain of Thought (COT) further augments the prompt with exemplar answers corresponding to different score levels. These demonstrations provide concrete interpretations of the abstract criteria in the rubric, helping the model better internalize grading standards and produce more accurate and consistent assessment outcomes. Specifically, in our evaluation, we include two demonstration examples for each score level. Finally, to balance computational cost and practical deployment constraints, we set the number of repeated grading samples per answer to five. 4.3 Effectiveness Result 4.3.1Overall Result. Due to variations in the number of score lev- els and the distribution of responses across questions, the scale of uncertainty values produced by the same method can differ sub- stantially across datasets and tasks. To avoid dominance by large- magnitude uncertainty values for specific questions, we rank each uncertainty method within each configuration (model, dataset, and generation strategy) and then aggregate results by averaging ranks across settings. Table 3 summarize the aggregated rankings of all uncertainty quantification methods in the four evaluation metrics. We observe that the CE method consistently attains the top rank across different questions, models, and generation strategies. This finding contrasts with prior uncertainty quantification studies that focus primarily on open-ended Q&A tasks [71]. In addition, other categorical-based methods, including MAR, FSD, and Numset, con- sistently outperform relation-based methods. This suggests that the discrete output space characteristic of most automatic assessment tasks makes categorical uncertainty measures more compatible with practical deployment. Within the group of relation-based methods, the average ranking follows the order JS < NLI < Embed. This in- dicates that, despite its simplicity, Jaccard similarity ( JS) effectively captures surface-level agreement among responses. In contrast, although embedding- and NLI-based methods leverage more so- phisticated pretrained encoders for sentence-pair comparison, they may require additional adaptation or task-specific calibration to achieve optimal performance in uncertainty quantification. Notably, this observation differs from trends reported in Q&A-centric bench- marks, where more advanced relation graph constructions (e.g., NLI-based similarity) often demonstrate superior performance. 4.3.2 Perspective-specific Discussion. In this section, we examine the robustness of our previously drawn conclusions under varia- tions in questions, models, and generation strategies. Specifically, we compute average ranks for each uncertainty quantification method under each evaluation metric and summarize their distribu- tions across different models, questions, and generation strategies. Figure 1 compares these distributions across the three perspectives. From the figure, we make the following observations. First, the vari- ability of method rankings differs across perspectives: generation Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. (a) Models. (b) Questions. (c) Generation strategies. Figure 1: Distribution of various in-group average rank for each model, question and generation strategies variants. Table 3: Average ranks of methods across questions, datasets, and generation strategies on effectiveness metrics. MethodAUROC C-index AUARC AUERC CE4.715.444.915.50 FSD5.205.675.245.61 MAR4.825.505.015.51 Numset5.866.555.756.32 JS_NAD6.976.697.036.78 JS_GE8.007.408.047.56 JS_Eigen7.877.427.847.47 NLI_NAD7.667.747.908.15 NLI_GE8.188.108.178.22 NLI_Eigen8.238.108.218.34 NLI_DSE9.959.709.189.08 Embed_NAD8.328.158.518.23 Embed_GE8.958.629.008.57 Embed_Eigen9.038.699.078.61 strategies exhibit the smallest variance, followed by models, while questions show the largest variance. This indicates that method rankings are generally stable across prompting strategies, exhibit moderate variation across models, and are most sensitive to the specific questions being graded. Consequently, while a method may perform consistently under different controls, the optimal choice of uncertainty metric may depend on the particular task or ques- tion context. Second, when comparing across methods, categorical- based approaches (e.g., CE, MAR, FSD, and Numset) consistently achieve higher average ranks, with limited overlap with relation- based methods across perspectives. This suggests that, although categorical methods may not always be optimal for specific models or questions, they serve as a strong and reliable default choice in most scenarios. Finally, within the group of relation-based methods, the embedding-based approach exhibits relatively larger variance across perspectives than other relation-based methods on metrics such as AUROC and C-index. This implies that pretrained sentence embeddings are more sensitive to external factors (e.g., dataset char- acteristics and prompting strategies) and may require more careful calibration or task-specific adaptation when applied to uncertainty quantification. 4.3.3Case Studies. To provide the specific evidence to support our conclusions in section 4.3.2 we conduct case studies on the two open- source datasets ASAP and SemEval. We further select two repre- sentative models, stablelm-2-1-6b-chat and gemini-3-flash-preview, which exhibit the lowest and highest average grading performance across questions, respectively, and evaluate them under the best- performing generation strategy (few-shot + CoT). Specifically, we How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY (a) stablelm-2-1-6b-chat (AUROC)(b) stablelm-2-1-6b-chat (AUARC)(c) stablelm-2-1-6b-chat (AUERC) (d) gemini-3-flash (AUROC)(e) gemini-3-flash (AUARC)(f) gemini-3-flash (AUERC) Figure 2: Method Comparison between stablelm and gemini over different metrics on the ASAP question. (a) stablelm-2-1-6b-chat (AUROC)(b) stablelm-2-1-6b-chat (AUARC)(c) stablelm-2-1-6b-chat (AUERC) (d) gemini-3-flash (AUROC)(e) gemini-3-flash (AUARC)(f) gemini-3-flash (AUERC) Figure 3: Method Comparison between stablelm and gemini over different metrics on the SemEval question. illustrate the ROC, accuracy–rejection (ARC), and error–rejection (ERC) curves for both models on ASAP (Figure 2) and SemEval (Fig- ure 3). By comparing AUROC, AUARC, and AUERC values across models and datasets, we observe that the effectiveness of uncer- tainty methods depends on the underlying model. For instance, categorical-based methods consistently outperform relation-based methods for stablelm-2-1-6b-chat, whereas this trend is reversed for gemini-3-flash-preview. This observation aligns with our findings in the previous section that method rankings are influenced by factors such as model quality and dataset characteristics. Further inspec- tion of the generated grading responses reveals that higher-quality outputs from gemini-3-flash-preview lead to more semantically coherent and informative rationales, making relation-based uncer- tainty measures more reliable. In contrast, noisier and less accurate responses from lower-performing models limit the effectiveness of relational analysis, favoring categorical-based uncertainty mea- sures. Accordingly, we recommend prioritizing categorical-based methods when model performance is low, and shifting to relation- based methods once grading accuracy surpasses a reasonable thresh- old. Finally, comparing method rankings for the same model across the two datasets shows that lower-performing models exhibit rel- atively stable gains from categorical-based methods, whereas for higher-performing models, the effectiveness of relation-based meth- ods is more sensitive to how the relation graph is constructed. This Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. Table 4: Average ranks of methods across questions, datasets, and generation strategies on stability metrics. Method Delta SpearmanrMethodDelta Spearmanr CE5.304.51NLI_NAD8.507.01 FSD7.696.58NLI_GE8.059.93 MAR8.655.37NLI_Eigen12.1010.56 Numset2.476.26 NLI_DSE13.5712.84 JS_NAD2.843.86Embed_NAD6.004.92 JS_GE1.897.23 Embed_GE5.027.89 JS_Eigen10.917.37Embed_Eigen11.958.20 suggests that, when working with high-performing LLMs, careful design of the relation graph construction strategy is crucial for the effectiveness of relational uncertainty quantification. 4.4 Stability Result 4.4.1 General Result. Table 4 reports the stability evaluation re- sults across questions, models, and generation strategies. Consis- tent with the effectiveness analysis, we compare methods based on their aggregated ranks over the change-ratio and Spearman corre- lation metrics. In contrast to effectiveness results, categorical-based methods do not consistently dominate relation-based methods in terms of stability. Notably, relation-based approaches grounded in graph properties—such as NAD- and GE-based variants (e.g., JS_NAD and JS_GE)—exhibit more stable uncertainty estimates as the number of sampled responses increases. It suggests that while categorical methods may achieve stronger overall effectiveness in ranking performance, certain relation-based methods provide su- perior robustness with respect to sampling variability. Meanwhile, we observe that the categorical method CE, although not the most stable, performs competitively and remains close to the top-tier relation-based methods in terms of stability. 4.4.2 Perspective-specific Discussion. Similar to Section 4.3.2, we analyze the robustness of our conclusions regarding stability rank- ings under variations in three factors: model, question, and genera- tion strategy. Figure 4 presents the distribution of stability ranks for each method across these factors. From the figure, we observe that, in contrast to the four effectiveness metrics, stability metrics are relatively insensitive to variations in questions and genera- tion strategies. However, model choice significantly influences the relative stability rankings of different methods. This observation indicates that stability should be re-evaluated when applying uncer- tainty quantification methods to new models, as method robustness does not necessarily transfer across model families. 4.5 Correlation Result 4.5.1General Result. To analyze correlations among methods, we compute the Pearson correlation coefficient for each pair of un- certainty quantification methods and aggregate these correlations across questions, models, and generation strategies by averaging. The resulting aggregated correlation matrix is shown in Figure 5a. From the figure, we observe strong correlations among categorical- based methods, reflecting their reliance on similar frequency-based evidence. Among relation-based methods, those that share the same graph construction strategy exhibit particularly high correlations. In addition, methods derived from embedding-based and Jaccard Figure 4: Distribution of various in-group average rank in stability across model, question and strategy. similarity graphs show moderate correlation, as both capture se- mantic relatedness between responses, albeit at different represen- tational levels. 4.5.2 Perspective-specific Discussion. Similar to our analyses of effectiveness and stability, we examine the robustness of our corre- lation findings under variations in questions, models, and gener- ation strategies. As shown in Figures 5c, Figure 5b and Figure 5d, although variability exists across different models, questions, and strategies, the magnitude of this variance is relatively small (typi- cally < 0.2) compared to the average correlation values (generally > 0.5). This indicates that the overall correlation patterns reported in our general results are robust to these sources of variation. 5 Conclusion We present a benchmark of uncertainty quantification methods for LLM-based automatic assessment, and emphasize the reliability of prediction by investigating the stability of the uncertainty metrics. Although uncertainty estimation has been studied in many fields, the automated assessment of LLMs is still underexplored due to the ordinal output and rubric-driven setting. To address this issue, we systematically evaluate a diverse set of UQ metrics across datasets of different subjects, LLMs of different scales, and prompt strate- gies to elicit assessment responses. We analyze the effectiveness, stability, and the relations between the UQ methods. Our analy- sis results reveal several insightful patterns. The categorical-based How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY (a) General mean values.(b) Variance over models. (c) Variance over questions.(d) Variance over strategics. Figure 5: Mean and variance of aggregated pairwise correla- tion values among methods across questions, models, and generation strategies. methods demonstrate strong effectiveness in probing uncertainty in assessment judgment; as for relation-based methods, although the performance is not so competitive, their robustness to sampling variation is superior to support a stable prediction. This phenome- non shows a trade-off between assessment discriminative ability and stability. The inter-metric analyses show clear redundancy within method families and provide practical guidance for selecting complementary uncertainty metrics to achieve better uncertainty modeling efficiently. Overall, our findings highlight that no UQ metric is universally optimal, and selecting the proper UQ metric should take the LLM’s capability, assessment tasks, and the settings into consideration. We hope this benchmark facilitates principled evaluation of uncertainty estimation for auto-assessment tasks and supports the development of reliable and practical LLM-based as- sessment systems. References [1]Alberto Abadie, Susan Athey, Guido W Imbens, and Jeffrey M Wooldridge. 2020. Sampling-based versus design-based uncertainty in regression analysis. Econo- metrica 88, 1 (2020), 265–296. [2]Abbirah Ahmed, Arash Joorabchi, and Martin Hayes. 2022. On deep learning approaches to automated assessment: Strategies for short answer grading. (2022). [3]Kirsti M Ala-Mutka. 2005. A survey of automated assessment approaches for programming assignments. Computer science education 15, 2 (2005), 83–102. [4] Anthropic. 2025. Claude Sonnet 4.5. https://w.anthropic.com/news/claude- sonnet-4-5. Large language model. [5]Anthropic. 2025. Introducing Claude Haiku 4.5. https://w.anthropic.com/ news/claude-haiku-4-5. Large language model. [6] Anthony J Bishara and James B Hittner. 2012. Testing the significance of a corre- lation with nonnormal data: comparison of Pearson, Spearman, transformation, and resampling approaches. Psychological methods 17, 3 (2012), 399. [7]Nathan Brake and Thomas Schaaf. 2024. Comparing Two Model Designs for Clinical Note Generation; Is an LLM a Useful Evaluator of Consistency?. In Findings of the Association for Computational Linguistics: NAACL 2024. 352–363. [8]Henry Braun, Isaac I Bejar, and David M Williamson. 2006. Rule based methods for automated scoring: Application in a licensing context. Automated scoring of complex tasks in computer-based testing (2006), 83–122. [9] Saloua Chlaily, Debanshu Ratha, Pigi Lozou, and Andrea Marinoni. 2023. On measures of uncertainty in classification. IEEE Transactions on Signal Processing 71 (2023), 3710–3725. [10] Yucheng Chu, Peng He, Hang Li, Haoyu Han, Kaiqi Yang, Yu Xue, Tingting Li, Joseph Krajcik, and Jiliang Tang. 2025. Enhancing llm-based short answer grading with retrieval-augmented generation. arXiv preprint arXiv:2504.05276 (2025). [11]Yucheng Chu, Hang Li, Kaiqi Yang, Harry Shomer, Hui Liu, Yasemin Copur- Gencturk, and Jiliang Tang. 2024. A llm-powered automatic grading framework with human-level guidelines optimization. arXiv preprint arXiv:2410.02165 (2024). [12]Ben E Cline, Carlyle C Brewster, and Richard D Fell. 2010. A rule-based system for automatically evaluating student concept maps. Expert systems with applications 37, 3 (2010), 2282–2291. [13]Clayton Cohn, Nicole M. Hutchins, Tuan Le, and Gautam Biswas. 2024. A Chain- of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science. In AAAI Conference on Artificial Intelligence. https://api.semanticscholar.org/CorpusID:268553761 [14]Jeremy Cole, Michael Zhang, Dan Gillick, Julian Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023. Selectively answering ambiguous questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 530–543. [15]Yasemin Copur-Gencturk and Tammy Tolar. 2022. Mathematics teaching exper- tise: A study of the dimensionality of content knowledge, pedagogical content knowledge, and content-specific noticing skills. Teaching and Teacher Education 114 (2022), 103696. [16] Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie. 2019. Class- balanced loss based on effective number of samples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9268–9277. [17] Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. 2005. A tutorial on the cross-entropy method. Annals of operations research 134, 1 (2005), 19–67. [18] Haotian Deng, Chris Farber, Jiyoon Lee, and David Tang. 2025.Rubric- Conditioned LLM Grading: Alignment, Uncertainty, and Robustness. arXiv preprint arXiv:2601.08843 (2025). [19]Myroslava O Dzikovska, Rodney Nielsen, Chris Brew, Claudia Leacock, Danilo Giampiccolo, Luisa Bentivogli, Peter Clark, Ido Dagan, and Hoa Trang Dang. 2013. Semeval-2013 task 7: The joint student response analysis and 8th recognizing textual entailment challenge. In Second Joint Conference on Lexical and Compu- tational Semantics (* SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013). 263–274. [20] Emrah Emirtekin. 2025. Large Language Model-Powered Automated Assessment: A Systematic Review. Applied Sciences (2025). https://api.semanticscholar.org/ CorpusID:278790434 [21]Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023. Human-like summarization evaluation with chatgpt. arXiv preprint arXiv:2304.02554 (2023). [22]Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al.2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024). [23] Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al.2024. A survey on llm-as-a-judge. The Innovation (2024). [24]Ben Hamner, Jaison Morgan, lynnvandev, Mark Shermis, and Tom Vander Ark. 2012. The Hewlett Foundation: Automated Essay Scoring. https://kaggle.com/ competitions/asap-aes. Kaggle Competition. [25]Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024. LLM-Rubric: A Multidimensional, Calibrated Approach to Auto- mated Evaluation of Natural Language Texts. In Annual Meeting of the Associ- ation for Computational Linguistics. https://api.semanticscholar.org/CorpusID: 271923672 [26]Jia He, Mukund Rungta, David Koleczek, Arshdeep Sekhon, Franklin X Wang, and Sadid A. Hasan. 2024. Does Prompt Formatting Have Any Impact on LLM Performance? ArXiv abs/2411.10541 (2024). https://api.semanticscholar.org/ CorpusID:274131315 [27] Peng He, Namsoo Shin, and Joseph Krajcik. 2024. SCHOOL LEVEL. Handbook of Research on Science Learning Progressions (2024). [28]Colin A Higgins, Geoffrey Gray, Pavlos Symeonidis, and Athanasios Tsintsifas. 2005. Automated assessment and experiences of teaching programming. Journal on Educational Resources in Computing (JERIC) 5, 3 (2005), 5–es. [29]Pedram Hosseini, Jessica M Sin, Bing Ren, Bryceton G Thomas, Elnaz Nouri, Ali Farahanchi, and Saeed Hassanpour. 2024. A benchmark for long-form medical question answering. arXiv preprint arXiv:2411.09834 (2024). Conference acronym ’X, June 03–05, 2018, Woodstock, NYLi et al. [30]Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Fe- lix Juefei-Xu, and Lei Ma. 2023. Look before you leap: An exploratory study of un- certainty measurement for large language models. arXiv preprint arXiv:2307.10236 (2023). [31]Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al.2024. Openai o1 system card. arXiv preprint arXiv:2412.16720 (2024). [32]Gerd Kortemeyer. 2023. Performance of the pre-trained large language model GPT-4 on automated short answer grading. Discover Artificial Intelligence 4 (2023). https://api.semanticscholar.org/CorpusID:262045158 [33]Luke Korthals, Emma Akrong, Gali Geller, Hannes Rosenbusch, Raoul Gras- man, and Ingmar Visser. 2025. Towards Reliable LLM Grading Through Self- Consistency and Selective Human Review: Higher Accuracy, Less Work. (2025). [34]Jack Krolik, Herprit Mahal, Feroz Ahmad, Gaurav Trivedi, and Bahador Saket. 2024. Towards leveraging large language models for automated medical q&a evaluation. arXiv preprint arXiv:2409.01941 (2024). [35]Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664 (2023). [36]Abhishek Kumar, Sonia Haiduc, Partha Pratim Das, and Partha Pratim Chakrabarti. 2024. Llms as evaluators: A novel approach to evaluate bug re- port summarization. arXiv preprint arXiv:2409.00630 (2024). [37]Alexander Kurz, Katja Hauser, Hendrik Alexander Mehrtens, Eva Krieghoff- Henning, Achim Hekler, Jakob Nikolas Kather, Stefan Fröhling, Christof Von Kalle, and Titus Josef Brinker. 2022. Uncertainty estimation in medical image classifica- tion: systematic review. JMIR Medical Informatics 10, 8 (2022), e36427. [38]Dan Levi, Liran Gispan, Niv Giladi, and Ethan Fetaya. 2022. Evaluating and calibrating uncertainty prediction in regression tasks. Sensors 22, 15 (2022), 5540. [39]Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th annual meeting of the association for computational linguistics. 7871–7880. [40]Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2024. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar.org/CorpusID:274280574 [41] Hang Li, Yucheng Chu, Kaiqi Yang, Yasemin Copur-Gencturk, and Jiliang Tang. 2025. LLM-Based Automated Grading with Human-in-the-Loop. 2025 IEEE International Conference on Teaching, Assessment, and Learning for Engineering (TALE) (2025), 1–8. https://api.semanticscholar.org/CorpusID:277626898 [42]Hai Li, Chenglu Li, Wanli Xing, Sami Baral, and Neil T. Heffernan. 2024. Au- tomated Feedback for Student Math Responses Based on Multi-Modality and Fine-Tuning. Proceedings of the 14th Learning Analytics and Knowledge Conference (2024). https://api.semanticscholar.org/CorpusID:268272301 [43]Qingquan Li, Shaoyu Dou, Kailai Shao, Chao Chen, and Haixiang Hu. 2025. Evaluating Scoring Bias in LLM-as-a-Judge. ArXiv abs/2506.22316 (2025). https: //api.semanticscholar.org/CorpusID:280010889 [44]Alexander H Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, et al.2026. Ministral 3. arXiv preprint arXiv:2601.08584 (2026). [45] Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing. 2511–2522. [46]Daniel Mark Low, Kate H Bentley, and Satrajit S. Ghosh. 2019. Automated assess- ment of psychiatric disorders using speech: A systematic review. Laryngoscope Investigative Otolaryngology 5 (2019), 96 – 116. https://api.semanticscholar.org/ CorpusID:211829165 [47]Chang Lu and Maria Cutumisu. 2021. Integrating Deep Learning into an Auto- mated Feedback Generation System for Automated Essay Scoring. International Educational Data Mining Society (2021). [48]Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583 (2023). [49] Qing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, and Chris Callison-Burch. 2024. Calibrating Large Language Models with Sample Consistency. ArXiv abs/2402.13904 (2024). https://api.semanticscholar.org/CorpusID:267770526 [50] José Mena, Oriol Pujol, and Jordi Vitrià. 2021. A survey on uncertainty estima- tion in deep learning classification systems from a bayesian perspective. ACM Computing Surveys (CSUR) 54, 9 (2021), 1–35. [51]OpenAI. 2024. GPT-4o System Card. https://api.semanticscholar.org/CorpusID: 273662196 [52]José Carlos Paiva, José Paulo Leal, and Álvaro Figueira. 2022. Automated assess- ment in computer science education: A state-of-the-art review. ACM Transactions on Computing Education (TOCE) 22, 3 (2022), 1–40. [53] Vyas Raina, Adian Liusie, and Mark Gales. 2024. Is llm-as-a-judge robust? inves- tigating universal adversarial attacks on zero-shot llm assessment. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 7499–7517. [54]Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). 3982–3992. [55]Murat Sensoy, Lance Kaplan, and Melih Kandemir. 2018. Evidential deep learning to quantify classification uncertainty. Advances in neural information processing systems 31 (2018). [56]Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2024. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. https://api.semanticscholar.org/CorpusID:274776630 [57] Paolo Simeone, Claudio Marrocco, and Francesco Tortorella. 2011. Shaping the error-reject curve of error correcting output coding systems. In International Conference on Image Analysis and Processing. Springer, 118–127. [58]Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. 2025. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267 (2025). [59]Juhi Singh, Ziqiao Ao, and Sebastian Antinome. [n. d.]. Optimizing Prompt Refine- ment: Algorithmic Strategies for Large Language Model-Based Text Classification. Available at SSRN 5525975 ([n. d.]). [60]Hwanjun Song, Hang Su, Igor Shalyminov, Jason (Jinglun) Cai, and Saab Mansour. 2024. FineSurE: Fine-grained Summarization Evaluation using LLMs. ArXiv abs/2407.00908 (2024). https://api.semanticscholar.org/CorpusID:270869629 [61] Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng. 2024. Automated essay scoring and revising based on open-source large language models. IEEE Transactions on Learning Technologies 17 (2024), 1880–1890. [62]Harald Steck, Balaji Krishnapuram, Cary Dehing-Oberije, Philippe Lambin, and Vikas C Raykar. 2007. On ranking in survival analysis: Bounds on the concordance index. Advances in neural information processing systems 20 (2007). [63]Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al.2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023). [64]Qwen Team. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https: //arxiv.org/abs/2505.09388 [65]Stability AI Language Team. [n. d.]. Stable LM 2 1.6B. [https://huggingface.co/ stabilityai/stablelm-2-1.6b](https://huggingface.co/stabilityai/stablelm-2-1.6b) [66] TII Team. 2024. The Falcon 3 family of Open Models. [67] Mani Kumar Tellamekala, Shahin Amiriparian, Björn W Schuller, Elisabeth André, Timo Giesbrecht, and Michel Valstar. 2023. COLD fusion: Calibrated and ordinal latent distribution fusion for uncertainty-aware multimodal emotion recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 2 (2023), 805– 822. [68]Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine- tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5433–5442. [69]Roman Vashurin, Ekaterina Fadeeva, Artem Vazhentsev, Lyudmila Rvanova, Daniil Vasilev, Akim Tsvigun, Sergey Petrakov, Rui Xing, Abdelrahman Sadallah, Kirill Grishchenkov, et al.2025. Benchmarking uncertainty quantification meth- ods for large language models with lm-polygraph. Transactions of the Association for Computational Linguistics 13 (2025), 220–248. [70]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al.2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837. [71]Zhiqiu Xia, Jinxuan Xu, Yuqian Zhang, and Hang Liu. 2025. A survey of uncer- tainty estimation methods on large language models. In Findings of the Association for Computational Linguistics: ACL 2025. 21381–21396. [72]Zhuyang Xie, Yan Yang, Jie Wang, Xiaorong Liu, and Xiaofan Li. 2024. Trustwor- thy multimodal fusion for sentiment analysis in ordinal sentiment space. IEEE Transactions on Circuits and Systems for Video Technology 34, 8 (2024), 7657–7670. [73]Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv preprint arXiv:2306.13063 (2023). [74] Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V. Chawla, and Xiangliang Zhang. 2024. Justice or Prejudice? Quantifying Biases in LLM-as-a- Judge. ArXiv abs/2410.02736 (2024). https://api.semanticscholar.org/CorpusID: 273098639 How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic AssessmentConference acronym ’X, June 03–05, 2018, Woodstock, NY [75]Caiqi Zhang, Fangyu Liu, Marco Basaldella, and Nigel Collier. 2024. Luq: Long- text uncertainty quantification for llms. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 5244–5262. [76] Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019). [77] Chenyan Zhao, Mariana Silva, and Seth Poulsen. 2025. Language models are few-shot graders. In International Conference on Artificial Intelligence in Education. Springer, 3–16. [78]Tianjie Zhao, Sheng Wang, Chaojun Ouyang, Min Chen, Chenying Liu, Jin Zhang, Long Yu, Fei Wang, Yong Xie, Jun Li, et al.2024. Artificial intelligence for geoscience: Progress, challenges, and perspectives. The Innovation 5, 5 (2024). [79] Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, An- drew M Dai, Quoc V Le, James Laudon, et al.2022. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems 35 (2022), 7103–7114. [80]Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Weirong Ye, Neil Zhenqiang Gong, Yue Zhang, and Xingxu Xie. 2023. PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. Proceedings of the 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis (2023). https://api.semanticscholar. org/CorpusID:259095572 [81] Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631 (2023). Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009