Paper deep dive
Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/24/2026, 4:38:14 AM
Summary
This study evaluates the performance of GPT-5.5-based multimodal AI in grading handwritten physics assessments, comparing AI scores against official human examiner scores across three high-stakes contexts: a national Physics Olympiad theory exam (OE1), a final Olympiad selection camp (OE2), and a university quantum mechanics exam (QM). The AI achieved high total-score correlations (0.91–0.97) with human marks. Notably, the AI correctly identified the same five-student team for the International Physics Olympiad as the official grading. A second grading round (RII) with refined, page-by-page instructions improved agreement, particularly in reducing over-awarding and improving partial-credit accuracy, though exact partial-credit grading in experimental work remains challenging. The study concludes that reliable AI grading requires detailed rubrics and should serve as a second reader or audit tool.
Entities (7)
Relation Signals (6)
GPT-5.5 → graded → Physics Olympiad
confidence 95% · This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions... across three assessments: a national Physics Olympiad theory examination...
GPT-5.5 → graded → OE2
confidence 95% · OE2 is a final Olympiad selection camp... This study evaluated GPT-5.5-based grading... across three assessments... OE2...
GPT-5.5 → graded → QM
confidence 95% · QM is an end-of-semester university examination in quantum mechanics... This study evaluated GPT-5.5-based grading... across three assessments... QM...
GPT-5.5 → identifiedteamfor → International Physics Olympiad
confidence 90% · For the final Olympiad selection, AI recovered the same five-student team as official grading.
OE1 → partof → Physics Olympiad
confidence 90% · OE1 is a national Olympiad theory examination used to identify a top cohort for the next stage...
OE2 → partof → Physics Olympiad
confidence 90% · OE2 is a final Olympiad selection camp... used to select the final team for the International Physics Olympiad.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions by 416 unique candidates or students across three assessments: a national Physics Olympiad theory examination, the final Olympiad selection camp with theory and experiment components, and a university quantum-mechanics examination. Each submission was graded twice by AI using the official rubrics. The second round used revised page-by-page and evidence-location instructions developed after first-round disagreement analysis. During grading, AI did not see official human marks or AI--human comparisons. Total-score correlations with official marks were high (0.91--0.97). For the final Olympiad selection, AI recovered the same five-student team as official grading. The second round improved aggregate question-part agreement, especially where first-round disagreements were larger. The main difficulty remained exact partial-credit grading, especially in experimental work. Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control.
Tags
Links
- Source: https://arxiv.org/abs/2608.20521v1
- Canonical: https://arxiv.org/abs/2608.20521v1
Trouble viewing inline? Open PDF directly →
Full Text
57,953 characters extracted from source content.
Expand or collapse full text
Large-scale AI grading of handwritten physics assessments: Score agreement and Olympiad team selection outcomes Praveen Pathak Email: praveen@hbcse.tifr.res.in Affiliation: Homi Bhabha Centre for Science Education–TIFR, Mumbai, India Affiliation: Lawrence Livermore National Laboratory, Livermore, CA, USA Siddharth Tiwary Email: siddharthtiwary@berkeley.edu Affiliation: University of California, Berkeley, CA, USA Charudatt Kadolkar Affiliation: Indian Institute of Technology, Guwahati, India Vijay Singh Affiliation: Centre for Excellence in Basic Sciences, Mumbai, India David Rakestraw Affiliation: Lawrence Livermore National Laboratory, Livermore, CA, USA Shirish Pathare Affiliation: Homi Bhabha Centre for Science Education–TIFR, Mumbai, India Anwesh Mazumdar Affiliation: Homi Bhabha Centre for Science Education–TIFR, Mumbai, India August 20, 2026 Abstract Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10 36410\,364 scanned pages from 520 handwritten submissions by 416 unique candidates or students across three assessments: a national Physics Olympiad theory examination, the final Olympiad selection camp with theory and experiment components, and a university quantum-mechanics examination. Each submission was graded twice by AI using the official rubrics. The second round used revised page-by-page and evidence-location instructions developed after first-round disagreement analysis. During grading, AI did not see official human marks or AI–human comparisons. Total-score correlations with official marks were high (0.91–0.97). For the final Olympiad selection, AI recovered the same five-student team as official grading. The second round improved aggregate question-part agreement, especially where first-round disagreements were larger. The main difficulty remained exact partial-credit grading, especially in experimental work. Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control. I Introduction Grading handwritten student responses in physics involves much more than checking final answers. Recent large language models (LLMs) make this problem testable. Vision-capable language models can read handwritten work, compare it with a rubric, and produce comments. Automated scoring predates generative AI, and reviews of short-answer and text-based assessment document a transition from hand-engineered features to learned language representations Burrows et al. 2015; Gao et al. 2024. Earlier work showed that AI can grade responses to introductory physics problems at a useful level Kortemeyer 2023. An independent study of university-level physics problems likewise found promising aggregate grading performance, but also question- and model-dependent errors Mok et al. 2025. A more demanding study on a high-stakes handwritten thermodynamics exam found that AI could support grading, while diagrams were harder than derivations and final grading still required human review Kortemeyer et al. 2024. A follow-up study argued that AI could handle some confident cases while humans review uncertain cases Kortemeyer and Nöhl 2025. Similar conclusions appear in mathematics. Recent studies of handwritten calculus and university mathematics report strong agreement when OCR (optical character recognition), rubrics, and human verification are controlled Liu et al. 2026; Kortemeyer et al. 2025; Vanhoyweghen et al. 2026. Studies of calculus submissions and national mathematics examinations also show strong automated-scoring performance alongside remaining item-level limitations Gandolfi 2025; Morris et al. 2025. Work on handwritten graphs and mathematical OCR also shows that visual representation remains a separate difficulty beyond text recognition Parsaeifard et al. 2025; Nath et al. 2025; Seong et al. 2026. Broader benchmarks such as MathVista make the same point for mathematical reasoning in visual contexts Lu et al. 2024. Recent studies also show that automatic-scoring performance depends materially on the prompt design, rubric and item context, and choice of model Latif and Zhai 2024; Lee et al. 2024; Jiang and Bosch 2024; Pečuchová et al. 2025. These studies show real progress and leave open an important question: Can AI grading help in handwritten physics examinations where the purpose includes a high-stakes outcome such as identifying a cohort or selecting a team? Physics Olympiad exam submissions are a useful test case because small score differences can affect candidates’ rankings, selection outcomes, and medal awards. Olympiad questions are long and often require judgment beyond matching a textbook answer. A correct final answer may be reached for the wrong reason. An incorrect final answer may still show a sound method. Human examiners therefore judge both the answer and the route taken to reach it. This study tests AI grading across a deliberately mixed set of physics assessments. OE1 is a national Olympiad theory examination used to identify a top cohort for the next stage from about 10 00010\,000 first-stage participants. OE2 is a final Olympiad selection camp comprising challenging theoretical and experimental examinations spanning several areas of high school physics and used to select the final team for the International Physics Olympiad. QM is an end-of-semester university examination in quantum mechanics and quantum computation. Together these assessments encompass the full spectrum of work physics examiners evaluate: theory, experiment, derivation, data analysis, diagrams, and conceptual reasoning. The central question is how closely AI scores agree with official examiner scores, and whether the agreement is reliable enough for the relevant assessment outcome. Established automated-scoring frameworks treat validity as a property of the proposed use of the scores. The analysis must therefore include the decision consequences of the assessment system along with average agreement with a reference score Williamson et al. 2012; Kane 2013; Bennett and Bejar 1998. A second question is how significantly the result depends on the details of how the grading instructions are written. The rubrics in this study were already detailed, including partial-credit rules, common errors, alternative valid approaches, and carried-forward errors. This rubric detail was central to the strong agreement reported here. AI grading was done after the exams were completed and after the Olympiad selection results and university grades had already been released. The released official results remained unchanged by this retrospective AI analysis. After describing the assessments and grading procedure, the analysis evaluates agreement at the total-score, selection, question-part, and question-type levels. It then examines representative examples, experimental grading, and recurring disagreements before assessing rubric refinement and confidence-based human-review strategies. The two grading rounds are compared within those results, but the primary focus is agreement with official grading for handwritten physics exam submissions and the conditions under which AI may become useful in future grading. I Data and methods I.1 Assessments and reference scores All exam submissions in this study were handwritten. OE1 and OE2 come from different years and stages of the Indian Physics Olympiad program. For context, this program is a multistage selection process. In the first stage, typically 40 00040\,000–50 00050\,000 students participate nationally, though participation was lower in the year of the OE1 examination used here (10 29210\,292 students). About 350 students are then selected for the next written Olympiad stage, represented here by OE1. From this group, about 35–40 students are selected for the final Olympiad camp, represented here by OE2, and the final team of five students is selected from that camp. The OE1 and OE2 datasets used in this study are from different years because of submission availability and scanning history. They therefore test AI grading at two different high-stakes stages of the same selection system. QM consists of 40 handwritten exam submissions from an end-of-semester university examination in quantum mechanics. In total, 10 36410\,364 scanned pages were processed for AI grading. Table 1 summarizes the datasets. Table 1: Datasets used in the study. An exam component means one separately graded OE2 theory or experiment test section. An exam submission means one student’s answer booklet for one component. OE1 and QM each had one component per student, while OE2 candidates completed five components. Dataset Candidates/ students Exam submissions Exam components Scanned pages Main role OE1 350 350 1 5600 Top-cohort identification OE2 26 130 5 4164 Final team selection QM 40 40 1 600 University class grade Total 416 520 7 10 364 – The human examiner scores used here are the final official scores for each exam submission. Each question was graded by one examiner and checked by a second examiner before release. When the checking identified a concern, the mark was reconsidered and settled before scores were released. Students also had review mechanisms: OE1 allows regrading requests after scores are released, while in OE2 and QM students can inspect their graded submissions and discuss the points awarded with examiners. We therefore treat the final official scores, after these checks, as the reference scores for this study. Differences between official scores and AI scores are reported as disagreements with the official score. Because OE1 and OE2 are selection-oriented assessments, rank and top-group agreement are part of the grading question. For OE2, the official final ranking followed the International Physics Olympiad theory–experiment weighting: theory contributes 30 marks and experiment contributes 20 marks, a 60:40 ratio. The two theory components were each originally out of 80 and were prorated to 120, so theory contributed 240 marks and the three experiment components contributed 160 marks, giving a combined score out of 400. In normalized analyses this is equivalent to combining the theory and experiment totals as 0.6T+0.4E0.6T+0.4E. I.2 AI grading procedure and grading rounds The AI grading was performed both in browser-based runs and through Codex using the OpenAI API. The AI grader was given the exam questions, solutions, official rubrics, and anonymized handwritten exam submissions. Visible human scores and comments on the submissions were erased before AI grading. The same official rubrics were used for human and AI grading. For human examiners, the rubrics served as scoring guidelines to be interpreted case by case, including responses that did not fit a listed solution path. The rubrics specified partial credit, common errors, alternative valid solution paths, and error-carried-forward rules. The AI prompt asked for these rules to be applied to the full written solution, including reasoning and intermediate steps, and to avoid repeated penalties for the same carried-forward error. The aim was to follow the International Olympiad exam grading standards. Table 2 summarizes the model and effort settings used in the reported analyses. The main comparison between the two full grading rounds uses GPT-5.5 Thinking at high effort. The 5.5 Pro runs were used to check whether the OE2 outcome was stable and to review selected exam submissions; the main OE2 selection outcome was unchanged in these checks. The analysis is organized around two full grading rounds and one focused refinement stage explained below: • Round I (RI): the initial full AI grading. RI had no access to human scores, official totals, ranks, or selection status. The AI was provided the exam submissions, exam questions, solutions, rubrics, and the general grading instructions described above. The output included scores and comments for each question part or subpart. • Focused refinement: carried out after inspecting Round I score differences and AI comments. Three places with repeated or large disagreements, one each from QM, OE1, and OE2, were tested with more explicit instructions. These diagnostic runs tested whether stating the intended physics and scoring conditions more clearly could reduce AI-official score differences. The examples are discussed in Sec. V.1. • Round I (RII): a new full grading run using instructions revised after the Round I disagreement analysis. During RII, the AI did not see RI scores, RI comments, human scores, official totals, ranks, selection status, or any AI–human comparison. It differed from RI in two main ways. First, it included the refined rubrics for the three parts discussed in the focused refinement stage. The remaining parts used the same rubrics as RI. Second, the AI was asked to grade long submissions page by page. A single submission in our exam set could exceed 30 scanned pages. RII therefore asked the AI to identify where the credited evidence appeared. The output also recorded confidence labels and review flags, discussed in Sec. V.2. Table 2: AI models and effort settings used in the study. Model Effort Exams/runs included GPT-5.5 Thinking High Main RI/RII grading: OE1, OE2, QM GPT-5.5 Pro Standard OE2 full comparison and outcome check GPT-5.5 Pro Extended OE2 pilot involving selected submissions I.3 Evaluation measures For analyses of individual official question parts, one comparison means one student response to one official part or subpart of a question. In these exams, questions were divided into parts and subparts, each with its own maximum mark; the grading scheme then specified how marks within that part were awarded or deducted. This gives 7058 comparisons between AI and official scores across OE1, OE2, and QM. Table 3 defines the quantities used below. Table 3: Analysis parameters used in the study. Unless raw marks are explicitly stated, D and MAD are reported as percentages of the relevant maximum possible score. Quantity What it measures Purpose r Pearson correlation between human and AI scores Score tracking ρ Spearman rank correlation Rank-order agreement D Mean of AI minus human scores. Positive values mean AI awarded more Direction of over- or under-awarding MAD Mean absolute difference between AI and human scores Size of the grading difference Top-k overlap Fraction of human top-k submissions also in the AI top-k Top-group and team-selection checks d=|AI−Human|d=|AI-Human| Raw point difference for one official question part Difference bands for individual question parts Confidence flag AI’s self-reported confidence or request for human review Prioritizing human review in RII I Agreement and selection results I.1 Total-score agreement At the level of total scores, AI and official human scores track each other strongly across all three assessments. Fig. 1 shows scatter plots of AI score against human score for OE1, QM, and the OE2 theory and experiment components. Each point represents one student. The points lie close to the equal-score diagonal in both rounds, although RII is generally less shifted toward positive AI–human differences than RI. The Pearson correlations are high in RI (r=0.91r=0.91–0.970.97) and remain high in RII (r=0.93r=0.93–0.960.96). RII was the full revised workflow described above: page-by-page checking, evidence notes, stricter checking of permitted scores, confidence/review fields, and clearer instructions for selected questions. The three focused refinements discussed later were worth only about 3.6% of the combined official rubrics, so RI–RII differences should be interpreted as effects of the full revised workflow rather than those three questions alone. With detailed rubrics and instructions, current multimodal AI can reproduce the overall score distribution of handwritten physics assessments reasonably well. Strong total-score agreement still leaves grading differences. In RI, D was positive in all four total-score comparisons, meaning that AI generally awarded more points than human examiners. RII reduced this positive shift in OE1 and QM, as seen from the smaller D values reported in Fig. 1. MAD gives the size of the remaining score difference. The clearest reductions in total-score MAD occurred for OE1 and QM: OE1 MAD fell from 7.1% to 4.8%, and QM MAD fell from 9.4% to 3.8%. OE2 already had strong total-score agreement in RI, and the RI–RII changes are smaller because OE2 combines several theoretical and experimental examinations. Similar over-awarding concerns have been reported in earlier AI-grading studies Kortemeyer et al. 2024; Kortemeyer and Nöhl 2025. RI RII005050100100r=0.93r=0.93ρ=0.92ρ=0.92D=+5.9%D=+5.9\%MAD=7.1%=7.1\%AI score (%)OE1 RI (n=350)005050100100r=0.97r=0.97ρ=0.97ρ=0.97D=+4.4%D=+4.4\%MAD=5.0%=5.0\%OE2 theory RI (n=26)005050100100r=0.91r=0.91ρ=0.90ρ=0.90D=+2.7%D=+2.7\%MAD=4.6%=4.6\%OE2 experiment RI (n=26)005050100100r=0.94r=0.94ρ=0.94ρ=0.94D=+9.3%D=+9.3\%MAD=9.4%=9.4\%QM RI (n=40)005050100100005050100100r=0.95r=0.95ρ=0.95ρ=0.95D=+2.2%D=+2.2\%MAD=4.8%=4.8\%Human score (%)AI score (%)OE1 RII (n=350)005050100100005050100100r=0.96r=0.96ρ=0.93ρ=0.93D=+4.5%D=+4.5\%MAD=5.1%=5.1\%Human score (%)OE2 theory RII (n=26)005050100100005050100100r=0.93r=0.93ρ=0.92ρ=0.92D=+4.0%D=+4.0\%MAD=4.2%=4.2\%Human score (%)OE2 experiment RII (n=26)005050100100005050100100r=0.96r=0.96ρ=0.96ρ=0.96D=−0.1%D=-0.1\%MAD=3.8%=3.8\%Human score (%)QM RII (n=40) Figure 1: AI versus human total scores in RI and RII, expressed as percentages of the assessment maximum. Each point represents one student. The diagonal line is perfect agreement. Blue points show RI and green points show RII. r, ρ, D, and MAD are shown inside each panel. To check the direction and spread of the differences, Fig. 2 plots AI minus human score against the human total score, pooled across the four total-score comparisons. Positive values mean AI awarded more than the human examiner. RI points are more often above zero, while RII reduces this upward shift. The differences appear across the score range rather than only among low-scoring submissions or near a cutoff. The fitted lines are included as a visual guide; because the vertical axis already contains the human score, the slope should not be over-interpreted as showing that AI is better or worse for high-scoring students. 005050100100−30-30−15-150015153030slope =−0.050=-0.050R2=0.02R^2=0.02Human total score (%)AI minus human (%)RI pooled (n=442)005050100100−30-30−15-150015153030slope =−0.035=-0.035R2=0.01R^2=0.01Human total score (%)RII pooled (n=442) Figure 2: Residual total-score plots pooled across OE1, OE2 theory, OE2 experiment, and QM. The vertical axis is AI minus human total score, expressed as a percentage of the assessment maximum. Blue circles show RI and green squares show RII. The fitted line is a least-squares visual guide, and the horizontal zero line is perfect agreement. The figure mainly shows the direction and spread of AI–human differences: RI is more often positive, while RII is less shifted upward. The fitted slopes should be read cautiously because the human score is part of the plotted difference. I.2 Top-cohort and selection agreement For assessments that classify or select candidates, average-score agreement is only one part of the evaluation. The practical question is whether AI also matches the relevant outcome. OE1 identifies a cohort for the next stage, OE2 selects a final team of five students, and QM assigns a course grade. Classifications based on test scores have their own accuracy and consistency properties Livingston and Lewis 1995, so these outcome checks are part of the grading question. For OE1, AI recovered most of the larger top group. Each round recovered 31 of the human top 40 and 40 of the human top 50. At smaller cutoffs RII improved the overlap. For example, the top-10 overlap increased from 3/10 in RI to 7/10 in RII. This level of agreement is useful for identifying the top group, while human grading remains necessary for exact ranks near a cutoff. For OE2, the key question is whether AI matches the selected group for India’s IPhO (International Physics Olympiad) team. In both RI and RII, the AI top five contained the same five students as the human top five, although the AI ranked them in a different order. This comparison is based on five exam components per candidate, so the final ranking combines more evidence than a single examination. At k=10k=10, both rounds recovered 9 of the human top 10, and at k=20k=20 both recovered 19 of the human top 20. At the OE2 top-five cutoff, overlap was complete in both rounds. QM has a different outcome: a course grade. The released grades span nine categories, from FP to AS. We therefore compared the AI-based course grades with the released official grades. RI exactly matched 26 of 40 official grades, while RII exactly matched 34 of 40, and all 40 RII grades were within one grade step on this nine-category scale (Fig. 3). AS is the highest grade. RI RII55101015152020252530303535404045455050005050100100Top-kkOverlap (%)OE155101015152020005050100100Top-kkOE2 FPFPDDDDCDCDCCCCBCBCBBBBABABAAAAASASHuman gradesAI grades (RII)QM grades1661431211311 Figure 3: Outcome agreement. The OE1 and OE2 panels show the percentage of human top-k submissions or students also present in the AI top-k group; blue circles show RI and green squares show RII. The QM panel shows RII course-grade agreement across the nine released grade categories. Green bubbles show exact grade matches and orange bubbles show different grade pairs. The number inside each bubble is the number of students; AS is the highest grade. The OE2 outcome was also checked for dependence on AI mode. The Pro Standard run gave a similar result: it recovered the same human top-five group as the main Thinking High run, although the internal order changed. At broader cutoffs, Pro Standard recovered 9 of the human top 10 and 19 of the human top 20. A Pro Extended pilot was run only on six selected OE2 submissions with large earlier AI–human disagreements. It modestly reduced some disagreements, especially in theory, while the same pattern of over-awarding remained. Although the setting and model configuration differ from ours, earlier work has also reported nontrivial run-to-run variation when AI graded the same responses Jauhiainen and Garagorry Guerra 2025. Overall, the OE2 team configuration was stable across these AI-mode checks. This mode comparison was limited to OE2 and was not repeated for OE1 or QM. I.3 Question-part agreement and partial credit Many questions in the exams were divided into parts and subparts. Scores on these official question parts therefore give a stricter test than total scores. In Table 4, each official question part is compared separately, using d=|AI−Human|d=|AI-Human| in raw points. The usual scoring increment was 0.5, so the table separates exact matches, differences up to 0.5 points, differences between 0.5 and 1 point, and differences larger than 1 point. Because the maximum mark differs across parts, the d bands should be read together with the normalized MAD values and the zero/partial/full-credit analysis below. Across all 7058 official question parts, RI matched the human score exactly in about 63% of parts. RII increased exact agreement to about 70%, and reduced parts differing by more than one point from about 13% to 7%. Table 4: Agreement between AI and human scores for individual question parts. Here d=|AI−Human|d=|AI-Human| in raw points. Values are rounded percentages. Bands sum to 100% before rounding. d=0d=0 0<d≤0.50<d≤ 0.5 0.5<d≤10.5<d≤ 1 d>1d>1 Set Question parts RI RII RI RII RI RII RI RII OE1 4200 61 71 15 12 9 9 15 8 OE2 theory 1482 70 75 13 13 9 6 8 6 OE2 experiment 416 45 56 21 17 9 13 25 14 QM 960 63 69 21 18 8 9 8 4 All official question parts 7058 63 70 16 14 9 9 13 7 The OE2 experiment row reports the official experiment question parts. Because some official experiment parts combine several judgments into one mark, RII was also checked at a more detailed experimental-component level for analysis only. In that check (Sec. IV.2), exact agreement was 77%, with 16% in 0<d≤0.50<d≤ 0.5, 4% in 0.5<d≤10.5<d≤ 1, and 3% in d>1d>1. Separating question parts by the official human score gives a more useful view of overall exact agreement. Fig. 4 shows two levels of agreement. At the broad level, AI often placed responses in the correct zero-, partial-, or full-credit band. In RII, it kept 80% of human-zero parts at zero, 86% of human-partial parts in the partial-credit range, and 87% of human-full parts at full credit. This band-level agreement matters because partial-credit parts account for 40% of the available points. AI = zero AI = partial AI = fullHuman = zero (n=2870n=2870)Human = partial (n=1834n=1834)Human = full (n=2354n=2354)RIRIIRIRIIRIRII69%29%80%18%83%16%86%12%16%84%11%87% Figure 4: AI behavior by official human score group. Each bar shows the distribution of AI scores within question parts where the official human score was zero, partial credit, or full credit. The partial-credit parts contain 40% of the available points. The harder task is the exact amount of partial credit. For human-partial parts, exact agreement rose from 24.6% in RI to 32.8% in RII. Among cases where both the human examiner and AI gave partial credit, MAD fell from 0.92 to 0.72 raw marks. When AI moved a human-zero part into the partial-credit range, the average award was about one raw mark, and this happened less often in RII than RI. Thus, RII improved both broad category recognition and the calibration of partial credit, while exact partial-credit scoring remained the most difficult case. To understand what lies behind these aggregate question-part results, we next examine representative successes, experimental grading, differences across question types, and recurring disagreement patterns. IV Strengths and disagreements in AI grading IV.1 Examples of physics-specific grading In many cases the AI comments were aligned with the physics conditions in the rubric. The model often found the relevant solution in lengthy, multipage exam submissions, even when the final summary answer and the detailed work appeared on different pages. It interpreted handwritten derivations, credited equivalent mathematical forms, and sometimes identified specific physics errors. These examples show that the total-score agreement was supported by physics-specific grading comments, with the AI often identifying the evidence relevant to the score. One example is a transcription error between the working pages and the summary answer box. In Fig. 5, the copied final expression in the summary box is off by a factor of two, but the detailed working contains the correct self-energy factor and the correct final result. The AI credited the detailed working despite the transcription error in the summary box. This matters in Olympiad-style grading, where students often copy final results into a summary answer space and a copying error is judged in the context of the full submission. (a) (b) Figure 5: Example of AI using the full working to interpret a summary-box transcription error. (a) The transferred expression in the summary answer box is off by a factor of two. (b) The detailed working of the same submission contains the correct self-energy factor and final expression. The AI comment stated: “The summary box misses a factor, but the working sheet correctly derives Q02/(8πϵ0)(1/R0−1/R)Q_0^2/(8π _0)(1/R_0-1/R).” This matched the official full-credit mark. A related example involved visual grading. The AI was often able to read qualitative sketches and identify physically relevant features. In one question, examinees were required to draw the effective-potential diagram, VeffV_ eff versus z. In one submission, the required sketches appeared on later pages outside the designated answer box. The AI found them and awarded full credit. In another case, it correctly gave only limited credit because the sketch showed curves meeting the zero line at the endpoints, but lacked the double-well and critical shapes required by the rubric (Fig. 6, with exact AI comments in the caption). This kind of interpretation matters in Olympiad grading because small score differences can affect ranking. (a) (b) (c) Figure 6: Examples of visual reasoning in AI grading. (a) Correct reference shapes from the solution. (b) Full-credit example where the relevant sketches were outside the designated answer box. AI found them on later pages and commented: “Detailed sketches show the required below-critical, critical, and above-critical effective-potential shapes with the expected qualitative behavior.” (c) Limited-credit example where AI awarded the endpoint behavior specified in the rubric but withheld credit for the missing double-well and critical-transition shapes. The AI comment stated: “Provides some qualitative potential curves and endpoint behavior, but the subcritical double-well and critical transition are not represented correctly.” IV.2 Experimental grading The OE2 experimental submissions are a useful case because they require reading tables, calculations, graphs, fit lines, uncertainty estimates, and written conclusions. Some official experiment parts were broad, with several of these judgments contributing to the same mark. RII also recorded them more separately, for example at the level of table quality, graphing, fit or slope extraction, uncertainty, and conclusion. These experimental components were compared with the corresponding human markings recorded in the official grading scheme. This more detailed experimental check is summarized below Table 4. The AI often recognized transformed tables, plotted points, fit and limiting lines, slopes, and reported values. The weaker cases were those where the score depended on whether the data range, fit, uncertainty, and final conclusion supported each other. (a) (b) (c) Figure 7: Examples from masked OE2 experimental submissions. Panels (a) and (b) belong to the same data-table and graph-analysis task, and the AI and official scores agreed on both marking components. AI comments for these two components were: “The transformed 1/a+1/b1/a+1/b values are tabulated for all n values and the linearization is written” and “Graph is present with plotted points, best-fit and limiting lines, and slope calculations for the equivalent 1/a+1/b1/a+1/b plot.” Panel (c) shows a partial-credit graph. Both graders awarded partial credit. AI comment for this component was: “Five points are plotted on an R-versus-h h graph, but no usable best-fit line, slope or uncertainty construction is shown.” The examples in Fig. 7 show that the scoring went beyond graph presence. AI identified fit and limiting lines when they were part of the grading evidence, and it agreed with partial credit when plotted points alone fell short of the full graph-analysis requirement. These examples indicate that multimodal AI can often read and score experimental evidence. Human review remains important when the mark depends on whether the table, graph, fit, uncertainty, and conclusion are physically consistent with each other. IV.3 Agreement across question types To compare OE1, OE2, and QM using the same labels, each official question part was assigned to one primary question type. These labels are broad, but they capture the main type of judgment the grader has to make (Table 5). For OE2 experimental sections, the label refers to the main task in the official question part. Table 5: Question-type labels used across the three assessments. Question type Description Conceptual written reasoning Written physical explanation, evaluation of an answer choice, interpretation, or validity check. Figure-diagram Circuits, free-body diagrams, qualitative sketches, ray diagrams, or graph shapes where the visual structure is the answer. Numerical Calculation or substitution where the quantity being assessed is primarily a value, probability, time, energy, or angle. Derivation Symbolic proof, formula derivation, operator derivation, or multi-step algebraic reasoning. Figure-data Experimental measurements, data tables, plotted data, graph fitting, slope extraction, uncertainty, or data-based inference. This question-type analysis shows that agreement extends beyond numerical work. In derivations, numerical answers, conceptual reasoning, diagrams, and data/graph tasks, AI scores generally rise when human scores rise. The differences lie in how large the score differences are on individual parts. Conceptual written reasoning shows the greatest over-scoring in RI, while figure-data and figure-diagram tasks are sensitive to whether the visual or experimental evidence satisfies the specific grading condition. Fig. 8 and Table 6 summarize the question-type results. MAD is lower in RII for every question type. The reviewed cases suggest why: page-by-page checking and clearer credit conditions help most when the score depends on evidence in a specific part of the student’s response. RI RIIConceptualDiagramNumericalDerivationFig.-data005510101515MAD% Figure 8: MAD by question type in RI and RII for the 7058 official question parts. MAD% is the mean absolute difference as a percentage of the maximum possible score for each part. Table 6: Question-type metrics for official question parts. N is the number of question-part comparisons. D and MAD are percentages of the maximum possible score for each part. Question type N DID_I DIID_I MADI MADII rIr_I rIIr_I Conceptual written reasoning 706 +13.3 +4.3 16.6 10.7 0.83 0.89 Figure-diagram 1082 +3.8 +2.5 10.6 7.7 0.84 0.89 Numerical 1978 +4.7 +0.7 11.1 7.6 0.85 0.88 Derivation 2876 +4.8 +2.3 11.4 9.1 0.88 0.91 Figure-data 416 +3.3 +4.4 9.3 6.6 0.92 0.96 The question-type results also show that RII gains extended beyond the explicitly refined questions. Figure-data and figure-diagram question parts contain several of the largest improvements, but numerical, conceptual, and derivation parts improved as well. The central concern is whether the table, graph, diagram, or derivation satisfies the specific physical criterion required by the rubric. IV.4 Recurring sources of disagreement The larger disagreements usually had identifiable causes, summarized in Table 7. These patterns motivated the focused refinement tests and the later revised full grading. Table 7: Recurring sources of disagreement between AI and official scores. These summarize patterns seen in reviewed cases; individual score differences are interpreted against the official scoring. Source Description Awarding credit too easily AI sometimes awarded partial credit to work to which human examiners assigned zero points, especially when the answer contained plausible symbols, text, tables, or diagram features. Over-valuing the final answer A correct-looking final answer sometimes received too much credit even when the reasoning had serious errors. Treatment of carried-forward errors AI could recognize an earlier error but sometimes applied carried-forward-error rules differently from the human examiner. Missing a required diagram feature AI often read the diagram but could miss a specific feature required by the rubric, such as relative placement, curvature, scale, or a required label. Experimental evidence chain AI often identified tables, graphs, and calculations, but disagreements arose when the score depended on whether the data, fit, uncertainty, and conclusion supported each other. Permitted scoring increments In a few cases, AI awarded increments smaller than the intended scoring resolution. This should be stated clearly in the prompt. These patterns lead to the two checks below: clearer scoring conditions and confidence flags for human review. V Improving and reviewing AI grades V.1 Focused rubric refinement Reviewing RI disagreements identified cases where the rubric stated a broad goal while leaving the specific physics needed for credit implicit. A focused test was therefore run on three problem cases: QM question 3, part b, OE1 Q2, and OE2 T2 Q1(f) (Table 8). Each was regraded with clearer question-specific instructions to test whether the AI-official score difference could be reduced. In these cases, making the required physics and scoring conditions explicit reduced several RI disagreements. Table 8: Focused refinement targets selected after reviewing RI disagreements. Target Points Responses Main issue clarified QM Q3b 3 40 Grover diagram and operator conditions OE1 Q2 8 72 A thermodynamic reasoning based question OE2 T2 Q1(f) 4 26 T–S graph shape, labels, and curvature In OE2, one part of a theory question asked students to draw a thermodynamic cycle on a T–S plot (Fig. 9). The original rubric allotted 0.5 points for the correct shape, but left the intended curvature implicit. The AI often awarded the 0.5 points for the shape even if the overall shape or curvature was physically wrong, though it had access to the correct intended shape in model solutions. The refined instruction stated the conditions directly: “The branches 1→21→ 2 and 3→43→ 4 had to be vertical isentropic branches. The branch 2→32→ 3 had to be an isobaric heating curve, convex upward on a T–S plot with slope increasing left-to-right. The branch 4→14→ 1 had to be an isobaric cooling curve traversed right-to-left. The temperatures had to satisfy T3T_3 highest, T1T_1 lowest, and T2=T4T_2=T_4.” (a) (b) Figure 9: Refinement of the rubric for the OE2 question. (a) The intended T–S cycle from the solution. (b) Student sketch with labels and arrows but with the wrong curvature. In RI, AI marked this as a correct shape. The refined rubric emphasized the physical conditions of the cycle, with less reliance on labels and arrows alone. With the refined rubric, AI stopped awarding shape credit in the incorrect-curvature cases we had identified. For the submission in Fig. 9, the AI comment in this round said that labels and values were present, but “the upper branch had the wrong curvature” and one process arrow was inconsistent. Another example is QM question 3, part b, which involved the Grover-rotation diagram. The original rubric provided only broad criteria for scoring the diagram and mathematical description. In several submissions, the AI awarded credit for a diagram that looked plausible but lacked the required elements. Fig. 10(b) shows one such case. The response contained a circuit-style sketch and a qualitative oracle/diffusion description, but lacked the required two-dimensional rotation diagram and the operator action expected in the solution. |x0⟂⟩|x_0 |x0⟩|x_0 |S⟩|S |S′⟩|S |ψ⟩|ψ |ψ⟩O|ψ O|ψ⟩DO|ψ θ 2θ2θ Choose |x0⟂⟩|x_0 such that |S⟩=cosθ|x0⟂⟩+sinθ|x0⟩. |S = θ|x_0 + θ|x_0 . Let |ψ⟩=cosα|x0⟂⟩+sinα|x0⟩|ψ = α|x_0 + α|x_0 , with α measured from |x0⟂⟩|x_0 . Then O|ψ⟩ O|ψ =cosα|x0⟂⟩−sinα|x0⟩, = α|x_0 - α|x_0 , DO|ψ⟩ DO|ψ =cos(α+2θ)|x0⟂⟩ = (α+2θ)|x_0 +sin(α+2θ)|x0⟩. + (α+2θ)|x_0 . (a) (b) Figure 10: Grover-rotation example from the QM exam. (a) Reference rotation diagram redrawn from the official solution. (b) Student response. In RII, the AI identified that the circuit-style sketch was not the required rotation diagram. The rubric was then rewritten as explicit items (Table 9). In this example, RI awarded 2 out of 3 points. RII matched the official 0.5 out of 3 points by crediting the qualitative oracle/diffusion idea and rejecting the sketch as the required Grover-rotation diagram. Table 9: RI and RII rubric/comment comparison for the Grover example in Fig. 10. RI RII Rubric Correct diagram with all symbols explained: 1 point. Detailed mathematical description of the Grover operator: 2 points. Determine an appropriate point allocation when the response is incomplete. 0.5 points for axes labeled as |x0⟩|x_0 and |x0⟂⟩|x_0 . 0.5 for showing rotation by 2θ2θ under G. 0.5 for defining |x0⟂⟩|x_0 . 0.5 for the action of O. 1 point for the action of G=DOG=DO. Comment by AI Awarded 2/3. Correctly describes Grover as oracle plus diffusion and notes the oracle phase flip, with a supporting circuit-style diagram. Deducted 1 because the diffusion/reflection and amplitude-rotation geometry are not fully explained. Method partly follows the official scheme. Awarded 0.5/3. Gives a qualitative oracle/diffusion description and a circuit-like oracle sketch, but not the required two-dimensional rotation diagram or mathematical action of O and G on cos(α)|x0⟂⟩+sin(α)|x0⟩ (α)|x_0 + (α)|x_0 . For this question, the refinement reduced MAD from 1.1 to 0.4 marks, increased Pearson’s r from 0.70 to 0.82, raised exact agreement from 15% to 45%, and raised agreement within 0.5 mark, including exact matches, from 35% to 80%. Fig. 11 shows the AI-minus-human score distributions for this case and the OE1 Q2 refinement example. Round I Refined−1-1001122330020204040AI minus human scoresSubmissions (%)QM Q3b (/3, n=40)−2-2−1-10011223344556677880020204040AI minus human scoresOE1 Q2 (/8, n=72) Figure 11: AI-minus-human score distributions for two focused refinement examples, shown as percentages of submissions. The horizontal axis gives the difference between AI and human scores. The Round I bars show the original scores, and the Refined bars show scores from the question-specific refinement run. OE1 Q2 was an 8-point conceptual thermodynamics question requiring reasoning for each option. Only the OE1 Q2 submissions with large Round I AI-official score differences (n=72n=72) were regraded to test whether clearer wording reduced these disagreements. The refined rubric required the relevant physical criterion for each option and limited credit for unsupported conclusions. Within this subset, MAD fell from 38% to 17% of the maximum possible score for the part, and Pearson’s r increased from 0.58 to 0.83. Taken together, the focused refinements support the practical requirement that reliable AI grading depends on a detailed, physics-specific marking rubric. Before AI grading is attempted, the rubric should state as explicitly as possible the conditions for awarding marks, common deductions, acceptable alternative solutions, and carried-forward-error rules. V.2 Confidence and human-review flags Confidence and human-review flags for each question part were available in RII. Fig. 12 compares high-confidence question parts with medium- or low-confidence question parts. In every comparison where the flag was available, high-confidence parts had much lower MAD. For example, OE1 RII had MAD% values of 6.0% for high-confidence parts and 17.0% for medium- or low-confidence parts. QM RII had 9.8% versus 28.2%. In OE2, theory parts showed 5.3% versus 27.1%, while experiment parts showed 6.9% versus 25.6%. High confidence Medium/low confidenceOE1OE2 TheoryOE2 ExptQM00101020203030MAD% Figure 12: Confidence flags in RII grading. Medium- and low-confidence question parts have larger differences from official scores, so confidence helps identify parts for human review. Confidence flags still miss some important disagreements. For the OE2 confidence analysis, the 1898 official question parts consisted of 1482 theory parts and 416 experiment parts. Accepting every high-confidence part with no review flag would have accepted 1724 parts; 87 of these still differed from the official human score by more than one point. These results support review triage as the appropriate use: confidence flags help decide what humans should inspect first, while final acceptance still needs examiner judgment. This conclusion aligns with prior work on model confidence and AI-assisted scoring Tian et al. 2023; Li et al. 2025. It also resembles systems that send uncertain cases to a human expert; the present study evaluates the AI’s own review flags, with no separately trained routing model Mozannar and Sontag 2020. VI Discussion and conclusions The results show strong performance, but also clear limitations. AI grading showed strong agreement with the official total scores. For OE1, it identified most of the larger top group in both rounds. For OE2, both RI and RII recovered the same top-five group as the human examiners, although the AI ranked them differently. For QM, RII reproduced 34 of 40 released course grades exactly, and all 40 grades were within one grade step in a nine-step grading system. These findings support the use of current multimodal AI as a useful aid for physics grading in high-stakes examinations. This study used moderated official scores as the human reference. These scores had already been checked by more than one examiner and were available for student review. A separate human–human regrading experiment was outside the scope of this analysis. Reported human-rater data for handwritten physics constructed responses show why this matters: for 20 physics constructed responses scored by four instructors, Tang, Ambrose, and Cheng reported human-only intraclass correlations (ICCs) of 0.88 with a holistic rubric and 0.94 with a checklist rubric, and only 0.28 for mid-level responses under the holistic rubric, rising to 0.89 with the checklist Tang et al. 2026. Partial-credit scoring is therefore difficult for human examiners too, and explicit credit conditions are the shared remedy. Prior automated-scoring work treats human scoring evidence, and in some cases human–human agreement, as the benchmark for interpreting model performance Williamson et al. 2012; Morris et al. 2025. A useful next step would be independent second-human grading on a representative subset of these same submissions, so that AI–official and human–official differences can be compared under the same conditions. The question-part results explain where caution is still warranted. AI is often effective at reading handwritten work, recognizing correct physics, and identifying relevant evidence within large amounts of text and diagrams. It can credit correct working despite a transcription error in the summary box, accept equivalent forms, and identify important diagram features. It is also more willing than human examiners to award partial credit for incomplete or visually plausible work, particularly in conceptual reasoning, diagrams, and experimental graph or data tasks, where human examiners rely more heavily on context and judgment. The refinement experiments show that useful prompt detail specifies the physics conditions required for awarding credit. This improved grading in the Grover diagram question in QM, where the diagram requirements had to be made explicit, and in OE1 Q2, where the expected reasoning for each option had to be clearly stated. The full RII results extend this pattern most clearly for OE1 and QM. OE2 total-score agreement was already high in RI, so its RI–RII changes were smaller and mixed, although question-part and experimental analyses still improved. Because the RII instructions were revised after inspecting RI disagreements in these submissions, the RII gains should be confirmed on an untouched examination or a holdout set. The OE2 results also demonstrate that total-score agreement alone can be misleading, since similar totals may conceal differences in the marks awarded to individual parts. Future prompts should specify the conditions for awarding marks explicitly. Confidence and review fields provided a useful way to prioritize manual review. High-confidence responses showed smaller average score differences across all examinations. Some high-confidence responses still differed from the official human markings, so these measures are best used to prioritize review effort, with examiner checking before final acceptance. A practical lesson from RII is that the unit of grading matters. Asking the AI to grade page by page and identify the evidence supporting each score appeared especially useful for long handwritten submissions, where relevant work may appear outside the expected answer space. This approach reduces the burden of processing an entire submission at once, improves auditability, and was associated with better agreement with human grading, although it also increased grading time. The human grading workload was also substantial. OE1 grading of about 350 submissions typically involves roughly 20 graders over about 2.5 days, with individual graders often working more than 12 hours per day. OE2 grading is a still larger distributed effort: about 25 people work intensively over four days, with the team operating nearly round the clock during final scoring and scrutiny. Because the study combined browser- and API-based workflows, comparable token usage and processing times were incompletely recorded. The study also lacked human grading-time logs suitable for direct comparison, so these staffing figures should be read as workload context rather than a measured time comparison. Future work should measure these costs directly. The practical implication is that expert examiners remain central. Reliable agreement with official grading requires a detailed rubric that makes the award and deduction of marks as explicit as possible. AI shifts part of examiner effort toward defining the rubric carefully, stating the conditions for credit, and reviewing flagged or borderline cases. For practical use, the more intensive Pro modes fit best as targeted audit tools. The main Thinking High run already reproduced the OE2 top-five outcome, and the Pro Standard and limited Pro Extended checks left the grading picture largely unchanged. Pro modes may still be useful for targeted audits of difficult cases, and in this study their outcome-level performance was comparable to Thinking High. Overall, AI is best used as a second reader, an audit tool, and an aid to selection decisions. It can flag submissions with large differences from official scores, provide an additional set of comments, and help examiners check grading consistency. Data and materials statement Anonymized scores by question part, task categories, prompts, and analysis scripts may be shared subject to approval from the relevant examination authorities and institutions. Raw handwritten submissions remain confidential because of student privacy and examination confidentiality. Acknowledgments The authors thank the human examiners and the team that scanned the exam submissions. We thank Mamatha Maddur for helping to digitize and organize the submissions. P.P. acknowledges support from the Fulbright Program administered by the Institute of International Education (IIE). This work was partially performed under the auspices of the U.S. Department of Energy by Lawrence Livermore National Laboratory under Contract DE-AC52-07NA27344 (D.R.). The Physics Olympiad program in India is supported by the Government of India, Department of Atomic Energy, under project identification number RTI4001. References Burrows et al. (2015) S. Burrows, I. Gurevych, and B. Stein, The eras and trends of automatic short answer grading, International Journal of Artificial Intelligence in Education 25, 60 (2015). Gao et al. (2024) R. Gao, H. E. Merzdorf, S. Anwar, M. C. Hipwell, and A. R. Srinivasa, Automatic assessment of text-based responses in post-secondary education: A systematic review, Computers and Education: Artificial Intelligence 6, 100206 (2024). Kortemeyer (2023) G. Kortemeyer, Toward AI grading of student problem solutions in introductory physics: A feasibility study, Phys. Rev. Phys. Educ. Res. 19, 020163 (2023). Mok et al. (2025) R. Mok, F. Akhtar, L. Clare, C. Li, J. Ida, L. Ross, and M. Campanelli, Using large language models for grading in education: an applied test for physics, Physics Education 60, 035006 (2025). Kortemeyer et al. (2024) G. Kortemeyer, J. Nöhl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study, Phys. Rev. Phys. Educ. Res. 20, 020144 (2024). Kortemeyer and Nöhl (2025) G. Kortemeyer and J. Nöhl, Assessing confidence in AI-assisted grading of physics exams through psychometrics: An exploratory study, Phys. Rev. Phys. Educ. Res. 21, 010136 (2025). Liu et al. (2026) T. Liu, J. Chatain, L. Kobel-Keller, G. Kortemeyer, T. Willwacher, and M. Sachan, AI-assisted automated short answer grading of handwritten university-level mathematics exam, Teaching Mathematics and its Applications: An International Journal of the IMA 45, 84 (2026). Kortemeyer et al. (2025) G. Kortemeyer, A. Caspar, and D. Horica, Artificial-intelligence grading assistance for handwritten components of a calculus exam (2025), arXiv:2510.05162 [cs.CY] . Vanhoyweghen et al. (2026) A. Vanhoyweghen, V. Holst, M. Mobini, L. Van de Voorde, T. Vanleke, B. Verbruggen, B. Verbeken, A. Algaba, S. Verboven, M.-A. Guerry, F. Van Droogenbroeck, and V. Ginis, Human-in-the-loop LLM grading for handwritten mathematics assessments (2026), arXiv:2603.13083 [cs.CY] . Gandolfi (2025) A. Gandolfi, GPT-4 in education: Evaluating aptness, reliability, and loss of coherence in solving calculus problems and grading submissions, International Journal of Artificial Intelligence in Education 35, 367 (2025). Morris et al. (2025) W. Morris, L. Holmes, J. S. Choi, and S. Crossley, Automated scoring of constructed response items in math assessment using large language models, International Journal of Artificial Intelligence in Education 35, 559 (2025). Parsaeifard et al. (2025) B. Parsaeifard, M. Hlosta, and P. Bergamin, Automated grading of students’ handwritten graphs: A comparison of meta-learning and vision-large language models (2025), arXiv:2507.03056 [cs.LG] . Nath et al. (2025) O. Nath, H. Bathina, M. S. U. R. Khan, and M. M. Khapra, Can vision-language models evaluate handwritten math? (2025), arXiv:2501.07244 [cs.CV] . Seong et al. (2026) J. Seong, W. Liermann, M. Kim, J.-h. Shin, and S. Lim, When VLMs ’fix’ students: Identifying and penalizing over-correction in the evaluation of multi-line handwritten math OCR (2026), arXiv:2604.22774 [cs.CY] . Lu et al. (2024) P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, MathVista: Evaluating mathematical reasoning of foundation models in visual contexts, in International Conference on Learning Representations (ICLR) (2024). Latif and Zhai (2024) E. Latif and X. Zhai, Fine-tuning ChatGPT for automatic scoring, Computers and Education: Artificial Intelligence 6, 100210 (2024). Lee et al. (2024) G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, Applying large language models and chain-of-thought for automatic scoring, Computers and Education: Artificial Intelligence 6, 100213 (2024). Jiang and Bosch (2024) L. Jiang and N. Bosch, Short answer scoring with GPT-4, in Proceedings of the Eleventh ACM Conference on Learning @ Scale, L@S ’24 (Association for Computing Machinery, New York, NY, USA, 2024) p. 438–442. Pečuchová et al. (2025) J. Pečuchová, Ľ. Benko, and M. Drlík, Automated grading of open-ended questions in higher education using GenAI models, International Journal of Artificial Intelligence in Education 35, 3813 (2025). Williamson et al. (2012) D. M. Williamson, X. Xi, and F. J. Breyer, A framework for evaluation and use of automated scoring, Educational Measurement: Issues and Practice 31, 2 (2012). Kane (2013) M. T. Kane, Validating the interpretations and uses of test scores, Journal of Educational Measurement 50, 1 (2013). Bennett and Bejar (1998) R. E. Bennett and I. I. Bejar, Validity and automated scoring: It’s not only the scoring, Educational Measurement: Issues and Practice 17, 9 (1998). Livingston and Lewis (1995) S. A. Livingston and C. Lewis, Estimating the consistency and accuracy of classifications based on test scores, Journal of Educational Measurement 32, 179 (1995). Jauhiainen and Garagorry Guerra (2025) J. S. Jauhiainen and A. Garagorry Guerra, Generative AI in education: ChatGPT-4 in evaluating students’ written responses, Innovations in Education and Teaching International 62, 1377 (2025). Tian et al. (2023) K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. Manning, Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback, in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, edited by H. Bouamor, J. Pino, and K. Bali (Association for Computational Linguistics, Singapore, 2023) p. 5433–5442. Li et al. (2025) Y. Li, M. Raković, N. Srivastava, X. Li, Q. Guan, D. Gašević, and G. Chen, Can AI support human grading? Examining machine attention and confidence in short answer scoring, Computers and Education 228, 105244 (2025). Mozannar and Sontag (2020) H. Mozannar and D. Sontag, Consistent estimators for learning to defer to an expert, in Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 119, edited by H. Daumé I and A. Singh (PMLR, 2020) p. 7076–7087. Tang et al. (2026) X. Tang, G. A. Ambrose, and Y. Cheng, Designing reliable LLM-assisted rubric scoring for constructed responses: Evidence from physics exams (2026), arXiv:2604.12227 [cs.AI] .