Paper deep dive
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment
Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fröhlich, Tim Schlippe
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/4/2026, 4:55:27 AM
Summary
This study investigates the alignment between human thesis supervisors' evaluation priorities and an AI-based assessment system (RubiSCoT). By surveying 84 supervisors across four disciplines, the authors derived empirical criterion weights for 35 assessment criteria. Comparing these to RubiSCoT's default weights revealed substantial divergences, particularly in Introduction and Literature Review chapters. Integrating supervisor-derived weights into RubiSCoT via various calibration configurations (global, degree-level, discipline-level) reduced the mean relative deviation between AI and human scores from 11.18% to 10.85%, though this improvement was not statistically significant. The study concludes that criterion-weight calibration alone does not substantially improve alignment between AI and human assessments, as human supervisors showed stronger inter-agreement (4.44% deviation) than AI did with humans.
Entities (17)
Relation Signals (11)
thesis supervisors → provide → supervisor-derived criterion weights
confidence 95% · We surveyed 84 thesis supervisors... and collected weighting data for 35 thesis assessment criteria.
RubiSCoT → uses → default criterion weights
confidence 95% · Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences
supervisor-derived criterion weights → divergefrom → default criterion weights
confidence 92% · Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights.
RubiSCoT → evaluates → thesis components
confidence 90% · RubiSCoT is an LLM-based thesis assessment system that evaluates theses using a structured analytic rubric. The system assesses thesis components separately
Discipline-Level Calibration → isatypeof → calibration configuration
confidence 90% · Discipline-Level Calibration: We calculated separate criterion-weight distributions for technical and non-technical disciplines.
Global Calibration → isatypeof → calibration configuration
confidence 90% · Global Calibration: We calculated one shared supervisor-derived criterion-weight distribution using all survey responses.
Degree-Level Calibration → isatypeof → calibration configuration
confidence 90% · Degree-Level Calibration: We calculated separate criterion-weight distributions for bachelor-level and master-level supervision.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.
Tags
Links
- Source: https://arxiv.org/abs/2608.00717v1
- Canonical: https://arxiv.org/abs/2608.00717v1
Trouble viewing inline? Open PDF directly →
Full Text
39,493 characters extracted from source content.
Expand or collapse full text
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fröhlich and Tim Schlippe IU International University of Applied Sciences, Germany. Email: thorsten.froehlich@iu.org; tim.schlippe@iu.org; Abstract Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although lit- tle empirical evidence exists regarding how thesis supervisors actu- ally prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis super- visors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default cri- terion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration con- figurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substan- tially improve alignment between AI-generated and human assessments. Keywords: AI in education, AI-supported assessment, thesis assessment, criterion weighting, rubric calibration, LLM grading, educational measurement, RubiSCoT 1 arXiv:2608.00717v1 [cs.AI] 1 Aug 2026 1 Introduction Rubrics are widely used in higher education to clearly define how student work is evaluated, support transparency, and improve consistency in assess- ment [2]. In analytic rubrics, student work is evaluated across multiple criteria rather than through a single holistic judgment, thereby supporting more detailed feedback and more transparent assessment decisions [2]. In thesis assessment, rubric criteria commonly include dimensions such as problem for- mulation, literature engagement, methodological rigor, argumentation quality, interpretation of results, and reflection on limitations. Despite the formalization provided by rubrics, thesis assessment remains a complex evaluative process that cannot be fully reduced to written criteria alone [3, 4]. Supervisors interpret rubric criteria through disciplinary expec- tations, assessment experience, and implicit academic standards developed through practice [4, 5]. In both human and AI-supported assessment systems, criterion weights determine how strongly individual rubric criteria contribute to assessment out- comes and which evaluation criteria are prioritized during assessment [1, 6]. However, these weights are typically defined through expert judgment or implementation assumptions rather than empirical evidence regarding actual supervisor weighting behavior. As a result, AI assessment systems may apply rubrics consistently while still diverging systematically from human assessment behavior [1, 7]. Previous research has shown that both human supervisors and AI- supported assessment systems may apply shared rubric criteria differently due to interpretive variability, implicit standards, and calibration limitations [7– 11]. While calibration approaches have been proposed to improve alignment between AI-generated and human evaluations [7], little research has system- atically examined how thesis supervisors prioritize individual rubric criteria when evaluating bachelor’s and master’s theses. Without empirical evidence regarding supervisor weighting behavior, AI assessment systems risk relying on default criterion weights that may not reflect actual supervisory priorities. To address this gap, our study presents an empirical analysis of criterion weights reported by thesis supervisors during thesis assessment. Based on a survey of 84 thesis supervisors, we derive criterion weights for rubric criteria distributed across five common thesis components: Introduction, Literature Review, Research Design, Results and Discussion, and Conclusion. We then compare these empirically derived criterion weights with the default criterion weights used in the AI assessment system RubiSCoT [1]. To evaluate the practical implications of the observed divergences, we inte- grate the empirically derived criterion weights into the AI assessment system and evaluate multiple calibration configurations on a corpus of 80 German- language theses. For each configuration, we measure the relative deviation between AI-generated and supervisor-assigned assessments. The contributions of this study are as follows: 2 • We provide the first empirical dataset of criterion weights reported by thesis supervisors for thesis assessment across five thesis chapters and four aca- demic disciplines, and release the underlying survey data for reproducibility and future research. 1 • We present a criterion-level analysis of divergences between empirically derived and default AI criterion weights. • We evaluate multiple supervisor-derived weighting configurations to analyze their alignment with supervisor-assigned assessments. The remainder of this paper is structured as follows: Section 2 reviews related work on rubric-based thesis assessment, human assessment behavior, and AI-supported assessment systems. Section 3 describes the survey design, corpus construction, and experimental methodology. Section 4 presents the supervisor-derived criterion weights, human–AI weight divergences, and assess- ment alignment results across weighting configurations. Section 5 discusses implications, limitations, and interpretation of the findings. Section 6 concludes the paper and outlines directions for future research. 2 Related Work This section reviews prior research on rubric-based thesis assessment, human assessment behavior, and AI-supported assessment systems. 2.1 Rubric-Based Thesis Assessment Rubrics are widely used in higher education to support transparency, con- sistency, and structured evaluation [2]. In thesis assessment, analytic rubrics commonly evaluate multiple dimensions of academic work, including prob- lem formulation, literature engagement, methodological rigor, argumentation quality, interpretation of results, and reflection on limitations. Despite the formalization provided by rubrics, thesis assessment remains partly interpretive and shaped by disciplinary expectations, assessment expe- rience, and implicit academic standards [3–5, 9]. Consequently, assessment decisions may still differ substantially despite shared rubric structures. While prior research has examined rubric reliability, assessment consis- tency, and rubric design, comparatively little work has investigated how supervisors prioritize individual rubric criteria during thesis assessment. To our knowledge, no previous study has systematically quantified criterion- level weighting behavior across thesis chapters and academic disciplines or examined how empirically derived criterion weights affect alignment between AI-generated and human assessments. 1 https://github.com/thorstenfroehlich/RubiSCoT_Validation_Study_Data 3 2.2 Human Assessment Behavior and Criterion Weighting Research on assessment behavior has shown that assessors frequently rely on implicit standards and personal benchmarks when evaluating student work [4, 9]. Consequently, supervisors may differ substantially in how strongly they prioritize specific criteria despite using the same rubric. Previous studies have shown that assessors may vary in their evaluation of analytical depth, originality, methodological rigor, and theoretical con- tribution [8]. Such differences are particularly relevant in thesis assessment because thesis evaluation involves multidimensional and interpretive judgment. Prior research on AI-supported assessment has similarly reported substantial variability between human evaluators. Existing research has primarily focused on assessment reliability, rubric interpretation, and evaluator consistency rather than explicitly measuring criterion-level weighting behavior. As a result, little empirical evidence exists regarding how supervisors allocate importance across rubric criteria when assessing bachelor’s and master’s theses. Understanding these weighting deci- sions is particularly relevant for AI-supported assessment systems that rely on predefined criterion weights. 2.3 AI Assessment Systems and Calibration AI-supported assessment systems have increasingly been applied to rubric- based evaluation tasks in higher education [6, 10, 11]. Recent advances in large language models have further expanded the capabilities of AI-supported assess- ment for evaluating complex written assignments, including essays, reports, and theses [10, 11]. Most AI assessment systems use criterion weights to determine how strongly individual evaluation criteria contribute to assessment outcomes. However, these weights are typically defined through expert judgment or implementation assumptions rather than empirical evidence regarding actual supervisor weighting behavior [1, 7]. Recent research has emphasized the importance of calibration approaches that align AI-generated assessments more closely with human judgment [7]. Existing calibration work has primarily focused on prompt engineering, score normalization, and supervised fine-tuning to improve agreement between AI- generated and human evaluations [7, 10]. In contrast, comparatively little attention has been paid to calibration through criterion-weight adjustment. This represents an important research gap as criterion weights determine how criterion-level evaluations are aggregated within rubric-based assessment systems. Without empirical evidence regarding supervisor weighting behavior, AI assessment systems may rely on weighting schemes that do not reflect actual supervisory priorities. Consequently, even technically consistent AI assessment systems may systematically diverge from human assessment behavior. 4 3 Methodology This section describes the survey design, weighting analysis, evaluation corpus, and calibration configurations used in the study. 3.1 Study Design Figure 1 summarizes the overall study design. We first conducted a super- visor survey in which participants allocated criterion weights within thesis assessment rubrics. Based on these responses, we derived aggregated criterion weights across rubric categories. The study then proceeded along two complementary evaluation paths. First, we conducted a weight analysis by comparing the criterion weights reported by thesis supervisors with the default criterion weights used in the AI assessment system. This analysis quantified where human assessment priorities and default AI criterion weights diverged. Second, we integrated the empirically derived criterion weights into the AI assessment system and evaluated whether recalibrated weighting configu- rations produced chapter-level scores that were closer to supervisor-assigned assessments. For this purpose, we executed multiple calibration configurations on a corpus of thesis submissions and compared the resulting AI-generated scores with the corresponding supervisor assessments. Fig. 1 Overview of the study design and evaluation workflow. 5 3.2 Supervisor Survey We conducted an online survey to collect criterion weights reported by thesis supervisors for thesis assessment rubrics. The survey was implemented using Qualtrics and replicated the rubric structure used in the AI assessment system. The rubric consisted of 35 evaluation criteria distributed across five core thesis components: Introduction, Literature Review, Research Design, Results and Discussion, and Conclusion. These components represent common func- tional elements of academic theses rather than fixed chapter titles. While individual theses may use different headings or organize content differently, the corresponding content can generally be mapped to these five areas during assessment. Throughout the remainder of this paper, these components are referred to as thesis chapters for simplicity. To measure relative criterion importance, we used a constant-sum allo- cation design [12]. Participants distributed 100% across the criteria within each chapter rubric, forcing explicit prioritization decisions and avoiding ceil- ing effects commonly observed in Likert-scale importance ratings. All criterion labels and descriptions matched those used in the AI assessment system to ensure direct comparability. Using the original rubric wording also minimized wording differences between the survey instrument and the AI assessment system. After each weighting section, participants could optionally provide qual- itative explanations for their weighting decisions. Demographic information included discipline, supervision experience, grading language, degree levels supervised, and AI tool usage. Participants were recruited through aca- demic mailing lists, professional networks, and social media channels between March 13 and March 25, 2026. The final sample of 84 thesis supervisors included supervisors from Business and Economics (n = 32), Social Sci- ences (n = 25), STEM disciplines (n = 19), and Humanities (n = 6). Two additional participants selected “Other” as their primary discipline. Most participants reported extensive thesis supervision experience and primarily assessed German-language theses. RubiSCoT is an LLM-based thesis assessment system that evaluates theses using a structured analytic rubric. The system assesses thesis components sep- arately, generates criterion-level scores and explanations, and aggregates these scores into chapter-level assessments using predefined criterion weights. In this study, RubiSCoT serves as the AI assessment system whose default criterion weights are compared with supervisor-derived criterion weights. Figure 2 illustrates RubiSCoT’s hierarchical rubric structure. While the system contains both chapter-level and criterion-level weights, the present study focuses exclusively on criterion-level weights since chapter-level weights are often institutionally predefined and cannot easily be modified within oper- ational assessment frameworks. The figure shows the default criterion-weight distribution for the Introduction chapter as an example. The hierarchical weighting structure forms the basis for the calibration configurations evaluated in this study. The default criterion weights used in 6 Fig. 2 Hierarchical structure of RubiSCoT’s rubric weighting system. Shown weights are for illustration only. the AI assessment system were originally defined during the development of RubiSCoT based on expert judgment and iterative rubric design decisions [1]. These weights originate from the production version of RubiSCoT and are not reported in the original system description. The complete default criterion- weight vector for all five rubrics is available in our accompanying repository (see Section 1). They reflect the assumed relative importance of rubric cri- teria during thesis assessment but had not previously been validated against empirical supervisor weighting data. 3.3 Weight Analysis We aggregated the supervisor responses to derive mean criterion weights for all rubric criteria. For each criterion, we calculated the mean weight, standard deviation, and coefficient of variation across all supervisors. These statistics describe both the average importance assigned to each criterion and the degree of agreement among supervisors. Unlike standard deviation alone, the coeffi- cient of variation normalizes variability relative to the mean weight, thereby enabling direct comparison across criteria with different average weights. Lower values indicate stronger agreement among supervisors. To quantify differences between supervisor-derived criterion weights and the default criterion weights used in the AI assessment system, we calculated the mean absolute error (MAE): MAE (%) = 1 n n X i=1 |w supervisor,i − w default,i |× 100(1) 7 where n denotes the number of criteria within a rubric, w supervisor,i represents the aggregated supervisor-derived criterion weight, and w default,i represents the default criterion weight used in the AI assessment system. Lower MAE values indicate stronger alignment between supervisor-derived and default criterion weights. We used MAE rather than relative percentage deviation because criterion weights sum to 100% within each rubric. Relative deviations can be mislead- ing for small weights, where modest absolute differences may produce large percentage changes. MAE therefore provides a more stable and interpretable measure of divergence between two criterion-weight distributions. This analysis enabled a criterion-level comparison between human assess- ment priorities and the default criterion-weight distribution of the AI assessment system. 3.4 Evaluation Corpus The evaluation corpus consists of 80 German-language theses from IU Inter- national University of Applied Sciences. The corpus includes 20 BA, 20 B.Sc., 20 MA, and 20 M.Sc. theses and covers both technical and non-technical disci- plines. The theses varied substantially in length, ranging from 47 to 335 pages (average: 110.1 pages; median: 99.5 pages). Each thesis was previously assessed by one or two human supervisors whose evaluations serve as the reference for comparison. For each thesis, chapter- level reference scores were available from the rubric-based assessment forms used during thesis evaluation. These chapter-level scores served as the human reference evaluations for all alignment analyses. The survey was anonymous and distributed through public channels, whereas the corpus was graded by supervisors of IU International University of Applied Sciences during routine thesis assessment without reference to RubiSCoT or its criterion weights. Both data sources therefore reflect thesis assessment practice, but they do not model the preferences of identical individual supervisors. For the subset of 70 theses with two available supervisor evaluations, we additionally calculated inter- supervisor chapter-level deviations to establish a human assessment baseline for comparison with AI-generated assessments. Across all chapter evaluations, the mean relative deviation between first and second supervisors was 4.44%. 3.5 Calibration Configurations We evaluated five weighting configurations. 3.5.1 Default Configuration The original default criterion weights of the AI assessment system were applied to the complete thesis corpus. 8 3.5.2 Global Calibration We calculated one shared supervisor-derived criterion-weight distribution using all survey responses. This configuration was applied to the complete thesis corpus. 3.5.3 Degree-Level Calibration We calculated separate criterion-weight distributions for bachelor-level and master-level supervision. The bachelor-level configuration was applied only to BA and B.Sc. theses, while the master-level configuration was applied only to MA and M.Sc. theses. 3.5.4 Discipline-Level Calibration We calculated separate criterion-weight distributions for technical and non- technical disciplines. The technical configuration was applied only to B.Sc. and M.Sc. theses, while the non-technical configuration was applied only to BA and MA theses. 3.5.5 Combined Calibration We calculated four specialized criterion-weight distributions: BA, B.Sc., MA, and M.Sc. Each configuration was applied exclusively to the corresponding thesis group. 3.6 Chapter-Level Alignment Evaluation For each calibration configuration, we executed the AI assessment system on the corresponding subset of theses. The system generated chapter-level assessment scores using the respective criterion-weight distribution. To isolate the effect of criterion-weight calibration, we focused the primary evaluation on chapter-level score alignment rather than final thesis grades. Chapter-level scores provide a more direct measure of the impact of criterion weights, whereas final thesis grades additionally depend on institution-specific chapter weights and grading systems. We therefore compared the AI-generated chapter-level scores with the corresponding supervisor-assigned chapter-level scores across the five the- sis chapters. Evaluation focused on the mean relative deviation between AI-generated and supervisor-assigned chapter scores: Relative Deviation (%) = |s AI − s human | s human × 100(2) where s AI denotes the AI-generated chapter-level score and s human denotes the corresponding supervisor-assigned chapter-level score. Lower deviation indicates stronger alignment with human scoring behavior. To evaluate whether the improvement of the best-performing calibration configuration over the default configuration was statistically significant, we 9 conducted significance testing on thesis-level aggregated chapter deviations. Since the same theses were evaluated across configurations, all tests used paired observations. We additionally calculated Cohen’s d to estimate effect sizes. Statistical significance was evaluated at the α = 0.05 level. To avoid treating chapter evaluations from the same thesis as statisti- cally independent observations, statistical testing was conducted on thesis-level aggregated deviations rather than on individual chapter-level deviations. For each thesis, the mean relative chapter-level deviation across all five chapters was calculated separately for each calibration configuration. 4 Results This section presents the supervisor-derived criterion weights, human-AI weighting divergences, and chapter-level alignment results across calibration configurations. 4.1 Human-AI Weight Divergence We compared the criterion weights reported by thesis supervisors with the default criterion weights used in the AI assessment system. Table 1 summa- rizes the mean absolute error (MAE) between both weighting schemes across the five thesis rubrics. The reported MAE values represent the average abso- lute difference between supervisor-derived and default criterion weights within each rubric. Lower MAE values therefore indicate stronger agreement between human and AI weighting priorities. Table 1 Mean Absolute Error Between Supervisor-Derived and Default Criterion Weights ChapterMean Absolute Error (%) Largest Criterion Difference (%) Introduction6.92Academic Positioning (−16.0) Literature Review6.72Analytical Synthesis (−15.5) Research Design1.70Reflexivity (+4.8) Results and Discussion2.58Derivation (−4.9) Conclusion3.59Synthesis (−8.4) Overall4.30 The largest divergences occurred in the Introduction and Literature Review rubrics, with MAE values of 6.92% and 6.72%, respectively. In both cases, the default AI assessment system emphasized higher-order analytical criteria more strongly than supervisors did. Within the Introduction rubric, the default AI assessment system allocated substantially more weight to the Academic Positioning criterion (−16.0%) than supervisors. Supervisors instead assigned greater importance to problem framing and relevance of literature sources. 10 Similarly, within the Literature Review rubric, the AI assessment system emphasized Analytical Synthesis (−15.5%) more strongly than supervisors, who instead prioritized relevance and coverage of literature sources. In contrast, the Research Design (MAE = 1.70%) and Results and Dis- cussion (MAE = 2.58%) rubrics showed comparatively strong alignment, emphasizing methodological quality, analytical rigor, and evidence-based interpretation of results. Overall, the findings indicate that the largest human–AI weighting diver- gences occurred in more interpretive evaluation criteria, whereas methodolog- ical and technical criteria showed substantially stronger alignment. 4.2 Chapter-Level Alignment Results We evaluated whether supervisor-derived criterion weights improved the align- ment between AI-generated and supervisor-assigned chapter-level assessments across the five thesis chapters. Table 2 summarizes the evaluated calibration configurations, while Table 3 reports the corresponding mean relative chapter-level deviations. Table 2 Overview of the Evaluated Calibration Configurations Configuration Supervisor-derived Weights Evaluated Theses DefaultDefault AI system weightsFull Corpus GlobalAll supervisorsFull Corpus Degree-LevelBachelor supervisorsBA + B.Sc. Master supervisorsMA + M.Sc. Discipline-Level Technical supervisorsB.Sc. + M.Sc. Non-technical supervisorsBA + MA CombinedNon-technical BA supervisors BA Technical B.Sc. supervisorsB.Sc. Non-technical MA supervisors MA Technical M.Sc. supervisorsM.Sc. Table 3 Chapter-Level Deviations Between AI-Generated and Supervisor-Assigned Scores ChapterDefaultGlobal Degree-Level Discipline-Level CombinedSupervisors Introduction9.949.769.909.559.912.51 Literature Review11.9612.3612.7812.3012.143.78 Research Design11.5011.5011.1711.9511.205.23 Results and Discussion 10.9110.6710.3510.379.814.84 Conclusion11.61 11.5212.1811.1611.175.85 Overall11.1811.1611.2811.0710.854.44 11 4.2.1 Results with Global Calibration The Global configuration applied one shared supervisor-derived criterion- weight distribution across the complete thesis corpus. Compared to the Default configuration (11.18%), the Global configura- tion achieved a slightly lower overall mean relative chapter-level deviation of 11.16%. Improvements were primarily observed for the Introduction, Results and Discussion, and Conclusion chapters, while deviations in Literature Review increased slightly. 4.2.2 Results with Degree-Level Calibration The Degree-Level configuration used separate criterion-weight distributions for bachelor-level and master-level theses. The bachelor-level configuration was applied to BA and B.Sc. theses, while the master-level configuration was applied to MA and M.Sc. theses. The Degree-Level configuration produced an overall mean relative chapter- level deviation of 11.28%, slightly higher than the Default configuration. Improvements in Research Design and Results and Discussion were offset by larger deviations in Literature Review and Conclusion. 4.2.3 Results with Discipline-Level Calibration The Discipline-Level configuration used separate criterion-weight distributions for technical and non-technical disciplines. The technical configuration was applied to B.Sc. and M.Sc. theses, while the non-technical configuration was applied to BA and MA theses. The Discipline-Level configuration achieved an overall mean relative chapter-level deviation of 11.07%, representing the second-best result among all evaluated configurations. The strongest improvements were observed for Introduction, Results and Discussion, and Conclusion, whereas Research Design showed slightly larger deviations than under the Default configuration. 4.2.4 Results with Combined Calibration The Combined configuration used separate criterion-weight distributions for BA, B.Sc., MA, and M.Sc. theses. Each configuration was applied exclusively to the corresponding thesis group. The Combined configuration achieved the strongest overall alignment, reducing the mean relative chapter-level deviation from 11.18% to 10.85%. This corresponds to a relative improvement of 2.95%. The strongest improvement was observed for Results and Discussion, where the relative deviation decreased from 10.91% to 9.81%, corresponding to a relative improvement of 10.1%. Additional improvements were observed for Research Design and Conclusion. 12 4.2.5 Overall Configuration Comparison Across all evaluated configurations, the Combined calibration achieved the low- est overall chapter-level deviation (10.85%), followed by the Discipline-Level (11.07%) and Global (11.16%) configurations. The Degree-Level configuration produced the largest overall deviation (11.28%). Despite these improvements, all AI configurations remained substantially less aligned with supervisor assessments than human supervisors were with each other. The mean inter-supervisor relative deviation was 4.44%, compared to 10.85% for the best-performing AI configuration. Our paired t-test (α = .05) comparing the Default and Combined configura- tions found no statistically significant difference in mean relative chapter-level deviation (t(79) = 0.75, p = .458, d = 0.08). Although the Combined configuration achieved the lowest overall devia- tion, the observed reduction remained within the range of sampling variability (95% CI [-0.57%, 1.26%]). These results indicate that the observed improve- ment is small relative to the variability across theses and therefore cannot be distinguished from random variation at the α = .05 significance level. 5 Discussion This section interprets the findings and discusses their implications for AI- supported thesis assessment. 5.1 Human-AI Weight Divergence The comparison between supervisor-derived and default AI criterion weights revealed a clear pattern. Divergences were largest in the more interpretive Introduction and Literature Review rubrics, whereas stronger alignment was observed in Research Design. This pattern suggests that human supervisors and AI assessment systems differ most strongly in areas requiring interpretive judgment, whereas methodological dimensions appear easier to operationalize consistently. The findings further indicate that the default AI weighting scheme places greater emphasis on higher-order analytical criteria, whereas supervi- sors assigned greater importance to problem framing, source relevance, and methodological quality. This suggests that default AI weighting schemes may reflect assessment priorities that differ from those of human supervisors. 5.2 Impact of Empirical Weight Alignment The results show that empirically derived criterion weights did not pro- duce a detectable improvement between AI-generated and supervisor-assigned assessments. The Combined configuration performed best, reducing the mean relative chapter-level deviation from 11.18% to 10.85%. However, the improvement was small, not statistically significant, and sub- stantially lower than the agreement observed between human supervisors. The 13 mean inter-supervisor relative deviation was 4.44%, compared to 10.85% for the Combined configuration. These findings suggest that reducing criterion-weight divergence alone is insufficient to substantially improve alignment between AI-generated and human assessments. One possible explanation is that differences arise less from criterion weights than from how human supervisors and AI systems interpret criteria, evaluate evidence, and form assessment judgments. The strongest improvements were observed in the Results and Discussion chapter. This may indicate that weighting adjustments are most effective in chapters where evaluation criteria are well defined and closely connected to explicit analytical tasks. In contrast, the stable deviations observed in Litera- ture Review suggest that factors beyond criterion weighting play a larger role in chapters requiring broader interpretive judgment. 5.3 Implications for AI-Supported Assessment Systems The findings have two implications for AI-supported assessment systems: First, the observed divergences indicate that default criterion weights should not be assumed to reflect actual supervisory priorities without empirical validation. Second, the limited impact of weight calibration suggests that AI-supported assessment systems should not be calibrated through criterion weights alone. Future calibration should also address rubric interpretation, evidence eval- uation, and assessment reasoning, for example through improved prompts, reasoning traces, or human-in-the-loop calibration. 5.4 Limitations Several limitations should be considered when interpreting the findings: First, the study relies on self-reported supervisor weights rather than directly observed assessment behavior. Although the constant-sum allocation design forced supervisors to make explicit trade-offs between criteria, actual assessment decisions may involve more dynamic weighting processes. In addi- tion, self-reported weights may be influenced by social desirability bias or normative assumptions regarding good assessment practice. Second, the supervisor sample was geographically concentrated in Germany and primarily involved German-language thesis assessment. As assessment cul- tures and thesis expectations vary across national and institutional contexts, the findings may not fully generalize to other higher education systems. Third, some subgroup analyses were based on relatively small samples, particularly for specialized combinations in the Combined configuration, which may reduce the stability of highly specialized criterion-weight distributions. Fourth, the study evaluates alignment at the chapter level rather than indi- vidual rubric criteria. While chapter-level assessments provide a meaningful 14 basis for comparing human and AI evaluations, criterion-level assessment data from human supervisors were not available for the thesis corpus. Finally, the study focused exclusively on criterion weighting and did not evaluate other components of AI-supported assessment systems, such as rubric interpretation, textual reasoning quality, or cross-chapter consistency analysis. 6 Conclusion and Future Work This section summarizes the findings and discusses possible future work. 6.1 Conclusion This study presented an empirical investigation of criterion weighting in thesis assessment. Using a survey of 84 thesis supervisors across four academic disci- plines, we collected criterion weights for 35 rubric criteria distributed across the five core thesis components: Introduction, Literature Review, Research Design, Results and Discussion, and Conclusion. The findings revealed substantial divergences between supervisor-derived and default AI criterion weights, particularly in the Introduction and Liter- ature Review rubrics, whereas methodological criteria showed substantially stronger alignment. To evaluate the practical implications of these divergences, we inte- grated supervisor-derived criterion weights into the AI assessment system and evaluated multiple weighting configurations on 80 German-language theses. The best-performing configuration reduced the mean relative chapter-level deviation from 11.18% to 10.85%. However, the improvement was not sta- tistically significant, and human supervisors showed substantially stronger agreement with each other (4.44%) than with AI-generated assessments. Over- all, the findings indicate that criterion-weight calibration alone is insufficient to substantially improve human–AI alignment in thesis assessment. This study contributes an empirical dataset of supervisor-derived crite- rion weights, a reproducible methodology for human–AI weighting comparison, and an experimental framework for evaluating assessment alignment in AI-supported thesis assessment systems. 6.2 Future Work Several directions for future research emerge from these findings. First, future studies could investigate whether criterion-weight distribu- tions differ across countries, institutions, or disciplinary cultures, and whether AI-supported assessment systems require context-specific weighting strategies. Second, future research should investigate additional sources of human–AI divergence, including differences in rubric interpretation, evidence evaluation, and assessment reasoning. 15 Third, future research should investigate whether limited variation in criterion-level scores constrains the effect of criterion-weight calibration on chapter-level scores. Fourth, future work could explore dynamic weighting approaches in which criterion importance adapts to thesis characteristics or intermediate evaluation signals rather than relying on fixed weighting configurations. Finally, future research should investigate how calibration strategies influ- ence fairness, transparency, and trust in AI-supported assessment systems. Similar calibration challenges may also emerge in other AI-supported assess- ment domains, including automatic short-answer assessment [13]. Acknowledgments We thank the 84 thesis supervisors who contributed their expertise by partici- pating in the survey. This research was conducted as part of the FAIRGRADE project at the Research Institute Artificial Intelligence (AI) of IU International University of Applied Sciences. References [1] Fröhlich, T., Schlippe, T.: RubiSCoT: A Framework for AI-Supported Academic Assessment. In: Schlippe, T., Cheng, E.C., Wang, T. (eds.) Artificial Intelligence in Education Technologies: New Development and Innovative Practices. AIET 2025, vol. 279. Springer, Singapore (2026). https://doi.org/10.1007/978-981-95-4423-3_20 [2] Brookhart, S.M.: Appropriate Criteria: Key to Effective Rubrics. Frontiers in Education3, 22 (2018) [3] Mullins, G., Kiley, M.: ’It’s a PhD, not a Nobel Prize’: How Experienced Examiners Assess Research Theses. Studies in Higher Education27(4), 369–386 (2002) [4] O’Donovan, B., Price, M., Rust, C.: Know What I Mean? Enhancing Student Understanding of Assessment Standards and Criteria. Teaching in Higher Education9(3), 325–335 (2004) [5] Bloxham, S., Boyd, P.: Accountability in Grading Student Work: Secur- ing Academic Standards in a Twenty-First Century Quality Assurance Context. British Educational Research Journal38(4), 615–634 (2012). https://doi.org/10.1080/01411926.2011.569007 [6] Shermis, M.D., Burstein, J.: Handbook of Automated Essay Evaluation: Current Applications and New Directions. Routledge, New York (2013) [7] Hashemi, M., Savelka, J., Olshefski, A., Ungar, L., Cohn, J., Rose, C.: LLM-Rubric: A Multidimensional, Calibrated Approach to Automated 16 Evaluation of Natural Language Texts. In: The 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), p. 5895–5911 (2024). https://doi.org/10.18653/v1/2024.acl-long.321 [8] Lundström, M., Lindberg, O., Olsson, U.: Different Profiles for the Assess- ment of Student Theses in Teacher Education. Higher Education83, 1183–1198 (2021) [9] Gingerich, A., Regehr, G., Eva, K.W.: Sources of Variability in Assess- ment. BMC Medical Education22(1), 670 (2022). https://doi.org/10. 1186/s12909-022-03724-2 [10] Pack, A., Maloney, J.: Using LLMs for Automated Feedback and Scor- ing of Student Writing: A Case of Overpromise and Caution. Computers and Education: Artificial Intelligence6, 100212 (2024). https://doi.org/ 10.1016/j.caeai.2024.100212 [11] Mathews, S., Sundaram, D.: Evaluating LLM-Generated Feedback on Programming Tasks: A Comprehensive Study in STEM Education. Computers and Education: Artificial Intelligence6 (2025). In press [12] Skedgel, C., Regier, D.A.: Constant-Sum Paired Comparisons for Eliciting Stated Preferences: A Tutorial. The Patient8(2), 155–163 (2015). https: //doi.org/10.1007/s40271-014-0077-9 [13] Schlippe, T., Sawatzki, J.: Cross-Lingual Automatic Short Answer Grad- ing. In: The 2nd International Conference on Artificial Intelligence in Education Technology (AIET 2021), Wuhan, China (2021) 17