Paper deep dive
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/20/2026, 4:22:26 AM
Summary
This paper introduces THPT-Ladder, a benchmark designed to evaluate large language models on Vietnam's 2025 National High School Graduation Examination using its specific convex marking scheme. The study highlights that standard accuracy metrics fail to capture the 'partial-credit gap' inherent in the exam's non-additive grading (Part II), where identifying three out of four correct statements yields only 0.50 points instead of the proportional 0.75. By scoring eight models against official keys and human cohort distributions, the authors demonstrate that this gap significantly alters model rankings and apparent competence, with penalties ranging from 0.020 to 0.159 points per question.
Entities (7)
Relation Signals (5)
THPT-Ladder → evaluates → Vietnam's National High School Graduation Examination
confidence 95% · We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students.
Standard Accuracy Metrics → failstocapture → Partial-Credit Gap
confidence 93% · This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme.
Part II → uses → convex marking scheme
confidence 92% · In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex...
Qwen3.5-27B → experiences → Partial-Credit Gap
confidence 90% · For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile...
Decision 764/QĐ-BGDĐT → defines → Part II
confidence 85% · Decision 764/QĐ-BGDĐT [5] now defines three question formats... four true/false statements marked together as one question...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.
Tags
Links
- Source: https://arxiv.org/abs/2608.18336v1
- Canonical: https://arxiv.org/abs/2608.18336v1
Trouble viewing inline? Open PDF directly →
Full Text
36,657 characters extracted from source content.
Expand or collapse full text
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam’s 2025 Convex Marking Scheme Nguyen Quoc Hung Affiliation: Corresponding author Nguyen Dang Minh Le Nhu QuynhTran Khanh Linh, Nguyen Kieu LinhPosts and Telecommunications Institute of Technology, Hanoi, Vietnamnguyenquochung.workvn@gmail.com Abstract When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam’s National High School Graduation Examination demonstrates the cost of this substitution. In Part I of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part I accounts for 4.00 of the exam’s 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part I question than proportional credit. This shortfall changes a model’s apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model’s accuracy does not predict this penalty. At Claude Sonnet 5’s accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify. Keywords: educational assessment, automated assessment and feedback, partial credit, marking scheme, benchmark, large language models, learning analytics, Vietnamese 1 Introduction The evaluation of language models heavily relies on examinations written for human candidates, ranging from multitask suites assembled out of school exams [3, 4] to single-country benchmarks [1, 2]. However, almost all of these benchmarks keep the questions but discard the marking scheme. They score each item right or wrong and report accuracy. This substitution assumes that partial knowledge is always worth proportional credit. In reality, an examination’s marking scheme encodes exactly how much partial knowledge is worth. Replacing that scheme with flat accuracy credits a model for every fragment it answers correctly, allowing it to appear highly competent on an exam it would have actually failed. 0246810mark out of 10.00 (282,519 candidates)pass 5.00SDSS1.758 models6.50–9.50 Figure 1: Marks of the 282,519 candidates who sat Economics & Law in 2026, with all eight models and a fixed answer string placed on the same scale. Each marker is drawn at its mark, with height proportional to the number of candidates who obtained that mark. On this exam the best fixed string is SDSS, which earns 1.75 without reading any question. Scoring under the official rules locates a model within the population the examination was written to assess, an outcome that standard accuracy cannot achieve. Vietnam’s National High School Graduation Examination (Kỳ thi tốt nghiệp trung học phổ thông, THPT) makes the cost of ignoring the marking scheme measurable. The examination was sat by 1.13 million candidates in 2025. The Ministry of Education and Training publishes the exams, the answer keys, and the binding rules that dictate exactly how each format is marked. These rules changed in 2025. The ministry restructured the examination, and Decision 764/QĐ-BGDĐT [5] now defines three question formats: a four-option item worth 0.25 points, four true/false statements marked together as one question, and a short-answer item with no options. The exams print these as PH`N I, PH`N I and, where a subject carries it, PH`N I, and we refer to them as Part I, Part I and Part I. No existing Vietnamese benchmark covers the two new formats. The largest existing dataset, VNHSGE, draws on exams from 2019 to 2023, predating the reform entirely. Part I makes the value of partial knowledge visible because it is the only format that uses a non-additive grading scheme. The number of a question’s four statements judged correctly indexes a convex step function ladder=(0,0.10,0.25,0.50,1.00)ladder=(0,0.10,0.25,0.50,1.00). Getting three out of four statements correct pays 0.50 points rather than proportional 0.75 points (Fig. 3). Where this format appears, it carries 4.00 of an exam’s 10.00 points. Additive credit would pay 0.25aq0.25a_q where aqa_q of the four statements are judged correctly, while the state pays ladder[aq]ladder[a_q]. We define this difference as the partial-credit gap, or shortfall. This gap measures the shape of a model’s knowledge. A model that is broadly but imperfectly right will fall into the punished zone, while a model that either sweeps a question entirely or fails it completely will not. The shortfall is therefore a calibration signal, not just a second accuracy metric (Section 4). Reporting this shortfall requires a baseline, and the standard approaches fail. Marking the four statements additively produces a score the ministry would never award and erases the shortfall by construction. Alternatively, calling 0.25 the chance level ignores what a convex scheme does to a guesser. Flipping a coin for every statement has an expected value of ∑k=04(4k)2−4ladder[k]=4.9016=0.30625, _k=0^4 4k2^-4\,ladder[k]= 4.9016=0.30625, (1) compared to 0.25 for a traditional four-option item. The ladder does suppress lucky full marks, as all four statements must land correctly, a 1 in 16 chance. However, it actually raises the guesser’s expected payout by 22.5% above 0.25. If we propagate Eq. 1 through each subject’s published structure, the whole-exam random baseline varies. It runs from 19.75% of the total scale in Mathematics to 27.25% in the five subjects that contain 24 Part I items and no Part I. A single “better than chance” threshold therefore misrepresents both the value of random guessing and the varying structure of the exams (Section 3). We resolve these errors by applying the ministry’s own published rules. We introduce THPT-Ladder, named after the school level trung học phổ thông. The corpus contains 632 items from 21 exams across 11 subjects, with every answer drawn directly from the ministry’s official key (Section 3). Because an exam’s difficulty fluctuates between years, we report every score as a percentile among the candidates who sat that exact exam (Fig. 1, Section 4). This cohort is the population the examination was designed for. Since the state publishes the score distribution, we can interpret a model’s mark on the very scale the examination uses. We evaluated eight models under this rule to measure the substitution cost. The test included three open-weight models and five closed models from two vendors. Qwen3.5-27B judged 92.6% of Part I statements correctly but completed only 81.0% of Part I questions. The ladder therefore paid it 0.884 points per question, whereas statement-by-statement credit would pay 0.926. Standard benchmarks report the higher of these two figures, while the ministry awards the lower. By using the published cohort marks, we can read this difference as a drop in standing rather than just a decimal. On the 2025 History exam, this penalty drops the model from the 90th percentile to the 77th. The shortfall runs from 0.020 to 0.159 points per question. It is tempting to read this ordering as the ladder heavily taxing knowledge that is broad but scattered. The closed models show that this interpretation is too simple. When ordered by statement accuracy, the shortfall is not monotone. Accuracy simply does not determine what a non-additive scheme pays (Section 5). This effect is not specific to any particular model. The shortfall follows mathematically from the ladder, and the guessing floor follows from how the keys are constructed. Any system scored under Decision 764 will encounter both. We make four contributions. First, we release THPT-Ladder, a benchmark of 632 items from 21 official exams containing every ministry key and marking rule, allowing models to be scored as candidates rather than by flat accuracy. Second, we formalize the shortfall, which measures how far a convex grading scheme’s award falls below the accuracy a benchmark would normally report, and provide cohort distributions to convert marks into human percentiles. Third, we show that the published keys are not statement-balanced. A fixed answer string can exploit this imbalance to earn 11.07% to 24.25% of an exam without reading a single question. Finally, we release the extraction pipeline used to build the corpus. This includes figure extraction, per-variant key reading, and parsers for the three formats introduced by Decision 764, ensuring that future exams can be added with minimal effort. 2 Related Work Two Vietnamese benchmarks are closely related to this work, but neither can evaluate models under a non-additive grading scheme. VMLU [1] measures subject knowledge across 58 subjects and four levels of education. It draws on assorted material rather than a single examination, and its multiple-choice component grades items strictly as right or wrong. VNHSGE [2] is built directly from the high school graduation exam, featuring over 19,000 multiple-choice items and 300 literary essays. However, these materials come from the curriculum that preceded the 2025 reform, and the benchmark evaluates two closed chat services rather than any open-weight models. Because both benchmarks describe the exam as it stood before 2025 and lack a marking scheme, neither can distinguish a model that gets three out of four statements correct from one that gets all four correct. Vietnam is not unusual in this respect. Curriculum-aligned benchmarks for other lower-resource languages, including LaoBench [6], SinhalaMMLU [7], and an assessment against Nepal’s K-10 curriculum [8], also score items as purely right or wrong. We are not aware of any public dataset that includes items in the new 2025 formats. Scoring models at a finer grain than the whole item is established practice in other domains. SteuerLLM [10] is the closest to our setting, marking German tax-law questions statement by statement. RadSEM [11] decomposes radiology reports into atomic findings and refuses credit when errors exist. PsyScore [12] applies a graded partial-credit item-response model, and CMPhysBench [13] awards non-binary credit over expression trees. However, none of these benchmarks uses a convex credit function where three statements out of four pay 0.50 instead of the proportional 0.75. We are not aware of any prior benchmark that scores language models against a non-additive grading scheme established by a national ministry, let alone one that evaluates the models against the actual human population that took the exam. Researchers usually test option-order sensitivity by synthetically permuting choices [14]. The Vietnamese examination instead distributes 24 or 48 official variants (mã đề) of every exam. These variants are not just reordered copies of each other. Out of the 20 subject-years where we hold the keys, 15 split into groups with entirely different content. We release every published key to support future research into these variations. 3 The Examination and the Corpus The examination’s published structure fixes both the shortfall and the answer-only baseline. Every candidate sits four subjects. Each exam is marked out of 10.00 and, since the 2025 reform, follows the three question formats defined by Decision 764. The ministry sets the marking scheme, and every answer in our corpus comes directly from the ministry’s published keys. Biology, upper-secondary graduation examination in Vietnam, 2025 — PH`N I (Part I) Câu 3. Hình bên thể hiện sự di truyền của 2 tính trạng bao gồm hội chứng nail-patella và hệ nhóm máu ABO ở một gia đình. … Hai gene (N, I) cùng nằm trên NST số 9 và có tần số hoán vị gene là 10%. a) Quần thể người có tối đa 8 kiểu hình liên quan đến 2 tính trạng này. b) Người I-1 tạo ra 2 loại giao tử mang gene quy định về 2 tính trạng này. c) Kiểu hình của người I-3 được quy định bởi 1 trong 5 loại kiểu gene về 2 tính trạng này. d) Cặp vợ chồng I-1 và I-2 sinh con đầu lòng, xác suất để người con này mắc hội chứng nail-patella và có nhóm máu AB là 22,5%. Question 3. The figure shows the inheritance of two traits, nail-patella syndrome and the ABO blood group, in one family. … The two genes lie on chromosome 9 with a recombination frequency of 10%. a) True. At most 8 phenotypes for the two traits exist in the population. b) True. Individual I-1 produces 2 gamete types for these two traits. c) True. The phenotype of I-3 arises from 1 of 5 genotypes. d) True. For the first child of I-1 and I-2, P(nail-patella and blood group AB) == 22.5%. What each model answered(a)(b)(c)(d)keyĐ Claude Opus 5Đ4 of 41.00InternVL3.5-8BĐSĐ3 of 40.50Claude Haiku 4.5ĐSĐS2 of 40.25Qwen3.5-9BSSSĐ1 of 40.10Qwen3.5-27B–0 of 40.00paidone question, every rung of the ladderTHPT-Ladder: 632 items, 21 official examsstmt. accuracymark earnedshortfallClaude Opus 50.024GPT-5.50.020Claude Sonnet 50.054Qwen3.5-27B0.0420.880.920.96per Part I question (of 1.00) Figure 2: One Part I question from the corpus, as the ministry prints and marks it. The Vietnamese original is on the left, and our English translation is on the right. All four statements are true (Đ = đúng). This single question demonstrates the convex grading ladder: judging three out of four statements correctly pays 0.50 points, not the 0.75 points that proportional credit would award. The right panel shows how this rule leaves all eight models short by 0.020 to 0.159 points per Part I question across the corpus. Four models are drawn, showing that their ranking by shortfall differs from their ranking by statement accuracy. The structure is uniform. Decision 764 sets the same three parts for every exam, and Circular 24/2024 [15] governs how students sit for them. Part I consists of four-option multiple choice questions worth 0.25 points each. Part I groups four true/false statements into one question graded on the convex ladder (Fig. 2 shows how the ministry prints and marks one). Part I requires short answers with no options provided. Table 1 details this structure per subject, the exact random baseline it creates, and our coverage. The 2026 exams share this structure because the ministry retained the decision for a second year. Part I appears in ten of the eleven subjects. Where it appears, its four questions account for 4.00 of the 10.00 total points. This means 40% of a candidate’s mark rests on just four out of the 22 to 28 questions on the exam. The random baseline for an entire exam is therefore not a flat 0.25, nor is it the same across subjects. It is lowest in Mathematics because its large Part I pays a guesser nothing. The THPT-Ladder corpus holds 632 items from 21 exams across 11 subjects over two years, covering 22 official variants (mã đề). Every item carries a direct ministry answer, and the dataset includes 80 figure crops for the 77 items that print a figure. We count a Part I question as one single item rather than four, since the ladder pays on the question level. Therefore, our Mathematics exams hold 22 items, even though the ministry’s summary reports 34. Informatics prints six Part I questions but only marks four (Table 1a) because it includes two common questions and two pairs from elective streams, ensuring its total also sums to 10.00 points. Table 1: Official structure (Decision 764) and corpus coverage. structure items Subject I I I rnd 2025 2026 Fig. Biology 18 4 6 2.350 28 28 28 Chemistry 18 4 6 2.350 28 28 9 Economics & Law 24 4 0 2.725 28 28 0 English 40 0 0 2.500 40 40 0 Geography 18 4 6 2.350 28 28 8 History 24 4 0 2.725 28 28 1 Informatics 24 4a 0 2.725 30 30 0 Mathematics 12 4 6 1.975 22 22 14 Physics 18 4 6 2.350 28 28 4 Technology (Agri.) 24 4 0 2.725 56 28 4 Technology (Ind.) 24 4 0 2.725 28 n.p. 12 total 632 items 80 I/I/I: questions per part, at 0.25, 1.00 and either 0.50 or 0.25 points; 10.00 total in 50 min (Mathematics 90). rnd: exact whole-exam random baseline, Eq. 1 on Part I and 0 on Part I. 2025/2026: items held, per year; a count above one exam’s total indicates two variants of the same exam, and n.p. that the ministry had not published the exam. Fig.: figure crops shipped. a6 printed, 4 answered. The chance baseline is therefore neither 25% nor one number: it runs from 19.75% of the scale in Mathematics, whose large Part I pays a guesser nothing, to 27.25% in the five subjects that ask 24 Part I items and no Part I. 4 Evaluation Methodology We score every model exactly as the ministry scores a candidate. A correct four-option item earns 0.25 points, a Part I question pays the convex ladder value, and a correct Part I item receives the subject’s short-answer value. A model’s headline figure in Table 2 represents the fraction of available points it successfully earned. Standard accuracy hides what the ladder penalizes, so we measure the penalty directly. Averaged over the Q Part I questions a model answered, the shortfall is SF=Q−1∑q=1Q(0.25aq−ladder[aq])≥0.SF=Q^-1 _q=1^Q (0.25\,a_q-ladder[a_q] )≥ 0. (2) The shortfall drops to zero only when a model gets every question completely right or completely wrong. It reaches its maximum penalty of 0.25 points on the questions a candidate only half-knows. Because of this, two models that judge the exact same number of statements correctly can earn entirely different marks. The shortfall characterizes the shape and distribution of a model’s knowledge, not just its total size. 17.6%72.8%012340.000.250.500.751.000.25 pt lost0.500.75statements correct, of fourpoints awardedofficial ladder (Decision 764)proportional credit, the counterfactualshare of the 672 Part I questions at each count Figure 3: The marking rule and where the answers actually fall. The state pays the step function, while proportional credit would pay the straight diagonal line. The gap is widest when three out of four statements are correct, paying 0.50 points instead of the implied 0.75 points. Notably, 17.6% of the Part I questions our models answered landed exactly on this rung. The shortfall in Eq. 2 represents the vertical distance between the two lines, weighted by how often each count occurs. Because the mass of answers sits where the gap is widest, the penalty is a real effect rather than a theoretical edge case. A mark on one exam cannot be compared directly to the same mark on another. Even with identical structures and marking rules, the two years of a subject were not equally difficult. For example, a score of 6.00 in Economics & Law beat only 7.99% of the field in 2025, but it beat 74.22% of candidates in 2026. On that specific exam, the number of candidates scoring a perfect 10.00 plummeted from 1,451 down to 2. Seven of the eleven subjects became harder between the two years. To handle this, we report every model’s mark as a percentile rank against the human candidates who sat the exact same exam. Our cohort records successfully reproduce the ministry’s published statistics on 310 out of 312 comparisons. (a)(b)(c)(d)Mathematics 20251.001.000.190.190.380.380.440.44Mathematics 20261.001.000.750.750.380.380.380.38Technology (Ind.) 20250.880.880.500.500.630.630.250.25Physics 20260.790.790.540.540.370.370.500.50Informatics 20260.670.670.580.580.830.830.170.17Economics & Law 20250.650.650.630.630.560.560.540.54Physics 20250.640.640.670.670.590.590.730.73Technology (Agri.) 20260.620.620.430.430.430.430.650.65Technology (Agri.) 20250.590.590.500.500.570.570.590.59Informatics 20250.580.580.500.500.420.420.580.58Chemistry 20260.580.580.240.240.440.440.620.62Biology 20250.560.560.750.750.440.440.690.69Chemistry 20250.550.550.520.520.560.560.500.50Technology (Ind.) 20260.500.500.750.750.500.500.250.25History 20250.480.480.690.690.410.410.420.42Geography 20250.450.450.520.520.400.400.510.51History 20260.450.450.480.480.590.590.480.48Biology 20260.380.380.630.630.750.750.250.25Economics & Law 20260.370.370.460.460.350.350.310.31Geography 20260.290.290.740.740.330.330.140.14statement position Figure 4: How often each Part I statement is true across the twenty subject-years where we hold the keys. In Mathematics, position (a) is always true in all 96 questions every year. However, in Geography 2026, position (a) is true in only 29% of questions and position (d) in 14%. This shows that the answer keys are not statement-balanced and that the imbalance changes direction depending on the subject. This imbalance allows a fixed answer string to beat random chance without reading a single question. 4.1 Answer-Only Baseline Eq. 1 assumes a candidate decides each of a question’s four statements independently. However, a respondent who simply submits the same four-value string to every question exploits the imbalance between true and false answers in the official key. We tested whether the published keys carry exploitable spurious features by scoring strings chosen without looking at the exam text. When we select the string DDSS based on the other 19 subject-years and apply it unseen, it wins on all 20 exams. It earns between 11.07% and 24.25% of a ten-point exam, beating independent guessing on 18 of them. It represents the most frequent of the sixteen possible patterns, appearing 14.5% of the time compared to the 6.25% a perfectly balanced key would produce. Furthermore, the imbalance runs in opposite directions by subject (Fig. 4). A Part I score at or below this level demonstrates no competence at all. 5 Experiments We evaluate eight models in Vietnamese, presenting one item at a time with its text and figures and no worked solution, under prompts identical across subjects and years. Three are open-weight checkpoints served locally at their publishers’ recommended sampling settings, capped at 16,384 generated tokens by the memory available; five are closed models from two vendors, and all but one of those rejects any sampling setting but its default, so the arms are not matched on decoding. Six open-weight responses reach that cap or cannot be read and score nothing, so the Part I comparisons are also reported over the 82 questions free of them. Every one of the 632 items is scored against the ministry’s key by the marking rule of Section 4. 5.1 A mark locates a model in the cohort, unevenly Because the ministry publishes the score distribution for every exam, we can read a mark as a true rank among the human candidates the examination was written for, rather than against an arbitrary scale. On the eighteen subject-years where we hold a complete exam, Qwen3.5-27B ranks above 99.95% of the 1,126,172 candidates who sat Mathematics 2025, but only above 76.97% of those who sat History 2025. InternVL3.5-8B fluctuates even more, running from the 17.8th percentile up to the 97.9th (Fig. 5). A single average score completely conceals a range this wide. Meanwhile, Claude Opus 5 and GPT-5.5 score 9.67 and 9.64 out of 10, respectively, with each taking full marks on eight of the eighteen exams. At this level of performance, the aggregate mark is nearly exhausted as a useful measurement instrument, but the partial-credit shortfall still differentiates them. 5.2 The exams a cohort finds hard are not the exams a model finds hard When we rank the eighteen exams by the mean mark the human cohort earned versus the mark each model earned (Fig. 6), the two orderings show no detectable agreement. The Spearman correlation (ρ) is 0.04 for Qwen3.5-27B, 0.02 for Qwen3.5-9B, and 0.35 for InternVL3.5-8B. None of these correlations are statistically significant, and the first two are indistinguishable from random noise. While previous studies have examined whether a model’s difficulty aligns with humans at the individual item level [9], our unit of measurement is the full exam, matching how the state defines and publishes the marks. This lack of correlation is not a statistical artifact of our sample size. The exact same eighteen exams detect strong agreement where it genuinely exists: Qwen3.5-27B and GPT-5.5 rank the exams almost identically (ρ=0.86ρ=0.86, p<0.001p<0.001), despite crossing both vendor and licensing boundaries. For human candidates, Mathematics 2025 proved to be the hardest exam in the corpus, averaging 4.78 out of 10, yet Qwen3.5-27B scored a perfect 10.00 on it. Conversely, History 2025 was much easier for candidates (averaging 6.52), but it is where Qwen3.5-27B stands lowest within its cohort (Fig. 5). The five closed models fare no better at mimicking human difficulty, with every correlation falling between −0.14-0.14 and 0.220.22 and none reaching significance. Since difficulty for a candidate and difficulty for a model act as nearly independent quantities on this examination, a model’s raw score on a subject does not tell us whether that subject is objectively hard. 0255075100estimated percentile lower bound among the candidates who sat the examMathematics 254.78Econ. & Law 265.02English 265.07Geography 265.10English 255.38Physics 265.56Mathematics 265.65Biology 255.78Tech. (Ind.) 255.79Biology 265.84Chemistry 256.06History 266.19Chemistry 266.28History 256.52Geography 256.63Tech. (Agri.) 266.96Physics 256.99Econ. & Law 257.69Qwen3.5-27BQwen3.5-9BInternVL3.5-8B50th = the median candidate Figure 5: Each open-weight model’s rank among the candidates who sat the same exam, for the eighteen subject-years where the corpus holds a whole 10.00-mark exam; the closed models’ standings are given in the text. Rows run from the exam the cohort found hardest to the one it found easiest, with the cohort’s own mean mark at the right. Percentiles are lower bounds: a mark on the ministry’s 0.25 grid ties with every candidate who earned it. If difficulty transferred from candidates to models the markers would drift rightwards down the figure; they do not, which indicates that a model’s standing is set by something other than what the exam cost a human. 4466881010equal marksMath 25E&L 25candidates’ average markmodel’s markQwen3.5-27B ρ=0.04ρ=0.04Qwen3.5-9B ρ=0.02ρ=0.02InternVL3.5-8B ρ=0.35ρ=0.35 Figure 6: The same eighteen exams, with the cohort’s mean mark against each open-weight model’s on one 0–10 scale; the closed models’ correlations are given in the text. The diagonal marks equal marks; each dashed horizontal is one model’s own mean. A cloud running parallel to the diagonal would mean difficulty transfers from candidates to models, and a cloud along a horizontal that it does not. Every cloud here is flat: a model’s mark is therefore near-independent of what the exam cost its candidates. No cohort averaged above 7.7, which is why the upper right is empty. Labelled: the cohort’s hardest exam and its easiest. Table 2: Eight models on the 632 scored items: three open-weight, then five closed. Model % avail. I I stmt I pts/q shf/q I Qwen3.5-27B 93.8 0.973 0.926 0.884 0.042 0.917 Qwen3.5-9B 90.7 0.965 0.896 0.830 0.066 0.850 InternVL3.5-8B 69.3 0.838 0.756 0.597 0.159 0.183 Claude Opus 5 97.0 0.988 0.973 0.949 0.024 0.950 GPT-5.5 97.0 0.988 0.967 0.948 0.020 0.950 Claude Sonnet 5 93.6 0.977 0.935 0.881 0.054 0.917 Claude Opus 4.8 92.9 0.975 0.920 0.867 0.052 0.917 Claude Haiku 4.5 82.9 0.934 0.842 0.735 0.108 0.567 % avail.: the marks a model earned as a percentage of the 220.00 the 22 variants are worth; Informatics is the mean over its two elective streams. I, I: the fraction of four-option and short-answer items answered correctly. I stmt: the fraction of individual true/false statements judged correctly. I pts/q: what the ladder paid, per Part I question, out of 1.00. shf/q: the shortfall SFSF of Eq. 2, the credit the ladder withheld (lower is better). The open-weight models are served locally at their publishers’ recommended sampling settings. The closed models are Anthropic’s Claude Opus 5, Opus 4.8, Sonnet 5 and Haiku 4.5 and OpenAI’s GPT-5.5, reached through Amazon Bedrock; all but Haiku 4.5 reject any sampling setting but their default. The shortfall is positive for all eight, which shows each earns fewer marks than its statement accuracy implies, and reveals that the shortfall is not ordered by that accuracy. 5.3 The shortfall follows the shape of a model’s errors Every model loses marks to the convex ladder (Table 2). If credit were awarded one statement at a time, Qwen3.5-27B would earn 0.926 points per question, but the ladder only pays it 0.884. Across all eight models, this shortfall runs from 0.020 to 0.159 points per question. Crucially, statement accuracy does not predict this penalty. When we rank models by accuracy, the shortfall does not decrease monotonically. GPT-5.5 knows 96.7% of statements compared to Claude Opus 5’s 97.3%, yet it faces a smaller penalty (0.020 versus 0.024). Claude Sonnet 5 knows 93.5% of statements compared to Qwen3.5-27B’s 92.6%, but it suffers a higher penalty (0.054 versus 0.042). If we hold a model’s overall accuracy constant and only vary how its errors cluster across questions, the resulting mark becomes a wide interval rather than a single number. For example, at Sonnet 5’s accuracy level, the ladder can pay anywhere between 0.869 and 0.932 points per question. Because of this, accuracy simply cannot recover the true mark. One specific comparison perfectly isolates the ladder’s effect. On four occasions, a model and the blind guessing string from Section 4.1 judged the exact same number of statements correctly on the same exam but received entirely different payouts. The marks depend on how the correct statements group together, not on the total count. On the Biology 2026 exam, Qwen3.5-27B and the fixed string each got 11 out of 16 statements right. However, the blind string earned 2.35 points while the model earned only 1.75. The model was uniformly close on every question (getting 3, 2, 3, and 3 statements correct), whereas the blind string swept two questions completely (getting 2, 4, 1, and 4 statements correct). The convex ladder explicitly pays for the second shape. 6 Implications for Assessment Practice Three constraints follow for anyone automating marking on this examination. First, a marking engine has to implement the ladder of Decision 764 rather than proportional credit: the two disagree by 0.020 to 0.159 points per Part I question, and on History 2025 that gap separates the 90th percentile from the 77th among 481,293 candidates. Second, an accuracy figure does not bound the mark a system would be awarded, because at a fixed statement accuracy the ladder pays between 0.869 and 0.932 points per question; procurement evidence should therefore be a mark under the published rules, read against the answer-only floor of Section 4.1. Third, model scores are not candidate difficulty data: over the eighteen exams no correlation between cohort mean marks and model marks reached significance (ρ from −0.14-0.14 to 0.350.35), so an item bank calibrated on model performance would not inherit the difficulty ordering its own students experience. 7 Limitations Percentiles are coarser than they initially appear. A single step on the ministry’s 0.25 grading grid moves a candidate several percentile places near the median. Therefore, this benchmark cannot separate two models that land within a quarter point of each other. The grid is set by the state and cannot be refined. Because the 2025 exams have been public for over a year, models may have encountered them during training. Data contamination would normally reveal itself if a model suddenly lost ground on the unseen 2026 exams, assuming we hold the human cohort’s change in difficulty constant. The intercept we measured is positive but statistically indistinguishable from zero across ten subjects. This acts as a null result rather than definitive proof of no contamination. Since every item in the corpus records its year, researchers can safely avoid this issue by restricting their comparisons exclusively to the 2026 exams, which were released on 19 June 2026. 8 Ethics and Data Statement Every exam, key and score distribution we release is a published document of Vietnam’s Ministry of Education and Training, redistributed as published and cited per item. The candidate score records are aggregate counts of marks and carry no personal identifier. We publicly release the corpus, the ministry keys and the code that extracts and scores it at [PENDING: release URL] under C-BY-4.0, with the marking rule of Decision 764 implemented so that a later year’s exams can be scored by the same command. 9 Conclusion THPT-Ladder scores language models on 632 items from 21 official exams using Vietnam’s published marking rules instead of flat accuracy. The ladder withholds 0.042 points a question from Qwen3.5-27B, a penalty that drops it thirteen percentile places on the 2025 History exam. It withholds 0.159 points from the weakest model, whose errors spread more thinly across questions. That statement accuracy and final marks are completely distinct quantities is a core property of the non-additive scheme. Even at a fixed statement accuracy, the mark the ladder pays spans a wide interval. With the strongest models earning full marks on eight of the eighteen exams, what separates them is exactly where their errors fall rather than how many they make. Convex partial credit is how an institution declines to pay for knowledge it cannot rely on. When a benchmark substitutes standard accuracy for these rules, it reports a competence the institution would never certify. References [1] C. T. Bui et al., “VMLU benchmarks: a comprehensive benchmark toolkit for Vietnamese LLMs,” in Proc. ACL, 2025, p. 11495–11515. [2] X.-Q. Dao et al., “VNHSGE: VietNamese high school graduation examination dataset for large language models,” arXiv:2305.12199, 2023. [3] F. Koto, N. Aisyah, H. Li, and T. Baldwin, “Large language models only pass primary school exams in Indonesia: a comprehensive test on IndoMMLU,” in Proc. EMNLP, 2023, p. 12359–12374. [4] F. Koto et al., “ArabicMMLU: assessing massive multitask language understanding in Arabic,” in Findings of ACL, 2024, p. 5622–5640. [5] Quyết định 764/QĐ-BGDĐT, Bộ GD&ĐT, 8 Mar. 2024. [6] J. Gao et al., “LaoBench: a large-scale multidimensional Lao benchmark for large language models,” in Proc. ACL, 2026, p. 23727–23743, arXiv:2511.11334. [7] A. Pramodya et al., “SinhalaMMLU: a comprehensive benchmark for evaluating multitask language understanding in Sinhala,” in Proc. EMNLP, 2025, p. 32943–32961, arXiv:2509.03162. [8] P. Acharya, P. Bharati, Y. Chapagain, I. S. Gauli, and K. Parajuli, “Assessing the pedagogical readiness of large language models as AI tutors in low-resource contexts: a case study of Nepal’s K–10 curriculum,” arXiv:2604.09619, 2026. [9] J. He-Yueya et al., “Psychometric alignment: capturing human knowledge distributions via language models,” arXiv:2407.15645, 2024. [10] S. Wind et al., “SteuerLLM: local specialized large language model for German tax law analysis,” Sci. Rep., vol. 16, art. 21640, 2026, arXiv:2602.11081. [11] Z. Yang et al., “RadSEM: a finding-by-finding metric for clinical consistency in radiology reports,” arXiv:2606.17062, 2026. [12] W. Xia, J. Wu, H. Shi, X. Wang, and C. Zheng, “PsyScore: a psychometrically-aware framework for trait-adaptive essay scoring and ZPD-scaffolded feedback,” in Findings of ACL, 2026, p. 7768–7786, arXiv:2606.20287. [13] W. Wang et al., “CMPhysBench: a benchmark for evaluating large language models in condensed matter physics,” arXiv:2508.18124, 2025. [14] W. Li, L. Li, T. Xiang, X. Liu, W. Deng, and N. Garcia, “Can multiple-choice questions really be useful in detecting the abilities of LLMs?,” in Proc. LREC-COLING, 2024, p. 2819–2834, arXiv:2403.17752. [15] Thông tư 24/2024/T-BGDĐT, Bộ GD&ĐT, 2024.