Source-linked AI summary
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
TL;DR
Human-exam benchmarks usually replace official marking schemes with flat accuracy, even when partial knowledge earns non-proportional credit. THPT-Ladder applies Vietnam’s 2025 rules to official exams and finds that eight models receive 0.020–0.159 fewer points per Part II question than proportional scoring, changing apparent human-cohort standing.
Problem
Most human-exam benchmarks discard official marking schemes and score answers as right or wrong, leaving Vietnam’s new question formats unevaluated.
Method
THPT-Ladder scores 632 items from 21 official exams across 11 subjects using ministry answer keys, marking rules, and exact-exam human percentiles.
Results
0.020 to 0.159 points per Part II question separate official-rubric scores from proportional credit across eight models, with Qwen3.5-27B falling from the 90th to 77th percentile on 2025 History.
Takeaways & Limitations
Flat accuracy can report competence that the examination institution would not certify because official marks depend on where errors fall across statements.
Takeaways & Limitations
The ministry’s 0.25-point grading grid cannot distinguish models whose scores fall within a quarter point of each other.
Abstract
from arXiv · showhide
When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.
1 Introduction
THPT-Ladder evaluates language models under Vietnam’s 2025 official examination rules rather than flat accuracy, making the cost of ignoring non-additive marking measurable. Its results show that statement accuracy alone does not determine official credit or model standing.
- Motivation: Most human-exam benchmarks discard marking schemes, score responses right or wrong, and assume partial knowledge deserves proportional credit.This substitution can report credit that an institution’s rules would not award.
- Examination reform: 2025 reforms introduced three formats, including Part II’s four true/false statements marked together as one question; existing Vietnamese benchmarks cover neither new format.Decision 764/QĐ-BGDĐT defines the formats as Parts I, II, and, where applicable, III.
- Convex grading: 0.50 rather than 0.75 points is awarded for getting three of four Part II statements correct, and Part II carries 4.00 of an exam’s 10.00 points.The official ladder is (0, 0.10, 0.25, 0.50, 1.00), defining the partial-credit gap against additive scoring.
- Benchmark: 632 items from 21 official exams across 11 subjects comprise THPT-Ladder, which scores models using ministry keys and reports percentiles among candidates from the exact exam year.This places model scores within the human cohort assessed by each examination.
- Findings: Qwen3.5-27B judged 92.6% of Part II statements correctly but received 0.884 points per question under the ladder versus 0.926 under statement-by-statement credit.The official rule therefore awards the lower figure than the standard benchmark report.
- Findings: Accuracy does not determine payment under the non-additive scheme: among closed models, shortfall is not monotone when models are ordered by statement accuracy.The shortfall follows mathematically from the ladder and applies to any system scored under Decision 764.
2 Related Work
Existing Vietnamese benchmarks use binary item scoring and therefore cannot evaluate non-additive grading, while other fine-grained benchmarks do not use convex credit functions. THPT-Ladder also treats official Vietnamese exam variants as substantive content variations rather than simple option reorderings.
- Vietnamese benchmarks: VMLU and VNHSGE use right-or-wrong scoring, preventing evaluation under a non-additive grading scheme.VMLU covers 58 subjects and four education levels; VNHSGE draws directly from the high school graduation exam.
- Fine-grained scoring: SteuerLLM, RadSEM, PsyScore, and CMPhysBench establish finer-grained scoring, but none uses a convex credit function.Their approaches include statement-level marking, atomic findings, graded item-response credit, and non-binary expression-tree credit.
- Exam variants: 15 of 20 subject-years with available keys split into groups with entirely different content across Vietnam’s 24 or 48 official exam variants.The variants are therefore not merely reordered copies, and every published key is released for future research.
3 The Examination and the Corpus
Vietnam’s 2025 examination uses a uniform three-part structure and a ministry-defined convex ladder for Part II true/false questions. THPT-Ladder counts ladder-graded questions as single items across 632 items from 21 official exams, preserving the ministry’s marking structure.
- Examination structure: Each exam is marked out of 10.00 points and follows three question formats established by Decision 764.The ministry sets the marking scheme, and corpus answers come from its published keys.
- Question formats: Part I uses four-option multiple-choice questions worth 0.25 points each, Part II groups four true/false statements under a convex ladder, and Part III requires option-free short answers.Circular 24/2024 governs how students sit the exams.
- Corpus construction: Mathematics exams contain 22 counted items despite the ministry’s summary reporting 34, while Informatics marks four of six printed Part II questions.Informatics retains two common questions and two elective-stream pairs so the exam totals 10.00 points.
- Corpus construction: 632 items from 21 official exams constitute the THPT-Ladder corpus.Part II questions are counted as single items because the ladder pays at the question level.
- Part II marking: Judging three of four Part II statements correctly pays 0.50 points rather than the 0.75 points proportional credit would award.The figure presents this as the convex grading ladder used by the ministry.
4 Evaluation Methodology
The benchmark applies Vietnam’s official marking rules to model responses and measures how convex partial credit changes scores beyond standard accuracy. It also accounts for cross-exam difficulty and tests whether answer-key patterns create exploitable baselines.
- Official scoring: Models are scored exactly like candidates: Part I earns 0.25 points, Part II follows the convex ladder, and Part III uses subject-specific short-answer values.Table 2 reports each model’s fraction of available points earned.
- Convex-credit penalty: The shortfall is zero for completely right or wrong questions and peaks at 0.25 points when a candidate only half-knows an item.Thus, identical statement-level accuracy can produce different official marks because the penalty depends on knowledge distribution.
- Cross-exam comparability: A score of 6.00 in Economics & Law beat 7.99% of candidates in 2025 but 74.22% in 2026, showing why marks require within-exam percentile normalization.The number of perfect 10.00 scores fell from 1,451 to 2, and seven of eleven subjects became harder between years.
- Random baselines: The unseen DDSS answer string earned 11.07%–24.25% of a ten-point exam and beat independent guessing on 18 of 20 exams.DDSS was selected from the other 19 subject-years and appeared 14.5% of the time versus 6.25% for a perfectly balanced pattern.
5 Experiments
Experiments evaluate eight Vietnamese language models under mostly standardized prompting and compare their official convex-ladder marks with human percentile ranks. The results show that the ladder materially changes scores, and that exam difficulty, statement accuracy, and total correct counts do not reliably determine official payouts.
- Experimental setup: Eight models were evaluated item-by-item in Vietnamese under identical prompts across subjects and years, though decoding settings were not fully matched.Three open-weight models used publishers’ recommended local settings with a 16,384-token cap, while five closed models came from two vendors and generally used defaults.
- Human-cohort ranking: Qwen3.5-27B ranked above 99.95% of Mathematics 2025 candidates but only 76.97% of History 2025 candidates.InternVL3.5-8B ranged from the 17.8th to the 97.9th percentile across complete subject-years.
- Human-cohort ranking: Spearman correlations between human and model exam rankings were 0.04 for Qwen3.5-27B, 0.02 for Qwen3.5-9B, and 0.35 for InternVL3.5-8B, none statistically significant.The first two correlations were indistinguishable from random noise.
- Convex-ladder effects: 0.020 to 0.159 points per question was the shortfall across eight models under the convex ladder versus statement-by-statement credit.Qwen3.5-27B would receive 0.926 points per question proportionally but earned 0.884 under the official ladder.
- Convex-ladder effects: Statement accuracy did not predict the penalty: Claude Sonnet 5 had 93.5% accuracy yet a 0.054 shortfall, versus Qwen3.5-27B’s 92.6% and 0.042.The penalty varies with how errors cluster across questions, because the official marks depend on grouping correct statements rather than their total count.
6 Implications for Assessment Practice
Assessment automation must reproduce Decision 764’s convex marking ladder rather than proportional credit, because accuracy alone cannot determine awarded marks. Model marks should be reported under the published rules and interpreted cautiously, since they do not track candidate difficulty across exams.
- Automated marking: Decision 764’s ladder should replace proportional credit: the two schemes differ by 0.020–0.159 points per Part II question, separating the 90th from the 77th percentile.On History 2025, this percentile gap is among 481,293 candidates.
- Automated marking: At fixed statement accuracy, the ladder awards 0.869–0.932 points per question, so procurement evidence should report marks under published rules.The passage recommends reading these marks against the answer-only floor of Section 4.1.
- Difficulty interpretation: Across eighteen exams, cohort mean marks and model marks showed no significant correlation, with ρ ranging from −0.14 to 0.35.Therefore, model scores should not be treated as candidate-difficulty data.
7 Limitations
The benchmark cannot distinguish models within a quarter point because the ministry’s 0.25 grading grid creates coarse percentile steps. Possible training-data contamination remains unresolved: the measured intercept is positive but statistically indistinguishable from zero across ten subjects, providing a null result rather than proof of no contamination.
- Percentile resolution: A 0.25-point step on the ministry’s grading grid can shift candidates several percentile places near the median.The state-set grid therefore limits percentile resolution.
- Percentile resolution: Models landing within a quarter point of each other cannot be separated by this benchmark.The grading grid cannot be refined because it is set by the state.
- Data contamination: Across ten subjects, the measured intercept was positive but statistically indistinguishable from zero, yielding a null result on contamination.Because the 2025 exams have been public for over a year, models may have encountered them during training; this result is not definitive proof of no contamination.
8 Ethics and Data Statement
The released materials consist of published Vietnamese ministry documents and anonymized aggregate score counts. The corpus, keys, and scoring code are publicly released under CC-BY-4.0, with Decision 764 implemented for future exams.
- Exam papers, keys, and score distributions are redistributed from published Ministry of Education and Training documents and cited per item.
- Candidate records contain aggregate mark counts without personal identifiers.
- The corpus, ministry keys, and extraction-and-scoring code are publicly released under CC-BY-4.0.
- Decision 764’s marking rule is implemented so later-year exams can be scored with the same command.
9 Conclusion
THPT-Ladder applies Vietnam’s published marking rules to 632 items from 21 official exams, showing that statement accuracy and final marks are distinct under the non-additive scheme. Its convex scoring penalizes Qwen3.5-27B by 0.042 points per question and makes error placement, rather than error count alone, decisive.
- Benchmark and scoring: 632 items from 21 official exams are scored using Vietnam’s published marking rules instead of flat accuracy.The benchmark uses the ministry’s marking scheme to evaluate language models.
- Benchmark and scoring: 0.042 points per question withheld from Qwen3.5-27B drops its 2025 History standing thirteen percentile places.The ladder’s penalty changes the model’s placement among human examinees.
- Benchmark and scoring: ρ = 0.04 for Qwen3.5-27B, ρ = 0.02 for Qwen3.5-9B, and ρ = 0.35 for InternVL3.5-8B.These reported correlations underscore the weak relationship between model marks and the examined quantity.
- Error placement: At fixed statement accuracy, the ladder’s mark spans a wide interval because convex partial credit depends on where errors fall.With strongest models earning full marks on eight of eighteen exams, error locations—not merely error counts—separate them.