Source-linked AI summary
Grading Needs a Rubric, Not Intelligence
Jhen-Ke Lin
TL;DR
Open-ended examination benchmarks need grading, but frontier-model judging is costly and biased. ANY-TO-BENCH extracts each exam’s rubric once with a frontier model, then tests cheaper models applying it repeatedly. Rubric-anchored grading makes judge intelligence largely irrelevant, with the official answer carrying most of the rubric’s value.
Problem
Open-ended benchmark questions require costly grading, while model judges retain documented preferences for answer length, position, and model family.
Method
ANY-TO-BENCH uses a frontier model once to extract questions and rubrics, then evaluates six configurations grading answers to 24 open-ended examination questions.
Results
Judge identity explains only 0.2 % of score variance, while removing the official answer collapses reliability and makes judge reasoning effort matter again.
Takeaways & Limitations
The rubric decouples grading from judge intelligence, with the official answer carrying nearly all the rubric’s weight.
Takeaways & Limitations
The study does not establish agreement between rubric-anchored scores and human examiner judgments.
Abstract
from arXiv · showhide
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.
1 Introduction
ANY-TO-BENCH shifts intelligence-intensive work to one-time rubric extraction, leaving repeated grading to lower-cost judges. Across 24 open-ended questions, judge capability had little effect on scores when grading was anchored to the extracted rubric, especially its official reference answer.
- Motivation: Open-ended examination items remain costly to benchmark because proofs, essays, translations, and drawings require a grader, while standard model judging uses a frontier call for every grade.The introduction identifies grading as the historically costly, unscalable component of converting examinations into machine-runnable benchmarks.
- Approach: ANY-TO-BENCH uses a frontier configuration once per exam to extract questions, answer formats, and complete scoring rubrics before lower-cost models perform grading.Each rubric includes the official reference answer and, where published, explicit criteria with defined levels.
- Evaluation: Six configurations from two model families at three reasoning-effort settings wrote answers and judged 24 open-ended questions from three Taiwanese national examinations.The same configurations served as writers and judges, allowing reasoning effort to be varied independently on the answering and grading sides.
- Findings: 0.2 % of score variance was attributed to judge identity, while writer reasoning effort changed earned scores by up to 0.143 of full marks and judge reasoning effort changed assigned scores by at most 0.006.Replacing the panel with its cheapest members changed per-answer scores by 0.019 on average, and frontier-tier judges reproduced the same scores.
- Evaluation: Agreement claims are conditioned on judges correctly identifying both anchors: the empty answer sheet bounds the scale below and published reference answers bound it above.The scale is anchored by construction rather than assumption.
2 Related work
Prior LLM-judge research mainly studies pairwise preferences and documents order sensitivity, whereas this work evaluates rubric-anchored point assignment using educational-measurement tools. Related examination-grading studies vary rubric detail or compare with human raters; this study holds the rubric fixed and varies judge intelligence.
- Prior LLM-judge research: Prior work mainly evaluates pairwise answer preferences, reporting over 80 % agreement between GPT-4 verdicts and human preferences while documenting answer-order sensitivity.This literature also includes form-filling evaluation and LLMs as replacements for human open-text evaluation.
- Rubric-anchored grading: This study assigns points against extracted criteria, defined levels, and official point values rather than comparing answers or rating free-form quality.It uses intraclass correlation for absolute agreement and variance decomposition over answers, judges, and their interaction.
- Positioning against direct grading work: Related examination-grading studies vary rubric detail or compare automated scores with human raters, whereas this study holds the rubric fixed and varies the judge.The central question is how much intelligence rubric application needs.
3 Experimental design
The experiment evaluates rubric-anchored grading on 24 questions drawn from a public corpus of Taiwanese national examinations, using six answer-writing and judging configurations plus frontier-tier checks. Agreement is measured with ICC(2,1), while same-question ablations isolate the contributions of rubric criteria, levels, and official answers.
- Experimental design: The public corpus contains 164 exam bundles and 7,121 questions from three Taiwanese national examinations spanning 2024–2026.A frontier model performed the one-time ingestion, extracting each question, answer format, and per-criterion grading rubric.
- Experimental design: The evaluation samples 24 questions across four guidance strata, 21 exam papers, three examination series, three years, eight subjects, and three answer formats.The questions are worth 1 to 25 points each, with six questions in every stratum.
- Experimental design: Six configurations—5.6 Luna and Sonnet 5 at low, medium, and high reasoning effort—both wrote answers and served as judges.Two anchors supplied reference endpoints: official-answer sheets intended to earn nearly full marks and empty sheets intended to earn nothing.
- Experimental design: The six judges graded eight sheets three times each, producing 3,456 complete per-question verdicts; six additional frontier-tier configurations graded each sheet once as a ceiling check.Statistics were computed with and without rubric-level snapping because snapping altered 1 of 3,456 verdicts.
- Agreement measure: Agreement uses ICC(2,1), a two-way random-effects measure of absolute agreement, with cluster-bootstrap confidence intervals and single-pass scores from the six model-written sheets.The two anchors are excluded because their easy separation would inflate agreement statistics.
- Ablations: Two same-question ablations remove rubric criteria and levels while retaining official answers, then remove official answers as well, isolating each component’s contribution.Both ablations use the twelve stratum-C and -D questions.
4 The scale is anchored
The grading scale was correctly anchored at both extremes: the empty sheet always received zero, while the reference sheet received nearly full marks. All six writers fell strictly between these anchors.
- Lower anchor: 0: the empty sheet scored exactly zero in all 432 gradings.This establishes the lower anchor without relying on inter-judge agreement alone.
- Upper anchor: 0.990 of full marks: the reference sheet’s mean score, with scores of at least 0.9 in 285 of 288 gradings.The reference sheet therefore functioned as a near-perfect upper anchor.
- Scale placement: All six writers sat strictly between the empty-sheet and reference-sheet anchors across the 3,456 verdicts.The figure shows the constructed anchors bracketing every writer’s scores.
5 Effort moves writers, not judges
Reasoning effort substantially changes scores earned by writers but not scores assigned by judges. Higher judge effort improves repeatability, while frontier-tier judges add no measurable change beyond the cheap panel.
- Writer versus judge effort: 0.143 of full marks separates low- and high-effort 5.6 Luna writers, whose mean scores rise from 0.790 to 0.934.Sonnet 5’s answers are flat in effort, so the writer-side response is carried by one family.
- Writer versus judge effort: ±0.012 bounds the six judges’ mean-leniency deviations from the panel mean across two families and three efforts.The judge side stays flat on the common scale used in Figure 2.
- Repeatability: 0.080 versus 0.038 is the within-judge standard deviation for low- versus high-effort 5.6 Luna judges, while Sonnet 5 falls from 0.038 to 0.026.The variation is centered on the same verdict and changes no comparison between answers across 3,456 verdicts.
- Frontier-tier check: 0.030 to 0.039 is the deviation of each frontier judge from the cheap panel’s consensus, within the cheap judges’ 0.023 to 0.042 leave-one-out range.The frontier check used six configurations—GPT-5.6 Sol and Claude Opus 5 at three efforts—to grade every sheet once.
6 The rubric carries the judgment
The rubric, especially its official answer, carries most of the judgment: removing criteria and levels changes little, while removing the official answer sharply reduces reliability and inflates scores. Without rubric guidance, judges have little signal to distinguish answers, making grading more dependent on judge capability.
- No-rubric contrast: Stratum A places six answers within a standard deviation of 0.043, versus 0.22–0.28 elsewhere, despite ordinary judge disagreement.The judges disagree about 0.080 in stratum A, compared with 0.092 and 0.085 in strata C and D; the collapsed signal leaves little to distinguish.
- Ablation: ICC remains 0.888 with the official answer alone versus 0.880 with the full rubric after criteria and levels are removed.Writer spread retains a median 114 % of its full-rubric value, while scores rise by 0.016 of full marks and mean judge spread increases from 0.122 to 0.145.
- Ablation: Reliability falls to 0.628 when the official answer is removed, while scores inflate by 0.074 of full marks.Median discrimination falls to 68 %, although the mathematics and one civics question retain 108 % and 102 % of their spread, respectively.
- Mechanism: The rubric decouples grading from judge intelligence, with the official answer carrying nearly all the weight and criteria and levels adding a margin of consistency.Removing the official answer makes grading turn back into answering, so capability matters again.
- Limitation: Essays remain within 0.04 of one another at every guidance level, so their similarity or judges’ inability to separate prose quality cannot be resolved without human raters.This limitation holds with a 26-level rubric, only a key, or no guidance.
7 Two biases that fail to appear
Rubric-anchored grading shows no meaningful length or same-family preference: longer answers can score higher within questions, but cross-writer scores remain essentially equal and family effects resemble noise.
- Length bias: Within-question length–score correlations are +0.60 to +0.74 for every judge, because longer answers cover more rubric criteria and are more complete.Length acts as a stand-in for completeness within a question, not as an independently rewarded feature.
- Length bias: 0.875 and 0.868 are the family means for short Luna and longer Sonnet answers, despite a 2.3× length difference.The three Luna sheets average near 230 characters, while the three Sonnet sheets average near 520; the full quality range, 0.790 to 0.934, occurs within the short regime.
- Length bias: At matched high effort, the 234-character writer outscores the 538-character writer by 0.055, showing quality can matter at nearly fixed length.Length varies by a factor of 2.3 without affecting scores, while quality varies at fixed length with full effect.
- Same-family bias: 0.874 and 0.886 are Luna judges’ scores for Luna and Sonnet answers, while Sonnet judges assign 0.861 and 0.864, with opposite-direction gaps.The small, opposing gaps are consistent with noise rather than same-family preference.
- Same-family bias: 0.021 within-family and 0.039 across-family judge disagreement are both small, indicating only a faint family boundary in pairwise disagreements.The reported family effects do not support a meaningful same-family preference under rubric anchoring.
8 One judge is enough
A single judge is sufficient: adding judges barely changes reliability, while repetition provides little additional benefit. Two judges achieve nearly the same average reliability as six, and repeated gradings are often identical.
- Panel size: 0.923 single-pass reliability with two judges versus 0.922 with six shows that larger panels add almost nothing on average.Averaged over every possible panel, the difference is 0.001.
- Panel size: 0.869 reliability for the worst two-judge panel indicates that even the weakest small panel remains reasonably reliable.Adding judges narrows the worst case rather than raising the mean.
- Repeated grading: 73 % of repeated gradings are exactly identical, with a pooled within-judge standard deviation of 0.050.Repetition therefore contributes little additional variation reduction.
9 Limitations
The study’s limitations concern untested capability floors, narrow examination and language scope, limited statistical power, and uncertainty about prose discrimination. Its validity is anchored and consistent rather than established against human examiners.
- Capability floor: The cheap-judge floor remains unknown: rubric application has not been tested farther down the capability scale.Frontier judges grade no better than cheaper judges, but the study only compares each with its frontier siblings.
- Scope: All 24 questions come from Taiwanese national examinations in Traditional Chinese, while a best-writer average of 0.934 under-represents discrimination among good answers.Other languages and examining traditions remain untested.
- Validity: Validity means anchored, consistent grading, not agreement with a human examiner.No human marker scored these answers.
- Power: 0.815 is the bootstrap lower bound supporting strong agreement, but six questions per stratum remains thin and writer-side effort results rest on one model family.No single question changes any stratum’s ICC by more than 0.17; the ablations are single-pass over twelve questions.
- Prose: The essays never discriminate at any guidance level, leaving unresolved whether competent models write similarly good prose or judges cannot distinguish prose quality.Human raters are needed to decide between these explanations, so free-form writing items may provide little benchmark signal.
10 Conclusion · A The 24 questions
ANY-TO-BENCH recovers existing examination judgment as a rubric using a frontier model once per exam, after which low-effort small judges grade answers like expensive counterparts. The study’s 24 questions are catalogued by source, format, value, rubric size, stratum, and official-answer availability.
- 10 Conclusion: ANY-TO-BENCH uses a frontier model once per exam to recover what to ask, what good answers contain, and how many points each part is worth.The experiment reports that this recovery works and that grading is the easier post-ingestion step.
- 10 Conclusion: 0.2 % of score variance is contributed by small judges at low effort, which grade the same answers to the same scores as their most expensive counterparts.This passage establishes the conclusion’s central grading result, though the supplied text ends mid-sentence.
- A The 24 questions: Table 3 lists all 24 questions included in the study.The table records each question’s stratum, exam paper, answer format, point value, and rubric size.
- A The 24 questions: Each question is identified by its stratum and the exam paper from which it comes.These are two of the study’s question-level catalogue fields.
- A The 24 questions: Each question is also classified by answer format and point value.Table 3 reports both fields for every question.
- A The 24 questions: Table 3 reports the size of each question’s rubric.The table’s rubric-size information is represented by the total number of defined levels across rubric criteria.
- A The 24 questions: Levels is the total number of defined levels across a rubric’s criteria.This is the table’s explicit definition of the Levels field.
- A The 24 questions: The eight questions without an official answer are excluded from reference-anchor statistics.Table 3 defines Key as whether the examining board publishes an official answer.
B Reproducibility
The paper’s numbers and figures are fully regenerable from stored grading reports and public exam materials. The repository includes the reports, tidy verdict data, and scripts needed to reproduce the analyses.
- Stored materials: Every number and figure regenerates from stored grading reports in the repository’s research/judge-reliability directory.The repository contains the materials used for regeneration.
- Stored materials: 144 main-study grading reports are stored as one JSON file per sheet, judge, and repeat.These reports cover the main study’s grading runs.
- Stored materials: 96 ablation reports, 48 frontier-judge reports, data.csv with 3,456 main-study verdicts, and analysis scripts are included.The scripts read only the stored reports listed in the repository materials.
- Public data: The public exam corpus contains the 24 selected questions and their rubrics, while each data.csv row represents one judge grading one question on one answer sheet in one repeat.Table 4 defines the data.csv columns.