Source-linked AI summary
LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
Veerendra Kumar Sunkavalli
TL;DR
LLM essay graders are usually evaluated by agreement with humans, but educational measurement indicates that severity, halo, and instability also matter. This paper treats LLM judges as raters and audits them across public bilingual essay corpora with a pre-registered, replicated psychometric design. It finds large severity and version effects, stable repetition without improved accuracy, and no credible excess halo over trained humans under a matched comparison.
Problem
Agreement statistics do not capture rater severity, halo, or instability, despite their relevance to automated writing decisions.
Method
The paper applies a pre-registered rater-effects battery to 2,377 essays scored by 12 judges across four providers, five version contrasts, replications, and released score data.
Results
Judge severity spans 219 ENEM points, all five version contrasts shift severity beyond the permutation null, replication yields φ≥.80 at k≤2 without human-level accuracy, and matched halo comparisons find no credible excess.
Takeaways & Limitations
LLM judges should be calibrated against human-anchored samples and monitored as mutable instruments rather than selected by agreement statistics alone.
Takeaways & Limitations
The findings are conditional on one platform, rubric prompts, decoding regime, collection window, and corpora without demographics for subgroup DIF analysis.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and "silent drift" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi>=.80 at k<=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.
1. Introduction
The paper argues that agreement statistics alone miss important rater effects in LLM essay grading. It treats LLM judges as raters and audits severity, halo, reliability, and version stability using a pre-registered, replicated design.
- Motivation: LLM judges can agree equally well with humans while differing systematically in severity, halo, and stability.The study applies classical educational-measurement concerns to LLM scoring rather than treating agreement as sufficient.
- Design: A fully crossed judges × essays grid with replications enables rater-effects analysis at a scale unavailable to human-rater studies.The design covers public ENEM and ASAP corpora, multiple judges, providers, languages, and version contrasts.
- Design: The study was pre-registered before data collection, with frozen hypotheses, decision rules, exclusions, and analysis specifications.Deviations were logged and disclosed, and confirmatory nulls received the same prominence as positive findings.
- Findings: Severity was the dominant observed rater effect, spanning 219 points on ENEM while correlations remained within a narrow .47–.56 band.The findings show why location-blind agreement statistics can miss consequential differences among judges.
- Contribution: The audit contributes a released score tensor and actionable guidance for calibrating, monitoring, and comparing LLM judges.The artifact includes raw captures, rubrics, canary logs, and analysis code for reuse as a benchmark.
2. Related Work
Prior work establishes agreement and variance-component perspectives but leaves a combined rater-effects audit of LLM essay judges underdeveloped. This study positions its contribution as the integration of crossed replication, psychometric effects, cross-lingual data, version contrasts, and a released tensor.
- Rater-effects foundations: Human educational measurement documents severity, halo, central tendency, and drift as distinct rater effects.MFRM separates rater severity from examinee ability and task difficulty, while generalizability theory partitions variance across raters, tasks, and occasions.
- Writing analytics: Automated writing evaluation has historically balanced agreement with validity for learning-support decisions.The paper argues that feedback, placement, and progress monitoring depend on rater properties, not rank agreement alone.
- LLM-judge literature: Prior LLM-judge studies separately address protocols, IRT, agreement reporting, G-theory, or human-label correction rather than combining the full battery.The cited threads do not jointly cross multiple providers, essays, replications, severity, halo, and versions.
- Positioning: Table 1 positions the study as a combination of a fully crossed judges×essays×replications design, classical rater-effects analyses, cross-lingual public data, version pairs, and a released tensor.The paper does not claim to be first in psychometric analysis, G-theory, or halo observation individually.
3. Method
The method uses operational measurement concepts to audit a crossed panel of LLM judges on public essay corpora. It combines multiple corpora, model versions, replications, canary instrumentation, and pre-registered hypothesis mapping.
- Measurement framework: Severity, halo, dependability φ(k), and family-wise permutation nulls are defined as operational measurement concepts for the audit.Dependability concerns reproducibility across repeated scoring, while the permutation null sets a multiple-comparison threshold.
- Corpora: The study samples 977 Portuguese ENEM essays and 1,400 English ASAP essays from public corpora.ENEM uses five competencies on a 0–1000 total scale; ASAP contributes four genuine essay-writing sets.
- Judges: The roster contains 12 judges from four providers, three price tiers, and five version contrasts.The contrasts include four within-family steps and one same-provider cross-generation step.
- Replication and instrumentation: Each judge scored every essay three times, with a deeper K=5 cell for the cheap tier on two ENEM prompts.Requests and responses were captured verbatim with metadata in an append-only store.
- Replication and instrumentation: Identity canaries repeatedly scored a fixed essay set at temperature 0 to detect mutable model identifiers and serving changes.The canaries detected one judge becoming provider-gated during the collection window.
- Pre-registration: The preregistered battery froze hypotheses, decision rules, exclusions, contamination rules, and analysis specifications before full collection.The protocol was amended once before confirmatory statistics and before English collection.
4. Results
LLM judges show large, judge-specific severity differences that correlation and rank-based agreement statistics largely miss. Severity adjustment did not produce leaderboard reversals beyond the family-wise permutation null, while English rubric paraphrase instability limits generalization.
- Severity: 219 points: judge severity spans this range on ENEM’s 0–1000 scale, with 11 of 12 judges exceeding the preregistered effect criterion.The judge effect is the second-largest variance component, accounting for 16.7% of total variance.
- Severity: 15–33% of score range: ASAP’s judge-panel severity spread contrasts with a 0.7–1.0% gap between two trained human raters.Matched comparisons put the judge severity SD at 7.6–14.6 times the SD implied by the trained pair.
- Severity: τ=.89 in Portuguese but τ=.56 in English: severity ordering was rubric-robust in Portuguese but not English.The English result is a scope condition indicating instrument fragility.
- Severity: .47–.56: all twelve judge–human Pearson correlations occupy this narrow band despite the 219-point severity spread.Location-invariant statistics are blind to severity by construction.
- Leaderboard test: 16 nominal pairwise reversals: severity adjustment changed held-out leaderboard ranks, but the largest reversal, .029, did not exceed the family-wise permutation null.The conservative design could detect only reversal magnitudes exceeding approximately .11 QWK.
4.3 Halo: the same-instrument check reverses the cross-instrument comparison (H3)
The matched same-instrument robustness check reverses the earlier cross-instrument halo comparison. It finds no credible evidence that LLM judges exceed trained humans in residual halo, while reliability remains high under limited replication.
- Cross-instrument comparison: .264: the cross-instrument judge mean residual halo is superseded by the matched same-instrument estimate.The cross-instrument comparison combines ENEM competencies with ASAP English traits and is calibration-sensitive.
- Same-instrument check: .07–.36 and .27–.58: judge residual halo ranges on ASAP sets 7 and 8, versus human ranges of .36/.47 and .53/.59.No judge exceeds the more halo-prone human rater on either instrument.
- Reliability: φ(k)≥.80 at k=1 for nine judges and k=2 for the remaining three, indicating high score reproducibility with little replication.At temperature 0, the cheap tier reaches φ(1)=.90–1.00 without replication.
- Design caveat: The variance decomposition is conditional on stratified essay sampling and only eight prompts, so component ratios may change under unstratified sampling.The paper notes that unstratified sampling would plausibly enlarge the essay component and shrink the judge share.
4.5 From diagnosis to treatment: calibration, cutscores, and monitoring (post-registration, exploratory)
The study tested calibration, threshold, monitoring, and cross-corpus remedies for judge effects. Calibration reduced offset but not all disagreement, version shifts were detectable, and severity did not transport cleanly across corpora.
- Calibration: 191 to 160 RMSE after 100-anchor calibration reduced held-out error, while linear equating reached 154 and no method approached zero.Calibration removed offset but not judge-specific disagreement; the ENEM platform reference is a scale reference, not error-free ground truth.
- Cutscores: 56% of essays passed the ENEM reference’s 600-point threshold yet failed under the most severe cheap judge, nova-lite.The result illustrates the decision consequences of unadjusted severity.
- Monitoring: All five version contrasts exceeded the family-wise permutation null, ranging from −133 points for sonnet-4.5→4.6 to +125 for nova-lite→nova-2-lite.Other shifts were −71, +45, and −19 points; sign and magnitude varied within providers and families.
- Cross-corpus functioning: Judge severity moved −0.84 to +0.42 human-SD between ENEM and ASAP, with 5/11 judges flagged as differentially functioning.Language, rubric, scale, population, genre, human-rater regime, and residual contamination differences are confounded, so magnitudes are exploratory.
5. Discussion and Implications for Learning Analytics
The discussion argues that agreement statistics should be supplemented with rater-effects audits and decision consequences. Crossed scoring makes such audits practical, while the paper’s nulls and limitations bound how results should guide deployment.
- Implications: Researchers should report rank and agreement statistics alongside severity and its decision-scale consequences, not instead of them.The battery costs approximately $460 in API calls and is fully scripted in the artifact.
- Design: The crossed design scored the same 977 essays by all twelve raters three times with zero missing cells, enabling direct checks of model-based claims.Severity orderings matched across MML, JMLE, and raw means at τ = 1.0.
- Nulls: The H2b null found no leaderboard reversal above the approximately .11-QWK detection floor, separating judge selection from calibration at practically consequential magnitudes.The H6 refutation likewise showed that severity shifts directly measure version change without relying on flat agreement.
- Practice: Practitioners should calibrate each judge on their own population and rubric, pin versions, run scheduled canaries, re-equate updates, and use replication for stability rather than validity.The study reports per-judge offsets of 23–189 points and version shifts up to 13% of the ENEM scale.
- Limitations: Findings are conditional on two educational corpora, one platform, one decoding regime, collection window, rubric prompts, and incomplete subgroup information.English severity orderings were prompt-conditional, learner-subgroup DIF was out of reach, contamination probes were weak, and transport beyond the platform was untested.
6. Conclusion
The conclusion frames LLM judges as raters with substantial severity and halo, plus provider-controlled re-instrumentation. These effects can be measured, anchored, and monitored with a released benchmark and established tools.
- Conclusion: LLM judges showed large idiosyncratic severity and halo without credible excess over trained raters, while hosted versions could change or disappear without announcement.The conclusion characterizes provider-controlled re-instrumentation as a failure mode humans do not have.
- Conclusion: Severity can be measured, anchored, and monitored with established measurement tools at a cost of a few hundred dollars.The released tensor and harness support auditing proposed judges, prompts, and calibration methods in the same crossed design.
Ethics Statement
The ethics statement describes reuse of public, anonymized student-writing corpora rather than collection of new identifiable learner data. Essay-BR and ASAP had distinct public-release and redistribution conditions.
- Data use: The study used pre-existing public corpora and created no new data about identifiable learners.Essay-BR was released under an MIT license, while ASAP was publicly released for anonymized research use in 2012.
Data Availability
The study releases a comprehensive score tensor and accompanying materials, including raw captures, rubrics, logs, probe outputs, and pinned analysis code.
- The artifact includes a score tensor with verbatim raw captures, bilingual rubrics, manifests, canary logs, probe outputs, and a deviation log.
- The release also contains analysis code with pinned environments and an OSF mirror of the pre-registration, amendment, and deviation log.
- The files were uploaded after collection, while protocol-before-data ordering rests on the freeze record and deviation log.
Appendix A. Measurement Model and Instrumentation Details
The appendices specify the many-facet measurement model, balanced generalizability design, instrumentation checks, monitoring treatment, and full call accounting.
- Measurement model: The MFRM models essay, judge, competency, replication, and score-category facets, with judge, competency, and replication effects constrained to sum to zero.
- Measurement model: ASAP calibrations are performed separately for each set, with cross-set quantities standardized before comparison.
- Generalizability studies: Generalizability estimates use balanced, complete ANOVA expected-mean-squares components for essay-by-prompt and judge effects with nested replications.
- Instrumentation: Identity canaries score twelve fixed essays in three temperature-0 sweeps, using response distributions as judge fingerprints and a halt-and-void trigger for distribution shifts.
- Monitoring design: The treatment arm uses 200 splits, 400 held-out essays per split, re-estimated calibrations, paired monitoring, and null-calibrated QWK alarms.
- Accounting: The study captured 110,571 calls, including 84,373 confirmatory-grid calls, 12,873 pre-registered auxiliary calls, and 13,325 post-registration calls.
Appendix B. Pre-Registration Deviations (complete)
The deviation log records design, estimation, baseline, robustness, and treatment changes, distinguishing pre-computation decisions from post-registration exploratory cells.
- Appendix B. Pre-Registration Deviations (complete): The amendment was drafted from independent design reviews before main-grid statistics were computed, although a 30-essay pilot had been seen.
- Appendix B. Pre-Registration Deviations (complete): The human baseline switched to the design-matched ASAP set 7–8 trained-rater statistic before any judge residual was computed.
- Appendix B. Pre-Registration Deviations (complete): Sonnet-4 was legacy-gated mid-study, so ENEM data were retained, no ASAP leg was collected, and analyses were reported with and without that judge.
- Appendix B. Pre-Registration Deviations (complete): The G-study estimator changed from REML to ANOVA-EMS because the design was balanced and complete, with definitions unchanged.
- Appendix B. Pre-Registration Deviations (complete): H1 was evaluated on the pre-registered bootstrap scale, while MML preserved the severity ordering with τ=1.0 on a compressed scale.
- Appendix B. Pre-Registration Deviations (complete): A raw-versus-residual halo mismatch was caught before the freeze, and the final analysis used one matched statistic for both sides.
- Appendix B. Pre-Registration Deviations (complete): The pre-writing audit foregrounded nulls, demoted H3, failed H6 as registered, specified per-instrument H5b, and disclosed the English paraphrase failure.
- Appendix B. Pre-Registration Deviations (complete): Post-registration cells tested same-instrument halo robustness, temperature-0 anchors, calibration, cutscore misclassification, version-boundary simulation, and monitoring power exploratorily.