Source-linked AI summary

Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale

Weiyue Li, Minda Zhao, Weixuan Dong, Jiahui Cai, Yuze Wei, Michael Pocress, Yi Li, Wanyan Yuan, Xiaoyue Wang, Ruoyu Hou, Kaiyuan Lou, Wenqi Zeng, Yutong Yang, Yilun Du, Mengyu Wang

arXiv:2601.03444v1cs.CLcs.AIcs.HC

TL;DR

LLM judges can be inconsistent, while the effect of grading scale on their consistency and agreement with humans has remained underexplored. The paper compares human and LLM ratings across three scales and six benchmarks using ICC-based reliability analysis. Across tasks, 0-5 produces the strongest human–LLM alignment, while subjective benchmarks and pooled reliability reveal important heterogeneity.

  • Problem

    The effect of grading scale on LLM-judge consistency and human–LLM agreement remains underexplored despite existing concerns about evaluation inconsistency.

  • Method

    The study collects human and LLM ratings across three scales and six objective, subjective, and mixed benchmarks, evaluating reliability with ICC and benchmark-level analyses.

  • Results

    Across diverse tasks, 0-5 maximizes average human–LLM absolute agreement, while 0-10 is consistently weakest; subjective benchmarks show lower inter-scale consistency.

  • Takeaways & Limitations

    Scale design and benchmark- and subgroup-level diagnostics should be treated as standard components of LLM-as-a-judge evaluation protocols.

  • Takeaways & Limitations

    The human raters are all graduate students, so their calibration and reliability may not represent broader annotation populations.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is altered. However, the effect of the grading scale itself remains underexplored. We study the LLM-as-a-judge problem by comparing two kinds of raters: humans and LLMs. We collect ratings from both groups on three scales and across six benchmarks that include objective, open-ended subjective, and mixed tasks. Using intraclass correlation coefficients (ICC) to measure absolute agreement, we find that LLM judgments are not perfectly consistent across scales on subjective benchmarks, and that the choice of scale substantially shifts human-LLM agreement, even when within-group panel reliability is high. Aggregated over tasks, the grading scale of 0-5 yields the strongest human-LLM alignment. We further demonstrate that pooled reliability can mask benchmark heterogeneity and reveal systematic subgroup differences in alignment across gender groups, strengthening the importance of scale design and sub-level diagnostics as essential components of LLM-as-a-judge protocols.

1 Introduction

LLM judges offer scalable evaluation but can produce inconsistent scores, and the grading scale itself has been largely overlooked as a source of variation. This study compares human and LLM ratings across scales and benchmarks to identify how scale choice affects reliability and human–LLM agreement.

  • Motivation: LLM-as-a-judge systems provide fast, low-cost, reconfigurable evaluation but can vary with prompt format, context, or judge model.Such inconsistency weakens result repeatability and raises concerns about judge reliability.
  • Research gap: Prior work mainly studies aggregate human correlations, prompting or training strategies, and consistency under resampling or format changes.The grading scale remains comparatively underexplored despite practical use of alternatives such as 0-5 and 0-10.
  • Approach: The study collects human and LLM ratings across multiple scales and six diverse benchmarks spanning objective questions and open-ended subjective judgments.It uses a fully crossed rater design and intraclass correlation coefficients to evaluate score reliability.
  • Main findings: Across tasks, 0-5 maximizes average human–LLM absolute agreement, whereas 0-10 is consistently the weakest choice despite high within-population reliability.The best scale can still depend on the benchmark, motivating per-task diagnostics rather than a one-size-fits-all choice.
  • Main findings: On subjective open-ended quality benchmarks, model–model agreement drops sharply, and both male and female annotator groups align more on 0-5 in the study’s pool.These findings support treating scale design as a controllable component of LLM-as-a-judge protocols.

2 Related Work

Prior research has expanded LLM-based evaluation and studied bias and self-consistency, but direct evidence on how grading scales affect consistency and human–LLM agreement remains limited. Existing work motivates examining scale choice as a distinct evaluation factor.

  • LLM-based evaluation: LLM judges are used across generation, translation, code, and dialogue evaluation, with specialized models and architectures proposed for robustness and rubric following.These frameworks offer scalable evaluation but do not eliminate concerns about judge behavior.
  • Known limitations: Documented judge biases include favoring longer responses or later options, prioritizing fluency over factual accuracy, and assigning high scores to same-model outputs.These issues complement concerns about consistency and alignment with human assessment.
  • Consistency: LLM self-consistency research examines whether evaluations remain stable under repeated scoring or altered prompts and formats, finding preference cycles and inconsistent criterion application.One study reports failures to maintain consistent preferences in roughly 25% of challenging cases.
  • Scale and reliability: Prior human-evaluation research links rating-scale design with reliability, while this paper applies intraclass correlation coefficients to assess absolute agreement and rater and item variance.The related literature therefore provides motivation for examining scale effects beyond aggregate correlation.
  • Human–LLM agreement: Earlier human–LLM studies typically fix one scoring scale and do not compare agreement across scales, leaving scale-invariant alignment an open question.The paper positions its study as the first to provide empirical evidence on how score-scale choice affects human–LLM alignment.

3 Methodology

The study measures rater reliability and human–LLM absolute agreement with ICC-based analyses, modeling item and rater effects through a two-way random-effects ANOVA. It compares human, LLM, and cross-population reliability across grading scales and benchmark subsets.

  • Reliability metrics: The two-way random-effects ANOVA treats items and raters as samples from broader task and rater populations.The model includes a grand mean, random item effects, random rater effects, and residual error.
  • Reliability metrics: ICC(A, 1) measures the reliability of one rater, whereas ICC(A, k) measures the reliability of the mean across k raters.Both statistics are derived from ANOVA mean squares for items, raters, and residual error.
  • Agreement metric: nMAE measures average absolute deviation between LLM scores and human reference scores after normalization by the rating-scale range.Lower nMAE indicates closer agreement, and the metric is intended to support comparisons across tasks with different score ranges.
  • Human reliability: Human reliability is computed for individual and averaged panels, separately by gender and benchmark subset.The human rating matrix contains all scores for each scale, enabling pooled and per-benchmark estimates.
  • LLM reliability: LLM reliability is computed for individual judges and ensemble averages, with ICCs recomputed for each benchmark subset.The LLM rating matrix contains scores from all judges for a given scale.
  • Human–LLM agreement: Human–LLM agreement compares the human consensus mean with either the LLM ensemble mean or an individual model using ICC(A, 1).The analysis uses an S × 2 matrix containing the human consensus and model score.

4 Experiments

The experiments evaluate inter-scale consistency and human–LLM agreement across six benchmarks using three grading scales. They use fully crossed human ratings, six LLM judges, pooled and benchmark-level analyses, and randomized scale blocks for humans.

  • Experimental design: The study conducts experiments on LLM inter-scale agreement and human–LLM agreement across six established benchmarks.The experimental pipeline summarizes these two experiment sets.
  • Data: Human–LLM agreement uses 150 sampled items, with 25 items drawn from each benchmark and analyses at pooled and benchmark levels.The benchmark subsets are treated as samples from a broader target-task distribution.
  • Raters: The human panel contains 12 graduate annotators, six female and six male, while the LLM panel contains six judges from diverse model families.Both groups score every item on every scale in the experimental setup.
  • Raters and scales: The experiments compare 0-5, 0-10, and 0-100 scales, allowing fractional scores while keeping each scale’s numeric range separate.Human scale blocks are randomized and item order is shuffled; LLMs receive a standardized rubric and target scale.
  • Analysis: The analysis computes pooled and per-benchmark human ICCs, LLM ICCs, and human–LLM ICCs for each grading scale.Inter-scale ICCs are computed after linearly mapping scores to [0,1].

5 Results

Across benchmarks, grading scale materially affects human–LLM agreement and inter-scale consistency, with the strongest alignment generally on 0-5. Reliability and alignment vary sharply by benchmark, model, subgroup, and decoding conditions.

  • LLM inter-scale agreement: Subjective open-ended benchmarks show substantially lower inter-scale agreement than objective or mixed benchmarks, especially for smaller or less capable judges.Scale changes can alter how LLMs execute the rubric rather than merely rescaling scores, making results across scales not necessarily comparable.
  • Scale-level comparison: 0-5 produces the strongest pooled human–LLM alignment, whereas 0-10 is the weakest despite highly reliable human and LLM panels within each scale.Changing the numeric range shifts rater calibration even when within-group reliability is near-perfect.
  • Benchmark heterogeneity and the reliability illusion: Pooled reliability can conceal benchmark heterogeneity: LLM panels are reliable on STS-B and ToxiGen but substantially less reliable on MT-Bench and SummEval, while human panels remain consistent.This reliability illusion motivates interpreting judge reliability at granular benchmark levels.
  • Per-benchmark human–LLM agreement: Human–LLM agreement is strongest on STS-B and ToxiGen, more moderate on MoralChoice, and consistently weak on SummEval, MT-Bench, and TruthfulQA.The 0-5 scale remains strongest in most benchmark comparisons despite benchmark-dependent patterns.
  • Model-wise agreement: GPT achieves the strongest pooled human alignment across scales, followed by Gemini, while most models show their strongest agreement on 0-5.The results indicate that both model choice and grading scale matter for human–LLM agreement.
  • Temperature robustness: Across temperatures from 0.1 to 1.0, agreement is largely insensitive to temperature and scale ordering remains stable, indicating conclusions are not driven by decoding strategy.Scale-dependent calibration differences dominate sampling randomness in this analysis.

6 Discussions

Across six benchmarks, grading scale materially affects LLM self-consistency and human-LLM agreement. Aggregated results favor 0-5, while pooled reliability can conceal benchmark-level differences.

  • 0-5 yields the strongest absolute agreement between LLM judges and human consensus across tasks, while 0-10 is consistently weakest.The ordering is stable under temperature perturbations for representative models.
  • On subjective open-ended tasks, LLM scores are not perfectly self-consistent across grading scales.
  • Human-LLM alignment shifts materially with the grading interface even when within-population panel reliability remains high.
  • Pooled reliability can mask substantial benchmark heterogeneity because objective-like tasks dominate aggregate variance.
  • Scale design and diagnostic reporting should be treated as standard components of evaluation protocols, especially for ambiguous, high-variance generations.

Limitations

The study’s generality is constrained by its graduate-student human sample and by difficult, under-specified items that can produce uncertainty in human ratings.

  • Human raters are all graduate students, so their calibration and reliability may not represent broader annotation populations.Educational background, LLM familiarity, and cultural context may influence subjective evaluations.
  • Under-specified or specialized items can produce uncertainty and differing rubric interpretations among human raters.This uncertainty may lower human-LLM agreement and make ICC estimates sensitive to benchmark composition.

Ethical considerations

Evaluative scores on normative and subjective benchmarks require cautious interpretation because they reflect context-dependent standards and specific annotator perspectives. The study also addresses annotator exposure to sensitive content through voluntary participation, informed consent, and anonymity.

  • LLM scores for toxicity and moral evaluation should not be treated as objective ground truth because they reflect context-dependent standards embedded in models.
  • Human ratings on normative benchmarks reflect a specific annotator population and may not generalize to broader moral or social norms.
  • Annotators participate voluntarily, provide informed consent, and receive anonymity despite possible exposure to sensitive or harmful content.

B Dataset Overviews

The study evaluates six datasets spanning objective-like and subjective open-ended judgments, using fixed prompts and a common annotation protocol across three scoring ranges. Results show strong LLM reliability on objective-like tasks but substantially lower reliability on subjective open-ended benchmarks.

  • Dataset coverage: The six datasets cover semantic similarity, toxicity, moral evaluation, truthfulness, instruction following, and summarization.They include STS-B, ToxiGen, MoralChoice, TruthfulQA, MT-Bench, and SummEval.
  • Dataset coverage: Both human annotators and LLM judges receive task-specific prompts formatted around inputs such as sentence pairs, question-answer pairs, and document-summary pairs.
  • Annotation protocol: Human annotations use randomized item order and task-appropriate criteria for assigning a single overall numerical quality score.
  • Annotation protocol: The evaluation protocol holds wording criteria and interface constant across datasets while varying only the numeric range: 0-5, 0-10, or 0-100.
  • Annotation protocol: Annotators may use fractional values within each scale, and repeated dataset evaluations use separate sessions and independently randomized item orders.
  • Interpretation and risk: Automated scores should be treated as supporting signals rather than definitive judgments, particularly in high-stakes settings.
  • Reliability pattern: The overall reliability pattern remains consistent across alternative scales: LLM panels are highly reliable on STS-B and ToxiGen but substantially less reliable on MT-Bench and SummEval.Human reliability remains strong across all benchmarks.

G Error Analysis Patterns Examples

Representative examples of poorly and well-aligned judgments are presented in Tables 16 and 17 to illustrate the reported error patterns.

  • Representative examples are provided to illustrate the paper’s error patterns.
  • Tables 16 and 17 contain the representative examples.
  • The examples cover both poorly aligned and well-aligned cases.

H The Use of LLMs

The paper documents its evaluation materials across multiple datasets, task types, prompts, and representative alignment cases. It also states that LLMs were used only for proofreading and grammar polishing, not for substantive research decisions.

  • LLM use: LLMs were used only for proofreading and grammar polishing, while authors made all substantive content decisions.
  • Datasets and tasks: The study includes datasets spanning semantic similarity, toxicity, moral evaluation, instruction following, summarization, and truthfulness.
  • Prompts: Separate prompt templates instantiate the evaluation tasks with rating scales and task-specific inputs.
  • Alignment examples: The materials include representative poorly aligned cases and well-aligned cases for comparison.
  • Poorly aligned cases: Poor alignment can result when judgment depends on a single specific factual detail or niche link.
  • Well-aligned cases: Well-aligned judgments can draw on multiple supporting facts or overlapping similarities.
Loading 2601.03444v1…