Source-linked AI summary
A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
TL;DR
The paper addresses whether agreement and robustness to surface changes are sufficient evidence that an LLM judge measures the intended construct. It formalizes and measures construct validity as invariance and sensitivity, finding high invariance but substantially weaker sensitivity, while showing that surface-only predictors can reproduce many public validation labels.
Problem
Prior judge evaluation measures invariance to construct-preserving changes, but does not establish sensitivity to minimal changes in the evaluated construct.
Method
The paper defines construct validity as a two-dimensional profile, measures it across judges and domains with human-directed interventions, and audits label-set recoverability by input mode.
Results
S = 0.945 versus R = 0.319 across 7 judges and 4 domains, with a +0.121 scope–strength sensitivity gap and 55%–67% surface-only label recovery across five public sets in paired mode.
Takeaways & Limitations
High judge agreement can coexist with weak sensitivity to construct changes, so invariance and sensitivity should be reported jointly and validation sets audited for surface leakage.
Takeaways & Limitations
Measured sensitivity is a lower bound and invariance an upper bound because human direction judgments and the limited control family constrain the profile estimates.
Abstract
from arXiv · showhide
LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.
1 Introduction
The paper argues that reliable evaluation is not necessarily valid evaluation: a judge must remain invariant to construct-preserving edits and respond to minimal construct-changing edits. It formalizes this two-dimensional profile, measures it across judges and domains, and audits whether validation labels can be reproduced from surface form alone.
- 1 Introduction: Construct validity requires both unchanged verdicts under construct-preserving edits and changed verdicts under minimal construct-changing edits.Prior judge evaluation primarily tests invariance, so passing that test does not establish validity.
- 1 Introduction: Scope edits expand what a claim applies to, whereas strength edits increase commitment while leaving the covered situations unchanged.The distinction explains why accuracy pressure can affect overgeneralization through different changes in scope and strength.
- 1 Introduction: A label set can reward provenance rather than construct understanding when generative conditions determine both labels and surface form.Such leakage leaves little headroom for sensitivity failures to appear during validation.
- 1 Introduction: The profile uses separate invariance and sensitivity coordinates because they are independent and cannot be fully ordered by one scalar.The paper argues that scalar summaries discard comparisons readers may need.
- 1 Introduction: The paper measures the profile across 7 judges and 4 domains, using human-established intervention direction and disjoint model families for generation, verification, and judging.It also audits five public sets in both input modes and releases interventions, labels, diagnostics, and a checklist.
2 Construct validity for an evaluator
The paper defines evaluator construct validity as a two-coordinate profile: invariance to construct-preserving edits and sensitivity to construct-changing edits. It argues these coordinates are independent, so validity comparisons require the pair rather than a scalar summary.
- Construct validity asks whether an evaluator measures the claimed theoretical construct rather than merely correlating with it.
- A judge maps an item and evaluation criterion to a verdict, while interventions are classified by whether they change the correct verdict.
- The validity profile V(J) = (S, R) records invariance over construct-preserving interventions and sensitivity over construct-changing interventions.
- Proposition 1 shows that S and R are independent, so strong invariance provides no bound on construct sensitivity.
- Degenerate judges can occupy opposite profile corners, with a constant judge at (1, 0) and an always-flipping judge at (0, 1).
- Because scalar summaries can identify degenerate judges or reverse comparisons, the paper compares judges at matched S or along their validity frontiers.
3 Related work
Prior work identifies widespread surface sensitivity and benchmark shortcuts, but often assumes reliability or label validity is sufficient. This paper instead tests whether judges detect construct changes and whether the scope–strength distinction should remain separate.
- A construct-blind predictor reproduces 55%–67% of labels across five public sets, showing why agreement with a validated set can be uninformative.
- Prior judge studies document sensitivity to length, formatting, ordering, metadata, and other properties unrelated to quality.
- Reliability audits and automatic stress tests typically assume that benchmark labels are sound rather than auditing the label sets themselves.
- Two nearby works deliberately unify scope and strength, making the distinction a testable hypothesis rather than an unclaimed gap.
- The paper differs from generation, multimodal counterfactual, and rubric-audit studies by applying construct-validity profiling to evaluators with human direction assignment.
4 Method: the interventions, the direction protocol, and the diagnostic
The method evaluates judges as instruments using construct-changing and construct-preserving edits, with human-verified directions and separate model families. It operationalizes overreach through scope and strength axes while auditing both intervention quality and validation-set headroom.
- 4 Method: The study measures evaluator validity as invariance under construct-preserving edits and sensitivity under minimal construct-changing edits.
- 4.1 Instantiating the two intervention classes: Scope edits grow the referent set while holding commitment fixed, whereas strength edits hold the referent set fixed while increasing commitment.
- 4.1 Instantiating the two intervention classes: Mixed-axis edits are rejected because they cannot isolate whether scope or strength produced the changed verdict.
- 4.1 Instantiating the two intervention classes: The intervention design uses seven construct-changing types and surface-only controls covering hedging, elaboration, register, ordering, verbosity, and format.
- 4.2 Who decides that the verdict should change: Three annotators assign intervention direction, with a preregistered 0.75 agreement bar and disjoint generation, verification, and judging model families.
- 4.2 Who decides that the verdict should change: Low-yield or low-agreement slots are reported as not measurable and excluded, distinguishing generator failures from unclear constructs.
- 4.3 Judges and domains: The profile spans multiple judges and domains, while validation power measures how much discriminative room the label set leaves for judge sensitivity to appear.
5 Experiments: the profile, the asymmetry, and what the rulers could have seen
The experiments measure judge validity as a profile and find strong invariance but weak, structured sensitivity. Audits of public label sets show that surface predictors often recover labels, limiting what agreement demonstrates.
- The profile: S = 0.945 and R = 0.319 on average across 7 judges at matched invariance S ≥0.90.The strongest judge reaches R = 0.561 while maintaining S = 0.945.
- The asymmetry: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges.The gap remains under preregistered bootstrap and permutation checks, including +0.125 among 217 items with unanimous annotator direction.
- The asymmetry: The +0.121 asymmetry is a lower bound because scope edits were longer and the measured length bias worked against the observed effect.Median length changes were 5.0% for scope and 5.3% for strength; the gap also persists when stratifying by whether edits lengthened or shortened text.
- The asymmetry: Sensitivity varies substantially by intervention: quantifier_strengthening and domain_extension reach 0.445, while hedging_removal reaches 0.210.The result indicates that nominally similar overreach edits are not equally detected.
- What the rulers could have seen: Surface-only predictors reproduce 55%–67% of labels across five public sets in paired mode, while provenance-inherited labels reach 99.6%.MT-Bench is especially recoverable at 67.4%, whereas RM-Bench is closest to chance at 54.7%.
- What the rulers could have seen: Per-item and paired validation modes reverse the apparent ordering of public sets, so neither mode dominates and reporting the lower number is not automatically conservative.The paper also notes that the predicted per-axis strength-validation result remains untested because the public sets lack scope/strength labels.
6 Discussion and limitations
The discussion recommends auditing label provenance and reporting impoverished predictors, human ceilings, nulls, and both construct-changing and surface controls. It also bounds the method’s interpretation through measurement, modelling, elicitation, diagnostic, and scope limitations.
- Recommendations: Label provenance should be stated, including whether labels are functions of generative conditions.
- Recommendations: Evaluation reports should include impoverished-predictor agreement, human ceilings with intervals, validation power, label-permutation nulls, and probes in both directions.
- Limitations: R is a lower bound and S an upper bound because human direction errors, coarse verdict spaces, and limited control types constrain measurement.
- Limitations: The scope/strength split is a modelling choice: it explains η2 = 0.40 of slot-level R variance, while a post-hoc grouping explains 0.50.The authors retain the registered split because distinguishing the alternatives requires a construct designed to separate them.
- Limitations: The frontier depends on the judge’s elicitation, since clustered grades yield fewer distinguishable operating points.Table 5 therefore reports realised operating-point counts and regrade agreement alongside each profile.
- Limitations: Validation power is an upper bound tied to a fixed predictor family, and a third evaluation mode cannot be deduced from the two tested modes.The paper also cannot separate whether MT-Bench reflects an incomplete provenance definition or substantial human length preference.
- Limitations: The profile generalises only to the studied domains and constructs whose minimal edits are well posed; the paradox result is correlational.
- Artifacts: The diagnostic harness, blinded labels, edit taxonomy, acceptance gates, and frozen judge versions are released.
7 Conclusion
The paper argues that judge validity requires both invariance to construct-preserving edits and sensitivity to minimal construct-changing edits. Across judges, sensitivity is substantially weaker than invariance, varies systematically by edit axis, and motivates reporting profiles and auditing validation labels.
- Conclusion: S = 0.945 versus R = 0.319, showing judges are markedly more invariant than sensitive to construct changes.The profile covers 7 judges and 4 domains, with direction set by 3 humans at agreement 0.852.
- Conclusion: +0.121 separates scope from strength sensitivity at matched invariance, with the same direction for all 7 judges.The paper reports this gap as structured rather than random and notes that it runs against the only measurable length bias.
- Implications: High agreement is not evidence of valid evaluation because invariance and sensitivity are independent coordinates of the validity profile.The paper further warns that interventions improving agreement have no predicted effect on sensitivity.
- Implications: Threshold comparisons can misidentify validity because a judge dominated across its frontier may still win at default thresholds.The paper therefore reads results from a common frontier rather than vendor-shipped operating points.
- Recommendations: The proposed practice is to report (S, R), compare judges with the best impoverished predictor, and disclose validation-label provenance.The paper describes this as a small reporting change rather than a new benchmark.
- Open questions: Whether strength blindness generalizes beyond this construct remains open and requires a second construct with well-posed minimal edits and inexpensive human verification.The released interventions, direction protocol, and frozen predictor family are intended to transfer to that future test.
A Extended definitions and proofs
The section defines estimators and reporting choices for construct sensitivity, emphasizing equal weighting of base items and dependence-aware uncertainty. It also clarifies that invariance and sensitivity are probabilities over different product spaces.
- Definitions: S and R are defined over different product spaces, so they are not coordinates of one joint distribution and have no defined covariance.In practice, D is the empirical distribution of base items surviving Protocol 1.
- Estimation: The natural estimator weights base items equally rather than interventions equally.Pooling interventions would over-weight items admitting more edits.
- Estimation: Pooling interventions can let hedge-dense sentences dominate the axis most relevant to them because admissibility is not random.Sentences with several hedges admit multiple strength edits, while bare sentences admit none.
- Uncertainty: Confidence intervals resample whole base items with replacement, never interventions, because interventions sharing an item are not independent draws.Nulls use label permutations over the same grouping.
- Agreement: Chance-corrected agreement is supplemented by linearly weighted agreement for ordered verdict spaces, while degenerate constant-rater cases are reported as 0.Weighted agreement distinguishes adjacent-class from opposite-class errors.
B The intervention families, in full
The intervention taxonomy separates construct-changing treatment edits from surface-form controls and freezes their mapping to axes. Generation and verification enforce factual, stylistic, and scope constraints before pairs enter the relevant arm.
- Taxonomy: Table 8 assigns every intervention type to a class and axis, with the taxonomy frozen in code to prevent table–implementation drift.An assertion fails if a type lacks an axis.
- Control arm: The control arm includes register manipulations plus verbosity and formatting controls because prior work implicated them as verdict-moving surface properties.Hedge padding was removed because adding a hedge changes commitment rather than only register.
- Treatment arm: Construct edits are instructed to make claims exceed their evidential support while changing nothing else.The treatment prompt supplies a well-calibrated sentence, evidence context, and exactly one edit type.
- Control arm: Surface-form edits must preserve the claim, scope, factual content, provenance status, and register while staying within a fixed length tolerance.The requested edit changes exactly one scope slot when applicable.
- Verification: The verifier gates entry into the sensitivity arm on factual identity, anchor freedom, register comparability, and the requested slot change; controls additionally require unchanged scope.Human annotators, not the verifier, assign intervention direction, with generation and verification separated by model family.
C Annotation protocol and codebook
The annotation protocol uses blinded pairwise judgments, explicit evidence quotations, and majority labels from three independent annotators. Reliability thresholds determine whether an intervention slot or axis is measurable and eligible for sensitivity analysis.
- Protocol: Human annotators determine which pair member is over-scoped because delegating direction to a model could turn sensitivity into agreement between models sharing a blind spot.The labels identify which member a calibrated reader should prefer.
- Protocol: Pairs are shown in randomized order with evidence context, while source identity, requested slot, axis, and model identities are withheld.Presentation order is recorded so first-position preference remains detectable.
- Labels: Annotators record direction, an exact evidence basis, and whether the sentences preserve the same facts.Cannot-tell pairs are reported separately, while fact-different pairs are discarded as generator failures.
- Labels: Three independent annotators label every pair, and the majority of three becomes the label without an automatic tie-break.All three labeled the full batch despite an earlier two-plus-arbitrator plan.
- Measurability: A slot or axis requires raw pairwise agreement of ≥0.75; otherwise it is reported as not measurable and excluded from R and axis comparisons.The threshold is preferred to κ because κ can be degenerate when one answer dominates.
- Validation: Span-based validation normalizes typography and permits multiple noncontiguous loci so formatting differences and multi-locus edits do not silently reject correct labels.A stricter validation rule can reject correct work without measuring the intended property.
D Judge versions and decision protocols
The paper fixes judge versions, routing, prompts, and evaluation procedures, then audits whether surface-only predictors can reproduce labels. It also distinguishes usable grading resolution from self-reproducibility and evaluates predictors with group-aware cross-validation.
- Judge versions and decision protocols: Aggregating APIs can silently route one model identifier across providers, quantisations, and sampling implementations, so pinning a name does not pin an instrument.The paper treats undisclosed quantisation as a specific hazard because different serving conditions may alter (S, R).
- Judge versions and decision protocols: Every judge runs at temperature 0 on a frozen prompt, with the provider pinned through the routing layer and resolved per call.
- Judge versions and decision protocols: Usable operating points span 2 to 9, while regrade agreement ranges from 0.412 to 1.000, showing that grading resolution and repeatability are distinct.A judge can emit many usable grades yet reproduce its own grade infrequently, or be perfectly repeatable while effectively offering one decision.
- Judge versions and decision protocols: Surface-only validation uses three fixed predictor families ordered by how much response surface they inspect: length, numeric features, and character or word n-grams.Their success localizes whether labels are separable by length, formatting or numeric density, or lexical style.
- Judge versions and decision protocols: Because predictors see response text but not evidence context, reproducing labels indicates that the label set may not require the intended claim-evidence relation.
- Judge versions and decision protocols: Scores use pooled out-of-fold κ from five folds, with GroupKFold by source document when possible and fixed-fold refitting of predictors.Grouping prevents document-level topic, register, or phrasing from leaking across folds; no hyperparameter search is used.
F Register of predicted values
The prediction register records pre-measurement values, their rationale, and refutation conditions so predictions remain falsifiable commitments rather than post hoc descriptions. Several entries concern the strength-axis gap and the possibility that public label sets lack sensitivity to it.
- Register of predicted values: Predicted values are documented with the reasoning that produced them and the measurement that would refute them, making them falsifiable as a set.The register distinguishes predictions from measurements and is intended to prevent one-at-a-time adjustment after results appear.
- Register of predicted values: The register is generated from a source file and the build fails when a rendered predicted value lacks its explanation, preventing transcription drift.
- Register of predicted values: Predictions are recorded for mechanical figure-rendering reasons, while the paper states that no headline claim argues from a predicted value.
- Register of predicted values: RstrengthAx 0.24 predicts that hedge and condition deletion can remove markers of what changed, making strength edits difficult for judges to detect lexically.
- Register of predicted values: The predicted axis gap is 0.33, with refutation defined as a gap below 0.10 or failure to reach significance under paired bootstrap testing.
- Register of predicted values: vpStrength 0.03 predicts that no public set has headroom on the strength axis, preventing detection of the paper’s asymmetry.
- Register of predicted values: humanAgree 0.81 exceeds the protocol’s pre-set agreement bar of 0.75.
G What is not measured here
This section identifies a deliberately unmeasured condition for interpreting weak labels: if agreement falls below 0.75, the affected axis is not treated as measurable.
- What is not measured here: Agreement below 0.75 makes the affected axis reported as not measurable rather than measured with weak labels.
H Secondary experiments and background work
Secondary analyses separate headline claims from robustness checks, examine verifier behavior, and document alternative decompositions that were rejected. These analyses support interpretation without replacing the main results.
- Secondary experiments and background work: The paper explicitly excludes secondary analyses from headline claims, using them to rule out alternatives, bound design choices, or document rejected paths.
- Secondary experiments and background work: Three specified robustness checks remain unrun, including remedies expected to raise S without moving R and finer verdict-granularity analysis.
- Secondary experiments and background work: A verifier notices 24.1% of scope edits versus 8.8% of strength edits, a factor of 2.7, and assigns the same axis as the authors 82% of the time when it detects change.
- Secondary experiments and background work: The verifier is not ground truth because it performs a different task and human annotators did not assign direction, so its result is reported only as an independent instrument.
- Secondary experiments and background work: A verifier-based gate rejected all 12 verified strength edits, illustrating why direction was determined by humans rather than filtered by model detectability.
- Secondary experiments and background work: The authors rejected alternative decompositions based on a single graded magnitude, a three-way quantifier-temporal split, and evidence type.