Source-linked AI summary
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
TL;DR
Multilingual LLM judges can reverse evaluator-backbone rankings across prompt languages, motivating calibration without expensive human annotations. The paper uses double-centering to estimate and remove the language-backbone interaction, with formal guarantees including unbiasedness under task-language effects. CBC improves held-out cross-task consistency and external agreement with human gold preferences, while its strongest claims remain bounded by observed-language coverage and two-way-model assumptions.
Problem
Multilingual judge rankings vary across prompt languages, with 7 of 15 backbone pairs showing significant rank reversal and no universal winner across the eight-language benchmark.
Method
CBC double-centers shared multi-backbone score matrices to estimate and subtract the language-backbone interaction without human labels.
Results
CBC raises held-out cross-task rank consistency τ from 0.650 to 0.902 and raises M-RewardBench agreement with public human gold preferences from 68.7% to 76.6%.
Takeaways & Limitations
The paper supports CBC as a label-free post-hoc calibrator for multilingual LLM judges when a shared evaluator ranking across languages is desired.
Takeaways & Limitations
CBC does not solve zero-shot transfer to unseen languages and cannot separate genuine three-way task-language-backbone effects from the estimated interaction.
Abstract
from arXiv · showhide
Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multilingual judge score decomposes additively into task difficulty, backbone skill, and a language-backbone interaction term, the last of which is recoverable without human labels by double-centering the cell-mean score matrix. We make this estimator (\textbf{Consensus-Based Calibration}, CBC) explicit, give an $O(1/\sqrt{n})$ finite-sample concentration bound with variance constant $(1-\tfrac{1}{m})(1-\tfrac{1}{k})$, and show that it is unbiased even when task-language interactions are present. Across 7{,}920 judge runs (6 backbones, 8 languages, 55 tasks, 3 frameworks), CBC raises held-out cross-task rank consistency $τ$ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100\% of per-language decisions versus 68.5\% raw; these are consistency diagnostics, not human-grounded correctness measures. On a separately collected M-RewardBench panel (7 languages, 1{,}500 items per language, 10{,}500 language-item instances, 5 evaluators), panel agreement with the public human gold preferences rises from 68.7\% to 76.6\% (gain 7.9 percentage points, 95\% CI $[6.0, 9.9]$), our strongest external evidence of downstream usefulness. The estimator is the standard two-way ANOVA interaction-recovery operation under sum-to-zero contrasts; our contribution is its application as a label-free post-hoc calibrator for multilingual LLM judges, an explicit finite-sample concentration bound, and an unbiasedness result that holds even under task-language misspecification.
1 Introduction
Multilingual judge rankings can reverse across prompt languages, creating an unresolved calibration problem without expensive human annotations. The paper applies double-centering to recover the language-backbone interaction and evaluates CBC as a label-free calibrator.
- Research gap: Existing approaches use human judgments, ensembles, or pairwise models, but none directly targets the additive language-backbone interaction in multi-evaluator multilingual benchmarks.Hada et al. calibrate against 20,000 human judgments, while Fu and Liu propose an ensemble.
- Approach: CBC double-centers the multi-evaluator score matrix to estimate the language-backbone interaction under a standard two-way additive layout.The model includes task difficulty, backbone skill, language-backbone interaction, task-language effects, and noise.
- Empirical motivation: 7 of 15 backbone pairs exhibit statistically significant pairwise rank reversal, and the top-ranked backbone alternates across all 8 languages.The benchmark contains 7,920 judge runs across 6 backbones, 55 tasks, and 3 frameworks.
- Guarantees: The estimator requires no human labels and is accompanied by an O(1/√n) concentration bound and an unbiasedness result under task-language interactions.The explicit variance constant is (1−1/m)(1−1/k), and the unbiasedness result concerns task-language effects that do not depend on the backbone.
- Scope: CBC is intended for language-invariant evaluator rankings, while language-specific specialization and shared language-wide shifts remain outside its correction target.Shared language-level shifts require an external anchor, and removing meaningful specialization can be inappropriate.
2 Related Work
The paper builds on multilingual judge evaluation, label-free aggregation, two-way ANOVA, and agentic code-evaluation research. Its formal novelty is applying standard interaction recovery to label-free multilingual judge calibration and extending the guarantees beyond correctly specified two-way models.
- Multilingual LLM-judge evaluation: Prior multilingual judge studies report score inflation, weak cross-lingual consistency, and linguistic or cultural bias in benchmark construction.Fu and Liu report Fleiss’ κ≈0.3 across 25 languages.
- Label-free calibration and aggregation: Existing label-free methods estimate judge reliability, model pairwise discrimination, or address position bias, but do not directly recover the additive language-backbone interaction.The related work includes Bradley–Terry–Luce extensions, Dawid–Skene methods, learning-from-crowds, and Item Response Theory.
- Two-way layouts: Double-centering under sum-to-zero contrasts is the standard two-way ANOVA interaction-recovery operation.Classical references assume correct model specification for their standard inference procedures.
- Novelty: The paper adds a finite-sample uniform bound and unbiasedness under an unmodeled task-language interaction to that standard operation.These properties are identified as non-standard relative to the cited classical ANOVA treatment.
3 Score Model and Identifiability
The score model separates task difficulty, backbone skill, language-backbone interaction, task-language effects, and noise. Under balanced shared-task data and sum-to-zero normalization, the target interaction is identifiable from population cell means, but genuine three-way effects are not separable.
- 3.1 Setup and Notation: The judge score S(t, ℓ, b) is modeled as task difficulty plus backbone skill, language-backbone interaction, task-language effect, and mean-zero noise.Framework effects cancel when they do not depend jointly on language and backbone.
- 3.1 Setup and Notation: CBC targets β(ℓ, b) when deployment requires a shared evaluator ranking across languages, not when language-specific evaluator capability is itself the objective.In the latter setting, CBC is best treated as a sensitivity analysis.
- 3.2 Rank Reversal: Pairwise rank reversal occurs when two per-language backbone gaps have an opposite-sign product at least as strong as −δ.The criterion rules out universal dominance between the two backbones.
- 3.2 Rank Reversal: 7 of 15 backbone pairs satisfy the reversal condition, and the empirical top-ranked backbone alternates across the eight languages.The observed data therefore contain no universal winner.
- 3.3 Identifiability Without Human Labels: Under a complete balanced panel, mean-zero noise, and sum-to-zero normalization, β(ℓ, b) is uniquely identifiable from population language-backbone cell means.The absolute task and backbone main-effect levels are not separately identified, but that indeterminacy does not affect β.
- 3.3 Identifiability Without Human Labels: A genuine task-language-backbone interaction lies outside the two-way model and cannot be separated from β by CBC.The estimated interaction can therefore mix evaluation variation with backbone-specific task adaptation.
4 Consensus-Based Calibration
CBC estimates and subtracts the language-backbone interaction using double-centering, requiring no human annotations. It has explicit concentration and consistency guarantees, while remaining unbiased under task-language interactions but vulnerable to genuine three-way misspecification.
- Estimator: CBC estimates β̂(ℓ, b) by subtracting the backbone and language marginal means from each cell mean and adding the grand mean.This is the constructive double-centering estimator under sum-to-zero normalization.
- Estimator: CBC calibrates scores as Ŝ(t, ℓ, b) = S(t, ℓ, b) − β̂(ℓ, b) and requires zero human annotations.It uses multi-backbone evaluation data already collected by labs.
- Finite-sample guarantees: For m=6, k=8, and n=55, the simultaneous high-probability radius is about 10.0 points at ε=0.05.This numerical radius is conditional on independent homoskedastic Gaussian noise and is not distribution-free.
- Convergence and consistency: Without Gaussian tails, CBC remains consistent under independent tasks and finite first absolute moments as n→∞.The O(1/√n) rate is the standard parametric rate for the sample-mean construction.
- Model misspecification: CBC is unbiased for β(ℓ, b) even when task-language interactions are present because those effects cancel in double-centering.This robustness holds because the task-language term does not depend on the backbone.
- Model misspecification: The robustness result does not cover genuine three-way task-language-backbone interactions, and simplified-model error bars may then fail to apply unchanged.Backbone-specific task adaptation can be absorbed into the estimated language-backbone interaction.
5 Experiments
Across internal and external evaluations, multilingual judge rankings vary by language, while CBC improves held-out rank consistency and model-based deployment decisions. External M-RewardBench validation also shows higher agreement with public human preferences, with important scope and interpretation limits.
- Empirical characterization: 7 of 15 backbone pairs exhibit statistically significant pairwise rank reversal, and the empirical top-ranked backbone alternates across all 8 languages.The reversals remain significant after Benjamini–Hochberg correction at FDR = 0.05.
- Model fit and limitations: The two-way model explains substantial variance, with R2=0.691 for the simplified fit versus 0.705 for the full model including task-language interactions.Task-language effects contribute approximately ΔR2≈0.015.
- Calibration results: CBC raises held-out cross-task rank consistency τ from 0.650 under Raw to 0.902 on the eight-language benchmark.The evaluation uses task-level bootstrap resampling and out-of-bag held-out tasks; the full-fit τ=1.000 alignment is only an in-sample sanity check.
- Calibration results: CBC significantly outperforms ComBat-EB, with τ_CBC−τ_ComBat positive in all 1,000 paired-bootstrap replicates (p<0.002).The comparator and controls are evaluated as continuous post-hoc methods on the same observed-language benchmark.
- Decision-level backbone selection: CBC agrees with the held-out additive-model oracle winner in 100% of language-decisions, compared with 68.5% for Raw.The oracle is an unattainable model-based reference, not a human-grounded correctness measure.
- External validation on M-RewardBench: 76.6% of CBC-aligned panel decisions agree with M-RewardBench human gold preferences, versus 68.7% for Raw, a gain of 7.9 points.The 95% CI for the gain is [6.0, 9.9]; this is supportive aggregate panel-level evidence rather than a per-language diagnosis.
- Model fit and limitations: These results do not establish universality across domains, languages, or evaluator families.The internal and external panels use different tasks, rubrics, evaluator routes, and overlapping language sets.
6 Analysis
The paper corroborates language-conditioned evaluator differences with conditional mutual information and identifies which backbone pairs drive the observed reversals.
- Methodological interpretation: CBC estimates the language-backbone interaction by double-centering the multi-evaluator score matrix without requiring human labels.The targeted interaction is the non-redundant signal represented by the conditional mutual information analysis.
- Empirical analysis: 0.178 nats of conditional mutual information separates Language from Score given Backbone and Task, with permutation interval [0.169, 0.187] and one-sided p < 0.002.The estimate averages KSG-style nearest-neighbor estimates over 330 task–backbone slices and uses 500 shuffles per slice.
- Empirical analysis: The seven verified pairwise reversals are exactly the cross-family pairs among GPT-4o, GPT-5.4, Sonnet, and DeepSeek.Qwen is Pareto-dominated in this panel, while Gemini is top- or second-ranked in every language and produces no sign change.
7 Discussion
The paper frames rank reversal as a practical obstacle to finding a language-neutral backbone and presents CBC as a label-free correction when language-invariant ranking is desired.
- Theory meets practice: Language-conditioned pairwise reversals make searching for a language-neutral backbone unreliable.The paper connects this observation to the need to estimate the language-backbone interaction from existing evaluator scores.
- Theory meets practice: CBC removes the language-backbone interaction using data already collected from multiple evaluator backbones scored on shared multilingual items.Its scope is continuous pointwise scores and the two-way language×backbone structure, unlike categorical-label reliability models such as Dawid–Skene.
- Cost and applicability: CBC requires zero human annotations, compared with 20,000 native-speaker annotations used by Hada et al.Reported compute costs are $54.93 for the three-language extension and $195.3 for the M-RewardBench panel.
- Cost and applicability: The same two-way decomposition and calibrator apply to settings with continuous multilingual judge scores beyond agentic code evaluation.The passage names RLHF reward modeling and safety classification as examples.
8 Conclusion
The paper concludes that CBC can recover and remove observed language-backbone interactions, while its claims remain bounded by model misspecification, observed-language coverage, benchmark design, and uncertainty.
- Conclusion: CBC raises held-out cross-task rank consistency τ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100% versus 68.5% raw.These are consistency and model-reference diagnostics rather than human-grounded correctness measures.
- Conclusion: Agreement with public M-RewardBench human gold preferences rises from 68.7% to 76.6%, a gain of 7.9 percentage points with 95% CI [6.0, 9.9].The paper describes this as supportive aggregate evidence and its strongest external evidence of downstream usefulness.
- Conclusion: CBC is the standard two-way ANOVA interaction-recovery operation applied as a label-free multilingual LLM-judge calibrator.The paper’s additional contributions are an explicit finite-sample concentration bound and unbiasedness under task-language misspecification.
- Limitations: Genuine three-way task-language-backbone effects can make the estimated residual mix stable evaluator interaction with task-conditional performance heterogeneity.In that regime, simplified-model error bars need not carry over unchanged, so the estimate is interpreted at benchmark level rather than as universal model-family behavior.
- Limitations: CBC’s strongest results concern observed languages; LOLO agreement improves only from 0.649 to 0.740 under heuristic extrapolation.The paper does not claim that CBC solves zero-shot transfer to unseen languages.
- Limitations: The closed-form estimator and guarantees assume a complete, balanced panel with every shared item scored in every language-backbone cell.Naive double-centering need not identify the interaction with partial or unbalanced coverage.
- Limitations: CBC does not correct a language effect shared uniformly across backbones; addressing such a shift requires an external human or trusted-reference anchor.The small human-anchor experiment is downstream validation, not part of the CBC estimator.
- Limitations: The internal benchmark contains 55 DevAI tasks from one agentic code-evaluation family, while the external panel focuses on one preference-style benchmark family.Together the panels broaden evidence but do not establish universality across domains, languages, or evaluator families.
A Proof Details and Identifiability Scope
Double-centering exactly recovers the centered language-backbone interaction under the two-way model, while identifiability is limited to observed cells and excludes separable three-way effects. CBC remains strong with small backbone subsets, though larger subsets reduce variability.
- Identifiability: Double-centering exactly recovers the centered interaction matrix and guarantees uniqueness under the stated zero-sum constraints.
- Identifiability scope: No human gold labels are needed for observed language-backbone cells, but unseen cells and three-way task-language-backbone effects are not identified.
- Concentration: CBC’s estimation error is a linear combination of centered cell-average noise terms, enabling the Gaussian concentration proof.
- Consistency: The consistency proof uses task-level independence and finite first moments so cell means converge in probability, with linear transformations preserving convergence.
- Backbone ablation: With m=2 backbones, mean rank consistency is 0.901, while standard deviation is 0.206; larger subsets mainly reduce variability.
B.2 Ablation: Number of Tasks (𝑛)
The task-count ablation shows that CBC improves and stabilizes as more tasks are available, consistent with its O(1/√n) convergence behavior. At n=55, estimation error is substantially lower than at n=10.
- Task-count ablation: CBC’s convergence bound scales as O(1/√n), motivating empirical evaluation across task-panel sizes from 10 to 55.
- Task-count ablation: 2.01 at n=10 falls to 0.67 at n=55 for mean absolute interaction-estimation error, consistent with the convergence story.
- Practical power check: At about 52 tasks, the simultaneous high-probability radius falls below half the largest observed interaction magnitude; by n=30, mean pairwise τ is 0.853.
- Requirement-type decomposition: CBC improves rank stability for operational requirements from τ=0.770 to τ=0.789 and semantic requirements from τ=0.737 to τ=0.840.
C External Validation Details
The external validation constructs CBC-ready evaluator tables from multilingual preference margins and compares CBC with a judge-aware Bradley–Terry–Luce baseline. CBC’s external diagnostic is reported alongside panel-construction and implementation boundaries.
- Collection pipeline: The M-RewardBench pipeline converts chosen-minus-rejected margins from fixed 1–5 rubric scores into task×language×evaluator tables for CBC.
- Collection pipeline: 52,500 judged preference pairs correspond to 105,000 pointwise evaluator calls because chosen and rejected responses are scored separately.
- Comparator: The adapted judge-aware Bradley–Terry–Luce baseline reaches τ=0.405 with 95% CI [0.352, 0.467], close to Raw and below CBC.
- Diagnostics: An in-sample full-fit CBC diagnostic reaches τ=1.000, but the paper treats it as a sanity check rather than an independent generalization estimate.
- Panel limitations: The completed validation panel excludes Qwen3 because its provider run produced incomplete chosen/rejected outputs and did not reach a clean terminal state.
- Variant analysis: The weighted CBC variant matches uniform CBC at τ=0.902 on the observed-language bootstrap/OOB metric and is omitted as a separate main-paper method.
D.3 Appendix Note on Dawid–Skene EM
The Dawid–Skene EM ablation treats language-backbone pairs as annotators after binarizing outcomes, but its bootstrap summary is highly unstable for this continuous score-matrix setting.
- Method: Dawid–Skene EM binarizes requirement outcomes and treats each language-backbone pair as an annotator.
- Result: τ=0.446 with 95% CI [−0.042, 1.000] indicates a highly unstable bootstrap summary spanning nearly the full possible outcome range.
- Interpretation: The paper does not treat Dawid–Skene EM as a useful main-table baseline for this continuous score-matrix benchmark.