Source-linked AI summary
IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages
Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ricardo Rei, Frédéric Blain, André F. T. Martins, Marco Turchi, Matteo Negri, Rajen Chatterjee, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
TL;DR
Indic QE and APE evidence is fragmented across releases, limiting unified training and evaluation across tasks and language pairs. INDICQE-APE consolidates these resources into one benchmark and shows that only conflicts between holistic and token-level quality signals remain difficult after matched controls.
Problem
Indic QE and APE data remain fragmented across releases, complicating unified training and evaluation across tasks and language pairs.
Method
The paper unifies Indic QE and APE data into a 126,754-instance benchmark with aligned labels and a difficulty-stratified test set.
Results
Only signal conflict survives controls matched on language and human score, while within-language QE correlation does not ensure cross-lingual comparability.
Takeaways & Limitations
QE difficulty and score comparisons should use language- and score-matched controls rather than rely on within-language accuracy alone.
Takeaways & Limitations
The benchmark has heterogeneous provenance, inferred domain labels, model-generated explanations, and post-edits rather than independent references throughout.
Abstract
from arXiv · showhide
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes. On it, we benchmark six prompted LLMs and three COMET metrics on segment-level QE, and three systems on APE. Two of the axes are defined partly on the direct assessment and select a compressed slice of it, so each axis is compared against a control drawn from the same language pair with the same score distribution. Only one survives that control: segments whose holistic and token-level quality signals conflict are ranked worse than equally-scored segments of the same language, for all nine systems and all seven pairs that carry the axis. Annotator disagreement, which looks second-hardest without the control, has no effect with it. Few-shot prompting costs every model $\leq$ $3.4$B both correlation and output-format compliance. Within-language accuracy does not make scores comparable across pairs: of the three trained metrics, the one with the best within-language correlation loses most when the pairs are pooled. The benchmark and code will be released.
1 Introduction
INDICQE-APE unifies fragmented Indic QE and APE data into a benchmark with aligned labels, broad language-pair coverage, and task-specific baselines. It shows that cross-lingual comparability differs from within-language skill, while only signal conflict remains a controlled difficulty effect.
- Motivation: Fragmented Indic QE and APE releases previously required corpus reconstruction because labels, schemas, splits, and provenance were incompatible.The fragmentation affected direct assessments, post-edits, word-level tags, and error explanations.
- Benchmark: 126,754 instances across nine directional pairs unify WMT 2020–2024 data with an extended English→Malayalam resource and four aligned label types.Instances can be sliced by task and language pair; the benchmark also includes difficulty stratification, MQM, domain labels, provenance, and duplication and leakage auditing.
- Findings: Within-language correlation and cross-lingual comparability diverge: CometKiwi-XL has the highest within-language correlation but the weakest cross-lingual agreement among trained metrics.Pooling language pairs costs CometKiwi-XL more than the other trained metrics because its per-language offset moves against quality.
- Findings: Of four difficulty axes, only holistic–token-level signal conflict remains harmful after matching language and human-score distributions; four-shot prompting costs every model at most 3.4B in correlation and format compliance.Annotator disagreement appears second-hardest without the control but has no controlled effect; the benchmark also studies frozen-LLM layer probing and a COMET regression head dominated by encoder initialization.
- Evaluation: Six prompted LLMs and three COMET metrics are evaluated on QE over nine pairs, while three systems are evaluated on APE with output-format compliance reported alongside correlation.The benchmark provides segment-level QE and APE baselines across the stated task and pair coverage.
2 The INDICQE-APE Benchmark
INDICQE-APE consolidates Indic QE and APE resources into one test-first, difficulty-stratified benchmark of 126,754 instances across nine directional pairs, with up to four aligned label types. Its provenance, deduplication, source-disjoint splits, and derived word tags support controlled evaluation across tasks and language pairs.
- Label types and domain completion: Up to four aligned label types—direct assessment, human post-edit, word-level OK/BAD tags, and explanation—serve QE, APE, and explainable-QE without re-alignment.The release contains 8,951 human Translation Quality Remarks and 14,644 model-generated error descriptions.
- Consolidation and lineage: 126,754 instances across nine directional pairs unify WMT QE, APE, domain-specific Indic QE, and an extended English→Malayalam resource in one release.Every instance records its source corpora and a human-readable cadence_id.
- Difficulty-stratified design: 59.8% of DA-bearing challenge segments carry at least one difficulty axis, compared with 30.6% outside the test, with enrichment in all nine pairs.The four axes ship as overlapping flags; only 3,265 of 7,616 flagged challenge segments sit on exactly one.
- Provenance, deduplication and split hygiene: 234,446 corpus contributions collapse to 126,754 rows through content-hash deduplication, with 77,505 rows carrying more than one source.Named configurations are exact views of the master table, and splits are source-disjoint.
- Word-level tags: 0.963 Matthews correlation reproduces released word-level tags from MT and post-edit over 633,988 tokens.The tags are deterministic post-edit alignments, and agreement is lowest for en-ml and en-mr and highest for X→en pairs.
3 QE Baselines
The baselines evaluate prompted LLMs and trained COMET metrics on a common challenge test, while separate analyses show that within-language correlation does not ensure cross-language comparability. Matched controls identify holistic–token quality disagreement as the only robust difficulty axis, whereas few-shot prompting and template changes generally hurt or modestly affect performance.
- Setup: 13,032 segments form the common challenge test, with 12,730 direct assessments used for every correlation; zero-shot prompted models score 12,552–12,730 segments.Trained metrics score the full test, while prompted-model coverage depends on usable parsed outputs.
- Cross-language comparability: 0.671 within-language Spearman for CometKiwi-XL contrasts with cross-language agreement of r = 0.34, versus 0.91 for XCOMET-XL.The comparison uses the full-population pass over six en→Indic pairs, not the challenge-test results elsewhere in the section.
- Cross-language comparability: 0.117 is CometKiwi-XL’s loss when nine pairs are pooled, from 0.671 macro to 0.554 pooled Spearman; XCOMET-XL loses 0.001, from 0.498 to 0.497.Pooling can credit metrics for between-pair mean differences rather than within-pair ordering.
- Where models fail: −0.146 is A4’s matched-control effect, with negative changes on all seven carrying pairs and for all nine systems; A1’s effect is −0.003 [−0.047, +0.037].Matched controls use the same language pair and direct-assessment score histogram, removing apparent difficulty caused by score-range compression.
- Where models fail: 1,149 A4 segments across four pairs have below-par holistic scores but little editing, producing contrast −0.070 and negative effects on every pair.The surviving phenomenon is disagreement between holistic and token-level quality signals, rather than edit effort or word error alone.
- Prompting and templates: 0.069 is the macro-Spearman difference associated with shot count, versus 0.026 for template choice; SQM is better on 51 of 94 qualifying cells.Few-shot prompting lowers macro Spearman for four of six prompted models and also reduces compliance, with comparisons based on differently scored segments.
4 Layer-Probing Quality Estimation
A trained single-layer probe reads quality from frozen tiny-aya hidden states across six English–Indic pairs, with the strongest average signal at layer −20. It outperforms prompting the same backbone and approaches GPT-5.5, but layer-specific predictions are unavailable for difficulty-axis analysis.
- Layer selection: 0.527 is the best average Spearman layer score, achieved at −20; the sweep endpoints are weaker at −5 (0.501) and −24 (0.496).The best layer is otherwise pair-dependent, while selecting a layer per pair raises the macro to 0.539.
- Comparison with prompting: 0.527 probe macro Spearman beats tiny-aya GEMBA-DA’s 0.202 across the same six pairs.The probe is ahead on every pair in that comparison.
- Comparison with prompting: 0.051 is the gap between the trained read-out and GPT-5.5, which scores 0.578 on the same six pairs.The probe nevertheless exceeds GPT-5.5 on en-te, scoring 0.282 against 0.255.
- Limitation: The probe is reported only at pair level because the layer sweep saved aggregates rather than per-instance predictions.Consequently, it is the only study here that cannot be joined to the difficulty axes.
5 A Lightweight QE Head on COMET
A lightweight regression head over a COMET/CometKiwi encoder reaches near-parity with zero-shot COMET baselines, with gains driven primarily by the encoder backbone rather than added features. Its axis analysis agrees with the panel on three of four axes, while diverging on A1.
- Method: The head regresses per-pair z-scored direct-assessment targets from source–hypothesis pairs using a Huber loss over upper encoder layers.Training uses six English-to-Indic pairs with 72,806 segments and evaluation uses a matched 9,730-segment challenge test.
- Backbone effects: 0.140 is the largest substitution gain: CometKiwi reaches 0.580, ahead of COMET-DA at 0.491 and XLM-R-large at 0.440.The ordering reflects reference-free QE pretraining, reference-based pretraining, and no QE pretraining, respectively; Williams t=20.5, p < 0.001, n=8,160.
- Performance: ρ = 0.596 for the trained head versus 0.588 for zero-shot CometKiwi-DA and 0.608 for CometKiwi-XL on the six-pair lane.The 0.008 margin over zero-shot CometKiwi-DA is not interpreted as an improvement.
- Axis analysis: The head reproduces three of four axis findings: A4 is negative and A2 and A3 are positive.For A4, the estimates are −0.176 funnel-only and −0.143 with features against a direct-assessment-matched control, negative on every pair for both.
- Axis analysis: A1 diverges from the panel: the head estimates −0.084 and −0.050 versus the panel’s −0.003.The head’s A1 effect is concentrated in en-gu at −0.41 and en-ml at −0.24, compared with +0.16 on ne-en.
6 Automatic Post-Editing
The benchmark supports zero-shot automatic post-editing evaluation on four English→Indic pairs, but unedited MT is best on three pairs under reference-based metrics. Reference-free scoring reverses this pattern, while the negative result is limited to prompting because available training triples were unused.
- Evaluation setup: Three systems—sarvam-m, sarvam-t, and tiny-aya—are evaluated zero-shot on four en→Indic pairs using human post-edits as references.sarvam-m and tiny-aya edit source-plus-MT inputs, whereas sarvam-t retranslates from the source alone.
- Reference-based results: On three of four pairs, do-nothing MT is best under char-TER and chrF++, including char-TER 26.1 versus 36.5 on en-mr.The exceptions are en-hi, where a system improves over leaving the MT unchanged.
- Metric dependence: CometKiwi ranks every system above unedited MT in 11 of 12 cells and orders them sarvam-t above sarvam-m above tiny-aya on all four pairs.chrF++ matches that order on en-hi but reverses it on the other three pairs.
- Metric dependence: The metric disagreement arises because reference-based surface scores measure residual repair effort against post-edits, while CometKiwi evaluates without a reference.Editing MT makes the MT close to its post-edit reference by construction; sarvam-t instead retranslates from the source.
- Scope and limitation: The negative result applies only to prompting: the 61,032 benchmark training triples were unseen, and trained systems could learn the conservative editing rewarded by the references.The do-nothing row is therefore a baseline to beat, not a ceiling on APE performance.
7 Conclusion
IndicQE-APE unifies Indic QE and APE data into a leakage-checked benchmark spanning 126,754 instances and nine directional pairs. Its results show that cross-pair QE scores are not directly comparable, while APE evaluations require a do-nothing baseline and several benchmark limitations remain open.
- Benchmark contribution: 126,754 instances across nine directional pairs form a unified Indic QE/APE benchmark with aligned labels, difficulty-stratified testing, verified provenance, and no train/test leakage.The benchmark consolidates the WMT 2020–2024 lineage with an extended English→Malayalam resource.
- Cross-pair comparability: Raw reference-free QE scores are not comparable between language pairs, and pooling pairs can reverse conclusions drawn from within-language correlations.Among five systems with per-language offsets, the best within-language correlator is the only one whose offset runs against quality and the only trained metric losing more than rounding when pairs are pooled.
- APE evaluation: APE results against post-edit references require a do-nothing row because unedited MT beats every tested post-editing system on three of four pairs under all four surface metrics.A reference-free metric ranks the post-editing systems above unedited MT, illustrating why reference-based results alone can mislead.
- Limitations and future work: Open issues include missing baselines for word-qe and explainable-qe, inadequate whitespace alignment for sub-word tagging in agglutinative targets, and unresolved human-score/surface-evidence disagreements.English→Telugu is the lowest-scoring pair for seven of nine systems, with its narrow score range explaining only part of that result.
Limitations
The benchmark’s conclusions are constrained by heterogeneous and partly inferred annotations, incomplete or coarse labels, and limited measurement designs. Model and difficulty analyses also have restricted coverage, comparability, and statistical identification.
- Data: Data provenance is heterogeneous, domain labels are partly inferred, explanations are model-generated outside en-ml, and post-edits serve as references throughout.The classifier inferred 3,756 of 12,730 challenge-item domain labels at 0.826 accuracy; X→en labels were inferred out of domain.
- Measurement: One translation per source segment prevents system-level ranking, while the cross-lingual coefficient has at most nine weakly identified points.Prompted models use non-identical populations, so correlation and compliance rates must be interpreted together.
- Labels: MQM labels cover only en-hi, fewer than one-third of DA-paired segments have trustworthy spans, and derived word tags are validated rules rather than human judgments.The rule reaches 0.976 and 0.987 token accuracy against real labels, but derived tags agree least for en-ml and en-mr and are coarser for agglutinative targets.
- Difficulty axes: Difficulty axes partly encode annotation properties; matched controls remove score-range differences but not label noise, and attenuation correction covers only three pairs.Only A4’s larger arm is readable, and its size halves without en-ml although its sign remains unchanged.
- Models: The probe uses one layer over one backbone without saved predictions, its cross-lingual coefficient is nonsignificant, and prompted APE covers four of seven pairs without training.The APE setup establishes a do-nothing baseline rather than reachable quality; the probe ablation predates a late data fix.
Ethics Statement · A Related Work
IndicQE-APE consolidates dispersed WMT QE and APE resources into a traceable benchmark, while documenting the provenance and annotation process of its English→Malayalam data. It builds on established reference-free metrics, prompted LLMs, and representation-based QE methods, while addressing incompatibilities and missing annotations across prior releases.
- Ethics Statement: The dataset combines publicly released shared-task corpora with an English→Malayalam resource produced through documented commercial annotation and quality-control procedures.Corpus-level provenance will allow each instance to be traced to its source.
- Ethics Statement: The English→Malayalam data was annotated under specified guidelines by annotators recruited and paid by a commercial agency rather than through per-task crowdsourcing.Appendix L records who annotated the data, the guidelines used, and the quality-control process.
- A Related Work: Reference-free COMET and CometKiwi estimate quality from source and hypothesis without references and return scalar scores without exposing per-language behavior.These learned metrics are trained on human direct assessments and represent current practice in WMT metrics and QE tasks.
- A Related Work: WMT QE shared tasks supply the benchmark’s direct-assessment data across X→en, English–Marathi, en-hi, en-gu, en-ta, and en-te pairs.MLQE-PE introduced multilingual QE; later WMT editions added the Indic pairs and new test sets.
- A Related Work: WMT APE shared tasks contribute en-mr data from 2022 and 2023 and en-hi/en-ta data from 2024, while the benchmark extends a recent English→Malayalam release.The cited APE editions provide the post-editing lineage consolidated by IndicQE-APE.
- A Related Work: Prior resources are separated by task and edition, use incompatible schemas, and omit combinations of explanations, MQM labels, difficulty structure, or post-edits.MLQE-PE provides direct assessments, post-edits, and word-level tags but no explanations, MQM, or difficulty structure; WMT QE and APE editions have complementary gaps.
- A Related Work: Instruction-tuned LLM prompting provides the prompted-QE baseline, while adaptive layer optimisation selects and weights hidden layers of a frozen LLM for QE.The benchmark uses both approaches as related methodological baselines or components.
B Construction Detail · C QE Meta-Evaluation
The benchmark construction prioritizes empirically validated tokenisation, explicit tag provenance, source-level domain completion, and transparent annotation reliability. Its meta-evaluation separates difficulty from annotator agreement and reports results against matched controls and per-pair compliance.
- B Construction Detail: IndicNLP tokenisation reaches 0.960 exact-sequence agreement with shipped mt_tokens, versus 0.660 for Moses and 0.313 on Malayalam.Agreement is 1.000 on en-ml, and released tag sequences retain the tokenisation they index.
- B Construction Detail: 33,119 of 38,616 tagged segments, covering 633,988 tokens, are rederived from MT and post-edit alone.Of the 5,497 omitted segments, 5,473 lack a post-edit and 24 en-ml rows have mismatched tag and tokenisation lengths.
- B Construction Detail: On a shared 500-segment sample, the classifier scores 0.804 versus 0.566–0.618 for three prompted GPT configurations.The advantage is concentrated in provenance-defined classes, with F1 gains of +0.19 to +0.34; health and legal gaps are +0.03 and +0.01.
- B Construction Detail: Released provenance distinguishes upstream, mlqe-peupstream, and derived tags, while 98.6% of X→en rows retain MLQE-PE imports.The remainder and 351 conflicting duplicate upstream segments are left to derivation or to no tag.
- B Construction Detail: Direct-assessment scores are available for 111,545 of 126,754 instances, averaging 3.5 annotators per instance, with recomputable agreement statistics.Agreement ranges from α = 0.955 for new en-ml to α = 0.080 for en-mr, whose mean reliability is 0.247 across 30,746 instances.
- B Construction Detail: en-te has near-ceiling agreement yet is lowest-scoring for seven of nine systems, while its challenge-test DA standard deviation is 14.0 versus 25.1 for en-ml.These results indicate that agreement does not predict difficulty and that score range is a more likely explanation.
- C QE Meta-Evaluation: The QE meta-evaluation presents each difficulty axis against both controls, reports per-pair four-shot compliance, and provides per-pair correlations and cross-lingual fits.These results are organized in Tables 11 and 12, Appendix F, and Table 5.
D The Axes Pair by Pair … J Duplication and Split-Hygiene Checks
Pairwise analyses show that only A4 has a consistent negative matched contrast, while thresholds, direction pooling, and partial coverage materially affect interpretation. The release also supplies per-pair diagnostics, benchmark distributions, and verification checks for composition and duplication hygiene.
- D The Axes Pair by Pair: A4 is negative on all seven pairs, with −0.598 on en-ml versus −0.061 to −0.088 on three en→Indic pairs; removing en-ml changes the mean from −0.146 to −0.071.The paper emphasizes sign consistency rather than magnitude because one small cell dominates the aggregate.
- D The Axes Pair by Pair: A1 has no explained aggregate effect: its mean is zero because pairwise contrasts range from −0.12 on en-gu and −0.11 on en-ml to +0.09 on ne-en.Pairwise splitting reveals cancellation rather than a robust shared effect.
- D The Axes Pair by Pair: A4’s barely-edited arm covers 1,149 segments across en-hi, en-mr, en-ta, and et-en, with a matched contrast of −0.070 on every pair.The converse arm covers 123 segments across en-ml, ne-en, and si-en and is negative on two of three pairs, with en-ml doing all the work.
- E Axis Thresholds by Pair: Pair-specific thresholds matter: en-hi enters its bottom 30% at direct assessment 79.2, whereas ne-en enters at 25.0; en-ml’s A4 lower-arm cut is unreachable at −0.08.On en-ml, 44.4% of post-edits leave MT untouched, producing mean TER 0.204 and standard deviation 0.563.
- F Per-Pair QE Decomposition: Telugu has the lowest column mean, 0.119 versus 0.274–0.364 for other en→Indic pairs, and is lowest for seven of nine systems; X→en scores higher on average.BharatGPT-3B is lowest on si-en at −0.091, while Llama-3.2-3B is lowest on en-mr by 0.006.
- G Direction Scope in Cross-Lingual Correlation: Pooling directions inflates cross-lingual coefficients: aya-expanse-32B falls from +0.92 pooled to −0.22 scoped, while only sarvam-m and XCOMET-XL exclude zero when scoped.Six of ten systems differ by more than 0.4 between pooled and scoped estimates, and four differ by more than 0.85.
- H Direct Assessment on the Full Population: The full-population pass covers 81,315 of 111,974 DA-bearing rows, with all three metrics evaluated on identical rows; MetricX-24’s mean error is 3.0–3.7 on en→Indic versus 6.9–10.3 on X→en.Zero-edit rows are included, but the pass is not a measurement over the entire release.
- I Benchmark Distributions / J Duplication and Split-Hygiene Checks: The challenge test’s axis examples and domain distribution are documented, while verification scripts rederive per-pair and per-label counts and audit exact task views plus MQM’s 2,327 DA-paired and 2,163 MQM-only rows.The MQM-only rows sit outside the 126,754-instance benchmark, motivating separate checking.
K Layer-Probing and COMET-Regression: Protocol and Full Results … N APE under Word-Level Metrics
The paper combines layer probing, COMET regression, new English–Malayalam annotation, strict evaluation settings, and APE comparisons. Results show instability and difficulty-axis effects in trained metrics, while annotation reliability and word-level scoring do not straightforwardly predict task difficulty or system quality.
- K Layer-Probing and COMET-Regression: Protocol and Full Results: Six layers are swept for a 4-bit QLoRA layer probe trained on DA and evaluated on the matched challenge split.The swept layers are {−5, −7, −11, −16, −20, −24}.
- K Layer-Probing and COMET-Regression: Protocol and Full Results: COMET regression trains jointly on 64,416 segments from five en→Indic pairs and evaluates on an 8,160-segment challenge test.The hybrid head uses upper encoder layers, optional TSA/TSE and length-ratio features, and per-pair z-scored DA under Huber loss.
- K Layer-Probing and COMET-Regression: Protocol and Full Results: 0.006 is the pairwise accuracy of constant-score runs, matching the human gold tie rate, while Spearman cannot distinguish them from weakly ordering runs.Weak runs reach 0.52 to 0.53 pairwise accuracy, whereas Spearman places both kinds in the −0.001 to 0.097 band.
- K Layer-Probing and COMET-Regression: Protocol and Full Results: −0.095 to −0.129 is the A4 contrast across COMET-regression ablations, and every configuration is worse than its DA-matched control.The strongest overall-correlation min–max target still reaches −0.099, across all three pairs carrying A4.
- L The English–Malayalam Annotation: 10,000 English–Malayalam segments are released, with annotation produced under a revised-guidelines and iterative quality-control protocol.The protocol began with a 500-segment pilot, followed by main batches of 3,750 segments and weekly validation meetings.
- L The English–Malayalam Annotation: α = 0.955 and ICC(1, k) = 0.985 make en-ml the benchmark’s highest-agreement pair, yet it remains difficult for several evaluated systems.The pooled comparison is α = 0.690 and ICC(1, k) = 0.881; en-ml is second-lowest of nine for tiny-aya and CometKiwi-XL.
- M Experimental Settings: Greedy decoding uses temperature 0.0, top-p 1.0, at most 512 new tokens, and seed 0 across prompted models and templates.Few-shot examples come from the same language pair’s training split, while strict parsing rejects outputs containing formatting ambiguities rather than extracting the nearest number.
- N APE under Word-Level Metrics: 17.8 TER points is the unedited MT’s lead over the best system on en-ta under word-level scoring, versus 11.1 under char-TER.The unedited MT wins on en-mr, en-ta and en-ml under all four surface metrics and loses on en-hi; the difference reflects Tamil inflection tokenization.
O English–Hindi MQM Layer · P Axis Difficulty against Label Reliability
The English–Hindi MQM layer shows that pooled metric correlations can collapse because systems misorder annotation batches, not merely languages, while MQM remains complementary to direct assessment. Reliability analysis indicates that label noise does not explain the observed P-axis contrast.
- O English–Hindi MQM Layer: MQM annotations hold language fixed while varying annotation batch, enabling a test of whether pooled-correlation failures reflect group offsets rather than language effects.The layer combines EAMT re-annotations with the new WMT24 general-MT batch.
- O English–Hindi MQM Layer: 0.116 pooled Spearman for CometKiwi-XL falls below its 0.276 and 0.219 batch correlations, while XCOMET-XL similarly falls from 0.243 and 0.052 to 0.029.Both XL metrics score the lower-quality WMT24 batch above the higher-quality EAMT batch, reversing the human ordering.
- O English–Hindi MQM Layer: 0.330 pooled Spearman for CometKiwi-DA remains consistent with its 0.307 and 0.375 batch correlations, unlike the two XL metrics.The passage identifies CometKiwi-DA as correctly ordering the two batches and as well behaved across languages.
- O English–Hindi MQM Layer: The batch effect is a per-group offset: the same systems misplace whole segment groups against human quality whether groups are languages or annotation batches.Range restriction is not the proposed explanation; the near-ceiling EAMT batch has gold standard deviation 1.96, versus 7.54 for WMT24.
- O English–Hindi MQM Layer: Pearson 0.573, Spearman 0.518, and Kendall 0.366 show moderate agreement between DA and MQM on 2,327 paired segments, making MQM complementary rather than interchangeable.On 476 test segments, XCOMET-XL agrees more with MQM, whereas both CometKiwi variants agree more with DA.
- O English–Hindi MQM Layer: 549 of 2,327 scored segments have localized spans, while 1,618 sub-perfect segments lack recorded offsets, so an empty span list has two meanings.The release exposes mqm_spans_complete: true for the 709 segments whose span completeness is known.
- P Axis Difficulty against Label Reliability: 0.534 reliability implies a 0.731 attainable-correlation ceiling, yet systems reach only 0.234 on A4, so target reliability does not explain its contrast.Reliability is near 0.75 for A2 and A3, while A1 has negative coefficients because within-segment disagreement exceeds between-segment spread.
- P Axis Difficulty against Label Reliability: −0.009 is the maximum correction to the P-axis contrast after matching reliability ceilings on the three X→en pairs, and the contrast remains negative.On four en→Indic pairs, ICC(1, k) is unestimable for the flagged and matched-control sets because the compressed score range is too narrow.
Q Pair-Level Factors
Among the tested pair-level explanations, label reliability and training exposure do not account for QE difficulty or the X→en advantage. Target-score spread is the only factor with a weak association with mean correlation.
- Label reliability: −0.16 (p = 0.67) is the across-pair correlation between ICC(1, k) and per-pair Spearman, with no individual system significant.Thus, less reliable target labels do not consistently make pairs harder for every system.
- Label reliability: 0.976 is en-ml’s ICC(1, k), yet it ranks second lowest of nine pairs on CometKiwi-XL.The pair has the benchmark’s most reliable labels but does not correspondingly achieve high system correlation.
- Training exposure: +0.18, +0.17 and +0.27 are the X→en gaps for CometKiwi-DA, CometKiwi-XL and XCOMET-XL, versus +0.14, +0.15, +0.26 and +0.06 for GPT-5.5, aya-expanse-32B, tiny-aya and sarvam-m.The similar gaps in systems without MLQE-PE label exposure show that exposure cannot fully explain the advantage.
- Target spread: +0.53 (p = 0.14) is the correlation between pair-level DA standard deviation and mean correlation, compared with +0.41 using the interquartile range.Target spread is the only tested factor that survives, and the association is weak.