Source-linked AI summary

Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation

Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, Wolfgang Macherey

arXiv:2104.14478v1cs.CLcs.AIcs.LG

TL;DR

Human evaluation of high-quality MT lacks a commonly accepted standard and can produce unreliable conclusions. This paper applies professional, context-aware MQM error analysis to WMT 2020 outputs, finding changed rankings, a clear human-over-machine gap, and stronger performance from embedding-based metrics than crowd workers.

  • Problem

    High-quality MT is difficult to evaluate reliably because automatic metrics are limited and human procedures lack a commonly accepted standard.

  • Method

    The study uses MQM-based explicit error analysis with professional translators who evaluate WMT 2020 systems in two language pairs using full document context.

  • Results

    MQM produces substantially different system rankings from WMT crowd-worker evaluation, clearly separates human from machine translations, and shows embedding-based metrics can outperform crowd workers.

  • Takeaways & Limitations

    MQM-based professional evaluation questions prior crowd-worker conclusions about high-quality MT and provides a publicly released corpus for further research.

  • Takeaways & Limitations

    Automatic-metric comparisons use Kendall correlation rather than the official WMT Pearson correlation because the study evaluates relatively few systems.

Abstract

from arXiv · show

Human evaluation of modern high-quality machine translation systems is a difficult problem, and there is increasing evidence that inadequate evaluation procedures can lead to erroneous conclusions. While there has been considerable research on human evaluation, the field still lacks a commonly-accepted standard procedure. As a step toward this goal, we propose an evaluation methodology grounded in explicit error analysis, based on the Multidimensional Quality Metrics (MQM) framework. We carry out the largest MQM research study to date, scoring the outputs of top systems from the WMT 2020 shared task in two language pairs using annotations provided by professional translators with access to full document context. We analyze the resulting data extensively, finding among other results a substantially different ranking of evaluated systems from the one established by the WMT crowd workers, exhibiting a clear preference for human over machine output. Surprisingly, we also find that automatic metrics based on pre-trained embeddings can outperform human crowd workers. We make our corpus publicly available for further research.

1 Introduction

The paper argues that high-quality MT requires more reliable human evaluation and proposes explicit MQM error analysis as a foundation for comparison. Applying this approach to WMT 2020 reveals substantially different system rankings, a clear human-over-machine preference, and weaknesses in crowd-worker evaluation.

  • Motivation: High-quality MT is difficult to evaluate because correct translations are numerous and unknown, limiting automatic metrics and requiring costly human assessment.Human evaluation also faces unresolved questions about quantifying small inaccuracies and rater agreement.
  • Motivation: Current evaluation often uses isolated sentences, inexperienced raters, and single scores, risking noise and misleading human-parity claims as MT quality improves.The paper identifies inadequate practices as a source of erroneous conclusions about MT quality.
  • Method: MQM makes the implicit identification of translation errors explicit, producing a flexible error-based standard from which task-specific scores can be derived.The framework provides a hierarchy of errors that can be tailored to specific applications.
  • Method: The study annotates outputs from 10 top-performing systems, including human references, in English→German and Chinese→English using professional translators with full document context.Scalar ratings from professionals and crowd workers were also collected for comparison.
  • Results: MQM sharply revises the original WMT ranking, clearly preferring human translations over machine translations and promoting some low-ranked MT systems.The corpus includes over 100k human-translation and high-quality-MT segments and is publicly released.
  • Results: Crowd-worker evaluation correlates poorly with MQM, while embedding-based automatic metrics can outperform human crowd workers.The findings question conclusions based on previous crowdsourced evaluations and add evidence against human-parity claims.

2 Related Work

Prior MT evaluation research spans adequacy, fluency, rankings, professional assessment, comprehension-based tasks, and error marking. This literature also documents weaknesses in crowd-worker evaluations and motivates adaptable error-based frameworks such as MQM.

  • Historical evaluation methods: Early MT evaluation used intelligibility, fidelity, adequacy, fluency, comprehension, and later ranking-based approaches as core quality measures.These methods evolved from the ALPAC report and ARPA initiative through the first WMT campaigns.
  • Professional and crowd evaluation: Professional translators have been used for ranking, error classification, and post-editing, while studies find crowd workers more accepting of subtle translation errors.This work established differences between evaluator populations and task formats.
  • Professional and crowd evaluation: Multiple studies report that crowd-worker evaluations require filtering or fail to distinguish human from machine translations as reliably as professional judgments.Professional re-evaluations have also produced changed rankings better aligned with other evidence.
  • Alternative evaluation methods: Alternative manual methods include reading comprehension, gap-filling with machine-translated hints, and marking actual issues instead of assigning a single score.These approaches broaden evaluation beyond adequacy and fluency ratings.
  • MQM: MQM was developed as an adaptable methodology addressing shortcomings of previous quality-evaluation methods and has supported tailored error taxonomies and automatic-metric training.Examples include a Slavic-language taxonomy and MQM labels used to fine-tune COMET.

3 Human Evaluation Methodologies

The study compares WMT baseline ratings, 7-point SQM ratings, and MQM error-based evaluations, using document context and different annotator pools. MQM assigns weighted error scores to translation segments and aggregates them for document- and system-level comparison.

  • WMT baseline: WMT collects segment-level ratings with document context on a 0–100 scale using source- or reference-based evaluation and rater quality controls.The evaluation uses researchers or translators for translations out of English and crowd workers for translations into English.
  • Scalar Quality Metric: SQM presents source and translation segments in document rows and asks raters to assign a 0–6 quality rating, including intermediate levels.Raters can scroll through the other segments in the document while evaluating each row.
  • MQM: MQM annotators identify and label segment errors by category and severity, while limiting each segment to its five most severe errors.The hierarchy includes Accuracy, Fluency, Terminology, Style, Locale, and a Non-translation category for severely garbled segments.
  • MQM: MQM scoring weights Minor errors at 1, explores Major weights from 1 to 10, and selects a Major weight of 5 for ranking stability and discrimination.Minor Fluency/Punctuation errors receive weight 0.1, Non-translation receives weight 25, and segment scores range from 0 to 25 before averaging over annotators.
  • Experimental setup: The experiments re-annotate 1,418 English→German segments and 2,000 Chinese→English segments across 10 systems per language pair, using professional and crowd-worker evaluations.Professional SQM and MQM used disjoint translator pools; all professional annotators were native speakers of the target language, and documents were assigned across rater sets.

4 Results

MQM-based professional evaluation substantially changes the ranking of WMT 2020 translations, clearly favoring human output over machine output. Results also characterize error patterns, annotator agreement, rating allocation, and metric correlations.

  • Overall System Rankings: All human translations ranked first under both pSQM and MQM for both language pairs, with MQM placing them ahead by a larger margin.Professional translators also ranked paraphrased Human-P above every MT output, unlike crowd workers.
  • Overall System Rankings: WMT human ratings showed low system-level correlation with MQM for English→German and negative correlation for Chinese→English, with very low segment-level correlation in both.The results question the reliability of WMT human ratings for these systems.
  • Error Analysis: Accuracy-based or major-error-only MQM scores produced nearly the same system ranking as full MQM, indicating that major accuracy errors drive most quality differences.Most major errors are accuracy errors, while fluency errors are generally judged minor.
  • Error Analysis: Accuracy/mistranslation errors constituted the majority of major errors, and Chinese→English contained more absolute errors than English→German.Human translations had fewer errors across nearly all categories, while systems recorded 4x more error points for English→German accuracy/mistranslation categories.
  • Document-error Distribution: MT systems shared difficulties on the same subset of documents, showing that output quality depended strongly on the input sentences as well as the system.Some documents brought MT close to human performance, while others showed a clear gap.
  • Rating Allocation: With 900 segment-level ratings, system scores became more accurate when ratings covered three consecutive sentences per document across more documents.Comparisons improved further when corresponding systems used the same documents, sentences, and, when possible, raters.
  • Automatic Metrics: MQM correlations exceeded WMT correlations at system and segment levels, although segment-level correlations were generally much lower than system-level correlations.Adding human translations caused a large metric-performance drop, especially for MQM, because MQM rated human outputs unambiguously higher than MT.

5 Conclusion

The study proposes MQM-based professional evaluation as a platinum standard and releases its ratings for further research. It reports that crowd-worker rankings diverge from MQM-based rankings, while MQM reveals a substantial human–machine quality gap.

  • The study uses professional-translator MQM ratings from WMT 2020 as a platinum standard for comparing simpler human and crowd-worker evaluation methods.
  • The released ratings support further research on human and automatic machine-translation evaluation.
  • MQM ratings sharply revise WMT crowd-worker rankings and show a clear preference for human over machine translations.
  • Embedding-based automatic metrics outperform crowd-worker human evaluation in the study.

A MQM Summary

MQM provides an adaptable framework for identifying translation issues and aggregating their severities into evaluation scores. Its research guidance covers hierarchy design, expert annotation, context, and analysis choices.

  • MQM defines a generic, adaptable methodology centered on a hierarchy of translation issues that annotators identify at a suitable granularity.
  • The MQM standard combines an issue vocabulary, scoring mechanism, XML metric formalism, selection guidelines, and mappings from legacy metrics.
  • Researchers should tailor the issue hierarchy and text-unit granularity to their questions, adding needed issues and pruning irrelevant ones.
  • MQM guidance recommends expert translators, three annotators per text item, training, guidelines, and calibration when possible.
  • Annotations may be conducted in short segments with time adjusted for difficulty, at an estimated cost of approximately 1 USD per segment.
  • Analysis can aggregate issue scores using severity levels none, minor, major, and critical, with recommended severity weights of 0, 1, 10, and 100.

Annotation

The study adapts MQM for broad-coverage MT by annotating segment-level errors with document context and a modified hierarchy. It simplifies severity handling and relies on experienced professional raters with concise instructions.

  • Annotation: The broad-coverage hierarchy is applied at segment level with document context and was modified with expert translators, including added Locale sub-categories and Non-translation.
  • Annotation: The scheme drops Critical severity because its context-specific distinction from Major was considered highly subjective for broad-coverage MT.
  • Annotation: Locale sub-categories were introduced after piloting, although collapsing them was considered an alternative and arguably preferable strategy.
  • Annotation: The annotation design includes the MQM Core issue hierarchy as its conceptual basis for organizing translation issues.
  • Annotation: Instructions remained minimal because professional raters already had translation-quality and MQM experience, leaving subtle contextual decisions to their judgment.

Scoring

The scoring scheme converts annotated error categories and severities into segment, document, and system scores. It selects a Major-error weight of 5 for ranking stability and discrimination while downweighting minor fluency and punctuation issues.

  • Scoring: Minor errors receive weight 1, while Major-error weights from 1 to 10 were evaluated using resampling-based ranking stability analysis.
  • Scoring: A Major-error weight of 5 provided the best balance between ranking stability and system discrimination.
  • Scoring: Minor Fluency/Punctuation errors receive weight 0.1 because they mainly affect appearance, are easy to correct algorithmically, and do not change meaning.
  • Scoring: The scoring model ignores error-span length because severity and category are judged to provide the relevant scoring information.
  • Scoring: A segment score sums its errors, averages across annotators, and ranges from 0 for perfect to 25 for maximally bad; document and system scores average segment scores.
  • Scoring: The weighting scheme is summarized in Table 1, while the supplied tables describe MQM hierarchy and severity levels.

C Analysis of Metric Performance

The paper compares metric correlations with MQM and WMT scoring at system and segment levels under standard, paraphrased, and human-inclusive evaluation settings. Correlations are generally lower at the segment level, while differences between MQM- and WMT-based correlations persist.

  • System-level correlations: System-level analyses report Pearson and Kendall correlations for WMT 2020 metrics against MQM and WMT scoring.Figures 11 and 12 provide the corresponding Pearson and Kendall plots.
  • System-level correlations: Paraphrased references substantially change metric ranking and performance for English→German under Kendall correlation.This comparison is shown in Figure 13.
  • Human-inclusive evaluation: Including human outputs produces lower correlations than MQM gold scores and much lower correlations than WMT gold scores.The comparison is shown in Figure 17 and its corresponding human-inclusive analysis.
  • Segment-level correlations: Segment-level correlations use a WMT Kendall-like measure that discards rankings with missing annotations or raw-score differences below 25.MQM correlations use a threshold of 0 because no significance procedure was available.
  • Segment-level correlations: Segment-level correlations are much lower than system-level correlations, but differences between WMT and MQM correlations remain similar.Figures 15, 16, and 17 cover standard references, paraphrased references, and human outputs, respectively.
Loading 2104.14478v1…