Source-linked AI summary
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation
Samuel Läubli, Rico Sennrich, Martin Volk
TL;DR
Claims that neural machine translation reaches professional human parity may depend on evaluating isolated sentences rather than whole documents. The paper compares sentence- and document-level pairwise judgments by professional translators and finds stronger human preference with document context, motivating document-level evaluation.
Problem
Recent research claims parity between neural machine translation and professional human translation, but the paper questions whether sentence-level evaluation adequately detects quality differences.
Method
The authors conduct a 2 × 2 evaluation with professional translators, comparing pairwise adequacy and fluency judgments for isolated sentences and entire documents.
Results
Document-level evaluation yields a stronger preference for HUMAN than sentence-level evaluation, with significant human preference for document-level adequacy and for fluency at both levels.
Takeaways & Limitations
As machine translation improves, document-level evaluation can expose discourse-related errors that remain hard or impossible to detect in isolated sentences.
Takeaways & Limitations
Direct human–machine comparison is hard to interpret when the English target is the original document rather than a human translation into English.
Abstract
from arXiv · showhide
Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese--English news translation task. We empirically test this claim with alternative evaluation protocols, contrasting the evaluation of single sentences and entire documents. In a pairwise ranking experiment, human raters assessing adequacy and fluency show a stronger preference for human over machine translation when evaluating documents as compared to isolated sentences. Our findings emphasise the need to shift towards document-level evaluation as machine translation improves to the degree that errors which are hard or impossible to spot at the sentence-level become decisive in discriminating quality of different translation outputs.
1 Introduction
Neural machine translation has prompted claims of parity with professional human translation, but the authors argue that such claims warrant further scrutiny. They investigate whether sentence-level evaluation may miss quality differences that document-level context reveals.
- Recent studies report neural machine translation approaching or reaching professional human translation quality on some test sets.
- Statistical ties between machine systems are less consequential than ties between machine translation and professional human translation.
- The study independently evaluates the professional and machine translations judged equal by Hassan et al. (2018).
- The authors focus on whether missing document-level context explains differences in evaluation outcomes.
- Their hypothesis is that professional human translation remains superior, while raters’ sensitivity to quality differences depends on evaluation protocol.
2 Human Evaluation of Machine Translation
The paper evaluates human and machine translations with professional translators using pairwise ranking across sentence and document units, with separate adequacy and fluency conditions. It applies this protocol to Chinese–English WMT 2017 news translations.
- The experiment uses a 2 × 2 mixed factorial design crossing source-text availability with sentence versus document evaluation.
- Granularity of Measurement: Raters pair a professional translation with machine translation and choose the better output, with ties allowed.
- Granularity of Measurement: Adequacy judgments show source texts and translations, whereas fluency judgments show translations alone.
- Raters: Professional translators with at least three years of experience and positive client reviews serve as raters.
- The evaluation samples 55 documents and 2×120 sentences from the native-Chinese portion of the WMT 2017 test set.
- The recruited translators averaged 13.7 years of experience and 8.8 positive ProZ client reviews.
3 Results
Document-level evaluation produces stronger evidence favoring human translation than isolated-sentence evaluation, especially for adequacy. Fluency favors human translation at both granularities, while document context increases ties and reduces machine preference.
- Adequacy: Adequacy showed no significant HUMAN–MT difference for sentences (x=86, n=189, p=.244), but favored HUMAN for documents (x=104, n=178, p<.05).
- Adequacy: Adequacy preference for MT fell from 50% with sentences to 37% with documents, while ties remained similar.
- Fluency: Fluency significantly favored HUMAN for sentences (x = 106, n=172, p<.01) and documents (x=99, n=143, p< .001).
- Fluency: In fluency evaluation, preference for MT fell from 32% to 22% with document-level context, alongside more ties.
- Inter-annotator agreement ranged from Cohen’s κ=0.13 to 0.32 despite statistically significant results.
4 Discussion
The evaluation protocol strongly affects whether human and machine translation quality differences are detected. Document-level assessment reveals distinctions that sentence-level judgments often miss, including cohesion, coherence, and ambiguous-word errors.
- Sentence-level adequacy judgments by professional translators did not show a statistically significant preference for HUMAN or MT.This held despite using pairwise ranking rather than direct assessment.
- Document-level adequacy judgments showed a statistically significant preference for HUMAN over MT.Individual raters also tended to rate HUMAN more favorably on documents than on isolated sentences.
- Document-level evaluation exposed lexical-coherence errors, including three inconsistent MT translations of the same app name within one article.HUMAN consistently translated the name as “WeChat Move the Car,” while MT used “Twitter Move Car,” “WeChat mobile,” and “WeChat Move.”
- Document-level context may reveal mistranslations of ambiguous words and errors in textual cohesion and coherence that are hard or impossible to spot sentence by sentence.The authors also observed more appropriate discourse connectives in HUMAN, leaving detailed investigation for future work.
- Fluency raters showed a stronger preference for HUMAN than adequacy raters, contrary to expectations based on NMT’s previously reported fluency strength.The authors caution that L1 interference in bilingual ratings may favor MT’s more literal translations, while document context still strongly affects fluency judgments.
5 Conclusions
The study finds that document-level context produces a markedly stronger preference for human translations than isolated-sentence evaluation. It therefore questions whether current evaluation practices can detect discourse-level quality differences, while identifying important scope boundaries for future validation.
- Raters showed a markedly stronger preference for human translations when evaluating documents rather than isolated sentences.
- The authors argue that improving MT quality may make sentence-level distinctions harder to detect, increasing the value of document-level evaluation.Documents provide context for understanding source and translation and expose discourse-related errors invisible at sentence level.
- Whether accurate document-level judgments can also be elicited through crowdsourcing and direct assessment remains an open practical question.The authors suggest testing alternative protocols using released data as a test bed.
- The tested MT system operates at the sentence level and ignores wider context, which is one reason document-level evaluation may widen the human–machine quality gap.
- Narrowing this gap may require document-level training data, appropriate models, and discourse-aware automatic metrics.
A Further Statistical Analysis
The analysis examines rating agreement, quality control, and aggregation across the four experimental conditions. Despite low agreement in most conditions, majority voting produces clearer discrimination between MT and HUMAN except for sentence-level adequacy.
- Raters selected whether MT, HUMAN, or a tie was better for each item.Table 1 reports results for all four experimental conditions and individual raters.
- Cohen’s kappa was calculated from pairwise ratings using observed agreement and chance agreement.P(A) denotes the proportion of rater agreements, while P(E) denotes agreement expected by chance.
- Lower inter-rater agreement occurred in three of four conditions, partly because high-quality MT was difficult to distinguish from professional translation at sentence level.Differing interpretations of translation quality and tie assignments also contributed to lower agreement.
- 1 of 40 document-level and 5 of 128 sentence-level spam items were missed, and every miss was labeled as a tie.Five of nine raters missed no spam items.
- Majority voting produced clearer discrimination between MT and HUMAN in every condition except sentence-level adequacy.The aggregation results are reported as percentages in Table 2.