Source-linked AI summary

Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus

Lifeng Han, Jiahui Liang, Anna Latusek, Karim El Haff, Amal Haddad Haddad, Josua Höfgen, Kilian Evang, Min Ma, Maryia Zhyrko

arXiv:2609.06634v1cs.CLcs.AI

TL;DR

MT systems can achieve strong aggregate scores while still mishandling MWEs and other linguistically challenging phenomena. Using AlphaMWE in the WMT 2026 Test Suites shared task, the study combines three automatic metrics with HOPE-based human evaluation across multiple language pairs. It finds persistent figurative and MWE-related errors, disagreement among metrics, and language-specific problems revealed by human assessment.

  • Problem

    Limited evidence shows how diverse contemporary MT systems perform on a shared multilingual benchmark targeting MWEs and other challenging constructions, while aggregate metrics may overlook important errors.

  • Method

    The study evaluates 31 MT systems across six language variants using AlphaMWE, three automatic metrics, mean-rank Top3 selection, and HOPE-based human evaluation.

  • Results

    MWEs, idioms, metaphors, lexical ambiguity, and language-specific collocations remain error sources; metrics sometimes disagree, and human evaluation reveals language-specific issues hidden by aggregate scores.

  • Takeaways & Limitations

    Combining complementary automatic metrics with targeted human evaluation supports investigation of linguistically challenging translation phenomena.

  • Takeaways & Limitations

    Human-evaluation coverage and agreement were limited by uneven annotator participation, only 150 common cross-language segments, differing system participation, sentence-level context, and three automatic metrics.

Abstract

from arXiv · show

LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they are trained upon. To examine if Multiword Expressions (MWEs) still set a bottleneck for LLMs regarding language understanding and translation, we report the system performances from the WMT2026 Test Suites shared task, for which we used the publicly available multilingual parallel corpus AlphaMWE as the test suites. We received 31 MT systems' outputs covering English to Chinese (zh), Polish (pl), German (de), Arabic (ar) including Modern Standard Arabic (MSA) and two dialectal ones (Egyptian and Tunisian Arabic). We carried out automatic evaluations using BLEU, ChrF, BERT-score to select the Top3 systems per language pair, followed up with human evaluations on the selected systems. Our findings show that: figurative/MWE phenomena remain challenging; automatic metrics sometimes disagree; human evaluation uncovers language-specific errors hidden by aggregate scores.

1 Introduction

The paper examines whether MWEs and related non-compositional phenomena remain blind spots for contemporary MT systems. Using AlphaMWE in the WMT 2026 shared task, it combines automatic metrics with human evaluation across languages to expose errors that aggregate scores can miss.

  • Motivation: High overall translation quality does not ensure reliable handling of MWEs, idioms, metaphors, or other non-compositional constructions.These phenomena can require contextual interpretation beyond individual-word composition.
  • Research gap: The study addresses limited evidence from shared multilingual benchmarks by evaluating diverse contemporary MT systems on challenging constructions.Previous work often used controlled settings or focused on individual phenomena.
  • Benchmark and scope: AlphaMWE covers English–Chinese, English–German, English–Polish, and three Arabic variants across technical, news, and literary genres.The shared task attracted 31 MT systems.
  • Method: The evaluation applies SacreBLEU, chrF++, and BERTScore, then uses HOPE-based human evaluation on top-ranked systems for each language pair.This pipeline compares automatic and human assessment across languages and linguistic phenomena.
  • Findings: Idioms, verbal MWEs, metaphors, lexical ambiguity, and language-specific collocations remain important error sources despite strong automatic scores.Errors are especially evident in literary and expression-rich texts, while technical texts are translated more reliably.
  • Findings: Automatic metrics sometimes disagree on rankings, while human evaluation reveals language-specific issues largely invisible to aggregate scores.The findings motivate combining complementary automatic metrics with targeted human assessment.

2 WMT2026 Test Suites: AlphaMWE

AlphaMWE is used as a multilingual diagnostic test suite for evaluating how MT systems handle complex multiword and verbal expressions. The WMT 2026 pipeline constructs compatible suites, applies three automatic metrics for system ranking, and conducts HOPE-based human evaluation.

  • Pipeline: The shared-task methodology has three stages: test-suite construction, automatic evaluation, and human evaluation with analysis.These stages structure the full diagnostic workflow.
  • AlphaMWE: AlphaMWE is a multilingual parallel corpus with MWE annotations, machine-translation post-editing, and human-in-the-loop quality controls.It builds on sources including the PARSEME shared task.
  • AlphaMWE: The corpus originally covered English–German, English–Polish, and English–Chinese, later extending to Arabic.The Arabic edition was added subsequently.
  • Corpus coverage: The original corpus contained 747 segments, with 150 segments for each Arabic variant and 597 for English–Polish because annotation was incomplete.The remaining files contained around 150 segments each.
  • Automatic evaluation: SacreBLEU, chrF++, and BERTScore provide n-gram, character-level, and contextual-semantic assessments.The metrics are used to characterize different aspects of translation output.
  • System selection: Systems are ranked per language pair and selected using average rank because the three metrics have different numerical scales.The rankings are averaged across the tested metrics rather than their raw scores.
  • Human evaluation: HOPE evaluates human-centered post-editing quality through eight error types and severity levels from 0 to 16.The error types include terminology, mistranslation, style, proofreading, and proper-name errors.

3 Systems Received

The paper analyzes outputs submitted by WMT 2026 Test Suites participants within a methodology that proceeds from AlphaMWE construction through automatic ranking to human and qualitative analysis.

  • System coverage: WMT 2026 participants submitted MT outputs for the shared-task evaluation.The number of participating systems varied by language pair because not all participants submitted every direction.
  • Methodology overview: Figure 1 presents the workflow from AlphaMWE test-suite construction to automatic evaluation, Top3 selection, HOPE-based human evaluation, and subsequent analyses.The figure summarizes the study methodology rather than a single system result.

4 Automatic Evaluation

Automatic evaluation ranks the Top3 systems for each language pair by aggregating rankings from three metrics. Rank standard deviation shows where those metrics agree or disagree, including a tied Chinese ranking with different agreement levels.

  • Selection procedure: Top3 systems are selected using the mean of the three automatic metric rankings for each language pair.Rank standard deviation is reported as an agreement indicator, not as a selection criterion.
  • Metric agreement: Half of the selected systems had Rank-SD 0, while the other half showed differing rankings across metrics.This indicates that metric agreement was not uniform across language pairs.
  • Metric agreement: Rank-SD 0 means all three metrics ranked a system identically, whereas larger values indicate stronger metric disagreement.Values above 2 indicate substantial disagreement.
  • Chinese ranking: For en→zh, SCIR-TG-MT and GoogleTranslate tied at aggregate rank 1, but SCIR-TG-MT had lower SD than GoogleTranslate.Their mean rank was 1.667; SD was 0.471 for SCIR-TG-MT and 0.943 for GoogleTranslate.

5 Human Evaluation

Human evaluation shows that translation quality differences are driven primarily by semantic, idiomatic, lexical, and stylistic errors rather than surface grammaticality. Across language pairs and genres, literary and expression-rich texts remain harder than technical texts, while language-specific analysis exposes issues that aggregate scores can miss.

  • English–Arabic: Arabic system severity scores were 57 for TRIVE, 84 for Google Translate, and 284 for Gemma 4-4B across 150 segments.TRIVE had errors in fewer sentences, whereas Gemma 4-4B showed the weakest performance.
  • English–Arabic: Across Arabic systems, mistranslation was the largest error category, while genuine grammatical violations were comparatively rare.Arabic evaluation also distinguished true grammar errors from stylistic or collocational unnaturalness.
  • English–Arabic: Arabic systems frequently struggled with idioms, lexical ambiguity, terminology, transliteration, calques, and source-convention adaptation.TRIVE correctly conveyed the idiomatic meaning of “this broke everybody up,” while other systems translated it literally.
  • English–Chinese: For English–Chinese, SCIR-TG-MT had the lowest total penalty at 52, compared with 98 for GoogleMT and 167 for Tower-9B.SCIR-TG-MT also handled the idiom “lifting a little finger” correctly, while the other systems translated it literally.
  • Cross-language observations: Literary and expression-rich texts were more difficult than technical texts, with idiomatic and metaphorical language creating persistent translation gaps.English–Polish technical translations benefited from established terminology, whereas literary texts remained considerably more challenging.
  • English–Polish: English–Polish errors commonly involved style, literal MWE translation, verbal aspect, grammatical gender, and collocations.The evaluation also noted that sentence-level judgments can make multiple translations appear acceptable without wider document context.

6 Conclusions and Future Work

The study finds that linguistically challenging phenomena remain difficult for MT systems despite strong automatic performance, while human evaluation reveals language-specific errors and metric disagreement. Future work targets context-aware evaluation, phenomenon-specific metrics, and broader human annotation.

  • Conclusions: High automatic performance does not eliminate errors involving idioms, verbal MWEs, figurative expressions, lexical ambiguity, and language-specific collocations.These weaknesses are especially visible in literary and expression-rich text, while technical and news-oriented sentences are generally handled more reliably.
  • Conclusions: Human evaluation exposes language-specific translation issues that aggregate automatic scores can miss.English–Chinese and English–Arabic prominently show literal interpretations of idiomatic and figurative expressions, while English–Polish shows difficulties in other linguistic phenomena.
  • Conclusions: No single automatic metric provides a complete account of translation quality.The metrics often agree on the strongest systems, but non-zero rank standard deviations reveal disagreement, while human annotation identifies errors involving semantic scope, idiomaticity, terminology, and context.
  • Limitations: Sentence-level evaluation can penalize legitimate alternatives when English source sentences are underspecified without wider discourse context.This affects distinctions such as gender, verbal aspect, number, and formal or informal forms of address in target languages.
  • Future Work: Future work will extend evaluation to context-aware and document-level translation, phenomenon-specific frameworks, and larger-scale multilingual human annotation.These directions aim to improve analysis of figurative expressions and inter-annotator agreement.

Limitations

The study reports limitations in annotation coverage, cross-language comparability, sentence context, and metric breadth.

  • Annotation: The human evaluation has uneven annotator coverage and incomplete inter-annotator agreement reporting across languages.Arabic received one annotation round, Chinese used a double-checking procedure, Polish annotation remained incomplete, and German agreement was still to be reported.
  • Evaluation Scope: Only 150 common segments were used for cross-language human comparison.
  • Evaluation Scope: Different systems participated in different language pairs, limiting direct cross-language system comparisons.
  • Evaluation Scope: Sentence-level context limitations can make multiple target translations equally valid.Reference-based evaluation may penalize legitimate alternatives when English does not encode distinctions required by the target language.
  • Metrics: Automatic evaluation was limited to three metrics, with additional metrics left for future work.The paper names h/LEPOR, COMET, and MetaHOPE as possible additions.

A Language Identifiers

The paper directs readers to Table 4 for the language identifiers used throughout the paper and the WMT 2026 Test Suites Shared Task.

  • Language Identifiers: Table 4 lists the language identifiers used throughout the paper and the WMT 2026 Test Suites Shared Task.

B Additional English–Polish Translation Examples

The English–Polish examples show that translation quality depends on aspect, idiomatic meaning, collocational naturalness, terminology, and discourse context. Several outputs remain understandable or plausible even when they differ from the reference or contain linguistically important errors.

  • Examples: The examples compare source sentences, references, and outputs from selected systems using qualitative Polish analysis.
  • Literal translation of an idiomatic MWE: DeepSeek V4 Pro translates give me the time of day literally as dało mi pory dnia, losing the intended idiomatic meaning.TRIVE and Dubformer instead render the expression idiomatically, although the reference differs from those machine translations.
  • Cross-example interpretation: The examples illustrate that literal lexical correspondence can produce severe MWE mistranslation, while other deviations are primarily stylistic or remain plausible.
  • Verbal aspect in Polish: DeepSeek V4 Pro uses imperfective ukrywane where Polish context calls for a resultative or perfective reading, producing a less natural interpretation.
  • Context-dependent translations: Multiple Polish renderings of an English expression can be plausible because sentence-level context does not determine a single valid formulation.The example contrasts dopiąć swego, ubić interes, and dobić targu, which introduce slightly different semantics.
  • Collocational naturalness: Dubformer’s postawną klatkę piersiową is understandable but collocationally unnatural because postawny normally describes a person’s build rather than a body part.

C Additional English–Chinese Discussion

The English–Chinese discussion shows that translation errors can arise from ambiguous references, while human severity judgments may change after closer comparison with the reference and system outputs.

  • Severity judgments: Annotators revised one error from major to lower-severity unnaturalness after recognizing that the translation preserved the reference meaning despite different word order.The initial MIS severity was 8, but discussion changed it to severity level 2.
  • Severity judgments: The discussion identifies disagreement over error severity as a larger annotation issue, rather than disagreement about whether an error exists.This pattern is reported as similar to findings from Liang and Han (2026).
  • Translation errors: A second example contrasts a reference describing eyes closing with SCIR-TG’s output describing a door closing, while GoogleMT preserves the eye reference.The three outputs make the source of the system distinction explicit.
  • Translation errors: SCIR-TG and Tower-9B translated “eyes closed” as “door closed,” whereas GoogleMT translated the intended expression correctly.The example concerns a referential ambiguity in the Chinese output.

D Detailed Automatic Scores Per System Per Language Pair

The paper reports automatic scores and corresponding rankings for submitted systems across language pairs, using three metrics and highlighting the top three systems selected for human evaluation.

  • Automatic scores: SacreBLEU, chrF++, and BERTScore results are listed for submitted systems by language pair.The detailed scores appear in Tables 5, 6, and 7.
  • System rankings: System rankings corresponding to the automatic scores are reported separately for SacreBLEU, chrF++, and BERTScore.Tables 8, 9, and 10 provide the rankings, with rank 1 denoting the best-performing system.
  • Human-evaluation selection: The top three systems per language pair are identified with their original metric scores for subsequent human evaluation.Figure 5 presents the selected systems and their original scores.
Loading 2609.06634v1…