Source-linked AI summary

Large language models effectively leverage document-level context for literary translation, but critical errors persist

Marzena Karpinska, Mohit Iyyer

arXiv:2304.03245v3cs.CL

TL;DR

Document-level literary translation by LLMs lacks reliable evaluation, especially for literary paragraphs where discourse context and authorial voice matter. The paper evaluates GPT-3.5 paragraph-level translation against sentence-level strategies and Google Translate across 18 language pairs using bilingual human annotators. Paragraph-level translation is preferred and has fewer mistranslations, grammatical issues, and stylistic inconsistencies, but omissions and other critical errors remain, so human intervention is still necessary.

  • Problem

    Reliable evidence about LLM performance on paragraph- and document-level literary translation is limited because evaluation is costly and automatic metrics are unreliable.

  • Method

    The study uses bilingual translators to annotate errors and preferences for GPT-3.5 paragraph-level and sentence-level translations across 18 literary language pairs, with Google Translate as a comparison.

  • Results

    GPT-3.5 paragraph-level translations are higher quality than sentence-level GPT-3.5 translations and Google Translate, with improved coherence, literary style, and context handling.

  • Takeaways & Limitations

    Paragraph-level discourse context helps GPT-3.5 produce more coherent literary translations with fewer mistranslations and grammatical issues than sentence-by-sentence translation.

  • Takeaways & Limitations

    GPT-3.5 paragraph-level translations still contain critical mistranslations and occasional content omissions, and human intervention remains necessary to preserve authorial voice.

Abstract

from arXiv · show

Large language models (LLMs) are competitive with the state of the art on a wide range of sentence-level translation datasets. However, their ability to translate paragraphs and documents remains unexplored because evaluation in these settings is costly and difficult. We show through a rigorous human evaluation that asking the Gpt-3.5 (text-davinci-003) LLM to translate an entire literary paragraph (e.g., from a novel) at once results in higher-quality translations than standard sentence-by-sentence translation across 18 linguistically-diverse language pairs (e.g., translating into and out of Japanese, Polish, and English). Our evaluation, which took approximately 350 hours of effort for annotation and analysis, is conducted by hiring translators fluent in both the source and target language and asking them to provide both span-level error annotations as well as preference judgments of which system's translations are better. We observe that discourse-level LLM translators commit fewer mistranslations, grammar errors, and stylistic inconsistencies than sentence-level approaches. With that said, critical errors still abound, including occasional content omissions, and a human translator's intervention remains necessary to ensure that the author's voice remains intact. We publicly release our dataset and error annotations to spur future research on evaluation of document-level literary translation.

1 Introduction

This paper examines whether GPT-3.5 can use paragraph-level discourse context to improve literary translation, addressing the difficulty of evaluating document-level translation. Human evaluation across 18 language pairs compares paragraph-level translation with sentence-level strategies and Google Translate.

  • The study evaluates literary paragraphs because preserving authorial voice and contextual nuances often requires stylistic or content-based changes across sentence boundaries.At least 55% of reference target paragraphs split or merge source sentences, underscoring the limits of sentence-level pipelines.
  • The evaluation collects span-level error annotations, preference judgments, and justifications from translators comparing multiple GPT-3.5 strategies and Google Translate.The study gathers annotations on 720 translated-paragraph pairs across 18 language pairs and three target languages.
  • GPT-3.5 paragraph-level translation produces significantly higher-quality literary translations than sentence-level GPT-3.5 methods and Google Translate.The comparison is based on annotated translation errors and translator preference judgments.
  • Paragraph-level translation improves coherence, literary-style preservation, and handling of context-dependent expressions, as illustrated by fewer word-choice and pronoun errors.Figure 2 contrasts paragraph-level and sentence-level Japanese-to-English translations.
  • Despite these gains, GPT-3.5 paragraph-level translations still contain numerous critical mistranslations and require human intervention to preserve the author’s voice.The remaining errors show that context-rich literary translation remains an open challenge.

2 Background

Prior document-level machine translation work incorporated discourse context through statistical and neural architectures. Recent LLM translation research spans several tasks, but this paper focuses specifically on literary paragraph translation.

  • Earlier machine translation systems modeled discourse context using concatenation, hierarchical, and multi-pass architectures.
  • Document-level translation research uses the term “document-level” for both multi-sentence passages and complete documents.
  • Recent LLM translation studies examine paragraph-level post-editing, sentence-level translation, hallucinations, evaluation, and prompt engineering.

3 Data & methods

The study builds a multilingual literary-paragraph dataset from recently published novels and compares GPT-3.5 prompting strategies with and without discourse-level context. Its design spans diverse languages, genres, paragraph structures, and few-shot demonstrations.

  • 3.1 Dataset collection: The dataset contains aligned literary paragraphs from 18 recently published novel translations, covering eight source languages and English, Polish, and Japanese targets.The collection includes 20 paragraphs from each translation and both dialogue and narrative texts.
  • 3.1 Dataset collection: Paragraphs contain at least two sentences, with most ranging from four to nine sentences, and are manually sentence-tokenized for comparison.The mean paragraph length is 7.45 sentences with a standard deviation of 4.14.
  • 3.1 Dataset collection: The language selection spans Indo-European, Sino-Tibetan, and Japonic families with varied morphological and writing-system properties.The study deliberately includes linguistically similar and dissimilar source-target pairs.
  • 3.2 Translation with large language models: Few-shot prompts use five manually curated literary demonstrations for each of the 18 language pairs, including both dialogue and narrative examples.The demonstrations come from novels outside the translation dataset.

4 Evaluating document-level literary translation

Because automatic metrics are unreliable for literary document-level translation, the study uses bilingual translators to annotate errors and compare candidate translations. The evaluation combines span-level MQM-inspired labels, omissions and additions, preferences, and written justifications.

  • Automatic metrics are considered unreliable for literary document-level inputs, motivating a detailed human evaluation by translators fluent in source and target languages.Existing document-level metrics may depend on sentence alignments or be available only for English.
  • Translators compare PARA with SENT, PARA_SENT, and Google Translate for each source paragraph.The annotation pipeline evaluates three candidate pairs per paragraph.
  • Annotators mark error spans using categories including mistranslation, grammar, untranslated text, inconsistency, register, and format.The categories are based on a predefined schema inspired by MQM.
  • Literary translation errors require contextual judgment because poetic license can make some changes acceptable rather than straightforwardly incorrect.The evaluation recognizes that literary translators may alter details to improve readability or enjoyment.
  • After span annotation, translators identify significant content additions or omissions, select the better translation, and explain their preference in two to five sentences.They also indicate whether the preferred translation is significantly superior or comparable in quality.

5 Results

Across automatic metrics and human judgments, paragraph-level GPT-3.5 translation (PARA) outperforms sentence-level and competing methods across the evaluated literary language pairs, while remaining imperfect.

  • 5 Results: PARA outperforms competing methods across all evaluations and language pairs, showing that GPT-3.5 effectively leverages paragraph-level context.The evaluation combines automatic metrics with aggregate human-evaluation statistics.
  • 5.1 Automatic metrics favor PARA: PARA significantly surpasses SENT and GTR on COMET, BLEURT, and COMET-QE, and surpasses GTR on BERTSCORE (p<.001).These paragraph-level automatic metrics were not explicitly designed for paragraph outputs and should be interpreted cautiously.
  • 5.2 Human evaluation also favors PARA: PARA is preferred over SENT by 71.1% of translators and over GTR by 82.8%, with both differences statistically significant at p<.001.Preference remains significant after excluding unsure votes: 78.5% versus SENT and 88.0% versus GTR.
  • 5.2 Human evaluation also favors PARA: PARA is slightly preferred over PARA_SENT at 66.1%, while PARA_SENT produces comparable mistranslation, grammar, and inconsistency counts but leaves around 22% more words untranslated.PARA also avoids the occasional sentence repetitions reported for PARA_SENT.
  • 5 Results: Models with paragraph-level context commit fewer translation errors overall, but the evaluation still identifies substantial errors and constrained error categorization.Some omissions or additions were annotated only as a binary serious-error decision because of time restrictions.
  • 5.2 Human evaluation also favors PARA: PARA translations are praised for smoother literary style, stronger rhetoric, better reflection of content and style, and greater within-paragraph consistency.Translators also note that PARA uses more poetic license, while both systems can still fall short.

6 Analyzing translation errors

GPT-3.5 paragraph-level translation produces fewer and better-resolved errors than sentence-level approaches, but substantial language-specific and context-dependent mistakes remain.

  • Overall findings: PARA translations are favored overall because they make fewer errors and better handle context than SENT, PARA_SENT, and GTR.The analysis links these gains to improved coherence, literary-style preservation, and handling of context-dependent expressions.
  • 6.1 Language-specific grammatical errors: English outputs contain fewer grammatical mistakes, especially incorrect articles, while Japanese systems struggle with particles and Polish systems with gender, case, and prepositions.For Japanese, PARA and SENT make twice as many particle errors as PARA_SENT and GTR; Polish GPT-3.5 outputs show 55 errors for PARA, 86 for SENT, and 64 for PARA_SENT in the cited comparison.
  • Overall findings: PARA generates the fewest grammatical errors across Japanese and Polish, with 97 errors versus 136 for SENT, 101 for PARA_SENT, and 122 for GTR.None of the systems is free of grammatical inaccuracies, including in English.
  • 6.2 Context-related errors: Paragraph context resolves pronoun ambiguities that sentence-level translation mishandles, including Russian gender agreement and Japanese second-person reference.PARA changes Russian “he” to “it” and changes “Furukura” to “you” when broader context identifies the intended referent.
  • 6.2 Context-related errors: Context also improves translations of ellipsis, ambiguous predicates, repeated wording, and situation-dependent expressions.PARA correctly translates Polish ellipses as “wash,” attributes “unaware” to the mother, preserves the repeated German equivalent of “bad,” and renders a clerk’s honorific as “right away.”

7 Limitations

Paragraph-level GPT-3.5 translation improves coherence and reduces several error types, but critical failures remain, including omissions, subject confusion, context misinterpretation, and loss of authorial voice.

  • PARA occasionally omits source content, and omissions appear more prominent than in SENT or GTR.
  • PARA translations contain fewer mistranslations and grammatical errors than SENT or GTR, while offering greater coherence and better literary-style preservation.
  • PARA can merge sentences with distinct subjects, confuse context-dependent pronouns, and mistranslate omitted subjects in Japanese.
  • In Japanese-to-English and Polish examples, PARA incorrectly assigns a second sentence’s subject to Miho instead of preserving the narrator’s perspective.
  • The study’s languages are mid- or high-resource, so performance may be worse for low-resource languages such as Zulu or Armenian; GPT-4 also does not resolve all issues.
  • Even accurate or enjoyable output may lose the original author’s voice, so human intervention remains necessary for literary translation.

8 Conclusion

The paper finds that paragraph-level LLM translation improves literary paragraph quality over sentence-level approaches and releases resources for better document-level evaluation. It also identifies full-novel coherence as an unresolved challenge.

  • Paragraph-level translations are more coherent and enjoyable than sentence-by-sentence translations and contain fewer mistranslations and grammatical issues.
  • Professional translators prefer PARA over GPT-3.5 sentence-level translations and Google Translate, while the released dataset and annotations support new evaluation methods.
  • Future work will integrate paragraphs into cohesive chapters and eventually full novels.
  • Providing three additional preceding paragraphs enabled accurate translations from both GPT-3.5 and GPT-4 in one analyzed example.

Ethical considerations

The paper situates LLM translation within broader ethical concerns about bias, toxicity, and literary translation’s multiple stakeholders. Human-evaluation procedures received IRB approval and translator consent.

  • LLMs encode biases and toxicity, and unconstrained prompting can exacerbate these behaviors.
  • Literary machine translation raises ethical concerns involving authors, translators, and other stakeholders.
  • The human-translator experiments were IRB-approved, and translators consented to disclosure of their annotations, comments, and preferences.

A The Dataset

The dataset samples literary paragraphs and considers practical alignment, intelligibility, and linguistic ambiguity during construction. It also motivates paragraph-level evaluation by showing that literary meaning can depend on context and translation choices.

  • Selecting paragraphs from novels: Sampling prioritized varied dialogue and narrative, human intelligibility without extra context, and feasible source–translation alignment without major cross-paragraph rearrangement.
  • Selecting paragraphs from novels: Japanese first-person texts can leave narrator gender indeterminate, so translators accepted either gender when consistent within the paragraph.
  • A note on literary translation: Literary translation requires choices beyond word equivalence because style, emotion, and meaning may depend on sentence-spanning context.
  • Paragraph length: Figure 7 reports the sentence-count distribution of sampled paragraphs, which were manually sentencized for comparison.

C Human Evaluation

The study uses bilingual human translators to evaluate paragraph-level literary translation and classify its errors. The evaluation addresses annotation subjectivity, sentence alignment, translator qualifications, and omissions.

  • Annotation considerations: Error annotation is inherently subjective because some translation choices permit multiple interpretations or error labels.The study distinguishes mistranslation from grammatical error based on whether the source was misunderstood or the resulting translation was ungrammatical.
  • Data considerations: About 55% of the data potentially lacks one-to-one sentence correspondence because translations split or merge sentences.This complicates sentence-level comparisons and alignment across the source and translations.
  • Annotation considerations: The annotation protocol used a minimal strategy for inconsistency and register errors, while quotation-mark format errors were manually corrected for comparability.The correction applied to SENT and PARA_SENT translations prevented simple quotation-mark heuristics from dominating comparisons with PARA.
  • Observed errors: PARA occasionally omits storyline-critical details and appears more omission-prone than SENT and GTR, while PARA_SENT reduces but does not eliminate this issue.PARA_SENT also introduces some repetition issues.
  • Human evaluators: The translators were highly proficient in the source language, mostly native in the target language, and instructed to evaluate paragraphs without prior book knowledge.Translators were recruited through Upwork; non-native target-language annotations were verified by native speakers.

D Pivot Pilot

The pilot tests whether translating through English improves paragraph-level translation for language pairs without English. The results find no apparent general benefit, so subsequent experiments use direct translation.

  • Pilot design: The pilot evaluated 20 passages for each non-English language pair and used 200 pairwise preference judgments in total.Language pairs involving English did not require pivoting.
  • Pilot setup: The PARA_PIVOT process provided both the original source and its English translation to help preserve information that English may omit, such as grammatical gender.Adding the source helped overcome this potential loss, but English pivoting itself showed no clear gain.
  • Pilot result: Pivoting appeared to help only for Polish-to-Japanese translation, although the reason is unclear and may relate to archaic expressions in the Polish novel.The authors suggest English training exposure might help with these difficult phrases, but present this as a possibility.
  • Pilot result: No apparent gains from English pivoting were observed, and the direct-translation setup was retained because pivoting also reduces the number of prompt examples.The reported comparison had p=0.62 with a 95% interval of [0.448, 0.591].

E Automatic Metrics

The study compares automatic metrics with human judgments and analyzes translation setups using mixed-effects models. COMET shows the strongest reported agreement with the human evaluation.

  • Metric correlation: The correlation analysis evaluates automatic metrics against both all human judgments and judgments where annotators were confident that one translation was clearly better.Agreement is measured using accuracy and Kendall’s Tau.
  • Metric correlation: COMET achieved the highest agreement with human judgments: 64.04% accuracy and Kendall’s Tau 0.341 overall, rising to 72.78% and 0.456 for confident votes.The evaluation reports both accuracy and Kendall’s Tau for all judgments and the confident-vote subset.
  • Statistical analysis: Linear-mixed-effects models analyze metric scores with translation setup as a fixed effect and source paragraph effects represented through random effects.The setups are PARA, SENT, PARA_SENT, and GTR.
  • Statistical analysis: Post hoc pairwise comparisons are reported for BLEURT, COMET, COMET-QE, and BERTSCORE.The comparisons use the emmeans package after the mixed-effects analyses.
Loading 2304.03245v3…