Source-linked AI summary

AI translation of literary texts is "fine", but readers still prefer human translations

Yves Ferstler, Adam Podoxin, Ty Brassington, Roman Grundkiewicz, Maite Taboada, Marzena Karpinska

arXiv:2606.26040v1cs.CL

TL;DR

Readers’ experience of AI-translated literature is poorly captured by existing evaluations. This paper compares human and agentic LLM translations through immersive and close reading, finding that readers generally prefer human translations despite judging machine translations reasonable and failing to identify them reliably.

  • Problem

    Existing literary-translation evaluations poorly capture readers’ immersion, literary experience, and responses to rhythm, emotion, voice, and deliberate word choices.

  • Method

    The study asks 15 avid readers to compare professional human and agentic LLM translations through uninterrupted 8K-word reading and side-by-side close reading across 15 French, Polish, and Japanese novels.

  • Results

    Readers consistently preferred human translations, especially in close reading, while finding machine translations reasonable, failing to detect them reliably, and automatic metrics favoring machine translations.

  • Takeaways & Limitations

    The findings support reader-centered evaluation for literary machine translation and future research on its effects beyond conventional automatic metrics.

  • Takeaways & Limitations

    The main evaluation covers only French, Polish, and Japanese translated into English, so larger studies are needed to assess other language pairs.

Abstract

from arXiv · show

AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers experience it in terms of immersiveness and literary effect, aspects poorly captured by automatic machine translation metrics or human evaluation targeting fluency and adequacy. We ask 15 avid readers to compare recently published human translations (HT) to machine translations (MT) generated with an agentic large language model (LLM)-based pipeline, for 15 recent novels in French, Polish, and Japanese and translated into English. Readers evaluated approximately 8K-word excerpts in two conditions: immersive reading of the whole excerpt (30 comparisons) and close reading of 386 aligned HT-MT chunk pairs (772 comparisons), with two readers per book and in alternating order of presentation. Overall, readers find MT "fine", but prefer HT (slightly at excerpt-level 19/30, more clearly at chunk-level 522/772) for its ease, clarity, and immersive nature. Readers' highlights show that MT's quality varies more within one book than HT's does. Crucially, readers cannot reliably tell the two apart (17/30 guess correctly) and tend to prefer the version they believe to be human. Automatic metrics, including LLM-as-a-judge approaches, fail to recover reader preferences and favor MT. We release LAIT (Literary AI Translation), a reader-centered evaluation dataset with 1K reader comments, 2K judgments and preference ratings, and 7.2K span-level annotations, along with our evaluation protocol and supporting interface.

1 Introduction

This study evaluates literary AI translation through readers’ immersive and close-reading experiences, finding that MT can be readable and difficult to detect but readers consistently value HT more. It introduces LAIT, a reader-centered dataset and protocol for assessing literary translation beyond conventional metrics.

  • Methodology: Readers compare professional human translations with an agentic LLM-based pipeline across 8K-word excerpts and 300-word aligned chunks from French, Polish, and Japanese novels translated into English.The study recruits 15 avid readers and uses uninterrupted full-excerpt reading followed by side-by-side close reading.
  • Contributions: The authors release a reader-centered protocol combining immersive and close reading, together with guidelines and annotation software to support reproducibility and future research.The protocol explicitly collects ratings, preferences, highlights, preference justifications, and suspected MT explanations.
  • Findings: 522/772 close-reading judgments and 19/30 immersive comparisons favor HT, although readers often find MT readable and sometimes prefer it.The preference for HT is slight at excerpt level but clearer when readers inspect aligned chunks.
  • Findings: MT quality varies more within individual books than HT quality, and MT success depends more on the book than on the source language.Reader comments praise HT particularly for immersion and related literary effects.
  • Findings: 17/30 readers identify MT correctly after reading both excerpts, while automatic metrics, including LLM-as-a-judge methods, favor MT rather than recovering reader preferences.Readers are misled by supposed AI cues such as em-dashes or assumptions that AI avoids swear words.
  • Contributions: LAIT contains 15 aligned FR/PL/JA→EN novel openings evaluated by two readers, with 2K ratings and preferences, 1K comments, and 7.2K span-level annotations.It also includes an unannotated 16-book development set with 1.7K paragraph-level alignments across five candidate MT systems.

2 Data & Methods

The study constructs LAIT from recent French, Polish, and Japanese fiction translated into English, pairing professional human translations with agentic machine translations. It evaluates these versions through immersive excerpt reading and close reading of aligned chunks across 15 books.

  • Dataset: LAIT contains 31 book-opening excerpts from recently published French, Polish, and Japanese novels, with 15 books reserved for human evaluation and 16 for MT-pipeline selection.Each human-evaluation excerpt pairs a published human translation with a corresponding machine translation.
  • Evaluation design: 15 excerpts of approximately 8K words were evaluated as whole reading experiences and as aligned approximately 300-word HT–MT chunk pairs.The excerpts averaged 7,623 words (SD = 449), while the chunked dataset yielded 386 aligned pairs.
  • MT pipeline: The selected MT system was P3, an AutoFiction-inspired agentic pipeline using Claude Code and Codex, chosen after a blind five-way preference comparison on 16 development books.P3 was chosen most often but only by a small margin, so adoption was a practical choice rather than evidence of decisive superiority.
  • MT pipeline: The pipeline generated translation guidelines, translated 1K-token chunks, applied local review-and-revision cycles, and then performed excerpt-level quality and cross-chunk consistency review.Up to three local cycles and two global cycles were allowed, with acceptance gates triggering revision when thresholds were not met.
  • Evaluation design: 60 excerpt-level evaluations and 772 chunk-level comparisons were collected, with two readers evaluating each book and each aligned chunk pair.The excerpt-level task used immersive reading, followed after a one-day break by side-by-side close reading of the aligned chunks.

3 Results & Analysis

Readers show a slight but consistent preference for human translation (HT), especially in close reading, while machine translation (MT) usually remains readable and sometimes wins. MT also varies more within books, whereas readers’ literary judgments and evidence favor HT more strongly than basic readability alone.

  • Translation preferences: 522 of 772 close-reading judgments favored HT, compared with 250 for MT, a significant order-adjusted preference for HT.At the excerpt level, 19 of 30 judgments favored HT versus 11 for MT, but this difference was not statistically significant.
  • Translation preferences: 54% of MT excerpt responses indicated willingness to continue reading, compared with 66% for HT, and roughly one-third of chunk choices favored MT.MT was not rejected across the board: 7 of the 11 excerpt-level MT preferences were clear preferences.
  • Variation across books and languages: MT preference varied from 4% to 88% across books but remained stable across languages, with French at 32%, Japanese at 34%, and Polish at 31%.Book identity was strongly associated with chunk-level preference (p<.001); Hooked was preferred in MT 88% of the time.
  • Reader evidence and annotations: 107.8 versus 68.5 positive-highlighted words per 1K words favored HT, while negative evidence was higher for MT at 100.7 versus 42.9.The higher-net-highlight version matched stated preference in 691 of 755 non-tied judgments (91.5%).
  • Within-book variation: 13 of 15 books showed greater variability in poor-wording highlights for MT than HT, while net quality was more variable for MT in 12 of 15 books.The paired sign tests were significant for poor-wording variability (p=.0074) and net-quality variability (p=.0352).
  • Reader comments and agreement: Both translations received praise for readability, but HT was praised more often for smoothness, with 19/30 positive mentions versus 11/30 for MT.The gap widened for more literary and experiential qualities, beyond MT’s basic readability bar.

4 Related Work

Prior evaluations of literary machine translation rely mainly on surface-level automatic metrics, error annotations, ratings, or short/local judgments. These approaches provide weak signals about literary quality and do not directly capture readers’ sustained reading experience.

  • Automatic evaluation: Automatic metrics and trained neural evaluators perform poorly on literary machine translation and weakly reflect perceived literary quality.LLM-as-a-judge and long-document evaluations remain indirect because they do not measure reader experience.
  • Human evaluation: Human evaluations commonly use error annotations or crowdsourced ratings, which do not capture reading experience.These methods are widely used but are not designed to represent how readers experience literary texts.
  • Human evaluation: Most literary machine-translation evaluations use short passages or local judgments, missing literary translation’s purpose as sustained reading.This limitation motivates evaluation protocols centered on longer, immersive reading.

5 Discussion & Conclusion

Across 15 recent French, Polish, and Japanese novels, readers found agentic LLM machine translation reasonable and difficult to detect, yet consistently preferred professional human translation, especially in close reading. The study releases LAIT and its evaluation pipeline to support future research on literary machine translation.

  • Study design: The evaluation compared human and machine translations across 15 recent French, Polish, and Japanese novels translated into English in immersive and close reading conditions.The machine translations were produced by a strong agentic LLM pipeline.
  • Findings: Readers preferred professional human translations over agentic LLM machine translations, especially during close reading, although they found MT reasonable and could not reliably identify it.Preferences varied mostly by book rather than source language, and readers often preferred the version they believed to be human.
  • Contribution: The study releases LAIT and its evaluation pipeline to support future research on literary machine translation.LAIT and the pipeline are presented as resources for continued investigation of literary MT.

Limitations

The study’s conclusions are limited by excerpt-based evaluation, a narrow and selective corpus, restricted language directions, and a small reader sample. These constraints reduce generalizability and statistical power, especially for exploratory multilingual results and chunk-level agreement.

  • Excerpt scope: 8K-word excerpts, rather than full books, may miss changes in literary quality, pacing, narrative voice, and book-spanning context.Financial limitations prevented translating and evaluating complete novels, although 8K words already matches the length of short stories.
  • Corpus selection: 2025–2026 critically acclaimed fiction with existing professional English translations limits corpus diversity and may not represent literature more broadly.The selection was intended to mitigate training-data contamination.
  • Languages and directions: The main evaluation covers only French, Polish, and Japanese translated into English, with five books per source language.The multilingual case study used two additional books and four target languages but had a much smaller, mixed-recruitment sample, so its results are exploratory and not generalizable.
  • Sample size and agreement: Two Upwork-recruited readers per book limit statistical power and leave chunk-level inter-reader agreement modest in places.A within-subject design reduces individual variability and requires fewer annotators than a between-subjects design.

Ethics Statement

The study obtained ethics approval, informed consent, fair compensation, and anonymized participant identities. It limits copyrighted data sharing, discloses AI assistance, and discusses corpus bias and risks of literary machine translation.

  • Human participants: The Simon Fraser University Research Ethics Board approved the study, and participants gave informed consent, could withdraw with prorated compensation, and received fair payment.Participants were paid $110 USD per book, approximately $27.50 per hour for about four hours, plus completion bonuses; identities were anonymized by participant codes.
  • Copyright and data sharing: The researchers purchased source books for research and retain no more than 10% of any individual work.The dataset will be released for research purposes only under fair-dealing or fair-use principles.
  • Corpus and societal impact: The corpus overrepresents internationally published, commercially successful fiction, while the study aims to improve evaluation practice and publishing transparency rather than replace translators or deceive readers.The authors report that readers often prefer human-authored translations.
  • Risks of literary machine translation: Literary machine translation can appear acceptable while altering readers’ experiences through problems with phrasing, context, voice, and potentially biased or harmful wording.These risks are especially concerning when AI-translated fiction reaches readers with little human review.
  • AI use disclosure: The authors used coding agents for plots and table-layout edits and language models as writing assistants, but no part of the paper was generated from scratch.This disclosure distinguishes assistance with presentation and writing from fully automated paper generation.

A Dataset · B Machine Translation · C Agentic pipeline runs

The appendix documents the evaluation dataset, compares five machine-translation setups—including an agentic pipeline—and details the selected pipeline’s reviews, acceptance gates, and run artifacts. The agentic system processed hundreds of source chunks across evaluation, development, and multilingual runs, while some translations remained acceptable despite failing enforced limits.

  • A Dataset: The dataset appendix lists evaluation and development books with bibliographic metadata and reports statistics for books, aligned chunks, and collected comments.These materials support the English human evaluation and the selection of an optimal machine-translation pipeline.
  • B Machine Translation: Three pipelines were tested in five setups: full-excerpt translation, chunk-level translation with excerpt-level post-editing, and an agentic pipeline.The first two used GPT-5.4 or Gemini 3.1 Pro; the agentic pipeline used Claude Code (Opus 4.6) and Codex (GPT-5.4).
  • B Machine Translation: The agentic pipeline combined data preparation, chunk-level translation, and excerpt-level review, using metrics and guard gates to trigger revisions when issues were detected.Its design was motivated by long-context and multi-document handling capabilities.
  • B Machine Translation: P1 translated each complete excerpt in one call, whereas P2 translated paragraph-preserving chunks targeted at a maximum of 1K source tokens before excerpt-level post-editing.P1 exposed both stages to full discourse context; P2 used smaller first-pass source chunks.
  • B Machine Translation: 25 yes/no/maybe quality questions and a four-level literary severity review governed chunk acceptance, requiring no NO judgments, at most five MAYBE judgments, and no MEDIUM+ findings.A chunk failed the local gate if either review component failed and was sent to another revision cycle.
  • C Agentic pipeline runs: Even when final translations failed to pass within enforced limits, increasing those limits did not yield justifiable improvements, and the failed translations were often acceptable.Run details include the numbers of chunks and excerpts passing acceptance gates within the enforced limits.
  • C Agentic pipeline runs: 35 completed agentic runs processed 441 source-language chunks and produced 1,059 chunk-cycle translation or revision artifacts.The runs comprised 15 evaluation-book, 16 development-book, and 4 multilingual target-language runs; job-level logs covered 31 runs and recorded 2,930 agent jobs, including 19 failed jobs.

D Human evaluation

The human evaluation recruited avid, native-level English readers through Upwork and assessed human and machine translations using questionnaires embedded in immersive-reading and close-reading interfaces.

  • Participant recruitment: Participants were recruited through Upwork and limited to avid readers who read at least one English book monthly and had native-level English command.Recruitment also collected reading habits and used an entry questionnaire.
  • Questions: The evaluation used separate and comparative questionnaires for immersive reading, plus short comparison questionnaires for each close-reading chunk pair.Immersive reading covered each text separately and then both translations together.
  • Interface: Immersive-reading presentation order was counterbalanced, while close-reading translation display order was randomized for side-by-side local comparison.The interfaces also supported holistic comparison and chunk-level difference detection.
  • Display normalization: Lightweight normalization standardized quotation marks, guillemets, apostrophes, and ellipses in the evaluation interface without modifying source files.The same normalized representation was used for span highlights and participant offsets.

E Human evaluation results

The human-evaluation appendix reports preference, agreement, rating, span-highlight, comment, AI-detection, and post-study interview analyses. Readers were often pleasantly surprised by MT’s quality, yet generally remained more favorable toward HT and could not detect AI translation reliably.

  • Overall results: The appendix breaks down book-level and chunk-level preferences, source-language patterns, highlighted spans, AI guesses, and participant-specific preference direction and strength.These analyses are presented across Figures 19–27.
  • Agreement: 0.307 overall span overlap and 0.217 label-aware overlap measured agreement in close-reading highlights.Among readers agreeing on the preferred translation, overlap increased to 0.322 and label-aware overlap to 0.263.
  • Comment analysis: Readers’ comments were inductively coded to compare positive and negative aspects of HT and MT and reasons for preferences at excerpt and chunk levels.Separate analyses identify why readers preferred HT or MT and visualize comments about the other candidate.
  • AI-detection guesses: Readers who often use AI were not better at detecting AI translation, with no apparent relationship between AI experience and detection accuracy.The appendix compares single-reading guesses with final comparison guesses and reports additional correctness and confidence flows.
  • Post-study interviews: Participants initially expressed skepticism about LLM translation but were generally pleasantly surprised by MT’s quality, despite a general preference for HT.This pattern also held for some participants who correctly discriminated HT from MT.

F Case study results

Across Spanish, French, Polish, and Japanese case studies, readers strongly preferred human translations in direct comparisons and close reading, while span annotations favored HT and disfavored MT. MT identification was reliable for Polish, Spanish, and Japanese in the final excerpt comparison, but not for the French reader.

  • MT identification: Readers reliably identified MT for Polish, Spanish, and Japanese in the final excerpt comparison, but the French reader was incorrect.The Spanish reader was initially incorrect in single-reading but later distinguished the translations reliably during comparison.
  • Preference judgments: 4 out of 5 readers preferred HT at the book level, and all 5 preferred it in close-reading judgments across target languages.Preferences shifted strongly toward HT after direct comparison and aligned-chunk evaluation.
  • Span annotations: Readers marked mostly positive evidence in HT and mostly negative evidence in MT across target languages.Median span lengths indicated how localized these judgments were.

G Statistical analysis

The analysis matches each response type with an appropriate mixed-effects model and uses order adjustment or simpler models when crossed random-effects structures are not estimable. Reader and book random intercepts are included whenever estimable.

  • Binary preferences: Binary HT-versus-MT choices use logistic GLMMs, with HT coded 1 and MT coded 0, and reader and book random intercepts when estimable.These models estimate whether readers are more likely to choose HT than MT.
  • Dialogue and word choice: Three-level dialogue and word-choice responses use signed preferences and order-adjusted fixed-effect linear models because crossed reader-and-book models were singular.HT is coded +1, no difference 0, and MT -1.
  • Preference strength: Preference-strength choices become signed scores analyzed with linear mixed-effects models when reader-and-book random-effects structures are estimable.HT preferences are positive, MT preferences negative, and larger absolute values indicate stronger preferences.
  • Single-reading ratings: Five-point Likert ratings use cumulative-link mixed models, which account for ordered categories without assuming equal distances between adjacent scale points.The ratings include acceptability, smoothness, immersion, and willingness to continue reading.
  • AI detection: AI-detection correctness is analyzed with logistic GLMMs, coding correct identification as 1 and incorrect identification as 0, with reader and book intercepts when estimable.In single-reading tasks, labels must correctly identify MT as machine-translated or HT as human-translated.
  • Model specification: Across analyses, singular mixed-effects fits are replaced by simpler models, and the reported formulas and outputs are provided in Table 35.Analyses were conducted in R using lme4, lmerTest, emmeans, and ordinal.

H Automatic Evaluation · 5 What, if anything, stood out positively in the wording or style of this version?

The appendix describes automatic evaluations comparing five MT setups across development and evaluation books, and testing whether metrics recover human preferences. It uses paragraph-level alignment and metrics plus chunk-level scoring on HT and agentic/P3 MT.

  • H Automatic Evaluation: Two automatic evaluations compare five MT setups and test whether metrics recover readers’ preferences against professional HT and agentic P3 MT.The setup comparison covers 16-book development and 15-book evaluation sets.
  • H Automatic Evaluation: Paragraph-level alignment uses shorter units because several metrics are constrained by context-window size.Source paragraphs are paired with Microsoft Translate output as an English pivot, then candidate English versions are evaluated.
  • H Automatic Evaluation: COMET-22 and METRICX use professional HT as a reference, whereas COMETKIWI and METRICX-QE are reference-free.COMET is higher-is-better, while METRICX is lower-is-better.
  • H Automatic Evaluation: Metric input limits are 512 tokens for COMET and 1,536 tokens for METRICX, with over-limit paragraph pairs skipped.Reference-based metrics are not reported for the HT row because they would score the reference against itself.
  • H Automatic Evaluation: Book-and-system results report the mean of all remaining paragraph scores for each metric.These paragraph-level automatic-evaluation results are summarized for all systems in Table 39.
  • H Automatic Evaluation: Chunk-level evaluation scores approximately 300-word aligned HT and agentic/P3 MT chunks with METRICX-QE and LITRANSPROQA.It also maps paragraph-level COMETKIWI and METRICX-QE scores onto the close-reading chunks.
  • 5 What, if anything, stood out positively in the wording or style of this version?: The evaluation materials include participant ratings of fluency, literary quality, immersion, willingness to continue, and perceived AI authorship.Paired-reading forms record whole-excerpt preferences, while per-chunk forms record local preferences and reasons.
  • 5 What, if anything, stood out positively in the wording or style of this version?: Participant questionnaires ask how acceptable, smooth, immersive, and continuation-worthy each translation felt.The response scales range from negative judgments to fully acceptable, very smooth, strongly immersion-supporting, and definitely yes.

6 What, if anything, stood out negatively in the wording or style of this version?

The section operationalizes negative wording and style through participants’ free-text comments, with negative codes applied to Single Q6 and to the non-preferred translation in comparative judgments. The supplied passages identify the annotation framework and contextual dimensions considered, but do not report specific negative wording patterns or frequencies.

  • Reported findings: The supplied passages do not provide counts or concrete examples of which wording or style problems stood out negatively.The listed figure and table descriptions indicate where such analyses or examples may appear, but do not state their findings.
  • Annotation of negative aspects: Negative comments were systematically annotated in Single Q6 and in comparative or chunk-level justifications for the non-preferred translation.The annotation scheme distinguishes positive codes for preferred translations from negative codes for non-preferred translations.
  • Contextual evaluation: Readers’ explanations could evaluate wording against genre, period, cultural setting, character voice, and reading effort.These contextual factors appear in examples of excerpt-level and chunk-level preference explanations.
Loading 2606.26040v1…