Source-linked AI summary
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, Sara Rajaee, Pinzhen Chen, Jannis Vamvas, Sara Papi, Ona de Gibert, Bhavitvya Malik, Eliya Habba, Orfeas Menis Mastromichalakis, Patrícia Schmidtová, Michelle Wastl, Sheriff Issaka, Leshem Choshen, Stella Biderman, Antonis Anastasopoulos, Jan Niehues, Rico Sennrich, Mrinmaya Sachan, Ondřej Bojar, Kenton Murray, Jörg Tiedemann, Alham Fikri Aji, Philipp Koehn, Christof Monz, Alexandra Birch, Sowmya Vajjala, Chalamalasetti Kranti, Cristina España-Bonet, Nobin Sarwar, David Kaczér, Shunta Asano, Malik Marmonier, Daban Q. Jaff, Vaisakhi Mishra, Hend Al- Khalifa, Gabriele Sarti, Sourajit Saha, Nils Rehlinger, Juan Daniel Cuervo Villa, Jonathan Tonglet, Saugata Purkayastha, Dominik Macháček, Jagannathan Ramanujam, Heejin Do, Zuzana Nadova, Fred Philippy, Fabian Retkowski, Maria Lymperaiou, Silvia Casola, Hanna Yukhymenko, Shubhashis Roy Dipta, Sangwon Ryu, Andrés Jerez, Ron Keinan, Shuaib Shuaib Yusuf, Avantica Vempati, Maria Carmen Staiano, Sukannya Purkayastha, Adrian Cosma, Vitalii Babenko, Erivan Inan, Aviral Nigam, Wafa Aissa, Fatima Haouari, Venkata Prasanth Kumar Gummadi, Mehdi Jafarzadeh, Valentin Scourneau, Lukas Edman, Kaiser Sun, Shaomu Tan, Mohammad Sadegh Gholizadeh, Johannes-Rudolf David, Dipankar Srirag, Javier García Gilabert, Ruta Binkyte, Manar Ali, Ana-Maria Bucur, Sabry E. Farrag, Youssef Saber, Yihong Liu, Jean Maillard, Cojocaru Nicoleta, Xiaochuang Yuan, Sina Ahmadi, Philipp Mondorf, Kaustubh Dhole, Roman Wixinger, Shenbin Qian, Manuel Tuor, Sergey Troshin, Jonathan Yahav, Fida Mohammad Thoker, Amir Arsalan Rezapour, Lance Calvin Lim Gamboa, Manon Reusens, Kätriin Kukk, Koel Dutta Chowdhury, Giuseppe Gallipoli, Christian Hoang, Shaswati Saha, Seth Aycock, Jan Kocoń, Bo Chen, Linh Vu, Vatsal Venkatkrishna, Arafat Ahsan, Luan Thanh Nguyen, Hassan Soliman, Daryna Dementieva, Theresia Veronika Rampisela, Ngoc Quynh Tram Do, Marius Huber, Kazuki Egashira, Azmine Toushik Wasi, Vladislav Poritski, Mike Zhang, Deep Shah, Paul Gavrikov, Luis Frentzen Salim, David Africa, R. Damanhuri, Bello Umar Bello, Anumit Garg, Gengyu Rao, Pawan Sasanka Ammanamanchi, Kamile Dementaviciute, Andrianos Michail, L D M S Sai Teja, Dawei Zhu, Yi Fan, Wei Liu, Farhan Farsi, Elias Herranen, Sankalan Pal Chowdhury, Karen Sanchez, Farzad Shami, Ashok Urlana, Zimu Wang, Tomasz Limisiewicz, Priyaranjan Pattnayak, Marii Ojastu, Hongbin Na, Emilian Radoi, Chenyi Zhao, Carlos Hinojosa, Andrea Gregor de Varda, Zaid Alyafeai, Reem Alzahrani, Nehal Kathrotia, Alex Flückiger, Ulysses Sekai Tully Carr, Jimson Paulo Layacan, Guy Kaplan, Ritwik Tiwari, Rishit Dagli, Oksana Volchek, Isaac R Caswell, Bowen Yi, Blanka Kövér, Amir Hossein Yari, Aicha Chorana, Zhengxiang Wang, Selja Keränen, Samuel Simko, Joy Olusanya, Jenny Chim, Enzo Doyen, Vivek Harsha Lakkamaneni, Sophia Conrad, Pouya Sadeghi, Panayiotis Panayiotou, Luis Lara, Jannatul Nayem, Eran Yahav, Debanshu Das, Antonia Karamolegkou, Anmol Goel, Aishik Mandal, Tommaso Cerruti, Raoyuan Zhao, Mykola Haltiuk, Thura Aung, Naser Almousa, Amir Hossein Kargaran, Rachel Bawden, Qiaoyuan Zheng, Mateusz Lango, Beni Egressy, Fidel Rodríguez Velásquez, Natchapon Jongwiriyanurak, Minh Ngoc Do, Marco Gaido, Lena Libon, Dzmitry Kuzmin, Badal Nyalang, Antoine Taroni, Andrei Niculae, Abdulaziz Nura Kani, Rushikesh Zawar, Marek Šuppa, Beatrice Savoldi, Andreas Simons, Rayyan Merchant, Ilai Yaron Levy, Francesco Pinto, Ziyi Yang, Yolanda Xavier, Samuel Frontull, Muhammad Ravi Shulthan Habibi, Kenneth Enevoldsen, Harris Abdul Majid, Francesca Padovani, Tim Graf, Tatiana Bielakova, Sharifa Djurabaeva, Shaoxiong Ji, Raia Abu Ahmad, Pavel Stepachev, Jirui Qi, Ayush Sunil Munot, Alireza Pakniat, Ayla Rigouts Terryn, Yuxing Lu, Yurii Paniv, Xiyan Fu, Tosin Adewumi, Sunisth Kumar, Stéphane J. P. S. Thunus, Shree Harsha Bokkahalli Satish, Shayan Bali, Prakhar Gupta, Papa Abdou Karim Karou Diallo, Matija Akrap, Marko Culjak, Kristýna Onderková, Joseph Attieh, Esrael Teferi Tensay, Elisabeth Fittschen, Benoît Sagot, Jingwei Ni, Yu Fan
TL;DR
Existing translation benchmarks and evaluation methods do not reliably expose important failures in strong models. The Last Translation Benchmark addresses this with human-authored, peer-reviewed multimodal challenge examples and example-specific verification rules. Its reported cases include universal failures involving a sign-language distractor, low-resource translation, and a polysemous contraction, while the benchmark remains a non-exhaustive stress test.
Problem
Standard translation benchmarks are approaching saturation, while automatic and human evaluations can be unreliable, opaque, non-reproducible, and difficult to scale.
Method
The benchmark crowdsources and peer-reviews difficult text, image, audio, and video examples, pairing each with handcrafted verification rules for automatic checking.
Results
All models fail on a Chinese Sign Language→Chinese example with a Peking University distractor, a Sandawe→Arabic example prompts refusal and hallucinations, and all models mistranslate a Yoruba→English polysemous contraction.
Takeaways & Limitations
The live benchmark provides a long-term goalpost for evaluating and diagnosing frontier translation failures through concrete, example-specific criteria.
Takeaways & Limitations
The benchmark is a stress test rather than an estimate of typical translation quality, and its verification rules and taxonomy are not exhaustive.
Abstract
from arXiv · showhide
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.
1 Introduction
The paper argues that existing translation benchmarks and evaluation methods fail to expose important errors in strong models. It introduces the Last Translation Benchmark, which pairs difficult examples with specific verification rules for reproducible diagnosis.
- Standard benchmarks are nearing saturation and fail to distinguish between strong translation models or expose their failure modes.
- Existing automatic metrics and human evaluation are unreliable, opaque, subjective, expensive, or difficult to reproduce at high translation quality.
- The Last Translation Benchmark combines difficult text and multimodal examples with verification rules targeting the specific failure modes tested.
- Each translation must pass every example-specific rule, producing a verifier pass rate that serves as a calibrated evaluation scale.
- The benchmark is intended as a live, long-term goalpost for model benchmarking and diagnosis, with contributions continuing through future releases.
- In an English→Czech example, Google Translate, Gemini 3.1 Pro, and GPT-5.6 Sol each select feminine forms for male nurses, while the human translation passes.
2 Creating the Last Translation Benchmark
The benchmark is built through crowdsourced submissions of difficult textual or multimodal examples, followed by rule-based automatic checking and human review. Contributors must show that references pass the rules while most displayed model translations fail them.
- Contributors submit difficult text, image, audio, or video inputs with correct reference translations through a custom online platform.
- The platform displays translations from several models, which contributors inspect to identify failure modes and formulate verification rules.
- An LLM checks candidate translations against every verification rule, and accepted submissions require a passing reference plus failures from most automatic translations.
- Reviewers check guideline compliance, fairness to expert human translators, and whether rule-detected errors are perceptible and significant.
- Multimodal content is translated directly when primary, but serves as disambiguating context when attached to text.
3 Analyzing the Last Translation Benchmark
LTBv1 evaluates difficult translation examples with verifier pass rates and reveals failures spanning linguistic, extralinguistic, and atypical phenomena. Its results indicate that verification rules provide more objective and interpretable evaluation than generic judges and metrics.
- Dataset and benchmark results: LTBv1 contains 3456 accepted examples across 109 main languages, with 94% textual submissions averaging 19 words or 104 characters.Language pairs involving English dominate, while non-English→non-English examples comprise 13% of the dataset.
- Dataset and benchmark results: Verifier pass rate measures the percentage of examples whose translations satisfy all verification rules, making the benchmark difficult across 29 evaluated models.Each example is difficult for at least all but two listed models.
- Evaluation with verification rules: Human translations and rule-based LLM verifiers generally agree, while generic LLM judges and translation metrics can disagree with them.The gap between humans and the second-best model widens when annotators consider the verification rules.
- Evaluation with verification rules: Verifier pass rate produces more stable rankings across evaluator choices and agrees better with itself on a small data subsample.This indicates greater evaluation efficiency and stability than the compared approaches.
- Evaluation with verification rules: Verification-based evaluation shows much smaller model self-bias than typical LLM-as-a-judge evaluation.The authors suggest that verification rules may provide more objective assessment criteria.
- What is difficult to translate?: The source-oriented taxonomy describes translation difficulty through linguistic, extralinguistic, non-compositional, and atypical phenomena, including polysemy, cultural knowledge, metaphors, and language-variant specifics.It also identifies less widely addressed challenges such as meta-reasoning, internet cultural artifacts, and phonological phenomena.
4 Conclusion
The benchmark exposes difficult translation failures across languages and modalities, including distractor-driven errors, refusals, hallucinations, ambiguity loss, and instruction-related information loss. These examples support evaluation and analysis of model weaknesses.
- The benchmark includes hard examples that break strong models, supporting evaluation and further linguistic analysis.
- An English→German example requires paraphrasing under a length constraint, causing information loss in translation.
- All models fail on a Chinese Sign Language example when a textual distractor leads them away from the meaning “home.”
- Sandawe→Arabic produces refusal and hallucinations on a simple low-resourced example.
- Yoruba→English models translate a polysemous contraction into questions, contrary to the intended translation.
- Latin→German models generally fail to preserve syntactic attachment ambiguity, while Claude adds “will you die?” to clarify it.
Ethics Statement
The dataset release removes personally identifiable information, filters problematic submissions through peer review, and uses a CC BY 4.0 license.
- Personally identifiable information is removed before dataset release.
- Peer review filters explicit or otherwise problematic submissions.
- The benchmark is released under a CC BY 4.0 license.
- Third-party materials require source citation and license compatibility.
A Instructions
Contributors are asked to find difficult translation inputs, create accurate references, and specify pass/fail rules that define acceptable translations.
- Instructions for Contributors: Contributors find an input in any source and target language.
- Instructions for Contributors: The input should expose inadequate translations from major AI translation tools.
- Instructions for Contributors: Examples may involve tools such as ChatGPT or Google Translate.
- Instructions for Contributors: Contributors provide pass/fail rules for evaluating translation accuracy.
- Instructions for Contributors: The rules specify requirements that any accurate translation must satisfy.
- Instructions for Contributors: The workflow connects difficult examples with concrete verification criteria.
Step-by-step
The contributor workflow is presented as a sequence from language selection through submission.
- The workflow proceeds through selecting languages, providing input, writing a perfect translation, translating, writing verification rules, verifying, and submitting.
Inputs
Contributions may cover any subject and modality, but inputs should remain realistic and include a perfect translation that passes the verification rules. Ambiguous inputs should be clarified with context.
- Inputs can concern any subject, including news, social media posts, and website instructions.
- Realistic inputs are required so the benchmark tests failures where they matter.
- A perfect translation must demonstrate that the input can be translated and pass the verification rules.
- Context should disambiguate inputs such as “We looked for a match” by specifying whether the goal was starting a fire or playing a game.
Verification rules
Verification rules instruct an AI judge to assess concrete translation failures. They should generalize beyond a single erroneous output while covering observed or anticipated difficulties.
- Verification rules should test the observed failure generally rather than target a particular wrong translation.
- Rules should target tricky aspects of the input and cover issues observed in tested models or other credible risks.
- Each rule is written as a short instruction for an AI judge.
Examples
The examples show verification rules catching semantic mistranslations involving connotation, relational meaning, and lexical sense. Rules specify the intended distinction rather than merely flagging an incorrect wording.
- The Chinese rule requires “撒娇” to mean acting cute and adorable, without flirtatious or sexual connotations.
- The Hindi rules distinguish two slaps lexically and require their difference to express intensity rather than vertical position.
- The English example identifies “paper” as a domain-specific term that was mistranslated as a physical piece of paper.
Multimodal input and additional context
The benchmark supports multimodal inputs, additional disambiguating instructions, broad language coverage, and an ongoing contribution and review process. Contributors may use outside materials with attribution, but automated outputs are not accepted directly.
- Multimodal input and additional context: Images, audio, and video can provide translation inputs or supporting context, while additional instructions can specify properties such as target style.
- Multimodal input and additional context: Contributions are welcome in all languages, dialects, lesser-known languages, and specified scripts or regional variants.
- Contribution process: Contributors can become co-authors after 10 accepted submissions.
- Contribution process: Translation and verification actions consume credits, with possible increases for contributors whose submissions are good.
- Contribution process: Existing materials are permitted with source attribution, and the planned benchmark license is CC BY 4.0.
- Contribution process: Internet sources such as social media, messages, and advertisements can inspire difficult-to-translate inputs.
- Contribution process: Generative AI may provide ideas, but directly submitting LLM outputs is not accepted because of their poor quality.
- Contribution process: The benchmark name refers hyperbolically to saturation in existing machine-translation benchmarks and their limited guidance for further research.
B Prompts and Annotator Guidelines
The benchmark uses standardized prompts for translation, verification, generic scoring, and human error annotation. Its guidelines distinguish error severity, omissions, hallucinations, language errors, and consistency problems.
- Translation prompt: The translation prompt supplies source and target languages, optional multimodal context, source instructions, and requires translation-only output.The context may be an image, audio, or video.
- Generic scoring: The generic LLM judge assigns a single translation-quality score from 0 to 100 using defined qualitative bands.The bands range from complete meaning transfer and naturalness to partial transfer, frequent errors, and confusing omissions.
- Human annotation: Human annotators highlight translation errors, label severity, and score competing outputs while checking omissions, unsupported additions, language, and consistency.Major errors include meaning confusion or misrepresentation; missing text is marked separately.
- Verification prompts: The verifier generates concise, specific, true-or-false verification rules and returns only pass or fail for each criterion.Rules are intended to target concrete translation requirements rather than vague quality judgments.
C Expanding Last Translation Benchmark by transferring across languages
The benchmark can be expanded across target languages by updating rules and human translations, then filtering and manually reviewing transferred examples. Transfer works best when difficulty lies in the source, but target-language-specific rules often become invalid.
- Transfer procedure: For a new target language, an LLM updates the verification rules and creates a new human translation before automatic acceptance filtering.The human translation must pass, while at most two automatic translations may pass the benchmark criteria.
- Transfer quality: 48% of 100 examples transferred across Czech, Chinese, Farsi, Italian, and Hebrew survived the automatic filters and were then manually reviewed.Most manual rejections involved verification rules that no longer applied in the new target-language context.
- Transferability: Source-based difficulties such as words, idioms, constructions, garden-path sentences, cultural artifacts, and wordplay transfer more successfully across target languages.These examples preserve the underlying interpretation challenge when the target language changes.
- Expansion scope: 31% of transferred examples succeeded for individual target languages, while 17% succeeded across language pairs and 11% across groups of three languages.Pairwise overlap exceeded chance, supporting the finding that some examples are intrinsically more transferable.
- Transfer example: The Hindi idiom “unees bees ka phark” transfers successfully when its negligible-difference meaning is preserved in English and Czech.Models instead translated the idiom literally as a numerical difference in both target languages.