Source-linked AI summary
DGT-TM: A freely Available Translation Memory in 22 Languages
Ralf Steinberger, Andreas Eisele, Szymon Klocek, Spyridon Pilos, Patrick Schlüter
TL;DR
The European Commission and Joint Research Centre are expanding public access to multilingual translation data beyond its primarily legal source texts. The paper describes how DGT-TM Version 2011 was produced and reports its scale, uses, and evaluation boundaries.
Problem
Existing European multilingual text collections are primarily legal, but their broader practical and language-technology uses motivate making sentence-aligned translations publicly available.
Method
The paper describes extending DGT-TM with sentence-aligned translations from EU documents, using an alignment tool tuned to European Commission and EU document features.
Results
About 38 million translation units were produced across 22 official EU languages, with alignment quality found very good and only few errors in production testing.
Takeaways & Limitations
DGT-TM can support translation practice and language-technology applications, while its free reuse enables dissemination for commercial and non-commercial purposes.
Takeaways & Limitations
Exact numbers for alignment quality evaluation are unavailable, so the reported quality assessment is based on production testing and translator-reported errors.
Abstract
from arXiv · showhide
The European Commission's (EC) Directorate General for Translation, together with the EC's Joint Research Centre, is making available a large translation memory (TM; i.e. sentences and their professionally produced translations) covering twenty-two official European Union (EU) languages and their 231 language pairs. Such a resource is typically used by translation professionals in combination with TM software to improve speed and consistency of their translations. However, this resource has also many uses for translation studies and for language technology applications, including Statistical Machine Translation (SMT), terminology extraction, Named Entity Recognition (NER), multilingual classification and clustering, and many more. In this reference paper for DGT-TM, we introduce this new resource, provide statistics regarding its size, and explain how it was produced and how to use it.
1. Introduction and Motivation
The European Commission is expanding public access to multilingual language resources through a larger annual DGT-TM release. The resource derives from sentence-aligned EU text collections and supports translation technology and related language applications.
- DGT-TM Release 2011 contains documents published from 2004 to 2010 and is twice as large as the 2007 release.
- The release complements publicly accessible EU resources including EUR-Lex, IATE, and EuroVoc.These resources provide access to EU law, inter-institutional terminology, and a multilingual thesaurus.
- DGT-TM originated as legal text and translation data but has practical uses beyond the legal domain.
- DGT’s information technology and language applications units sentence-aligned EU full-text collections and added them to the translation memory for public release.
- The paper surveys multilingual translation-memory uses, describes DGT-TM production, and provides practical usage details.
2. Possible uses of DGT-TM
Highly multilingual parallel resources support machine translation and a broad range of multilingual language-technology tasks. DGT-TM is especially valuable because such collections remain relatively scarce.
- Highly multilingual parallel collections are relatively scarce, with EuroParl, DGT-TM, and JRC-Acquis among the most multilingual resources.The 2011 EuroParl update had 21 languages, while DGT-TM and JRC-Acquis had 22 each.
- Parallel data is crucial for creating Statistical Machine Translation models.EuroParl enabled systems for up to 110 language pairs, and its multilingual data supported projects such as EuroMatrix.
- The JRC-Acquis publication enabled Statistical Machine Translation systems for 462 European language pairs.
- Parallel sentence collections support multilingual lexical and semantic resources, annotation projection, plagiarism detection, clustering, classification, and semantic-space creation.
3. Details on DGT-TM-2011
DGT-TM-2011 was constructed from officially translated EU legislation by segmenting and aligning documents into multilingual translation units. The release contains about 38 million translation units across 22 languages, with alignment quality judged very good in production testing.
- DGT imported official EU documents, automatically sentence-aligned their full texts, and evaluated translation quality, alignment, and corpus statistics.
- The dataset covers L-Series EU legislation published from 2004 through 2010, excluding documents already present in the 2007 release.
- The source translations were produced by specialized human translators and checked through linguistic, legal, service-level, and Publications Office review.
- Translation units include sentences, titles, headings, and sentence parts separated by colons or semicolons.Legal texts frequently use semicolons to separate longer sections.
- The alignment tool uses numbering to define zones, then strengthens alignments with numbers, images, and other non-linguistic clues.When external clues are absent, it uses character-count statistics for typical relative sentence lengths across languages.
- Exact alignment-evaluation numbers are unavailable, but translators tested the automatically aligned translations in production.
- Translators reported alignment errors for algorithm improvement, and the resulting alignment quality was judged very good with few errors.
- About 38 million translation units were produced across 22 official EU languages, averaging about 1.9 million per language.Bulgaria, Romania, and Malta have fewer units for stated accession and translation-obligation reasons; Irish was excluded.
4. Downloading and using DGT-TM
DGT-TM-2011 is distributed in TMX packages with accompanying extraction software for producing parallel sentence collections. Users must account for package completeness, extraction-order and redundancy issues, licensing conditions, and the absence of guarantees regarding data quality.
- Downloading DGT-TM: The resource is distributed as 25 zip packages containing TMX files identified by EUR-Lex document identifiers and plain-text language lists.The extraction program can access the data directly without requiring users to unzip the packages.
- Downloading DGT-TM: Users must download all packages to obtain the complete parallel corpus, because language data are distributed across the zip files.Downloading only a subset produces only a subset of the parallel corpus.
- Downloading DGT-TM: TMX files use UTF-16 Little Endian encoding and are compliant with TMX 1.4b despite headers that mention TMX 1.1 for backwards compatibility.TMX is described as a widely used translation-memory format.
- Extracting parallel sentences: TMXtract must be copied beside the zip files to extract parallel sentence collections for any of the 231 language pairs.The tool is available as graphical and command-line versions for different operating systems, and its sequence of extracted files may differ from the source documents.
- Conditions of use: DGT-TM may be freely reused and disseminated under the Commission’s re-use conditions, while the accompanying software is licensed under GPL 2.0.Re-users must state the source, update date, and the Commission’s retained ownership.
- Conditions of use: The database and software are provided without guarantees, and the Commission accepts no liability for alignment quality, data correctness, or reuse consequences.These conditions are stated alongside a reference to the DGT-TM website’s terms of use.
5. Summary and future work
DGT and JRC released a larger multilingual DGT-TM collection covering EU legislative documents from 2004–2010, with annual releases planned. Because sentence-level translation memories do not preserve full-document order, a full-text parallel version is also planned.
- Summary: The 2012 DGT-TM Version 2011 release contains sentences and translations in up to 22 languages from EU Official Journal L-Series legislation published between 2004 and 2010.The release is larger than the first version, and future TM releases are planned annually.
- Future work: Annual future releases of the translation memory are planned.
- Future work: Because translation memories contain individual sentences and fragments rather than document order, a full-text parallel release is planned for purposes requiring the original text sequence.