Source-linked AI summary
DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
Mariia Fedorova, Andrey Kutuzov, Khonzoda Umarova
TL;DR
Lexical semantic change research lacks large, permissively licensed diachronic corpora across many languages. DHPLT constructs standardized corpora from HPLT web crawls for 41 languages and three periods, adding pre-computed representations and frequencies. The resource demonstrates diachronic semantic trajectories while retaining crawl-time and representation-scope limitations.
Problem
Lexical semantic change research has appropriate diachronic corpora for only a small set of mostly high-resource languages, limiting multilingual coverage.
Method
DHPLT samples standardized corpora from HPLT web crawls across three periods for 41 languages, using crawl timestamps and providing pre-computed embeddings, lexical substitutions, and frequency counts.
Results
The reported English and Spanish examples show semantic trajectories for AI-related terms across the three periods, with the largest change observed for “ai” and corresponding Spanish findings.
Takeaways & Limitations
DHPLT provides an open multilingual resource for experimenting with lexical semantic change across languages, periods, target words, and representation types.
Takeaways & Limitations
Web crawl timestamps give only an upper bound on document creation time, so later-period subsets can contain text created earlier.
Abstract
from arXiv · showhide
In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time. The collection covers three time periods: 2011-2015, 2020-2021 and 2024-present (1 million documents per time period for each language). We additionally provide pre-computed word type and token embeddings and lexical substitutions for our chosen target words, while at the same time leaving it open for the other researchers to come up with their own target words using the same datasets. DHPLT aims at filling in the current lack of multilingual diachronic corpora for semantic change modelling (beyond a dozen of high-resource languages). It opens the way for a variety of new experimental setups in this field. All the resources described in this paper are available at https://data.hplt-project.org/three/diachronic/, sorted by language.
1 Introduction
DHPLT addresses the shortage of permissively licensed, multilingual diachronic corpora for lexical semantic change research. It provides standardized resources across 41 languages, with target-word representations for immediate experimentation and original texts for further study.
- Diachronic semantic change modelling requires texts associated with their creation dates so word usage can be compared across periods.
- Existing lexical semantic change resources cover only a small set of mostly high-resource languages, with limited public availability.The field has at most about a dozen languages, with Indo-European languages strongly over-represented.
- DHPLT provides standardized diachronic corpora for 41 languages spanning 12 language families and three time-dependent subsets.Each language has three subsets containing 1 million documents, extracted from cleaned and filtered HPLT web pages using crawl timestamps for temporal separation.
- DHPLT supplies static word2vec embeddings, token embeddings, and lexical substitutions for selected target words, while preserving the original texts for new target-word sets.The pre-computed representations reduce the compute needed to begin multilingual LSCD experiments.
2 Diachronic corpora out of HPLT
DHPLT constructs multilingual diachronic corpora from cleaned web crawls because standardized historical corpora are unavailable for many languages. It separates documents into three crawl-time periods while retaining timestamps and reproducibility, but crawl dates provide only upper bounds on document creation.
- Historical diachronic resources are nearly nonexistent in standardized form for most of the world’s languages, motivating the use of web data.HPLT processes web crawls with language identification, deduplication, and cleaning, and publishes them under CC0.
- Web crawl timestamps cannot guarantee document creation dates because a newly crawled page may contain older text.They do guarantee that the text was not created later than the crawl timestamp.
- The three periods are 2011-2015, 2020-2021, and 2024, chosen to support multipoint change analysis with comparable data volumes and temporal gaps.The periods are intended to capture linguistic innovation or semantic-change onset, while alternative splits remain possible.
- The selected 41 languages meet data and model-availability criteria, including at least 0.5 million documents per period where possible and a corresponding monolingual T5 model.The final language set and document counts are listed in Appendix Table 1.
- For each language and period, DHPLT randomly samples up to 1 million documents and publishes approximately 170 GB containing about 59 billion words.Documents are distributed as zstd-compressed JSONL files in HPLT format.
- The corpora are directly usable for multilingual LSCD, and the released representations illustrate experimental setups that can be built from DHPLT.
3 Target word selection
DHPLT narrows each language’s vocabulary to linguistically relevant target words for which semantic representations can be produced. The selection pipeline filters model vocabulary items and then groups related forms through lemmatization.
- Target-word selection narrows the full vocabulary while retaining words likely to interest lexical semantic change researchers.
- The pipeline removes word pieces and infrequent tokens from each language’s T5 vocabulary, retaining nouns, verbs, and adjectives in the main script.
- The resulting target-word sets contain approximately 18,600 words per language on average, compared with T5 vocabularies of 32,768 items.
- Lemmatization merges distinct word forms such as “thread,” “Thread,” and “threads” into a shared lemma for later representation grouping.
4 Target word representations
DHPLT provides contextual and static semantic representations, frequency counts, and aligned time-specific word spaces for selected target words. These resources support multiple LSCD workflows while reducing the computational effort required to recreate representations.
- The released representations can be directly used to evaluate or train semantic change models across the three DHPLT periods.
- DHPLT includes sampled contextual token embeddings from T5, XLM-R, and GPT-BERT models, plus lexical substitutes from GPT-BERT and XLM-R.The sampling schemes include 1,000 T5/XLM-R occurrences or 100 GPT-BERT occurrences per target word, and 100 GPT-BERT occurrences represented by top-15 substitutes.
- Lexical substitutes are included as an alternative contextual representation for semantic-change quantification and discovery tasks.
- Static word embeddings provide one vector per word type and remain useful because of their simplicity and modest training and inference costs.
- 4.3 Word type embeddings: Each language-period combination receives a 300-dimensional SGNS word2vec model with a 50,000-word vocabulary after document preprocessing.The models use window size 10, five epochs, and five negative samples.
- 4.3 Word type embeddings: Earlier-period static embedding spaces are aligned to the 2024 space with standard Procrustes alignment for cross-period similarity comparisons.
- 4.4 Frequency counts: Frequency counts across periods support initial usage-change inspection, frequency-effect control, and compute planning.
5 Sanity check
DHPLT sanity checks show coherent cross-period semantic drift for AI-related words in English, Spanish, and Russian, while encoder distances distinguish stronger and weaker changes.
- Semantic change examples: English ‘AI’ shifts from video-game characters in the 2010s to chatbots and machine learning in 2020-2021, then LLMs, ChatGPT, and generative AI in 2024-.The trajectory remains visible despite overlap caused by crawl timestamps.
- Semantic change examples: Spanish ‘IA’ similarly moves from gaming contexts to algorithms and technologies, then generative AI and ChatGPT.The 2020-2021 neighbours include semantically related English words such as ‘AI’ and ‘learning’.
- Semantic change examples: Russian DHPLT embeddings exhibit very similar semantic trends to the English and Spanish examples.
- Embedding-distance check: The largest T5 encoder-distance change is for ‘ai’, whereas conservative legal terms ‘legislative’ and ‘jurisdiction’ change least.‘Remote’ changes intermediately, with its largest shift between 2011-2015 and 2020-2021 as it becomes associated with remote work.
6 Conclusion
DHPLT provides an open, multilingual diachronic resource spanning 41 languages and three periods, supplemented with semantic representations for experimentation. It addresses the shortage of multilingual corpora available for lexical semantic change research.
- Resource: DHPLT is an open collection of large-scale diachronic corpora covering 41 languages from 12 language families.
- Resource: The corpus uses HPLT v3.0 web-crawl timestamps as its temporal signal across 2011-2015, 2020-2021, and 2024.
- Representations: DHPLT adds pre-computed token-level representations and aligned static word embeddings for language-specific target words and periods.
- Significance: The resource partially addresses the lack of multilingual diachronic corpora in lexical semantic change detection.
- Significance: DHPLT is intended to make historical language-change modelling richer and more diverse, especially for semantic change discovery.
Limitations
DHPLT’s main limitations concern approximate temporal boundaries, incomplete representation coverage, and the paper’s lack of full-scale semantic change discovery experiments.
- Temporal signal: Web-crawl timestamps provide only an upper bound on document creation time, so a document may contain text created earlier than its crawl period.Documents cannot contain text created later than their timestamp under the stated assumption.
- Representation coverage: DHPLT supplies only some representation types for selected target-word sets rather than all words, due to compute and storage constraints.
- Evaluation scope: The paper introduces the DHPLT dataset but leaves full-scale semantic change discovery on it for future work.
A DHPLT datasets
The DHPLT datasets organize HPLT web documents by crawl period and expose document metadata, segment quality scores, language statistics, and filtered target-word inventories for 41 languages.
- Dataset composition: Figure 1 reports document counts by crawl year for English and Georgian in HPLT v3.0.
- File structure: Each DHPLT file stores one document per line with an identifier, crawl timestamp, segmented text, and segment-level quality scores.The scores are assigned using HPLT’s Web Docs Scorer heuristics.
- Dataset statistics: DHPLT provides language, writing-system, family, and historical-period size statistics for its 41 languages.
- Target-word selection: Target-word selection starts from each language’s HPLT T5 vocabulary and excludes word pieces, non-words, and insufficiently frequent terms.A retained token must occur as a full word at least 10 times in each of the three periods.
- Target-word selection: The target set is restricted mainly to nouns, verbs, and adjectives, with language-specific tokenization and additional single-character and script filters.
- Target-word statistics: Figure 2 distributes target-word counts across languages overall and separately for nouns, verbs, and adjectives.
C T5 substitutions
This section explains why T5 models are unsuitable for generating lexical substitutions and describes the resulting representation choices.
- C T5 substitutions: Encoder-decoder T5 models duplicate compute usage because both encoder and decoder components are required.
- C T5 substitutions: Generated predictions may repeat sentence terms instead of representing the target word’s semantics.
- C T5 substitutions: T5 predictions can produce long continuations from real-world documents rather than concise lexical substitutes.
- C T5 substitutions: Tables 2 and 3 report average pairwise distances for English and Spanish target words using T5 encoder embeddings.
E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions
The GPT-BERT substitution sanity check shows semantic trajectories changing across periods, with AI-related terms moving from technical associations toward social and human-related ones.
- E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions: AI substitutions shift from non-technical, gaming, or automotive associations in 2011-2015 toward diverse technologies in 2020-2021.Examples in the middle period include IoT, NLP, robotics, and animation.
- E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions: In 2024, AI substitutions continue reflecting social consequences, with more human-related terms and fewer technical terms.Examples include elite, censorship, communism, scammers, capitalism, art, and healthcare.
- E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions: Remote substitutions move from geographic distance in 2011-2015 to virtual contexts in 2020-2021 and technology, jobs, and society in 2024.Later associations include skilled, flexible, professional, satellite, healthcare, climate, and rural.
- E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions: Jurisdiction and legislative substitutions remain related to law across all three time periods.
- E.2 Sanity check of HPLT 3.0 GPT-BERT substitutions: Contextualized representations capture more fine-grained semantic nuances than static word-embedding models because they are sensitive to prediction contexts.
E.3 Sanity check for SWEs
This sanity check presents static word-embedding trajectories for AI terms in English, Spanish, and Russian across the DHPLT periods.
- E.3 Sanity check for SWEs: Table 4 lists the top five cosine-similarity neighbours of English ‘AI’ for each DHPLT time period.Case is ignored in the table.
- E.3 Sanity check for SWEs: Table 5 lists the top five cosine-similarity neighbours of Spanish ‘IA’ across the three DHPLT time periods.‘IA’ stands for ‘inteligencia artificial’.
- E.3 Sanity check for SWEs: Table 6 lists the top five cosine-similarity neighbours of Russian ‘ИИ’ across the DHPLT time periods.The 2011-2015 model lacks ‘ИИ’ because its frequency was too low.