Source-linked AI summary
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger
TL;DR
Historical language-change research uses incompatible representations across linguistic levels, limiting direct comparison of their trajectories. ChronoLens provides a unified framework for measuring change across levels, languages, and periods, finding comparable within-language magnitudes but substantially different cross-language timing, magnitude, and direction.
Problem
Historical language-change studies use incompatible representations across morphology, syntax, semantics, and pragmatics, preventing direct comparison of whether these levels evolve together.
Method
ChronoLens uses feature-aligned crosscoders with frozen multilingual language models and post-hoc linguistic interventions to compare sparse features across languages, periods, and linguistic levels.
Results
Morphology, syntax, semantics, and pragmatics generally change by comparable amounts within languages, while languages differ substantially in timing, magnitude, and direction.
Takeaways & Limitations
Historical language change is structured and multidimensional, so meaningful cross-linguistic comparison requires measuring both distance and direction.
Takeaways & Limitations
The analysis is restricted to parliamentary discourse in five languages, and its four linguistic dimensions are operational simplifications that may obscure cross-level features.
Abstract
from arXiv · showhide
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce ChronoLens, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and apply it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026. The resulting sparse representations agree substantially more strongly with linguistic statistics than dense embeddings or a pooled sparse autoencoder ($ρ=0.72$ versus $0.29$ and $0.28$), and reveal that morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ markedly in when, how far, and in which direction they change. These findings show that historical language change is a structured, multidimensional process: similar magnitudes can conceal different trajectories, and meaningful cross-linguistic comparison requires measuring both distance and direction.
1 Introduction
ChronoLens addresses the difficulty of comparing historical change across linguistic levels by placing morphology, syntax, semantics, and pragmatics for five languages in a common analytical space. It combines a multilingual diachronic corpus with feature-aligned sparse representations to compare change across languages and periods.
- Motivation: Historical language-change studies use incompatible representations, datasets, and measurement units across linguistic levels, preventing direct comparison of their trajectories.
- Method: Sparse autoencoder features are non-canonical, so features learned in separate languages or periods have no guaranteed correspondence without an alignment method.
- Framework: ChronoLens compares morphology, syntax, semantics, and pragmatics across English, German, Italian, Polish, and Turkish within one analytical space.
- Method: Feature alignment combines frozen multilingual language models, crosscoders, and post-hoc probe interventions to compare sparse features across languages and periods without linguistic labels during feature learning.
2 Related Work
Prior computational studies of language change typically focus on individual linguistic levels, especially lexical semantics and syntax, while morphology is studied through productivity and structural complexity. Sparse autoencoders provide interpretable features, but crosscoders are needed to align features across models or checkpoints.
- Computational approaches to language change: Most computational work on diachronic language change focuses on lexical semantics, whereas syntactic change is commonly measured through dependency distance and structural complexity.The cited approaches include lexical-semantic studies and syntactic measures based on dependency distance and structural complexity.
- Computational approaches to language change: Morphological research has examined productivity, morphosyntactic complexity, and relationships involving morphological structure.
- Feature alignment: Sparse autoencoders recover interpretable, language-selective, and culturally selective features, while crosscoders learn shared feature indices across models or checkpoints.Independently trained dictionaries need not contain aligned features, motivating crosscoder-based alignment.
3 Dataset
ChronoLens uses a multilingual diachronic corpus of parliamentary and related political texts from five languages, standardized with shared temporal and document-level metadata. The corpus contains 44.98M documents and approximately 17.2B tokens spanning 1803–2026, with source-specific quality control to reduce OCR-related distortions.
- Sources and coverage: The corpus combines 22 open official and research sources across five languages, including parliamentary records and complementary political texts such as party manifestos.Its sources include long parliamentary records alongside smaller complementary collections.
- Sources and coverage: 44.98M documents and approximately 17.2B tokens span 1803 to 2026 in the final corpus.Shared temporal and document-level metadata enable metric comparison across languages.
- Quality control and density: OCR-derived Reichstag pages below the confidence threshold or with more than 15% low-confidence glyphs are removed, excluding about 8% of the oldest pages.Borndigital sources are screened with a character-n-gram gibberish detector because they lack OCR confidence scores.
4 CHRONOLENS
CHRONOLENS measures multilingual representation change across language, historical period, and linguistic level in a shared feature space. It combines frozen-model sentence representations, sparse crosscoder features, post-hoc linguistic attribution, and trajectory measures of magnitude and direction.
- Framework: CHRONOLENS connects language, historical period, and linguistic level to measure movement through shared feature space and compare aligned change trajectories.It measures displacement within languages, directional similarity between languages, and trajectory alignment across four linguistic levels.
- Linguistic attribution: Features receive linguistic-level assignments only when held-out ablations produce a sufficiently large dominant selective prediction drop; otherwise, they remain unassigned.Assignments summarize dominant effects rather than exclusive linguistic interpretations because levels are not mutually exclusive.
- Input construction: Matched sentence tuples are sampled by token length from parliamentary data, using 1950–2020 for cross-lingual comparisons and each language’s full record for within-language analyses.The matched sentences are neither translations nor paraphrases.
- Crosscoder training: The crosscoder learns shared sparse feature indices with separate language- or period-specific decoders from contextualized multilingual representations.Training uses 20k input tuples per decade and language, a dictionary twice the backbone hidden dimension, and a BatchTopK target active fraction of 0.10.
- Validation: CHRONOLENS evaluates reconstruction, trajectory quality, and linguistic specificity against direct embeddings, a pooled sparse autoencoder, and permuted-label probe controls.Linguistic specificity is interpreted relative to 0.25, which represents equal effects across the four linguistic levels.
- Historical change: Magnitude is the normalized Euclidean displacement from a language’s initial decade, while cosine similarity distinguishes parallel, opposing, and unrelated directions of change.A magnitude of 0 indicates no change; cosine values near 1, −1, and 0 indicate parallel, opposite, and unrelated directions, respectively.
5 Results
ChronoLens yields linguistically stronger representations than dense embeddings or a pooled SAE and reveals that historical change varies across languages in magnitude, timing, and direction. Equal magnitudes can mask distinct trajectories, while similar directions can occur despite different magnitudes or genealogical relationships.
- 5.1 Representation validity: Agreement with direct linguistic changes reaches ρ = 0.72 for the crosscoder, versus 0.29 for embeddings and 0.28 for the pooled SAE.Linguistic specificity also reaches 0.89, compared with 0.74 for embeddings and 0.73 for the pooled SAE.
- 5.1 Representation validity: Period-specific crosscoder features capture pragmatic parliamentary frames, including written Minister questions in P2 and adversarial oral questioning in P4.Feature 5236 activates across policy content, supporting a pragmatic frame rather than a migration subtopic.
- 5.2 Magnitude and temporal profile: German reaches the largest final magnitude at 0.89, English ends at 0.62, and Polish declines from approximately 0.53 around 1980 to 0.45 in the final decade.Italian and Turkish both end at 0.53 but follow different trajectories.
- 5.2 Magnitude and temporal profile: Across the common 1950–2020 interval, Turkish has the highest average magnitude at 0.61, followed by German at 0.43, English and Italian at 0.35, and Polish at 0.34.This ranking differs from the ranking based on each language’s full available record.
- 5.3 Magnitude and direction: Italian and Turkish share net magnitude 0.53 but differ in direction, whereas German and Turkish have different magnitudes, 0.89 and 0.53, but similar directions.English and Italian also have similar directions, while English and German are not the closest pair despite both being Germanic languages.
6 Conclusion
ChronoLens measures historical language change jointly across languages, periods, and linguistic levels. Its representations align more strongly with direct linguistic statistics, while findings show comparable within-language change across levels but substantial cross-language differences in magnitude and direction.
- 6 Conclusion: The crosscoder produces representations that agree more strongly with direct linguistic statistics while preserving stable historical trajectories.
- 6 Conclusion: Across five parliamentary traditions, morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ substantially in change magnitude and direction.
Limitations
ChronoLens’s four-way division into morphology, syntax, semantics, and pragmatics is an analytical simplification because linguistic levels overlap. Hard assignment prevents double counting but can obscure cross-level features, so trajectories represent operational dimensions.
- Analytical simplification: The four linguistic levels are not mutually exclusive, and individual phenomena or learned features may span several levels.This makes the morphology–syntax–semantics–pragmatics division an analytical simplification.
- Analytical simplification: Hard assignment uses a feature’s dominant probe effect to place it in one level and prevent double counting.The assignment is operational rather than a claim that each feature belongs exclusively to one linguistic level.
- Analytical simplification: Hard assignment can obscure genuinely cross-level features, so the resulting trajectories should be interpreted as changes along four operational dimensions.
Broader Impact
ChronoLens offers a common framework for comparing historical language change across languages and linguistic levels, while its parliamentary-records basis limits conclusions about whole populations.
- Research utility: ChronoLens may support research in computational linguistics, political science, history, and the digital humanities by enabling common comparisons across languages and linguistic levels.The framework is designed for comparing historical change within a shared analytical approach.
- Limitations: Parliamentary records capture institutional discourse produced by political actors rather than language use across entire populations.Therefore, cross-linguistic similarities should not be interpreted as essential properties of national communities.
Ethical Considerations
The analysis uses publicly available parliamentary and political texts, reporting aggregate language–period patterns rather than predictions about individual speakers. Because records may include identifiable speakers and sensitive issues, derived-data releases should preserve attribution, licensing, and applicable restrictions.
- The analysis relies on publicly available parliamentary and political texts from official and research sources.
- Reports aggregate language–period patterns rather than predictions about individual speakers.
- Derived-data releases should preserve source attribution, licensing conditions, and applicable restrictions because records may identify speakers or discuss sensitive political and social issues.This means derived outputs should not redistribute source material indiscriminately.
A Dataset Sources … C.2 Backbones, representations, and layer selection
ChronoLens builds cross-lingual and temporal comparisons from uniformly structured parliamentary records, carefully matched and masked concept sentences, and fixed multilingual-model representations. Its sampling, layer selection, and data-addressing rules are designed to make linguistic comparisons reproducible and reduce confounding variation.
- B Dataset examples: 22 sources are normalized into one 16-field record schema, giving language, year, and document type the same meaning across collections.This makes metrics comparable across parliamentary traditions because they operate on the same kind of record.
- B Dataset examples: Granularity distinguishes speeches, segments, and sessions, while inherited source identifiers are document identifiers rather than primary keys.Session-structured sources may assign one sitting identifier to every speech, so records must be addressed by row.
- C.1 Concept matching and sentence sampling: 13 concepts are divided into target, general political-discourse, and frequency-matched control groups to distinguish concept-specific change from broader parliamentary or corpus effects.Migration and gender are targets; road and water serve as controls.
- C.1 Concept matching and sentence sampling: Language-specific lowercased stem lexicons use longest-first prefix matching, capturing inflected forms while risking unintended stem matches.The full lexicons are released and ambiguous terms are inspected during error analysis.
- C.1 Concept matching and sentence sampling: Each sentence is assigned to at most one concept, with the first matched concept winning when multiple concepts occur.Thus concept counts partition harvested sentences rather than estimate total corpus frequency.
- C.1 Concept matching and sentence sampling: Concept expressions and inflectional endings are masked before policy-frame labeling and matching, and excluded from activation pooling.These operations prevent the concept’s surface form from determining either matching strata or model representations.
- C.1 Concept matching and sentence sampling: Cross-lingual and temporal tuples are matched within shared token-length and policy-frame strata, while sampling retains at most 400 sentences per concept, language, and decade.Reservoir sampling uses a fixed seed, and decade conditions require at least 250 distinct matched sentences.
- C.2 Backbones, representations, and layer selection: Four frozen multilingual backbones provide mean-pooled residual-stream representations, with concept tokens excluded when present.The backbones are Qwen3-8B, Llama-3.1-8B, Mistral-Nemo-2407, and EuroLLM-9B-2512; extraction uses one no_grad forward pass and fp16 storage.
C.3 Crosscoder training and calibration … D.1 Cross-language agreement
ChronoLens trains and calibrates crosscoders across languages and periods, then combines conservative feature interventions, probe diagnostics, and trajectory analyses to compare observable and representational language change. The results show substantial cross-language agreement in pragmatics but weaker or direction-dependent agreement in morphology, semantics, and syntax.
- C.3 Crosscoder training and calibration: Crosscoders train on matched multilingual or historical tuples, using 20,000 samples split into 18,000 training and 2,000 validation tuples.Each tuple contains one representation from every language or available period, and sampled and distinct-sentence counts are recorded.
- C.3 Crosscoder training and calibration: Crosscoder conditions are standardized separately, and shared-versus-specific features are calibrated with decoder-norm rules, latent-scaling checks, held-out FVU, and a condition-shuffled null.A real crosscoder resolves conditions only when it has at least twice as many specific features as the null and meets minimum feature coverage.
- C.4 Feature attribution and interventions: Feature attribution uses the 1,000 highest-activation-mass candidates and conservative intervention criteria, while diagnostics remain evidence rather than filters for the main trajectories.The main thresholds are an absolute probability decrease above 0.01 and at least 1.2 times the second-largest decrease; threshold sensitivity preserves the assigned inventory and convergence findings.
- C.5 Probe diagnostics: Probe diagnostics report majority, accuracy, and selectivity separately because class imbalance makes raw accuracy incomparable across language–task pairs.Speech-act accuracy is 0.95–0.97 while its majority class accounts for 0.84–0.96, illustrating why selectivity is necessary.
- C.6 Additional trajectory and statistical details: Trajectory analysis averages shared-feature activations within adequately populated cells and distinguishes convergence from directional alignment using net-displacement and stepwise cosine measures.Cells require at least 25 sentences, while stepwise alignment labels pairs parallel above 0.20 and anti-parallel below −0.20.
- D Observable linguistic changes: Across 90 observable language–measure trends from 1950–2020, 37 remain significant after Benjamini–Hochberg correction, with personal deixis increasing in all five languages.Other measures reverse direction across languages, including long dependencies, left-branching, and coordination.
- D.1 Cross-language agreement: Cross-language agreement is strongest for pragmatics, with mean cosine similarity 0.92 and positive bootstrap intervals for all ten language pairs, while morphology and semantics are weaker at 0.21 and 0.29.Pragmatic agreement is driven mainly by personal deixis, although leave-one-measure-out means remain positive from 0.53 to 0.99.
- D.1 Cross-language agreement: Syntax has mean cosine similarity 0.00, with four significant positive and five negative pairwise similarities, and language-family membership alone does not explain the broader pairwise patterns.For Indo-European versus Turkish-involving pairs, syntax means are 0.12 and −0.18, whereas pragmatics means are 0.89 and 0.95.
D.2 Long-window results
Long-window analysis reports the largest significant trends over each language’s complete historical interval, extending beyond the balanced 1950–2020 comparison. Because intervals differ, the fitted percentage-point changes per century are supplementary and should not be compared as total changes.
- Long-window results: The common 1950–2020 interval enables direct cross-language comparison but omits available earlier records for English, German, Italian, and Polish.Table 13 addresses this limitation by using each language’s complete available interval.
- Long-window results: Table 13 reports each language’s largest significant observable changes across its complete historical record.All listed trends have q < 0.05.
- Long-window results: The long-window values are fitted percentage-point changes per century, not total historical changes.Unequal historical intervals make this analysis supplementary to the balanced 1950–2020 comparison.