Source-linked AI summary
How Much Does Corpus Choice Change Dependency-Distance Estimates?
Sirui Chen
TL;DR
This study tests whether dependency-distance estimates generalize across independently sourced treebanks of the same language. Using agreement analyses and a twelve-specification design, it finds that corpus choice changes ordinal rankings while dependency-length minimization remains consistent across corpora.
Problem
Single-corpus dependency-distance estimates are routinely treated as representative of a language, although their generalizability across independently compiled treebanks remains untested.
Method
The study compares same-language UD treebank pairs using concordance correlation, Bland–Altman analysis, sampling-stability checks, and a twelve-specification multiverse design.
Results
Nearly 40% of pairwise language rankings reversed across treebanks, while all 76 treebanks had normalized ratios below 1.0, confirming dependency-length minimization.
Takeaways & Limitations
MDD is more consistent with a corpus-conditioned composite shaped by grammatical, register, annotation, and source factors than with a stable language-level parameter.
Takeaways & Limitations
The design cannot separate annotation practice, register, time period, and translation status, and the sample is dominated by Indo-European languages.
Abstract
from arXiv · showhide
Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, yet this assumption has not been tested across independently compiled corpora. We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, using concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design. Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance. This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications. Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1). The data are more consistent with MDD as a corpus-conditioned composite of grammatical, register, and annotation factors than as a stable language-level parameter: the qualitative DLM universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.
1 Introduction
The study asks whether dependency-distance estimates generalize across independently sourced treebanks of the same language. It tests agreement, sampling variation, sentence-length effects, and preprocessing sensitivity to distinguish stable language properties from corpus-conditioned estimates.
- Motivation: Single-corpus dependency-distance estimates may not represent a language because MDD varies with sentence length, genre composition, and annotation conventions.The study therefore treats corpus choice as a substantive measurement issue rather than mere convenience.
- Motivation: If MDD reflects stable grammatical constraints, distinct treebanks should yield concordant estimates; if it is composite, agreement should depend on shared sampled components.Cross-treebank agreement discriminates between these competing predictions.
- Research gap: Split-half stability cannot establish cross-corpus agreement because subsamples share source selection, genre composition, and annotation history.The closest precedent did not focus on MDD.
- Research questions: The study evaluates agreement across same-language treebanks, disagreement relative to within-treebank sampling variation, sentence-length matching, and sensitivity to preprocessing choices.The research questions cover raw and normalized estimates and declared choices about punctuation, sentence length, and aggregation.
- Approach: The design treats distinct-source treebanks as repeated measurements and combines concordance correlation, Bland–Altman analysis, and a specification-curve design.The paper reports a reproducible pipeline and tests pre-declared hypotheses without changing the estimand after null or contrary results.
- Contribution: Cross-treebank agreement is moderate at best, with disagreement exceeding within-treebank sampling error and reversing nearly 40% of pairwise language rankings.The paper interprets MDD as a corpus-conditioned composite whose qualitative properties survive corpus substitution but whose ordinal rankings do not.
2 Data
The dataset combines independently sourced Universal Dependencies treebanks after eligibility filtering and source-overlap auditing. The resulting sample contains 38 language-level pairs, but its family composition is heavily skewed toward Indo-European languages.
- Data sources: The study uses Universal Dependencies v2.18 and Glottolog CLDF 5.3, retrieving treebanks from a checksum-verified official archive.Each data source is recorded with its URL, version, license, size, and cryptographic digests.
- Eligibility: Eligibility screening requires sentences to meet structural and length criteria, including 5–40 non-punctuation syntactic words, one basic-dependency root, integer token identifiers, and connected acyclic trees.The initial metadata screen requires at least 500 sentences and at least two treebanks per language.
- Source audit: Source independence is operationalized using normalized sentence-text hashes and sentence identifiers to group potentially overlapping same-language treebanks.Pairs enter the same source group at 5% text-hash overlap or 80% sentence-ID overlap with a sentence-count ratio between 0.95 and 1.05.
- Source audit: The source-independence procedure is a screen rather than proof that corpora differ in register, time period, or annotation guidelines.The 5% overlap threshold is treated as an operational default and varied from exact overlap through 20%.
- Final sample: The final dataset contains 38 language-level pairs comprising 76 treebanks from distinct source groups across 10 language families.Indo-European languages account for 24 of 38 pairs, limiting genealogical balance.
3 Measures and Analysis
The analysis defines raw and normalized dependency-distance measures and evaluates agreement, sampling stability, sentence-length composition, and preprocessing sensitivity. Multiple complementary statistics and pre-declared specifications are used to separate corpus disagreement from finite-sample and analytic variation.
- Dependency-distance metrics: Raw MDD averages absolute head–dependent position differences across retained dependency arcs.It is the study’s primary unnormalized dependency-distance summary.
- Dependency-distance metrics: Normalized dependency distance divides MDD by the exact random-order expectation, (n + 1)/3, to separate structural signal from sentence-length baseline.The analytic expectation avoids Monte Carlo error.
- Agreement framework: Agreement is assessed with CCC, ICC, Bland–Altman limits, and rank correlation, with each statistic addressing a different aspect of paired treebank estimates.CCC measures absolute agreement, ICC addresses range restriction, and Bland–Altman limits summarize practical disagreement.
- Sampling stability: Sampling stability draws 50, 100, 250, and 500 sentences repeatedly to quantify how much each treebank estimate fluctuates under finite samples.Complete-document sampling is used when document metadata are sufficiently available to preserve shared topic, register, and syntactic priming effects.
- Sentence-length matching: Sentence-length matching compares unmatched absolute disagreement with disagreement from 1,000 draws balanced across four sentence-length bins.The matched comparison is capped at 20 words and can exclude pairs lacking adequate bin coverage.
- Sensitivity analyses: The pre-declared multiverse crosses punctuation treatment, 20-, 40-, or 80-word caps, and sentence- or token-weighted aggregation into twelve specifications.Additional diagnostics vary sample composition, source-overlap thresholds, genre comparability, and treebank substitution.
4 Results
Cross-treebank dependency-distance agreement was moderate at best and reflected source-level differences beyond sampling error or sentence-length composition. Corpus choice altered language rankings substantially, while the qualitative dependency-length-minimization result remained robust.
- 4.2 Cross-treebank agreement: CCC was 0.391 for raw MDD and 0.508 for the normalized ratio under the canonical specification, indicating moderate cross-treebank agreement.Mean differences were 0.194 words for raw MDD and 0.075 for the normalized ratio.
- 4.4 Genealogical and sampling-frame diagnostics: The normalized ratio was below 1.0 for all 76 primary treebanks, ranging from 0.26 to 0.76, so corpus choice changed rankings but not the qualitative minimization result.The agreement estimates were also robust to overlap-threshold changes and residual shared sentences, but sensitive to genealogical weighting and treebank representation.
- 4.3 Decomposing disagreement: sampling error and sentence length: 0.166 raw-MDD units and 0.067 normalized-ratio units separated observed disagreement from 500-sentence joint sampling error, with both intervals excluding zero.Finite-sample noise therefore did not account for the between-treebank differences.
- 4.3 Decomposing disagreement: sampling error and sentence length: 0.040 normalized-ratio units of disagreement were removed by sentence-length matching across 30 language pairs, while the raw-MDD reduction of 0.033 had an interval including zero.The normalized-ratio reduction was roughly half the baseline ratio MAE of 0.075.
- 4.3 Decomposing disagreement: sampling error and sentence length: 29% of between-group raw-MDD variance was attributable to treebank level, while normalized-ratio treebank variance became effectively zero after sentence-length adjustment.Residual disagreement remained consistent with register, source population, annotation practice, or their interaction.
- 4.4 Genealogical and sampling-frame diagnostics: 39.5% of raw-MDD and 28.7% of normalized-ratio pairwise language orderings reversed when substituting the second-largest treebank.The raw-MDD mean inter-treebank shift of 0.194 words was comparable to the between-language interquartile range of approximately 0.28 words.
5 Discussion
The discussion interprets moderate cross-treebank agreement as evidence that MDD is corpus-conditioned, while normalization removes sentence-length variation without resolving register and annotation effects. Consequently, DLM remains stable within languages, but cross-linguistic rankings are sensitive to corpus and analytical choices.
- Why normalization does not clearly improve agreement: Normalization removed sentence-length variation but left register and annotation effects intact, so it could not resolve disagreement across corpora.After length adjustment, normalized-ratio treebank variance effectively vanished, whereas raw-MDD treebank variance remained 0.017; matching lengths reduced ratio disagreement by 0.040.
- Sensitivity to preprocessing specifications: Preprocessing choices shifted agreement estimates substantially, with specification spreads of roughly 0.16 CCC units for raw MDD and 0.26 for the normalized ratio.No specification produced CCC values above 0.52, and punctuation treatment reversed direction at the 80-word cap for raw MDD.
- Implications for comparative research: Every one of the 76 treebanks yielded a normalized ratio below 1, preserving evidence for dependency-length minimization across corpus substitutions.The discussion characterizes DLM as an ordinal fact about each language individually rather than a cross-language comparison.
- Implications for comparative research: Treebank choice reversed nearly 40% of pairwise language rankings, and treebank-level variance accounted for roughly 29% of between-group variance in raw MDD.Mean inter-treebank shifts were comparable to the between-language interquartile range, concentrating reversals among closely spaced languages.
- Limitations: The composite interpretation remains circumstantial because the design confounds grammar, register, annotation, time period, and translation status.A factorial design crossing annotation team with register within a language would be needed to decompose these sources directly.
- Implications for comparative research: Comparative studies should use multiple treebanks per language, report estimate ranges, and pair treebank-derived conclusions with specification-curve analyses.The paper recommends Bland–Altman limits as an empirical prior on corpus-conditioned variation when only one treebank is available.
6 Conclusion
Treebank substitution preserves the qualitative dependency-length-minimization pattern but destabilizes raw cross-linguistic rankings. The findings favor MDD as a corpus-conditioned composite rather than a stable language-level parameter.
- Nearly 40% of pairwise raw-MDD language rankings reversed when one UD treebank was substituted for another.
- All 76 treebanks confirmed dependency-length minimization, with normalized ratios below 1.
- MDD is better characterized as a corpus-conditioned composite shaped by register, annotation, source characteristics, and grammatical constraints.
- Single-corpus estimates should not be treated as representative of a language without justifying that corpus-specific factors are negligible for the comparison.
Data and Code Availability
The reproduction package provides the materials needed to inspect and rerun the study, subject to corpus-license restrictions on raw text redistribution.
- The reproduction package includes source metadata, versions, licenses, checksums, inclusion decisions, derived metrics, analysis code, and result objects.
- Raw corpus text is redistributed only where its license permits.