Source-linked AI summary

Systematic Inequalities in Language Technology Performance across the World's Languages

Damián Blasi, Antonios Anastasopoulos, Graham Neubig

arXiv:2110.06733v1cs.CL

TL;DR

NLP progress has been concentrated in a small subset of the world’s languages, leaving systematic inequalities in language-technology performance insufficiently measured. The paper estimates global utility and demand across representative NLP tasks and examines related societal, economic, and academic factors. It finds immense inequality, with English and a small group of Western European and other major languages dominating development, while economic strength appears more influential than demographic demand.

  • Problem

    NLP performance has grown substantially but remains restricted to a minuscule subset of the world’s over 6,500 languages, leaving language-based inequalities insufficiently measured and understood.

  • Method

    The paper estimates normalized language-technology utility and alternative demographic or linguistic demand across representative NLP tasks, then analyzes publication-based societal, economic, and academic correlates.

  • Results

    English and a small group of Western European and other major languages dominate NLP development, while preliminary evidence suggests users’ economic strength drives development more than demographic demand.

  • Takeaways & Limitations

    Global coverage metrics can identify large underserved languages and provide common ground for coordinating cross-linguistic technology development and resource allocation.

  • Takeaways & Limitations

    User-facing tasks depend tightly on large data and computational resources, and the study does not replace in-depth evaluation of need for individual groups and languages.

Abstract

from arXiv · show

Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown enormously over the last decade, this progress has been restricted to a minuscule subset of the world's 6,500 languages. We introduce a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. Our analyses involve the field at large, but also more in-depth studies on both user-facing technologies (machine translation, language understanding, question answering, text-to-speech synthesis) as well as more linguistic NLP tasks (dependency parsing, morphological inflection). In the process, we (1) quantify disparities in the current state of NLP research, (2) explore some of its associated societal and academic factors, and (3) produce tailored recommendations for evidence-based policy making aimed at promoting more global and equitable language technologies.

1 Introduction

NLP has become broadly useful, but its rapid progress has primarily benefited a small subset of the world’s languages. The paper frames this imbalance as a systematic inequality requiring measurement and investigation of its societal, economic, and academic correlates.

  • NLP has advanced from a technical niche into a fundamental tool across domains involving language data.Applications include health, bias analysis, consumer-behavior prediction, information access, translation, and speech commands.
  • Standardized benchmarks, evaluation metrics, incentives, and resources have concentrated progress around optimizing performance on established tasks.
  • Performance drops systematically across user and language dimensions, including gender, racial identity, and language varieties.The paper notes that these biases can arise throughout NLP development, from training data to algorithms.
  • Only a handful of the world’s over 6,500 languages are systematically represented in academia and industry, and transferring NLP advances across languages is nontrivial.Performance depends on properties such as morphology, word order, phonological repertoire, and data availability.
  • The study develops global estimates to measure how language-technology utility is distributed across languages and populations and traces inequalities to societal, economic, and academic correlates.

2 Methodology

The methodology compares language-technology utility and demand across representative user-facing and linguistically focused NLP tasks. It combines normalized performance measures, alternative demand assumptions, manually collected benchmark evidence, and publication-based analyses of academic and economic correlates.

  • Quantifying utility and demand: Utility is defined as task-and-language performance normalized by the best possible performance, using the best-performing language when the theoretical maximum is unattainable.
  • Quantifying utility and demand: Demand is modeled between demographic demand proportional to speaker counts and linguistic demand that treats the approximately 6,500 languages equally.
  • Quantifying utility and demand: The global metric Mτ summarizes how much language-technology demand is met, with τ = 1 representing demographic demand and τ = 0 representing linguistic demand.Mτ is bounded between 0 and 1, and increases when a language’s utility improves.
  • NLP tasks: The study evaluates user-facing tasks including question answering, machine translation, and text-to-speech, alongside morphological inflection, dependency parsing, and natural language inference.
  • Correlates of NLP utility: The field-level analysis uses major international NLP-conference publications to examine citation-language diversity and contrasts worldwide users with associated GDP as predictors of papers per language.These regression analyses use a Bayesian generalized mixed-effects framework.
  • Data and analyses: Task-performance data are manually collected from multilingual benchmarks, shared tasks, and published conference results, with demographic and linguistic information used to estimate demand.

3 Results and Analysis

NLP utility and coverage vary sharply across tasks and languages: user-facing benchmarks often serve few languages, while broader coverage can still include substantial quality gaps. Priorities shift depending on whether demographic reach or linguistic inclusion is emphasized.

  • General observations: Most NLP tasks perform substantially better on demographic than linguistic utility.
  • General observations: Text-to-speech covers more than 630 languages, but measured speech quality for most is about half as good as the exceptionally good English system.
  • General observations: NLI and QA benchmarks cover up to 15 and 17 languages, respectively, producing very low linguistic-utility scores.
  • Historical progression: Inflection shows significant utility improvements over seven years, while machine translation from English has improved linguistic coverage as the field follows demographic and socioeconomic priorities.
  • Within-language variation: Arabic vernaculars have small utility differences but lag behind Modern Standard Arabic, while Coastal Swahili utility in Tanzania is about 10% lower than in Kenya.
  • Priorities in NLP development: Priority languages shift from populous Mandarin and Hindi toward under-served Asian, African, indigenous, and potentially endangered languages as linguistic utility receives greater emphasis.
  • Societal, economic, and academic correlates: Publication language reporting is often unclear, and because English dominates NLP study, this usually indicates English-only systems.
  • Societal, economic, and academic correlates: Approximate GDP predicts the number of published papers with substantially smaller error than number of language users, although the predictors are substantially collinear.

4 Discussion

The study finds profound inequality in language-technology development, with English and a small set of other languages dominating. Economic strength appears more influential than demographic demand, while task characteristics and academic incentives shape opportunities for broader coverage.

  • 4 Discussion: English, several Western European languages, and a few non-Indo-European languages dominate language-technology development.The dominant non-Indo-European languages are primarily Chinese, Japanese, and Arabic.
  • 4 Discussion: Economic prowess of language users appears to drive technology development more than sheer demographic demand.This is presented as a preliminary investigation rather than a definitive causal finding.
  • 4 Discussion: Inflection shows year-over-year improvement in both demographic and linguistic utility.Small, highly curated datasets can support reasonably accurate solutions for this task.
  • 4 Discussion: Bottom-up initiatives such as Universal Dependencies and UniMorph have accelerated cross-linguistic development by pooling distributed linguistic expertise.Their contributions can rely more on expertise and academic incentives than on large resource investments.
  • 4 Discussion: User-facing tasks remain tenuously connected to more technical tasks and depend heavily on large datasets, computational resources, and financial means.This difference complicates efforts to use technical-task progress as a proxy for user-facing technology coverage.
  • 4 Discussion: Global-coverage measures can identify large underserved languages and provide common ground for coordinating cross-linguistic efforts.They do not replace detailed evaluations of the needs of individual groups and languages.

A Materials

The materials combine publication records, multilingual benchmarks, shared tasks, and task-specific metrics with demographic and economic indicators. The analysis also documents important measurement constraints, including reliance on automatic metrics and simplifying assumptions about language communities.

  • A Materials: The publication analysis uses ACL Anthology papers and an automatic pipeline to detect language mentions and assign research areas.The pipeline searches English names, endonyms, and ISO or Glottolog codes, followed by post-processing for precision.
  • A Materials: Published automatic-metric improvements are not guaranteed to translate into user-perceived utility.Only a tiny fraction of NLP research diverges from automatic evaluation by using human evaluations.
  • A Materials: The analysis assumes standard language versions are comprehensible and acceptable to all identified speakers, despite limited fine-grained demographic data.The authors characterize this as an oversimplification.
  • A Materials: The study combines multilingual benchmarks, shared tasks, and published results across machine translation, question answering, speech synthesis, parsing, inflection, and language understanding.Examples include TyDi-QA and MLQA for extractive question answering, XNLI for understanding, and SIGMORPHON shared tasks for inflection.
  • A Materials: Machine-translation utility is based on normalized BLEU with Z = 70, the largest reported BLEU used as the attainable snapshot utility.The normalization uses the reported Serbian–Croatian BLEU score of 70.
  • A Materials: Task metrics include MCD for speech synthesis, LAS for dependency parsing, and exact-match accuracy for morphological inflection.MCD is a distortion measure where lower values are better; LAS measures labeled syntactic-attachment overlap.
  • A Materials: Population demand aggregates language variants and speakers using Ethnologue statistics, while economic demand uses GDP, trade, and import or export indicators.For Nahuatl, the associated GDP is estimated as 1.3% of Mexico’s GDP based on its population share.

B Methods

To estimate utility for languages and pairs absent from published evaluations, the paper extends observed results through structured prediction methods. For machine translation, graph-based pivoting selects the highest estimated multiplicative path between languages.

  • B Methods: Published results alone leave many languages and language pairs without quality or utility estimates.The paper addresses this missing-evaluation problem by predicting utility on unseen languages and pairs.
  • B Methods: A naive unseen-language estimate assigns utility 0, which can be plausible when systems lack relevant language or writing-system coverage.Yupik and Dhivehi are examples of languages absent from Wikipedia and using different writing systems.
  • B Methods: Machine-translation pivoting estimates an unobserved pair through existing intermediate translation systems rather than assigning zero utility.High-quality German–English and English–Chinese systems can support an estimate for German–Chinese translation.
  • B Methods: Cascaded translation systems require accounting for error propagation across stages.Two sequential systems with 80% accuracy each yield an expected accuracy of 64%.
  • B Methods: The method permits multiple-language paths when their combined quality exceeds a single pivot.A Catalan–Spanish–English–Chinese route may outperform a single-language pivot for Catalan–Chinese translation.
  • B Methods: A weighted directed graph represents languages as nodes and published normalized BLEU scores as edge weights, with missing edges set to 0.For each pair, estimated utility is the maximum cumulative multiplicative weight over available paths; absent paths produce 0.

C Bibliometric Analysis

The bibliometric analysis models normalized citation percentiles as a function of the number of languages associated with a publication, with area-specific effects. Comparing models with and without a smooth language-count term finds no major performance difference.

  • Citation analysis: Normalized citations are analyzed with Bayesian generalized additive mixed-effects models using a beta distribution.The models use weakly informative priors and four converged MCMC chains.
  • Citation analysis: The full model adds a smooth thin-plate-spline function of the number of languages associated with each paper.Area-specific random intercepts and slopes are also included.
  • Model comparison: The comparison tests the full model against a counterpart without the smooth language-count term using leave-one-out performance.The comparison is intended to assess support for the language-count relationship.
  • Model comparison: -0.9 (SE=0.6) is the difference in expected log pointwise predictive density between the two models.This implies no major performance difference between them.
  • Publication counts: Publication counts are modeled with a zero-inflated negative binomial distribution because the distribution contains many zero values.The analysis focuses on the expected publication count and the mixture probability.

D Machine Translation Case Studies

The machine-translation case studies distinguish utility from translating into English versus from English and extend estimates across all language pairs using reported results and pivoting. Demographic utility is relatively similar in English-centered settings, while language-level scores differ sharply across well-studied and underserved languages.

  • Translation involving English: Translation utility is calculated separately when a language is the source and when it is the target, because translation involves two language communities.The study uses each direction separately in its utility calculations.
  • Translation involving English: M = 0.25 from English and M1 = 0.27 to English are the demographic-based utilities, while the linguistic-diversity score is around 0.005.Published results cover 101 languages in these estimates.
  • Translation among all languages: The all-language analysis combines reported translation results with pivoting-based accuracy estimates for language pairs without reported results.The pivoting approach selects the best-performing translation path for such pairs.
  • Translation among all languages: English is the best and often only pivot in almost all cases, making each language’s final utility strongly dependent on its English translation systems.Scores averaged by demographics and languages are consequently similar to English-focused scores.
  • Translation among all languages: German (M1 = 0.356), Chinese (M1 = 0.232), and French (M1 = 0.309) have demographic-averaged utilities nearly double those of underserved languages.The passage identifies stark differences among language scores.
  • Research interests and demand: Table 2 compares machine-translation research interests in directions to and from English with the population-based demand model.The caption states that these interests do not match the model.
  • Bibliometric comparison: Figure 6 compares cumulative citations with the number of languages in publications according to topic.The caption specifies the two quantities and the topic-based organization.
Loading 2110.06733v1…