Source-linked AI summary

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, Monojit Choudhury

arXiv:2004.09095v3cs.CL

TL;DR

NLP technologies represent only a small share of the world’s languages, raising questions about their resource and typological coverage. The paper quantitatively analyzes language resources and conference representation, finding persistent disparities and uneven inclusion across language classes. It argues that the NLP community should examine whether methods extend across languages and consider inclusion in research practices.

  • Problem

    NLP systems are not truly language agnostic because training and evaluation concentrate on related languages from a few regions, leaving many typological phenomena underrepresented.

  • Method

    The paper combines quantitative analyses of language resources with a consolidated dataset covering 11 NLP conferences and analyzes their research trends.

  • Results

    The study finds uneven conference inclusion, typological gaps, and a persistent resource divide, while some low-resource languages have focused research communities and many others lack support.

  • Takeaways & Limitations

    The findings support asking whether language technologies apply or scale across languages and whether research contributions advance linguistic inclusion.

  • Takeaways & Limitations

    The resource analysis treats each dataset as one unit even though different NLP tasks require different amounts and types of data.

Abstract

from arXiv · show

Language technologies contribute to promoting multilingualism and linguistic diversity around the world. However, only a very small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications. In this paper we look at the relation between the types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time. Our quantitative investigation underlines the disparity between languages, especially in terms of their resources, and calls into question the "language agnostic" status of current models and systems. Through this paper, we attempt to convince the ACL community to prioritise the resolution of the predicaments highlighted here, so that no language is left behind.

1 The Questions

The paper examines how linguistic resources, typological diversity, and conference research coverage are distributed across languages. It asks whether NLP systems and the ACL community adequately represent the world’s linguistic diversity, and whether zero-shot methods can reduce resource-based disparities.

  • Motivation: NLP systems trained on related languages from a few regions create a typological echo-chamber that leaves many linguistic phenomena unexplored.The paper notes that systems described as language agnostic are trained and tested on a narrow set of languages.
  • Motivation: Neural systems increase the data divide because they require more training data, while recent massively multilingual zero-shot methods may reduce the need for labeled datasets.The passage presents zero-shot learning as a possible bridge, while noting its reliance on large unlabeled resources.
  • Research questions: The study asks how labeled and unlabeled resources are distributed across languages and how resource availability relates to native-speaker populations.It also asks what current and near-future NLP technologies may achieve for languages with different resource levels.
  • Research questions: The authors investigate which typological features current NLP systems have encountered and which remain unexplored because resources and data-driven research are scarce.This question directly connects resource availability with the linguistic phenomena represented in NLP systems.
  • Research questions: The paper evaluates whether ACL research has become more or less linguistically diverse over time and whether resource levels influence research questions and publication venues.It compares earlier periods with the 2000s and 2010s and examines venue-level inclusion.
  • Approach: The authors use a multi-pronged quantitative approach and classify the world’s languages into six resource-based classes with distinct ACL trajectories.They also consider the roles of individual researchers and communities in bridging the linguistic-resource divide.

2 The Six Kinds of Languages

The paper proposes a six-class taxonomy of languages based on labeled and unlabeled resource availability, then uses it to compare resource distributions and language-class trajectories. The findings show strong disparities: Class 0 languages lack representation, while Winners lead across resource rankings.

  • Taxonomy: The taxonomy classifies languages by the number of labeled and unlabeled resources available for them.These axes reflect the importance of both supervised datasets and unlabeled data for current data-driven and transfer-learning systems.
  • Resource measures: The authors count organized LDC and ELRA collections as labeled-resource indicators because they provide standardized datasets used in prior NLP work.Other repositories are excluded to keep the analysis focused on datasets used in computational-linguistics conferences.
  • Resource measures: Wikipedia pages serve as the measure of unlabeled resources because they provide freely accessible unsupervised training data and factual information.The measure is intended to represent digital resource availability rather than every possible source of unlabeled data.
  • Language classes: The six classes range from Left-Behinds with exceptionally limited resources to Winners with dominant online presence and major industrial and government investment.Intermediate classes differ in their balances of labeled data, unlabeled data, and research-community support.
  • Findings: Winners lead in all resource rankings, while Class 0 languages remain out of the race with no representation in any resource.The taxonomy’s visualizations place the classes into six distinct positions in the language-resource space.
  • Findings: Wikipedia coverage is more even for Classes 1, 2, and 3 than for Classes 4 and 5, whereas Web-resource distribution shows clear disparity.The comparison indicates that different kinds of unlabeled resources are distributed unevenly across language classes.
  • Findings: Class 0 contains the largest section of languages and represents 15% of speakers across classes.The paper warns that limited technological inclusion may draw speakers toward better-supported languages and exacerbate disparity.

3 Typology

The paper examines how typological diversity is represented in NLP resources and systems, finding that many features and languages remain underrepresented. These exclusions correspond to substantial performance differences for some languages.

  • Typological features: Typological databases classify languages by structural and semantic properties, enabling analysis of features represented in NLP.WALS provides typological information used to examine languages with limited resources.
  • Typological features: 549 of 1139 unique typological categories in resource-poor classes are absent from better-resourced classes, averaging 2.86 ignored categories per feature.The analysis uses WALS data for 2679 languages and 192 features.
  • Typological exclusions: Rare features such as 144E occur across regions but receive little representation, whereas common feature 83A has values documented for 1321 languages.The contrast illustrates uneven typological coverage in multilingual NLP resources.
  • Typological exclusions: English→Amharic has an error rate of 60.71 versus 7.8 for English→Arabic, while Amharic has nine ignored features and Arabic has none.Both languages belong to the Semitic family, but their typological coverage differs substantially.

4 Conference-Language Inclusion

The paper quantifies language inclusion across NLP conferences using a consolidated multi-conference dataset, entropy, and class-wise inverse MRR. Inclusion varies across venues and over time, with LREC and workshops standing out as especially inclusive.

  • Dataset: The study consolidates data from 11 NLP conferences and journals using ACL-ARC, Semantic Scholar, and ACL Anthology scraping.The dataset extends beyond ACL-ARC’s 2015 cutoff and includes non-ACL venues.
  • Language Occurrence Entropy: Language occurrence entropy summarizes how evenly languages are distributed in conference papers: higher entropy means broader spread, while lower entropy indicates skew.The measure is computed from language occurrences across papers for each conference and year.
  • Class-wise Mean Reciprocal Rank: Class-wise inverse MRR measures conference inclusion by language-resource class, with smaller values indicating greater inclusivity.Ranks are based on how frequently languages are mentioned in conference papers.
  • Findings: LREC and workshops are the most inclusive across language-resource classes and maintain this pattern over the years.This conclusion is supported jointly by entropy trends and MRR figures.
  • Findings: Entropy spikes in the 2010s for ACL, EMNLP, NAACL, and LREC, while later-starting conferences appear to have incorporated lessons about language inclusion.The proposed explanation for the entropy spike is increased interest in cross-lingual techniques.
  • Findings: Class 0 is left behind, with average ranks from 600 to 1000, and its disadvantage is sharper in CONLL, TACL, and SEMEVAL.The rank disadvantage is less pronounced in LREC and workshops.

5 Entity Embedding Analysis

The paper complements aggregate inclusion measures with jointly learned embeddings of conferences, authors, and languages. These representations reveal temporal and resource-related structure in NLP research communities.

  • Model: The proposed embedding method jointly represents conferences, authors, and languages in one space to uncover patterns beyond aggregate statistics.The entities are learned from their contextual distributions in research papers.
  • Model: The model predicts randomly sampled title and abstract words from an associated entity, using a Skipgram-like training objective.Titles and abstracts are selected as concise, lower-noise signals compared with full papers.
  • Model: The entity architecture maps an input entity through a hidden embedding layer to an output layer that predicts words.The learned input and output matrices contain entity and word embeddings, respectively.
  • Visualization: t-SNE projects the learned embeddings into two dimensions to visualize conference and language relationships, excluding noisy Class 0 projections.The visualization includes ACL, LREC, workshops, and CL, together with the other taxonomy classes.
  • Findings: The embedding layout shows a temporal or research-focus progression, with earlier theoretical work on the left and later data-driven work on the right.This pattern appears across several conferences, while CL embeddings remain mostly on the left.
  • Findings: Less-resourced language classes lie farther from the ACL trendline, whereas Class 5 and Class 4 are closer; LREC and workshops lie nearer the language clusters.The LREC cluster is positioned in the middle of the language clusters, especially in recent iterations.
  • Findings: Class 0 has the highest author-language MRR, while MRR generally decreases toward Class 5, indicating concentrated attention on some low-resource languages and popular languages.Japanese, Mandarin, Turkish, and Hindi have high MRR, whereas Burmese, Javanese, and Igbo have low MRR despite millions of speakers.

6 Conclusion

The paper finds persistent disparities in language resources and NLP representation, while identifying some inclusive venues and focused communities for low-resource languages. It recommends stronger language-inclusion checks in future research and conference review.

  • The taxonomical hierarchy recurs across resource availability, conference entropy, and embedding analyses, consistently revealing language disparity.
  • LREC and workshops are more inclusive across language classes than other examined venues, according to inverse MRR, entropy, and embedding analyses.
  • Typological feature 144E occurs across many resource-poor languages but is insufficiently represented in resource-rich languages, potentially reducing transfer-learning performance.
  • Newer conferences are more language-inclusive, whereas older conferences retained research themes that did not necessarily favor multilingual systems.
  • Low-resource languages have focused research communities, but languages such as Javanese and Igbo still lack comparable support.
  • The authors propose language-related diversity and inclusion questions in submission and reviewer forms to assess whether methods apply across languages.

A.1 Embedding Visualization

The paper provides an interactive browser visualization of conference and language embeddings to examine language inclusion over time.

  • The visualization projects conference and language embeddings into an interactive browser-based space for exploring language inclusion over time.
  • Users can combine or hide conference and language classes through clickable legends; legend numbers identify the respective classes.
  • The embedding space includes conferences and languages, allowing users to inspect how NLP research has progressed in language inclusion.

A.2 ACL Anthology Dataset Statistics

The ACL Anthology dataset covers main-track long and short papers while excluding several specialized tracks from language-usage trend measurement.

  • The dataset includes all long and short papers from main-track conference proceedings.
  • System demonstrations, tutorial abstracts, student research workshops, special issues, and other out-of-scope tracks are excluded.
  • The paper reports that the dataset and its documentation are being released.

A.3 Hyperparameter Tuning

The embedding model uses Word2Vec-matched hyperparameters and evaluates year prediction after an 80-20 split, with embedding size producing the relevant tuning difference.

  • The dataset is split 80-20, and a linear regression model predicts publication year from each paper’s embedding.
  • Embedding size is the only hyperparameter showing a significant difference in the reported tuning, with 75 dimensions selected as best.
  • 0.6 R2 and 4.04 MAE are reported for the 75-dimensional embeddings.

A.4 Cosine distance between conferences and languages

The analysis compares conference vectors with mean vectors for language-taxonomy categories using cosine distance, while tracking inclusion-class progress over time. LREC shows smooth forward progression, and the full taxonomy classification is released online.

  • Cosine distance quantifies how closely each conference vector aligns with the mean vector for each language category.Table 7 reports these conference-to-category distances.
  • The conference–language-category distances show differing proximity patterns across conferences and taxonomy categories.ACL is at an average distance of 0.291 from category 5 languages, according to the passage.
  • The authors release the complete language-taxonomy classification on the project website.
  • The analysis measures progress in the inclusion of defined language-taxonomy classes over the years.
  • LREC exhibits smooth forward progression in the inclusion of the defined taxonomy classes.
Loading 2004.09095v3…