Source-linked AI summary

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai, Bontu Fufa Balcha, Zayd Bashir, Angana Borah, Zara Burzo, Yubin Choi, Naihao Deng, Samika Gupta, Michel Faloughi, Claude Kwizera, Ziqiao Ma, Cynthia Yacel Fuertes Panizo, Ellie Seehorn, Hui Shen, Jiayi Tang, Zesen Zhao, Boyuan Zheng, Rada Mihalcea

arXiv:2608.30107v1cs.CLcs.AIcs.CY

TL;DR

NLP datasets rarely document which countries and populations they represent, and language metadata does not reliably resolve that gap. AtlasNLP builds a country-aware resource combining human curation with ACL-scale extraction, finding uneven country-task coverage, geographic asymmetry between representation and production, and no implication from language coverage to geographic coverage.

  • Problem

    Country-level NLP dataset representation is difficult to measure because geographic metadata is rarely available and language metadata is an incomplete geographic signal.

  • Method

    AtlasNLP combines a 1,480-entry human-curated reference set with automated extraction over ACL Anthology records, tracking represented countries separately from producer countries.

  • Results

    Dataset coverage is highly uneven across countries and tasks, production and representation are geographically asymmetric, and language coverage does not imply geographic coverage.

  • Takeaways & Limitations

    Reliable country-aware NLP evaluation requires explicit geographic metadata rather than relying on language metadata or producer geography alone.

  • Takeaways & Limitations

    Country attribution is incomplete and non-random, so low or zero coverage indicates limited documented representation in ACL literature rather than absence of relevant resources.

Abstract

from arXiv · show

Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.

1 Introduction

NLP dataset geography is difficult to assess because country and population representation is rarely documented, while language-level metadata provides only an incomplete geographic signal. AtlasNLP addresses this gap by tracking represented and producer countries across more than 13,000 dataset records and showing uneven coverage across countries and tasks.

  • A small number of countries account for many represented dataset records, while others are sparsely represented or absent.Coverage also varies by task, with countries often covered for some tasks but not question answering, reasoning, safety, or multimodal understanding.
  • Geographic metadata is rarely documented, leaving countries and populations represented in NLP datasets difficult to determine.
  • Language metadata is an incomplete proxy for geographic representation because languages span countries and countries contain multiple linguistic communities.Spanish-labeled datasets, for example, may represent only a subset of Spanish-speaking countries.
  • AtlasNLP records both represented countries and producer countries across over 13,000 records and 31 normalized NLP task categories.It combines AtlasNLP-Gold, a human-curated reference set, with AtlasNLP-Core, an ACL-derived large-scale collection.
  • 79.2% of country-task pairs contain no dataset records with documented country representation.The resource distinguishes representation from production and motivates more explicit country-level dataset documentation.

2 Related Work

Prior NLP research has documented geographic and population disparities, but studies commonly focus on individual datasets, models, tasks, languages, or regions. AtlasNLP extends this work by adding country-level structure to language-centered evaluation and dataset metadata.

  • Existing studies show that geographic and population representation is uneven across datasets, models, tasks, languages, and regions.
  • Representation differences concern not only dataset quantity but also the tasks, resources, and technologies available to languages, regions, and populations.
  • Multilingual benchmarks expanded evaluation beyond English, while surveys continue to document disparities across low-resource and non-Western languages.
  • Language is an incomplete proxy for geographic and cultural representation because languages span countries and countries contain multiple linguistic communities.
  • Dataset documentation emphasizes origins, collection procedures, intended uses, and biases, but geographic metadata remains underdeveloped.

3 Methodology

AtlasNLP combines a structured geographic annotation framework, human curation, and automated extraction over ACL Anthology papers. The pipeline separates represented-country evidence from producer geography and uses audits and reference-set comparisons to evaluate extraction quality.

  • Annotation framework: The annotation framework records represented countries, producer countries, task, language, modality, licensing, and attribution method.Task categories are normalized from ACL and EMNLP thematic areas.
  • Annotation framework: AtlasNLP distinguishes a represented country from a producer country, separating the population represented in a dataset from its creators’ institutional location.
  • Annotation framework: Country attribution is classified as explicit, inferred, or unattributed, with primary analyses using explicit attribution.Language alone is insufficient for explicit country attribution.
  • Human curation: AtlasNLP-Gold contains 1,480 human-curated entries collected and cross-validated to improve regional coverage and provide a reference set.
  • Automated expansion: 119,963 ACL records were filtered to 20,277 candidate dataset papers using ModernBERT-base-NLI, followed by schema-constrained extraction from full paper text.The extraction pipeline records page- and quote-level provenance and initially produced 18,035 successful paper-level extractions.
  • Audit and evaluation: A 400-record audit identified attribution errors and prompted stricter dataset-role and country-attribution criteria before re-evaluating the collection.
  • Audit and evaluation: 97.7% of matched Core records with explicit country attribution shared at least one country with Gold, while 94.7% of Core country assignments were Gold-supported.Exact task agreement was 79.9%.
  • Audit and evaluation: 94.7% of 187 incorrect country attributions were removed from the explicit layer after revision, and 97 of 98 unsupported cases were left unattributed.

4 AtlasNLP Overview

AtlasNLP combines a large ACL-derived collection with a human-curated supplement, preserving distinctions between dataset records, represented-country evidence, and task coverage. Its country-task view shows fragmented coverage across the NLP dataset ecosystem.

  • AtlasNLP-Core contains 13,462 paper-level dataset records, while AtlasNLP-Gold contains 1,480 human-curated entries grouped into 989 normalized dataset names.Core supplies large-scale ACL metadata, while Gold adds resources from underrepresented regions.
  • The collections use dataset records rather than unique datasets because related papers may introduce, extend, or compile the same underlying resource.
  • Explicit country attribution covers 2,447 Core records (18.2%), increasing to 3,506 records (26.0%) when inferred attribution is included.Records lacking supported country evidence remain in Core for analyses that do not require represented-country metadata.
  • Coverage across countries and tasks is highly fragmented, with most country-task pairs lacking represented dataset records and high-coverage countries concentrated in some tasks.Figure 2 displays only a subset of country labels for readability.

5 Dataset Geography Analysis

AtlasNLP reveals severe geographic and task-level unevenness in documented NLP dataset coverage, alongside asymmetric production patterns and the limits of language as a geographic proxy.

  • Country coverage: 2,447 Core dataset records contain explicit country attribution, yielding 4,421 country-record associations across 158 of 197 countries.The United States, China, India, the United Kingdom, and Germany account for 36.4% of these associations, while 39 countries have none.
  • Country coverage: 79.2% of country-task pairs have no explicitly attributed dataset records, remaining 75.0% empty under explicit+inferred attribution.Geographic imbalance concerns both the volume of represented records and the NLP tasks they cover.
  • Country coverage: Country attribution is incomplete: 70.7% of unattributed Core records include English, compared with 49.9% of explicitly attributed records.Low or zero coverage therefore indicates gaps in documented country representation within ACL dataset contributions, not necessarily absent resources.
  • Production and representation: Among countries meeting minimum-count thresholds, 39 of 62 (62.9%) have most represented records produced by institutions outside the country.The analysis expands each represented-country association across all associated producer countries, preserving multi-country representation and collaboration.
  • Production and representation: The United States accounts for 21.3% of expanded producer-representation associations, while the ten largest producer countries account for 61%.Production geography captures institutional locations rather than researcher identities and remains concentrated under explicit+inferred attribution.
  • Research infrastructure: Dataset availability is higher in high-income countries and positively associated with national university counts, with Pearson r = .65 and Spearman ρ = .61.The association strengthens under explicit+inferred attribution to r = .70 and ρ = .67, but countries with similar institutional capacity can still differ substantially.
  • Task coverage: Under explicit attribution, the median country spans 11 of 30 task categories, while its three most common tasks comprise 55.6% of represented records.Explicit+inferred attribution raises median breadth to 13 tasks, while concentration remains similar at 54.3%.
  • Language and geography: Chinese-language records are associated with China at 89.2%, whereas English records are more distributed but still have about half associated with the United States, United Kingdom, and India.Portuguese records concentrate 69% in Brazil and Portugal, while France, Switzerland, and Canada account for about 63% of French records.

6 Lessons Learned and Implications

AtlasNLP’s lessons emphasize documenting geography explicitly, separating represented populations from producer locations, and tracking task-level coverage. These distinctions make uncertainty and geographic gaps more visible than language-centered metadata alone.

  • Geographic metadata: Explicit and inferred geographic attribution should remain separate because plausible signals are not equivalent to documented provenance.This separation supports strict analyses based on explicit provenance and broader sensitivity analyses that include inferred attribution.
  • Geographic metadata: Language is necessary for organizing NLP resources but insufficient as a proxy for geographic representation.Languages can span multiple countries, while language-labeled resources may represent only a subset of the populations using them.
  • Representation and production: Separating represented geography from producer geography makes institutional asymmetries in dataset creation more visible.The two dimensions capture who or what a dataset represents versus where its production is institutionally located.
  • Task coverage: Task-level geographic coverage identifies whether dataset development is broad, narrow, or missing across the NLP task space.Some countries may have coverage for selected tasks while lacking coverage for other major areas.
  • Infrastructure: Standardized metadata for represented populations, geographic provenance, and producer geography would make geographic gaps easier to audit.AtlasNLP presents one schema for making these distinctions explicit within dataset infrastructure.

7 Conclusion

AtlasNLP provides a country-aware view of dataset representation, production, and task coverage by distinguishing represented populations from producer locations and explicit from inferred attribution. Its analysis finds uneven country-task coverage, geographically asymmetric production, and no guarantee that language coverage represents countries.

  • Conclusion: AtlasNLP distinguishes represented populations from institutional producer locations and explicit from inferred geographic attribution.These distinctions expose patterns that are difficult to observe in language-centered resource collections.
  • Conclusion: Dataset coverage is highly uneven across countries and tasks, while dataset production is geographically concentrated and asymmetric.The analysis treats representation and production as separate dimensions of dataset geography.
  • Conclusion: Language coverage does not imply geographic coverage.The conclusion therefore cautions against assessing geographic diversity from language metadata or producer geography alone.
  • Implications: The findings expose limitations in current dataset documentation and support making geographic metadata a first-class component of dataset documentation.The paper releases AtlasNLP to support work on transparency, geographic representation, and population-aware evaluation.

Limitations

AtlasNLP’s measurements are constrained by incomplete attribution, paper-level records, automatically extracted annotations, imperfect candidate-paper recall, institutional rather than identity-based producer geography, and ACL-centered scope.

  • Metadata and coverage: Country attribution is incomplete and non-random, so low or zero coverage indicates limited documented representation within ACL literature rather than no relevant resources.The precision-first strategy leaves records unattributed when provenance cannot be established confidently, especially for globally distributed languages.
  • Measurement scope: AtlasNLP measures documented paper-level dataset activity rather than unique datasets, dataset size, quality, or downstream evaluation capacity.Related records may describe extensions, compilations, or versions of the same underlying resource, and entity-level deduplication is not always straightforward.
  • Annotation quality: Most Core annotations are automatically extracted, and remaining errors may be concentrated in complex multi-country scope.AtlasNLP-Gold shares the annotation schema and is not a fully independent ground truth.
  • Coverage scope: The candidate-paper filter was not evaluated for population-wide recall across ACL Anthology, so AtlasNLP is not an exhaustive ACL census.The paper also avoids temporal claims requiring uniform recall across publication periods.
  • Producer geography: Producer geography reflects institutional affiliation rather than researcher nationality, identity, or intent.This distinction matters especially for diaspora researchers and cross-border collaborations.
  • Coverage scope: ACL-centered sourcing underrepresents datasets released only through platforms, repositories, government portals, community archives, or non-ACL venues.The normalized task schema also trades granularity for cross-country comparability.

Ethical Considerations

AtlasNLP’s ethical framing treats country-level geography as informative but incomplete, preserves uncertainty in attribution, and warns against interpreting underrepresentation as evidence of absent resources or justification for poorer performance. Its construction and schema support auditable distinctions between representation, production, and attribution confidence.

  • Ethical considerations: Country-level analysis should not be used to rank countries, essentialize populations, or treat national labels as complete representations of identity or culture.Substantial variation exists within national boundaries, and countries are not equivalent to cultures, languages, ethnic groups, or lived experiences.
  • Ethical considerations: Inferred country assignments indicate plausible geographic relevance, not definitive claims about contributor or population identity.AtlasNLP therefore preserves explicit versus inferred attribution and uses explicit provenance for primary geographic analyses.
  • Ethical considerations: Evidence of underrepresentation should motivate more inclusive and accountable data collection and evaluation, not normalize poorer performance for some regions.Low or absent AtlasNLP coverage is not evidence that a country or community lacks NLP resources.
  • Ethical considerations: Producer-country metadata describes institutional affiliation, not researcher nationality, identity, or intent.External production may reflect collaboration, diaspora scholarship, or resource sharing rather than a preference for domestic production.
  • Data construction: AtlasNLP-Gold was collaboratively curated to broaden geographic coverage and should not be treated as a natural sample of the global NLP dataset ecosystem.Its curation used shared guidelines, validation, and machine-assisted extensions aligned with human curation goals.
  • Data construction: AtlasNLP-Core is built from ACL Anthology papers through extraction, audit-driven refinement, and retained metadata documenting eligibility and geographic-attribution decisions.The pipeline yields 13,462 final primary dataset-contribution records, while Gold contains 1,480 human-curated entries across 989 normalized dataset-name groups.
  • Schema: The shared schema separates represented country from producer country and dataset identity from dataset relationship at the paper–dataset contribution level.The schema also captures task and language coverage and retains audit and normalization fields.
  • Attribution: AtlasNLP uses 197 geopolitical entities and classifies represented-country attribution as explicit, inferred, or unattributed.Explicit attribution requires direct country evidence, while inferred attribution is reserved for plausible indirect evidence and broader sensitivity analyses.

A.9 Country–Task Matrix Construction

AtlasNLP constructs country–task matrices over a fixed ontology and separates explicit from inferred geographic attribution. The resulting matrices preserve zero-count cases and show substantial country–task sparsity.

  • Matrix construction: 13,462 AtlasNLP-Core records are normalized into one of 30 task categories and analyzed over 197 geopolitical entities.Country names are normalized to the AtlasNLP ontology before analysis.
  • Matrix construction: Each multi-country record contributes one country-record association per supported represented country while remaining one underlying dataset record.
  • Attribution layers: 2,447 explicit records expand to 4,421 country-record associations across 158 countries, while explicit+inferred attribution yields 3,506 records and 6,001 associations across 168 countries.
  • Matrix construction: The matrices use the complete 197 × 30 analysis space, including zero-count rows and columns, so absent coverage remains analytically visible.
  • Coverage: 79.2% of explicit country-task cells are empty, compared with 75.0% under explicit+inferred attribution.The broader layer increases observed coverage but does not substantially change overall sparsity.
  • Geographic evidence: Language labels are retained and audited but do not directly establish represented countries; broad languages require additional geographic evidence.

C.2 Represented-Country and Task Validation

AtlasNLP validates represented-country and task metadata against curated references and targeted human audits. The checks support aggregate analysis while motivating conservative attribution and explicit handling of uncertainty.

  • Gold–Core comparison: 225 high-confidence Gold–Core dataset matches form the reference population for validation.Matching was established at the dataset level rather than from paper overlap alone.
  • Represented-country agreement: 97.7% of explicit matched records share at least one country with Gold, 66.7% exactly match the country set, and individual country-assignment precision is 94.7%.
  • Represented-country agreement: With explicit+inferred geography, 97.8% of assigned records overlap Gold, 68.8% exactly match the country set, and country-assignment precision is 94.0%.Agreement is computed conditional on Core making a geographic assignment; abstentions count as missing coverage.
  • Task agreement: 79.9% exact task agreement is obtained among 219 matched records meeting the single-task comparability criterion.
  • Post-revision audit: 94.7% of audited incorrect country assignments are removed from the explicit layer, and 97.6% of inferred cases are not promoted to explicit attribution.
  • Producer-country validation: 96.9% producer-country precision and 97.7% recall are achieved across 100 audited records, with the primary producer country recovered in all 100.

D.2 Dataset Search Procedure

AtlasNLP-Gold combines country-targeted human curation with cross-validation and a shared annotation schema, while Core and repository comparisons expose metadata-scale trade-offs. The resulting collection broadens coverage but remains distinct from a naturally occurring ecosystem sample.

  • Dataset search: Contributors searched academic, public, national, institutional, and community sources using a shared task–country matrix to reduce duplicate work.
  • Annotation: Gold records capture dataset identity, task, language, represented country, provenance, and geographic-attribution evidence under a shared schema.
  • Geographic annotation: Country attribution prioritizes direct evidence, permits conservative inference from specific geographic signals, and leaves unsupported cases unattributed.
  • Gold coverage: 1,480 curated entries provide explicit coverage for 192 of 197 entities, expanding to all 197 with conservative inferred attribution.
  • Scope: Gold’s deliberately broadened country distribution is intended for coverage, transparency, and validation rather than as a natural sample of the NLP dataset ecosystem.
  • Repository comparison: Hugging Face offers greater scale but lacks standardized geographic annotations, whereas ACL-derived records provide more consistent metadata with more limited coverage.

G.2 Additional Robustness Checks

Robustness checks show that the main geographic conclusions persist across attribution layers, record handling choices, and task-portfolio analyses. Temporal recovery variation remains a scope boundary for interpretation.

  • Broad multi-country records: Excluding records attributed to more than 20 countries removes 0.4% of attributed records but 5.3% of country-record associations.
  • Exact-name sensitivity: Collapsing exact normalized dataset-name keys leaves the top-10 represented-country set unchanged.
  • Temporal scope: Recovery varies across publication periods, so temporal composition is treated descriptively rather than as evidence of ecosystem diversification.
Loading 2608.30107v1…