Source-linked AI summary

The Tower of Babel Meets Web 2.0: User-Generated Content and its Applications in a Multilingual Context

B. Hecht, D. Gergle

arXiv:1904.01689v1cs.CLcs.HC

TL;DR

The paper asks whether Wikipedia’s language editions represent world knowledge consistently, then measures diversity across 25 editions at conceptual and sub-conceptual levels. It finds a minuscule shared conceptual core, substantial representational diversity, and significant effects on Wikipedia-based technologies, motivating culturally-aware and hyperlingual applications.

  • Problem

    Wikipedia-based technologies often assume that encyclopedic world knowledge is consistent across languages and cultures, although language editions may represent knowledge differently.

  • Method

    The study aligns concepts across 25 Wikipedia language editions, measures diversity at article and content levels, and tests its effect on Explicit Semantic Analysis.

  • Results

    The common encyclopedic core is around one-tenth of one percent of concepts, while knowledge diversity significantly affects technologies using Wikipedia as world knowledge.

  • Takeaways & Limitations

    Knowledge diversity can inform culturally-aware and hyperlingual applications, while Wikipedia-based applications should account for language and culture.

  • Takeaways & Limitations

    Unqualified information arbitrage can temper cultural diversity or introduce culturally irrelevant or culturally false information.

Abstract

from arXiv · show

This study explores language's fragmenting effect on user-generated content by examining the diversity of knowledge representations across 25 different Wikipedia language editions. This diversity is measured at two levels: the concepts that are included in each edition and the ways in which these concepts are described. We demonstrate that the diversity present is greater than has been presumed in the literature and has a significant influence on applications that use Wikipedia as a source of world knowledge. We close by explicating how knowledge diversity can be beneficially leveraged to create "culturally-aware applications" and "hyperlingual applications".

INTRODUCTION

Language fragmented Wikipedia’s single-page consensus ideal into separate language editions, challenging the assumption that encyclopedic knowledge is globally consistent. Across 25 editions, the study finds a tiny shared conceptual core, substantial sub-conceptual diversity, and effects on Wikipedia-based technologies.

  • Language fractured Wikipedia’s intended single neutral point of view across more than 250 separate language editions.
  • Current Wikipedia-based technologies often assume that encyclopedic world knowledge is largely consistent across cultures and languages.
  • Around one-tenth of one percent of concepts form the common encyclopedic core across the 25 examined language editions.
  • Sub-conceptual knowledge diversity is much greater than presumed, sharply contradicting the global consensus hypothesis.
  • Knowledge diversity affects technologies such as information retrieval and Explicit Semantic Analysis that use Wikipedia-based semantic relatedness.
  • The paper proposes culturally-aware and hyperlingual applications as design opportunities arising from a more realistic view of knowledge diversity.

BACKGROUND

Prior Wikipedia research has largely centered on English or treated language editions as globally convergent representations. This study instead situates multilingual differences within knowledge diversity and warns that unqualified information transfer can distort culturally situated knowledge.

  • Most HCI, CSCW, and AI research has focused on a single Wikipedia language edition, usually English.
  • Recent multilingual Wikipedia research has addressed language- and culture-related HCI problems but often lacks a full understanding of Wikipedia’s multilingual nature.
  • The global consensus hypothesis expects language editions to cover roughly the same concepts and represent them in nearly the same way.
  • Information arbitrage assumes that information missing from one language edition is inherently useful to another.
  • Applied heedlessly, information arbitrage risks tempering cultural diversity or introducing culturally irrelevant or culturally false information.
  • Prior work found that each language edition focuses content on the geographic culture hearth of its language, termed the self-focus of user-generated content.

OUR APPROACH

The study aligns equivalent concepts across Wikipedia language editions, measures diversity in both articles and their content, tests its effect on Explicit Semantic Analysis, and derives design implications. These stages support quantitative assessment of knowledge diversity and introduce culturally-aware and hyperlingual applications.

  • The study begins by reviewing stages needed to quantify knowledge diversity and test the global consensus hypothesis.
  • Data Preparation and Concept Alignment: CONCEPTUALIGN aligns equivalent concepts across language editions, such as World War II, Zweiter Weltkrieg, and Andre verdenskrig.
  • World Knowledge Diversity in Wikipedia: The diversity analysis compares language editions at both the conceptual article level and the sub-conceptual content level.
  • The Effect of Diversity on Technologies: The study tests whether cross-language knowledge diversity affects Explicit Semantic Analysis, a technology applied in HCI, AI, and NLP.
  • Implications for Design: The final stage discusses design implications and introduces culturally-aware and hyperlingual applications.

DATA PREPARATION AND CONCEPT ALIGNMENT

The study prepares a 25-language Wikipedia dataset and aligns cross-language articles into concepts using interlanguage-link graph components. It evaluates alignment quality and finds evidence that the resulting concept groups are reliable, while noting limits from missing or incomparable links.

  • Data preparation: 25 Wikipedia language editions were processed using WikAPIdia, which extracts article metadata, links, interlanguage links, disambiguation pages, and redirects.The study used the 25 largest editions available in mid-2009; the median edition contained 225,370 articles.
  • Data preparation: Around 52 million interlanguage links were parsed, but bot propagation may not make the collection exhaustive.This incompleteness creates the possibility of missing links affecting concept alignment.
  • Concept alignment: CONCEPTUALIGN treats interlanguage links as graph edges and articles as nodes, grouping all articles in each connected component as one concept.It ignores edge direction during breadth-first search and adds links so each concept component becomes fully connected.
  • Algorithm evaluation: CONCEPTUALIGN found links for 95.8% of the German/English missing-link dataset used for comparison with an SVM approach.The comparison is interpreted as showing that CONCEPTUALIGN and its underlying data are at least as effective as the SVM, although additional links were not measured on the same sample.
  • Evaluation limitations: The SVM comparison is constrained because 7.6% of its dataset was not comparable due to article-structure changes and included project pages.The authors therefore cannot report exact SVM precision figures or fully quantify CONCEPTUALIGN’s additional missing-link coverage.
  • Human evaluation: Human evaluations across English-Spanish, English-Japanese, and English-Italian pairs found low missing-link probabilities and no incorrect links.Together, the evaluations support relatively high concept-alignment quality and estimate the likely error from missing interlanguage links.

WORLD KNOWLEDGE DIVERSITY IN WIKIPEDIA

The paper tests the global consensus hypothesis at both the conceptual and sub-conceptual levels across Wikipedia language editions. It distinguishes which concepts editions cover from how articles describe shared concepts.

  • Conceptual level: At the conceptual level, the global consensus hypothesis predicts that language editions cover roughly the same set of concepts.Because English has many more articles, this is often expressed as the assumption that English nearly contains the concepts covered by other editions.
  • Sub-conceptual level: At the sub-conceptual level, the hypothesis predicts that articles about the same concept describe it in roughly identical ways across languages.The paper uses articles such as those about Psychology to illustrate this description-level assumption.
  • Analytical distinction: The higher conceptual level concerns which topic an article covers, whereas the sub-conceptual level concerns how that topic is defined.The sub-conceptual analysis presumes that the higher-level conceptual assumptions hold.

Concept Diversity

Across Wikipedia’s 25 language editions, conceptual coverage overlaps surprisingly little, while articles covering the same concepts also differ substantially in their descriptions. This diversity persists even among the largest editions and reflects cultural, linking, and descriptive differences.

  • Concept-level diversity: Over 74 percent of concepts are described in only one language, and more than 95.5 percent appear in six or fewer languages.These results refute the global consensus assumption at the concept level.
  • Concept-level diversity: Even among the three largest Wikipedias, around 80 percent of concepts remain single-language, while only 7 percent appear in all three.Using the six largest editions yields similar results: 77 percent single-language and only 1.5 percent shared across all six.
  • Concept-level diversity: Only 6,966 concepts, or 0.12 percent of all concepts, are covered in all 25 language editions.These concepts are treated as potentially globally relevant.
  • Pairwise coverage: English covers no more than approximately three-quarters of any other Wikipedia, and covers only slightly more than 50 percent of German concepts despite being over three times larger.Pairwise overlap therefore does not support an English-as-superset corollary.
  • Sub-concept diversity: Sub-concept diversity remains prominent because same-concept articles can differ through cultural focus, linking behavior, and seemingly random descriptive choices.For example, Spanish and German articles about Psychology link to different cultural content, while some differences reflect whether a hyperlink is included.
  • Sub-concept diversity: The mean overlap coefficient is only 0.41, meaning the longer article contains 41 percent of the shorter article’s outlinks on average.The coefficient controls for systematic differences in article length and linking behavior.

THE EFFECT OF DIVERSITY ON TECHNOLOGIES

The study tests whether multilingual knowledge diversity changes Wikipedia-based semantic relatedness technologies, using Explicit Semantic Analysis across ten languages. ESA outputs vary substantially across languages, indicating that the cultural and linguistic source of world knowledge affects its scores.

  • Research question: The study examines whether diversity in world-knowledge representations produces significantly different ESA scores for the same concept pairs.The authors frame this as an effect that could influence end-user applications.
  • Method: ESA represents concepts as vectors of abstract relationships derived from Wikipedia articles and assigns higher semantic-relatedness values to more similar vectors.The implementation was compared across Spanish, Hungarian, Norwegian, Portuguese, Romanian, English, German, French, Italian, and Danish.
  • Method: The first experiment used 8,264 perfect-clarity concepts shared across the ten languages, while the second used 10,000 randomly selected articles from each language.The second design incorporates both concept-level and sub-concept-level diversity.
  • Results: Mean correlation was only 0.13 in the first experiment, showing that ESA generated very different semantic-relatedness values across language-based knowledge sources.Excluding identical concept pairs raises the mean correlation only to 0.16.
  • Results: For Germany and Saxony-Anhalt, most languages assign high ESA scores, whereas Italian and Danish perceive no relation because their articles do not mention the concepts together.Thus, relatedness depends on which language’s article network supplies the world knowledge.
  • Results: The second experiment produced mean r = 0.16, compared with rMULT = .16 versus rENG16 = .77 for English-only simulations.The English-only mean was significantly higher than the multilingual mean, with p < 0.001.

DISCUSSION

Wikipedia knowledge diversity constrains applications that assume globally shared representations, while also enabling culturally-aware and hyperlingual designs. The discussion connects these implications to concept-level gaps, sub-conceptual cultural variation, and multilingual access to otherwise unavailable knowledge.

  • DISCUSSION: The lack of global consensus places boundary conditions on Wikipedia-based technologies, including information arbitrage and semantic relatedness systems.Concepts may be absent from some language editions, while sub-concept diversity can encode culture-specific information that should not automatically be propagated.
  • DISCUSSION: Wikipedia-based ESA can bias applications toward the world knowledge represented in the language edition used.For multinational groups, conversation clustering may differ depending on whether English, German, or another language edition supplies the semantic relatedness metric.
  • Implications for Design: Culturally-aware applications swap between representations of world knowledge as context demands.For semantic relatedness involving Romanian immigrants in the United States, a system could substitute Romanian Wikipedia for English Wikipedia.
  • Implications for Design: Hyperlingual applications extend culturally-aware design by considering multiple languages simultaneously.A weighted combination of participants’ native languages is one proposed extension for multilingual conversation analysis.
  • DISCUSSION: Hyperlingual access can expose world knowledge unavailable in any single language edition, including English.Articles outside the intersection of language editions far outnumber those in the intersection, creating opportunities to use multiple representations simultaneously.
  • Implications for Design: Future applications include culturally adapted writing support and systems for understanding concept and sub-concept diversity in Wikipedia.The proposed writing application identifies parochial or region-specific references and suggests alternatives for international audiences.

CONCLUSION

The paper identifies large cross-Wikipedia knowledge diversity, demonstrates effects on technology, and proposes culturally-aware and hyperlingual applications as design implications.

  • CONCLUSION: The paper’s four contributions quantify Wikipedia knowledge diversity, demonstrate its technological effects, census language’s impact on UGC, and introduce culturally-aware and hyperlingual applications.The authors present these contributions as a basis for future multilingual Wikipedia applications.
Loading 1904.01689v1…