Source-linked AI summary
Wiktionary as a Crowdsourced Lexicon for English Dialects
Sidney Wong
TL;DR
Dialect-responsive NLP still lacks sustainable resources for low-resource varieties and remains vulnerable to geographic bias. This paper evaluates Wiktionary across 12 English varieties against the OED and geo-referenced Reddit data, finding broad lexicon coverage and accurate regional-dialect identification when register and preprocessing are considered.
Problem
Dialect-responsive NLP lacks sustainable resources for low-resource varieties, while web-corpus reliance leaves systems vulnerable to geographic bias.
Method
The paper descriptively analyzes Wiktionary across 12 national English varieties, benchmarks it against the OED, and evaluates dialect features on geo-referenced Reddit data.
Results
Wiktionary broadly matches or exceeds OED coverage beyond American and British English, and identifies regional dialect usage accurately when register and preprocessing context are considered.
Takeaways & Limitations
Wiktionary is a viable, responsive complement to curated dictionaries and should supplement, not replace, existing corpora.
Takeaways & Limitations
The exploratory study relies heavily on descriptive analysis, and short unigrams can overlap with cross-linguistic homonyms or lose dialect distinctions without diacritics.
Abstract
from arXiv · showhide
This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects. We took a two-phase approach, providing an in-depth descriptive analysis of the crowdsourced lexicon for 12 national varieties of English before applying the lexicon to geo-referenced, country-level social media language data to examine the real-world performance of this crowdsourced dialect lexicon. We demonstrate that Wiktionary matches or exceeds the coverage of traditional dictionaries, such as the Oxford English Dictionary (OED), for regional and Outer-Circle varieties. Our dialect-specific case study on New Zealand English found high alignment between Wiktionary and the OED based on word-formation patterns (R = 0.883). Similarly, we observed high alignment between the dialect lexicon and geo-referenced social media language. While this paper found that Wiktionary has broad coverage of lexical properties, it also highlighted some of the macro-challenges involved in evaluating dialect-responsive language resources and tools, such as the role of language contact in dialects and register effects in web-based corpora.
1 Background and Motivation
Dialect-responsive NLP remains difficult because web-based data can misrepresent geographic variation and non-standard dialect lexicons are often excluded or unsustainably maintained. The section motivates Wiktionary as an open, structured resource for underserved dialects and frames evaluation against traditional lexicons and social-media usage.
- Motivation: Dialect-responsive NLP remains challenging for low-resource varieties because systems often rely on web corpora that may not represent dialect variation accurately.Geographic metadata from social media can support dialect analysis, but geographic information does not necessarily align with real-world linguistic variation.
- Wiktionary as a resource: Wiktionary is presented as an ethically crowdsourced, open-source resource that provides essential dialect lexical information in pluricentric English contexts.Its collaborative entries can include etymology, pronunciation, definitions, usage notes, derived forms, related terms, alternative forms, and translations.
- Motivation: Non-standard dialects risk exclusion from standard-language dictionaries, while traditional dialect lexicons can vary in size and quality.The New Zealand Dictionary Centre database contained approximately 42,000 entries before the centre became defunct.
- Motivation: Dialect lexicons remain important for NLP because out-of-vocabulary items can reduce model performance, and standard tokenisation and instruction tuning do not improve reasoning on non-standard English varieties.This sustains the need for structured lexicons and dictionary-based features for underserved dialects.
- Research questions: The study asks whether Wiktionary’s regional English coverage compares with the OED and whether Wiktionary-derived features identify regional dialect usage across social-media registers.These questions connect lexicon coverage with real-world dialect identification in social-media text.
2 Methodology
The study uses a two-phase design to compare Wiktionary with the OED and validate the crowdsourced lexicon against geo-referenced Reddit language data. It focuses on twelve national English varieties, with New Zealand English as a detailed case study.
- Lexicon Construction: Wiktionary’s variety metadata organizes English entries into geographic categories and national-variety subcategories, including broader and more specific classifications.The study uses these categories to construct dialect-specific lexicons.
- Corpus and Varieties: The Reddit sample comprises country-level communities representing twelve national English varieties, divided into six Inner-Circle and six Outer-Circle varieties.Corpus characteristics are summarized separately for submission posts and comments.
- Phase 1: Descriptive Analysis: The analysis first compares Wiktionary and OED coverage across national dialect varieties, then manually codes New Zealand English entries by word-formation process for cross-validation.Phase 1 addresses the first research question through descriptive comparisons and an NZE-specific case study.
- Phase 2: Corpus Validation: The corpus-validation phase compares relative word frequencies in Reddit data, measured as occurrences per 100,000 words (µi) to account for differing community sizes.This phase evaluates whether crowdsourced dialect entries align with real-world, geo-referenced language use.
- Data Processing: Wiktionary entries were filtered to remove items longer than three words and duplicates, while Reddit observations were stripped of punctuation and special characters before lowercasing.The procedure retained ampersands and dollar signs and considered NZE variation in diacritic use.
3 Results
Wiktionary broadly covers English dialect varieties, including all six Outer-Circle national varieties, while complementing the OED’s stronger coverage of American and British English. Its New Zealand English entries and country-level subreddit patterns show substantial but context-dependent alignment, varying between submission posts and comments.
- Dialect coverage: 13.1%; the OED is larger overall, with 37,679 entries versus Wiktionary’s 33,043.The OED also exceeds Wiktionary’s coverage of American and British English by 33.8% and 43.5%, respectively, while Wiktionary covers regional Inner-Circle varieties more strongly.
- Dialect coverage: Wiktionary covers all six Outer-Circle national varieties, whereas the OED lacks Kenyan and Pakistani English entries.The OED contains entries for South African, Indian, Malaysian, and Filipino English.
- New Zealand English: 1,541; Wiktionary’s two NZE-related categories contained 1,541 entries, reduced to 1,332 after data-processing steps.The te reo Māori borrowing category contained 324 entries, and 498 processed entries were unique to NZE.
- Social media alignment: Strong alignment appeared between each Inner-Circle variety and its associated country-level subreddit in Submission Posts, including r/usa despite its small subscriber base.The relationship did not hold for Comments; Irish, Canadian, and Australian English were the only varieties maintaining consistency across both types of content.
4 Discussion
The discussion finds that Wiktionary broadly complements the OED in dialect coverage and can identify regional social-media usage, provided register and preprocessing effects are addressed. It also highlights challenges from historical coverage, language contact, and differences between posts and comments.
- Lexicon coverage: Across six Inner-Circle varieties, Wiktionary contains 33,043 entries versus the OED’s 37,679, a 13.1% difference.The OED is more comprehensive for American and British English, while Wiktionary covers the other varieties more extensively.
- Social-media evaluation: Wiktionary-derived features accurately identify regional dialect usage on social media when register context and character-level preprocessing are controlled.The pipeline should prioritize titles over comment threads, preserve diacritics, and filter cross-linguistic homonyms.
- Lexicon coverage: Wiktionary complements the OED through accessibility and coverage, particularly for Outer-Circle varieties and regional loanwords.The OED maintains broader historical coverage of primary US and UK standard varieties.
- Lexical properties: In NZE idiomatic phrases, Wiktionary covered more expressions, but the OED better reflected their actual usage.Of nine multi-word expressions, only “suck the kumara” and “turn to custard” were absent from the OED.
- Language contact: NZE features were over-represented in r/Philippines comments, illustrating how borrowings, loanwords, and shared lexical roots complicate dialect analysis.The inspected features were mo, pa, para, eh, and po; te reo Māori and Tagalog both belong to the Austronesian language family.
- Register effects: Submission Posts and Comments diverged lexically because posts may signal regional context, whereas comments are interactive conversational discourse.This register difference contributed to closer alignment between national dialect features and corresponding subreddits, particularly for unigrams, in Submission Posts.
5 Conclusion
Wiktionary is presented as a viable, highly responsive complement to curated dictionaries for English varieties. The conclusion recommends using crowdsourced dialect lexicons to supplement, rather than replace, existing corpora while addressing challenges from language contact and register effects.
- Wiktionary serves as a viable, highly responsive complement to curated dictionaries for English varieties.
- Evaluating and validating the crowdsourced dialect lexicon reveals broader digital-corpus challenges involving language contact in dialects and register effects.
- Crowdsourced dialect lexicons should supplement existing corpora, not replace them.The conclusion also calls for future evaluation of crowdsourced lexicon performance.
Limitations
The paper is exploratory and relies heavily on descriptive analysis. It acknowledges methodological limitations in using Wiktionary as a crowdsourced dialect lexicon, including orthographic overlap between short dialect entries and high-frequency Tagalog words.
- The study is exploratory and relies heavily on descriptive analysis.
- Several methodological limitations affect the use of Wiktionary as a crowdsourced dialect lexicon.
- Short unigram entries such as mo and po can overlap orthographically with high-frequency Tagalog words.Wiktionary lists mo (‘moustache’) and po (‘a chamber pot’) under New Zealand English, while Tagalog uses them differently.
Ethics Statement
The paper frames Wiktionary as a resource for examining the changing role of lexicography, not as a means to overexploit crowdsourced language data or replace lexicographers with volunteers.
- Ethics Statement: The authors reject overexploiting crowdsourced language data and shifting lexicographical work to volunteers, instead presenting lexicons as support for changing lexicographic practice.This position is situated against the decline of the digital commons associated with proprietary frontier LLMs.