Source-linked AI summary
The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language
MohammadHossein Mortazavi, Mostafa Salehi, Hadi Veisi
TL;DR
The paper addresses whether Persian should be called low-resource when it has substantial raw text and many datasets but uneven annotation coverage and usability. It reviews Persian resources and combines web, speech, and task-matched comparative evidence. The resulting diagnosis is qualified: Persian has strong annotation islands, but reusable, domain-matched, interoperable, and well-documented supervision remains uneven across tasks and varieties.
Problem
Existing low-resource labels collapse distinct shortages, while Persian’s many datasets and substantial digital presence leave its annotation and resource adequacy insufficiently synthesized.
Method
The paper conducts a structured review of 34 Persian text resources with independent web measurements, a selective speech cross-check, and task-matched Persian-English annotation-density comparisons.
Results
The evidence shows that Persian is not uniformly annotation-poor: selected syntax and news NER resources are comparatively dense, while natural-language inference falls below the web-proportional baseline.
Takeaways & Limitations
Annotation-scarce is useful as a diagnostic of recurring gaps in task, domain, access, interoperability, documentation, and variety coverage rather than as a ranking of total labeled volume.
Takeaways & Limitations
The review is not a complete census, focuses mainly on standard Iranian Persian text, and uses selective speech coverage and changing web proxies.
Abstract
from arXiv · showhide
Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.
1 Introduction
Persian illustrates why “low-resource” is too coarse: substantial digital and annotated resources coexist with uneven coverage, usability, and standardization. The paper therefore evaluates annotation scarcity as a qualified, task- and resource-ecology description.
- 1 Introduction: Persian has broad resources across corpora, treebanks, parallel data, NER, sentiment, discourse, and language-understanding benchmarks.Its web presence is also measurable through W3Techs and Common Crawl, so raw pretraining volume alone is insufficient for assessment.
- 1 Introduction: 0.9% of websites with a known content language and 0.7039% of Common Crawl HTML pages were measured as Persian.The two measurements place Persian within roughly the top twenty languages under different web methodologies.
- 1 Introduction: Resource counts overstate readiness because many datasets share domains while differing in annotation schemes and access conditions.Licenses, formats, manuals, download locations, and version histories also affect reuse.
- 1 Introduction: The paper combines a 34-resource text inventory, web measurements, a selective speech cross-check, and a task-matched Persian-English comparison.It is a structured review and position paper with a quantitative comparative component, not an exhaustive census or controlled causal study.
- 1 Introduction: Annotation-scarce describes recurring shortages that are too small, narrow, inaccessible, or weakly documented for reliable training and evaluation.At language-resource level, the diagnosis applies when such shortages recur across important tasks, domains, and varieties.
- 1 Introduction: Persian’s main empirical scope is standard Iranian Persian text, while speech is treated as a selective cross-check and Dari and Tajik are not inventoried.Pooling related varieties would overstate coverage of the reviewed resources.
3 Related Work
Related work supplies broad low-resource concepts, documentation principles, task-specific Persian resources, and model evaluations, but not a unified ecosystem diagnosis. This paper synthesizes these strands to distinguish raw-data availability from annotation, access, interoperability, and coverage problems.
- 3 Related Work: Broad low-resource taxonomies support task-sensitive resource analysis but do not directly diagnose Persian’s uneven annotation ecology.Existing perspectives address multilingual resource inequality or labeled data for particular modeling setups.
- 3 Related Work: Dataset documentation treats provenance, coverage, annotation, intended use, licensing, and limitations as determinants of scientific reusability.The review adopts this broader adequacy perspective rather than evaluating resources by example count alone.
- 3 Related Work: Persian resources developed through corpus- and task-specific projects spanning information retrieval, morphosyntax, web and literary text, and translation.This history expanded both raw-text and annotated-resource coverage.
- 3 Related Work: Speech-resource development progressed from FARSDAT and multilingual telephone corpora to larger microphone, emotional, accented, child-speech, and text-to-speech collections.Recent public resources further increased scale across speaker verification, ASR, audio-visual speech, and crowdsourced recordings.
- 3 Related Work: Persian task supervision now covers NER, affect, pragmatics, web register, language understanding, inference, question answering, summarization, reasoning, and punctuation restoration.These resources remain distributed across particular tasks and construction traditions.
- 3 Related Work: Most publications introduce one resource or benchmark, leaving raw corpora, adjudicated labels, conversions, generated labels, and evaluation sets difficult to compare uniformly.They also differ in domain, licensing, maintenance, tokenization, and compatibility.
- 3 Related Work: Persian model capability is neither absent nor uniform, but benchmark performance alone cannot identify annotation scarcity as the cause of every error.Pretraining composition, tokenization, model scale, prompting, contamination, and domain shift also affect results.
- 3 Related Work: The paper’s contribution is a structured, task-sensitive synthesis explaining how Persian can have substantial datasets while remaining qualifiedly annotation-scarce.It connects broad taxonomies, dataset research, and model benchmarks into a shared resource profile.
4 Resource Review Method
The review builds a documented, representative inventory and supplements it with independent web measurements and a task-matched Persian-English comparison. These procedures preserve task boundaries and treat web shares as proxies rather than token counts.
- 4 Resource Review Method: The inventory draws on major NLP venues, repositories, Universal Dependencies, and publication-linked sources using Persian or Farsi task and resource searches.Included resources had to identify Persian as a target language and report enough information about task, construction, or scale.
- 4 Resource Review Method: Distinct corpus conversions are listed separately only when they change interoperability or public access.This rule applies, for example, to Universal Dependencies versions of Persian treebanks.
- 4 Resource Review Method: The review favors publicly described resources with traceable documentation but does not assume publication guarantees current downloadability, licensing clarity, or easy integration.Commercial, unpublished, and institution-internal datasets may alter practical resource availability.
- 4 Resource Review Method: W3Techs estimates Persian’s share among websites with identified content languages, while Common Crawl uses CLD2 to detect the primary language of HTML pages.Their denominators and collection processes differ, so the percentages are not averaged or treated as token counts.
- 4 Resource Review Method: Persian-English comparisons include only tasks with comparable source-reported units and retain tokens, sentences, documents, and task instances as separate measures.The resulting ratios are dimensionless, task-specific, illustrative, and comparator-sensitive.
5 Raw Language Resources and Corpus Diversity
Persian has substantial raw text, web presence, domain-specific, parallel, variation-oriented, and speech resources, but their coverage and supervision remain uneven across research needs.
- Persian in the global web ecosystem: 0.9% of websites and 0.7039% of HTML pages were measured as Persian in independent July 2026 web datasets.W3Techs counts websites, whereas Common Crawl counts crawled HTML pages; the measures have different denominators.
- General and domain corpora: 72.9 billion preprocessed and deduplicated tokens in Matina, alongside nearly 20 million hmBlogs posts, make raw-text absence difficult to defend.These corpora increase scale but do not resolve representativeness across clinical notes, contracts, private dialogue, regional usage, or documented colloquial text.
- General and domain corpora: 112 poets and 57,980 literary items provide diachronic, stylistic, authorship, and literary-language coverage unavailable from ordinary news or crawl data.The Corpus of Persian Literary Text includes partial rhetorical-figure annotation and spans the ninth to the twenty-first century.
- Parallel and variation-oriented corpora: 554,621 aligned subtitle lines and 1,011,085 Persian-English sentence pairs support translation but retain strong subtitle and translated-literature domain signatures.TEP contains approximately 3.71 million Persian words, while MIZAN is drawn mainly from literary translations.
- Parallel and variation-oriented corpora: Six million colloquial-formal sentence pairs support style transfer and robustness, but do not completely describe naturally occurring spoken-style Persian.HarfoSokhan is a parallel variation resource whose construction and task differ from independently written colloquial corpora.
- Interpreting raw-resource adequacy: Raw-resource adequacy depends on represented distributions, legal and technical reusability, and evaluation-set independence, not volume alone.A billion web tokens and a million literary translation pairs serve different research needs and should not be collapsed into one resource count.
- Speech resources as complementary evidence: 540,000 DeepMine recordings, almost 220 hours from Arman-AV, and 430.86 recorded Common Voice hours show substantial speech scale without uniform supervision.Speech resources differ in transcripts, speaker metadata, phoneme boundaries, prosodic labels, spontaneous-speech coverage, dialect labels, and task-specific evaluation.
6 Annotated Datasets and Evaluation Resources
Persian has substantial annotated datasets across linguistic, sequence-labeling, semantic, discourse, classification, and evaluation tasks. However, these resources differ in annotation provenance, compatibility, accessibility, and domain coverage, so published volume does not directly equal reusable supervision.
- Linguistic and sequence annotation: Persian dependency resources exceed 645,000 tokens across Persian-PerDT and Persian-Seraji, but their origins, genres, conversion histories, and tokenization differ.These differences complicate direct combination despite the substantial aggregate size.
- Entities, sentiment, pragmatics, and web register: PEYMA, ARMAN, and NSURL-2019 provide established news-domain named-entity resources spanning six or seven entity classes and more than one million NSURL-2019 tokens.NSURL-2019 contains 1,029,822 tokens across its training and test splits.
- Entities, sentiment, pragmatics, and web register: Persian sentiment, emotion, irony, and register resources extend annotation beyond product reviews to e-commerce, social media, tweets, and web genres.SentiPers, MirasOpinion, ArmanEmo, the Persian irony corpus, and ParsCORE cover multiple labeling targets and source types.
- NLU, question answering, and summarization: Persian evaluation resources now span NLU, inference, question answering, summarization, reasoning, knowledge, and punctuation restoration.Examples include ParsiNLU, FarsTail, PerCQA, pn-summary, PersianMMLU, FarsEval-PKBETS, PARSE, and PersianPunc.
- Semantic roles, discourse, coreference, and lexical meaning: PerPB and Persian discourse, coreference, and RST resources demonstrate substantial annotation beyond sentence-level classification, but their frameworks, availability, genres, and evaluation protocols remain fragmented.PerPB contains 29,982 manually annotated sentences, while discourse and coreference resources include approximately 30,000 sentences and 547 coreference documents.
- Annotation depth and practical availability: Automatic, converted, crowdsourced, synthetic, and manually adjudicated resources represent different kinds of supervision and should not be summed as equivalent annotation.Automatically sense-tagged data are not equivalent to adjudicated semantic gold standards, and synthetic corpora should not be treated as unquestioned independent tests.
- Annotation depth and practical availability: Published-resource counts overstate practical availability because links, licenses, preprocessing software, and source overlap can prevent retrieval and combination.For Persian NLP, discoverability and maintenance are part of the annotation bottleneck.
7 Linguistic and Technical Sources of Annotation Difficulty
Persian annotation difficulty arises from script, spacing, morphology, register, variety, and domain variation. These differences make explicit normalization, segmentation, and independently sampled data important for reliable coverage.
- Script normalization and tokenization: Arabic and Persian Unicode variants can be mixed across corpora, and normalization rules require documentation because normalization affects downstream modeling results.The distinction is operational rather than merely typographic.
- Script normalization and tokenization: ZWNJ and spacing variants can cause tokenizers to segment equivalent forms differently, while UD treebanks also differ in clitic and multiword-token representation.Combining corpora without reconciling these decisions can introduce systematic label noise.
- Morphology and ezafe: Annotation must address Persian-specific morphology, including verbal prefixes, agreement suffixes, enclitics, plural and possessive morphology, light verbs, and the ezafe linker.Ezafe is often not overtly written, so some annotation settings may need to recover unmarked information.
- Register, domain, and variety: Formal written Persian, colloquial Tehran Persian, and social-media Persian differ in grammar, vocabulary, orthography, spacing, script use, emoji, and creative spelling.Existing resources improve variation coverage but do not eliminate the need for independently sampled conversational data.
8 A Task-Level and Cross-Language Diagnosis
Persian’s annotation profile is uneven rather than uniformly deficient: several mature tasks are comparatively well supplied, while reuse is constrained by compatibility, access, documentation, and coverage gaps. Task-specific web-normalized comparisons therefore support a qualified, coverage- and usability-based diagnosis.
- 8 A Task-Level and Cross-Language Diagnosis: Persian has substantial annotation in several mature tasks, but important parts of its NLP ecosystem remain missing, mismatched, incompatible, inaccessible, or weakly documented.The review identifies cross-resource compatibility, public access, documentation, specialist domains, independent evaluation, and representation beyond standard Iranian Persian as recurring bottlenecks.
- 8 A Task-Level and Cross-Language Diagnosis: Dataset fragmentation, annotation cost, access constraints, and academic incentives are plausible contributors, but this review does not measure their independent effects.The paper states that causal claims about these factors would require interviews, funding and publication analyses, or practitioner surveys.
- 8 A Task-Level and Cross-Language Diagnosis: WNAD compares selected Persian and English annotation volumes within each task against their relative Common Crawl web presence.The measure is a resource-level diagnostic, not a universal quality score, and uses documented canonical resources rather than exhaustive ecosystem totals.
- 8 A Task-Level and Cross-Language Diagnosis: Dependency parsing, selected news NER, and news summarization lie above the web-proportional baseline, whereas FarsTail NLI lies below it.Because the ratios vary sharply across tasks, the paper reports no cross-task average.
- 8 A Task-Level and Cross-Language Diagnosis: The quantitative comparison shows that annotation-scarce should not mean that Persian’s total labeled volume is small relative to its digital footprint.For several mature tasks, selected Persian annotation volume is not below the web-proportional expectation.
9 Large Language Models and the Annotation Bottleneck
Persian-specific and multilingual language models have expanded the ecosystem, but benchmark evidence still shows uneven capability and continued need for reliable gold data. Synthetic resources may accelerate construction, yet they require human validation, provenance records, and separation from evaluation data.
- 9 Large Language Models and the Annotation Bottleneck: PersianMind and multilingual instruction-tuning projects such as Aya show that Persian-specific open models and community-contributed instruction data are part of the current ecosystem.PersianMind extends a Llama-family model with Persian vocabulary and nearly two billion Persian tokens.
- 9 Large Language Models and the Annotation Bottleneck: Benchmark results indicate uneven capability, with general-purpose models often behind Persian-fine-tuned systems and translation improving GPT-3.5 results on some Persian test items.PersianMMLU provides an original Persian evaluation set, while FarsEval-PKBETS reported average accuracy below 50 percent for three evaluated Persian-capable models.
- 9 Large Language Models and the Annotation Bottleneck: Synthetic examples, model-proposed labels, and generated questions can accelerate data construction, but they require human validation, provenance records, and separate development and evaluation data.A large synthetic corpus may improve training while remaining unsuitable as an independent gold benchmark.
10 Discussion
Persian has substantial raw and labeled resources, but annotation availability and usability vary sharply across tasks, domains, and modalities. The paper therefore treats annotation scarcity as uneven resource friction rather than uniformly low volume.
- Evidence and interpretation: 0.7039% of Common Crawl HTML pages were Persian, yet labeled volume scales unevenly across tasks relative to web presence.Selected Persian syntax, news NER, and summarization resources are large relative to the web baseline, whereas NLI is not.
- Evidence and interpretation: Hundreds of hours—and ParsVoice’s TTS subset exceeding 1,800 hours—show that raw speech volume does not guarantee phoneme, prosodic, dialect, or evaluation annotations.The cross-check extends the raw-versus-annotation distinction beyond text.
- Evidence and interpretation: Resource friction arises because datasets use different tokenization rules, label inventories, genres, licenses, and release practices.The reviewed landscape resembles valuable projects more than coordinated infrastructure.
- Evidence and interpretation: Large web and blog corpora support representation learning but cannot replace task-specific labels for coreference, semantic roles, specialist entities, or human preferences.Reasoning benchmarks may reveal failures without supplying the supervision needed to correct them.
- Implications: Large language models shift the bottleneck toward reliable evaluation, specialist annotation, preference data, safety judgments, and colloquial or regional coverage.Raw-text scale can reduce supervision needs for some applications but does not make gold data irrelevant.
- Implications: Annotation-scarce is a diagnostic description: Persian has strong annotation islands, but reusable, documented, domain-appropriate supervision remains uneven.The diagnosis indicates whether effort should target annotation, harmonization, licensing, maintenance, or evaluation.
11 Recommendations
The recommendations prioritize infrastructure that makes existing Persian resources discoverable, interoperable, documented, and reusable. New annotation should target measurable gaps rather than duplicate mature datasets.
- Infrastructure: A public registry should record each resource’s task, variety, domain, size, annotation method, guidelines, agreement, license, identifier, status, version, and overlap.The proposal is intended as infrastructure rather than another isolated list.
- Priorities: Semantic-role and discourse resources should be made easier to locate, license, convert, benchmark, and extend rather than recreated from nothing.New annotation should focus on multi-domain NER, clinical and legal text, natural conversation, regional varieties, stance, and cross-domain evaluation.
- Priorities: Mature tasks may benefit more from harmonization and carefully designed test sets than another dataset from the same news distribution.This recommendation follows the documented concentration of existing resources.
- Documentation: Manually labeled resources should release guidelines, training and adjudication procedures, and task-appropriate agreement statistics.Dataset splits, normalization rules, and tokenization versions should be fixed and versioned.
- Annotation practice: Active learning and model pre-annotation can reduce human labeling effort only when annotators can reject suggestions and evaluation sets are independently checked.Error analysis should test effects on rare entities, nonstandard spellings, minority varieties, and sensitive content.
- Governance: Projects involving sensitive Persian data should document lawful use, privacy, compensation, harmful-content exposure, and regional or social representation.Guidelines should permit disagreement when the task is subjective.
12 Limitations
The review provides a representative map of Persian resources, not a complete or independently verified census. Its scope, measurements, comparisons, and causal interpretations are deliberately bounded.
- Scope: The review inventories 34 representative text resources and a selective speech set, excluding proprietary and institution-internal datasets.Quality assessment relied on published documentation rather than re-annotation or reproducing every construction pipeline.
- Usability: Publication does not certify present-day usability because links, licenses, preprocessing code, formats, splits, duplicates, and overlaps were not independently verified.The inventory should be read as documented work, not frictionless access.
- Scope: The main scope is standard Iranian Persian text NLP, with speech treated as a selective cross-check and Dari and Tajik left outside the inventory.Web statistics are changing proxies based on websites or crawled HTML pages and automatic language identification, not clean training-token estimates.
- Quantitative comparison: A high WNAD indicates that a selected Persian resource is large relative to a matched benchmark and web baseline, not parity in coverage, quality, licensing, or the wider ecosystem.The comparison depends on the chosen English comparator and canonical datasets such as CoNLL-2003.
- Causal scope: The paper does not establish why particular resource gaps arose, because causal explanations require evidence beyond the reviewed resource counts.Possible influences such as funding or geopolitical restrictions remain untested.
13 Conclusion
Persian is annotation-scarce only in a qualified, coverage-based NLP sense: it has substantial digital and task-specific resources, but uneven reusable supervision. The label is diagnostic, not a ranking of total labeled volume.
- Conclusion: Persian ranks among roughly the top twenty web languages, while recent speech resources reach hundreds or thousands of hours.These measurements do not imply uniform annotation coverage.
- Conclusion: The annotation-scarce description identifies mismatches between digital presence and reusable, domain-matched, interoperable, well-documented supervision.Resource development should be planned across tasks, domains, modalities, construction methods, accessibility, and varieties.