Source-linked AI summary

A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study

Mullosharaf K. Arabov

arXiv:2608.23421v2cs.CL

TL;DR

Arabic NLP lacked a comprehensive quantitative account of its growth, topics, citation patterns, and task–dialect coverage. The study analyzes a large multi-source corpus with bibliometric, topic, regression, network, and geographic methods, finding rapid post-2020 expansion, concentration in foundational tasks, and persistent dialectal and evaluation gaps.

  • Problem

    Arabic NLP lacked a comprehensive bibliometric and topic-based meta-analysis to quantify field growth, research distribution, citation structure, and task–dialect gaps.

  • Method

    The study analyzes Arabic NLP publications using BERTopic alongside bibliometric, regression, network, and geographic analyses.

  • Results

    82% of examined papers appeared in the most recent period, while the largest theme covers text processing, speech, translation, and recognition at 41.3% of the corpus.

  • Takeaways & Limitations

    The findings provide quantitative validation of survey claims while identifying under-resourced dialects and gaps requiring attention in Arabic NLP research and evaluation.

  • Takeaways & Limitations

    Metadata incompleteness and source coverage may bias geographic, institutional, and publication analyses toward well-indexed venues and resource-rich countries.

Abstract

from arXiv · show

Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis remains absent. This study presents a bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from five platforms (arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex) plus an additional targeted OpenAlex subset. We employ BERTopic for topic modeling, regression analysis, social network analysis, and geographic mapping. Our findings show a significant publication surge after 2020, driven by transformer models and LLMs. Topic modeling identifies 19 themes, the largest centered on text, speech, translation, and recognition (2,942 papers). Citation analysis reveals a positive correlation between paper age and citations (r = 0.245, p < 0.001); regression (R^2 = 0.105) shows that indexing in OpenAlex or Semantic Scholar and institutional affiliation are associated with higher citations. Saudi Arabia, the United States, and Egypt lead in research output. A task-dialect gap matrix identifies understudied areas, including summarization for Maghrebi, Iraqi, and Sudanese dialects. The largest topic has the highest H-index (90), followed by sentiment analysis (57). Our quantitative approach complements existing qualitative surveys and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks.

1 Introduction

Arabic NLP has expanded rapidly, but the field lacked a comprehensive quantitative synthesis. This study addresses that gap with large-scale bibliometric, topic, citation, network, and geographic analyses.

  • The study analyzes Arabic NLP research quantitatively because existing claims about growth, topic dominance, and dialect neglect remained largely unverified.The authors position bibliometric methods as a complement to qualitative surveys.
  • 7,120 papers from 1960–2026 were analyzed using BERTopic, regression, social network analysis, and geographic mapping.The dataset combines arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex, and a targeted OpenAlex subset.
  • 19 research themes were identified, with text, speech, translation, and recognition forming the largest cluster at 2,942 papers.
  • Publication year, indexing source, and institutional affiliation significantly predicted citation counts, with OpenAlex, Semantic Scholar, and affiliation associated with higher averages.The reported average associations are approximately +11, +5, and +8.7 citations, respectively.
  • Saudi Arabia, the United States, and Egypt led research output, with King Saud University and Cairo University the most productive institutions.The country totals were 519, 463, and 266 affiliations; the institution totals were 142 and 67 papers.
  • Task–dialect analysis identified summarization for Maghrebi, Iraqi, and Sudanese Arabic as understudied areas.Dialect identification for Sudanese and Yemeni Arabic was also identified as understudied.

2 Literature Review

Existing Arabic NLP surveys provide detailed qualitative coverage, but they remain fragmented and do not quantify research distribution, topic growth, or network structure. The literature also highlights persistent challenges in dialectal coverage, data resources, evaluation, and culturally aligned Arabic LLM benchmarks.

  • Survey coverage: Existing surveys cover Arabic NLP foundations, preprocessing, dialects, sentiment, code-switching, LLMs, linguistic tasks, translation, question answering, speech, and document processing.
  • Dialectal NLP: Sentiment analysis dominates dialectal NLP literature at approximately 32%, followed by resource building at 21% and identification and code-switching at 10%.
  • Dialectal NLP: Arabic dialect research is concentrated on varieties such as Moroccan, Saudi, Egyptian, and Algerian Arabic, while many dialects remain under-resourced.
  • Large language models: Arabic LLM reviews identify limited transparency, concentration on MSA, insufficient multi-turn and temporal evaluation, and cultural misalignment in translated datasets.
  • Research gap: Most prior surveys are qualitative and do not measure research activity, citation impact, topic growth, collaboration, or citation-network structure.
  • Research gap: The literature lacks a unified, data-driven synthesis integrating tasks, dialects, and methodological paradigms.Fragmentation is especially consequential during rapid publication growth and shifting research priorities.
  • Study contribution: The study complements qualitative synthesis by supplying quantitative validation and revealing patterns that manual review may not show.

3 Methodology

The study builds and analyzes a curated Arabic NLP corpus through multi-source collection, filtering, deduplication, topic modeling, regression, and diagnostic checks. The methodology prioritizes broad coverage while addressing relevance, metadata quality, and citation-model limitations.

  • Data acquisition: Five platforms plus a targeted OpenAlex subset supplied records for a unified Arabic NLP dataset.Sources covered preprints, conference papers, journal articles, and workshop contributions.
  • Filtering and integration: Papers were retained when titles or abstracts contained at least two Arabic NLP-specific terms, with additional pedagogy-aware filtering to reduce educational noise.The term list covered dialectology, morphology, speech processing, translation, and language-model developments.
  • Filtering and integration: 16,380 relevant records were loaded, 12,507 remained after abstract filtering, and 7,120 papers remained after cleaning and date restrictions.Papers without abstracts or with abstracts shorter than 10 characters were removed, followed by a 1960-onward publication filter.
  • Filtering and integration: Two-stage DOI-based deduplication merged records by selecting richer metadata fields and aggregating authors, affiliations, concepts, and citation counts.The integration combined records from multiple sources into a unique master dataset.
  • Analytical methods: BERTopic was compared with LDA, and BERTopic outperformed the baseline on topic diversity and coherence.BERTopic uses contextual embeddings, whereas LDA provides a traditional bag-of-words comparison.
  • Analytical methods: The citation regression was statistically significant (F = 92.83, p < 0.001) but explained 10.5% of citation-count variance (R^2 = 0.105).Residuals were right-skewed, violating OLS normality assumptions; the authors identify count-based models as a future alternative.

4 Results

The final corpus contains 7,120 papers spanning 1960–2026, with publication activity accelerating sharply after 2015 and especially after 2020. AraBERT and the later ChatGPT/LLM boom are identified as major inflection points in this growth.

  • Corpus overview: 7,120 unique papers published between 1960 and 2026 formed the final longitudinal corpus.The dataset spans more than six decades of Arabic NLP research activity.
  • Corpus overview: Semantic Scholar contributed 31.4% of records and OpenAlex 24.2%, while 68.5% had citation data and 36.1% had institutional or country information.The mixed-source corpus broadens coverage, but citation and affiliation metadata remain uneven.
  • Publication trends: Annual output stayed below 50 papers for most years until approximately 2015, followed by exponential acceleration after 2020.The early period emphasized rule-based approaches, morphology, and early statistical methods.
  • Publication trends: 2018: AraBERT catalyzed growth by bringing Transformer architectures to Arabic NLP and expanding downstream task research.The reported downstream areas include sentiment analysis, named entity recognition, and machine translation.
  • Publication trends: 2023: The ChatGPT and LLM boom accelerated output, with research shifting toward instruction-tuned models, benchmarks, and safety studies.The publication curve reached an unprecedented peak in 2025.
  • Publication trends: Approximately 82% of papers were published after 2020, concentrating recent activity in modern deep-learning and generative-AI research.The concentration also raises questions about the sustainability of this growth.

4.3 Most Productive Authors

Arabic NLP productivity is concentrated among a small, collaborative core, while leading researchers show distinct trajectories spanning foundational resources, deep learning, dialectal work, and speech processing.

  • 151 papers place Nizar Habash first among the most productive authors, followed by Muhammad Abdul-Mageed with 95 and Mona Diab with 86.
  • Large annotation projects and benchmark series foster collaboration among high-output researchers and support dataset and standard creation.
  • Nizar Habash maintained sustained productivity across the period, spanning classical morphological resources and modern neural models.
  • Muhammad Abdul-Mageed reached 20 papers in each year of 2022–2023, coinciding with the transformer and LLM research surge.
  • Wajdi Zaghouani peaked at 12 papers in 2016, Mona Diab at 9 in 2013, and Ahmed Ali at 10 around 2020.

4.6 Citation Analysis

Citation impact reflects both accumulated paper age and topic-specific influence. Core text, speech, translation, and recognition research has the strongest aggregate impact, while specialized topics achieve higher average citations per paper.

  • 1,480 citations make the AraBERT paper the most cited work in the dataset.
  • r = 0.245, p < 0.001 indicates a positive, statistically significant correlation between paper age and citations.
  • The text, speech, translation, recognition topic has the highest H-index at 90, followed by sentiment analysis at 57 and patient-related studies at 42.
  • Topic 18 has the highest average citations at 29.90, followed by Topic 6 at 24.90 and Topic 2 at 21.84.

4.7 Regression Analysis: Factors Affecting Citation Impact

The citation regression identifies publication year, indexing source, institutional affiliation, matched-term coverage, and the extra OpenAlex subset as statistically significant correlates, while explaining a limited share of citation variation.

  • R^2 = 0.105 shows that the statistically significant OLS model explains 10.5% of citation-count variance.
  • -1.234 is the publication-year coefficient, indicating that older papers receive more citations when other factors are held constant.
  • +11.09 and +5.45 citations are associated with OpenAlex and Semantic Scholar indexing, respectively.
  • +8.73 citations are associated with institutional affiliation, while each additional matched term is associated with +1.21 citations.
  • -9.042 is the coefficient for is_OpenAlex_(extra), indicating fewer citations relative to papers not in that source.

4.8 Geographic Distribution

Arabic NLP research is geographically concentrated in Saudi Arabia, the United States, and Egypt, within a broader international and Francophone collaboration network. However, affiliation rankings cover only a metadata-available subset and may underrepresent several Arabic-speaking countries.

  • 519 Saudi Arabian affiliations, 463 United States affiliations, and 266 Egyptian affiliations make these the top three observed contributors.
  • Egypt, Jordan, the UK, the UAE, Morocco, France, and Algeria contribute to a global network that includes a distinct Francophone Maghrebi research cluster.
  • The UAE and Qatar rank among the top ten and include major institutions such as MBZUAI and Qatar University.
  • 36.1% of papers have available affiliation metadata, so rankings may favor wealthier institutions and underrepresent Iraq, Sudan, Yemen, and other regions.

4.9 Institutional Distribution

Arabic NLP output is concentrated in a small group of highly productive institutions, while dialect research remains dominated by MSA and Egyptian Arabic.

  • 142 papers place King Saud University first, followed by Cairo University with 67 and Columbia University with 64.
  • King Saud University accounts for nearly 6% of papers with affiliation data, associated with its active NLP group and governmental AI support.
  • The top institutions combine major global universities with regional powerhouses, indicating strong concentration among elite, well-resourced organizations.
  • Institutional research profiles span Arabic linguistics and rule-based NLP alongside statistical and neural methods, often through collaboration.
  • Dialect distribution: MSA consistently dominates dialect mentions, while Egyptian and Maghrebi show greater growth than substantially underrepresented dialects.
  • Dialect distribution: MSA aligns strongly with the largest text, speech, translation, and recognition topic, whereas dialect-specific NER remains limited.

4.11 Task–Dialect Gap Matrix

The task–dialect matrix shows extreme concentration on MSA and identifies zero-coverage dialect–task combinations, especially for specialised tasks such as summarisation.

  • The matrix is constructed by row-normalizing the percentage of abstracts mentioning each task–dialect pair.
  • MSA appears in over 50% of papers across all core tasks, while Egyptian Arabic has double-digit coverage across most tasks.
  • Hejazi and Hassaniya have zero coverage across all five core tasks, while Sudanese and Yemeni appear only sporadically.
  • Specialised tasks: Summarisation is concentrated in MSA at 76.9%, with Egyptian at 15.4% and Gulf at 7.7%; all other listed dialects have zero coverage.
  • Specialised tasks: OCR has 33.3% Maghrebi coverage but zero coverage for Gulf, Iraqi, Sudanese, Yemeni, Najdi, Hejazi, and Hassaniya.

4.12 Supplementary Statistical Analyses

Supplementary analyses confirm that Arabic NLP topics vary across decades, core topics continue growing, and matched-term breadth has only a weak relationship with citation impact.

  • χ2 = 75.04, df = 8, p < 0.001 shows that topic distributions differ significantly across the 1960–2009, 2010–2019, and 2020–2026 periods.
  • CAGR = 11.54% makes text, speech, translation, and recognition the fastest-growing top topic over the last five years.
  • Health-related topics grew at CAGR = 5.63%, sentiment analysis at 3.69%, while NER declined at CAGR = -5.26%.
  • The growth-rate snapshot is sensitive to the five-year window and topic-assignment threshold.
  • Concepts and terms: The controlled-vocabulary concepts emphasize linguistics, philosophy, Arabic, artificial intelligence, natural language processing, and speech recognition.
  • Concepts and terms: r = 0.036 indicates a weak positive correlation between matched-term count and citations, while the regression coefficient is +1.208.
  • Concepts and terms: Topic-specific matched terms support semantic coherence for Topic 0, sentiment analysis, and NER.

4.15 Co-authorship Network Analysis

The co-authorship network reveals extensive collaboration and a small set of researchers connecting otherwise separate sub-communities.

  • The analysis uses an undirected co-authorship network with degree and betweenness centrality measures.
  • Normalized degree centrality reaches approximately 0.16 among the most connected authors, indicating extensive direct collaboration.
  • Mustafa Jarrar has the highest degree centrality despite not being the most productive author, suggesting a connecting role across research groups.
  • Nizar Habash has the highest betweenness centrality at 0.025, reflecting intermediary status across network sub-communities.
  • The broader results describe a collaborative field whose progress is driven by core productive authors and institutions.

5 Discussion

The discussion shows Arabic NLP expanding rapidly while remaining concentrated in infrastructure, institutions, and MSA-oriented tasks. It also identifies quantitative gaps, methodological contributions, and evidence limitations that shape priorities for future research.

  • Publication growth: Approximately 82% of corpus papers appeared after 2020, reflecting a shift toward transformer-based architectures and large language models.This acceleration is especially consequential for Arabic because morphology, diglossia, and dialectal diversity have historically challenged computational processing.
  • Topic structure: Topic 0 accounts for 41.3% of the corpus, indicating that infrastructure-building in text, speech, translation, and recognition dominates Arabic NLP.The paper links this emphasis to constructing tools, datasets, and models, with reusable resources such as AraBERT, ARBERT, and MARBERT receiving high citations.
  • Topic evolution: 614 sentiment-analysis papers represented 8.6% of the corpus, growing rapidly from around 2017, peaking during 2020–2022, and then gradually declining.The early increase is associated with abundant Arabic social-media data and commercial interest in opinion mining.
  • Dialectal coverage: MSA appears in 1,553 papers, whereas Hejazi appears once, Hassaniya not at all, and Sudanese only 20 times, revealing severe dialectal under-representation.Summarization is 76.9% MSA-focused and absent for several dialects, while OCR includes 33.3% Maghrebi work, showing that gaps vary by task.
  • Geographic distribution: Saudi Arabia, the United States, and Egypt lead institutional output with 519, 463, and 266 affiliations, respectively.The paper connects this concentration with funding availability and established research groups, while raising concerns about equity and representativeness.
  • Contributions and limitations: The study contributes a multi-source BERTopic analysis, a task–dialect gap matrix, conceptual distinctions in citation impact, and an openly released 9,141-paper corpus with analysis code.It also reports that metadata completeness correlates with citation impact, while acknowledging uneven metadata, filtering trade-offs, limited model variance explained, and provisional topic labels.

6 Conclusion

This study shows that Arabic NLP has expanded rapidly while research remains uneven across dialects and geography. Its quantitative evidence identifies under-resourced areas and supports targeted future research, funding, and benchmarking.

  • Summary of Findings: 82% of examined papers were published in the most recent period, indicating especially rapid growth after 2020.The expansion coincides with a field that has evolved from a relatively small area into a broad research domain.
  • Summary of Findings: The largest thematic cluster covers text processing, speech, translation, and recognition, while sentiment analysis is the second-largest theme but appears more mature.The largest cluster represents 41.3% of the corpus, and sentiment analysis represents 8.6%.
  • Summary of Findings: Research output is concentrated in Saudi Arabia, the United States, and Egypt, alongside substantial imbalance in Arabic dialect coverage.Modern Standard Arabic dominates the literature, while several regional varieties receive little attention.
  • Contributions: The task–dialect gap matrix identifies combinations with little or no published work, helping expose systematic coverage gaps.The study frames these gaps as a basis for recognizing where future work may be most needed.
  • Implications: Researchers and funders are encouraged to prioritize under-resourced dialects, dialectal corpora, and evaluation benchmarks.Other proposed directions include stronger metadata standards and open dissemination of preprints.
  • Limitations: The findings are indicative rather than exhaustive because coverage, text selection, normalization, topic modeling, and citation measures have important limitations.The corpus may underrepresent Arabic-language venues, and titles and abstracts may incompletely capture tasks and dialects.

Supplementary Material

The supplementary materials provide the curated corpus, analysis code, and detailed supporting tables and figures. They are released under specified licenses to support reproducibility and further research.

  • Curated Dataset: The curated dataset contains 9,141 deduplicated papers before final filtering, with the analysis subset comprising 7,120 papers.It includes abstracts, unified metadata, citation counts, matched terms, countries, and institutions.
  • Analysis Code: The analysis code covers data collection, filtering, deduplication, topic modeling, regression, network analysis, and visualization.The scripts are publicly available in the project repository.
  • Supplementary Tables and Figures: Supplementary tables and figures provide the full task–dialect gap matrix, detailed regression results, topic evolution charts, and additional visualizations.These materials are available in the repository or from the corresponding author.
  • Licensing: The analysis code uses the MIT License, while the curated dataset uses CC BY 4.0 subject to source-specific metadata terms.The original metadata remain governed by the terms of arXiv, ACL Anthology, Semantic Scholar, Crossref, and OpenAlex.
Loading 2608.23421v2…