Source-linked AI summary

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas, Nada Almarwani, Samah Aloufi, Saad Ezzini, Maged S. Al-Shaibani, Doaa Dalaq, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Mohamed Mehdi Trigui, Dania Refai, Layan Refai, Mohamed Akrout, Mustafa Jarrar, Wasfi G. Al-Khatib, Alaa Dalaq, Darin El-Nakla, Samir Abdaljalil, Abdulrahman Al-Fakih, Nour El Imane Zeghib, Moussa Redah, Salmane Chafik, Mohamed El-Attar, Rima Grati, Sarah Kohail, Malak Alkhorasani, Khadijah Al Safwan, Ismail M. Mudhaffar, Ali Altam, Ahmed Al-Shaikh, Adnan Saeed, Hamzah Luqman

arXiv:2608.21950v1cs.CLcs.AI

TL;DR

Arabic ASR remains challenged by extensive dialectal variation, uneven resources, and limited coverage of accented MSA and classical Arabic. BULBUL addresses these gaps with a multi-dialect corpus and benchmarks ASR systems, finding OmniLLM-7B strongest across dialectal and accented MSA/CA speech.

  • Problem

    Existing Arabic ASR datasets cover few dialects and largely overlook accented MSA and classical Arabic, leaving several varieties comparatively low-resource.

  • Method

    BULBUL constructs an Arabic speech corpus spanning dialectal subsets and accented MSA/CA, with country-level regional grouping for analysis.

  • Results

    OmniLLM-7B achieves the best overall dialectal performance, with mean WER 41.7 and CER 12.7, and accented MSA/CA performance of 13.3/4.8.

  • Takeaways & Limitations

    BULBUL provides benchmarks for dialectal and accented Arabic ASR, while results show standardized MSA/CA is easier for the evaluated systems than dialectal speech.

  • Takeaways & Limitations

    Coverage is constrained by limited domain diversity in some country subsets, smaller MSA/CA subsets, and uneven gender representation across countries.

Abstract

from arXiv · show

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.

1 Introduction

Arabic ASR is hindered by extensive dialectal variation, uneven resource coverage, and limited evaluation of accented MSA and classical Arabic. BULBUL addresses these gaps with broad dialect coverage, accented formal-Arabic recordings, human verification, and benchmark baselines.

  • Arabic ASR must handle substantial phonetic and lexical diversity across dialects and subdialects, while speech resources remain limited and unbalanced.
  • Existing datasets primarily cover MSA, Egyptian, Levantine, Gulf, and Moroccan Arabic, leaving several dialects comparatively low-resource.
  • Accented MSA and classical Arabic have been largely overlooked, with no dedicated dataset systematically evaluating them across dialectal backgrounds.
  • BULBUL covers speakers from 11 Arab countries and includes fine-grained Saudi and Yemeni subdialects to improve dialectal coverage.
  • The dataset adds accented MSA and classical Arabic, uses two-level human verification, benchmarks diverse ASR systems, and releases demographic metadata.

2 Related Work

Arabic speech resources range from multilingual and uni-dialect corpora to broader multi-dialect datasets, but existing resources often trade dialectal breadth for domain or annotation limitations. BULBUL combines regional and sub-dialect coverage with accented formal-Arabic recordings and verified metadata.

  • Multilingual datasets support cross-lingual learning but typically provide limited Arabic dialectal diversity, while uni-dialect corpora offer deeper but geographically narrow coverage.
  • Multi-dialect datasets broaden regional coverage but often rely on broadcast, telephone, or web recordings with domain bias and inconsistent annotation quality.
  • BULBUL includes naturally diverse community speech, explicit sub-dialect annotations, accented MSA and classical Arabic from 11 dialect speakers, human-verified transcriptions, demographic metadata, and domain diversity.

3 BULBUL Dataset

BULBUL is a community-driven, multi-country Arabic speech dataset built through staged collection, dialect-specific organization, text preparation, recording guidance, and two-level verification. Its design combines broad regional coverage with selected sub-dialect and formal-Arabic resources.

  • BULBUL is a community-driven dataset collected from 275 participants across 11 Arab countries, organized through country-specific teams and coordination.
  • The collection expanded incrementally from Saudi Arabia and Egypt to 11 countries and includes dialectal speech plus CA and MSA in participants’ native accents.
  • Seven countries use a white dialect, while Syria and the UAE use Levantine and Abu Dhabi dialects, respectively.
  • Saudi data cover six subdialects, with Qatifi and Hassawi combined as Eastern, while Yemen includes Sana’ani and Ta’izzi.
  • Dialect-specific texts were collected from public datasets, social media, and private communication platforms, then organized by country and subdialect for recording.
  • Texts were manually revised and classified into 11 thematic domains before recording, while speakers and native reviewers performed self-verification and external verification.

4 BULBUL Analysis

BULBUL combines substantial dialectal and accented MSA/CA speech coverage with descriptive analyses of duration, demographics, domains, and evaluation splits. The analysis also reveals coverage imbalances and recording-condition differences that constrain interpretation.

  • 36,949 dialectal utterances span 61.46 hours, while 2,325 accented MSA/CA recordings span 10.42 hours.
  • Yemeni, Saudi, and Syrian dialects have the longest dialectal durations, whereas Sudanese speech remains below one hour.
  • 2.82 hours of accented MSA/CA recordings are Yemeni and 2.02 hours are Egyptian, compared with 0.02 hours for UAE and 0.11 hours for Tunisian.
  • Dialectal and MSA/CA temporal differences require caution because the subsets were recorded under different conditions, including read speech.
  • Gender representation varies across countries, including only male participants in Sudan and an approximately 89.5% female majority in the UAE.
  • The dataset’s domain distribution is concentrated in Social speech at 65.4%, followed by Economy at 15.7%.

5 Evaluation

BULBUL is benchmarked with diverse multilingual ASR models under zero-shot evaluation. The setup measures out-of-the-box generalization across dialects and accents without dataset-specific fine-tuning.

  • The benchmark evaluates Whisper, SeamlessM4T, MMS, OmniLLM, and OmniCTC models spanning multiple model scales.
  • Zero-shot evaluation uses the BULBUL test set without fine-tuning on BULBUL or any country-specific subset.
  • Default decoding configurations measure out-of-the-box generalization and cross-lingual and cross-accent robustness.

6 Results and Discussion

BULBUL evaluation shows substantial variation in ASR performance across dialects, with OmniLLM-7B strongest overall and standardized accented MSA/CA easier than dialectal speech. Error analysis identifies phonetic, phonological, and lexical failure patterns, while model scaling benefits LLM-decoder systems more consistently than CTC systems.

  • Overall Results: OmniLLM-7B achieves the best overall dialectal performance, with mean WER 41.7 and CER 12.7.The evaluation covers regional dialect clusters in a zero-shot benchmark.
  • Per-dialect Results: Palestinian and Saudi speech has the lowest error rates across most models, whereas Algeria and Sudan are generally the most difficult dialects.The Palestinian subset has the lowest WER for 9 of 12 evaluated systems, partly associated with limited domain diversity.
  • Accented MSA and CA: Accented MSA/CA substantially improves performance; OmniLLM-7B reaches 13.3/4.8 mean WER/CER, while Sudanese-accented speech remains challenging above 20% WER.The dialectal-to-accented robustness gap is largest for North African and Yemeni dialects.
  • Model Scale: LLM-decoder models improve almost monotonically with model size, but CTC scaling is unstable and omniASR-CTC-1B outperforms larger 3B and 7B variants.The CTC result indicates that increasing parameter count alone does not guarantee gains under that architecture.
  • Error Analysis: Qualitative analysis finds phonetically similar substitutions, dialect-driven phonological mismatches, and distortions of rare or dialect-specific lexical items.Character-level errors often reflect phoneme-to-grapheme mismatches, while word-level errors are more common for proper nouns and low-frequency vocabulary.

7 Conclusions

BULBUL is a community-driven Arabic speech corpus designed to represent linguistic diversity beyond dominant dialects and MSA. Its broad geographic and sub-dialect coverage, accented MSA/CA recordings, and benchmark evaluations support more inclusive Arabic speech technology research.

  • Dataset Scope: BULBUL covers 11 countries and 16 dialects through fine-grained sub-dialect annotations, providing a challenging benchmark under diverse regional and demographic conditions.Saudi and Yemeni dialects receive additional sub-dialect division.
  • Accent Coverage: BULBUL includes accented MSA and CA recordings for studying accent transfer and pronunciation variation across standardized speech forms.These recordings extend evaluation beyond ordinary dialectal speech.
  • Benchmark Findings: Benchmark results reveal substantial performance disparities across countries and dialects, showing uneven support for underrepresented Arabic varieties.Lexical-overlap and speaking-rate analyses further demonstrate heterogeneity across Arabic speech communities.
  • Research Uses: BULBUL is intended to support dialect-aware ASR, dialect identification, accent adaptation, and speech generation for low-resource Arabic varieties.The corpus is positioned as a resource for more inclusive Arabic speech research.

8 Limitations

BULBUL has scope limitations arising from controlled recording conditions, geographic and gender imbalance, uneven domain coverage, and limited MSA/CA participation. These constraints affect applicability to noisy or emotional speech and caution against broad generalization from some subsets.

  • Recording Conditions: Emotionally neutral, low-noise recordings limit applicability to emotion recognition and noisy, unconstrained speech processing.The evaluation is therefore more closely associated with dialectal understanding than emotion or environmental-noise robustness.
  • Participant Balance: Some countries are underrepresented and gender distributions are not perfectly balanced because participation was voluntary.The dataset provides statistical analysis to characterize these imbalances, but they cannot be fully corrected retrospectively.
  • Domain Coverage: Limited domain diversity in some subsets, especially Palestine’s economy-focused data, may make test data more homogeneous and easier for ASR models.Cross-country performance differences should account for this domain imbalance.
  • MSA and CA Coverage: Small MSA and CA subsets from few participants reduce diversity in speaking styles, accents, and acoustic conditions, limiting generalization.Findings for CA and MSA should therefore be interpreted cautiously.

9 Ethics and Data Statement

BULBUL combines consented, anonymized speech collection with research-use restrictions and supporting documentation on data preparation, verification, and benchmarking. The dataset acknowledges demographic and regional imbalance as a boundary for model performance and fairness evaluation.

  • Data collection and consent: Participants provided informed consent for recordings and transcripts, including voluntarily contributed private chat conversations used for some conversational content.The corpus includes prompted speech, natural conversation, and anonymized contributions from WhatsApp and Telegram conversations.
  • Privacy protection: Personally identifying information was removed, while only age and gender metadata were retained; original private chat logs are not distributed.The released material consists of anonymized speech recordings and corresponding transcripts rather than original chat records.
  • Use restrictions: The dataset is restricted to non-commercial academic research and must not support identification, personal-information reconstruction, surveillance, or harmful applications.Users are expected to follow applicable privacy regulations and ethical guidelines.
  • Scope and limitations: Demographic and regional imbalances may affect performance across speaker groups, motivating future fairness evaluation and dialect-aware benchmarking.The authors also discourage privacy-compromising or harmful surveillance applications.

B.4 Lexical and Semantic Analysis

The analysis examines lexical and semantic variation across BULBUL’s countries and selected sub-dialects, while the dataset pipeline uses participant recording, review, and verification procedures. Results show substantial lexical differences alongside generally high cross-dialect semantic similarity.

  • Lexical diversity: Yemen and Saudi contain the largest corpora, whereas smaller subsets show inflated lexical-diversity measures associated with limited data.Yemen has 4,078 sentences and Saudi 3,082; Sudan and UAE have the highest reported TTR values among the cited small subsets.
  • Lexical diversity: Palestine has the lowest reported lexical diversity, consistent with its text being restricted to economically oriented utterances.The cited passage reports TTR = 0.184 and a hapax proportion of 0.550.
  • Semantic similarity: Cross-country semantic similarity is generally greater than 0.75, with a prominent cluster spanning Jordan, Egypt, Saudi Arabia, Syria, Tunisia, and Yemen.Morocco is slightly less similar to Eastern dialects, while the analysis uses repeated random sampling of 48 sentences per country to reduce corpus-size bias.
  • Sub-dialect variation: Saudi sub-dialect differences often involve phonetic or phonological variation rather than lexical substitution.The cited example contrasts gahwa in Najdi speech with ghawa in southern Saudi speech despite the same orthographic word.
  • Recording and verification: BULBUL records were collected through a web interface allowing review, trimming, repetition, and skipping, followed by structured secondary-speaker verification.Reviewers could accept or reject samples using predefined reasons such as unclear recording, text–audio mismatch, prolonged silence, or dialect mismatch.

E Dataset Split Statistics

The split statistics summarize utterances, duration, and speaker counts for development and test data across dialectal and accented datasets. Accented MSA/CA recordings have one speaker per country because suitable speakers were limited.

  • Dataset split statistics: Table 9 reports utterance counts, total minutes, and speaker counts for development and test splits across evaluated dialect datasets.Development splits were primarily used for the paper’s error analysis.
  • Accented MSA/CA: Accented MSA/CA recordings use one speaker per country because few speakers were available to record those samples.Table 9 also summarizes the accented MSA/CA dataset separately.

F Benchmarking Experiment Setup

The experiments use isolated software environments and multiple GPU platforms to run inference with pretrained ASR systems. Evaluation applies consistent model weights, settings, and transcript normalization across environments.

  • Computing environment: Experiments ran on Google Colab and institutional servers using NVIDIA T4, A100, H100, RTX 3090, and RTX A6000 GPUs.Hardware selection depended on experiment scale and resource availability.
  • Reproducibility setup: Separate Anaconda environments and Colab notebooks were maintained for the three ASR model families to manage differing software dependencies.The environments were configured separately for compatibility across model families.
  • Evaluation protocol: The benchmark evaluates pretrained ASR models using inference only, without additional training or fine-tuning.Hardware differences therefore primarily affected execution time and resource utilization rather than transcription quality or evaluation outcomes.
  • Evaluation protocol: Reference and predicted transcripts undergo the same normalization pipeline before CER and WER calculation.The pipeline includes Unicode normalization, removal of diacritics and Tatweel, numeral and Alef normalization, punctuation removal, and whitespace collapsing.
Loading 2608.21950v1…