Source-linked AI summary
EDRAC: Benchmarking Arabic Dialect Reading Comprehension
Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
TL;DR
Dialectal Arabic lacks broad, high-quality benchmarks for machine reading comprehension and generative QA, especially from naturally spoken language. EDRAC introduces a five-dialect benchmark built from spoken interactions through human–LLM generation and verification, and its evaluation reveals a mismatch between semantic quality and dialectal fidelity. The findings show that standard metrics can misrepresent dialectal proficiency.
Problem
Dialectal Arabic has limited high-quality MRC and QA resources, while existing benchmarks often focus on MSA, formal text, or multiple-choice tasks.
Method
EDRAC constructs a five-dialect benchmark from naturally occurring spoken interactions using iterative LLM generation, LLM evaluation, and human verification.
Results
Standard metrics reveal substantial gaps between semantic answer quality and dialectal fidelity, with GPT-5.4 falling from first to joint fifth under CAMeLBERTScore-F1.
Takeaways & Limitations
EDRAC provides a realistic benchmark for evaluating dialectal Arabic MRC and generative QA and highlights limitations of current evaluation metrics.
Takeaways & Limitations
EDRAC covers five dialects but not the full linguistic diversity of Arabic, including finer-grained regional and sociolectal variation.
Abstract
from arXiv · showhide
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.
1 Introduction
EDRAC addresses the scarcity of high-quality dialectal Arabic resources for open-ended reading comprehension and generative QA. It introduces a five-dialect benchmark built through a human–LLM pipeline and evaluates Arabic-centric and multilingual models.
- Dialectal Arabic has fewer high-quality resources because naturally spoken dialects are underrepresented in written data.
- Existing dialectal MRC resources largely use multiple-choice QA, whereas EDRAC supports open-ended answers requiring descriptive and explanatory responses.
- EDRAC introduces a large-scale generative QA benchmark covering Egyptian, Emirati, Moroccan, Syrian, and Saudi Arabic from naturally occurring spoken interactions.
- The dataset uses a scalable human–LLM pipeline combining iterative generation, LLM-based evaluation, and human verification.
- The study benchmarks Arabic-centric and multilingual LLMs and reports substantial gaps in dialectal fidelity, establishing a benchmark for future research.
2 Related Work
Prior Arabic QA benchmarks primarily use formal MSA or constructed written data, leaving naturally occurring dialectal speech underrepresented. EDRAC fills this gap with authentic conversational speech and human-in-the-loop QA generation.
- Existing Arabic QA and reading-comprehension benchmarks primarily focus on MSA and formal written text.
- Recent dialectal benchmarks cover multiple varieties, but most rely on paraphrasing, translation, or controlled prompting rather than naturally occurring speech.
- EDRAC captures orthographic inconsistency, code-switching, and disfluent discourse from naturally occurring dialectal speech.
- Unlike translated conversational QA, EDRAC uses authentic conversational speech to evaluate comprehension under realistic dialectal conditions.
- EDRAC applies a human-in-the-loop framework in which LLMs generate candidate QA pairs from spoken transcripts while human review supports data quality.
3 EDRAC
EDRAC is a five-dialect generative QA benchmark grounded in spoken interactions. Its two-stage pipeline combines passage curation with iterative LLM-assisted QA generation and human verification.
- EDRAC releases 499 passages and 4,977 QA pairs across Egyptian, Emirati, Moroccan, Syrian, and Saudi Arabic.The release contains approximately 1,000 QA pairs per dialect.
- The end-to-end pipeline consists of passage curation in the top panel and QA generation in the bottom panel.
- The dataset relies on unscripted and partially scripted real-world speech to preserve linguistic and cultural subtleties without artificial modification.
- Passage curation processes dialectal YouTube videos through diarization, transcription alignment, text restoration, passage extraction, and native-speaker review.
- QA pairs are generated with an iterative Gemini 2.5 pro framework using multiple refinement steps to improve quality and reduce hallucinations.
4 Passage Curation
EDRAC curates dialectal reading-comprehension passages from diverse YouTube speech using automated preprocessing followed by native-speaker correction and review. The process prioritizes representative genres and conversational authenticity.
- Existing dialectal speech corpora were designed mainly for ASR or dialect identification and were unsuitable for EDRAC’s reading-comprehension objectives.
- YouTube was selected to provide diverse audio genres across five dialects and to leverage pre-existing automated transcriptions.
- The collection targeted 100 videos per dialect, sourced roughly 150 per dialect for filtering, and extracted the first six minutes of each video.Videos were restricted to publications from June 2024 through January 2026, and those without existing transcriptions were excluded.
- Preprocessing used speaker diarization, timestamped transcription, and text-structure restoration to preserve speaker-separated conversational material.
- Native speakers performed passage review, while annotators corrected transcriptions and the final passages underwent internal sanity checks.
5 Question-Answer Pairs Generation
EDRAC uses an iterative generate–judge–revise pipeline, followed by human-in-the-loop validation, to create dialectal Arabic QA pairs across five dialects. The process addresses hallucinations, dialectal naturalness, correctness, and annotation consistency through automated checks and native-speaker review.
- An initial two-week pilot with in-house linguists was too difficult and time-consuming, motivating the use of LLMs for QA generation.
- Initial single-prompt generation caused hallucinations, excessive MSA usage, and irrelevant content, so the pipeline added iterative refinement while preserving dialect characteristics.
- The generation pipeline produces 10 QA pairs per passage, evaluates them with an LLM-as-a-judge, and iteratively revises failed pairs using judge feedback.Steps 2 and 3 repeat for up to five iterations or until all quality checks are passed.
- Annotator disagreement in the Emirati pilot led to clearer guidance on naturalness, dialectal variation, and code-switching, with separate guidelines for each dialect.The disagreement was primarily attributed to regional linguistic variation and annotators’ daily linguistic habits.
- Human reviewers assessed generated pairs for relevancy, naturalness, and correctness using dialect-specific guidelines refined through pilot testing.The criteria covered passage relevance, naturally spoken dialect representation, and answers supported by or inferable from the passage.
- Inter-annotator agreement was assessed with Cohen’s Kappa on 20 passages per dialect, with low Kappa attributed to skewed label distributions known as the Kappa Paradox.
- 209 of 5,000 QA pairs were flagged during manual review, and 23 were discarded, leaving 499 passages in the released dataset.Flagged issues included incorrect answers, question misinterpretation, gender confusion, partial correctness, and hallucinated answers.
- Naturalness judgments remain challenging because regional preferences and unfamiliarity with dialect sub-varieties can bias annotators against legitimate variation.The guidelines instructed annotators to look past sub-variety bias and treat variation as a natural linguistic phenomenon.
6 Experimental Setup
The experiments benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic evaluation metrics. Models are ranked across dialects by aggregate performance.
- Model Evaluation: The study evaluates Arabic-centric and multilingual LLMs on EDRAC for dialectal reading-comprehension question answering.The evaluated models include dialect-focused, open-weight multilingual, and closed-weight multilingual systems.
- Evaluation Metrics: Model outputs are assessed with ROUGE, BERTScore, and CAMeLBERTScore.BERTScore uses mDeBERTa, while CAMeLBERTScore uses CAMeLBERT fine-tuned for Arabic dialects.
- Evaluation Metrics: ROUGE measures exact lexical overlap using whitespace tokenization, whereas BERTScore-F1 is used to measure semantic similarity.The semantic evaluation uses encoder-only language models, including multilingual mDeBERTa and dialect-oriented CAMeLBERT.
- Reporting: Results in Figure 4 order models by average performance across dialects under the three metrics, with the best-performing model shown first.The ordering is based on descending average performance across dialects.
7 Results and Analysis
Across dialects, model rankings depend on the evaluation metric, and automatic scores diverge from human judgments. The results expose a gap between semantic answer quality and dialectal naturalness or fidelity.
- Automatic Evaluation: GPT-5.4 achieved the highest average performance across dialects according to ROUGE-L and BERTScore-F1.Figure 5 reports aggregate performance across the five dialects and three metrics.
- Model Comparisons: Atlas-Chat-9B showed solid overall performance, with exceptionally strong results on Moroccan Arabic.The Moroccan performance aligns with the model’s Moroccan-focused training data.
- Model Comparisons: Nile-Chat-4B outperformed larger Arabic-centric models including Jais-13B-Chat, SILMA-9B, and Hala-9B despite having 4B parameters.The passage characterizes this as remarkable efficiency.
- Model Comparisons: Gemma4-31B delivered strong results but was outperformed by the smaller Atlas-Chat-9B in most scenarios, reducing its deployment practicality.Qwen3.6-Plus also degraded relative to its predecessor in this evaluation setting.
- Metric Analysis: GPT-5.4 fell from first to joint fifth with Qwen2.5-7B under CAMeLBERTScore-F1, despite semantically accurate responses.The result indicates that standard semantic and lexical metrics do not consistently capture dialectal constructions and lexical overlap.
- Human Evaluation: Human evaluation found that ROUGE penalizes lengthy accurate answers, while BERTScore fails to capture non-dialectal text.Annotators evaluated accuracy and naturalness on outputs from GPT-5.4 and Hala-9B across the five dialects.
8 Conclusion and Future Work
EDRAC establishes a large-scale benchmark for dialectal Arabic reading comprehension and generative QA across five dialects. The authors identify metric limitations and outline extensions to broader dialectal and spoken-language settings.
- Conclusion: EDRAC is a large-scale benchmark for dialectal Arabic machine reading comprehension and generative QA across five dialects.It is grounded in naturally occurring spoken interactions and built through iterative generation, LLM-as-a-judge evaluation, and human verification.
- Conclusion: Experiments reveal substantial gaps between semantic answer quality and dialectal fidelity, exposing limitations in current evaluation metrics.The benchmark is intended to support future research on dialectal Arabic NLP.
- Future Work: Future work will expand EDRAC to additional dialects, conversational and multi-hop QA, dialect-sensitive metrics, and other low-resource spoken languages.These directions are explicitly listed as planned extensions.
Limitations
EDRAC covers five major dialects but not the full regional and sociolectal diversity of Arabic. Its QA pipeline, model selection, and benchmark trends also remain bounded by documented methodological and temporal constraints.
- Scope: EDRAC does not capture finer-grained regional and sociolectal variation across the Arabic-speaking world.The dataset covers five major dialects but not the full linguistic diversity of Arabic.
- Data Sources: Some source content may contain partial scripting or speaker self-monitoring typical of online media despite being grounded in spoken interactions.This qualifies the naturalness of the source material.
- Pipeline and Models: LLM-assisted QA generation and evaluation may introduce residual biases or stylistic artifacts despite human verification and quality control.The authors also limited evaluated Arabic-centric models to a representative and diverse sample.
- Benchmark Stability: Future language models may show different performance trends because the model landscape is rapidly evolving.This limits the durability of current benchmark comparisons.
Ethics and Broader Impact
EDRAC is released for non-commercial evaluation while incorporating safeguards for sensitive content and underrepresented spoken varieties. The authors acknowledge that naturally occurring dialectal data and LLM-assisted generation may retain bias and subjectivity.
- Data use: EDRAC is released under CC BY-NC-SA 4.0 for non-commercial use and should be excluded from pretraining and fine-tuning data.The dataset is intended solely as an evaluation benchmark.
- Risk mitigation: Sensitive, toxic, or explicit videos were excluded, with multiple layers of human review applied throughout annotation.These safeguards were introduced to reduce potential harms.
- Broader impact: EDRAC supports research on underrepresented spoken Arabic varieties that are often overlooked by current language technologies.The authors connect this goal to dialect-aware technologies and more robust spoken-language evaluation.
- Limitations: Regional, social, and cultural biases may remain in dialectal data despite native-speaker annotation and quality-control procedures.Annotation subjectivity and LLM-assisted generation may also introduce residual biases or stylistic artifacts.
A Results
The results section reports model evaluation across dialects and describes the evaluation and transcript-structuring procedures. It uses lexical and semantic scoring while restoring punctuation and paragraph structure through a three-stage pipeline.
- Model evaluation: Table 6 reports model evaluation results across dialects using ROUGE-L and CAMeLBERTScore-F1.RL denotes ROUGE-L with whitespace tokenization, while BFD 1 denotes CAMeLBERTScore-F1.
- Model evaluation: Figure 5 summarizes average performance per dialect across all evaluated models.The figure is organized around dialect-level average performance rather than individual model results.
- Model evaluation: Model rankings across the five Arabic dialects are sorted by average rank, with lower rank values considered better.The ranking aggregates performance across dialects.
- QA evaluation: The QA-generation prompts require short, third-person, non-opinionated questions and answers grounded in the passage’s language.They also require exact supporting quotes and character indices, while avoiding repetition of previous questions.
- QA evaluation: Generated QA pairs are judged for non-opinionation, answer clarity, bias, answerability, relevance, third-person form, and linguistic or grammatical terminology.The judge returns critical issues and recommendations for revision.
- Text restoration: Raw transcripts are restored in three stages: utterance-level punctuation classification, structured-text generation, and merging punctuation markers with original tokens.The final stage aligns generated tokens with original tokens and inserts only punctuation and paragraph markers into the original stream.
- Text restoration: Generated structured text may differ lexically from the original through added words, dropped particles, or paraphrases.These differences are treated as text noise for later cleaning.
F Additional Details on the Experimental Setup
The experimental setup evaluates a list of language models whose sizes are reported in Table 8, using different inference arrangements depending on platform availability.
- Inference setup: Most models were run through OpenRouter at a total inference cost of approximately $200.Arabic-centric models unavailable on OpenRouter were run on a single NVIDIA RTX 6000.
- Evaluated models: Table 8 lists the evaluated language models and their sizes.The sizes of GPT-5.4 and Qwen3.6-Plus are not publicly disclosed.
G Final dataset stats
The final dataset combines dialectal video material, corrected transcriptions, and reviewed QA-pair annotations. Its guidelines define procedures for transcription, dialectal orthography, and QA-quality assessment.
- Dataset statistics: Table 9 records the total number of QA pairs before and after review for each dialect.The table is specifically a post-review reconciliation of QA-pair counts.
- Dataset statistics: Figure 7 shows genre distribution across Egyptian, Moroccan, Saudi, Syrian, and Emirati dialects.The figure compares genre composition by dialect.
- Transcription guidelines: Annotators reviewed automated transcriptions while listening to the source audio, skipping passages with large missing segments and applying minimal post-editing.The workflow emphasized readability, comprehension, and dialectal pronunciation.
- Transcription guidelines: Annotators were instructed to preserve dialectal words rather than replace them with default MSA forms and to retain spoken filler words.Onomatopoeic sounds could instead be represented by descriptions in curly brackets.
- Orthographic policy: The guidelines distinguish acceptable spelling variants from unacceptable ASR errors requiring correction.Acceptable variants do not affect understanding, whereas unacceptable errors hinder readability or comprehension.
- Orthographic policy: A complete list of acceptable variants is impossible because context can make otherwise interchangeable spellings convey different meanings.The guidelines therefore treat some spellings as context-dependent.
- QA annotation: Automatically highlighted passage spans are intended to help locate answers but may be inaccurate, so answers can occur elsewhere in the passage.Annotators were told that answers may also require combining multiple sentences or interpreting the passage overall.
- QA annotation: QA evaluation covers question relevance, QA-pair naturalness, and answer correctness for dialectal passages.Naturalness concerns how well generated pairs represent Arabic as naturally spoken by native speakers.