Source-linked AI summary

Tadabur: A Large-Scale Quran Audio Dataset

Faisal Alherran

arXiv:2604.18932v1cs.SDcs.AI

TL;DR

Qur’anic speech datasets have been limited in scale, diversity, and annotation depth, while Qur’anic recitation presents distinctive modeling challenges. Tadabur addresses this gap with a large, diverse corpus and automated curation and alignment pipeline. The dataset and alignment analysis provide a substantial resource, with SILMA achieving 96.63% average coverage versus 86.03% for fuzzy matching, while some reciters lack recordings for every ayah.

  • Problem

    Existing Qur’anic speech datasets are limited in scale, reciter diversity, recording variability, and annotation depth, while Qur’anic ASR must model distinctive phonological and prosodic characteristics.

  • Method

    Tadabur is built through automated collection, LLM metadata normalization, Whisper/WhisperX alignment to canonical verses, boundary refinement, content filtering, and deduplication.

  • Results

    96.63% average coverage was achieved by SILMA embedding-based alignment, compared with 86.03% for fuzzy text matching across reciters and ASR models.

  • Takeaways & Limitations

    Tadabur provides a large and varied Qur’anic speech resource intended to support ASR, speech modeling, speaker and style analysis, prosody, tajwīd, robustness, and transfer learning.

  • Takeaways & Limitations

    Some reciters lack recordings for every ayah because source recordings were limited or the processing pipeline failed to match audio correctly, mainly due to speech-recognition errors.

Abstract

from arXiv · show

Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more than 1400+ hours of recitation audio from over 600 distinct reciters, providing substantial variation in recitation styles, vocal characteristics, and recording conditions. This diversity makes Tadabur a comprehensive and representative resource for Quranic speech research and analysis. By significantly expanding both the total duration and variability of available Quran data, Tadabur aims to support future research and facilitate the development of standardized Quranic speech benchmarks.

1 Introduction

Tadabur addresses the limited scale, diversity, and annotation depth of existing Qur’anic speech datasets by providing a large, varied, richly annotated corpus. It is intended to support Qur’anic speech research and standardized benchmarking.

  • Existing Qur’anic speech datasets are limited in scale, reciter diversity, audio quality, and annotation depth.These limitations affect research on ASR, tajwīd-aware modeling, reciter identification, and prosodic analysis.
  • Tadabur contains more than 1400+ hours of audio from over 600 reciters, covering 113 surahs and thousands of verses.It includes varied recitation styles, speaking rates, recording conditions, acoustic qualities, automatically derived metadata, and temporal annotations.
  • Tadabur is positioned as a comprehensive resource for ASR, speaker and style analysis, prosody, tajwīd, robustness, and transfer-learning research.The authors also describe it as a standardized benchmark foundation for domain-adapted speech technologies.
  • The paper combines dataset introduction, automated Qur’an data curation, and machine-readable word-level alignments with structured verse-level metadata.The curation pipeline integrates LLM-based metadata extraction, Whisper/WhisperX-based alignment, and ASR-driven content filtering.

2 Related Work

Prior Qur’anic audio datasets support several research tasks but generally lack the scale, diversity, and annotation richness needed for robust Qur’anic speech modeling. Tadabur is motivated by the domain’s distinctive phonological, prosodic, and acoustic challenges.

  • Existing Qur’anic datasets remain limited in scale, reciter and speaker diversity, recording variability, and linguistic or phonetic annotations.These constraints persist across datasets developed for ASR, pronunciation assessment, and computer-assisted recitation.
  • One publicly available classification dataset contains 6,689 files from 12 reciters but lacks transcriptions and time-aligned metadata.Its stated utility is therefore restricted to audio classification tasks.
  • Modern ASR has progressed from HMM–GMM systems through neural, CTC, attention-based, and Transformer architectures.Transformer models are described as dominant because they model long-range dependencies and scale to massive datasets.
  • Self-supervised models such as wav2vec 2.0, HuBERT, and Whisper learn transferable acoustic representations from large unlabeled speech collections.Their performance across languages and acoustic conditions highlights the role of large, diverse datasets in robust ASR.
  • Qur’anic ASR must handle prolonged phonemes, tajwīd rules, melodic articulation, speaker-dependent styles, and substantial recording variability.General ASR benchmarks based on conversational or read speech do not adequately represent these characteristics.
  • A highly variant Qur’anic speech dataset can support modeling of reciter, style, and pronunciation variation alongside domain-adapted ASR systems.This motivation connects Tadabur to the gap between general-purpose ASR and Qur’anic audio.

3 Dataset Overview

Tadabur is constructed through a fully automated pipeline that collects diverse Qur’anic audio, normalizes metadata, aligns recitations to canonical verses, refines boundaries, and removes invalid or duplicate recordings. The resulting process produces curated, temporally precise ayah-level data.

  • Dataset Overview: Tadabur targets variation across reciters, styles, chapters, acoustic environments, and recording qualities in a structured, prosodically rich speech domain.Its construction comprises data collection, metadata extraction, verse-level alignment, content cleaning, and validation.
  • Data Collection: Audio collection maximizes diversity in reciter identity, style, recording conditions, formats, and surah coverage while standardizing audio format and sampling rate.Long-form recordings are preserved for verse-level and continuous-speech modeling.
  • Metadata Extraction: LLMs infer and normalize surah and reciter metadata from unstructured titles, descriptions, tags, and file information.Gemini 2.5 Flash classifies whether files represent valid Qur’anic surah recitations and extracts structured fields.
  • Verse-Level Alignment: Whisper Large v3 and WhisperX transcribe recordings and extract word-level timestamps for alignment with canonical Qur’anic text.This produces synchronized verse audio and ground-truth text through the Ayah Alignment Module.
  • Verse-Level Alignment: SILMA embeddings match each canonical verse to transcription segments using cosine similarity and a predefined threshold.Successful matches provide start and end timestamps for verse-level segmentation.
  • Boundary Correction: Recitation-boundary detection refines preliminary endpoints so ayah segments end at the reciter’s natural stopping point.A 5-second buffer is appended before segmentation inference to address possible boundary underestimation.
  • Dataset Curation: Dataset curation combines LLM metadata validation, ASR-based canonical-verse verification, and deduplication.The pipeline filters mislabeled or irrelevant recordings, excludes non-recitation content that fails alignment, and removes duplicate groups using EAT embeddings, cosine similarity, and connected components.
  • Deduplication: Duplicate recordings are grouped by reciter and verse, compared using EAT embeddings, and merged when cosine similarity exceeds 0.9.A union–find procedure identifies connected duplicate components, retaining one representative recording per group.

4 Pipeline Quality Evaluation

Tadabur evaluates its Ayah Alignment Module across alignment methods and ASR backbones using unseen recordings from five reciters. Semantic embedding alignment with a domain-adapted ASR model provides the strongest coverage and underpins the dataset pipeline.

  • Evaluation Setup: Coverage is evaluated on independently collected, deduplicated recordings in which each ayah appears exactly once.The recordings were unseen during fine-tuning, and coverage therefore measures unique-ayah identification and segmentation without repeated-recording inflation.
  • Alignment Method: 96.63% average coverage: SILMA Embedding with the Tadabur model outperforms fuzzy matching at 86.03%.The difference exceeds 10 percentage points across the five-reciter evaluation.
  • Alignment Method: SILMA Embedding outperforms fuzzy text matching across all three ASR backbones.Semantic matching is reported as more robust to phonological variation, elongated phonemes, and recitation-specific disfluencies.
  • ASR Model: 82.57% under SILMA Embedding and 72.80% under Fuzzy Matching: Whisper Small has the lowest coverage without domain adaptation.The paper attributes this degradation to Qur’anic transcription errors propagating into alignment.
  • ASR Model: 96.63% under SILMA Embedding: Tadabur marginally exceeds Whisper-Quran at 95.50% among domain-adapted models.Under Fuzzy Matching, Whisper-Quran reaches 87.23% versus Tadabur’s 86.03%.
  • Summary: SILMA Embedding with the Tadabur fine-tuned model is selected as the adopted configuration for full dataset construction.The configuration achieves the highest alignment coverage in the evaluated settings.

5 Dataset Statistics

Tadabur is presented as a large and diverse Quranic recitation dataset, with extensive annotated audio and reciter coverage. Its diversity includes both cross-reciter differences and within-reciter variation across recordings, styles, and acoustic conditions.

  • Dataset Statistics: Recording distributions are non-uniform across reciters because the dataset reflects natural variation in publicly available audio sources.This qualifies interpretation of reciter-level counts.
  • Dataset Size: More than 1400+ hours of verse-level annotated audio, over 600 distinct reciters, and automatically generated word-level alignments define the final dataset.The dataset also includes structured metadata.
  • Reciter Diversity: Reciter coverage spans varied ages, dialects, and recitation traditions.These dimensions contribute to the dataset’s stated diversity.
  • Dataset Size: Tadabur offers substantially larger total audio duration and reciter diversity than previously available Quranic datasets.Table 2 provides the quantitative comparison with widely used public datasets.
  • Reciter Diversity: Multiple recordings of the same surah and ayah capture within-reciter variation in sessions, pace, melodic choices, and acoustic environments.The dataset therefore represents both cross-reciter diversity and repeated-performance variation.

6 Models Evaluation Against Tadabur

Tadabur is used to benchmark eight ASR models with standardized WER and CER evaluation. Results show that Qur’anic domain adaptation matters more than model size.

  • Evaluation Setup: Eight publicly available ASR models spanning different architectures, parameter scales, and Arabic or Qur’anic adaptation levels are evaluated on Tadabur.The evaluation uses WER and CER after normalizing diacritics, punctuation, and Uthmani orthographic variants against canonical Qur’anic text.
  • Evaluation Setup: Table 3 reports WER and CER for all evaluated models, sorted by WER.
  • Results: 8.7% WER and 6.5% CER make Whisper-Quran the best model despite its 74M parameters.It substantially outperforms much larger general-purpose models.
  • Results: 11.2% WER for Cohere Transcribe and 15.1% for Voxtral Mini are the strongest results among the larger general-purpose models.Cohere Transcribe achieves this result without Qur’anic training data.
  • Results: 51.1% WER for MMS 1B and 57.4% for Wav2Vec2 XLSR-53 Arabic show poor generalization to Qur’anic recitation.
  • Analysis: Domain adaptation outweighs model size in Qur’anic ASR.

7 Licensing and Ethical Considerations

Tadabur is openly released with audio, annotations, and metadata to support Arabic speech research. Its use is framed around respectful and beneficial educational, accessibility, and academic applications.

  • Licensing: Tadabur is published as a freely accessible open-source dataset for Arabic audio and speech research.
  • Licensing: Released metadata files support exploration, analysis, and reuse alongside the audio and annotations.
  • Ethical Considerations: Users are expected to apply Tadabur in respectful and beneficial contexts, particularly education, accessibility, and academic research.The guidance excludes mocking, distortion, or otherwise disrespectful manipulation of Qur’anic recitation.

8 Limitations

Tadabur has coverage and alignment limitations despite its scale. Some reciters lack recordings for every ayah, and word-level timestamps are not always precise.

  • Coverage: Tadabur’s first limitation is incomplete ayah coverage for some reciters.This reflects either limited available recordings or failures in matching audio to the correct ayah.
  • Coverage: Speech-recognition errors mostly caused some audio-matching failures during processing.
  • Alignment: Tadabur’s second limitation is that word-level timestamps are not always precise.The alignment model was not originally built for Qur’anic audio and struggles with its pronunciation and recitation style.
Loading 2604.18932v1…