Source-linked AI summary
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
Yuhang Dai, Ziyu Zhang, Shuai Wang, Longhao Li, Zhao Guo, Tianlun Zuo, Shuiyuan Wang, Hongfei Xue, Chengyou Wang, Qing Wang, Xin Xu, Hui Bu, Jie Li, Jian Kang, Binbin Zhang, Lei Xie
TL;DR
Sichuanese dialect speech lacks large, diverse open-source resources, limiting speech-technology development for a major dialect community. The paper introduces a 10,000-hour richly annotated corpus, a dedicated processing pipeline, and manually verified ASR/TTS benchmarks. Models trained on the corpus achieve state-of-the-art performance among open-source systems and perform comparably to commercial systems.
Problem
Sichuanese dialects lack sufficiently large and diverse open-source resources for robust ASR and TTS, despite substantial linguistic distinctiveness and approximately 120 million speakers.
Method
The paper introduces WenetSpeech-Chuan, Chuan-Pipeline, and manually verified ASR/TTS benchmark sets for Sichuanese dialectal speech.
Results
Models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and perform comparably to commercial systems across ASR and TTS.
Takeaways & Limitations
WenetSpeech-Chuan provides a large open-source resource and standardized benchmarks for advancing Sichuanese dialectal speech processing.
Takeaways & Limitations
The commercial TTS comparison uses a single fixed speaker, without considering speaker similarity.
Abstract
from arXiv · showhide
The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constructed using our novel Chuan-Pipeline, a complete data processing framework for dialectal speech. To facilitate rigorous evaluation and demonstrate the corpus's effectiveness, we also release high-quality ASR and TTS benchmarks, WenetSpeech-Chuan-Eval, with manually verified transcriptions. Experiments show that models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and demonstrate results comparable to commercial services. As the largest open-source corpus for Sichuanese dialects, WenetSpeech-Chuan not only lowers the barrier to research in dialectal speech processing but also plays a crucial role in promoting AI equity and mitigating bias in speech technologies. The corpus, benchmarks, models, and receipts are publicly available on our project page.
1. INTRODUCTION
Sichuanese dialect speech lacks large, diverse open-source resources despite being spoken by approximately 120 million people and differing substantially from Standard Mandarin. WenetSpeech-Chuan addresses this gap with a richly annotated corpus, Chuan-Pipeline, and ASR/TTS benchmarks.
- Approximately 120 million people speak Sichuanese dialects, whose tonal system, vocabulary, and grammar differ substantially from Standard Mandarin.
- Existing open-source Sichuanese resources are limited to two small MagicData corpora with 4.53 and 6.4 hours, while KeSpeech contains accented Mandarin rather than dialectal speech.
- These resources are inadequate for robust Sichuanese ASR and TTS because of their limited scale and narrow coverage.
- WenetSpeech-Chuan provides over 10,000 hours of richly annotated Sichuanese speech from diverse domains, including short videos, entertainment, and live streams.
- The project also develops Chuan-Pipeline and releases benchmark sets with fine-grained manual corrections for rigorous ASR and TTS evaluation.
- Models trained on WenetSpeech-Chuan achieve state-of-the-art recognition accuracy and synthesis quality for Sichuanese dialects among open-source systems.
2. CHUAN-PIPELINE
Chuan-Pipeline transforms raw, unlabeled Sichuanese audio into a richly annotated corpus through acquisition, segmentation, speaker and quality processing, transcription correction, and multimodal punctuation prediction. Its LLM-GER transcription stage combines multiple ASR outputs with Qwen3 [14], improving transcription accuracy by approximately 15% on average.
- 2. CHUAN-PIPELINE: Chuan-Pipeline systematically transforms raw, unlabeled audio into a richly annotated corpus suitable for ASR and TTS research.
- 2. CHUAN-PIPELINE: Audio acquisition begins with online metadata mining and manual dialect verification, followed by 5–25 second VAD-based segmentation and single-speaker selection.
- 2. CHUAN-PIPELINE: The pipeline assigns speaker IDs using CAM++ embeddings [11] and adds gender, age, and seven-category emotion annotations.
- 2. CHUAN-PIPELINE: Quality assessment computes WVMOS from duration, SNR, and timestamp-aligned speech, discarding low-quality samples while retaining varied quality levels.
- 2.3. LLM-GER Processing: LLM-GER merges candidate transcriptions from FireRed-ASR, SenseVoice-Small, and TeleASR with Qwen3 [14], which corrects errors without changing original semantics or token length.
- 2.3. LLM-GER Processing: Approximately 15% average transcription-accuracy improvement over individual ASR systems demonstrates the benefit of combining multiple ASR outputs with LLM-based dialect normalization.
- 2.4. Punctuation Prediction: Multimodal punctuation prediction combines force-aligned audio pauses with text features, refining pause thresholds through human feedback so punctuation follows speech timing.
3. THE WENETSPEECH-CHUAN CORPUS
WenetSpeech-Chuan is a large-scale, multi-label, multi-domain Sichuanese corpus with confidence-based partitions and dedicated ASR and TTS evaluation sets. Its data spans diverse sources and balances clean with real-world acoustic conditions.
- 3.1. Data Size and Confidence: 10,013 hours of raw audio comprise 3,714 hours of Strong Label data and 6,299 hours of Weak Label data partitioned by transcription confidence.Strong Label segments exceed 0.90 confidence, while Weak Label segments fall between 0.60 and 0.90.
- 3.2. Domain Distribution: Short videos contribute 52.83% of the corpus, followed by entertainment at 20.08% and live streams at 18.35%.Documentaries, audiobooks, interviews, news, reading, and drama form smaller source categories.
- 3.3. Quality Distribution: Audio quality scores concentrate between 2.5 and 4.0, peaking at 3.0–3.5 according to the WVMOS-based metric.The distribution combines relatively clean recordings with real-world acoustic conditions.
- 3.4. WenetSpeech-Chuan Eval Benchmark: WSC-Eval-ASR is a manually refined benchmark partitioned into Easy and Hard subsets for fine-grained analysis across source domains and acoustic environments.The 9.7-hour set also includes speaker attributes such as age, gender, and emotional state.
- 3.4. WenetSpeech-Chuan Eval Benchmark: WSC-Eval-TTS contains easy dialect-word sentences and hard long or LLM-generated sentences spanning tongue twisters, folk sayings, and emotional speech.Ten speakers, evenly split by gender, each provide 200 audio-prompt sentences.
4. EXPERIMENTS
Experiments evaluate ASR and TTS systems on Sichuanese benchmarks, including models trained or fine-tuned with WenetSpeech-Chuan. These evaluations show improved dialect recognition and competitive synthesis quality, with further gains from additional fine-tuning.
- 4.1. Automatic Speech Recognition: Paraformer achieves a state-of-the-art average CER of 13.38% across all test sets after additional fine-tuning with 1000 hours of internal data.The ASR table reports CER results on Sichuanese datasets; the caption identifies fine-tuned systems and foundation-model evaluation.
- 4.2. Speech Synthesis: The TTS evaluation measures CER, speaker similarity, intelligibility, speaker quality, and accent naturalness using objective metrics and listener ratings.AMOS ratings include ten native Sichuanese raters and ten non-expert listeners across 30 speech samples.
- 4.2. Speech Synthesis: CosyVoice2-WSC reaches 4.28% CER on the easy split and 8.78% on the hard split, while maintaining stronger perceptual quality than Qwen-TTS.On the easy split, its CER is close to Qwen-TTS’s 4.13%; on the hard split, Qwen-TTS reaches 7.35%, while CosyVoice2-WSC maintains SIM above 62%.
- 4.2. Speech Synthesis: CosyVoice2-WSC-SFT obtains 4.08% CER and 78.84% SIM on the easy split, then lowers hard-split CER to 7.22%.The fine-tuned system also achieves leading MOS-family scores and the best AMOS on the hard split.
5. CONCLUSION
WenetSpeech-Chuan provides a large, richly annotated Sichuanese corpus, a processing toolkit, and manually verified ASR and TTS benchmarks. Models trained on it achieve state-of-the-art open-source performance and results comparable to commercial systems.
- 5. CONCLUSION: WenetSpeech-Chuan comprises over 10,000 hours of Sichuanese speech with multi-dimensional annotations, constructed using the Chuan-Pipeline.The paper presents it as the largest open-source corpus for Chinese Sichuanese dialects.
- 5. CONCLUSION: The authors establish ASR and TTS evaluation benchmarks with manually verified transcriptions to address the lack of standardized testing.
- 5. CONCLUSION: Models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and perform comparably to commercial systems.