Source-linked AI summary

ICASSP 2026 URGENT Speech Enhancement Challenge

Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang, Kohei Saijo, Samuele Cornell, Yihui Fu, Zhaoheng Ni, Tim Fingscheidt, Shinji Watanabe, Yanmin Qian

arXiv:2601.13531v1eess.AScs.SD

TL;DR

The paper addresses the limited generalizability of speech enhancement systems under diverse distortions, domains, languages, and input conditions. It presents the ICASSP 2026 URGENT Challenge with universal enhancement and speech-quality-assessment tracks, broad datasets, baselines, and evaluation protocols. Over 80 teams registered and 29 submitted valid entries, while results highlighted hybrid generative-discriminative systems and curated high-quality data.

  • Problem

    Speech enhancement systems often struggle with unseen environments, speaker characteristics, and distortion types because prior evaluations commonly use matched conditions and limited distortions.

  • Method

    The challenge evaluates universal speech enhancement and MOS prediction using diverse speech conditions, curated datasets, baseline systems, objective metrics, and subjective human ratings.

  • Results

    Hybrid generative-discriminative systems consistently outperformed standalone generative or discriminative baselines across intrusive and non-intrusive metrics, while Track 2 leaders achieved high correlation with human ratings.

  • Takeaways & Limitations

    The results underscore the importance of high-quality training data and hybrid generative-discriminative models for diverse distortions and languages.

Abstract

from arXiv · show

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge is divided into two complementary tracks. Track 1 focuses on universal speech enhancement, while Track 2 introduces speech quality assessment for enhanced speech. The challenge attracted over 80 team registrations, with 29 submitting valid entries, demonstrating significant community interest in robust SE technologies.

1. INTRODUCTION

The challenge targets speech enhancement systems that generalize beyond matched conditions to diverse distortions, speakers, and acoustic environments. It organizes this effort around expanded data diversity, curation, and multi-stage evaluation, attracting broad participation.

  • Speech enhancement systems often struggle with unseen acoustic environments, speaker characteristics, and distortion types because prior evaluations use matched conditions and limited distortions.
  • The challenge emphasizes curated data selection and expanded diversity spanning age, accents, whispered and singing voices, and emotional expression.
  • 80 teams registered for Track 1 and 77 for Track 2, with 29 teams submitting valid entries across both tracks.
  • Track 1 blind-test evaluation advances the top six objective-ranking systems to subjective testing using ITU-T P.808 ACR and CCR ratings.

2. CHALLENGE DESCRIPTION

The challenge's first track requires one speech enhancement system to operate across varied distortions, domains, formats, sampling frequencies, and acoustic conditions without knowing the input type.

  • Track 1: Universal Speech Enhancement: Track 1 requires a single system to handle diverse distortions, domains, and formats while adapting to sampling frequencies from 8 to 48 kHz.
  • Track 1: Universal Speech Enhancement: Systems must process inputs across acoustic conditions without prior knowledge of the input type.
  • Distortions and evaluation: The task covers additive noise, reverberation, clipping, bandwidth limitation, codec distortion, packet loss, and wind noise.
  • Distortions and evaluation: Evaluation measures whether enhanced speech matches reference quality and intelligibility, or remains high quality for real recordings without references.

3. DATASETS

The datasets combine broad speech, noise, and room-response sources with curated training data and challenging blind-test conditions. Track 1 also evaluates generalization to five unseen languages.

  • Track 1 Datasets: Track 1 training data combine public speech corpora covering audiobooks, reading, multiple languages, children, elderly speakers, singing, and emotional speech.
  • Track 1 Datasets: Noise sources include Audioset, WHAM!, FSD50K, and Free Music Archive, alongside public speech and room impulse-response sources.
  • Data curation: The challenge encourages selecting high-quality data from large corpora instead of indiscriminately using all available recordings.
  • Data curation: The baseline curation method selects 700 hours of high-quality speech from the original 2500-hour URGENT 2025 training set.
  • Evaluation data: The Track 1 blind test contains 360 simulated and 480 real-world samples, including unseen Hindi, Korean, Arabic, Japanese, and Italian.

4. BASELINE SYSTEMS

The challenge provides discriminative and generative baselines for universal speech enhancement, plus a multi-metric baseline for speech quality assessment. An additional leaderboard baseline is reported for comparison.

  • Track 1: Track 1 uses BSRNN with adaptive STFT for sampling-frequency-independent processing and FlowSE with conditional flow matching to generate clean speech.
  • Track 2: Track 2 uses Uni-VERSA-Ext, which incorporates multi-metric supervision for MOS prediction.
  • Additional baseline: URGENT-PK is included in the final leaderboard as an additional comparison baseline.

5. EVALUATION AND RANKING METHODS

Track 1 uses a two-stage ranking process that combines objective multi-metric qualification with subjective human evaluation. The protocol covers speech quality, intelligibility, linguistic and speaker-related properties, and statistical significance.

  • Track 1 Evaluation: Stage 1 ranks submissions with a Friedman-test-inspired algorithm using non-intrusive, intrusive, and downstream-task metrics.Metrics include DNSMOS, NISQA, UTMOS, SCOREQ, PESQ, ESTOI, POLQA, SpeechBERTScore, LPS, speaker and emotion similarity, language identification, and character accuracy.
  • Track 1 Evaluation: The top six systems advance to Stage 2, where human listeners use ITU-T P.808 ACR and CCR to assess speech quality.Listeners evaluate submissions on Amazon Mechanical Turk.
  • Track 1 Evaluation: Final rankings use MOS and CMOS from the blind test, with statistical significance testing to assess ranking robustness.This follows earlier ITU-T standardization efforts.

6. RESULTS

Track 1 results favored hybrid generative-discriminative systems and quality-focused data practices, while Track 2 systems achieved high agreement with human speech-quality ratings.

  • Track 1: 23 Track 1 submissions competed, and the top six advanced to final subjective evaluation.Leading systems predominantly used hybrid generative-discriminative designs.
  • Track 1: Hybrid pipelines consistently outperformed standalone generative or discriminative baselines across intrusive and non-intrusive metrics.Top entries commonly combined generative restoration with discriminative signal preservation in dual-branch or multi-stage designs.
  • Track 1: Top-ranked teams emphasized MOS-based filtering, pretrained restoration models for data cleaning, and tailored augmentation to improve generalization.These practices align with the challenge’s emphasis on data efficiency and quality-aware training.
  • Track 2: Track 2’s top-ranked systems achieved high correlation with human ratings at both system and utterance levels.They used Uni-VERSA-Ext-style training and improved model components; URGENT-PK achieved the highest system-level correlation metrics outside the competition.

7. CONCLUSION

The challenge broadened universal speech enhancement and speech-quality assessment through diverse data, curation, and evaluation across two tracks. Its results highlight hybrid modeling and high-quality training data, while detailed analysis remains deferred.

  • Conclusion: The results underscore high-quality training data and hybrid generative-discriminative models for diverse distortions and languages.The challenge builds on previous URGENT editions in universal speech enhancement and quality assessment.
  • Conclusion: A more detailed analysis and comprehensive overview of the challenge are deferred to future work because of space limitations.The conclusion identifies scalable and multimodal technologies as directions for future research.
Loading 2601.13531v1…