Source-linked AI summary
ICASSP 2021 Deep Noise Suppression Challenge
Chandan K A Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan
TL;DR
The challenge addresses degraded speech quality from background noise and the remaining gap between current real-time enhancement and ideal perceptual quality. It expands datasets, introduces real-time denoising and personalized tracks, and evaluates systems subjectively; results show substantial room for robust performance across conditions.
Problem
Background noise degrades speech quality and intelligibility, while prior DNS results remained about 1.4 DMOS from the ideal MOS of 5 on a realistic noisy test set.
Method
The challenge expands diverse speech and noise resources, provides over 118,000 room impulse responses, defines denoising and personalized tracks, and uses crowdsourced ITU P.808 subjective evaluation.
Results
The best real-time denoising model achieved 0.53 DMOS over noisy speech with absolute MOS 3.38, while the best personalized model achieved 0.14 DMOS.
Takeaways & Limitations
Noise suppression remains far from robust across singing, emotional, multilingual, and other challenging conditions, and personalized DNS is still in its infancy.
Takeaways & Limitations
Personalized enhancement is constrained to using two minutes of a particular speaker’s speech and must improve the same speaker’s noisy test segment over enhancement without speaker information.
Abstract
from arXiv · showhide
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020. We open sourced training and test datasets for researchers to train their noise suppression models. We also open sourced a subjective evaluation framework and used the tool to evaluate and pick the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we are expanding both our training and test datasets. There are two tracks with one focusing on real-time denoising and the other focusing on real-time personalized deep noise suppression. We also make a non-intrusive objective speech quality metric called DNSMOS available for participants to use during their development stages. The final evaluation will be based on subjective tests.
1. INTRODUCTION
The challenge targets perceptual speech-quality improvement for real-time noise suppression amid growing communication needs and diverse background noise. DNS Challenge 2 expands the datasets and introduces a personalized track focused on speaker information.
- Motivation: Remote communication needs high speech quality because background noise degrades perceived speech quality and intelligibility.The motivation also covers hearing aids and smart devices.
- Motivation: Earlier DNS results remained about 1.4 DMOS from the ideal MOS of 5 on the challenge test set.The earlier challenge used subjective evaluation on a realistic noisy test set.
- Challenge expansion: DNS Challenge 2 increases clean training speech by 50% to over 760 hours and adds diverse speech, noise, and room impulse-response resources.The expanded data include singing, emotion, Chinese, and over 118,000 room impulse responses.
- Challenge expansion: The challenge adds a personalized DNS track alongside the real-time denoising track, using speaker information to seek better perceptual quality.The real-time denoising track remains similar to the earlier challenge.
2. CHALLENGE TRACKS
The challenge is organized around two tracks. The supplied passage identifies the tracks but does not describe their individual requirements.
- Challenge tracks: The challenge has two tracks.
- Challenge tracks: The passage introduces the tracks as separate evaluation areas.
- Challenge tracks: Track-specific details are presented after the introductory track statement.
1. Track 1: Real-Time Denoising track requirements
The real-time track constrains processing latency and requires personalized systems to use speaker information on the same speaker’s noisy test segment. Speaker-informed enhancement must outperform enhancement without that information.
- Track 1: Real-Time Denoising track requirements: Real-time processing must complete each frame within the stride time on an Intel Core i5 quad-core machine or equivalent.The requirement is tied to frame size T and stride time Ts.
- Track 1: Real-Time Denoising track requirements: Total algorithmic latency, including frame size, stride, and look-ahead, must be at most 40ms.A 20ms frame with a 10ms stride gives 30ms algorithmic latency and satisfies the example requirement.
- Track 2: Personalized Deep Noise Suppression (pDNS) track requirements: Personalized DNS systems can use two minutes of a particular speaker’s speech to adapt enhancement for that same speaker’s noisy test segment.
- Track 2: Personalized Deep Noise Suppression (pDNS) track requirements: Speaker-informed enhancement must have better quality than enhancement without speaker information.
3. TRAINING DATASETS
The challenge releases a larger, more diverse training resource spanning speech varieties, balanced environmental noise, reverberation, and acoustic metadata. Its design addresses missing emotional, singing, and non-English speech coverage from the earlier dataset.
- Clean speech: 760.53 hours of clean speech combine read speech, singing, emotion data, and Chinese Mandarin, expanding beyond the 562.72-hour earlier set.The dataset also adds other non-English languages.
- Clean speech: Clean speech is divided into read, singing, emotional, and non-English subsets sourced from multiple public corpora.The read-speech subset includes recordings from 11,350 speakers.
- Noise: The noise dataset contains about 150 balanced audio classes, 60,000 clips, 10,000 additional clips, and 181 hours of noise.Speech activity detection removes clips containing speech, while sampling ensures each class has at least 500 clips.
- Room acoustics: The resource provides 3,076 real and approximately 115,000 synthetic room impulse responses for reverberant speech generation.Participants can perform dereverberation and denoising on noisy reverberant speech.
- Room acoustics: Training speech and room impulse responses include reverberation time T60 and clarity C50 metadata for selecting data subsets.The RIR metadata also identify whether each response is real or synthetic.
4. TEST SET
DNS Challenge 1 used a mixed real-and-synthetic test set, primarily comprising English noisy speech clips sampled at 16 kHz.
- The DNS Challenge 1 test set combined 300 real recordings with 300 synthesized noisy speech clips.Synthetic clips were divided into reverberant and less reverberant conditions.
4.1. Track 1
Track 1 evaluated general real-time denoising on diverse real and synthetic noisy speech, including multiple languages, emotional speech, and singing. Its blind test set contained 700 clips spanning five categories.
- Track 1 combined realistic recordings with synthetic scenarios that were difficult to collect under realistic conditions.The test design used synthetic clips mainly for scenarios unavailable through real-data collection.
- The real recordings included English and non-English speech, with non-English clips covering Portuguese, Russian, Spanish, Mandarin, Cantonese, Punjabi, and Vietnamese.The non-English segment included both non-tonal and tonal languages.
- The synthetic test set contained 200 noisy clips created by mixing clean non-English, emotional, and singing speech with noise.The synthetic set included 100 non-English, 50 emotional, and 50 singing clips.
- The blind test set contained 700 noisy speech clips: 650 real recordings and 50 synthetic noisy singing clips.The categories were emotional, English, non-English, tonal languages, and singing.
4.2. Track 2
Track 2 evaluated personalized denoising using clean adaptation speech for each primary speaker, with test conditions involving neighboring speakers, background noise, or both.
- Personalized DNS models were expected to use speaker-aware training and speaker-adapted inference.Clean adaptation data supported speaker modeling and adaptation to the primary speaker.
- Clean adaptation speech was also intended to provide accurate speech-activity labels and generate reverberant or noisy adaptation data.These uses supported both speaker modeling and multi-conditioned speaker adaptation.
- The development set included 100 real recordings from 20 primary speakers across neighboring-speaker, background-noise, and combined-noise scenarios.Each primary speaker had noisy clips for all three scenarios.
- The synthetic set included 500 noisy clips from 100 primary speakers with varying levels of neighboring speakers and noise.Primary speech came from VCTK, neighboring speakers from VoxCeleb2, and each speaker had two minutes of clean adaptation data.
- The blind test set contained 500 real noisy speech recordings from 80 unique speakers, mostly with a secondary speaker as the noise source.Each primary speaker also received two minutes of clean English speech for adaptation.
5. CHALLENGE RESULTS
The challenge used subjective P.808 evaluation for both tracks and found that robust denoising remained difficult, especially for emotional, singing, and personalized conditions.
- Evaluation Methodology: Subjective P.808 evaluation was used because objective metrics did not correlate well with subjective quality in background noise.The evaluation used five raters per clip and achieved a 95% confidence interval of 0.03.
- Track 1: Track 1 results showed that most methods struggled with singing and emotional clips, and only half of submissions outperformed the baseline.The results also indicated a need for balanced training data across English, non-English, and tonal languages.
- Track 1: 0.53 DMOS was the best Track 1 improvement over noisy speech, corresponding to an absolute MOS of 3.38.The authors concluded that robust performance across nearly all conditions remained distant.
- Track 2: Only two teams participated in Track 2, whose best model achieved 0.14 DMOS.The authors characterized speaker-information-based adaptation as still being in an early stage.
- Track 2: Track 2 results were interpreted as evidence that personalized noise suppression using speaker information remained an immature research area.The track was described as the first of its kind with limited prior work.
6. SUMMARY & CONCLUSIONS
The challenge advances real-time noise suppression for perceptual quality in difficult noisy conditions through open resources and broad participation, while personalized DNS remains nascent.
- The challenge targets real-time noise suppression optimized for human perception in challenging noisy conditions.
- It open sourced large, inclusive, diverse training and test datasets with supporting scripts.
- Many industry and academic participants found the datasets useful and submitted enhanced clips for final evaluation.
- Only two teams entered the personalized DNS track, indicating that this research area remains nascent.