Source-linked AI summary
Voices Obscured in Complex Environmental Settings (VOICES) corpus
Colleen Richey, Maria A. Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, Paul Gamble, Jeff Hetherly, Cory Stephenson, Karl Ni
TL;DR
Speech research lacks accessible datasets that capture real-world acoustics, while synthetic mixtures do not fully represent dynamic noise and room conditions. VOICES addresses this gap with an open corpus recorded using distant microphones in furnished rooms with concurrent background noise, providing data for speech and acoustic research.
Problem
Speech datasets are often expensive, limited, or inaccessible, and synthetic mixtures do not accurately represent real-world acoustics and dynamic noise.
Method
VOICES records foreground speech with concurrent distractor noise in two furnished rooms with different acoustic profiles using 12 distant microphones.
Results
The corpus provides audio recorded under realistic noisy conditions that better represent real-use situations for speech and acoustic research.
Takeaways & Limitations
The publicly available corpus can support development and evaluation of robust acoustic models for speech, speaker, and related acoustic-processing tasks.
Takeaways & Limitations
The foreground speech and distractor noise were selected from sources permitting data derivatives and commercial use under public-domain or Creative Commons licenses.
Abstract
from arXiv · showhide
This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions. Publicly available speech corpora are mostly composed of isolated speech at close-range microphony. A typical approach to better represent realistic scenarios, is to convolve clean speech with noise and simulated room response for model training. Despite these efforts, model performance degrades when tested against uncurated speech in natural conditions. For this corpus, audio was recorded in furnished rooms with background noise played in conjunction with foreground speech selected from the LibriSpeech corpus. Multiple sessions were recorded in each room to accommodate for all foreground speech-background noise combinations. Audio was recorded using twelve microphones placed throughout the room, resulting in 120 hours of audio per microphone. This work is a multi-organizational effort led by SRI International and Lab41 with the intent to push forward state-of-the-art distant microphone approaches in signal processing and speech recognition.
1. Introduction
VOICES addresses the need for realistic, publicly available speech data recorded with distant microphones in noisy, reverberant environments. The corpus is intended to support speech, acoustic, and signal-processing research with source audio, retransmitted audio, transcriptions, and speaker labels.
- VOICES provides speech data recorded in acoustically challenging reverberant environments with concurrent background noise.
- VOICES is intended to support speaker identification, speech recognition, event and background classification, source separation, localization, noise reduction, enhancement, and acoustic-quality research.
- The corpus includes source and retransmitted audio, orthographic transcriptions, and speaker labels.
- Existing speech datasets are often expensive, limited, paywalled, or based on isolated speech and synthetic reverberation that poorly represent real-world acoustics and dynamic noise.
- The authors recorded under realistic noisy conditions to support development and evaluation of acoustic algorithms intended for field deployment.
2. Dataset Collection
The VOICES collection combines furnished rooms with different acoustic profiles, concurrent foreground speech and distractor noise, and distributed distant microphones. Multiple sessions and source materials were organized to provide varied recording conditions and reusable audio data.
- Recordings used two furnished rooms with different acoustic profiles, 12 distant microphones, and four sessions per room covering three distractor noises plus foreground speech alone.
- Foreground speech came from 300 English-speaking LibriSpeech speakers, with at least three minutes selected per speaker for speaker-identification tasks.
- The four recording conditions were ambient room noise only, television, music, or overlapping speech played concurrently with foreground speech.
- Babble consisted of nine overlapping voices played simultaneously from three noise-dedicated loudspeakers.
- The setup used 7 cardioid studio microphones and 5 omnidirectional lavalier microphones placed throughout the rooms.
- A robotic platform rotated the foreground loudspeaker by 10 degrees per hour across 180 degrees to emulate non-static conversational sources.
3. Data Statistics
Corpus statistics indicate consistent segmentation and playback levels while showing that realistic room distance and distractor noise reduce signal-to-noise ratio. The measurements characterize amplitude, energy, duration, and recording conditions across subsets.
- Statistics were calculated for duration, amplitude, RMS energy, and SNR using SoX and SRI in-house utilities.
- Average and median file durations were 15.62s and 15.97s, respectively, with a 1.91s standard deviation, supporting direct comparison between noisy and source files.
- Average RMS levels ranged from -27.0 to -27.5 dBFS across subsets, indicating consistent playback volume.
- Recorded-file minimum and maximum amplitudes ranged from -0.5 to 0.5 across subsets, while source audio averaged -0.93 and 0.91.
- Room-1 and room-2 recordings averaged 22.19 dB and 19.50 dB SNR, respectively, with SNR degrading as microphone distance increased and further under distractor noise.
4. Model Baselines
Baseline ASR and speaker-identification systems both show substantial degradation under distant-microphone conditions, with further effects from distractor noise and room acoustics.
- Automatic speech recognition (ASR): WER is 19.0% for distant microphones without added distractor noise, more than double the source-audio WER, and babble produces the worst performance.Added distractor noise further degrades ASR, while speech-like babble most strongly confuses the system.
- Automatic speech recognition (ASR): ASR WER increases with microphone distance, while differences between comparable microphone distances across rooms reflect room-specific acoustic environments.The comparison uses microphones 02, 04, and 06 with the foreground loudspeaker centered at 90°.
- Speaker identification (SID): Speaker-identification EER rises from 5.72% on source audio to 10.7%-10.9% for close microphones and 15.1%-16.6% for far microphones.Enrollment used clean source audio and testing used distant microphones without distractor noise.
- Speaker identification (SID): Distractor noise increases speaker-identification EER by 2% absolute for music and television and 3.5% absolute for babble.Babble may be more damaging because it is speech-like and was played from three separate loudspeakers.
5. Conclusions and Future Work
VOICES provides realistic distant-microphone recordings with background noise and reverberant acoustics for speech and acoustic research. The corpus is intended to support robust models and will be expanded with additional rooms and more challenging distractors.
- Conclusions and Future Work: VOICES provides realistic distant-microphone audio with background noise and reverberant room acoustics for speech and acoustic research.The corpus supports development and evaluation across speech and acoustic-processing tasks.
- Conclusions and Future Work: Phase II will add more rooms and more challenging distractor-noise profiles to the phase I data collection.The planned expansion broadens the environmental conditions represented by the corpus.