Source-linked AI summary
Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement
Alessia Milo, Georg Götz, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
TL;DR
Single-channel speech enhancement often trains on simplified RIRs, leaving the effect of acoustic realism underexplored. The paper compares complete ISM and higher-fidelity Hybrid RIR pipelines with unchanged DeepFilterNet3 training, evaluating unseen measured environments. Hybrid training consistently improves objective enhancement metrics and substantially lowers ASR word error rates, while the study cannot attribute gains to individual acoustic factors.
Problem
The effect of RIR fidelity on DeepFilterNet3 generalization remains underexplored because enhancement systems commonly use simplified ISM RIRs that do not fully capture real acoustic conditions.
Method
The study compares complete ISM and higher-fidelity Hybrid RIR generation pipelines while keeping DeepFilterNet3 architecture and training procedure unchanged.
Results
Hybrid-trained models consistently achieve modestly better objective enhancement metrics and substantially better downstream ASR performance on unseen measured environments than ISM-trained models.
Takeaways & Limitations
Increasing overall synthetic RIR realism improves DeepFilterNet3 generalization across speech enhancement and downstream recognition tasks.
Takeaways & Limitations
Because geometry, furnishing, materials, and simulation method vary simultaneously, the experiments cannot attribute gains to any individual acoustic modelling component.
Abstract
from arXiv · showhide
We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
1. INTRODUCTION
Single-channel enhancement remains difficult because devices lack spatial information and commonly used ISM-generated RIRs do not fully capture real acoustics. This study tests whether replacing simplified RIRs with a more realistic complete simulation pipeline improves DeepFilterNet3 generalization.
- Single-microphone enhancement is essential in constrained consumer and embedded devices but lacks spatial information under adverse acoustics.
- Deep-learning enhancement systems commonly rely on simplified image-source-method RIRs that do not fully capture real acoustic conditions.
- The impact of RIR fidelity on enhancement performance remains underexplored despite DeepFilterNet enabling real-time denoising and dereverberation.
- The paper compares complete conventional ISM and higher-fidelity hybrid RIR pipelines while keeping DeepFilterNet3 training and architecture unchanged.
- Evaluation tests whether greater overall simulation realism improves objective enhancement metrics and downstream ASR on unseen measured RIRs.
2. ACOUSTIC SIMULATION
Room-acoustic simulation spans geometrical and wave-based methods with different fidelity and computational properties. Hybrid methods combine wave-based low-frequency modeling with geometrical-acoustics modeling at higher frequencies.
- Geometrical-acoustics methods approximate propagation through specular reflections and energy transport assumptions that are most valid at higher frequencies.
- Wave-based methods solve the acoustic wave equation directly, modeling diffraction, modal behavior, and frequency-dependent boundary conditions especially relevant at low and mid frequencies.
- Hybrid approaches use wave-based solvers below a crossover frequency and geometrical acoustics above it.
- Figure 1 presents histograms of Sabine RT60 values for the matched Hybrid and ISM datasets.
- Figure 2 shows the level of detail for a room simulated in the Hybrid dataset.
3. EXPERIMENTS
The experiments compare matched ISM and higher-fidelity Hybrid RIR datasets while holding DeepFilterNet3 unchanged and varying training conditions. The Hybrid dataset incorporates richer room, material, and simulation properties, making the comparison holistic rather than component-specific.
- Experiments: The study compares an ISM RIR dataset with a higher-fidelity Hybrid dataset using a fixed DeepFilterNet3 training and evaluation protocol.
- Acoustic datasets: The DNS4-based ISM dataset contains 60 000 RIRs across small, medium, and large shoebox-room categories.
- Acoustic datasets: Hybrid and ISM rooms are paired using Sabine RT60 estimates to match dataset size and reverberation-time distributions.
- Acoustic datasets: The Hybrid dataset contains furnished living rooms, classrooms, and restaurants with diverse geometries and structural complexity.
- Acoustic datasets: Hybrid materials use frequency-dependent complex surface impedances assigned to physical counterparts such as glass, wood, plastic, gypsum, and concrete.
- Acoustic datasets: Hybrid simulation combines wave-based modeling up to a room-dependent 1–2 kHz crossover with geometrical acoustics at higher frequencies up to 12 kHz.
- Experimental scope: Geometry, furnishing, materials, and simulation method differ jointly between datasets, so the experiment evaluates overall realism rather than individual factors.
- Training configurations: Training configurations vary data size, epoch count, and decay augmentation to test robustness across training conditions.
4. EVALUATION
Evaluation uses unseen speech, measured RIRs, and held-out noise to assess objective enhancement quality and ASR. Across configurations, Hybrid training improves objective metrics and yields a systematic advantage over ISM.
- Evaluation protocol: Evaluation uses 827 unseen VCTK utterances, 284 measured ACE and MIT RIRs, and held-out AudioSet and Freesound noise clips.
- Evaluation protocol: Objective evaluation covers PESQ, SI-SDRi, STOI, and SRMR, while enhanced files are additionally transcribed for ASR.
- Objective results: Across all configurations, Hybrid has positive mean gains with 95% bootstrap confidence intervals strictly positive for all four objective metrics.
5. RESULTS
Training with the Hybrid RIR dataset consistently improves objective enhancement metrics over ISM and produces larger gains in downstream ASR. These results support higher overall simulation realism as a generalization benefit, while not identifying any individual acoustic factor as responsible.
- Objective evaluation: Hybrid improves PESQ, SI-SDRi, STOI, and SRMR over ISM across training configurations.In the small 30-epoch setting, the gains are +0.18, +0.21, +0.014, and +0.18, respectively.
- Objective evaluation: Positive mean gains with strictly positive 95% bootstrap confidence intervals occur for all four objective metrics.The intervals represent average paired improvements pooled across configurations, not variability within individual settings.
- ASR evaluation: Hybrid achieves lower WER than ISM in every matched ASR configuration.Relative improvement over the noisy baseline ranges from 12.0% to 23.8% for Hybrid, compared with 0.5% to 13.1% for ISM; the same pattern holds in small-data models.
- ASR evaluation: ASR gains are substantially larger than the modest improvements in conventional enhancement metrics.The results suggest that higher-fidelity training better preserves speech characteristics relevant for recognition than conventional perceptual metrics alone reflect.
- Interpretation and limitations: The comparison demonstrates an overall realism advantage rather than the contribution of any single acoustic modelling component.Geometry, furnishing, frequency-dependent materials, and simulation method differ jointly between the complete pipelines.
- Interpretation and limitations: Future work should use controlled ablations and complementary perceptual evaluation to disentangle simulation factors and assess perceptual significance.Proposed evaluations include listening tests and learned quality metrics such as DNSMOS or UTMOS.
6. CONCLUSION
The study finds that more realistic synthetic RIR training data improves DeepFilterNet3 generalization across objective enhancement and downstream ASR evaluations. The conclusion supports practical use of realistic simulation pipelines while reserving individual-factor attribution for future controlled studies.
- Conclusion: Higher-fidelity synthetic RIR training yields modest objective-metric gains and substantially larger ASR improvements on unseen measured environments.This pattern holds across all evaluated training configurations.
- Conclusion: Increasing synthetic training-data realism improves model generalization without establishing which acoustic modelling component produces the gains.The finding complements studies that isolate individual simulation factors from a practical dataset-generation perspective.
- Conclusion: Future studies should vary room geometry, material modelling, and frequency-dependent acoustic simulation independently, alongside broader perceptual evaluation.Listening tests and learned quality metrics are identified as complementary evaluation methods.