Source-linked AI summary
ABSE-NET: A Lightweight Neural Model for Active Binaural Speech Enhancement in Open-Fit Hearing Aids
De Hu, Xue Du, Qingying Zhao, Qintuya Si
TL;DR
Open-fit hearing aids improve wearing comfort but suffer acoustic leakage that degrades binaural speech enhancement. ABSE-NET cascades BMVDR with a lightweight neural network to cancel leakage and compensate for distortion, and experiments report superiority over state-of-the-art methods.
Problem
Open-fit hearing aids introduce acoustic leakage into the ear canal, limiting the effectiveness of existing binaural speech enhancement methods designed mainly for closed-fit devices.
Method
ABSE-NET cascades a BMVDR beamformer with a lightweight encoder–decoder neural network containing frequency-time dependency learning and convolutional attention blocks.
Results
ABSE-NET outperforms state-of-the-art approaches while exhibiting remarkable computational efficiency.
Takeaways & Limitations
The framework jointly suppresses acoustic leakage and compensates for BMVDR-induced distortion while preserving spatial cues in open-fit hearing aids.
Takeaways & Limitations
BMVDR performance remains vulnerable to acoustic-transfer-function and spatial-covariance estimation errors, which can distort speech and spatial cues and degrade noise reduction.
Abstract
from arXiv · showhide
Open-fit hearing aids have attracted growing attention due to their superior wearing comfort. However, the open-fit design inevitably causes acoustic leakage into the ear canal, degrading the performance of existing binaural speech enhancement (BSE). To this end, we propose ABSE-NET, an active BSE framework integrating active noise control (ANC) with BSE to jointly enhance target speech and suppress acoustic leakage. The ABSE-NET pipeline cascades a binaural MVDR (BMVDR) with a lightweight neural network (LNN). The former achieves a coarse BSE, whereas the latter simultaneously cancels acoustic leakage and compensates for BMVDR-induced distortion. The LNN uses an encoder-decoder with a feature fusion module, which includes frequency-time dependency learning and convolutional attention blocks. Unlike traditional BSE+ANC solutions via adaptive filtering, ABSE-NET needs no in-ear microphone in practical deployment. Experiments validate its superiority over state-of-the-art methods. Code repository: https://github.com/Bream101/ABSE-NET.
1. Introduction
Open-fit hearing aids reduce occlusion-related discomfort but allow external noise to leak into the ear canal, degrading binaural speech enhancement. ABSE-NET addresses this challenge by cascading BMVDR beamforming with a lightweight neural network that jointly suppresses leakage and compensates for distortion.
- 1. Introduction: Open-fit hearing aids use vents to relieve ear-canal pressure and avoid the occlusion effect, but external noisy signals inevitably leak into the ear canal.The leakage degrades BSE performance and corrupts the enhanced signal.
- 1. Introduction: Existing BSE methods largely target closed-fit hearing aids and therefore cannot handle acoustic leakage in open-fit configurations.This limitation affects both model-driven and data-driven approaches.
- Active BSE: Active BSE combines BSE with ANC by having the internal loudspeaker play an anti-leakage signal that cancels leakage through destructive interference.The paper describes both cascaded and parallel BSE–ANC architectures.
- 1.2. Contribution: ABSE-NET cascades a BMVDR beamformer with a lightweight neural network, combining model-driven coarse enhancement with data-driven leakage cancellation and distortion compensation.The BMVDR preserves spatial cues, while the neural network refines its output.
- 1.2. Contribution: ABSE-NET is presented as a lightweight active BSE framework that does not require an in-ear error microphone during practical deployment.The contribution passage also reports superiority over state-of-the-art approaches and computational efficiency.
2. Preliminaries
The preliminaries model open-fit hearing-aid signals and identify leakage and BMVDR distortion as the central problems. ABSE-NET therefore uses a data-driven post-filter to cancel leakage while compensating for beamformer-induced distortion, with training-time error-microphone information reducing deployment dependence.
- 2. Preliminaries: The open-fit model represents each hearing aid with M/2 microphones, where M is even, and processes multichannel signals in the STFT domain.Time and frequency indices are used in the formulation and later omitted for notational simplicity.
- 2. Preliminaries: The microphone observation model combines target speech, I interferers filtered by acoustic transfer functions, background noise, and sensor self-noise.The target and interferer components use acoustic transfer functions a(k) and b_i(k), respectively.
- 2. Preliminaries: BMVDR beamforming preserves the target signal at each reference microphone while minimizing interferer-plus-noise power, enabling fast closed-form computation.Its filter coefficients depend on the estimated interferer-plus-noise covariance and target acoustic transfer function.
- 2. Preliminaries: In open-fit hearing aids, leakage enters through the vent and combines with the BMVDR output, reducing SINR and potentially producing comb-filtering artifacts.The leakage path is characterized by d_L, with g_L representing the loudspeaker-to-error-microphone secondary path.
- 2. Preliminaries: ATF and SCM estimation errors distort target speech and spatial cues while degrading noise reduction performance.These errors create mismatch in the BMVDR constraint and impair its interference-plus-noise estimate.
- 2. Preliminaries: The proposed data-driven post-filter uses the noisy left reference signal as an auxiliary input while canceling leakage and compensating for BMVDR distortion.The auxiliary signal helps generate anti-leakage output without excessive noise suppression.
- 2. Preliminaries: The error microphone is mainly used for signal modeling and model training, reducing the need for real-time error feedback during inference and deployment.This design supports practical operation without continuous in-ear error-microphone feedback.
3. Proposed Method
ABSE-NET combines coarse binaural MVDR enhancement with a lightweight neural network that refines features, suppresses acoustic leakage, and compensates for BMVDR distortion. Its design emphasizes frequency–time dependency modeling, lightweight attention, causal processing, and a loss balancing waveform quality with intelligibility.
- Pipeline: ABSE-NET independently processes each hearing aid by cascading BMVDR coarse enhancement with an LNN that receives the BMVDR output and noisy reference signal.The decoder output is transformed to the time domain, emitted by the loudspeaker, and propagated through the secondary path to interfere destructively with leakage.
- Network architecture: The LNN uses an encoder, repeated feature-augmentation modules, and a decoder to transform latent features back into an STFT-domain signal.Each feature-augmentation module contains an F-TDL block followed by a ConvAtt block.
- Loss Function: The loss combines negative SI-SDR with a weighted negative STOI term to jointly optimize waveform reconstruction quality and perceptual intelligibility.The clean speech u is the ground-truth target, and λ controls the relative contribution of the two objectives.
- F-TDL Block: The F-TDL block separates frequency and time dependency learning, processing frequency information independently across time and temporal information causally across frequency bins.Its design uses residual structures and avoids computationally expensive multi-head attention; causal convolutions restrict temporal context to past and present frames.
- F-TDL Block: The FDL multi-scale convolution can be fused into a single inference-time convolution while retaining the multi-scale structure’s modeling capacity.Zero-padding aligns kernels before fusion, reducing inference complexity without changing the stated modeling capacity.
- ConvAtt Block: ConvAtt refines features through sequential channel and frequency–time attention maps instead of a full 3D attention tensor.The frequency–time attention uses intermediate variables G(Φ′) and ρ(Φ′), while the module preserves a lightweight design.
4. Experiments
Experiments evaluate ABSE-NET on simulated open-fit hearing-aid conditions, compare it with model- and data-driven baselines, and analyze its architecture, efficiency, and robustness to ATF errors.
- Experimental Setup: The dataset combines Librispeech speech, NOISEX-92 noise, and Hearpiece HRIRs across 24 directions, producing 43,200 two-second samples at random -5 dB to 0 dB SNR.Training, validation, and testing sets use an 8:1:1 split with strict mutual exclusivity.
- Comparison Study: ABSE-NET achieves the best or second-best speech quality across evaluated metrics, while ASE-TM reaches the highest reported SI-SDR at 10.45 dB.The comparison includes BMVDR variants, FxMWF, DeepANC, ASE-TM, and ABSE-NET.
- Comparison Study: BMVDR performance degrades sharply with acoustic leakage, with SI-SDR falling from 5.216 dB to 0.878 dB and PESQ from 3.437 to 2.196 in open-fit conditions.FxMWF is designed for open-fit hearing aids but remains slightly below BMVDR without acoustic leakage, while its linear filtering may not fully eliminate leakage in complex scenes.
- Effectiveness of RMB-Conv1D: Heterogeneous RMB-Conv1D kernel sizes achieve the best overall performance, capturing coarse- and fine-grained acoustic features while maintaining comparable parameter counts and FLOPs.Smaller kernels model local spectral variations, whereas larger kernels capture broader contextual dependencies; multi-scale convolutions are fused into one convolution at inference.
- Effectiveness of FA Module: The F-TDL block outperforms SpatialNet with 1.06 dB higher SI-SDR and seven times fewer FLOPs, while removing either FDL or TDL substantially degrades enhancement.The ablation supports retaining both frequency- and time-dependency sub-blocks in the feature augmentation design.
- Effectiveness of FA Module: Removing ConvAtt reduces SI-SDR to 7.441 dB, PESQ to 2.939, and STOI to 0.928, despite slightly lowering parameters and computational cost.Under a 15° DOA mismatch, ABSE-NET maintains 6.434 dB SI-SDR and 3.001 PESQ, with the smallest ILD and IPD errors across tested conditions.
5. Conclusion
The conclusion presents ABSE-NET as a lightweight open-fit hearing-aid framework that combines BMVDR and a neural network to enhance speech, suppress leakage, and preserve spatial cues without an in-ear microphone.
- 5. Conclusion: ABSE-NET cascades BMVDR coarse enhancement with an LNN that cancels acoustic leakage and compensates for BMVDR-induced distortion.Its LNN uses an encoder, feature augmentation module, and decoder to process frequency-time representations.
- 5. Conclusion: The framework combines model-driven and data-driven approaches while requiring no in-ear microphone for practical deployment.The conclusion reports superior binaural enhancement performance over model-driven approaches and better performance and computational efficiency than existing data-driven methods.
- 5. Conclusion: Future work will target further computational-efficiency improvements and deployment on practical open-fit hearing aids.
6. Generative AI Use Disclosure
The disclosure states that Generative AI tools were used only as auxiliary support and that the author verified the generated content and remains responsible for the final thesis.
- 6. Generative AI Use Disclosure: Generative AI tools assisted with literature-review sorting, language optimization, and brainstorming without replacing the author’s independent thinking or conclusions.
- 6. Generative AI Use Disclosure: The author verified and corrected AI-generated content and accepts full responsibility for the thesis’s final content.
- 6. Generative AI Use Disclosure: The disclosure states that the AI use complies with academic-integrity requirements.