Source-linked AI summary
Enhanced Robot Audition Based on Microphone Array Source Separation with Post-Filter
Jean-Marc Valin, Jean Rouat, François Michaud
TL;DR
The paper addresses how mobile robots can separate simultaneous, non-stationary sound sources in noisy environments. It combines a microphone-array Geometric Source Separation stage with a multi-channel post-filter that adapts to new sources and interference. Experiments report reductions in log spectral distortion and increases in signal-to-noise ratio, with acceptable distortion for most listeners.
Problem
Mobile robots need to separate simultaneous environmental sound sources, but the paper focuses on developing artificial hearing capable of handling interfering and non-stationary sounds.
Method
The system combines a microphone array, frequency-domain Geometric Source Separation, and a multi-channel post-filter using stationary-noise and leakage estimates.
Results
Up to 11 dB lower log spectral distortion and 14 dB higher signal-to-noise ratio were obtained versus noisy signal inputs, with acceptable distortions for most listeners.
Takeaways & Limitations
The proposed separator and post-filter provide rapid adaptation to new sources and non-stationarity while improving separation of simultaneous sound sources.
Abstract
from arXiv · showhide
We propose a system that gives a mobile robot the ability to separate simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation and a post-filter that gives us a further reduction of interferences from other sources. We present results and comparisons for separation of multiple non-stationary speech sources combined with noise sources. The main advantage of our approach for mobile robots resides in the fact that both the frequency-domain Geometric Source Separation algorithm and the post-filter are able to adapt rapidly to new sources and non-stationarity. Separation results are presented for three simultaneous interfering speakers in the presence of noise. A reduction of log spectral distortion (LSD) and increase of signal-to-noise ratio (SNR) of approximately 10 dB and 14 dB are observed.
I. INTRODUCTION
The paper targets mobile robots that must detect, localise, separate, and interpret simultaneous environmental sounds. It proposes a two-step separation system combining linear Geometric Source Separation with a multi-channel post-filter for rapidly changing sources.
- Mobile robots need to detect, localise, separate, and process simultaneous environmental sound sources.
- The paper frames source separation as a robotic version of the cocktail party effect, requiring separation of all sources present at a given time.
- The system uses eight low-cost microphones and separates sources in two stages.The first stage performs linear separation; the second enhances its outputs with a post-filter.
- The linear stage uses a simplified Geometric Source Separation method with faster stochastic-gradient estimation and shorter estimation frames.
- The multi-channel post-filter adaptively estimates stationary noise and interfering sources, including leakage between initial separation channels.
- The system is designed for sources that may appear, disappear, or move, while maintaining low delay and real-time processing complexity.
III. LINEAR SOURCE SEPARATION
The linear separator operates in the frequency domain, estimating a separation matrix that combines decorrelation with geometric constraints to suppress interference and preserve desired-source gain.
- The proposed linear source separation method adapts Geometric Source Separation for mobile-robot use and supports relatively low-complexity implementation.
- The frequency-domain model represents microphone signals as mixtures of unknown sources transformed by a frequency-dependent transfer matrix.
- The separated output is computed as y(k, ℓ) = W(k, ℓ)x(k, ℓ), where W(k, ℓ) is the estimated separation matrix.
- The method enforces output decorrelation and the geometric constraint W(k)A(k) = I.The geometric constraint provides unity gain toward the source of interest and zeros toward interferences.
- The separation matrix is updated iteratively using a gradient-based rule with adaptation rate µ and energy normalisation factor α(f).
B. Stochastic Gradient Adaptation
The stochastic-gradient adaptation replaces multi-second correlation estimates with instantaneous estimates to reduce computational cost while retaining reported separation accuracy.
- The algorithm uses instantaneous estimates of Rxx(k) and Ryy(k) instead of estimating correlations over several seconds.
- The approximation is analogous to the Least Mean Square adaptive filter and reduces the required computation to matrix-by-vector products.
- Second-order statistics are treated as sufficient for ensuring independence under the assumption of non-stationary sources.
- Instantaneous correlation estimation showed no reduction in accuracy and eased real-time integration.
C. Initialisation
The initialization scheme is designed to accommodate sources that appear or disappear while preserving acceptable separation before adaptation and avoiding disruption to existing sources.
- Source appearance or disappearance requires initialization of the separation matrix W(k) to support new sources.
- A new-source initialization must provide initial weights and acceptable separation before adaptation.
- When a source appears or disappears, the other sources should remain unaffected.
- The proposed initialization sets the column of W(k) corresponding to a new source using a delay-and-sum beamformer.
IV. MULTI-CHANNEL POST-FILTER
The post-filter enhances geometric source separation by estimating both stationary background noise and leakage from other channels. It then uses these combined estimates to suppress interference, including non-stationary components.
- The frequency-domain post-filter extends MMSE estimation to suppress interference remaining after geometric source separation.It treats transient corruption as leakage from other channels during separation.
- For each channel, stationary and transient interference components are combined into a single noise estimator.The stationary component is estimated with MCRA, while leakage represents residual source interference.
- The filter assumes non-background interferences are localized sources and that inter-channel leakage is constant.Leakage may arise from reverberation, localization error, microphone-response differences, and near-field effects.
- The estimated noise variance combines stationary noise and source leakage before computing the suppression gain.The gain multiplies the separated output to produce a cleaned signal spectrum.
- The leakage estimate assumes the separation algorithm reduces interference from other sources by typically −10 dB to −5 dB.The estimate uses a smoothed spectrum of the separated source.
B. Suppression rule in the presence of speech
The suppression rule estimates speech amplitude in the loudness domain using an MMSE framework. Its gain calculation incorporates posterior and prior SNR estimates while accounting for uncertainty about speech presence.
- The suppression rule uses MMSE estimation of spectral amplitude in the loudness domain, |X(k)|^1/2.The loudness domain was chosen because it produced better results under speech-presence uncertainty than spectral- or log-spectral-amplitude domains.
- The loudness-domain estimator uses α = 1/2 and a spectral gain conditioned on speech presence.The gain under speech presence is denoted GH1(k).
- The gain formulation uses the confluent hypergeometric function together with a posteriori and a priori SNRs.γ(k) is the a posteriori SNR, ξ(k) is the a priori SNR, and υ(k) combines them.
- The a priori SNR is estimated recursively using modifications that account for uncertainty about speech presence.This recursion follows the modifications proposed in.
C. Optimal gain modification under speech presence uncertainty
The gain is modified by weighting speech-present and speech-absent hypotheses according to the estimated probability of speech presence. Allowing zero minimum gain enables stronger attenuation when speech is absent.
- The loudness-domain estimator combines speech-present and speech-absent conditional estimates using the probability of speech presence.The two hypotheses are denoted H1 and H0.
- The optimal gain is modified according to the speech-presence probability and a minimum allowed gain.The gain under speech presence is GH1(k).
- Setting Gmin = 0 removes an arbitrary attenuation limit and lets the gain tend toward zero when speech is certainly absent.This is especially relevant when the interference is also speech, because residual babble noise produces musical noise.
- The speech-presence probability is computed from an a priori speech-presence probability defined for each frequency.The supplied passages identify the probability calculation and its a priori component but do not provide the complete expressions.
D. Initialisation
When a new source appears, post-filter state variables are initialized, mostly to zero. The initial stationary-noise estimate is handled separately because MCRA needs several seconds to produce a source-specific estimate.
- Most post-filter state variables for a new source can safely be initialized to zero.The stationary-noise state λstat.m(k, ℓ0) is the exception.
- MCRA requires several seconds to produce its first source-specific noise estimate, so an interim background estimate is needed.The interim estimate uses microphone noise estimates under delay-and-sum initialization of the weights.
- The microphone-level noise estimate xn(k) represents the noise estimate for microphone n.
V. RESULTS
The evaluation compares microphone inputs, beamforming, GSS, and post-filter configurations using SNR, LSD, and qualitative inspection. The proposed multi-channel post-filter improves separation over plain GSS and single-channel filtering, including under non-stationary speech interference.
- Evaluation setup: The evaluation uses SNR and log spectral distortion (LSD) to compare separated outputs.The experiments compare unprocessed inputs, delay-and-sum beamforming, GSS, and two post-filter configurations.
- Quantitative comparisons: The proposed multi-channel post-filter significantly improves results over both plain GSS and the single-channel post-filter.The compared configurations include unprocessed microphone inputs, delay-and-sum, GSS, single-channel post-filtering, and multi-channel post-filtering.
- Qualitative results: The post-filter removes most interference without excessive distortion despite non-stationary interference sharing the signal’s frequency content.The observation concerns separation of the first female source and is supported by signal-amplitude and spectrogram inspection.
- Qualitative results: The separated speech receives positive subjective assessments for both quality and intelligibility.The paper reports informal subjective evaluation alongside visual inspection of the separated signals.
VI. CONCLUSION
The paper presents a microphone-array separator combining simplified GSS with a multi-channel post-filter for simultaneous sound sources. Experiments report lower LSD and higher SNR, while future work targets speech-recognition optimization and automatic leakage adaptation.
- Conclusion: The system combines a microphone-array linear separator with a frequency-domain multi-channel post-filter for simultaneous sound sources.The separator uses a simplified GSS approach, while the post-filter estimates stationary noise and GSS leakage.
- Conclusion: The post-filter is sufficiently general to supplement most linear source-separation algorithms.Its noise estimate combines stationary noise with estimated leakage from geometric source separation.
- Conclusion: 11 dB maximum LSD reduction and 14dB SNR increase are reported versus noisy signal inputs.Preliminary perceptive tests and spectrogram inspection found the introduced distortions acceptable to most listeners.
- Future work: Future work includes optimizing separation directly for speech-recognition accuracy and automatically adapting the GSS leakage coefficient η.These are proposed next steps rather than evaluated capabilities of the reported system.