Source-linked AI summary

Under-determined reverberant audio source separation using a full-rank spatial covariance model

Ngoc Duong, Emmanuel Vincent, Remi Gribonval

arXiv:0912.0171v2stat.ML

TL;DR

The paper addresses under-determined convolutive source separation in reverberant environments, where narrowband models inadequately represent spatial spread. It models source images with Gaussian spatial covariances, estimates model parameters using EM with initialization and frequency alignment procedures, and reports improved separation from the full-rank unconstrained model in realistic reverberation.

  • Problem

    Under-determined convolutive separation is difficult in reverberant environments because reverberation creates spatial spread and the narrowband approximation may fail.

  • Method

    The paper models each source image as a zero-mean Gaussian variable with spatial covariance, considers four covariance models, and estimates their parameters using EM with initialization and frequency-wise source alignment.

  • Results

    The full-rank unconstrained model improves separation over rank-1 models and state-of-the-art algorithms in realistic reverberant environments.

  • Takeaways & Limitations

    Full-rank spatial covariance matrices better account for reverberation and support more effective separation than rank-1 models in the reported experiments.

  • Takeaways & Limitations

    The rank-1 anechoic and full-rank direct+diffuse models receive no detailed EM derivation because they appeared to provide lower performance in a semi-blind comparison.

Abstract

from arXiv · show

This article addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial characteristics of the source. We then consider four specific covariance models, including a full-rank unconstrained model. We derive a family of iterative expectationmaximization (EM) algorithms to estimate the parameters of each model and propose suitable procedures to initialize the parameters and to align the order of the estimated sources across all frequency bins based on their estimated directions of arrival (DOA). Experimental results over reverberant synthetic mixtures and live recordings of speech data show the effectiveness of the proposed approach.

Under-determined reverberant audio source separation using a full-rank spatial covariance

This report was authored by Ngoc Q.K. Duong, Emmanuel Vincent, and Rémi Gribonval and issued as INRIA research report 7116 in December 2009.

  • The report contains 19 pages.
  • It is an INRIA Rennes–Bretagne Atlantique research report.

S´eparation de m´elanges audio r´everb´erants sous-d´etermins l’aide d’un mod`ele de covariance

The paper concerns reverberant audio source separation using a full-rank spatial covariance model, with convolutive separation and EM among its keywords.

  • The paper addresses separation of reverberant audio mixtures using a full-rank spatial covariance model.
  • Its approach models reverberant recording environments in the context of under-determined source separation.
  • The listed keywords include convolutive source separation, under-determined mixtures, spatial covariance models, EM, and the permutation problem.

1 Introduction

The introduction frames under-determined reverberant separation as a limitation of narrowband and sparse time-frequency methods, then motivates full-rank covariance modeling and blind parameter estimation.

  • Convolutive mixing models each source’s spatial image through an acoustic filter vector linking the source to all microphones.
  • Blind source separation recovers source signals or spatial images from multichannel mixtures, including under-determined cases where I < J.
  • Most time-frequency methods use a narrowband approximation and exploit source sparsity, including binary masking and ℓ1-norm minimization.
  • Reverberation limits these methods because long mixing filters violate the narrowband approximation and produce spatial spread.
  • The proposed framework models source images with phase-invariant multivariate distributions whose covariances encode spatial position and spread.
  • The article addresses blind parameter initialization and frequency-wise source permutation while evaluating the models against state-of-the-art techniques.

2 General framework and spatial covariance models

The framework models source spatial images as zero-mean Gaussian variables whose time-varying variances and frequency-dependent spatial covariance matrices capture spectro-temporal power, position, and spatial spread. It defines rank-1 and full-rank covariance models for reverberant mixtures, estimates parameters by maximum likelihood, and reconstructs spatial images with multichannel Wiener filtering.

  • General probabilistic framework: Each source spatial image is modeled as a zero-mean Gaussian vector, while its covariance factors into a time-varying variance and a spatial covariance matrix.The variance encodes spectro-temporal source power, whereas the spatial covariance encodes spatial position and spread.
  • General probabilistic framework: Under uncorrelated sources, the mixture STFT vector is also zero-mean Gaussian, with covariance determined by the source variances and spatial covariance matrices.The likelihood of the observed mixture STFT coefficients is consequently parameterized by these variance and spatial covariance sets.
  • General probabilistic framework: Source separation first estimates variance and spatial parameters by maximum likelihood, then obtains all source spatial images by MMSE multichannel Wiener filtering.This separates parameter estimation from spatial-image reconstruction.
  • Rank-1 convolutive model: The narrowband convolutive model yields rank-1 spatial covariance matrices based on the frequency-domain mixing vectors, with the anechoic case parameterized by propagation delays and gains.The rank-1 model uses the Fourier-transformed mixing filters, while anechoic filters reduce to delay-and-gain combinations determined by source–microphone distances.
  • Full-rank direct+diffuse model: Reverberation invalidates the single-position interpretation of narrowband mixing because echoes create spatial spread and full-rank source covariance matrices.The direct-plus-reverberant model treats the spatial image as two uncorrelated parts and sums their covariances.
  • Full-rank direct+diffuse model: The direct-plus-diffuse model assumes equal reverberant power across microphones with correlations determined by microphone distances, but diffuse reverberation is rarely satisfied in practice.Early echoes can be spatially concentrated; the diffuse covariance relation was observed to hold on average across many source positions, not generally for each source independently.
  • Full-rank unconstrained model: The unconstrained full-rank model leaves covariance coefficients unrelated a priori, providing more flexible mixing-process modeling than the structured full-rank alternatives.The paper presents this flexibility as potentially improving separation performance for real-world convolutive mixtures.

3 Blind estimation of the model parameters

The paper estimates model parameters from mixtures using EM, with hierarchical clustering for initialization and DOA-based alignment across frequency bins. It also derives updates for several covariance models, including full-rank unconstrained separation.

  • Parameter estimation: EM estimates model parameters from the observed mixture, replacing an earlier quasi-Newton approach because it converges faster in practice despite requiring more iterations.The procedure begins by initializing model parameters, then applies EM-based updates.
  • Initialization: Hierarchical clustering initializes mixing vectors and spatial covariance matrices by grouping normalized mixture STFT coefficients within each frequency bin.Clusters are linked bottom-up until a threshold K, after which the J largest clusters are retained.
  • Initialization: The proposed initialization uses average distances between phase-normalized coefficients and improves initial mixing-parameter estimates over alternatives tested in the experiments.Random and DOA-based initialization produced slower convergence and poorer separation performance.
  • Rank-1 convolutive model: For the rank-1 convolutive model, EM uses a noisy mixture with stationary spatially uncorrelated Gaussian noise and alternates Wiener-filter inference with parameter updates.The E-step computes conditional source statistics; the M-step updates source variances, the mixing matrix, and noise covariance.
  • Full-rank unconstrained model: For the full-rank unconstrained model, EM operates directly on source spatial images and updates each source variance and spatial covariance independently by frequency bin.Its derivation avoids the fixed-initial-mixing-vector issue encountered in the rank-1 formulation.
  • Permutation alignment: DOA-based permutation alignment orders independently estimated frequency-bin parameters, using normalized mixing vectors or PCA components of full-rank spatial covariances.In a three-source stereo recording with RT60 = 250 ms, source order was globally aligned for most frequency bins after alignment.

4 Experimental evaluation

The experiments compare spatial covariance models and separation methods across semi-blind, blind, real-world, and moving-source conditions. The full-rank unconstrained model performs especially well under reverberation and small source movements.

  • The evaluation compares four covariance models in semi-blind mixtures, then studies the rank-1 convolutive and full-rank unconstrained models in blind synthetic and real-world mixtures.Performance is measured with SDR, SIR, SAR, and ISR.
  • 1.8 dB, 2.3 dB, and 2.9 dB are the full-rank unconstrained model’s SDR improvements over the rank-1 convolutive model, binary masking, and ℓ1-norm minimization, respectively.The rank-1 anechoic model performs worst because it accounts only for the direct path; the full-rank direct+diffuse model decreases SDR by 0.6 dB relative to the rank-1 convolutive model.
  • At T60 = 50 ms, the rank-1 convolutive model provides the best SDR and SAR because the direct path contains most received energy.
  • At T60 ≥130 ms, the full-rank unconstrained model outperforms the rank-1 model and binary masking in SDR and SAR, while its SIR is close to binary masking.At T60 = 500 ms, its SDR is 2.0 dB, 1.2 dB, and 2.3 dB higher than the rank-1 convolutive model, binary masking, and ℓ1-norm minimization, respectively.
  • 0.1 dB and 0.5 dB are the full-rank unconstrained algorithm’s SDR improvements over the best reported algorithms for three- and four-source SiSEC 2008 mixtures, respectively.
  • With a 5° source rotation, SDR drops 0.6 dB for the full-rank unconstrained model and 1 dB for the rank-1 convolutive model.The full-rank model accounts for both spatial spread and spatial direction, making small movements within the spread less disruptive.

5 Conclusion and discussion

The article presents a spatial-covariance framework with four models and efficient parameter estimation for convolutive source separation. Experiments indicate that the full-rank unconstrained model better accounts for reverberation and improves separation in realistic environments, while several extensions remain future work.

  • The framework models convolutive source separation through spatial covariance matrices and includes four specific covariance models.
  • Full-rank models overcome the narrowband approximation used by rank-1 models.
  • Future work includes modeling diffuse or semi-diffuse sources, improving small-data parameter robustness, addressing permutation probabilistically, and combining spatial covariance with source-spectrum models.
Loading 0912.0171v2…