Source-linked AI summary

Geometric Ceilings on Time-Frequency Masking for Single-Channel Separation

Maxime Baelde

arXiv:2609.03481v1eess.SPcs.SDeess.AS

TL;DR

The paper asks what real-gain masking forbids and whether prior modelling can produce estimates outside that constraint. It derives geometric and operator-class ceilings, then evaluates a non-circular Gaussian-mixture posterior mean. The estimator leaves the class but remains below the per-frame and fixed-class ceilings, with the supplied limitation evidence highlighting multi-source partition constraints and an overparameterized pilot.

  • Problem

    Existing oracle masks do not bound the full real-gain class, and prior work does not connect prior assumptions to membership in that class.

  • Method

    The paper derives the real-gain optimum and residual, organizes estimators into four nested real-linear operator classes, and evaluates a closed-form Gaussian-mixture posterior mean.

  • Results

    The fitted separator remains 9.4 to 12.0 dB under held-out measured ceilings, with no crossing across 189 per-excerpt figures.

  • Takeaways & Limitations

    A phase-symmetric posterior returns the MMSE estimate to the real-gain line, while non-zero means, non-circularity, and coupling provide the prior freedoms associated with leaving it.

  • Takeaways & Limitations

    Bounded multi-source outputs can break mixture partitioning, and an overparameterized covariance fit can reproduce its training excerpt rather than model it.

Abstract

from arXiv · show

Most single-channel separators estimate a source by applying a real gain to the mixture in each time-frequency bin. The optimum of that format, which the oracle masks used as bounds do not attain, is the orthogonal projection of the source onto the line spanned by the mixture, its residual set by the angle between them. Locating an estimator reduces to the block structure of a real-linear operator on stacked spectra, giving a chain of four nested classes whose three larger terms match three assumptions on the prior: zero means, circularity and absence of inter-frequency coupling. Held fixed the chain is a cascade of four orthogonal projections; refitted per frame it collapses onto its first term, attributing the whole residual to one missing real parameter per bin, the phase. When the phase posterior is symmetric about the mixture direction, the minimum mean-square estimate falls back onto the line, with gain the posterior mean of the oracle gain and excess error its variance. On MUSDB18 a posterior mean under a non-circular Gaussian-mixture prior leaves the class yet stays 11.44 dB under the per-frame ceiling, which four times as many components and 7.5x the data do not close; a closed-form gate attributes some 70% of it, in decibels, to the predicted variance. The widest fixed class stays 6.70 dB under the same ceiling. Leaving the class and minimising squared error are conflicting requests: the barrier lies in the criterion rather than in the prior.

Highlights

The paper identifies phase symmetry as the condition that returns the MMSE estimate to the real line, while multi-source mask partitions can fail under range constraints.

  • Multi-source masks can lose their partition when the mask range is bounded.
  • A symmetric phase posterior pulls the MMSE estimate back onto the real line.

1. Introduction

Real-gain masking is geometrically constrained because every estimate lies on the mixture line, leaving an angle-dependent residual. The paper characterizes this constraint with nested operator classes and tests whether a prior can produce estimates outside it.

  • Real-gain masking estimates each time-frequency source value as a non-negative scalar multiple of the mixture, preserving mixture phase.
  • The best real-gain estimate is the orthogonal projection of the source onto the mixture line, leaving a residual determined by their angle.
  • Oracle ratio and Wiener masks are members of the real-gain class but can remain several decibels below its true optimum.
  • Four nested real-linear operator classes locate estimators by progressively allowing similitudes, arbitrary planar maps, and inter-frequency dependence.
  • A non-circular Gaussian-mixture prior provides a closed-form posterior mean that can leave the real-gain class, while phase symmetry returns it to the line.

2. Related Work

Prior work supplies masking and operator constructions but does not establish a class-wide ceiling or connect prior assumptions to class membership. This paper frames that missing connection through non-zero means, non-circularity, and inter-frequency coupling.

  • The phase-sensitive mask is the constrained minimizer of squared reconstruction error over real gains, providing the class ceiling used here.
  • Existing oracle-mask benchmarks bound tried candidates rather than the full masking class, so beating an oracle does not establish that the format was exceeded.
  • Open-Unmix belongs to the bounded-mask subclass because it reconstructs magnitude estimates with the mixture phase.
  • Complex masking and widely linear estimation add operator degrees of freedom, but prior work does not ask whether they pass the exact real-gain ceiling.
  • Gaussian-mixture separation uses closed-form conditional expectations, but its conventional priors are zero-mean, circular, and bin-wise, leaving phase modelling unaddressed.
  • The paper identifies non-zero means, non-circularity, and inter-frequency coupling as the three prior properties linked to leaving the real-gain class.

3. Problem Statement and Geometry

The paper treats real-gain separation as a geometric constraint: each bin’s estimate lies on the mixture line, and the orthogonal residual is an irreducible ceiling determined by source–mixture angle. It then places broader estimators in nested real-linear classes and distinguishes the resulting per-frame and constrained subclass limits.

  • Per-bin geometry: A real-gain estimator outputs each bin on the line ℝx, so its optimum is the orthogonal projection of s onto that line.The residual is the source component orthogonal to the mixture.
  • Per-bin geometry: The irreducible residual has squared magnitude |s|^2 sin^2 θ and depends on the data, not on training, parameters, or compute.This makes the ceiling a property of the estimator class rather than a particular construction.
  • Phase dependence: The residual vanishes for in-phase or antiphase source and interference, but a matching-amplitude bin can lose the source entirely at antiphase.The maximum phase-dependent residual occurs at cos Δφ = −min(r, 1/r), reaching one when |n| = |s|.
  • Subclass ceilings: The widest real-gain class has nested bounded and non-negative subclasses whose constrained optima are projections of the unconstrained gain onto their admissible intervals.Non-negative masks lose bins with phase disagreement beyond π/2; masks bounded by one clip gains above one and incur the interference energy.
  • Multi-source constraints: For multiple sources, unconstrained optimal gains partition the mixture, but output restrictions can break that partition; the reported experiments use two sources and leave more-than-two-source hierarchical priors open.Non-negative gains already break the partition with two sources, while bounded gains can break it with three.
  • Operator hierarchy: Real-linear operators form a hierarchy of four nested classes, and each refinement adds directions or cross-bin dependence while lowering the residual.The estimator’s fitted operator identifies the smallest containing class through its block structure.

4. Signal Model and Proposed Prior

The paper constructs a posterior-mean separator from Gaussian-mixture priors whose covariance structure determines its operator class. The model combines source priors exactly, yielding a closed-form responsibility-weighted estimator while exposing data and runtime constraints.

  • Prior and estimator: The estimator is a posterior mean from a Gaussian-mixture prior fitted independently on isolated sources, so its operator class follows from the prior.The construction is generative rather than an output-stage design choice.
  • Prior and estimator: Nonzero means, non-circular covariances, and full covariance across frequencies relax the three phase-invariant prior assumptions.Cross-frequency covariance can encode partials from the same harmonic series rising and falling together.
  • Model capacity and data: The covariance structure selects the operator class, but full covariance requires d(d + 1)/2 parameters per component versus d for diagonal covariance.With d = 2F, shorter windows improve the training-frames-per-dimension ratio but alter the representation and ceiling.
  • Source combination: Independent Gaussian sources produce a Cartesian-product Gaussian mixture whose component weights multiply, while means and covariances add.For two sources the combined prior has K1K2 terms, and the exact closure extends hierarchically to subsets of sources.
  • Posterior construction: Conditioning the mixture on the observation yields pairwise Wiener estimates combined by responsibilities, with posterior covariances independent of the observation.The posterior mean depends on the observation through both the component estimates and their weights.
  • Scope and implementation: The unconstrained mean and covariance define a learned phase prior without an attached signal model, but the estimator is not competitive on runtime.Cholesky factors are reused across frames, yet the stated complexity remains O(K1K2d^3) once and O(K1K2d^2) per frame.

5. Theoretical Analysis

The analysis locates estimators within nested real-linear operator classes and shows that squared-error estimation can pull even an operator outside the real-gain class back onto the mixture line. It also identifies posterior asymmetry, prior support, and model-family realism as boundaries on when leaving the class yields an advantage.

  • 5.1. The Class Containing the Estimator: A responsibility-weighted mixture of Wiener masks remains in the bounded real-gain class and cannot exceed SDR⋆.The estimator is a mixture of Wiener-mask terms with observation-dependent weights.
  • 5.1. The Class Containing the Estimator: Four nested real-linear operator classes correspond to assumptions about prior means, circularity, and inter-frequency independence.Each reinstated assumption removes one step of the operator-class chain.
  • 5.1. The Class Containing the Estimator: Outside the class is necessary but insufficient for crossing the ceiling, because an estimator can leave the class yet perform arbitrarily poorly.The ceiling compares an estimator using a fitted prior with an oracle computed from the true source.
  • 5.2. Pull of the Squared-Error Criterion: When symmetry holds, the minimum mean-square estimate is a real gain whose value is the posterior mean of the oracle gain, with shortfall governed by posterior variance.This makes leaving the operator class and minimizing squared error competing demands.
  • 5.2. Pull of the Squared-Error Criterion: A posterior mean leaves the real-gain class only when the phase posterior is asymmetric about the mixture direction.Symmetry cancels the odd phase moment and forces the conditional mean onto the mixture line.
  • 5.3. Behaviour Away From the Training Support: Away from training support, degradation is attributed to changing component responsibilities rather than the regression term, with errors expected to break where the dominant component switches.The predicted break under continuous deformation provides a falsifiable test of the responsibility-based explanation.
  • 5.4. The Exact Posterior as a Reference for Samplers: The exact posterior provides sampler benchmarks for mean bias, covariance calibration, and pointwise posterior scoring without Monte Carlo error on the reference side.Its usefulness is limited by the realism of the Gaussian-mixture prior.
  • 5.5. Invariance to the Component Law: Richer component laws preserve the operator-class and symmetry results while potentially changing the deficit through a smaller posterior variance.The magnitude of that improvement remains empirical rather than settled by the propositions.

6. Experimental Conditions

The experiments measure deficits to a resynthesized, paired ceiling while varying evaluation material, frame length, covariance structure, model capacity, and training volume. The protocol also verifies how the time-frequency ceiling transports through resynthesis and controls edge, window, and silent-excerpt effects.

  • Dataset and configurations: The study evaluates posterior-mean estimators on MUSDB18 vocals and accompaniment using paired excerpts and two training-volume caps.The corpus includes 100 training and 50 test tracks; configurations vary material, frame length, and covariance structure.
  • Model configurations: Diagonal covariance places the main estimator in M3, while full covariance at L=256 tests M4 by adding inter-frequency coupling.The full-covariance comparison holds components, seed, iterations, floor, tracks, and frame cap fixed against the diagonal configuration.
  • Evaluation metric: The reported metric is time-domain SDR after resynthesis, measured without scale invariance and paired against the measured ceiling SDR⋆.The ceiling is measured after resynthesis rather than taken from its analytic time-frequency counterpart.
  • Transport through resynthesis: Resynthesis projects modified spectrograms onto the consistent subspace, so the measured ceiling is a lower bound on the true time-domain class ceiling.A crossing cannot be caused by transport, but a shortfall remains genuine only up to the transport slack.
  • Controls and exclusions: The protocol excludes silent-source excerpts above 60 dB and fixes the periodic window because window choice changes the fitted estimator substantially while barely changing the oracles.A symmetric-window refit moves the estimator by 1.8 dB while every oracle moves by at most 0.03 dB.

7. Results

Results compare fixed operator classes and posterior means against measured ceilings on MUSDB18. The largest fixed-class gain comes from inter-frequency coupling, while the fitted estimator remains substantially below its class ceiling.

  • Ceilings: 16.62 dB is the measured M1 ceiling at L=1024, compared with 15.49 dB for its analytic counterpart on nine held-out excerpts.The measured and analytic ceilings have the predicted ordering after resynthesis.
  • Mask subclasses: 1.25 dB for vocals and 1.26 dB for accompaniment is lost when the real mask is restricted to the bounded [0,1] subclass.The pooled subclass ceilings are 16.62, 15.98, and 15.37 dB for M1, M+, and M[0,1].
  • Fixed operator cascade: 2.68 dB is the corrected gain from M3 to M4 at L=1024, whereas the earlier cascade steps contribute only 0.0013 and 0.020 dB in the reported rows.At L=256, the corresponding M3-to-M4 gain is 1.14 dB.
  • Adaptivity versus class widening: 6.70 dB from frame adaptivity exceeds the 2.68 dB gain from widening to M4 on this material.The comparison is made against the same per-frame ceiling after correcting the M4 in-sample residual by a degrees-of-freedom heuristic.
  • Caveat: The corrected M4 result is heuristic because its raw residual is optimistic and the fit noise is neither Gaussian nor independent across frames.The correction is therefore reported as an order of magnitude of optimism rather than an unbiased corrected value.

7.3. Deficit Under the Ceiling and Effect of the Number of Components

The diagonal-covariance posterior mean remains far below its measured ceiling across component counts, training volumes, and configurations. Increasing components or data barely closes the gap, while posterior witnesses place the residual in predicted variance rather than mean displacement.

  • Deficit under the ceiling: 9.63 to 12.01 dB is the posterior mean’s deficit below the measured ceiling across 144 M3 measurements, with no zero crossings.The range spans three configurations, three component counts, and two training volumes.
  • Number of components: More than 50 further component doublings would be needed to close the 11.44 dB deficit at K=32 if the observed slope continued.The extrapolated component count is of order 10^19, so counting components does not reach the class ceiling.
  • Training volume: 7.5x more training data changes the deficit by only 0.02 to 0.19 dB across nine cells, with the larger volume moving toward the ceiling in eight.At K=32 held out, the two arms differ by 0.02 dB despite 19.5 versus 146.2 frames per dimension.
  • Frame length: The short-frame configuration has deficits of 9.63, 9.85, and 9.83 dB across K, without an ordering in K.The comparison across frame lengths is not paired because the ceiling and excerpt subset differ.
  • Mechanism of the deficit: A better-fitted posterior tightens around the real line, while the remaining squared error is consistent with posterior variance rather than off-line mean displacement.With 7.5x more data, phase rotation falls by 2.7x and out-of-class gains by 3.5x, while the deficit changes by 0.17 dB.

7.4. Effect of the Covariance Structure

Allowing full covariance increases apparent in-support capacity, but held-out performance is governed mainly by regime and does not benefit from leaving the real-gain class.

  • 8.53, 8.01 and 7.64 dB: the full-covariance deficit decreases monotonically with component count, while the diagonal fit decreases from 9.22 to 8.59 to 8.15 dB.Held-out full-covariance deficits are 9.54, 9.38 and 9.55 dB, versus 9.63, 9.85 and 9.83 dB for the diagonal fit.
  • 1.01, 1.37 and 1.91 dB: the generalisation gap grows with component count under full covariance, compared with 0.41, 1.26 and 1.68 dB under diagonal covariance.The covariance structure adds only 0.60, 0.11 and 0.23 dB, without a trend of its own.
  • At the training volume, none of the 12 cells crosses the exact class ceiling, and none of the 60 per-excerpt figures does either.A separate pilot crosses its ceiling because its covariance parameter count exceeds the available fitting frames.
  • 3.6 to 4.4 degrees: full covariance produces substantially larger median phase rotations than the diagonal fit’s 0.33 to 0.52 degrees.Full covariance also yields 3.4 to 4.4% of active bins above unit gain, versus 0.2 to 0.8% for the diagonal fit, without improving the criterion.
  • 24 to 26 degrees: a posterior draw leaves the class clearly, but loses 1.54, 1.64 and 1.66 dB across component counts.The loss is nearly component-count independent and is compatible with posterior variance dominating the excess term.
  • The posterior mean and hard-assignment read-outs agree within 0.06 dB and 0.14 degrees, but hard assignment leaves the class two to three times as often.Their agreement reflects near-degenerate per-frame responsibilities, while their different class departures expose averaging effects.

7.6. Share of the Deficit Attributable to Posterior Variance

The measured deficit above the geometric ceiling is separated from the ceiling residual and predicted from posterior variance, which accounts for roughly 70% in decibels but is not an energy decomposition.

  • The measured ceiling is the residual of the best real-gain estimate, whereas the tabulated deficit is the estimator’s additional loss above that residual.The additional loss is identified with |x|^2 Var(m⋆|x), a property of the fitted prior.
  • The closed-form prediction is available because the oracle gain is a linear functional of the stacked source at fixed mixture, making its posterior a scalar mixture.The variance decomposes into within-pair and between-pair contributions through the law of total variance.
  • 6.87/9.87 and 6.37/8.93: posterior variance accounts for 69.6% at K=8 and 71.4% at K=32 of the deficit in decibels.These are quotients of decibel deficits, not shares of error energy; the corresponding energy share at K=8 is 44.3%.
  • 96 to 98% of the predicted variance lies in the within-pair term, and no excerpt in 30 rows has a realised-to-predicted ratio below one.The lowest ratio is 1.42, while the model underestimates error especially in the lowest-predicted-variance quarter.
  • The gate does not bound other priors, while Gaussian scale mixtures preserve the symmetry result and make their own deficit a posterior variance.The scope is therefore a statement about the fitted prior and the in-class component of error, not a universal bound across priors or synthesis.
  • 0.91 to 0.97 dB: time-domain deficits exceed the corresponding spectral deficits on the same bins without retraining.The difference arises in reconstruction, with overlap-add plausibly reducing the ceiling residual’s off-line component.

7.7. Diagnostics of the Prior

The diagnostics identify weak non-circularity as the limiting prior property, while off-support errors are concentrated at sharp responsibility boundaries rather than distributed smoothly.

  • ρ tests phase structure, κ tests tails, the Hill estimator tests tail index, and γ tests shared latent-scale dependence between adjacent bins.Each statistic is interpreted against a synthetic null satisfying the corresponding assumption.
  • Positive γ indicates non-Gaussian dependence rather than merely heavy marginal tails, because the rank-Gaussianised statistic vanishes for any Gaussian copula.A common scale variable is the mechanism that produces the detected dependence.
  • 2.2 and 3.5: non-circularity exceeds its noise floor by these factors for vocals and accompaniment, respectively, but remains the weak point.The paper links this limited broken circularity to the small phase rotations of the posterior mean.
  • Figure 6 compares error increases along three spectral-tilt paths, marking arg-max responsibility switches against the between-switch trajectory.The plotted curves show slow hundredths-of-a-decibel movement between switches and roughly decibel-scale jumps at switches.
  • 209 switches across 36 paths: the median absolute error step is 1.00 dB at a switch versus 0.03 dB elsewhere, a pooled ratio of 30.The largest switch step is 8.50 dB, compared with a 0.35 dB 95th percentile away from switches.
  • 46 of 209 switches are small because their responsibility assignments are not sharp, with median maximum responsibility 0.85 versus 0.997 for the others.The break therefore depends on a sharp cell boundary, not merely on a changing arg-max label.

7.9. Position of the Estimator and Scope of the Measurement

The estimator remains far below the per-frame real-gain ceiling, and leaving that class under squared-error estimation worsens performance rather than closing the gap.

  • 11.44 dB: the estimator’s smallest held-out deficit is 5.18 dB below the ideal ratio mask and 8.50 dB below it on the same pooled comparison.At L=1024, the ideal ratio mask is 13.68 dB against a 16.62 dB ceiling, while the estimator is 11.44 dB below that ceiling.
  • 16.62 dB: the real-gain ceiling on nine retained held-out MUSDB18 excerpts at L=1024 is a class-level data statistic that no class member exceeds.Separate material and frame-length settings yield separate ceiling values rather than a single cross-setting comparison.
  • The supported measurement concerns both a class ceiling and a model’s position relative to it, which must not be conflated.The estimator can depart from the real line without approaching the class optimum.
  • At most 0.19 dB: the large arm’s deficit changes little with data volume despite providing 144 to 581 frames per dimension.The paper treats this as a reached plateau, while the pilot is sensitive to volume when data are scarce.
  • 1.6 dB: the read-out that visibly leaves the class is worse by a component-count-independent margin than the posterior mean.The paper attributes the conflict to squared-error estimation and the symmetry of the phase posterior, rather than to insufficient prior quality alone.

8. Conclusion

The paper establishes geometric ceilings for real-gain separation, explains when conditional means can leave that class, and measures the resulting gap on MUSDB18. It also provides an exact posterior construction and identifies limits and promising extensions of the evidence.

  • Conclusion: The real-gain ceiling is determined by angular disagreement between source and mixture, while conditional-mean membership depends jointly on the prior and read-out.A phase posterior symmetric about the mixture direction returns the estimate to the mixture line, regardless of the prior.
  • Conclusion: The spectral-to-time-domain transport is supported by a numerically checked inequality rather than a proof, and time-domain deficits are 0.91 to 0.97 dB above spectral counterparts.The difference is reported as attributable to synthesis but is not decomposed by an isolated experimental arm.
  • Conclusion: The exact posterior offers normalized densities, posterior means, observation-independent covariances, and pointwise log-densities for evaluating generative samplers without Monte Carlo error on the reference side.Its limitation is the measured realism of the prior.
  • Conclusion: Three proposed extensions target heavier tails, local inter-frequency coupling, and consecutive-frame phase advance, with the banded covariance intended to retain coupling while reducing overfitting.The diagnostics motivate these extensions in increasing order of the analysis's expected gains.

CRediT authorship contribution statement

Maxime Baelde is credited across the paper’s conceptual, methodological, analytical, implementation, data, writing, and visualization work.

  • Maxime Baelde contributed to conceptualization, methodology, formal analysis, software, investigation, data curation, writing, and visualization.

Data availability

The experiments use only the public MUSDB18 corpus, and the code for every figure and table is publicly available.

  • The experiments use the public MUSDB18 corpus and no other data.
  • Code producing every figure and table is public at the masking-ceiling repository, tag v1.0.0.
  • The measured separator is the thesis implementation imported from the generative-audio-source-models repository rather than reimplemented.
Loading 2609.03481v1…