Source-linked AI summary

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

Kai Li, Jintao Cheng, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu

arXiv:2601.22599v2cs.SDcs.HC

TL;DR

Query-based universal sound separation remains vulnerable to residual interference because in-the-wild datasets contain weak labels and co-occurring events. The paper introduces a pipeline that mines high-purity single-event segments and synthesizes semantically consistent mixtures into Hive. Hive-trained models achieve competitive separation quality against SAM-Audio at ∼0.2% of its data scale, while the authors note real-world variability and model-bias boundaries.

  • Problem

    Query-based USS suffers residual interference because weak labels and event co-occurrence in in-the-wild datasets misalign supervision.

  • Method

    The paper mines high-purity single-event segments and synthesizes mixtures using a semantically consistent pipeline with ontology reconstruction and alignment.

  • Results

    ∼0.2% of SAM-Audio’s data scale: Hive-trained models achieved competitive separation accuracy and perceptual quality against the million-hour-scale baseline.

  • Takeaways & Limitations

    Prioritizing purity of supervised signals offers a data-efficient alternative to brute-force scaling for query-based USS training.

  • Takeaways & Limitations

    Synthetic mixtures may not fully reflect real-world acoustic variability, and multimodal-model relabeling and compatibility estimates may propagate biases.

Abstract

from arXiv · show

Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex acoustic scenes. This performance limitation stems largely from a data bottleneck: in-the-wild datasets contain weak labels and severe co-occurrence of events. These flaws induce models to learn spurious correlations between background noise and target categories instead of robust acoustic features. To address this, we propose an automated pipeline that eliminates co-occurrence of events by mining high-purity single-event segments from in-the-wild datasets via a semantically consistent synthesis protocol. Utilizing this pipeline, we constructed Hive, a high-quality synthetic dataset comprising 2.4k hours of raw audio. Experimental results demonstrate that, compared with the state-of-the-art model SAM-Audio which was trained on a huge dataset $\sim$500 times larger than Hive, certain open-source models trained on Hive achieve competitive separation accuracy and perceptual quality. Moreover, these models exhibited remarkable zero-shot generalization on out-of-distribution evaluation benchmarks. These findings highlight that prioritizing purity of supervised signals enables significant data efficiency, offering a new paradigm for training robust auditory foundation models with reduced computational costs. Code and dataset are available at https://cslikai.cn/Hive.

1. Introduction

Query-based universal sound separation targets arbitrary sources in complex mixtures, but residual interference persists because weak labels and event co-occurrence misalign supervision. The paper addresses this data bottleneck with high-purity synthesis and reports competitive performance using far less data than million-hour-scale training.

  • Query-based universal sound separation isolates arbitrary environmental and mechanical sources from mixtures for machine-hearing and audio-editing applications.
  • Residual interference persists across query-based USS methods as background noise or concurrent events leak into separated outputs.
  • Weak labels and frequent event co-occurrence in in-the-wild datasets create label-signal misalignment, such as rain recordings containing wind or traffic.
  • Scaling dataset size and model capacity delivers performance but imposes prohibitive computational costs that limit reproducibility and accessibility.
  • ∼0.2% of SAM-Audio’s data scale: Hive-trained AudioSep and FlowSep achieved competitive perceptual quality with the million-hour-scale baseline.
  • The authors attribute the result to prioritizing supervised-signal purity as a data-efficient alternative to brute-force scaling.

2. Related Work

Query-based USS extends source separation beyond fixed-output blind separation by using prompts, yet weakly labeled in-the-wild data still causes interference. The paper therefore emphasizes event-aligned, high-quality training data alongside architectural advances.

  • Blind source separation is limited by fixed source counts and the permutation problem, motivating query-based USS with auxiliary prompts for arbitrary targets.
  • Query-Based Universal Sound Separation Methods: Discriminative USS methods estimate target signals directly, often through time-frequency masking, and include LASS-Net, AudioSep, and CLIPSep.
  • Query-Based Universal Sound Separation Methods: Generative USS formulates separation as conditional synthesis, including flow-matching FlowSep and training-free methods that use pretrained priors.
  • Query-Based Universal Sound Separation Methods: Weakly labeled in-the-wild training data still produces severe interference because label-signal misalignment remains unresolved.
  • In-the-wild datasets contain co-occurring events that associate background artifacts with target categories, limiting separation purity.
  • Naive random mixing ignores semantic consistency, while existing resources such as Scaper presuppose curated isolated recordings rather than solving segment-level purification.

3. Single-Event Data Collection Pipeline

The data-collection pipeline converts heterogeneous in-the-wild audio into semantically aligned, acoustically singular sources. It combines taxonomy reconstruction, polyphony filtering, coarse-to-fine labeling, and sampling-rate standardization.

  • The pipeline has three stages: ontology reconstruction and preprocessing, single-event semantic-acoustic alignment, and super-resolution-based standardization.
  • AudioSet’s taxonomy is reconstructed by merging overlapping labels and pruning background textures, producing a separation-oriented space of 283 leaf nodes.
  • Audio is segmented with 10-second windows and 5-second overlap; segments below RMS energy 5 × 10^-4 and short residual tails are discarded.
  • A cascaded alignment framework targets acoustic singularity and precise mapping to mutually exclusive ontology leaves.
  • Metadata filtering and a Qwen3-Omni polyphony detector reject multi-event or interfering segments before semantic refinement.
  • Coarse-to-fine classification uses a 0.7 confidence threshold for parent nodes before Qwen3-Omni assigns restricted candidate leaf labels.
  • Source Curation and Properties: ∼0.9M unique clips totaling ∼2,442 hours were aggregated from 12 public datasets after cleaning and alignment.

4. Dataset Construction

Hive synthesizes training mixtures from purified single-event segments while constraining event combinations for semantic plausibility. The resulting dataset provides millions of mixtures and a substantial unmixed evaluation pool.

  • Naive random mixing can create implausible combinations and incorrect contextual priors, motivating a semantic compatibility protocol.
  • Semantically Consistent Mixing Strategy: The binary matrix M encodes semantic validity, and mixtures sample 2–5 sources while adding events compatible with all previously selected events.
  • Semantically Consistent Mixing Strategy: Valid mixtures undergo duration normalization and energy unification before additive superposition with randomly sampled SNRs from [−5, 5] dB.
  • 19.6 million mixtures (∼22.4k hours) are partitioned into 17.5M training, 1.75M validation, and 350k test samples.
  • Hive reserves 292 hours of distinct unmixed sources for validation and testing, exceeding the evaluation-set durations reported for AudioSep and SAM-Audio.

5. Experiments Settings

The experiments evaluate Hive-trained and baseline separation models using zero-shot benchmarks, standard quality metrics, dataset statistics, and controlled training settings.

  • Datasets and Benchmarks: Original pretrained checkpoints were tested on Hive without fine-tuning, while AudioSep and FlowSep were trained from scratch on Hive for out-of-domain comparisons.The evaluation includes zero-shot generalization and assesses Hive as a training resource.
  • Baselines and Metrics: Models span discriminative and generative query-based separation paradigms, including AudioSep, FlowSep, and SAM-Audio.The comparison covers nine listed baselines across both model families.
  • Dataset Statistics: Hive label statistics are reported through an overall word cloud, the ten most frequent labels, and the ten least frequent labels.Word-cloud token size is proportional to mixture count.
  • Baselines and Metrics: Performance is measured through signal-fidelity, perceptual-semantic, and reference-free quality metrics.Metrics include SDR, SI-SDR, FAD, LPAPS, CLAP similarity, MUSHRA, and SAM-Audio Judge scores.
  • Implementation Details: Experiments use PyTorch on 8× NVIDIA A100 GPUs and preserve the original model architectures and hyperparameter settings for fair comparison.AudioSep and FlowSep are trained for approximately 3M steps, with outputs resampled to 44.1 kHz during evaluation.

6. Results

Results show that Hive’s purified and semantically consistent supervision improves separation, efficiency, shortcut robustness, scaling behavior, and out-of-distribution performance.

  • Label Purity Validation: 98.0% agreement with human consensus was achieved by Qwen3-Omni on the 4-AFC relabeling task, exceeding Gemini 3.1 Pro at 95.0% and GPT-Audio at 90.0%.The human consensus used 67 valid participants, with high inter-rater reliability (Fleiss’ κ = 0.843).
  • Semantic Consistency Ablation: 1.0 dB SDR improvement for AudioSep resulted from enforcing semantic consistency over random mixing, with additional perceptual and reference-free gains for both models.Both regimes used 175k Hive mixtures and the same purified source pool.
  • Hive Test Set: 68.4 MUSHRA for AudioSep (Hive) exceeded SAM-Audio at 62.6 and AudioSep (Orig.) at 60.9, while FlowSep (Hive) improved over FlowSep (Orig.) by 7.1 points.AudioSep also achieved the strongest prior-system SDR at 2.37 dB; performance declined as mixtures increased from two to five sources.
  • Efficiency Analysis: 0.02 s GPU time and approximately 1.1 GB memory for AudioSep contrasted with SAM-Audio’s 8.2 B parameters and approximately 32 GB memory.Hive-trained AudioSep narrows the perceptual-quality gap against SAM-Audio without changing AudioSep’s efficiency profile.
  • Third-Party Test Sets: 2.29 dB USS-Bench SDR versus −1.86 dB was achieved by AudioSep (Hive) using 2.4k hours, roughly one sixth of AudioSep’s original 14.1k-hour training data.AudioSep (Hive) also improved USS-Bench OQ from 2.97 to 3.56 and MUSDB18-HQ SDR from −1.01 to 1.36 dB, while VGGClean SDR declined despite higher OQ.
  • Controlled Shortcut Analysis: 0.39 dB was Hive’s shortcut gap versus 1.41 dB for the original AudioSet/VGGSound-trained AudioSep under paired, controlled evaluation.FlowSep’s OQ gap decreased from 0.23 to 0.10 and its CLAP-T gap from 0.04 to 0.02.
  • Training Data Scale: 0.85 to 0.80 FAD improvement accompanied increasing Hive training scale, while CLAP-A rose to 0.65.The authors report stable gains across signal-fidelity and perceptual-quality metrics without saturation or overfitting at the tested scale.
  • Training Data Scale: 4.96 dB SDR was achieved with 875k Hive samples, compared with 2.37 dB for original AudioSep trained on 14,100 hours of in-the-wild data.The comparison supports the paper’s claim that supervisory-signal purity can matter more than data volume for general separation.

7. Conclusions

The paper presents Hive as a high-fidelity synthetic dataset and a data-efficient alternative to brute-force scaling for query-based universal sound separation.

  • Conclusion: The proposed pipeline combines ontology reconstruction, multimodal semantic filtering, high-purity single-event mining, and semantically consistent mixture synthesis.The resulting Hive-trained models achieve competitive separation and perceptual quality with approximately 0.2% of SAM-Audio’s data volume.

Impact Statement

The paper contributes a data-cleaning and synthesis pipeline and Hive dataset for training query-based universal sound separation models with less data and compute. It also recognizes misuse, bias, and realism risks requiring safeguards and future evaluation.

  • The work combines a data-cleaning and synthesis pipeline with Hive to support robust open-domain separation using substantially less data and compute.The pipeline integrates purified sources and semantically consistent synthesis; Hive contains approximately 2.4k hours of raw audio.
  • The paper warns that separation systems may enable sensitive-speaker extraction, deceptive editing, or misuse of copyrighted content.It also identifies bias propagation through multimodal language models and mismatch between synthetic mixtures and real acoustic variability.
  • AudioSet label refinement merges synonyms and fine-grained categories while excluding abstract acoustic attributes to create a separation-oriented taxonomy.Examples include merging “Acoustic guitar” into “Guitar” and excluding labels describing environments or acoustic attributes.
  • The pipeline uses multimodal-model relabeling, majority voting, quality auditing, and coarse-to-fine semantic refinement to improve annotation and source purity.Qwen3-Omni performs repeated annotation and audits single-event clips, while AudioTag predictions constrain leaf-node classification.
  • The integrated source pool combines licensed datasets with complementary acoustic characteristics to construct a long-tailed space for universal sound separation.Sources include ClothoV2, Voicebank DEMAND, AVE, and FSD50K, alongside other public datasets.

E. Full Class Frequency Statistics of the Hive Dataset

Hive contains 19.6 million synthetic mixtures across 283 classes, with pronounced class imbalance concentrated in several frequent natural-sound and human-activity categories.

  • Hive comprises 19.6 million synthetic mixtures organized into training, validation, and test splits across 283 classes.The dataset is described as a large-scale benchmark with high semantic consistency and mixtures containing two to five sources.
  • Bird, Male speech, man speaking, and Boat are among the most frequent classes, with approximately 3.115 million, 2.205 million, and 1.523 million instances respectively.These head classes establish the primary supervision signal for model training.

F. Details of Out-of-Domain Datasets

The paper evaluates transfer across music, wide-domain speech-and-instrument, and in-the-wild separation benchmarks, using established baselines and a standardized human-agreement audit. These datasets differ acoustically and semantically from Hive, while some retain residual co-occurrence or target specific source configurations.

  • Out-of-domain benchmarks: Zero-shot evaluation uses MUSDB18-HQ, USS-Bench, and VGGClean eval without fine-tuning to test robustness beyond Hive’s training distribution.The benchmarks differ substantially in acoustic characteristics and semantic distributions from Hive.
  • MUSDB18-HQ: MUSDB18-HQ contains 150 stereo songs with four isolated stems—drums, bass, vocals, and others—for music separation evaluation.The benchmark uses uncompressed 44.1 kHz WAV audio and museval metrics including SDR.
  • USS-Bench: USS-Bench combines Mandarin speech and instrument audio into four-source mixtures, including both simulated and real-world recordings.It is designed to test adaptability in idealized and realistic acoustic settings.
  • VGGClean eval: VGGClean eval contains 1,000 mixtures formed from manually selected VGGSound clips at approximately 0 dB average SNR.The released split is evaluated using Hive transfer without additional fine-tuning, although residual co-occurring events may remain in source clips.
  • Baseline methods: LASS-Net separates queried targets by conditioning a ResUNet spectrogram-mask predictor on BERT-derived natural-language embeddings.Its query-based design supports arbitrary text queries beyond a predefined label set.
  • Baseline methods: AudioSep and FlowSep are used for Hive training comparisons as reproducibly trainable representatives of discriminative and generative separation paradigms.AudioSep is selected for signal fidelity, while FlowSep has a complete open training framework.

I. Details of Zero-Shot Results

Zero-shot performance deteriorates as mixtures become denser, with discriminative models losing signal fidelity and generative models losing semantic adherence. Hive-trained models are more robust in high-interference settings, while computational efficiency remains strongly dependent on model paradigm.

  • Performance across acoustic densities: Performance degrades significantly across baselines as concurrent sources increase, exposing failures in acoustically dense mixtures.Table A3 evaluates two- through five-source mixtures and isolates failure modes under text-only queries.
  • Discriminative baselines: AudioSep’s SDR drops from high performance in two-source mixtures to negative values in five-source settings.The sign inversion indicates difficulty suppressing interference in dense scenes.
  • Generative baselines: Generative baselines maintain lower FAD but lose CLAP-Text similarity in four- and five-source mixtures.This reflects a trade-off between acoustically plausible output and semantic fidelity under increasing mixture complexity.
  • Hive-trained models: Hive-trained AudioSep maintains positive signal fidelity in five-source scenarios, demonstrating greater robustness than the evaluated baselines.The authors attribute this resilience to reduced co-occurrence noise and label-signal misalignment in training data.
  • Computational efficiency: Discriminative models are substantially more efficient, whereas SAM-Audio requires 8.2 B parameters and approximately 32 GB of GPU memory.AudioSep and LASS-Net run within tens of milliseconds on GPU and below 1.1 GB of GPU memory.

K. Controlled Shortcut Experiment Details

The controlled experiment separates co-occurrence effects from mixture complexity by pairing matched mixtures that differ only in statistical relations between target and interferers. Across models and source counts, Hive training reduces shortcut reliance, while the analysis remains bounded by class coverage, label bias, and unexamined shortcut sources.

  • Experimental design: The paired experiment holds mixture complexity, target identity, and SNR vector fixed while varying co-occurrence relations.This design distinguishes gains from semantically consistent mixtures from gains caused by reduced reliance on co-occurrence shortcuts.
  • Co-occurrence measurement: PMI measures whether two events co-occur more or less often than expected from their marginal frequencies.Positive PMI indicates statistical dependence, near-zero PMI approximate independence, and negative PMI avoidance.
  • Interferer selection: Co-occurring interferers have high PMI with the target, whereas decorrelated interferers have PMI near zero under shared semantic-compatibility constraints.Candidate targets must have at least four eligible interferers in both pools.
  • Results: Across both AudioSep and FlowSep, Hive training consistently shrinks the performance gap between co-occurring and decorrelated conditions.The absolute reduction grows as mixtures become more complex, according to the per-source-count analyses.
  • Scope and limitations: The analysis excludes extremely rare tail events and inherits biases from AudioSet’s labeling, especially background under-labeling.It isolates statistical co-occurrence but does not rule out shortcuts from room acoustics or recording-channel artifacts.

L. Scaling Law Analysis

Hive supports different scaling behaviors across discriminative and generative separation models. AudioSep improves continuously in signal fidelity as clean data grows, while FlowSep gains perceptual quality early and requires greater scale for semantic consistency.

  • Discriminative scaling: AudioSep’s Overall SDR rises from 4.12 dB at 175k samples to 5.67 dB at 17.5M samples without saturation.In the 5-mix setting, SDR improves from −0.05 dB to 1.46 dB across those scales.
  • Discriminative scaling: At 875k samples, AudioSep reaches an Overall SDR of 4.96 dB, surpassing the official baseline trained on millions of in-the-wild samples.The result is presented as evidence that clean supervision can outweigh raw data volume for discriminative alignment.
  • Generative scaling: FlowSep’s LPAPS and FAD improve rapidly in the low-data regime but saturate early, with FAD stabilizing around 0.85 by 1.75M samples.Semantic adherence and overall quality show a separate critical-mass behavior.
  • Cross-paradigm interpretation: Hive supports efficient training for both paradigms through continuous signal-fidelity gains for discriminative models and usability-threshold gains for generative models.The authors distinguish these mechanisms as reflecting different data requirements for signal estimation and conditional synthesis.
Loading 2601.22599v2…