Source-linked AI summary
CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
Zhiyuan Zhu, Han Wang, Wenxiang Guo, Yu Zhang, Changhao Pan, Rui Yang, Zhou Zhao
TL;DR
Spatial audio vocoders must convert multi-channel mel-spectrograms into waveforms without losing inter-channel cues, while meeting real-time constraints. CSAVocoder combines pose- and mel-conditioned spatial modeling with spatially informed adversarial training and a strictly causal stateful generator. Experiments report improved spatial fidelity over channel-wise baselines while maintaining strong audio quality and real-time performance.
Problem
Most neural vocoders target monaural audio, and direct spatial extensions can degrade spatial quality by ignoring inter-channel cues and real-time constraints.
Method
CSAVocoder is a causal GAN-based spatial vocoder that conditions on multi-channel mel-spectrograms and source-listener pose, uses spatial losses and discrimination, and supports stateful streaming.
Results
CSAVocoder outperforms channel-wise vocoder baselines in spatial fidelity while maintaining strong audio quality and real-time performance.
Takeaways & Limitations
Within the binaural and FOA settings studied, one unified architecture supports multiple spatial formats for immersive audio applications.
Takeaways & Limitations
The study focuses on binaural and FOA formats, leaving higher-order ambisonics, loudspeaker layouts, object-based audio, and personalized HRTF rendering for future work.
Abstract
from arXiv · showhide
Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We present CSAVocoder, a causal GAN-based spatial audio vocoder that jointly optimizes waveform fidelity and spatial rendering. Our framework introduces a Spatial Adaptor that fuses multi-channel mel-spectrograms with dynamic source-listener pose information, together with a spatial consistency discriminator that supervises inter-channel cues. To meet real-time requirements, we design a strictly causal, stateful generator that supports efficient streaming inference with constant memory overhead. Experiments on large-scale spatial audio datasets show that CSAVocoder improves spatial fidelity at competitive audio quality and real-time performance.
1 Introduction
Spatial audio vocoders must preserve inter-channel spatial cues while producing high-fidelity waveforms with low-latency causal streaming. CSAVocoder addresses these requirements by conditioning on multi-channel mel-spectrograms and source-listener pose.
- Spatial audio provides immersive three-dimensional rendering by modeling directional and distance cues used by human auditory localization.
- Most existing vocoders target single-channel audio, and direct spatial extensions can degrade quality by ignoring inter-channel cues.
- Relative source-listener pose affects spatial rendering, including the acoustic behavior associated with source position and motion.
- Real-time applications require low-latency spatial audio, making vocoder inference speed important because it is the final waveform-generation stage.
- CSAVocoder maps multi-channel mel-spectrograms and relative pose to spatial waveforms through a causal GAN-based vocoder supporting chunk-wise streaming inference.
2 Related Work
Prior work spans spatial audio generation, neural waveform synthesis, and real-time streaming, but these components are not consistently integrated. Existing vocoders largely target mono or stereo audio without explicitly modeling spatial cues.
- Spatial Audio Rendering: Binaural audio models ear-canal signals for headphone playback, whereas FOA represents sound fields using spherical harmonics and is common in VR, AR, and 360° video.
- Spatial Audio Generation: Spatial generation methods use visual, textual, multimodal, diffusion, and autoregressive conditions to synthesize binaural or FOA audio.
- Spatial Audio Generation: Many spatial systems operate spectrally and rely on separate vocoders or reconstruction stages, adding latency and often injecting spatial information only implicitly.
- Neural Vocoders: GAN-based neural vocoders are prominent because they offer favorable quality-efficiency trade-offs, alongside structured-domain, diffusion, and flow-based alternatives.
- Neural Vocoders: Existing vocoders primarily target monophonic or stereophonic audio and do not explicitly model spatial cues, limiting spatial-rendering effectiveness.
- Real-time Speech Synthesis: Causal vocoders enable streaming, but autoregressive sample-by-sample generation is too slow, while noncausal fast vocoders access future frames.
3 Method
CSAVocoder extends HiFi-GAN into a causal spatial vocoder that conditions waveform generation on multi-channel acoustics and pose. Its spatial discriminators and format-aware losses supervise inter-channel and sound-field cues, while stateful streaming blocks preserve causality.
- 3.1 Task Definition: The task is to learn G mapping a multi-channel mel-spectrogram M and source-listener pose sequence P to a multi-channel waveform y.M has dimensions C×F×T, P records time-varying position and orientation, and y has dimensions C×L.
- 3.2.1 Generator: The generator extends HiFi-GAN with causal convolutional upsampling and residual blocks redesigned for strictly causal stateful streaming inference.
- 3.2.1 Generator: Shuffle upsampling folds extra channels into time through tensor reordering, preserving causality without temporal mixing.
- 3.2.1 Generator: Streaming residual blocks use dilated causal convolutions and buffers storing the left context required for subsequent chunks.
- 3.2.2 Discriminator: The architecture combines waveform and spectral discriminators with a Spatial Consistency Discriminator that models temporal and cross-channel relationships.
- 3.2.3 Training Objectives: The composite generator objective combines adversarial, feature-matching, mel, STFT, and format-aware spatial losses.
- 3.2.3 Training Objectives: Binaural spatial loss supervises IPD and ILD, while FOA loss supervises intensity-vector direction, magnitude, and diffusion-related sound-field descriptors.
- 3.2.1 Generator: The channel-free generator uses an attention-based mel adaptor for inter-channel fusion and projects the representation to exactly the requested number of waveform channels.
4 Experiments
Experiments evaluate CSAVocoder against vocoder and spatialization baselines using objective, subjective, qualitative, and ablation analyses. The model improves spatial preservation while remaining competitive in audio quality and supporting real-time generation, with causal inference introducing a minor quality trade-off.
- Evaluation Protocol: The evaluation combines objective audio and spatial metrics, subjective MOS tests, qualitative spectrogram comparisons, and component ablations.Objective metrics cover waveform, spectral, temporal, perceptual, angular, and distance similarity.
- Quantitative Comparison: CSAVocoder improves spatial metrics over all vocoder baselines, including diffusion models, while remaining competitive on audio metrics.The principal audio-quality trade-off appears in PESQ, where non-causal Vocos and WaveFM score higher.
- Efficiency: RTF = 0.1587 on a single NVIDIA RTX 4090 GPU, indicating that the causal streaming architecture supports real-time generation.Its real-time factor remains competitive with other GAN-based baselines while providing substantially better spatial preservation.
- Spatialization Baselines: 62.11/77.05 ANG/DIS COS with ground-truth mel-spectrograms exceeds DSP at 26.65/51.07 and BinauralGrad at 50.83/72.13.With ISDrama-generated mel inputs, ISDrama + Ours reaches 47.73/71.44 versus 45.77/67.70 for ISDrama + HiFi-GAN.
- Qualitative Evaluation: Qualitative spectrograms show harmonic stacks, formant trajectories, and left-right spectral patterns that closely match ground truth, with fewer band-wise artifacts and a cleaner noise floor.The comparison places ground truth first, CSAVocoder second, and baselines in the remaining columns.
- Subjective Evaluation: CSAVocoder achieves the highest MOS-P score and a high MOS-Q score, although causal generation can be slightly weaker in audio quality than non-causal models.The MOS tests use a 1-to-5 scale for spatial position and overall audio quality.
- Ablation Study: Removing the Mel Adaptor clearly reduces spatial metrics, removing the Position Adaptor causes moderate degradation, and the Spatial Consistency Discriminator improves spatial metrics.The Spatial Consistency Discriminator requires careful adversarial-weight tuning because otherwise training can become unstable.
5 Conclusion
CSAVocoder jointly targets high-fidelity waveform synthesis and accurate spatial rendering through explicit spatial modeling and causal streaming. Experiments report stronger spatial fidelity than channel-wise baselines while maintaining strong audio quality and real-time performance.
- CSAVocoder jointly addresses high-fidelity waveform synthesis and accurate spatial rendering.
- The Spatial Adaptor fuses multichannel mel-spectrograms with dynamic pose information to capture inter-channel relationships.
- The Spatial Consistency Discriminator explicitly supervises spatial cues.
- The strictly causal, stateful generator supports streaming inference with constant memory overhead.
- CSAVocoder outperforms channel-wise vocoder baselines in spatial fidelity while maintaining strong audio quality and real-time performance.Within the studied binaural and FOA settings, the unified architecture supports multiple spatial formats.
- Explicit spatial modeling and causal streaming provide a foundation for future real-time spatial audio generation.
Limitations
The work identifies limitations in causal baseline comparisons, supported spatial formats, and pose conditioning.
- Fair causal-baseline comparisons remain challenging because buffering strategies and runtime optimizations affect latency and quality.The authors leave a standardized causal-baseline suite with matched latency budgets and consistent objective measurements for future work.
- The evaluation focuses on binaural and FOA formats rather than higher-order ambisonics, multichannel loudspeaker layouts, object-based audio, or personalized HRTF rendering.Higher channel counts alter inductive bias, adversarial-training stability, and computational cost.
- The model conditions on pose, while alternative or complementary representations may be more robust or expressive.
Ethical Considerations
CSAVocoder’s ethical considerations concern integration into upstream speech systems, data governance and privacy, misuse risks, and limitations related to bias and compute-intensive training.
- Integration with upstream TTS/VC systems means both model-related and data-related risks must be considered.
- Public speech and spatial-audio corpora may involve licensing constraints, personally identifying information, and sensitive attributes.
- Low-latency spatial speech generation may enable impersonation, live spoofing, private-conversation re-synthesis, and privacy exposure through logged pose trajectories.
- Potential misuse includes covert surveillance, harassment, social engineering, misleading evidence, and non-consensual downstream profiling.
- Recommended mitigations include acceptable-use terms, provenance signals, deployment restrictions, minimized retention, and transparent reporting of limitations and failure modes.
- Under-represented languages, accents, environments, and accessibility-related speech characteristics may lead to uneven performance, while training remains compute-intensive.
C.2 Simulated Spatial Data from SoundSpaces (MP3D)
The study augments real spatial-audio corpora with SoundSpaces and Habitat-Sim simulations in MP3D indoor environments, covering static and moving source configurations for binaural and FOA audio.
- Simulated data use SoundSpaces and Habitat-Sim with MP3D indoor scenes and binaural or FOA audio sensors.
- Static simulations sample receiver and source positions subject to distance and height constraints, then query binaural or FOA room impulse responses.
- Dynamic simulations keep the listener fixed while the source follows a navigation-mesh path sampled across time steps.
- Clean LibriSpeech speech is convolved with simulated BRIRs or RIRs to create multichannel training audio with normalization to avoid clipping.
- The combined corpus contains approximately 600 hours of binaural data and 900 hours of FOA data, with about 220k and 70k synthesized samples respectively.
- After bandwidth matching at 7.8 kHz, CSAVocoder retains leading spatial consistency, while causal variants outperform Vocos variants on spatial metrics.
D.3 Pose Perturbation Robustness
Pose robustness is evaluated by adding Gaussian noise and randomly dropping pose entries during inference. Performance degrades gradually as perturbations strengthen, while moderate perturbations remain tolerated.
- The evaluation perturbs pose inputs with Gaussian noise or randomly dropped pose entries during inference.
- The model remains robust to moderate pose noise and dropout, with gradual degradation under stronger perturbations.
E FOA Results
FOA experiments evaluate both waveform quality and spatial consistency using established audio and spatial metrics. The supplied passages define the evaluation framework but do not report specific FOA outcomes.
- FOA synthesis is evaluated with PESQ, MRSTFT, and MCD for audio quality, plus Corr_all and AUC_j_all for spatial consistency.
F Latency Evaluation
The latency evaluation defines algorithmic and compute latency alongside real-time factor, then benchmarks streaming inference across chunk sizes. Results show approximately 15 ms/chunk mean compute latency and real-time operation across all tested settings.
- Latency measures: Streaming latency is evaluated using algorithmic latency, compute latency, and real-time factor.Algorithmic latency captures design-inherent delay, compute latency measures forward-pass time per chunk, and RTF compares generation time with audio duration.
- Latency measures: For this model, Tlookahead = 0 and Toverlap = 0, indicating no future-context or overlap-add waiting under the stated latency formulation.
- Latency setup: Chunk sizes of 40/60/80/100 ms correspond to 6/9/12/15 mel frames, respectively.
- Latency setup: The benchmark uses batch size 1, disabled gradients, synchronized GPU timing, warm-up, repeated iterations, and p50/p90/p99 latency statistics.Unless stated otherwise, compute latency includes model inference but excludes feature extraction and file operations.
- Results and discussion: Approximately 15 ms/chunk mean compute latency remains stable across chunk sizes, while RTF improves with larger chunks and all settings achieve RTF < 1.The reported behavior is attributed to fixed GPU overheads being amortized more effectively by larger chunks.
G.2 MUSHRA-style Expert Evaluation
The expert evaluation uses a MUSHRA-style protocol to assess quality and spatial perception across CSAVocoder and five baselines. CSAVocoder obtains the highest spatial-perception and overall scores, while its quality score is slightly below two baselines.
- Protocol: The MUSHRA-style test follows ITU-R BS.1534 and evaluates 120 dynamic spatial-audio results with 9 spatial-audio experts.The evaluated systems are CSAVocoder, Vocos, WaveFM, DiffWave, PriorGrad, and FastDiff.
- Protocol: Each trial includes an explicit reference, hidden reference, and 3.5 kHz low-pass/dual-mono anchor, with ratings on a 0 to 100 scale.
- Metrics: MUSHRA-Q measures quality, while MUSHRA-P measures spatial perception across direction, distance, externalization, and image stability.
- Results: 88.0 MUSHRA-P and 85.5 overall score are achieved by CSAVocoder, the highest values among the evaluated systems.Its MUSHRA-Q is slightly below Vocos and PriorGrad, while its spatial-perception score exceeds all baselines by a large margin.
- Results: Hidden-reference scores close to 100 and anchor scores close to the lower bound support the reliability of the evaluation protocol.