Source-linked AI summary

Stable Audio Open

Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons

arXiv:2407.14358v2cs.SDcs.AIeess.AS

TL;DR

Many text-to-audio models are inaccessible because they lack public weights and transparent training-data licenses. Stable Audio Open releases an open-weights latent diffusion model trained on Creative Commons audio, reporting competitive results and potential for high-quality 44.1kHz stereo synthesis.

  • Problem

    Many text-to-audio models lack public weights or fully documented training-data licenses, limiting their usefulness for research and artistic creation.

  • Method

    Stable Audio Open is a latent diffusion model with an autoencoder, T5 text conditioning, and a diffusion transformer, trained on Creative Commons audio.

  • Results

    The model performs competitively with state-of-the-art systems across reported metrics and shows potential for high-quality stereo sound synthesis at 44.1kHz.

  • Takeaways & Limitations

    Public weights, code, data attributions, and consumer-grade GPU operation support academic and artistic use cases.

  • Takeaways & Limitations

    Limited high-quality music in the Creative Commons training data leaves the model noncompetitive with state-of-the-art music models.

Abstract

from arXiv · show

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

1. INTRODUCTION

Stable Audio Open addresses the limited accessibility and transparency of existing text-conditioned audio models by releasing an open-weights model trained on Creative Commons audio. The model targets open research and artistic creation while reporting high-quality stereo generation at 44.1kHz.

  • Existing text-conditioned audio models are often private or lack public weights, limiting their use for research and artistic creation.
  • The model is trained only on Creative Commons-licensed audio.
  • Public model weights, code, and data attributions are released to facilitate open research.
  • Stable Audio Open targets state-of-the-art sound-quality generation in 44.1kHz stereo.
  • The release emphasizes evaluation, data transparency, accessibility, and operation on consumer-grade GPUs.

2. ARCHITECTURE

Stable Audio Open uses a latent diffusion architecture that generates variable-length stereo audio from text prompts. An autoencoder compresses waveforms into latents, while a T5-conditioned diffusion transformer generates audio in latent space.

  • The model generates variable-length stereo audio up to 47 seconds at 44.1kHz from text prompts.
  • Its three components are a waveform autoencoder, T5-based text embedder, and transformer-based diffusion model.
  • 2.1. Autoencoder: The autoencoder uses convolutional downsampling and upsampling around a variational bottleneck with latent size 64.
  • 2.2. Diffusion-transformer (DiT): The DiT combines attention, gated MLPs, skip connections, cross-attention conditioning, and memory-saving computation techniques.
  • 2.2. Diffusion-transformer (DiT): Conditioning signals encode text, timing for variable length, and the current diffusion timestep.
  • 2.2. Diffusion-transformer (DiT): Variable-length generation fills a specified window with silence, which can be trimmed for shorter outputs.

3. TRAINING DATA

The training dataset combines Creative Commons recordings from Freesound and the Free Music Archive after procedures intended to exclude copyrighted content. Metadata is transformed into text prompts for conditioning.

  • The dataset contains Creative Commons recordings from Freesound and the Free Music Archive, screened for copyrighted content.
  • The Free Music Archive filtering process produced 8,967 CC-BY and 4,907 CC0 tracks after metadata matching and human review.
  • The final dataset contains 486,492 recordings totaling 7,300 hours under CC-0, CC-BY, or CC-Sampling+ licenses.
  • Audio was sampled in 5-second chunks, with additional high-fidelity material from FMA and selected stereo, full-band Freesound recordings.
  • Training audio is paired with descriptions, titles, tags, and music metadata used to generate text prompts.

4. EXPERIMENTAL SETUP

The authors train the autoencoder and diffusion transformer with AdamW-based optimization and evaluate generation using established audio-quality and text-audio metrics. Training uses separate objectives and schedules for the autoencoder, decoder, and DiT.

  • All models use AdamW with weight decay of 0.001 and a learning-rate scheduler with exponential ramp-up and decay.
  • The autoencoder uses reconstruction, adversarial feature-matching, and KL-divergence losses.
  • The encoder trains for 183 hours and the decoder for 273 hours on 32 A100 GPUs, using 44.1kHz audio chunks of approximately 1.5 seconds.
  • The DiT predicts noise increments from noised ground-truth latents using the v-objective and is trained for 338 hours on 64 A100 GPUs.

5. EVALUATION

The evaluation measures generation quality on sound and instrumental-music benchmarks, alongside autoencoder reconstruction, memorization, and inference speed. Stable Audio Open performs strongly for sounds and field recordings, is weaker for music than Stable Audio but slightly better than MusicGen, and shows no observed memorization beyond similarities in well-defined sounds.

  • 5.1. Generative model evaluation: The model is evaluated with FDopenl3, KLpasst, and CLAPscore on AudioCaps for sound generation and Song Describer for music generation.Lower FDopenl3 indicates plausible similarity to references, lower KLpasst indicates semantic correspondence, and higher CLAPscore indicates text adherence.
  • 5.1. Generative model evaluation: Stable Audio Open outperforms comparable baselines on AudioCaps, particularly on FDopenl3, indicating potential for realistic sounds and field recordings.AudioCaps contains sound rather than music; comparisons with AudioLDM2 and AudioGen may be affected because they were trained on AudioSet, of which AudioCaps is a subset.
  • 5.1. Generative model evaluation: Stable Audio Open performs worse than Stable Audio for instrumental music but slightly better than MusicGen, described as the best open model.The evaluation uses 586 instrumental-music captions and 446 tracks from the Song Describer Dataset.
  • 5.2. Autoencoder evaluation: Autoencoder reconstruction quality is assessed on sound and instrumental-music datasets using STFT distance, MEL distance, and SI-SDR.The evaluation compares ground-truth and reconstructed audio, with results reported in Tables 3 and 4.
  • 5.3. Memorization (exact copies) analysis: Memorization analysis found no memorization among candidates from 11,000 random training prompts or outstanding generations beyond similarities involving sounds such as “storm” or “1000Hz”.The dataset analysis identified 3,693 repeated Freesound recordings and 856 repeated FMA recordings before generation-level listening tests.
  • 5.4. Inference speed: Inference runs at 8 steps/sec on an RTX-3090, 11 steps/sec on an RTX-A6000, and 20 steps/sec on an H100.The reported hardware configurations have 24GB, 48GB, and 80GB of VRAM, respectively.

6. CONCLUSIONS

Stable Audio Open is released as an accessible, transparent text-to-audio model for stereo sound synthesis. Its main limitations concern connectors, speech, music quality, and prompting scope.

  • Contributions: The model is accessible to everyone, can run on consumer-grade GPUs, and is intended for academic and artistic use.It synthesizes high-quality stereo sounds at 44.1kHz and includes a continuous autoencoder operating at 21.5Hz.
  • Audio generation: It struggles with prompts containing connectors, sometimes omitting sounds, and cannot generate intelligible speech.Speech is not intelligible because the model is not conditioned for spoken word.
  • Audio generation: CLAPscore improves when connector- and speech-related prompts are removed.
  • Music generation: Limited high-quality Creative Commons music data makes the model not competitive with state-of-the-art music models.The training focus on CC data constrained the available music data.
  • Prompting: Prompt engineering may be required for best results, and performance is not expected to be strong in languages other than English.

A. ADDITIONAL SONG DESCRIBER DATASET RESULTS

The appendix reports additional results on the complete Song Describer Dataset, complementing the main evaluation on instrumental prompts.

  • Evaluation scope: The main paper reports Song Describer results on an instrumental-music subset to ensure fair evaluation.
  • Additional results: The appendix adds results for the entire Song Describer Dataset using the same metrics as the main tables.
  • Evaluation scope: The full-dataset evaluation broadens coverage beyond the no-singing subset used in the main body.

A.1. Generative model evaluation

On the complete Song Describer Dataset, Stable Audio Open is weaker than Stable Audio for music generation but slightly better than MusicGen among open models.

  • Dataset: The full dataset contains 1,106 captions for 706 tracks, producing 1,106 generated audios.
  • Generative model results: Stable Audio Open performs worse than Stable Audio at generating music on the complete dataset.
  • Generative model results: Stable Audio Open performs slightly better than MusicGen, identified as the best open model in this comparison.

A.2. Autoencoder evaluation

Full-dataset autoencoder results indicate that lower latent rates generally worsen reconstructions, while Stable Audio Open is slightly below Stable Audio 2.0 under the comparable setting.

  • Evaluation setup: The reconstruction evaluation uses 706 tracks from the complete Song Describer Dataset.
  • Reconstruction results: Lower latent rates generally yield worse audio reconstructions.
  • Reconstruction results: Compared with Stable Audio 2.0 at the same latent rate, the autoencoder shows slightly worse performance.The comparison is made because Stable Audio 2.0 is continuous and uses the same latent rate.

B. VRAM CONSUMPTION DURING INFERENCE

Inference memory use differs substantially between diffusion and decoding: the DiT uses 5.9 GB VRAM, while waveform rendering raises RAM usage to 14.5 GB.

  • 5.9 GB VRAM is used during diffusion, while decoding waveforms from latents increases RAM usage to 14.5 GB.

B.1. Chunk decoding

Chunk decoding reduces decoder memory usage by splitting latent sequences into overlapping chunks and reassembling the separately decoded waveforms without changing the resulting audio.

  • Chunk decoding splits the latent sequence into overlapping chunks, decodes them separately, and reassembles the final audio.The resulting audio remains the same when overlaps include the decoder’s receptive field of 16 latents on each side.
  • The decoder’s receptive field requires overlaps of 16 latents on each side to preserve the same resulting audio.
Loading 2407.14358v2…