Source-linked AI summary

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, Wenwu Wang

arXiv:2303.17395v2eess.AScs.CLcs.MMcs.SD

TL;DR

Audio-language research lacks large datasets because audio collection is costly and time-consuming. The paper introduces WavCaps, using a three-stage ChatGPT-assisted pipeline to turn noisy web descriptions into captions. WavCaps supports state-of-the-art results across the evaluated audio-language tasks and has influenced subsequent research.

  • Problem

    Audio-language research is limited by the scarcity of large datasets, while human-annotated audio captioning data is expensive and time-consuming to collect.

  • Method

    WavCaps collects audio and raw descriptions from web sources and AudioSet, then uses pre-filtering, ChatGPT-based filtering and rewriting, and post-processing.

  • Results

    WavCaps evaluation achieves new state-of-the-art performance across multiple audio-language multimodal learning tasks.

  • Takeaways & Limitations

    WavCaps is intended to facilitate audio-language research and demonstrate how ChatGPT-like language models can enrich academic research.

Abstract

from arXiv · show

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which are limited in size. To address this data scarcity issue, we introduce WavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration is for the WavCaps dataset we have proposed to facilitate research in audio-language multimodal learning and demonstrate the potential of utilizing ChatGPT to enhance academic research. Our dataset and codes are available at https://github.com/XinhaoMei/WavCaps.

I. INTRODUCTION

Audio-language research is constrained by scarce, costly datasets, motivating WavCaps: a large weakly-labelled corpus built from web audio descriptions and processed with ChatGPT. Experiments across multiple tasks report substantial gains over prior benchmarks.

  • Audio-language datasets remain limited because audio collection is more laborious, costly, and time-consuming than visual-data collection.
  • WavCaps gathers audio clips and raw descriptions from three web platforms and AudioSet to address audio-language data scarcity.
  • The three-stage pipeline pre-filters data, uses ChatGPT to filter and rewrite descriptions, then post-processes outputs.
  • WavCaps contains about 400k audio clips with paired captions and is presented as the first large-scale weakly-labelled audio captioning dataset.
  • Experiments on multiple audio-language tasks achieve new state-of-the-art results on most tasks, surpassing previous benchmarks by significant margins.
  • The work contributes WavCaps, ChatGPT-based caption processing, and downstream evaluations demonstrating dataset effectiveness.

II. RELATED WORKS

Prior vision-language datasets demonstrate large-scale web harvesting, while audio-language research remains constrained by limited human-annotated resources. WavCaps adapts web harvesting but uses ChatGPT to filter and rewrite noisy audio descriptions.

  • Automatically annotated vision-language datasets harvest millions of image-caption pairs from online platforms.
  • WavCaps follows web harvesting but replaces complex predefined filtering rules with ChatGPT-based filtering and rewriting of raw descriptions.
  • Audio-language research is limited by scarce datasets, with major tasks relying on expensive and time-consuming human-annotated AudioCaps and Clotho data.
  • AudioCaps contains about 50k clips, while Clotho contains around 6k clips, illustrating the smaller scale of human-annotated audio captioning resources.
  • Earlier web-harvested resources include SoundDescs, WavText5K, and LAION-Audio-630K for audio-language tasks.
  • Subsequent work has used WavCaps for new datasets, text-to-audio generation, and large audio-language models.

III. WAVCAPS DATASET

WavCaps combines heterogeneous audio sources with a three-stage pipeline designed to retain data while converting noisy descriptions into usable captions. ChatGPT performs semantic filtering and rewriting, followed by quality-control processing.

  • A. Data Sources: WavCaps collects audio from FreeSound, SoundBible, and AudioSet’s strongly-labelled subset, combining diverse web and dataset sources.
  • A. Data Sources: FreeSound descriptions are noisy because they may be unrelated, fragmentary, overly specific, or repetitive.
  • B. Data Processing: The pipeline simplifies filtering and transformation to retain more harvested audio-description pairs than stringent image-dataset processing procedures.
  • B. Data Processing: Pre-filtering removes clips shorter than one second and high-frequency FreeSound descriptions.
  • B. Data Processing: ChatGPT prompts require concise, accurate single-sentence captions while excluding unrelated names, locations, devices, emotions, and opinions.
  • B. Data Processing: ChatGPT transforms fragments into sentences, removes irrelevant specificity, condenses lengthy descriptions, and can output “Failure” for non-audio content.
  • B. Data Processing: Post-processing addresses residual prompt-following failures, including retained numbers, names, locations, and some irrelevant descriptions.

C. Dataset Analysis

WavCaps processing substantially reduces noisy raw descriptions while retaining diverse audio sources, durations, vocabulary, and caption expressions. The resulting dataset is much larger than human-labeled alternatives, but its captions remain weakly labeled and generally coarse-grained.

  • Data filtering: More than half of FreeSound samples were filtered out during processing, while only a small number from BBC Sound Effects and SoundBible were removed.Samples shorter than one second were excluded, increasing the average duration of FreeSound samples.
  • Caption transformation: ChatGPT-augmented captions show low lexical overlap with raw descriptions, indicating substantial transformation and word deletion.Jaccard similarity measures shared words divided by the total unique words across the raw description and augmented caption.
  • Caption quality: The processed dataset contains meaningful sound-related vocabulary, whereas harvested raw descriptions are largely devoid of meaning.The analysis reports that ChatGPT extracted sound-related information and removed unrelated information.
  • Caption diversity: WavCaps contains 330 609 unique captions, including 311 242 that appear only once.Captions recurring more than five times average 4.9 words, compared with 7.8 words for all captions.
  • Dataset scale and diversity: WavCaps is an order of magnitude larger than human-labeled audio captioning datasets and contains more diverse content.The dataset retains long audio clips, producing a longer total duration than comparable datasets.
  • Scope and limitations: Because processing uses text metadata rather than audio, captions inherit omissions or inaccuracies in the original descriptions and generally lack spatial-temporal details.The design intentionally focuses on coarse-grained descriptions, although AudioSet strongly-labeled captions may include temporal relationships.

D. Human Validation

Human validation found that ChatGPT-augmented captions generally correspond well to the audio, while often omitting events or specificity present in human annotations. These omissions support WavCaps’s weakly-labeled designation.

  • Caption accuracy: 3.89 was the mean accuracy score for ChatGPT-augmented captions, and more than 40% received the perfect score of 5.The MOS scale ranges from 1 for no correspondence to 5 for an exact match.
  • Comparison with human captions: Most evaluators selected category (B), meaning captions partially covered the audio while omitting some events or specificity.The prompt intentionally generated more general captions by discarding specific details.
  • Weak labeling: Lower scores occurred when raw descriptions failed to cover all sound events, because augmented captions preserved those omissions.This limitation is a principal reason the dataset is labeled weakly.
  • Comparison with human captions: Category (C) received a mean accuracy score of 3.17, partly reflecting differing interpretations in Clotho’s audio-only annotation process.Category (C) denotes captions judged significantly different from human-annotated captions.

IV. EXPERIMENTS

The experiments evaluate WavCaps across audio-language retrieval settings using two-tower acoustic-semantic models with alternative audio encoders. The evaluation covers zero-shot, pretraining, and fine-tuning regimes on AudioCaps and Clotho.

  • Audio-language retrieval: Audio-language retrieval maps paired audio clips and captions closer in embedding space while separating non-paired examples.The study evaluates zero-shot, pretraining, and fine-tuning settings.
  • Models: The retrieval models use a two-tower architecture with separate audio and language encoders.The audio encoders are CNN14 or HTSAT, and the language encoder is pretrained BERTbase.
  • Evaluation: The experiments report retrieval results on AudioCaps and Clotho test sets, with higher scores indicating better performance.Table V distinguishes AudioCaps, LAION-Audio-630K, zero-shot, pretraining, and fine-tuning conditions.
  • Training objective: The contrastive training strategy uses batch size B and temperature hyperparameter τ.This strategy is also known as contrastive language-audio pretraining (CLAP).

2) Experimental Setup:

The experiments compare WavCaps-trained models with existing systems and larger-data baselines across audio-language retrieval settings. Results indicate strong zero-shot generalization and substantial gains after fine-tuning, while model performance varies by dataset and encoder.

  • Experimental Setup: Zero-shot WavCaps models outperform previous SOTA models on Clotho retrieval and achieve comparable results on AudioCaps.The models generalize across both evaluation datasets despite training without overlapping samples.
  • Experimental Setup: 16.7% improvement in R@1 on AudioCaps and 5.4% improvement in R@1 on Clotho are achieved for audio-to-text retrieval after fine-tuning.Fine-tuned models also improve text-to-audio retrieval and outperform existing methods by significant margins.
  • Experimental Setup: WavCaps-trained models outperform LAION’s models on most metrics for both datasets despite using less data.LAION’s training set contains approximately 2.63 million examples, about six times larger than the compared training set.
  • Experimental Setup: CNN14 underperforms HTSAT on AudioCaps but surpasses HTSAT on Clotho.The authors associate this difference with Clotho’s variable-duration clips and possible information loss from HTSAT’s 10-second random crops.

B. Automated Audio Captioning

The automated audio captioning experiments use encoder-decoder models with CNN14 or HTSAT audio encoders and a pretrained BART-based decoder. WavCaps pretraining and fine-tuning produce strong zero-shot and state-of-the-art captioning results on AudioCaps and Clotho.

  • B. Automated Audio Captioning: Audio captioning generates a natural-language sentence describing an audio clip, primarily focusing on environmental sounds.The task is evaluated on the AudioCaps and Clotho datasets.
  • B. Automated Audio Captioning: The captioning architecture combines a CNN14 or HTSAT audio encoder with a pretrained BART-based language decoder.The encoder extracts audio features, which the decoder uses to generate captions.
  • B. Automated Audio Captioning: Zero-shot models demonstrate robust captioning capabilities, supporting WavCaps’s suitability for the task.Overlapping AudioCaps and Clotho samples are excluded during zero-shot training.
  • B. Automated Audio Captioning: Fine-tuned models achieve new SOTA performance on both Clotho and AudioCaps, surpassing existing methods.Pretraining on WavCaps significantly improves final performance over baseline systems on both datasets.
  • B. Automated Audio Captioning: CNN14 outperforms HTSAT on Clotho but performs worse on AudioCaps.This encoder-dependent pattern aligns with the retrieval experiments and may relate to variable audio duration in Clotho.

1) Models:

The audio classification setup reformulates classification as audio-language retrieval and evaluates zero-shot transfer on three audio event datasets. WavCaps-trained models achieve SOTA zero-shot results, with performance varying by dataset granularity and data leakage controls.

  • 1) Models:: Class probabilities are obtained by comparing audio embeddings with prompt-free text embeddings for each class label.Similarity scores between audio and class embeddings are normalized into a probability distribution.
  • 2) Experimental Setup:: Zero-shot evaluation uses ESC-50, UrbanSound8K, and VGGSound, with top-1 accuracy as the metric.The datasets contain 50, 10, and 310 classes, respectively.
  • 2) Experimental Setup:: The comparison excludes overlapping evaluation samples from the merged WavCaps, AudioCaps, and Clotho training set.This setup is intended to prevent overlap between training data and evaluation samples.
  • 3) Results and Analysis:: WavCaps-trained models achieve SOTA zero-shot results on all three classification datasets.They significantly outperform other models on ESC-50 and UrbanSound8K.
  • 3) Results and Analysis:: WavCaps models outperform LAION and BLAT using less data, while an AudioSet–VGGSound overlap creates a reported leakage issue for LAION.The authors identify the leakage because overlapping samples were not excluded in that comparison.
  • 3) Results and Analysis:: Zero-shot results are close to supervised SOTA on ESC-50 and UrbanSound8K but show a considerable margin on VGGSound.The authors suggest VGGSound’s 310-class granularity may make generalization more difficult.

D. Text-based Sound Generation

The text-based sound generation experiments adapt AudioLDM with CLAP encoders trained on WavCaps, with optional AudioCaps fine-tuning, and compare against prior generation systems. Evaluation uses distributional, diversity, and sample-level similarity metrics on AudioCaps.

  • D. Text-based Sound Generation: Text-based sound generation produces speech, music, or sound effects from textual information and is evaluated on AudioCaps.The experiments follow prior work using the AudioCaps dataset for training and evaluation.
  • D. Text-based Sound Generation: AudioGen comparisons may be unreliable because its pretrained model and evaluation data are not open-sourced.This caveat is stated in the Table VIII caption.
  • D. Text-based Sound Generation: AudioLDMWavCaps replaces the original AudioLDM CLAP encoder with one trained on WavCaps, while AudioLDMWavCaps-FT additionally fine-tunes CLAP on AudioCaps.The original AudioLDM uses a CLAP model trained on approximately 2.6 million audio-text pairs.
  • D. Text-based Sound Generation: The diffusion objective predicts Gaussian noise from a noisy VAE latent conditioned on text-derived embeddings.The forward process transforms the original latent z0 into the n-th noisy latent zn according to a noise schedule.
  • D. Text-based Sound Generation: Models are trained for 400k steps and selected using the best Frechet Audio Distance on the AudioCaps test set.Evaluation checkpoints are assessed every 50k training steps.
  • D. Text-based Sound Generation: FAD and FD measure distribution similarity, IS measures diversity and target-distribution similarity, and KL measures sample-level similarity.These metrics provide complementary views of generated-audio quality.

3) Results and Analysis:

The experiments show that WavCaps processing improves retrieval and captioning, while WavCaps-based AudioLDM remains competitive despite using substantially less data. ChatGPT-based caption refinement is especially beneficial for automated audio captioning.

  • AudioLDM evaluation: Comparable AudioLDM performance is achieved with WavCaps using only 15% of AudioLDM_LAION’s dataset size.WavCaps performs better on FAD and inception score with text embeddings, but performs less well on IS with audio embeddings.
  • Ablation design: The ablation analysis merged ChatGPT-based transformation and post-processing to compare raw descriptions with ChatGPT-augmented captions.The FreeSound subset was selected because of its complexity and the first pipeline step’s focus on filtering it.
  • Zero-shot retrieval: 561 124 FreeSound samples were reduced to 257 040 after pre-filtering, yet all zero-shot retrieval scores improved.Adding ChatGPT caption augmentation produced slight further improvements across all scores, while additional resources improved results further, especially on AudioCaps.
  • Overall outcome: Across multiple audio-language tasks, systems trained on WavCaps achieved new state-of-the-art performance.The conclusion presents this as the overall evaluation outcome for the dataset.
Loading 2303.17395v2…