Source-linked AI summary

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Thorsten Müller, Piotr Syga, Philip Sperl, Konstantin Böttinger

arXiv:2401.09512v11cs.SDeess.AS

TL;DR

Audio deepfake detection needs broader evidence because laboratory performance may not transfer to real-world settings and existing spoof data is heavily English-focused. The paper introduces MLAAD, a multilingual synthetic-speech dataset, and evaluates it with state-of-the-art detectors across multiple datasets. MLAAD outperforms InTheWild and FakeOrReal as a training resource, while MLAAD and ASVspoof 2019 each achieve the best accuracy on four of eight cross-dataset tests.

  • Problem

    Audio deepfake detection remains difficult in practical environments, while English-dominated data limits applicability to non-English speakers.

  • Method

    The paper constructs MLAAD from authentic speech and synthetic audio across 54 languages and trains deepfake detectors to evaluate cross-dataset performance.

  • Results

    Across eight cross-dataset tests, MLAAD and ASVspoof 2019 each achieve the highest accuracy on four datasets, while MLAAD outperforms InTheWild and FakeOrReal.

  • Takeaways & Limitations

    MLAAD complements ASVspoof 2019 and provides an openly available multilingual resource and accessible trained models for audio anti-spoofing.

Abstract

from arXiv · show

This paper presents the Multi-Language Audio Anti-Spoofing Dataset (MLAAD), version 11: a dataset of synthetic audio to train and evaluate audio deepfake detection models. It features 205 Text-to-Speech (TTS) models, comprising a total of 1153.6 hours of synthetic voice in 54 different languages. To evaluate this dataset, we train three state-of-the-art deepfake detection models with MLAAD and observe that it demonstrates superior performance to comparable datasets like InTheWild and FakeOrReal when used as a training resource. Moreover, compared to the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, each excelling on four datasets. By publishing the dataset and making a trained model accessible via an interactive webserver, we aim to democratize anti-spoofing technology, making it accessible beyond the realm of specialists, and contributing to global efforts against audio spoofing and deepfakes.

I. INTRODUCTION

Audio deepfakes and spoofs create risks for fraud, misinformation, and biometric systems, while detection remains difficult across real-world conditions and languages. MLAAD addresses these challenges with a large multilingual synthetic-audio dataset and publicly accessible models.

  • TTS advances support beneficial applications but also enable deepfakes that can facilitate fraud, misinformation, and fake news, plus spoofs that can compromise biometric identification.
  • Reliable detection remains an active research need because models trained in controlled laboratory settings may fail in practical environments.
  • English-dominated deepfake datasets limit applicability and may discriminate against non-English speakers, motivating multilingual detection.
  • MLAAD provides 1153.6 hours of synthesized speech across 54 languages, generated by 205 state-of-the-art TTS models.
  • The authors open-source MLAAD and provide interactive access to trained models for non-technical users through deepfake-total.com.

A. Speech Synthesis

Speech synthesis primarily comprises text-to-speech and voice conversion, while anti-spoofing models classify audio as authentic or fake using raw signals or transformed representations.

  • Text-to-speech generates speech from text, with voice, accent, and pitch characteristics derived from training data.
  • Voice conversion generates an utterance from separate linguistic and vocal-characteristic inputs.
  • Neural voice anti-spoofing classifies audio files as authentic, or bona-fide, versus fake, or spoof.
  • Anti-spoofing models either process raw waveforms directly or analyze transformed representations such as mel-spectrograms and cepstral coefficients.

C. Datasets for Deepfake Detection

Existing anti-spoofing datasets broaden generation and recording conditions but remain limited in language diversity. MLAAD extends authentic multilingual speech with synthetic audio generated across many languages and TTS configurations.

  • Existing datasets target weak generalization by covering generation methods, codecs, noises, data quality, partial spoofs, real-world deepfakes, and multimodal examples.
  • Most existing datasets focus predominantly on English, leaving a gap in the language diversity of spoofed utterances.
  • MLAAD extends M-AILABS with authentic speech in eight original languages and computer-generated audio in 54 languages.
  • For each language–TTS-model combination, MLAAD selects 1000 source instances and translates English text when the target language is absent from M-AILABS.
  • Each baseline file receives a corresponding synthetic version, optionally using a randomly selected speaker reference for multi-speaker TTS models.

B. Formatting

MLAAD organizes generated audio by language and model type, with metadata describing each file’s provenance, language, duration, training data, model, architecture, transcript, and reference speaker.

  • Each output directory contains synthesized audio and a meta.csv file with standardized descriptive fields.
  • Metadata records file paths, baseline-file links, language and translation status, duration, TTS training data, model name, architecture, transcript, and reference speaker.

C. Statistics

MLAAD contains 1153.6 hours of synthesized speech across 54 languages, generated by 205 TTS models spanning 127 architectures. The dataset distributes synthesized samples, while original audio remains available through M-AILABS.

  • 1153.6 hours of synthesized speech span 54 languages and 205 state-of-the-art TTS models across 127 architectures.
  • MLAAD distributes synthesized samples, while users must obtain the original audio from the M-AILABS dataset.
  • Tables IV and V summarize the TTS architectures and languages constituting MLAAD.

A. Training Setup

The evaluation trains three anti-spoofing models using audio segments, augmented data, balanced sampling, and early stopping. Accuracy is used instead of EER because some test datasets contain only spoof samples.

  • Three state-of-the-art voice anti-spoofing models—RawGat-ST, SSL-W2V2, and WhisperDF—are selected for evaluation.
  • RawGat-ST and SSL-W2V2 process 5-second raw-audio clips, whereas WhisperDF processes 30-second segments converted into time-frequency representations.
  • Training data is augmented with randomly selected noise or music, overlaid with a 5% probability per sample.
  • Random undersampling provides equal fake and authentic samples per epoch, and models train for 100 epochs with early stopping based on training accuracy.
  • Accuracy replaces EER because some test datasets, including Voc.v, contain only spoof samples.

B. Cross-Dataset Evaluation

Models are trained separately on four datasets and evaluated across eight datasets to measure cross-dataset generalization. MLAAD and ASVspoof19 each lead four cases, while some learned features transfer poorly or become detrimental.

  • Four training datasets are evaluated across eight datasets, with later releases not re-evaluated in the reported results.
  • Five repeated trials report mean accuracy and standard deviation to quantify cross-dataset generalization for real-world applicability.
  • ASVspoof19 and MLAAD each achieve the highest accuracy in four of eight cross-dataset test cases, while FakeOrReal and InTheWild top none.
  • The shared success of ASVspoof19 and MLAAD suggests that the datasets complement each other in training effectiveness.
  • SSL-W2V2 trained on ASVspoof19 reaches 51% accuracy on WaveFake, while MLAAD-trained models are the only ones not significantly below 50%.

V. CONCLUSION

MLAAD provides multilingual voice spoofs from 205 TTS models and complements existing anti-spoofing datasets. Across eight datasets, MLAAD and ASVspoof 2019 each excel on four, while multilingual evaluation remains future work.

  • MLAAD contains voice spoofs in 54 languages generated by 205 TTS models spanning 127 architectures.
  • MLAAD consistently outperforms InTheWild and FakeOrReal for the trained models’ cross-dataset generalization capability.
  • Across eight datasets, MLAAD and ASVspoof 2019 each achieve the best performance on four datasets, indicating complementary training resources.
  • The multilingual anti-spoofing capability of MLAAD remains for future evaluation because no comparable multilingual datasets are available.
Loading 2401.09512v11…