Source-linked AI summary
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
Xin Wang, Héctor Delgado, Hemlata Tak, Jee-weon Jung, Hye-jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu, Md Sahidullah, Tomi Kinnunen, Nicholas Evans, Kong Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Ge Zhu, Yongyi Zang, You Zhang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Cheng Gong, Hanjie Guo, Liping Chen, Vishwanath Singh
TL;DR
Speech deepfakes pose a growing security concern, while existing databases have limited acoustic diversity and speaker coverage. ASVspoof 5 addresses these gaps with a crowdsourced, diverse database and validates it using speaker-verification and spoof/deepfake detectors, finding that generative attacks remain challenging for robust detection.
Problem
Speech deepfakes threaten communication and biometric systems, while earlier ASVspoof databases were constrained by studio-quality data and approximately 100 speakers.
Method
The paper designs, crowdsources, and validates ASVspoof 5 using diverse MLS speech, substantially more speakers, varied TTS/VC and adversarial attacks, shortcut-artefact reduction, surrogate-model optimisation, and baseline detectors.
Results
Generative technology still poses a grave threat to automatic-system reliability, while baseline spoof/deepfake countermeasures achieve pooled development-set EERs of 29.49% and 17.83%.
Takeaways & Limitations
ASVspoof 5 provides a more diverse benchmark for evaluating automatic speaker verification and spoof/deepfake detection, with most resources freely available for reproducible research.
Takeaways & Limitations
The paper’s visualisations rely on oracle speaker and attack labels that are unavailable during detection.
Abstract
from arXiv · showhide
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ~2,000 speakers (cf. ~100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
1. Introduction
ASVspoof 5 addresses the need for realistic, representative speech-deepfake detection data by replacing studio-quality sources with diverse crowdsourced recordings and many more speakers. It introduces contemporary attack generation, surrogate-model tuning, and validation resources for spoofing and deepfake detection.
- Speech-deepfake detectors require realistic, representative data because their reliability depends strongly on the data used for implementation and training.
- Earlier ASVspoof databases relied mainly on studio-quality VCTK recordings, limiting how reliably their results estimated performance in practical acoustic conditions.
- ASVspoof 5 uses crowdsourced MLS English speech from diverse recording settings and thousands of speakers, enabling more flexible protocols and surrogate-model tuning.
- The database combines distinct TTS and VC attacks, optional adversarial attacks, and speaker-disjoint data partitions for training, development, and evaluation.
- The paper covers database design, crowdsourced attack collection, shortcut-artifact reduction, attack visualisation, and validation with CM and ASV baselines.
2. Database generation
ASVspoof 5 is generated from the diverse MLS English database through speaker-disjoint partitions, separate attack-training resources, and protocols supporting both detection and spoofing-robust speaker verification. Its attack pipeline covers TTS, VC, adversarial generation, adaptation-data variation, and optional surrogate-model tuning.
- Attack generation uses disjoint source partitions and held-out adaptation or input utterances, with separate training resources for training, development, and evaluation attacks.
- The MLS English source contains recordings from approximately 2,400 female and 2,300 male speakers collected in varied environments and with different devices.
- MLS data are divided into seven speaker-disjoint subsets supporting ASVspoof 5 sets, TTS/VC and adversarial attack training, and surrogate CM and ASV development.
- Eight TTS/VC systems generate the ASVspoof 5 training set, while eight generate development data and sixteen TTS/VC systems plus adversarial attacks generate evaluation data.
- Adaptation configurations vary the amount and collection conditions of target-speaker data to test their influence on spoofed-speech generation and surrogate ASV performance.
3. Spoofing attacks
ASVspoof 5 includes attacks spanning legacy and contemporary TTS/VC technologies, zero-shot and few-shot approaches, and adversarial perturbations. Its attack sets combine diverse generation strategies, including classical unit-selection and modern neural systems.
- Attack scope: ASVspoof 5 includes both recent and legacy TTS/VC algorithms because successful attacks need not use the latest or highest-quality synthesis.The database also introduces adversarial attacks designed to increase CM or ASV error rates.
- Attack types: TTS attacks synthesize target-speaker speech from text and adaptation utterances, using zero-shot voice cloning or few-shot speaker-adaptive approaches.The acoustic decoder produces acoustic features, which a vocoder transforms into a waveform.
- Attack types: VC attacks convert a non-target speaker’s voice into a target speaker’s voice while preserving the input utterance’s linguistic content.VC systems use adaptation utterances containing the target speaker’s voice.
- Dataset partitions: The training set uses zero-shot voice-cloning TTS, while the development set mixes zero-shot and few-shot TTS and VC systems.The training attacks are summarized in the top block of Table 2, and development attacks in its second block.
- Evaluation attacks: A19 uses classical MaryTTS unit selection, selecting speech units that match input phonemes while considering concatenation distortion.The paper notes that MaryTTS attacks can threaten ASV reliability and may be difficult to detect.
- Evaluation attacks: A17 includes 2.4k training utterances from 10 target and 2 non-target speakers, yet was not more effective against seen than unseen targets.The overlapping speakers represent less than 2% of the evaluation speakers and simulate a worst-case training scenario.
4. Post Processing
ASVspoof 5 varies evaluation data through codec, compression, and bandwidth conditions while post-processing aims to suppress shortcut artefacts. The resulting distributions are substantially more aligned between bona fide and spoofed speech.
- Encoding and compression: Evaluation subsets apply encoding or compression conditions to both bona fide and spoofed utterances to test detection under bandwidth and channel variation.The evaluation set is treated to keep the database manageable while users choose augmentation strategies for training and development.
- Encoding and compression: Conditions C00-C07 use 16 kHz data, whereas C08-C11 use an 8 kHz narrow-band setting before upsampling to the common 16 kHz distribution rate.C01-C11 apply lossy encoding or compression; C04 and C07 use Encodec, with C07 adding prior MP3 compression.
- Shortcut artefacts: Post-processing targets five potential shortcut artefacts: peak amplitude, leading and trailing non-speech durations, utterance duration, and average energy.It is applied to development and evaluation sets but not the training set.
- Shortcut artefacts: The pipeline scales peak amplitude, randomly trims non-speech segments, and selects a random 4.0-to-10.0-second speech chunk.Trimming indices are sampled with higher probability at low-amplitude waveform locations.
- Shortcut artefacts: Post-processing substantially reduces shortcut discrepancies, producing near-identical distributions for spoofed/deepfake and bona fide utterances.The resulting overlap indicates that the five artefacts are unlikely to provide detection-relevant cues.
5. Visualisation
ASVspoof 5 visualisations use ASV speaker embeddings, t-SNE, and hierarchical clustering to expose relationships among bona fide and spoofed/deepfake utterances. They show greater overlap under encoding/compression and clustering patterns linked to shared generative components, while relying on oracle labels unavailable during detection.
- t-SNE visualisation: t-SNE represents ASVspoof 5 utterances as embedding points, with grey bona fide data and coloured points for 32 attacks.The visualisation samples 10% of utterances from training, development, and evaluation data, with encoding/compression variants shown for evaluation.
- t-SNE visualisation: Encoding/compression increases overlap between bona fide and spoofed/deepfake utterances, indicating a greater detection challenge.The comparison uses paired t-SNE plots with and without encoding/compression.
- t-SNE visualisation: Attacks A07, A11, and A14 fall outside the bona fide confidence contour, while A13 forms a low-variance cluster suggesting limited inter-speaker variation.Compared with ASVspoof 2019, a greater proportion of attacks lie within the confidence contour.
- Hierarchical clustering: Attacks using the same generative technology cluster together, including Glow-TTS attacks A01–A03, Grad-TTS attacks A04–A06, and Toucan systems A09–A10.The dendrogram similarly groups A01–A03 and A04–A06, while the two groups have low inter-group similarity.
- Hierarchical clustering: Shared acoustic decoders, text encoders, and vocoders may explain within-group similarity, whereas differing vocoders produce more substantial differences than differing prosody predictors.The visualisations are derived from pairwise cosine similarities between ASV speaker embeddings and use oracle speaker and attack labels.
6. TTS and VC system optimisation using surrogate models
The paper optimises selected TTS and VC attacks against surrogate ASV and countermeasure models, measuring progress with equal error rates. The results show that surrogate models can produce stronger attacks, although adoption was limited by optimisation cost and uncertainty about modifying quality-oriented systems.
- Surrogate models: Surrogate models comprise one ECAPA-TDNN ASV system and three CM systems used to optimise selected TTS and VC attacks.The ASV surrogate uses cosine scoring and is trained with VoxCeleb data plus MUSAN noise augmentation.
- Evaluation: Figure 8 measures optimisation progress through ASV and CM EERs across multiple rounds for attacks A10, A11, and A24.The four plots correspond to one ASV and three CM surrogate models.
- Evaluation: ASV EER uses one negative class, whereas a-DCF treats bona fide non-target and spoofed data as independent negative classes.MOS EER is computed like CM EER but uses a MOS estimator output.
- Optimisation procedure: Contributors selected checkpoints using the highest ASV EER, highest overall surrogate CM EER, or manually adjusted speaker-encoder training configurations.The selection strategy differed across A10, A11, and A24.
- Findings: Surrogate systems successfully implemented stronger attacks, but data-provider use remained modest because optimisation was costly and system modifications were difficult to determine.Most TTS and VC systems were designed to maximise quality-based measures rather than EERs.
7. Experimental validation
The validation experiments use ASV, CM, and SASV baselines to assess the ASVspoof 5 protocols and database. Results show substantial variation across attacks, degraded reliability on evaluation data and compression conditions, and persistent difficulty for automatic detection.
- Systems and validation: The open-sourced validation suite comprises one ASV system, two CM systems, and two SASV systems, rather than a baseline comparison.CM systems support Track 1, while SASV systems support Track 2 through fusion or end-to-end approaches.
- ASV results: 5.22% ASV EER on the evaluation set exceeds 1.88% on development, with encoding/compression raising EER from 2.20% to 6.00%.Compression reduces score-distribution differences and increases overlap.
- ASV results: 37.10% ASV EER makes development attack A12 the most disruptive, while 14 of 16 evaluation attacks produce EERs above 20%.Pre-trained TTS attacks A17, A28, and A29 each produce ASV EERs above 31%.
- ASV results: Adversarial attacks designed to compromise ASV increase EER over their underlying TTS/VC attacks, including A26 15.77% → A27 34.68%.Comparable increases are reported for A22→A31 and A25→A32; Malafide attacks are generally ineffective except A23.
- CM results: 29.49% and 17.83% pooled CM EERs on development rise to 36.04% and 29.12% on evaluation, showing limited reliability across attacks.The evaluation set contains more advanced algorithms and substantial encoded/compressed data.
- SASV results: 0.3156 and 0.2254 pooled development min a-DCF values rise to 0.6806 and 0.5741 on evaluation, while SASV baselines struggle with A12, A19, and A28.Lower a-DCF indicates better performance; values are higher under encoding/compression.
8. Conclusions
The paper presents ASVspoof 5 as a more diverse and challenging database for spoofing and deepfake detection. Validation shows that robust, generalisable detection remains difficult, especially for encoded and compressed data.
- ASVspoof 5 uses wild-collected source data with greater acoustic diversity and many more speakers than previous ASVspoof databases.
- The database introduces new challenges in generating spoofed and deepfake speech and in detecting those attacks.
- Validation demonstrates that robust, generalisable detection is especially difficult for data encoded and compressed with recent neural codecs.
- Legacy and contemporary generative technologies continue to threaten automatic-system reliability, motivating further progress in spoofed and deepfake detection.
- Most resources are freely available for challenge analysis, reproduction, and other spoofing or deepfake detection research.
Appendix A. Encoding and compression condition C11
Appendix A identifies the C11 encoding and compression condition and refers to its configuration table.
- Appendix A concerns encoding and compression condition C11.
- The appendix includes a table titled “C11 configurations.”
- The table presents configurations associated with condition C11.
Appendix B. Dendrogram of ASVspoof 5 data
The appendix constructs a pairwise similarity matrix for attack and bona fide speaker-embedding collections, then uses it for hierarchical clustering.
- Figure 7 uses a 33×33 pairwise similarity matrix across attack and bona fide utterance collections.
- Hierarchical agglomerative clustering uses Ward’s minimum variance criterion as its objective.
Appendix C. DET Curves
Figure C.10 presents DET curves for ASV and countermeasure baselines across evaluation-set attacks, distinguishing non-adversarial from adversarial conditions.
- The figure shows DET curves for ASV and countermeasure baselines across different evaluation-set attacks.
- Solid lines represent zero-effort non-target-speaker trials and non-adversarial attacks.
- Dashed lines represent adversarial attacks, including Malafide, Malacopula, and their combination.
- Dashed-line colours match the corresponding non-adversarial attack counterparts.
- The figure marks EER operating points for non-adversarial, Malafide, Malacopula, and combined attacks.
Appendix D. Results with and without encoding/compressiong
Table D.7 reports ASV and CM baseline performance using EER (%) and SASV baseline performance using a-DCF on evaluation set A17-A32, separated by encoding/compression condition.
- EER (%) measures ASV and CM baseline performance in Table D.7.
- a-DCF measures SASV baseline performance in Table D.7.
- The results are broken down into conditions with and without encoding/compression.The upper and bottom subtables report the two respective conditions.