Source-linked AI summary
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
Xin Wang, Hector Delgado, Hemlata Tak, Jee-weon Jung, Hye-jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu, Md Sahidullah, Tomi Kinnunen, Nicholas Evans, Kong Aik Lee, Junichi Yamagishi
TL;DR
Speech spoofing and deepfake detection must be evaluated under broader, more realistic attacks and acoustic conditions. ASVspoof 5 addresses this with crowdsourced data, two tracks, surrogate-optimized and adversarial attacks, and new metrics; submissions generally outperform baselines, although calibration remains weak in top systems.
Problem
Existing spoofing challenges need detection evaluation across more diverse crowdsourced speech, stronger attacks, adversarial attacks, and both stand-alone detection and spoofing-robust speaker verification.
Method
ASVspoof 5 defines two challenge tracks, introduces a crowdsourced database with surrogate-optimized and adversarial attacks, and adopts track-specific evaluation metrics.
Results
Most challenge submissions outperform the baselines, while spoofing attacks substantially compromise baseline systems.
Takeaways & Limitations
The challenge provides a broader benchmark for stand-alone spoofing detection and spoofing-robust automatic speaker verification under variable acoustic conditions.
Takeaways & Limitations
Top systems obtain actDCF values close or equal to 1.0 because their scores are normalized rather than calibrated as log-likelihood ratios.
Abstract
from arXiv · showhide
ASVspoof 5 is the fifth edition in a series of challenges that promote the study of speech spoofing and deepfake attacks, and the design of detection solutions. Compared to previous challenges, the ASVspoof 5 database is built from crowdsourced data collected from a vastly greater number of speakers in diverse acoustic conditions. Attacks, also crowdsourced, are generated and tested using surrogate detection models, while adversarial attacks are incorporated for the first time. New metrics support the evaluation of spoofing-robust automatic speaker verification (SASV) as well as stand-alone detection solutions, i.e., countermeasures without ASV. We describe the two challenge tracks, the new database, the evaluation metrics, baselines, and the evaluation platform, and present a summary of the results. Attacks significantly compromise the baseline systems, while submissions bring substantial improvements.
1. Introduction
ASVspoof 5 expands the challenge to crowdsourced, adversarially attacked speech and evaluates both stand-alone detection and spoofing-robust speaker verification with new metrics.
- Challenge tracks: ASVspoof 5 combines logical access and speech deepfake detection into two tracks: stand-alone countermeasures and spoofing-robust automatic speaker verification.Track 1 assumes attackers generate victim-like speech for scenarios such as social-media defamation; Track 2 models synthetic or converted speech injected into telephony.
- Database and attacks: The new database uses crowdsourced source data and attacks with greater acoustic variation, more speakers, contemporary TTS and VC, and first-time adversarial attacks.Attacks are optimized against surrogate ASV and countermeasure systems, rather than ASV alone.
- Challenge conditions: Both tracks introduce an open condition allowing external data and pre-trained speech foundation models, provided training data does not overlap evaluation data.This contrasts with the traditional closed condition, which restricts participants to the specified data protocol.
- Evaluation: Track 1 uses minDCF as its primary metric, while Track 2 uses architecture-agnostic DCF, with calibration and tandem metrics as complements.The metric suite is designed to assess stand-alone spoofing detection and spoofing-robust speaker verification.
- Paper scope: The paper describes the tracks, database, metrics, baselines, evaluation platform, and results from systems submitted by 54 challenge participants.Its scope covers both challenge design and comparative system performance.
2. Database
The ASVspoof 5 database broadens speaker and acoustic diversity, strengthens attacks through surrogate-model optimization, and evaluates systems under separated speakers, attacks, codecs, and participation conditions.
- Source data: The database uses MLS English data from more than 4k speakers recorded with diverse devices, replacing earlier VCTK data from around 100 speakers recorded in an anechoic chamber.This broadens evaluation beyond studio-quality source conditions.
- Attacks: Spoofing attacks use contemporary TTS and VC methods and are optimized to fool both ASV and countermeasure surrogate systems.Adversarial attacks apply Malafide and Malacopula filters; codecs are also applied to bona fide and spoofed data.
- Construction: The database is constructed from disjoint MLS partitions, with one contributor group building training TTS systems and another creating development and evaluation TTS and VC attacks.Surrogate ASV and CM systems gauge attack effectiveness during construction.
- Data separation: Training, development, and evaluation speakers are disjoint, and attacks across those sets are also disjoint.The evaluation set excludes speakers overlapping Librispeech because open-condition participants may use Librispeech-pretrained models.
- Evaluation conditions: Evaluation includes codecs and compression conditions spanning 16 kHz and 8 kHz narrow-band settings, while Track 2 omits some codec conditions.C0 represents unencoded data; narrow-band data are downsampled, codec-processed, and upsampled.
- Participation conditions: Closed-condition participants use the prescribed training and development sets, whereas open-condition participants may use external data and pre-trained foundation models without challenge-data overlap.The two tracks share the same evaluation utterances except for several codec conditions ignored by Track 2.
3. Performance measures
ASVspoof 5 replaces Track 1’s primary EER comparison with normalized detection costs and introduces a-DCF as Track 2’s primary SASV metric, alongside complementary calibration and tandem measures.
- 3.1. Track 1: from EER to DCF: The Track 1 DCF combines bona fide miss and spoof false-alarm rates as functions of the detection threshold.The specified costs and spoofing prior are Cmiss = 1, Cfa = 10, and πspf = 0.05, yielding β ≈1.90.
- 3.1. Track 1: from EER to DCF: Track 1 uses normalized DCF, with minDCF as its primary metric and actDCF evaluating performance at a fixed Bayes threshold.minDCF uses an oracle threshold, whereas actDCF uses τBayes = −log(β) and is meaningful when scores are calibrated log-likelihood ratios.
- 3.1. Track 1: from EER to DCF: Cllr complements Track 1’s DCF measures by assessing detection-score calibration and discrimination when scores are interpreted as log-likelihood ratios.Lower Cllr indicates better-calibrated and more discriminative scores; EER is also reported.
- 3.2. Track 2: from SASV-EER to a-DCF: Track 2’s primary metric is minimum architecture-agnostic DCF, or min a-DCF, for spoofing-robust speaker verification.The metric accounts for ASV misses and false alarms from both non-target speakers and spoofing attacks, using specified costs and priors.
- 3.2. Track 2: from SASV-EER to a-DCF: Track 2 additionally reports t-DCF and tandem EER for submissions with identifiable ASV and CM subsystems.t-EER measures the equal-error point across ASV misses and non-target and spoofing false alarms in a tandem architecture.
4. Common ASV, surrogate systems, and challenge baselines
The organisers evaluate common ASV, surrogate systems, and baseline countermeasures spanning standalone spoof detection and integrated or fused SASV architectures.
- 4.1. Common ASV system by organisers: The common ASV uses an ECAPA-TDNN encoder with cosine-similarity scoring, followed by s-norm score normalisation.It is trained on VoxCeleb 1 and 2; bona fide target-versus-non-target EER is 5%, while spoofing attacks produce much higher EERs.
- 4.2. Baseline systems: Track 1 baselines are RawNet2 (B01) and AASIST (B02), both end-to-end countermeasures operating directly on four-second raw waveforms.RawNet2 uses sinc filters, residual blocks, and gated recurrent units, while AASIST integrates spectro-temporal representations with graph attention.
- 4.2. Baseline systems: Track 2 baselines comprise a fusion system (B03) combining common ASV and AASIST scores, and an integrated MFA-Conformer system (B04) producing one SASV score.B04 is trained through speaker-classification pre-training, copy-synthesis training, and an adapted SASV loss.
- 4.2. Baseline systems: Figure 2 counts Track 1 and Track 2 submissions separately across the CodaLab progress and evaluation phases.The chart concerns submissions during the last three days of the evaluation phase as well as the progress phase.
- 4.2. Baseline systems: Surrogate evaluation systems include ECAPA-TDNN with PLDA scoring for ASV and AASIST, RawNet2, and LFCC-based LCNNs for countermeasures.The surrogate countermeasures do not encounter development or evaluation attacks during training.
5. Evaluation platform
ASVspoof 5 used CodaLab for score submission and result reporting across progress and evaluation phases, with different submission limits and evaluation coverage.
- 5. Evaluation platform: During the progress phase, participants could submit up to four times daily and receive results on a progress subset of the evaluation set.Participants could choose to place their results on an anonymised leaderboard.
- 5. Evaluation platform: The evaluation phase lasted only a few days and allowed participants to make a single submission.That submission was evaluated using the whole evaluation set.
- 5. Evaluation platform: Track 1 received comparable submission amounts in its closed and open conditions, whereas Track 2 received considerably more open-condition submissions.The reported pattern demonstrates the need for additional training data for SASV systems.
6. Challenge results
ASVspoof 5 baseline systems performed poorly on challenging Track 1 data, while most submissions improved substantially; Track 2 results likewise favored submitted systems, especially fused approaches.
- Track 1: Baseline Track 1 systems exceeded 0.7 minDCF and 29% EER despite using architectures effective on earlier challenge databases.The authors associate this poor performance with non-studio-quality MLS data and more advanced spoofing attacks.
- Track 1: Top-five closed-condition Track 1 submissions achieved minDCF below 0.5 and EER below 15%, about 50% relative improvement over baselines.Most closed-condition submissions outperformed the baselines in minDCF.
- Track 1: Open-condition Track 1 submissions generally obtained lower minDCF and EER than closed-condition systems, with many strong systems using wav2vec 2.0 features.The open condition allowed external data and pre-trained speech foundation models.
- Track 1: Top Track 1 systems had actDCF values close or equal to 1.0, while Cllr results indicated that score calibration remained weak.Outputs normalized between 0 and 1 were not calibrated to approximate log-likelihood ratios, although the primary minDCF metric is calibration-agnostic.
- Track 2: All top Track 2 submissions fused ASV and CM sub-systems, although baseline results do not establish that fusion is inherently superior.The integrated B04 system performed better than B03, whose CM sub-system provided no useful spoofing-detection information relative to a random-guessing reference.
- Track 2: Most Track 2 submissions outperformed baselines, with the top three closed-condition systems achieving 50% relative improvement on min a-DCF values.Open-condition submissions reached lower metrics, and SSL-based features were common among top submissions.
7. Conclusions
ASVspoof 5 expands challenge complexity through crowdsourced, variable-condition data, stronger attacks, adversarial attacks, and evaluation of both stand-alone detection and SASV. Submissions generally surpassed weak baselines, while score calibration emerged as an important deployment issue and further analyses were deferred to future work.
- Conclusions: ASVspoof 5 evaluates both stand-alone speech spoofing and deepfake detection and spoofing-robust speaker verification.It combines a new task structure with two challenge tracks.
- Conclusions: The challenge uses crowdsourced data collected under variable conditions, contemporary attacks optimized against surrogate ASV and CM systems, and new adversarial attacks.These changes make the fifth edition considerably more complex than its predecessors.
- Conclusions: Baseline detection performance was relatively poor, whereas most challenge submissions outperformed the baselines, sometimes by a substantial margin.The conclusion attributes this context to lower-quality data used to create spoofs and deepfakes.
- Conclusions: Score calibration is identified as an essential consideration for deploying detection solutions in practical scenarios.The results reveal calibration as an issue that had previously received little attention in this context.
- Conclusions: More detailed analyses were deferred to the ASVspoof 5 workshop and future work because of the particularly tight schedule.The paper explicitly limits the present analysis in this respect.