Source-linked AI summary

ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild

Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas Evans, Andreas Nautsch, Kong Aik Lee

arXiv:2210.02437v3cs.SDcs.CRcs.MMeess.AS

TL;DR

The paper addresses whether spoofing and deepfake countermeasures generalize beyond clean laboratory data to realistic transmission, physical, and compressed-audio conditions. It summarizes the ASVspoof 2021 challenge across three tasks and finds modest robustness to transmission and compression, but weaker generalization across real acoustic environments and source datasets.

  • Problem

    Prior ASVspoof data was relatively free of real-world distortions and often derived from the same source corpus, limiting evidence about robustness and cross-domain generalization.

  • Method

    The paper analyzes ASVspoof 2021 datasets, three task-specific evaluations, 54 teams' results, hidden subsets, data factors, system descriptions, and post-challenge studies.

  • Results

    Across tasks, logical-access systems showed modest degradation from real telephony transmission, deepfake systems showed modest compression impacts but poor source-data generalization, and physical access was most challenging.

  • Takeaways & Limitations

    The findings support continued benchmarking under more realistic conditions, with data augmentation common among top logical-access and deepfake systems.

  • Takeaways & Limitations

    ASVspoof 2021 still lacked background noise in LA, human-talker recordings in PA, and real social-media data in DF.

Abstract

from arXiv · show

Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof.

I. INTRODUCTION

ASVspoof evolved from clean laboratory benchmarks toward realistic spoofing conditions, culminating in the three-task ASVspoof 2021 challenge. The paper describes its datasets, results, analyses, and future directions.

  • Motivation: ASV systems are vulnerable to impersonation, voice conversion, text-to-speech synthesis, and replay attacks.Voice conversion, text-to-speech, and replay can be mounted with consumer devices and readily available software.
  • Prior challenges: ASVspoof progressed from voice-conversion and text-to-speech detection in 2015, to replay detection in 2017, and combined logical- and physical-access tasks in 2019.The 2019 logical-access task covered VC and TTS, while physical access covered replay attacks in simulated acoustic environments.
  • Motivation: Earlier challenge data largely lacked noise, encoding, compression, and transmission artifacts, limiting alignment with real-world application conditions.The paper notes that prolonged reliance on clean data may promote countermeasures that fail to generalize in the wild.
  • ASVspoof 2021: ASVspoof 2021 comprised separate logical-access, physical-access, and deepfake sub-challenges that exposed countermeasures to more realistic conditions.The tasks covered telephony coding and transmission, acoustic propagation in physical spaces, and manipulated speech intended for online use.
  • Paper scope: The article analyzes dataset design, challenge results, influential data factors, hidden subsets, participating systems, post-challenge studies, limitations, and future research directions.It also reports newly released metadata for the three task databases.

A. Logical Access (LA)

The supplied passages describe how ASVspoof 2021 constructs evaluation data and conditions to test generalization across transmission channels, physical environments, replay devices, and compressed speech.

  • Logical Access (LA): The 2021 logical-access task transmitted bona fide and spoofed speech through real VoIP and PSTN systems to test robustness to channel variation.The task focused on compression, packet loss, bandwidth, transmission infrastructure, and bitrate effects.
  • Logical Access (LA): Balanced speakers and spoofing attacks across LA conditions allow detection differences to be attributed reliably to encoding and transmission variation.Training used earlier clean ASVspoof 2019 data, while evaluation introduced unknown channel conditions.
  • Physical Access (PA): The physical-access design evaluates replay detection across real rooms, microphones, distances, attacker recording conditions, and replay devices.The design covers 162 acoustic environments and uses a non-exhaustive policy because all 1,458 factor combinations are too numerous.
  • Deepfake (DF): The DF task evaluates bona fide and spoofed speech after lossy codec processing, including conditions designed to study transcoding distortions.Its application scope includes compressed audio used in television, news websites, and social media platforms.
  • Physical Access (PA): PA evaluation reserves selected replay-factor combinations exclusively for evaluation to test generalization beyond the progress subset.The resulting progress and evaluation sets are utterance-disjoint and include the full speaker set.

C. Deepfake (DF)

The DF task evaluates detection of bona fide and spoofed speech after lossy media-codec processing, while testing generalization across codecs, source datasets, and spoofing attacks.

  • Task design: DF evaluation data uses bona fide and spoofed utterances processed with lossy codecs, and EER is the task’s primary metric.Audio is encoded and decoded to recover uncompressed samples, introducing codec-dependent distortions.
  • Task design: Two additional source datasets, VCC 2018 and VCC 2020, extend the ASVspoof 2019 LA source data used in the evaluation.The added databases contain 26 additional speakers and spoofed utterances generated by previously unused VC algorithms.
  • Task design: The evaluation database contains more than 100 spoofing attack algorithms and analyzes performance by broad vocoder categories.The paper focuses on vocoder categories rather than individual VC systems because many approaches share similar vocoders.
  • Task design: Nine evaluation conditions vary codec type and variable bit rate, including no codec, mp3, m4a, ogg, and dual-codec processing.Conditions C2-C7 use paired lower- and higher-VBR configurations for mp3, m4a, and ogg; C8-C9 apply two codecs successively.

III. ASVSPOOF 2021 CHALLENGE RESULTS

ASVspoof 2021 evaluates countermeasures across LA, PA, and DF using task-specific primary metrics and compares progress, evaluation, and baseline performance. Results show strong LA performance under transmission variation, while analyses expose condition- and attack-dependent effects.

  • Full challenge results: The three tasks use min t-DCF as the primary metric for LA and PA, but EER as the primary metric for DF.EER is also shown for all tasks, while the ASV floor provides a reference for LA and PA.
  • Full challenge results: LA systems generally show modest progress-to-evaluation gaps, and top systems approach the ASV floor in t-DCF.Some systems overfit known conditions C1-C4 and perform worse on unknown evaluation conditions C5-C7.
  • Logical access: Wideband LA conditions yield lower t-DCF distributions than narrowband conditions, indicating the importance of higher-frequency information.The comparison covers no codec, G.722, and OPUS versus a-law, PSTN, µ-law, and GSM conditions.
  • Logical access: Transmission routes across LAN, France–Italy, and France–Singapore have similar t-DCF distributions and little apparent impact on countermeasure performance.The analysis suggests that future challenges could use simpler LAN routes only.
  • Logical access: For attacks A18 and A17, median t-DCF varies substantially across conditions, from approximately 0.4 in C1 to over 0.8 in C3 for A18.Figure 5 encodes median t-DCF by color and inter-quartile range by circle radius.
  • Top-performing systems: Top-performing systems commonly use data augmentation, ensembles, short-term spectral or raw-waveform inputs, and ResNet-family classifiers.Fusion commonly uses weighted averaging with uniform or empirically selected weights.

B. Physical Access (PA)

PA systems can detect replay attacks in real physical spaces, but performance remains substantially below the ASV floor and is sensitive to attacker and recording factors. DF analyses further show codec- and source-dataset-dependent difficulty.

  • Physical access: Almost half of PA systems outperform baseline B01, yet every system remains substantially worse than the ASV floor of 0.12.Progress and evaluation performance differ only modestly, suggesting stable but high error rates across rooms and devices.
  • Physical access: The PA difficulty may reflect differences between simulated replay training data and real-space evaluation data, including room acoustics and noise conditions.The result indicates potential detection in real spaces but limited transfer from simulated environments.
  • Physical access: Higher-quality attacker microphones and replay devices, together with shorter attacker-to-talker distances, produce higher min t-DCF values.The closest attacker-to-talker position, Da = c4, has the highest min t-DCF values.
  • Top-performing systems: PA top systems universally use ensembles, with varied classifiers and relatively little variation in front-end features.The top-1 system combines frame-level and temporal-level features with a parallel VAE-based architecture.
  • Deepfake: DF evaluation EERs exceed 15% for all systems, whereas 23 of 33 systems achieve below 10% on the progress subset.The best progress-subset system has EER below 1%, but 18 systems still outperform baseline B04 on evaluation data.
  • Deepfake: DF EERs are substantially higher for VCC source datasets than for the ASVspoof 2019 LA database, indicating limited generalization to mismatched source data.The paper attributes this pattern likely to use of the ASVspoof 2019 LA database only in the progress set and consequent over-fitting.
  • Deepfake: For a given codec, neural vocoders produce higher EERs than traditional vocoders or waveform concatenation.The comparison pools VCC 2018 and VCC 2020 source data.

3) Top-performing systems:

Top systems rely heavily on augmentation and fusion, while hidden-subset analyses test whether performance depends on non-speech segments or simulated replay artifacts. These analyses reveal important evaluation-scope boundaries.

  • Top-performing systems: DF top submissions nearly all use ensemble systems, all use data augmentation, and all employ some form of media-codec augmentation.Their acoustic features are diverse, classifiers range from CNNs to MLPs and GMMs, and all use score averaging for fusion.
  • Hidden subsets: Hidden subsets remove non-speech segments or withhold simulated replay data to measure dependence on data characteristics absent from challenge rankings.Participants were unaware of hidden-subset inclusion, and hidden results were excluded from challenge rankings.
  • The role of non-speech: Removing non-speech segments makes PA performance notably worse, although non-speech content may provide legitimate spoofing cues that an adversary can easily remove.The PA hidden subset controls other factors by using one talker-to-ASV distance.
  • Real versus simulated replay: Simulated PA data has a much higher median min t-DCF than real evaluation data, with four of the top ten systems exceeding 0.99.The gap suggests simulated data omits useful artifacts from real recording and replay environments.
  • Real versus simulated replay: Simulated evaluation data may misestimate real-space performance when it fails to reproduce room acoustics and device frequency responses.The paper specifically cautions against using such simulations for real-environment estimates unless these characteristics are faithfully reflected.

C. Performance gap for DF progress and evaluation subsets

The DF evaluation gap is linked to previously unexposed VCC source corpora and shifts in bona fide score distributions, indicating poor cross-source generalization. Analyses implicate source-specific speech and silence characteristics in this mismatch.

  • The DF progress–evaluation performance gap relates partly to two previously unexposed VCC source corpora.The paper examines class- and source-conditional score distributions to investigate this gap.
  • Bona fide scores for the two VCC corpora are consistently lower than scores for ASVspoof 2019 LA data, increasing bona fide–spoof overlap.Spoofed-score distributions are reasonably aligned across sources, whereas bona fide distributions differ substantially.
  • Training only on ASVspoof 2019 LA data causes over-fitting and poor generalization to VCC evaluation data.The resulting shift produces greater confusion between bona fide and spoofed trials and degraded detection performance.
  • EERs are approximately 15%–34% when either VCC subset is included on the bona fide side, but below 1% for T23 when both are excluded.The lower EERs persist regardless of whether spoof trials include VCC 2018 or VCC 2020 data.
  • Bona fide silence distributions are separated between ASVspoof 2019 and VCC data, suggesting that models may learn source-specific nonspeech cues.The paper links this cue to failures when unseen bona fide data has a different silence distribution.
  • Post-challenge work emphasizes end-to-end architectures, self-supervised front-ends, and data augmentation, while noting protocol constraints and fragility to unseen or lower-quality audio.Reported examples include SSL-based improvements and the emergence of the ADD challenge for more difficult audio conditions.

VI. LIMITATIONS AND FUTURE DIRECTIONS

The paper identifies limitations involving realism, attack and data diversity, evaluation scope, training policy, and scenario coverage. It proposes broader data, stronger adversaries, richer assessment, and integrated task designs as future directions.

  • Future evaluations should include background noise for LA, human-talkers recordings for PA, and real social-media data for DF.These additions are identified as remaining steps toward realistic audio conditions.
  • Non-speech intervals can influence spoof detection, but their database-dependent length should not become a detection cue.The paper calls for studying non-speech generation or conversion in TTS and VC attacks.
  • A dual training policy could retain a fixed protocol while permitting a relaxed policy for larger models trained with external data.Post-evaluation results show benefits from external-data training and semisupervised learning.
  • The 2021 LA and DF attacks used TTS and VC algorithms that were state of the art before 2020, motivating renewed attack collection.Future editions should explore vulnerabilities to more recent techniques.
  • ASVspoof data are English read speech from VCTK, while unexposed DF datasets revealed weak countermeasure generalization across bona fide source data.The paper recommends broader languages, environments, speaking styles, and source populations.
  • LA and PA evaluate countermeasures with a fixed ASV system, limiting study of jointly optimized ASV–CM architectures.The paper anticipates closer integration with SASV and corresponding metric development.
  • Future protocols could combine replay with LA conditions and VC/TTS with PA conditions, though this would increase protocol and data-collection complexity.The paper identifies integrated scenarios as an important future direction.
  • ASVspoof 2021 does not include partially spoofed utterances, although short manipulated segments can be harder to detect and may alter phrase meaning.The paper suggests segment-level bona fide/spoof labeling could help explain classifier decisions.

VII. CONCLUSIONS

ASVspoof 2021 benchmarked spoofing and deepfake detection under more realistic conditions, attracting 54 teams across three tasks. LA and DF showed modest degradation from transmission or compression, whereas PA was most challenging because of training–evaluation mismatch.

  • 54 teams submitted valid challenge scores across the three ASVspoof 2021 tasks, and 34 submitted system descriptions.The paper analyzes the challenge datasets, results, and data-related issues.
  • Real telephony transmission causes only modest degradation in LA spoofing detection, with LAN and geographically distant endpoint estimates similarly reliable.This conclusion covers transmission across real telephony systems and different endpoint distances.
  • Compression effects cause modest DF performance impacts, but detection lacks generalization to different source data.Data augmentation is common among top-performing LA and DF systems.
  • PA appears to be the most challenging task, likely because training uses simulated replay while evaluation uses mismatched real acoustic conditions.Difficulty also increases with higher-quality replay and recording devices or a lower-quality ASV microphone.
  • Future editions will use stronger adversaries, larger and more varied datasets, broader speaker populations, relaxed training policies, and joint ASV–CM evaluation.The roadmap includes merging ASVspoof with SASV and exploring alternative combination architectures.

A. Logical Access (LA)

The illustrated LA systems combine diverse acoustic frontends, neural or statistical classifiers, augmentation, and score fusion. Their architectures are presented through standardized processing blocks and ranked by minimum t-DCF.

  • T23 combines codec-augmented trimmed audio, spectral and SincNet frontends, LightCNN, ResNet, LSTM, and weighted score fusion.Its subsystem outputs are summed or fused using weights.
  • T35 applies pre-emphasis and alaw companding, extracts LFCCs, and averages outputs from two ResNet classifiers.The system uses two parallel ResNet-based classifiers.
  • T19 augments training with RIR, MUSAN, trimming, and additive noise before classifying mel spectrograms with a squeeze-and-excitation ResNet.Its output scores are averaged using predefined empirical weights.
  • T36 combines RawNet and ECAPA-TDNN classifiers using raw waveform, mel-spectrogram, and learnable LEAF features.Training includes additive noise and three-fold speed perturbation at 0.9, 1.0, and 1.1.

C. DeepFake (DF)

The DF analysis examines top-system designs and how hidden conditions, especially non-speech, affect spoofing detection. Hidden-track performance deteriorates substantially, while utterance-boundary non-speech carries more information than intervals between words.

  • Top-performing systems: Top DF systems combine diverse front-end features and classifiers, with all top-five systems using score averaging for fusion.The reported classifiers include convolutional networks, MLPs, and GMMs.
  • Hidden-track results: More than 20% EER was observed for top DF systems on the hidden track, compared with under 10% on the main track.The hidden track tests dependence on data characteristics including non-speech.
  • Role of non-speech: Removing non-speech at utterance boundaries makes LA and DF detection much more difficult than removing intervals between words.For LA B04, min t-DCF rose from 0.426 with non-speech to 0.908 after trimming the two ends, versus 0.471 after trimming between words.
  • Role of non-speech: On PA, non-speech at both utterance ends also degrades EER and min t-DCF, but differences across data sets are smaller than for LA and DF.The hidden-track construction removes all non-speech intervals, including leading, trailing, and between-speech regions.

C. Additional analysis on PA simulated hidden track

The PA simulated hidden-track analysis compares performance across simulated rooms, devices, and other factors. Most countermeasures perform worse than on real recorded data, consistent with limitations in impulse-response simulation.

  • Performance comparison: Most countermeasures perform worse on the simulated hidden track than on the real recorded evaluation data.Median min t-DCF is more similar across simulated rooms, devices, and other factors than across factors in the evaluation data.
  • Simulation limitations: Estimated impulse responses may not accurately represent room acoustics and device frequency responses across different conditions.The responses were estimated from sweep signals recorded by devices in rooms used for real-data collection.
  • Simulation limitations: Simulation may lose artifacts that discriminate real recorded bona fide speech from real replayed spoofed speech.Such losses may contribute to higher min t-DCF medians across simulated factors.
  • System variability: T16 and T01 rose to around 1.0 min t-DCF under most simulated conditions, whereas T07 and T16 degraded less and T04 improved.The comparison uses top-five systems from the real evaluation set and the simulated hidden track.

XI. ANALYSIS ON PRACTICALITY OF COMMON TECHNIQUES

The practicality analysis tests front ends, classifiers, losses, and augmentation under multiple ASVspoof conditions. Results show dataset-dependent effects: spectrograms can help, loss choices vary, and augmentation yields only modest post-hoc gains.

  • Feature extractors: Switching from LFCC to spectrogram degrades 2019 LA performance, while M02 achieves the best performance on both 2021 datasets among tested systems.The comparison is between M01 and M03 for the front-end change, with M02 serving as the stronger 2021 system.
  • Data augmentation: All tested augmentations improve 2019 LA EER and 2021 LA minimum t-DCF, but 2021 DF performs best without augmentation.The reported augmentations include room impulse response, MUSAN noise, and mp3 compression.
  • Data augmentation: Post-hoc analysis finds only modest augmentation improvements, likely because participants’ augmentation implementations differ.The paper recommends sharing augmentation-pipeline details to improve reproducibility in future editions.
  • Evaluation scope: The paper reports full progress and evaluation results for LA, PA, and DF, using t-DCF and EER for LA and PA and EER for DF.These results are presented in Tables XIV–XVI.
Loading 2210.02437v3…