Source-linked AI summary
The VoicePrivacy 2020 Challenge: Results and findings
Natalia Tomashenko, Xin Wang, Emmanuel Vincent, Jose Patino, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
TL;DR
Speech contains extensive personal information, yet voice-anonymization protection lacked a formal task, attack model, and common evaluation framework. The paper reports the first VoicePrivacy Challenge, defining the task and evaluation framework, introducing baselines, and analyzing submitted systems. Results show partial anonymization with privacy–utility trade-offs: x-vector systems perform best objectively on average, while signal-processing systems tend to yield higher subjective naturalness and intelligibility.
Problem
Speech carries personal information and speaker identity, while voice-anonymization protection lacked a formal task, attack model, common datasets, protocols, and metrics.
Method
The paper defines a voice-anonymization challenge with attack models, datasets, protocols, privacy and utility metrics, open-source baselines, and objective and subjective evaluation of submitted systems.
Results
X-vector systems provide the best objective results on average, whereas signal-processing systems tend to yield higher subjective naturalness and intelligibility; anonymization remains partial and trades privacy against utility.
Takeaways & Limitations
No single system performs best across all metrics, and system rankings depend on the attack model and the privacy–utility trade-off.
Takeaways & Limitations
The first challenge edition focuses on biometric identity and speech-recognition operability, while other sensitive information and authorship cues remain open concerns.
Abstract
from arXiv · showhide
This paper presents the results and analyses stemming from the first VoicePrivacy 2020 Challenge which focuses on developing anonymization solutions for speech technology. We provide a systematic overview of the challenge design with an analysis of submitted systems and evaluation results. In particular, we describe the voice anonymization task and datasets used for system development and evaluation. Also, we present different attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and provide a summary description of the anonymization systems developed by the challenge participants. We report objective and subjective evaluation results for baseline and submitted systems. In addition, we present experimental results for alternative privacy metrics and attack models developed as a part of the post-evaluation analysis. Finally, we summarize our insights and observations that will influence the design of the next VoicePrivacy challenge edition and some directions for future voice anonymization research.
1. Introduction
The VoicePrivacy 2020 Challenge addresses the unclear protection offered by speech anonymization by establishing a common task, attack models, datasets, protocols, and metrics. It focuses first on altering speaker identity while preserving other speech attributes, alongside reviewing existing approaches and introducing baseline systems.
- The challenge is motivated by the personal information carried by speech, including demographic, geographic, health, emotional, political, religious, and identity-related information.Speaker recognition systems can reveal the speaker’s identity, motivating the VoicePrivacy initiative.
- Existing speech privacy approaches include obfuscation, encryption, distributed learning, and anonymization, with differing effects on recoverability, computational complexity, and downstream use.Anonymization is presented as compatible with supervised machine learning and existing systems.
- Voice anonymization alters the speaker’s voice to hide identity while preserving traits, states, and spoken contents.The broader anonymization goal may also require changing other potentially identifying speech characteristics, but the first challenge focuses on the speaker’s voice.
- The challenge addresses the lack of a formal task definition, attack model, common datasets, protocols, and metrics for evaluating speech anonymization.It was designed to make privacy protection results comparable across anonymization solutions.
- The paper presents challenge design, datasets, attack models, objective and subjective evaluation methods, baseline and submitted systems, and comparative results.The challenge also introduces noise addition, speech transformation, voice conversion, and disentangled representation learning as approaches related to voice anonymization.
2. Challenge design
The challenge defines voice anonymization as hiding speaker identity while preserving other speech attributes and enabling downstream speech tasks. It evaluates systems under multiple attacker knowledge conditions using objective and subjective privacy and utility metrics.
- 2.1. Anonymization task and attack models: The privacy scenario models users sharing anonymized speech while attackers use available utterances, derived data, and prior knowledge to identify speakers.The scenario is specified by the data, protected information, downstream goals, attacker access, and attacker knowledge.
- 2.1. Anonymization task and attack models: Voice anonymization alters speaker identity while preserving traits, states, and spoken content, with consistent distinct pseudo-speakers across utterances.The task requires waveform output and supports ASR and multi-party conversation use cases.
- 2. Challenge design: The challenge combines privacy metrics with utility-oriented evaluation so anonymized speech can remain usable for automated processing and human conversation.The objective framework includes speaker-verification and speech-recognition assessment, while the subjective framework evaluates perceptual properties.
- 2.3.1. Objective metrics: Objective evaluation uses ASVeval for speaker verifiability and ASReval for linguistic-information preservation, with additional models trained on anonymized speech for post-evaluation analysis.ASVeval outputs log-likelihood-ratio scores, while ASReval outputs word sequences.
- 2.3.1. Objective metrics: Attack models vary by whether enrollment and trial utterances are original or anonymized, including ignorant, lazy-informed, and stronger semi-informed conditions.The semi-informed post-evaluation attack retrains automatic speaker verification using anonymized training data.
3. Anonymization systems
The challenge provided two complementary anonymization baselines and attracted diverse participant systems that modified x-vector anonymization, feature extraction, and speech synthesis components.
- 3.1. Baseline systems: B1 offers flexible pseudo-speaker selection and strong objective privacy and utility but demands substantial development effort and computational resources, whereas B2 is simpler with better subjective naturalness and intelligibility but weaker privacy.The baselines were designed to help participants explore a wide range of solutions for the new task.
- 3.1. Baseline systems: The primary baseline B1 combines x-vector, pitch, and bottleneck-feature extraction, x-vector anonymization, and neural speech synthesis using the anonymized x-vector with original F0 and BN features.Its x-vector and bottleneck extractors use TDNN-based models trained on VoxCeleb and LibriSpeech data.
- 3.1. Baseline systems: The secondary baseline B2 requires no training data and anonymizes speech by shifting LPC pole angles with the McAdams coefficient before time-domain resynthesis.Unlike frequency-warping approaches, it modifies the spectral envelope but not the pitch.
- 3.2. Challenge submissions: The challenge attracted 45 participants from 13 countries organized into 25 teams, with 16 successful eligible submissions summarized alongside their organizations and system identifiers.Each team could submit up to five systems and had to designate one primary system.
- 3.2. Challenge submissions: Participant systems explored distribution-preserving or dissimilarity-constrained x-vector generation, adversarial training, singular-value modification, component decomposition, and alternative acoustic models.Teams also examined pitch extraction and speech-synthesis or neural source-filter components, while some systems retained the baseline x-vector extractor.
4. Results
The paper reports evaluation results for the described systems, combining results from the official challenge with post-evaluation analyses.
- 4. Results: Evaluation results for the described systems include both official challenge results and post-evaluation analysis results, presented without distinction.The results section follows the system descriptions and reports their evaluation outcomes.
4.1. Objective evaluation results
Objective evaluations show that anonymization can raise speaker-verification error, but protection depends strongly on the attack model and involves utility costs.
- 4.1.1. Privacy: objective speaker verifiability: 53.37% EER was achieved by M1c1 versus 22.56% for M1c4 against ignorant attackers, with K2, A*, M1c1, M1, and B1 exceeding 50%.Higher EER indicates better privacy; the cited systems fully met the anonymization requirement against ignorant attackers.
- 4.1.1. Privacy: objective speaker verifiability: No single system was best in both attack conditions: A1, A2, M1, M1c1, and K2 outperformed B1 for ignorant attacks, whereas S2, S2c1, O1, and O1c1 did so for lazy-informed attacks.The differing winners demonstrate the difficulty of optimizing anonymization across attack scenarios.
- 4.1.2. Utility: speech recognition error: All anonymization systems increased WER, with relative increases of 40–217% on LibriSpeech-test and 14–120% on VCTK-test.The best post-anonymization LibriSpeech WER was 5.83% for I1, while x-vector systems related to B1 performed better on average across both datasets.
- 4.1.3. Using anonymized speech data to assess privacy: Retraining the ASV evaluator on anonymized data substantially decreased EER, making measured protection closer to—but still better than—original unprotected speech.Using an evaluator trained only on original speech therefore overstates privacy protection.
- 4.1.6. Relation between privacy and utility metrics: No system jointly maximized privacy and minimized WER across datasets: x-vector systems provided the best LibriSpeech privacy, whereas I1 yielded the lowest LibriSpeech WER.On VCTK-test, x-vector systems achieved better results for both metrics than I1.
4.2. Subjective evaluation results
Subjective evaluations indicate that anonymization reduces perceived speaker similarity but also degrades naturalness and intelligibility, with measurable effects on human linkability judgments.
- 4.2.1. Naturalness, intelligibility, and speaker verifiability: Anonymized samples were significantly inferior to original data in naturalness and intelligibility, with differences at p ≪0.01; I1 ranked highest among anonymization systems.The gap appeared for systems based on both B1 and B2.
- 4.2.1. Naturalness, intelligibility, and speaker verifiability: Anonymized trials were perceptually much less similar to the original enrollment voice than original trials, indicating effective human-perceived anonymization across systems.The subjective similarity evaluation separated same-speaker and different-speaker pairs.
- 4.2.2. Naturalness, intelligibility, and verifiability DET curves: DET curves placed anonymized speaker-similarity results near the top-right, making same- versus different-speaker decisions difficult after trial anonymization.For naturalness and intelligibility, anonymized curves instead indicated inferiority to original speech.
- 4.2.2. Naturalness, intelligibility, and verifiability DET curves: All systems concealed perceived speaker identity to some degree, but none matched original speech in naturalness and intelligibility.I1 degraded these qualities less severely than other systems, though it remained below original speech.
- 4.2.3. Speaker linkability: Non-native English evaluators showed a larger mean F1 difference than native evaluators, 0.26±0.02 versus 0.19±0.02, while B1 exceeded B2, 0.24±0.02 versus 0.21±0.02.The evaluator’s native language affected F1, but anonymization system and original speaker gender did not.
- 4.2.3. Speaker linkability: Original speech yielded 86.40% average clustering purity, compared with 61.68% and 62.58% for the anonymized panels.The original and anonymized purity distributions differed significantly for both female and male speakers, with p < 0.001.
4.3. Comparison of objective and subjective evaluation results
Objective and subjective measures agree that anonymization improves privacy while reducing utility, although the relationship varies across systems and datasets.
- 4.3. Comparison of objective and subjective evaluation results: Anonymizing trial speech increased objective EER and decreased same-speaker subjective similarity, while different-speaker similarity remained roughly unchanged.The precise impact depended on the anonymization system and test set.
- 4.3. Comparison of objective and subjective evaluation results: All systems degraded both objective and subjective utility measures, with I1 best on LibriSpeech-test and no single utility winner across VCTK-test metrics.WER was less consistent across datasets than subjective naturalness and intelligibility.
5. Conclusions
The first VoicePrivacy Challenge established a shared framework for evaluating voice anonymization and showed that privacy protection remains partial and traded against utility. Its findings motivate stronger systems, attack models, metrics, datasets, and broader privacy goals.
- 5.1. Summary and findings: X-vector systems provide the best objective results on average, while signal-processing systems tend to achieve higher subjective naturalness and intelligibility.The comparison covers anonymization systems evaluated on objective and subjective criteria.
- 5.1. Summary and findings: All systems degrade naturalness, intelligibility, and WER; x-vector systems achieve the best WER, whereas system I1 achieves the best intelligibility.The paper notes exceptions related to system I1 and WER on LibriSpeech.
- 5.1. Summary and findings: Anonymization is partial and costly: no system performs best across all privacy and utility metrics, regardless of attack model.The paper reports different privacy–utility trade-offs across systems under objective and subjective evaluation.
- 5.1. Summary and findings: Participants identified privacy leakage in x-vector embeddings, phonetic features, and pitch, while retraining downstream models on anonymized data can mitigate utility degradation.They also observed performance differences across datasets and speaker gender and proposed improvements over baselines in some test cases.
- 5.2. Open questions and future directions: Future challenge editions will strengthen baselines and attack models, revisit privacy and utility metrics, and add datasets or tasks for sensitive attributes beyond speaker identity.Proposed extensions include selective suppression of emotional state, age, gender, and accent.
- 5.2. Open questions and future directions: VoicePrivacy research should seek integrated designs that address privacy and security together while balancing privacy loss against utility gain.The paper frames privacy–utility threshold selection and joint optimization as open design questions.