Source-linked AI summary
An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
Berrak Sisman, Junichi Yamagishi, Simon King, Haizhou Li
TL;DR
Voice conversion seeks to change speaker identity while preserving linguistic content, but effective manipulation requires coordinating speech representations, mapping, prosody, and waveform generation. This paper surveys the field from statistical modeling through deep learning, including evaluation methods, challenges, and resources. Reported challenge results improved from 3.0 naturalness and about 70% target-speaker judgments in 2016 to 4.1 and about 80% in 2018.
Problem
Voice conversion must modify speaker-dependent characteristics while preserving linguistic content, yet it remains far from perfect.
Method
The paper provides a comprehensive overview of voice conversion technologies, evaluation methods, challenges, and resources from statistical approaches to deep learning.
Results
2018 challenge systems achieved 4.1 naturalness and about 80% target-speaker judgments, compared with 3.0 and about 70% in 2016.
Takeaways & Limitations
Deep learning advances mapping, neural vocoding, and end-to-end voice conversion, while voice conversion challenges provide evidence of improved system performance.
Abstract
from arXiv · showhide
Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this paper, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research.
I. INTRODUCTION
Voice conversion manipulates speaker identity while preserving linguistic content, using a pipeline that has evolved from statistical modeling toward deep learning. The paper surveys these techniques, their limitations, evaluation challenges, voice conversion challenges, and research resources.
- Voice conversion changes speaker identity without changing linguistic content, making it a challenging speech-processing problem.
- Statistical approaches progressed from parallel-data spectrum mapping to sparse representations and methods supporting non-parallel training data.Sparse representation techniques require less training data than parametric methods and address over-smoothing; alignment methods extend conversion beyond parallel data.
- Deep learning improves mapping through nonlinear, data-trained models, advances neural vocoding, and supports research beyond the traditional parallel/non-parallel data paradigm.Neural vocoders learn waveform reconstruction from acoustic features and can be trained jointly with mapping or analysis modules.
- Voice conversion has expanded from a niche speech-synthesis topic into a major research area with applications including personalized synthesis, communication aids, de-identification, mimicry, disguise, and dubbing.
- The paper reports voice conversion challenges and summarizes publicly available resources for researchers and engineers.The discussion includes challenge activities and their relationship to anti-spoofing speaker-verification studies.
II. TYPICAL FLOW OF VOICE CONVERSION
A typical voice conversion system analyzes source speech, maps its features toward a target speaker, and reconstructs an audible waveform. The pipeline enables manipulation but remains constrained by representation and reconstruction artifacts, vocoder assumptions, and data requirements.
- Voice conversion modifies speaker-dependent characteristics while carrying over speaker-independent speech content.Relevant characteristics include formants, F0, intonation, intensity, and duration.
- The analysis-mapping-reconstruction pipeline decomposes source speech into features, transforms them toward the target speaker, and re-synthesizes time-domain speech.The mapping function is typically applied to an intermediate speech representation.
- A. Speech Analysis and Reconstruction: Speech analysis derives an intermediate representation, while reconstruction operates on modified parameters to generate an audible speech signal.Examples include vocoder-based reconstruction and Griffin-Lim reconstruction from modified short-time Fourier transforms.
- A. Speech Analysis and Reconstruction: Model-based representations express speech frames with time-varying model parameters, whereas signal-based representations use controllable time- or frequency-domain elements without assuming a model.
- A. Speech Analysis and Reconstruction: Analysis and reconstruction inevitably introduce artifacts, and traditional parametric vocoders can produce robotic or buzzy speech under simplified assumptions.These problems become more serious when F0 and formant structure are modified together; neural vocoding offers a data-driven alternative but requires substantial training data.
3) WaveNet Vocoder:
WaveNet vocoding models waveform distributions conditioned on speech features and has been widely used for high-quality reconstruction. The section also situates neural vocoding among faster alternatives and the feature representations used in conversion.
- 3) WaveNet Vocoder:: WaveNet factorizes the joint probability of waveform samples into a product of conditional probabilities.With auxiliary features h, it models the conditional distribution p(x|h).
- 3) WaveNet Vocoder:: WaveNet vocoders generally reconstruct waveforms from intermediate speech representations rather than performing speech analysis.They learn relationships between input features and waveform output, including interactions among input features.
- 3) WaveNet Vocoder:: 100 times faster than WaveNet vocoder, neural source-filter waveform modeling reportedly achieves comparable voice quality on a large speech corpus.
- 3) WaveNet Vocoder:: Parallel WaveGAN generates high-fidelity speech with a compact non-autoregressive architecture trained using multi-resolution spectrogram and adversarial losses.The passage notes that generating coherent raw audio waveforms with GANs remains challenging.
- 3) WaveNet Vocoder:: Voice conversion features include spectral and prosodic components, but vocoding parameters may require further transformation because they are not always ideal for voice-identity conversion.Prosodic features can separate speaker-dependent from speaker-independent parameters during mapping.
C. Feature Mapping
Feature mapping modifies source-speaker acoustic features toward a target speaker within the analysis-mapping-reconstruction pipeline. Traditional approaches commonly learn mappings from parallel utterances, using alignment and statistical models to transform speech features at run time.
- Feature mapping: Feature mapping changes source timbre and prosody features toward the target speaker, with spectral mapping remaining central in voice conversion studies.The broader pipeline analyzes speech, maps acoustic features, and reconstructs time-domain speech.
- Parallel-data modeling: Traditional voice conversion commonly trains a mapping function on paired source and target utterances containing the same linguistic content.Parallel data provide paired speech vectors after frame alignment.
- Statistical approaches: Statistical approaches include parametric models such as GMM and non-parametric methods such as vector quantization.Vector quantization maps each source codeword to a corresponding target codeword, while parametric models represent feature relationships statistically.
- Parallel-data modeling: Dynamic time warping commonly aligns parallel source and target speech vectors before statistical mapping.Speech recognizers with phonetic knowledge can also perform model-based alignment.
- Statistical approaches: Maximum-likelihood trajectory modeling and JD-GMM address framewise conversion and over-smoothing by modeling dynamic spectral features and their variances.JD-GMM parameters are estimated with expectation-maximization and can be robust when training data are limited.
- Statistical approaches: A modulation-spectrum post-filter compensates global variance and is useful for reducing the inherent over-smoothing of statistical mapping.The GMM approach is described as a successful parametric solution for parallel training data.
B. Dynamic Kernel Partial Least Squares
Dynamic and exemplar-based feature-mapping methods address limitations of local statistical mappings, especially temporal discontinuities, over-smoothing, and lost spectral detail. DKPLS models nonlinear temporal structure, while sparse representations use spectrogram exemplars and constrained activations.
- Dynamic mapping: Local framewise mappings can create temporal discontinuities because each speech frame is transformed independently of neighboring frames.This motivates methods that incorporate temporal dependencies.
- Dynamic mapping: DKPLS concatenates adjacent frames and applies a kernel transformation to model nonlinear relationships and speech dynamics.The method was reported to outperform GMM in voice quality.
- Dynamic mapping: Simple F0 shifting and scaling is insufficient for high-quality prosody conversion, motivating hierarchical modeling across linguistic units and temporal scales.Wavelet-based DKPLS provides a platform for multi-scale prosody conversion.
- Over-smoothing: Parametric methods often over-smooth acoustic features because mean-square-error or maximum-likelihood objectives favor statistical averages.The resulting speech fails to capture desired temporal and spectral dynamics.
- Over-smoothing: Low-dimensional features such as MCEP and LSF lose spectral detail, and together with statistical averaging can produce muffled output speech.Frequency-warping methods instead transform high-resolution source spectra toward the target spectrum.
- Sparse representation: NMF-based methods can work with very limited training data, while exemplar-based sparse representation uses sparsity constraints to alleviate over-smoothing.NMF factorizes nonnegative matrices into dictionaries and activation matrices; sparse representations construct target spectrograms from exemplars.
- Sparse representation: Sparse voice conversion constructs coupled source-target dictionaries from aligned exemplars and uses a shared sparse activation matrix to transfer a source spectrogram.The activation matrix serves as the pivot for converting source utterance X to target utterance Y.
- Sparse representation: Phonetic sub-dictionaries selected according to runtime speech content consistently outperform a single dictionary in exemplar-based sparse representation.This strategy requires phonetic labels and, in phonetic sparse representation, a speech recognizer at run time.
2) Phonetic Sparse Representation:
Non-parallel voice conversion seeks mappings without paired source-target utterances, using alignment, retrieval, or speaker modeling to establish correspondences. INCA iteratively refines alignment and conversion, including for cross-lingual settings.
- Non-parallel data: Non-parallel training enables applications where paired source-target utterances are unavailable, but requires establishing correspondences between speakers’ speech.Conventional mapping functions are easier to train with parallel data.
- Non-parallel data: Unsupervised clustering and target-database retrieval can establish source-target mappings, but errors across processing steps may accumulate and harm parameter estimation.A target-speaker HMM approach achieved performance comparable to parallel-data methods but required training-utterance orthography.
- INCA: INCA iterates nearest-neighbor alignment with a parametric conversion step, using each intermediate converted voice to refine the next alignment.Training repeats until the intermediate voice is sufficiently close to the target voice.
- INCA: INCA was first combined with GMM to estimate a linear mapping function without phonetic or linguistic information.This permits use with non-parallel data and cross-lingual voice conversion.
- INCA: Cross-lingual INCA achieved similar performance to an intra-lingual counterpart trained on parallel data.The comparison is reported for the INCA implementation of a cross-lingual system.
- INCA: INCA-DKPLS uses INCA to find corresponding frames so DKPLS can learn a nonlinear mapping, producing high-quality voice comparable to parallel-data training on the same task.The cited report describes the comparison as being on the same task.
B. Unit Selection Algorithm
Unit selection converts speech by retrieving target-speaker feature sequences rather than modeling a parametric mapping. Speaker modeling offers another non-parallel-data route by representing speaker characteristics with models or latent factors, while deep learning supplies data-driven representations.
- Unit selection: Unit selection directly searches a target-speaker inventory for an output feature sequence instead of modeling conversion with parameters.It optimizes the retrieved sequence at the utterance level.
- Unit selection: Because unit-selection waveforms come directly from the target speaker, the approach is associated with high speaker similarity and voice quality.The method is widely used for natural-sounding speech synthesis.
- Unit selection: Dynamic programming selects target feature vectors by balancing source-target acoustic distance against concatenative costs between consecutive target vectors.Acoustic distance preserves similarity to the source, while concatenative cost encourages coherent multi-frame retrieval.
- Speaker modeling: Text-independent speaker characterization can use GMMs or i-vectors to model speakers for voice conversion with non-parallel data.Reference-speaker and eigenvoice approaches can avoid parallel data for the target conversion pair but still require parallel data from reference speakers.
- Speaker modeling: Factor analysis separates speaker-specific information into low-dimensional latent variables, enabling models learned from prior non-parallel data to require fewer training samples than conventional JD-GMM.Parallel utterances remain required for estimating the conversion function.
- Speaker modeling: I-vectors can disentangle speaker identity from linguistic content in an unsupervised manner, including when source and target speech differ in content or language.This expands possibilities for non-parallel data scenarios.
- Deep learning: Deep learning benefits voice conversion by exploiting abundant training data and learned embeddings for linguistic content and speaker identity.The paper presents deep learning as transforming the analysis-mapping-reconstruction pipeline and addressing parallel and non-parallel settings.
A. Deep Learning for Frame-Aligned Parallel Data
Deep learning replaces frame-level statistical mapping with neural transformations that model nonlinear feature relationships and, through recurrent or attention-based architectures, temporal alignment and duration.
- 1) DNN Mapping Function:: DNN voice conversion learns a frame-wise nonlinear mapping y = F(x) from source to target features.It can model feature dimensions with fewer restrictions than GMM and DKPLS approaches.
- 1) DNN Mapping Function:: Deep neural approaches can transform spectral features as well as acoustic features such as fundamental frequency and energy contours.
- 1) DNN Mapping Function:: DNN and LSTM systems typically require large parallel corpora from paired speakers and external frame alignment during training.At run time, the conventional frame-level pipeline preserves the source speech duration.
- 2) LSTM Mapping Function:: LSTM models temporal correlations across speech frames, improving the naturalness and continuity of converted speech.Its memory blocks and gates learn how much long-range contextual information to retain.
- 2) LSTM Mapping Function:: DBLSTM voice conversion outperforms DNN voice conversion even without dynamic features.The architecture stacks multiple hidden layers of bidirectional LSTM networks.
- B. Encoder-decoder with Attention for Parallel Data: Attention-based encoder-decoder models learn mapping and alignment jointly, allowing converted speech duration to differ from the source duration.The decoder attends to multiple speech frames rather than mapping frames individually.
C. Beyond Parallel Data of Paired Speakers
Deep learning extends voice conversion beyond parallel data by learning mappings between unpaired speaker domains, especially through adversarial and cycle-consistency objectives.
- C. Beyond Parallel Data of Paired Speakers: Recent deep learning work considers non-parallel paired-speaker data, TTS and ASR systems, and speaker–linguistic-content disentanglement.
- 1) Non-parallel data of paired speakers:: Cycle-consistency uses L1 loss to preserve contextual information and find an optimal pseudo pair through circular conversion.The L1 norm represents least absolute errors and is described as producing sharper spectral features.
- 1) Non-parallel data of paired speakers:: CycleGAN learns source-to-target and target-to-source mappings from non-parallel data without dynamic-time-warping or attention-based frame alignment.Its training combines adversarial, cycle-consistency, and identity-mapping losses.
- 1) Non-parallel data of paired speakers:: CycleGAN achieves comparable performance to a GMM system trained on twice as much parallel data.Adversarial training also effectively addresses over-smoothing, a major contributor to speech-quality degradation.
- 1) Non-parallel data of paired speakers:: CycleGAN has been applied to mono-lingual, cross-lingual, emotional, and rhythm-flexible voice conversion.
- 1) Non-parallel data of paired speakers:: Because CycleGAN does not explicitly model internal representations such as identity, duration, or emotion, it is better suited to a specific source–target pair.The authors nevertheless describe it as an important milestone toward non-parallel voice conversion.
2) Leveraging TTS systems:
Voice conversion can leverage TTS and ASR representations or disentangle speaker identity from linguistic content to improve linguistic control and reduce dependence on paired data.
- 2) Leveraging TTS systems:: TTS-based voice conversion leverages shared attention knowledge or decoder architectures from text-to-speech systems.Tacotron-style encoder-decoder models provide the main architectural basis.
- 2) Leveraging TTS systems:: Joint TTS and voice-conversion training improves voice conversion with or without text input at run-time.The multi-source model supports text-only synthesis as well as voice or hybrid TTS-and-VC inputs.
- 2) Leveraging TTS systems:: Cotatron uses a pretrained multi-speaker Tacotron to derive speaker-independent linguistic features from source speech, requiring a source transcription at inference.
- 2) Leveraging TTS systems:: TTS-based techniques usually require a large training corpus, although speaker-adaptive TTS bootstrapping has been proposed for limited-data systems.
- 2) Leveraging TTS systems:: ASR systems can supply posterior sequences or bottleneck features to guide sequence-to-sequence voice conversion.This approach reuses phonetic representations learned from large speech-recognition corpora.
- 2) Leveraging TTS systems:: Average modeling maps phonetic posteriograms to acoustic features, separating speaker-independent linguistic information from speaker identity.The framework has been extended to waveform conversion, speaker conditioning, cross-lingual conversion, and emotional conversion.
- 2) Leveraging TTS systems:: Auto-encoders and related techniques disentangle speaker identity from linguistic content so the identity can be changed independently.
4) Disentangling speaker from linguistic content:
Voice conversion separates speaker identity from linguistic content through latent representations and speaker-conditioned decoding. The section also covers evaluation methods and their perceptual limitations.
- Disentangled representations: Auto-encoders avoid requiring parallel training data by encoding speech into a latent code and reconstructing it with a decoder.The bottleneck is intended to retain speaker-independent linguistic content while discarding other information.
- Disentangled representations: VQ-VAE representations preserve the most linguistic content while being the most speaker-invariant among three compared auto-encoding networks.The comparison examines separation of speaker identity from linguistic content.
- Speaker-conditioned decoding: VAE-based conversion conditions decoding on both a latent code and a speaker code, forcing the encoder to capture information separate from speaker identity.Speaker codes may be one-hot vectors, i-vectors, bottleneck representations, or d-vectors.
- Speech quality: GANs were proposed to address the over-smoothed, buzzy-sounding speech produced by VAE decoders.The approach trains a generator to deceive a discriminator distinguishing real from generated data.
- Prosody and style: Sequence-to-sequence non-parallel conversion can explicitly transfer source rhythm, speaking style, and emotion to target speech.
- Evaluation: Objective metrics include MCD for spectrum and PCC or RMSE for prosody, but acoustic-feature distortions do not always correlate with human perception.Subjective measures such as MOS, preference tests, and best-worst scaling are therefore also used.
2) Prosody Conversion:
Prosody conversion is evaluated through duration, pitch, and energy measurements alongside subjective listening tests. Reference-based metrics provide quantitative comparisons, while learned and non-intrusive measures address practical constraints.
- Prosody metrics: Prosody evaluation covers phonetic duration, energy contours, and pitch contours, each requiring measurements of similarity to reference speech.
- Prosody metrics: PCC measures linear dependence between aligned converted and target prosody contours, with higher values indicating better F0 conversion.
- Prosody metrics: RMSE measures differences between converted and target F0 features, with lower values indicating better performance; the same measure applies to energy contours.
- Prosody metrics: GPE counts voiced frames whose pitch differs by more than 20%, while FFE additionally counts voicing-decision errors.
- Subjective evaluation: MOS rates converted-voice quality on a five-point scale, while AB, ABX, and XAB tests compare naturalness or similarity between samples.
- Subjective evaluation: Best-worst scaling uses a few randomly selected options per decision and can produce more discriminating voice-quality rankings than MOS and preference tests.
- Evaluation constraints: Subjective evaluation represents intrinsic naturalness and similarity but is time-consuming and expensive because it requires many listeners.
- Learned evaluation: Reference-based perceptual evaluation is restricted when reference speech is unavailable, motivating non-intrusive metrics such as Quality-Net and MOS predictors.MOSNet predicts human MOS ratings and is highly correlated with them at system level but fairly correlated at utterance level.
VII. VOICE CONVERSION CHALLENGES
Voice Conversion Challenges establish shared datasets and evaluation protocols for fair comparison across systems. From 2016 to 2018, the tasks expanded from parallel to non-parallel conversion while performance improved and spoofing evaluation was added.
- Challenge framework: VCC is a biannual event begun in 2016 in which participants use common data and organizers evaluate submitted converted speech.
- Challenge progression: VCC 2016 used parallel training data, VCC 2018 added non-parallel conversion, and VCC 2020 introduced cross-lingual conversion.
- Why challenges are needed: Shared databases and protocols reduce dependence on re-implementing systems with different data and enable direct comparison with released state-of-the-art outputs.
- Security relevance: VCC datasets also support anti-spoofing research by providing advanced converted speech for developing and testing speaker-recognition defenses.
- VCC 2016: VCC 2016 used subjective naturalness and speaker-similarity evaluation, with 17 participants and 200 native English listeners.
- VCC 2016: 3.0 MOS and about 70% same-speaker judgments were reported for the best VCC 2016 system, while a substantial gap from target natural speech remained.
- VCC 2018: VCC 2018 attracted 23 parallel-task participants, including 11 in the non-parallel task, and evaluated systems with 260 crowd-sourced listeners.
- VCC 2018: 4.1 MOS and about 80% same-speaker judgments were reported for the best VCC 2018 system in both parallel and non-parallel tasks.The best system had similar performance across the two tasks.
D. Overview of the 2020 Voice Conversion Challenge
The 2020 Voice Conversion Challenge evaluated non-parallel conversion within and across languages, using limited speaker data and multifaceted assessments. Results showed near-target natural-speech similarity in the first task but no human-level naturalness, while the harder cross-language task still produced best-system MOS scores above 4.0.
- The challenge comprised non-parallel training in same-language English and cross-language English-to-Finnish, German, or Mandarin tasks.
- Each task used up to 70 utterances per speaker, with the first task involving 16 source-target models across four speakers.
- 31 participants submitted first-task results and 28 submitted second-task results, with speaker-independent training and orthographic transcriptions permitted.
- The challenge combined traditional metrics with speech recognition, speaker recognition, and anti-spoofing evaluations of converted speech.
- Several first-task systems approached target-speaker natural-speech similarity, but none achieved human-level naturalness; in the harder second task, best-system MOS scores exceeded 4.0.