Source-linked AI summary
A survey of AI-generated voices and their detection
Chengzhe Sun, Tianle Yang, Siwei Lyu
TL;DR
Highly realistic AI-generated voices create risks including impersonation, fraud, and disinformation, while their detection remains difficult because audio differs fundamentally from visual deepfakes. This survey integrates generation mechanisms with human speech physiology, phonetics, auditory perception, and detection research. It concludes that robust detection must address domain shift and unseen synthesis methods while complementing detection with watermarking, provenance tracking, and preventative verification tools.
Problem
Reliable AI-generated voice detection remains challenging because audio signals differ fundamentally from visual signals, while phonetics, prosody, and auditory perception shape synthetic-voice authenticity.
Method
The survey synthesizes human voice production and perception, AI voice-generation technologies, synthetic-speech detection methods, benchmark datasets, and future directions in an integrated framework.
Results
The survey identifies robustness to real-world degradation, higher-level phonetic and prosodic evidence, multimodal verification, and generalization to unseen synthesis methods as central detection directions.
Takeaways & Limitations
Detection should be complemented by watermarking, provenance tracking, real-time spoof prevention, and privacy-preserving authenticity verification.
Takeaways & Limitations
The survey may not cover some newer work because the AI voice synthesis and detection literature is rapidly evolving.
Abstract
from arXiv · showhide
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.
ATSIP 15,1
AI-generated voice technology has progressed from early rule-based and statistical systems to neural models capable of highly realistic, expressive speech. This survey connects voice-generation mechanisms with human speech physiology, phonetics, auditory perception, and detection methods to address the distinct challenges of synthetic-voice detection.
- AI voice generation: Neural architectures including WaveNet, Tacotron, VALL-E, GANs, RNNs, transformers, and diffusion models have enabled increasingly natural voice synthesis and voice cloning.WaveNet captured subtle prosody and natural inflection, while later systems advanced text-to-speech and cloning capabilities.
- Risks and motivation: AI-generated voices facilitate impersonation, fraud, and disinformation, including a $243,000 CEO-impersonation transfer and robocalls mimicking President Biden.These incidents motivate reliable detection safeguards.
- Detection challenge: Voice detection lags behind visual deepfake detection because audio has fundamentally different signal characteristics requiring methods tailored to human voice production and perception.The survey treats generation and detection as a combined scientific and societal challenge.
- Survey perspective: The survey integrates AI-generation mechanisms with phonetic and articulatory constraints to inform detection strategies grounded in human speech physiology.It links technical generation pipelines with forensic phonetic analysis rather than treating synthesis and detection in isolation.
- Survey scope: The survey reviews human voice production, auditory perception, phonetics, signal processing, state-of-the-art generation and detection methods, benchmark datasets, and future directions.Its stated scope covers both current techniques and technologies beyond detection.
- Scope boundary: The survey may omit newer research because the literature on AI voice synthesis and detection is rapidly evolving.The authors state that they intend to augment the survey as new work emerges.
- Phonetic foundations: Phoneme-level analysis is useful for detection because synthesis systems may fail to reproduce the fine-grained articulatory patterns defining phonemic contrasts.Subtle deviations in phoneme characteristics can distinguish natural from synthetic voices.
- Phonetic foundations: TTS and voice-conversion systems can produce irregular accent patterns for rare or out-of-domain words and may not reliably capture accentual targets.Accentual and prosodic features influence listeners’ evaluations, while dialectal contrasts remain measurable.
ATSIP 15,1
AI-generated voices have become increasingly natural and difficult for unaided listeners to detect reliably. Detection depends on phonetic, prosodic and perceptual cues, while synthesis systems are shaped by design choices involving representation, temporal modeling, decoding and data.
- Human speech and detection cues: Models may regress toward population means and underrepresent socially and linguistically conditioned phonetic variation.These weaknesses can produce subtle segment-level and speaker-specific deviations that support synthetic-voice detection.
- Human auditory perception: 73%: listeners correctly spotted deepfakes in a large English-and-Mandarin experiment, with prior examples providing only slight improvement.Detectability did not differ between Mandarin and English.
- Human auditory perception: Around 80%: participants judged cloned voices to match the real speaker’s identity, while AI-generated classification accuracy was about 60%.The findings indicate that strong commercial cloning systems can fool listeners in both identity and authenticity judgments.
- Human auditory perception: Political framing, available modality and synthesis method can shift real-versus-fake judgments.Text-only conditions reduced discernment, and TTS-synthesized political speech was harder to identify than voice-actor performances.
- Human auditory perception: Short cue-focused training reduced uncertainty and improved classification on initially uncertain items.Relevant cues included pitch, pauses, stop bursts, audible breath and overall audio quality.
- Human–machine comparison: Human listeners and machine detectors share some weaknesses, but humans can outperform benchmark models on bona fide audio in some evaluations.Other studies report similar failure types, native-listener advantages and greater susceptibility among older participants.
- Implications: Unaided auditory perception is insufficient for reliable detection under realistic conditions, supporting combinations of perceptual awareness and automated safeguards.Audio quality, prior information and modality affect outcomes, while brief training offers only limited gains.
- TTS design choices: Modern TTS pipelines map text and optional prosodic or stylistic conditioning to acoustic representations or directly to waveforms.Key design choices concern acoustic representation, temporal modeling, decoder architecture and the data regime.
ATSIP 15,1
Speech synthesis has progressed from concatenating recorded units and statistical parameter generation to neural, parallel, multilingual and diffusion-based systems. Each generation improves naturalness, control or adaptability while retaining trade-offs in speed, data requirements, robustness or computation.
- Historical development: Concatenative synthesis joined recorded phones, diphones or syllables, offering predictable behavior but audible joins, mechanical prosody and limited expressiveness.PSOLA enabled small pitch and duration modifications but did not remove the basic limitations.
- Historical development: SPSS modeled acoustic parameters with HMMs and GMMs, enabling explicit pitch, duration and speaking-rate control but often producing over-smoothed, flat speech.Postfilters and variance compensation could not fully recover expressive prosody.
- Neural synthesis: End-to-end neural TTS learned text-to-acoustics mappings and alignments, capturing coarticulation and natural prosody while introducing autoregressive speed and stability limitations.Char2Wav, Tacotron and WaveNet reduced many vocoder artifacts and reached speaker similarity comparable to natural recordings.
- Neural synthesis: Parallel and integrated architectures improved generation efficiency and robustness through non-autoregressive decoding, explicit predictors, monotonic alignment and unified synthesis.FastSpeech, Glow-TTS, VITS and PortaSpeech illustrate these design directions, though fine-grained control can remain challenging.
- Personalization and multilinguality: Few-shot meta-learning and style conditioning made rapid voice cloning practical while remaining sensitive to recording quality, reference choice, accents and prosody.Meta StyleSpeech adapted quickly while preserving timbre and prosody, but difficult accents often required additional data or fine-tuning.
- Personalization and multilinguality: Zero-shot multilingual systems broadened speaker and language coverage, but data imbalance, style drift, code-switching and noisy prompts remain challenges.YourTTS demonstrated multilingual cloning and accent transfer without per-language engineering.
- Diffusion synthesis: Diffusion models improved expressive and robust synthesis through iterative denoising and rich linguistic or prosodic conditioning.NaturalSpeech, NaturalSpeech 2 and StyleTTS 2 reached strong quality, including zero-shot singing and near-natural ratings.
ATSIP 15,1
Recent synthesis systems increasingly combine diffusion, codec tokens and large language-model-style architectures to improve fidelity, controllability and scalability. These gains remain constrained by inference cost, prompt sensitivity and trade-offs among naturalness, efficiency, robustness and control.
- Current limitations: Diffusion systems still require multi-step inference and substantial computation, while few-step samplers and hybrids only mitigate latency.The computational burden remains a practical boundary for deployment.
- Codec and language-model synthesis: Codec- and language-model-based systems scale synthesis to massive corpora and support robust, near-human zero-shot cloning, but at high resource cost or with prompt sensitivity.BASE TTS used over 100k hours of speech; VALL-E and VALL-E 2 enabled near-human zero-shot cloning.
- Current limitations: The field is moving toward universal, real-time and personalized synthesis, while still facing unresolved scalability and expressiveness challenges.The survey frames this trajectory as a trade-off between fidelity, style control, speed and computational demands.
- Voice conversion: Voice conversion alters a source speaker’s identity characteristics toward a target while preserving linguistic content through disentanglement and resynthesis.Systems separate phonetic information from cues such as timbre, formant structure and pitch before waveform generation.
- Voice conversion: VC design varies in content–timbre representation, training-data assumptions and waveform-generation strategy.Choices span mel cepstra, phonetic posteriorgrams or self-supervised embeddings; parallel or nonparallel data; and parametric, neural or direct waveform generators.
ATSIP 15,1
The surveyed generation landscape spans representative TTS and voice-conversion approaches with varied languages, adaptation settings, representations and evaluation metrics. Across these systems, scalability and general-purpose deployment depend on balancing quality, inference efficiency, controllability and data requirements.
- Representative models: Representative TTS models are organized by model approach, language coverage, adaptation ability, evaluation metrics and publication year.The table is presented as a representative overview rather than a single benchmark comparison.
- Representative models: VALL-E 2 uses an improved neural codec language model for English zero-shot synthesis evaluated with MOS and speaker similarity.Its listing emphasizes neural codec modeling and zero-shot adaptation.
- Representative models: MaskGCT applies a masked codec-based transformer to multilingual zero-shot synthesis, including MOS and inference-efficiency evaluation.The entry highlights codec modeling and efficiency as central evaluation dimensions.
- Representative models: XTTS provides multilingual zero-shot TTS evaluated with MOS and speaker similarity, while SupertonicTTS uses flow-matching-based latent diffusion for English few-shot TTS.The table entries show differing language and adaptation regimes across recent systems.
- Model design: CM-TTS targets fast inference with a consistency model, whereas CLaM-TTS uses residual vector quantization codecs.These approaches represent distinct strategies for improving efficiency or codec-based representation.
- Model design: DiFlow-TTS and VITA-Audio represent newer directions using discrete flow matching and interleaved token generation, respectively.Their listings indicate continued exploration of discrete and token-based synthesis designs.
- Voice conversion: VC systems trade naturalness, controllability, efficiency and robustness while varying in alignment, data dependence and representation.The survey covers statistical, NMF, neural, autoencoder, adversarial and cycle-consistent approaches.
ATSIP 15,1
Voice conversion has progressed from raw-audio and autoregressive systems toward diffusion-based models that improve naturalness and control while increasing computational demands. Evaluation therefore combines human judgments with objective measures covering fidelity, prosody, intelligibility, identity, and efficiency.
- End-to-end voice conversion: End-to-end voice conversion bypasses hand-engineered features and classical vocoders by coupling content encoders with waveform decoders.Autoregressive decoders offer high fidelity but latency, whereas non-autoregressive and diffusion-inspired decoders improve speed.
- Controlled and multilingual conversion: Emotion, style, accent, and cross-lingual conversion rely on prosodic features, style tokens, universal content spaces, and speaker embeddings.Sparse data and the balance between expressiveness and intelligibility remain central challenges.
- Diffusion-based voice conversion: Diffusion-based voice conversion uses iterative denoising with content, speaker, and prosody conditioning for fine-grained timbre and expressiveness control.These systems provide strong naturalness, content preservation, and resilience to acoustic variation, but slower inference remains a drawback.
- Evaluation: Evaluation triangulates human MOS tests with objective measures for spectral fidelity, prosody, intelligibility, identity transfer, and real-time performance.Learned predictors such as MOSNet reduce evaluation cost but correlate imperfectly with human judgment.
- Open challenges: Open challenges include reducing latency and memory use, improving robustness to noise and channel variability, strengthening prosody control, and scaling across speakers, styles, and languages.These goals must be pursued without sacrificing fidelity.
ATSIP 15,1
Voice cloning evolved from data-intensive concatenative and statistical methods to neural, embedding-based, meta-learning, few-shot, and zero-shot systems. Newer diffusion and codec-based models improve adaptation, generalization, and control, while evaluation and safeguards remain necessary because automatic metrics imperfectly represent perception.
- Classical cloning: Classical unit-selection and statistical cloning methods produced stable, intelligible speech but required large speaker-specific datasets and offered limited expressiveness.They predicted acoustic parameters for vocoder rendering after matching linguistic and prosodic context.
- Neural cloning: Neural TTS with neural vocoders improved naturalness and similarity using modest speaker data and gradient-based fine-tuning.Tacotron-style systems reproduced timbre, coarticulation, and more natural prosody, but still required speaker adaptation.
- Embedding-based cloning: Speaker embeddings enabled cloning new voices from seconds of reference audio at inference time without retraining the model.D vectors and x vectors condition multi-speaker acoustic or waveform models, although noisy references and other sensitivities remain limitations.
- Few-shot and meta-learning: Meta-learning and related adaptation methods reduce data requirements and accelerate personalization while improving robustness to speaker and channel variability.Atypical accents, speaking styles, and training complexity remain challenges.
- Zero-shot cloning: Zero-shot systems synthesize unseen voices and languages from short prompts using multilingual models, verification encoders, or codec-token language models.These methods provide strong generalization and multilingual capability without per-speaker adaptation.
- Recent architectures: Diffusion decoders and large codec-based language models combine conditioning on content, timbre, and prosody with scale and improved inference strategies.ProDiff accelerates diffusion through adaptive step scheduling while maintaining high-quality mel-spectrogram generation.
- Evaluation and safeguards: Evaluation combines human listening tests with speaker, intelligibility, prosody, timbre, and efficiency metrics, while safeguards are increasingly integral.Automatic measures provide scale but remain imperfect proxies for perception.
4. Detection of AI-generated voices
AI-generated voice detection uses signal-level acoustic analysis to identify artifacts without requiring transcripts. These methods are efficient and model-agnostic but vulnerable to recording conditions, domain shift, and deliberate obfuscation.
- Signal-level analysis: Signal-level detectors analyze spectrogram statistics, phase irregularities, harmonic-noise balance, micro-jitter, shimmer, pitch, formants, and vocoder fingerprints.The features target acoustic traces associated with synthetic speech generation.
- Limitations: Signal-based methods are fast, model-agnostic, and transcript-free, but brittle under room impulse responses, domain shift, and intentional obfuscation.Band-limiting and re-recording are examples of obfuscation conditions that challenge detection.
ATSIP 15,1
Detection research combines phonetic and linguistic cues with benchmark-driven deep learning, raw-waveform modeling, and vocoder-aware detection. Phonetic cues can be more semantically grounded, while detector robustness remains constrained by transcription accuracy, channel mismatch, and accent variation.
- Phonetic and linguistic analysis: Phonetic detection examines phoneme duration, coarticulation, formant trajectories, prosody, disfluencies, and text–acoustic consistency.These cues are more semantically grounded and potentially more interpretable than purely signal-level traces, but require accurate linguistic analysis.
- Benchmarks: ASVspoof established standardized datasets, protocols, and metrics for logical- and physical-access scenarios, including EER and tandem DCF.The challenges promoted joint reporting of speaker-verification and anti-spoofing performance.
- Spectrogram-based detectors: Spectrogram CNN and RNN detectors learn local spectral micro-patterns and longer temporal dependencies, while AASIST adds spectro-temporal attention and graph reasoning.AASIST improved robustness across datasets and attack types according to the cited work.
- Raw-waveform detectors: Raw-waveform detectors learn time-domain filters and long-context representations directly, potentially capturing subtle fingerprints and generalizing to some unseen attacks.They remain sensitive to sample-rate and channel mismatch unless heavily augmented.
- Vocoder-aware detection: Vocoder-aware detectors exploit systematic artifacts imprinted by neural vocoders and can generalize more gracefully across generation systems.Rawnet2-vocoder jointly classifies real versus fake speech and identifies the underlying vocoder.
ATSIP 15,1
Detection methods have progressed from acoustic cues toward phonetic and vocoder-aware approaches, but robustness depends on attack coverage, domain conditions, and calibration. Phonetic analysis offers interpretable cues tied to articulatory constraints.
- Detection methods: Vocoder-aware systems improve cross-attack generalization but add complexity that complicates calibration in deployment.Coverage of novel decoders and codec-LM models remains important, while shallow fingerprints can be obfuscated.
- Phonetics-based detection: Phonetic analysis targets fine-grained segmental and suprasegmental features that current synthesis models may reproduce inconsistently.This perspective complements general acoustic and spectral cues as synthesis advances narrow broad acoustic differences.
- Phonetics-based detection: Vowel formants F1–F3 can distinguish real from synthetic speech more effectively than several global acoustic measures.The cited comparison includes long-term formant distributions, long-term F0, and speaker-level MFCCs.
- Phonetics-based detection: Reconstructing vocal-tract configuration achieves detection rates above 90% even on short utterances.The framework identifies inconsistencies between natural and synthetic speech through physiological constraints.
ATSIP 15,1
Phonetics-based detection connects articulatory, prosodic, consonantal, perceptual, and sociolinguistic evidence to synthetic-speech analysis. Shared benchmarks improve comparability, but benchmark overfitting and limited coverage constrain real-world transfer.
- Phonetic interactions: Missing consonant-induced F0 perturbations can diagnose synthetic speech, especially when TTS models fail to generalize beyond frequent words.The study interprets this failure as reliance on surface-form memorization rather than abstract articulatory-acoustic encoding.
- Consonant acoustics: Unvoiced fricatives and stops contain distinctive artifacts, while combining voiced and unvoiced cues can improve detection accuracy and interpretability.These artifacts are described as features current synthesis models fail to reproduce consistently.
- Sociolinguistics: Integrating linguistic expertise can strengthen technical detection and broaden societal resilience against deception.The cited sociolinguistic perspective emphasizes phonetic, phonological, and linguistic variation alongside listener education.
- Benchmarks: Shared datasets establish common protocols and metrics that make detector results reproducible and directly comparable.Later evaluation tools include tDCF, minDCF, and Cllr for system risk and score calibration.
- Benchmarks: Benchmark overfitting can make models exploit dataset-specific attack fingerprints that fail to transfer to real-world conditions.This limitation affects the interpretation of performance measured on shared resources.
- Benchmark evolution: ASVspoof 2015 established a unified spoofing testbed with roughly 260,000 utterances from 106 speakers.The corpus combined bona fide speech with classical voice-conversion and statistical-TTS fakes.
- Benchmark evolution: ReMASC provided approximately 54,700 clips from 50 speakers across devices, rooms, and playback-capture configurations.Its real acoustic channels, reverberation, and device effects supported physical-access stress testing.
ATSIP 15,1
Benchmark resources have expanded from replay and statistical spoofing toward multimodal, large-scale, vocoder-focused, and adversarial settings. Despite broader coverage, open-set evaluation still faces domain shift and gaps in emerging generators, conversational audio, and long-form conditions.
- Multimodal benchmarks: FakeAVCeleb combines manipulated video with synthetic or manipulated audio from 500 celebrities for multimodal deepfake analysis.Its design supports lip-audio synchronization, identity coherence, and cross-modal fusion studies.
- Large-scale benchmarks: ASVspoof 2021 unified logical-access, physical-access, and speech-deepfake tracks with over one million utterances.The expanded spoofing lists reduced the effectiveness of simple system fingerprinting.
- Large-scale benchmarks: ADD 2022 contains more than 500,000 audio clips and introduces demographic condition splits for industrial-scale detection benchmarking.Language and content concentration may leave gaps for low-resource settings.
- Vocoder-focused benchmarks: LibriSeVoc contains approximately 92,400 samples generated with six neural vocoders to isolate decoder-level fingerprints.The corpus includes 13,201 real and 79,206 synthetic samples and supports vocoder identification.
- Adversarial benchmarks: ASVspoof 5 adds crowd-sourced deepfakes and adversarial attacks while expanding evaluation to minDCF and Cllr beyond EER and tDCF.These additions target realistic threats and deployment-oriented calibration assessment.
- Benchmark evolution: Benchmark development progresses from statistical spoofing and replay realism to neural, multimodal, vocoder-focused, and adversarial resources.The progression increases scale, condition diversity, and alignment with real-world deployment scenarios.
- Open challenges: Current resources remain limited in diffusion- and codec-based generators, spontaneous noisy conversation, cross-modal and long-form testing, and domain-shift robustness.These gaps constrain conclusions from closed-set and open-set comparisons.
5. Conclusion and future directions
The survey concludes that increasingly realistic synthetic voices sustain an ongoing arms race between generation and detection. Future work emphasizes robust, generalizable detection alongside watermarking, provenance, real-time prevention, and privacy-preserving verification.
- Conclusion: Voice generation and detection are engaged in a continuing cat-and-mouse process, with advances on each side prompting countermeasures from the other.The survey frames this cycle as ongoing because generative models and detection methods continually become harder to defeat.
- Conclusion: The authors state that rapid progress in generative models means the arms race between synthesis and countermeasures remains unresolved.The conclusion presents the identified trends as directions for future research and broader protective strategies.
- Future generation: Zero-shot and few-shot voice conversion, expressive synthesis, and diffusion models are identified as directions for more data-efficient, flexible, natural speech generation.These approaches include capturing a speaker from seconds of audio, modeling prosody and affect, and iteratively refining waveforms from noise.
- Future detection: Future detectors must withstand compression and noise, exploit phonetic and prosodic cues, combine audio with video or metadata, and generalize to unseen synthesis methods.The survey notes that traditional classifiers can overfit known attack types, motivating meta-learning or anomaly detection.
- Preventative tools: Preventative safeguards include imperceptible watermarking, provenance tracking, real-time spoof prevention, and privacy-preserving authenticity verification.Deployment in live communication channels makes latency and computational cost important constraints, while verification must protect sensitive speaker data.
ATSIP 15,1
This supplied material is primarily a fragmented reference list rather than a substantive section narrative. The references indicate work spanning speech synthesis, voice conversion, spoof detection, multimodal deepfakes, and human discernment.
- Human evaluation: The references also address human ability to detect political speech deepfakes and AI voice clones through perception, naturalness, similarity, and discernment studies.The supplied entries include work on human detection across transcripts, audio, and video, as well as expert listening and naturalness ratings.
- Speech generation: The references include studies on text-to-speech, voice conversion, zero-shot synthesis, accent conversion, and diffusion-based speech generation.Examples include multilingual and zero-shot systems, custom voice synthesis, and diffusion models for speech.
- Detection: The cited detection literature covers audio anti-spoofing, phonemic features, vocal-tract reconstruction, causal linguistic cues, and multimodal audio-video deepfake datasets.These titles suggest complementary signal-level, linguistic, articulatory, and multimodal detection perspectives.