Source-linked AI summary
MakeItTalk: Speaker-Aware Talking-Head Animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, Dingzeyu Li
TL;DR
MakeItTalk addresses audio-driven facial animation from a single portrait by separating speech content from speaker information and using speaker-aware facial landmarks as an intermediate representation. It generates expressive animations across human and artistic portraits, with evaluations reporting higher overall quality than prior state-of-the-art methods. The authors identify speaker mood and certain bilabial and fricative sounds as remaining limitations.
Problem
Audio-to-facial-animation mapping is ambiguous because speakers show different expressions and head poses for the same content, while expressive dynamics are difficult to estimate from audio.
Method
MakeItTalk disentangles audio content and speaker information, predicts speaker-aware facial landmarks, and uses them to animate a single portrait image.
Results
The method generates expressive talking-head animations for unseen audio, faces, and voices, including photorealistic and non-photorealistic portraits, with higher overall quality than prior state-of-the-art methods.
Takeaways & Limitations
Disentangled content and speaker representations support lip synchronization, personalized expressions, head-motion dynamics, and unified animation across human faces and artistic imagery.
Takeaways & Limitations
Head motion and facial expressions do not fully model factors such as speaker mood, and bilabial and fricative sounds are not always captured well.
Abstract
from arXiv · showhide
We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating talking faces, our method first disentangles the content and speaker information in the input audio signal. The audio content robustly controls the motion of lips and nearby facial regions, while the speaker information determines the specifics of facial expressions and the rest of the talking head dynamics. Another key component of our method is the prediction of facial landmarks reflecting speaker-aware dynamics. Based on this intermediate representation, our method is able to synthesize photorealistic videos of entire talking heads with full range of motion and also animate artistic paintings, sketches, 2D cartoon characters, Japanese mangas, stylized caricatures in a single unified framework. We present extensive quantitative and qualitative evaluation of our method, in addition to user studies, demonstrating generated talking heads of significantly higher quality compared to prior state-of-the-art.
1 INTRODUCTION
MakeItTalk generates expressive talking-head animations from an unseen audio clip and a single unseen portrait image by disentangling speech content from speaker information and predicting speaker-aware facial landmarks. The method supports varied portrait types and achieves higher-quality, more plausible animations than prior work.
- MakeItTalk animates new portrait images from audio alone, producing facial expressions and head motions for faces and voices unseen during training.
- Disentangled speech content synchronizes lips and nearby regions, while speaker information controls personalized expressions and head-motion dynamics.
- Facial landmarks provide a compact intermediate representation for speaker-dependent dynamics, with tens of degrees of freedom rather than millions of pixels.
- The landmark representation supports animation of human faces, sketches, 2D cartoons, Japanese mangas, and stylized caricatures.
- The method achieves significantly more accurate and plausible talking heads than prior work in qualitative and quantitative evaluations, especially for unseen static face images.
2 RELATED WORK
Related work spans audio-driven landmarks, lip-sync synthesis, style-aware animation, portrait warping, evaluation metrics, and image-to-image translation. MakeItTalk combines landmark-based animation with speaker-aware disentanglement and generalization to unseen faces and non-photorealistic imagery.
- Audio-driven facial landmark synthesis: Prior audio-driven landmark methods synchronize facial landmarks, model expressions and poses, or decouple landmark prediction from rasterized video generation.
- Lip-sync facial animation: End-to-end lip-sync systems generate cropped lips or full human faces, while rigged portrait methods may require manual cartoon rigging and retargeting.
- “Style”-aware facial head animation: Earlier style-aware methods reproduced speaker motion dynamics but were limited to specific subjects or did not generalize to unseen speakers.
- Warping-based character animation: Portrait-warping systems animate images using videos, landmarks, templates, or motion capture, whereas this work synthesizes facial expressions and head pose from audio alone.
- Evaluation metrics: Existing evaluation often relies on subjective studies or pixel-level metrics, motivating metrics for high-level facial expression and head-motion dynamics.
- Image-to-image translation: Image-to-image translation supports natural human talking-head animation, while MakeItTalk additionally handles unseen faces without fine-tuning and non-photorealistic images.
- Disentangled learning: Disentangling content and style in audio draws on methods from the voice-conversion community, including speaker-identity embeddings and few-shot voice conversion.
3 METHOD
MakeItTalk disentangles speaker-agnostic speech content from speaker identity, predicts speaker-aware facial landmarks, and renders those landmarks into animated cartoon or natural face images.
- Audio and landmark prediction: MakeItTalk separates audio content for lip and nearby-region motion from speaker information for personalized expressions and head dynamics.Content is extracted with AutoVC, while speaker identity is encoded separately.
- Audio and landmark prediction: An LSTM followed by an MLP maps audio content and initial 3D landmarks to per-frame landmark displacements.The input landmarks q ∈ R68×3 are animated by predicted displacements Δq_t.
- Speaker-aware animation: The method differentiates conservative and active head-motion dynamics by conditioning landmark sequences on different speaker identities.Figure 3 compares predicted landmark sequences for two speakers with distinct head-motion tendencies.
- Speaker-aware animation: Speaker-aware prediction combines a speaker embedding, separately encoded content, and initial landmarks to model personalized expressions and head motion.A self-attention module captures longer dependencies because head motions last much longer than phonemes.
- Landmark-to-image synthesis: For cartoons, Delaunay-triangulated landmark regions warp the original image textures, while natural faces use a landmark-conditioned image-to-image translation network.The final animation takes an input image Q and predicted landmarks {y_t}; the two portrait types use different synthesis implementations.
4 TRAINING
Training uses separate data and objectives for content animation, speaker-aware dynamics, and image synthesis, with landmark losses supplemented by graph-structure and adversarial terms.
- Voice conversion training: The voice-conversion component is trained on VCTK using same-speaker utterances to reconstruct source speech from content and speaker embeddings.VCTK contains utterances from 109 native English speakers with varied accents.
- Speech content animation training: Content animation is trained on six hours of Obama Weekly Address video with extracted audio and facial landmarks.The dataset’s high resolution and consistent front-facing camera angle support accurate landmark detection.
- Speech content animation training: The content loss combines landmark-position distance with graph Laplacian-coordinate distance to preserve relative landmark placement and facial shape details.The Laplacian neighborhoods connect landmarks within each of eight predefined facial parts, with λ_c=1.
- Speaker-aware animation training: Speaker-aware dynamics are learned from a diverse VoxCeleb2 subset using a discriminator that evaluates landmark sequences together with audio content and speaker embeddings.The discriminator treats training landmarks as real and generated landmarks as fake under an LSGAN objective.
- Speaker-aware animation training: Generator training combines realism, absolute landmark position, and Laplacian-coordinate objectives, alternating optimization with discriminator training.The implementation sets λ_s=1 and μ_s=0.001 through hold-out validation.
- Image synthesis training: Natural-image synthesis is pretrained on paired VoxCeleb2 frames and fine-tuned on high-resolution video crops using pixel and perceptual reconstruction losses.The encoder/decoder receives a source frame concatenated with rasterized target landmarks.
5 RESULTS
MakeItTalk generates talking-head animations from a single portrait and audio, including facial expressions, head poses, and stylized faces. Results include qualitative galleries, comparisons with prior methods, and broader image-domain generalization.
- The method generates full facial expressions and dynamic head poses from a single input image and audio.
- It animates paintings, sketches, 2D cartoons, Japanese mangas, stylized caricatures, and casual photos using portrait images unseen during training.The method generalizes despite training only on human facial landmarks by using relative landmark displacements.
- Compared with prior video-generation methods, MakeItTalk captures facial expressions and head motion, whereas the baselines primarily predict lips on cropped faces.The authors also report better lip synchronization than Chen et al. for side-facing portraits, while noting background distortion artifacts.
- The image translation module also produces plausible animations for paintings, statue heads, and rendered 3D-model images.
- The supplementary comparison reports more accurate lip synchronization than Thies et al. and explicit head-pose generation rather than heuristic post-processing.
5.3 Evaluation Protocol
The evaluation uses a VoxCeleb2 test split and manually verified reference landmarks. Metrics assess lip accuracy, overall landmark accuracy, motion dynamics, and head pose.
- The test split contains 268 video segments from 67 speakers, whose test speech and video differ from training data.Speaker identities were observed during training.
- Reference landmarks are extracted from test clips and manually verified for evaluating synthesized landmarks.Each clip lasts 5 to 30 seconds.
- Jaw-lip landmark distance and velocity difference measure positional and temporal lip-motion accuracy using normalized Euclidean distances.
- Overall landmark distance and velocity difference measure facial landmark accuracy and motion dynamics normalized by face width.
- The evaluation protocol includes metrics for head rotation and position in addition to landmark-based measures.
5.4 Content Animation Evaluation
The evaluation compares MakeItTalk with landmark-synthesis and head-motion baselines on facial expression, head pose, and speaker-aware dynamics. The full method yields more faithful head-motion predictions than copying motion from other videos.
- Prior landmark-synthesis methods are compared after factoring out head motion because they cannot produce head motion directly.
- MakeItTalk produces more accurate facial landmark configurations than the compared methods, which show conservative mouth opening or inaccurate closed-mouth predictions.
- 2.7x less D-Rot error and 1.7x less D-Pos error are achieved than with the retrieve-random ID baseline.The full method also achieves much smaller errors than both head-motion retrieval baselines.
- Generated examples show head poses including nods and swings for cartoon and natural human faces.
- The model’s predicted head-motion dynamics lie closer to reference speaker dynamics than the evaluated alternatives.The comparison uses variance in Action Units, head pose, and position across 8 representative speakers.
5.6 Ablation study
Ablations show that separating speech content from speaker identity and using both branches improves lip synchronization and speaker-aware head motion. User studies additionally evaluate speaker awareness and perceived realism.
- Branch ablations: The no-speaker-branch variant has 1.6x larger head-pose errors and 1.3x larger head-position errors than the full method.It performs well on lip landmarks because those are synchronized with audio content.
- Branch ablations: The no-content-branch variant has 1.6x higher jaw-lip landmark error and 2.4x higher open-mouth-area error.Its head-pose and position performance is slightly worse than the full method, while lower-face dynamics are not synchronized well with audio content.
- Branch ablations: Using both content and speaker-aware branches provides the best performance across all evaluation metrics.
- Disentanglement: The no-separation variant has 1.5x, 2.4x, and 4.1x higher errors for jaw-lip position, velocity, and open-mouth-area difference.The authors hypothesize that entangled content and speaker information make the audio-to-landmark mapping difficult to disambiguate.
- Speaker identity: Random speaker-identity injection causes 3.6x more head-pose error, indicating that the model captures speaker-specific head-motion dynamics rather than random ones.
- User studies: The natural-human user study gathered 4680 responses, and MakeItTalk was voted most realistic and plausible by a large majority.
5.8 Applications
MakeItTalk supports dubbing and bandwidth-limited video conferencing by driving a target portrait from audio, including natural and cartoon profiles. It also supports text-to-video generation through speech synthesis and interactive pose editing.
- Given a target actor’s frame and another person’s dubbing audio, the method generates the actor speaking according to that speech.
- Audio-driven talking-head video can support bandwidth-limited conferencing when visual frames cannot be delivered at high fidelity or frame rate.The text emphasizes preserving facial expressions, especially lip motions, because they contribute to communication understanding.
- The supplementary video demonstrates natural-face video synthesis from text after converting the text to audio with a speech synthesizer.
- The synthesized heads can be edited interactively by rotating the predicted intermediate landmarks to change pose.
6 CONCLUSION
The method generates speaker-aware talking-head animations from audio and a single image, including unseen clips and portraits, by predicting landmarks from disentangled audio representations. It achieves higher overall quality than prior work, while limitations remain in phoneme capture, image translation, large motion, landmark sparsity, and user interaction.
- MakeItTalk generates speaker-aware talking-head animations from an audio clip and a single image, including new audio clips and portrait images unseen during training.Its key insight is predicting landmarks from disentangled audio content and speaker information.
- The approach produces more expressive animations with higher overall quality compared to the state-of-the-art.
- The speech-content animation does not always capture bilabial and fricative sounds, including /b/m/p/f/v/.The authors attribute this to short phoneme representations being missed during voice-spectrum reconstruction.
- The landmark-to-video translation can show background distortion, artifacts, and camera-motion impressions because it warps foreground and background together.
- Large head rotations or translations produce more artifacts because unseen neck, shoulder, and hair regions must be hallucinated from one image.
- Sparse landmarks can distort natural faces during large motion because they act as coarse proxies for head structure.Denser landmarks or morphable-model parameters are suggested alternatives, but zero-shot training remains challenging.
- The fully automatic pipeline does not yet incorporate user interaction for editing landmarks in selected frames and propagating edits through the video.
7 ETHICAL CONSIDERATIONS
The paper notes that synthetic talking-head videos may be misunderstood or misused, and includes a watermark to make generated videos identifiable as synthetic.
- Talking-head generation algorithms can be misused to spread misinformation or support other malicious acts.
- The authors include a watermark in generated videos to make their synthetic origin clear.
A SPEAKER-AWARE ANIMATION NETWORK
The described speaker-aware animation network uses a compact attention module and a generator architecture for natural-face image synthesis.
- The attention network contains two identical layers, each using two self-attention heads with model dimensionality 32 and a feed-forward layer.
- The attention network also includes a one-layer MLP embedding with hidden size 32 and a positional encoder.
- Table 4 presents the generator architecture for synthesizing natural face images.
B IMAGE-TO-IMAGE TRANSLATION NETWORK
The network architecture for natural human face synthesis is specified by feature-map resolutions and named block types.
- Table 4 defines ResBlock down, ResBlock up, and Skip operations used in the natural human face image-generation network.ResBlock down uses strided convolution and residual blocks; ResBlock up uses nearest-neighbor upsampling, convolution, and residual blocks.