Source-linked AI summary

EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Linrui Tian, Qi Wang, Bang Zhang, Liefeng Bo

arXiv:2402.17485v3cs.CV

TL;DR

Talking-head generation must model ambiguous audio-to-expression mappings while preserving identity and realistic motion. EMO directly synthesizes audio-synchronized portrait video from a single image using weakly conditioned diffusion, and reports improved expressiveness and realism over existing state-of-the-art methods. The framework also supports expressive speaking and singing videos, though its weakly constrained motion can sometimes generate incorrect body parts.

  • Problem

    Audio-to-head animation has an ambiguous one-to-many mapping, while prior methods often rely on constrained intermediate signals or motion priors that limit expressive generation.

  • Method

    EMO directly generates audio-synchronized video from a single portrait using diffusion, audio attention, temporal modules, and weak face-location and motion-speed conditions.

  • Results

    EMO outperforms existing state-of-the-art methods in expressiveness and realism, including on E-FID, and generates both speaking and singing videos.

  • Takeaways & Limitations

    Audio-driven weakly conditioned diffusion can generate lifelike portrait videos while maintaining consistent identity and expressive motion across frames.

  • Takeaways & Limitations

    Expressive audio can cause unintended body-part generation because the head-focused training data contains only 3% frames with hands.

Abstract

from arXiv · show

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of traditional techniques that often fail to capture the full spectrum of human expressions and the uniqueness of individual facial styles. To address these issues, we propose EMO, a novel framework that utilizes a direct audio-to-video synthesis approach, bypassing the need for intermediate 3D models or facial landmarks. Our method ensures seamless frame transitions and consistent identity preservation throughout the video, resulting in highly expressive and lifelike animations. Experimental results demonsrate that EMO is able to produce not only convincing speaking videos but also singing videos in various styles, significantly outperforming existing state-of-the-art methodologies in terms of expressiveness and realism.

1 Introduction

EMO addresses the difficulty of mapping audio to expressive head motion by directly generating synchronized portrait video from a single image without strong intermediate motion representations. It uses weak conditions and diffusion modeling to preserve identity while supporting expressive, realistic animation.

  • Motivation: Audio-to-head animation has an ambiguous one-to-many mapping, making facial expressions and head movements difficult to generate together.Existing methods often separate head motion from facial expression and use predefined poses or explicit 3D and landmark signals.
  • Approach: EMO transforms a single portrait image into an audio-synchronized expressive video without intermediate 3D representations or predefined motion templates.The framework directly captures audio-visual correlations with a diffusion-based generator.
  • Approach: Weak conditions guide face location and approximate movement speed without strictly constraining facial position or motion speed.This design is intended to stabilize generation while preserving expressiveness.
  • Data and evaluation: Over 250 hours of multilingual speaking and singing footage support training across varied expressions and vocal styles.The dataset includes speeches, film and television clips, and singing performances in Chinese and English.
  • Data and evaluation: EMO outperforms existing state-of-the-art methods on expressiveness and realism-related evaluations, including the introduced E-FID metric.The reported contribution covers talking-head video quality broadly rather than a single evaluation measure.

2 Related Work

Related work spans video-based and single-image audio-driven talking-head generation. Prior systems improve particular controls such as lip synchronization but often inherit fixed motion or limited 3D representations that constrain expressiveness and realism.

  • Approach categories: Audio-driven talking-head methods are broadly divided into video-based and single-image approaches.Video-based methods edit an input video, whereas single-image methods animate a reference photograph.
  • Video-based methods: Video-based systems such as Wav2Lip rely on a base video, which can fix head movements and limit generation primarily to mouth motion.Wav2Lip uses a discriminator for audio-lip synchronization.
  • Single-image methods: Single-image methods may separately learn blendshapes and head poses before using a 3D facial mesh to guide frame generation.This introduces an intermediate representation between audio-driven motion modeling and final video synthesis.
  • Limitations: The limited representational capacity of 3D meshes can constrain the expressiveness and realism of generated talking-head videos.The related-work discussion identifies this as a common issue across such approaches.

3 Method

EMO extends Stable Diffusion into an audio-driven video generator that combines reference features, audio attention, temporal modeling, and weak spatial and speed controls. Its staged training and motion-frame reuse support identity consistency, coherent transitions, and controllable long-duration generation.

  • System overview: Given one portrait and voice audio, EMO generates synchronized video while preserving identity, natural head motion, and expression.Cascaded clips support long-duration videos with coherent motion.
  • Preliminaries: The method builds on Stable Diffusion’s latent denoising framework and adapts its UNet-based architecture for video generation.Stable Diffusion encodes images into latent space, adds noise, and learns to denoise the resulting latent representation.
  • System overview: Frames Encoding extracts reference-image and motion-frame features, while the Diffusion Process denoises multi-frame noise conditioned on audio and a facial-region mask.ReferenceNet supplies features, and the Backbone Network performs denoising.
  • Network pipelines: Reference-Attention uses ReferenceNet features for identity consistency, while Audio-Attention injects nearby-frame audio features to drive motion.Audio features come from pretrained wav2vec blocks and include temporal context around each generated frame.
  • Network pipelines: Temporal self-attention models dynamics across frames, and motion frames from preceding clips are reused to improve cross-clip continuity.The temporal modules operate across the frame dimension at multiple resolutions.
  • Implementation: Reference features are extracted once and reused across repeated denoising iterations, avoiding substantial additional inference-time computation.The target image and motion frames enter ReferenceNet only once per generation.
  • Weak controls: Face Locator and Speed Layers provide weak spatial and velocity guidance, stabilizing motion without imposing strong pose constraints.Speed buckets approximate head rotation velocity because estimated speed labels are noisy.
  • Training strategies: Training proceeds through image pretraining, video training with temporal and audio layers, and a final stage that trains temporal and speed layers.The final stage omits audio-layer training because audio is correlated with expression, mouth motion, and head-movement frequency.

4 Experiments

Experiments evaluate EMO on HDTF and internet data through dataset protocols, qualitative comparisons, style and singing showcases, quantitative metrics, and a speed-layer ablation. The results emphasize expressive motion, identity preservation, video quality, and robustness across portrait styles and audio conditions.

  • Evaluation Setup: Experiments use HDTF and internet data, with HDTF split into 10% testing and 90% training and 1k internet clips of approximately 4 seconds.The HDTF split avoids character-ID overlap between training and test subsets.
  • Evaluation Setup: The evaluation compares EMO with Wav2Lip, SadTalker, DreamTalk, MakeItTalk, and Diffused Heads using qualitative and quantitative assessments.Diffused Heads is evaluated qualitatively because its released model was trained on CREMA and exhibited low resolution and error accumulation.
  • Qualitative Comparisons: EMO generates a greater range of head movements and more dynamic facial expressions than SadTalker and DreamTalk without direct blendshape or 3DMM motion control.The reported motions are directly driven by audio rather than intermediate motion signals.
  • Qualitative Comparisons: EMO produces approximately synchronized lip movements across realistic, anime, and 3D portrait styles using identical vocal audio inputs.Although trained only on realistic videos, the model is reported to animate a wide array of portrait types.
  • Qualitative Comparisons: Pronounced tonal audio elicits richer facial expressions and movements, while motion frames extend videos and preserve identity during substantial motion.The showcased singing clips last approximately 1 minute.
  • Quantitative Comparisons: EMO achieves lower FVD and improved FID scores, while E-FID reflects lively facial expressions and the 250-hour dataset further improves dynamics and expression variety.The model retains exceptional FVD and E-FID performance even when trained without the self-collected dataset.
  • Ablation Studies: The speed layers target consistency of head-motion frequency across contiguous clips, while speeds exceeding 1.5 can produce unnaturally rapid and jittery movements.Velocity variance measures average velocity variance across generated sequences.

5 Conclusion

The conclusion presents EMO as an audio-driven diffusion framework for expressive talking head videos. It reports improved expressiveness and identity consistency, with performance exceeding existing state-of-the-art methods.

  • Conclusion: EMO uses weakly conditioned audio2video diffusion to generate lifelike portraits without traditional intermediate signals.The framework drives animation directly from audio input.
  • Conclusion: EMO improves capture of human expressiveness and maintains consistent portrait identity across video frames.The conclusion attributes these properties to the framework's audio-driven animation process.
  • Conclusion: EMO outperforms existing state-of-the-art methods in the reported experiments, supporting its potential for rich audio-driven video content.

(Supplementary Material)

The supplementary material provides additional results and training-data discussion, followed by a candid discussion of EMO's limitations and future research directions.

  • Supplementary Material: The supplementary material expands the method's results and discusses its training dataset.
  • Supplementary Material: It also presents limitations of EMO and previews future research directions.

A.1 User study

The user study evaluates generated videos for lip synchronization and vividness, while experiments examine how denoising steps and sequential clip generation affect output quality and continuity.

  • User study: Twenty participants rated every method on the same image and audio for lip synchronization and vividness using a 1–5 scale.Participants were evenly split by gender, aged 20–60, and had diverse technical expertise.
  • Sequential generation: Motion frames improve consistency across sequentially generated clips, but transitions may still exhibit frame discontinuity without strong pose-sequence control.
  • Denoising steps: Fewer than 20 denoising steps produce temporal inconsistencies and frequent visual artifacts.
  • Denoising steps: Using 20–35 denoising steps somewhat reduces artifacts but can introduce noticeable jitter and instability.
  • Denoising steps: More than 35 denoising steps are evaluated as the preferred setting in the reported experiments.

B.1 Data overview

EMO is trained on diverse public and online video sources, with preprocessing designed to preserve coherent clips and facial annotations. The compiled data primarily improves facial expressions and video dynamics, while strong performance remains without the self-collected dataset.

  • Dataset composition: The training data includes 15.8 hours of HDTF, 16k VFHQ clips, and 250 hours of speech and singing data from online platforms.
  • Dataset contribution: EMO retains exceptional performance without the self-collected dataset, while the larger compiled dataset mainly enhances facial expressions and video content dynamics.
  • Preprocessing: Videos are segmented into shorter clips to maintain temporal coherence and scene consistency during training.
  • Preprocessing: Training clips are cropped around expanded facial bounding boxes, converted to 30 FPS, and labeled with facial regions, audio embeddings, and 6-DoF head pose.

C Limitations and future work

The paper identifies artifacts, controllability limits, and computational costs as limitations of EMO. Future directions include adding control signals for body parts, subtitles, and explicitly defined emotions.

  • Artifacts: Expressive audio can trigger unintended body-part generation because the head-focused training data contains only 3% frames featuring hands.
  • Artifacts: Internet videos with embedded subtitles can cause caption-like patterns to appear in generated frames.
  • Future work: Mask-like control inputs for face regions could be extended to guide body parts and subtitles.
  • Controllability: Audio-driven expression correlations may not match user expectations, motivating mechanisms that explicitly define emotions for more predictable results.
  • Efficiency: EMO generates 12 frames per 18 seconds on an A100 GPU under 40 denoising steps, making it more time-consuming than diffusion-free methods.
Loading 2402.17485v3…