Source-linked AI summary
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, Xun Cao
TL;DR
Existing audio-driven talking-face methods either neglect facial emotion or cannot be applied to arbitrary subjects. EAMM combines audio-driven unsupervised motion with linearly additive emotion displacements from an emotion source video, and experiments report expressive one-shot animations with realistic emotion patterns.
Problem
Existing audio-driven talking-face methods either neglect facial emotion or cannot be applied to arbitrary subjects, limiting one-shot expressive animation.
Method
EAMM combines an Audio2Facial-Dynamics module with an Implicit Emotion Displacement Learner using an emotion source video.
Results
Qualitative and quantitative experiments show more expressive animation results than state-of-the-art methods, with realistic emotion patterns on arbitrary subjects.
Takeaways & Limitations
Emotion dynamics can be formulated as transferable motion patterns and applied to arbitrary audio-driven talking faces.
Takeaways & Limitations
The method can produce unnatural transferred emotion patterns, neglects audio-emotion correlation, and generates weak mouth-region emotion dynamics.
Abstract
from arXiv · showhide
Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware Motion Model (EAMM) to generate one-shot emotional talking faces by involving an emotion source video. Specifically, we first propose an Audio2Facial-Dynamics module, which renders talking faces from audio-driven unsupervised zero- and first-order key-points motion. Then through exploring the motion model's properties, we further propose an Implicit Emotion Displacement Learner to represent emotion-related facial dynamics as linearly additive displacements to the previously acquired motion representations. Comprehensive experiments demonstrate that by incorporating the results from both modules, our method can generate satisfactory talking face results on arbitrary subjects with realistic emotion patterns.
1 INTRODUCTION
The paper targets one-shot expressive talking-face animation from a neutral identity image, audio, pose, and emotion source video. EAMM combines audio-driven facial dynamics with learned emotion-related displacements to control emotional motion on arbitrary subjects.
- One-shot talking-face methods often synthesize synchronized mouth shapes without modeling emotion, while long source videos are unavailable in many scenarios.
- Fixed emotion labels provide coarse discrete representations, while audio-only emotion inference can be ambiguous and unreliable for general speech.
- EAMM takes a neutral identity image, speech audio, predefined pose, and emotion source video as inputs for emotional talking-face generation.
- Audio2Facial-Dynamics: The Audio2Facial-Dynamics module maps audio and pose to unsupervised key points and first-order dynamics, then reconstructs talking-face frames.
- Implicit Emotion Displacement Learner: The Implicit Emotion Displacement Learner extracts emotion-related displacements from face-related motion representations and linearly combines them with audio-driven motion.
- The reported contributions are the Audio2Facial-Dynamics module, the Implicit Emotion Displacement Learner, and EAMM’s one-shot emotion-controlled talking-head animation.
2 RELATED WORK
Prior work spans audio-driven, emotional, and video-driven facial animation, but existing approaches face person-specific training costs, limited emotion control, or restricted generalization. EAMM instead uses an emotion source video for one-shot control.
- Audio-Driven Talking Face Generation: Audio-driven talking-face methods include person-specific and person-agnostic approaches, with person-specific systems requiring substantial training videos or time.
- Emotional Talking Face Generation: Prior emotional talking-face methods use learned lip-emotion relationships, temporal discriminators, emotion labels, or audio decomposition, but cited approaches lack semantic manipulation or generalization to unseen characters and audios.
- Emotional Talking Face Generation: EAMM uses a source video to disentangle emotion information and achieve emotion control in the one-shot setting.
- Video-driven Facial Animation: Video-driven facial animation reenacts facial motion and is related to audio-driven talking-face generation, while traditional approaches require prior target knowledge or manual labels.
3 METHOD
EAMM combines an Audio2Facial-Dynamics module for neutral audio-driven motion with an Implicit Emotion Displacement Learner that transfers emotion through face-related motion displacements.
- Audio2Facial-Dynamics Module: A2FD extracts identity, audio, and pose features, then an LSTM decoder predicts ten unsupervised key-points and their local affine jacobians over a sequence.The key-points are zero-order representations, while jacobians encode first-order local motion around each key-point.
- Audio2Facial-Dynamics Module: A pretrained key-point detector supplies initial motion representations, while self-supervised key-point and perceptual objectives train the audio-based module.The perceptual objective compares reconstructed and target frames using pretrained VGG features.
- Audio2Facial-Dynamics Module: The A2FD pipeline combines predicted motion with source-image representations to estimate dense warping fields, which an image generator converts into output frames.Masks weight multiple warping flows before the final dense field is passed with the source image to the generator.
- Implicit Emotion Displacement Learner: The face region is affected by three face-related key-points, whose relative displacements are treated as approximately linearly additive for transferring emotion across subjects.Adding emotional-minus-neutral displacements to another person’s motion transfers dynamics, but can also introduce artifacts around the face boundary and mouth.
- Implicit Emotion Displacement Learner: The Implicit Emotion Displacement Learner extracts emotion-related displacements for those face-related key-points and jacobians, complementing audio-driven mouth motion.The method addresses the need to decouple emotion from identity, pose, and speech information contained in an emotion source.
4 RESULTS
EAMM is evaluated against state-of-the-art methods using quantitative metrics, qualitative comparisons, a user study, and ablations. The results show strong audio-visual synchronization, emotional animation, and effectiveness of the proposed components.
- 4.1 Evaluation: EAMM is compared with ATVG, Speech-driven-animation, Wav2Lip, MakeItTalk, and PC-AVS on LRW and MEAD.The evaluation uses mouth synchronization, facial landmark, image-quality, and user-study measures.
- 4.1 Evaluation: EAMM achieves the highest score among all metrics on MEAD and most metrics on LRW, with satisfactory audio-visual synchronization.Wav2Lip obtains the highest SyncNet confidence on LRW, while EAMM remains comparable with ground truth.
- 4.1 Evaluation: EAMM generates vivid emotional animation with natural head movements and accurate mouth shapes, whereas competing methods show weaker or unstable emotion dynamics.Wav2Lip and PC-AVS produce competitive mouth motions, but they do not jointly provide the same emotion and pose behavior.
- 4.2 User Study: The user study evaluates lip synchronization, facial-expression naturalness, video quality, and emotion classification using participant ratings and muted-video classification.Twenty participants score three aspects from 1 to 5 and classify emotions across eight categories.
- 4.2 User Study: 58% emotion-classification accuracy is achieved by EAMM, while the method also obtains the highest scores over the three evaluated aspects apart from real data.The ablation study further evaluates emotion accuracy, M-LMD, F-LMD, and SSIM across five variants and EAMM.
- 4.3 Ablation Study: The Implicit Emotion Displacement Learner and its three components are effective, with data augmentation especially important for accurate emotional dynamics without sacrificing identity.The feature-based model produces unstable face shapes and less obvious emotion, indicating that emotion is not well disentangled at the feature space.
5 CONCLUSION
EAMM generates one-shot emotional talking faces by transferring emotion dynamics from a source video to arbitrary audio-driven faces. Qualitative and quantitative experiments report more expressive animation than state-of-the-art methods.
- EAMM transfers emotion dynamics from an emotion source video to arbitrary audio-driven talking faces.
- The method is intended for applications including video conferencing and digital avatars.
- Qualitative and quantitative experiments show more expressive animation results than state-of-the-art methods.
6 ETHICAL CONSIDERATIONS
The method targets emotional talking-face animation for digital entertainment and advanced teleconferencing, while acknowledging possible misuse on social media.
- The method is intended for digital entertainment and advanced teleconferencing systems.
- The authors acknowledge that the technology may be misused on social media, causing negative societal impacts.