Source-linked AI summary

Neural Head Avatars from Monocular RGB Videos

Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, Justus Thies

arXiv:2112.01554v2cs.CVcs.GR

TL;DR

Head-avatar reconstruction from monocular RGB video must recover consistent 3D geometry and appearance despite complex facial dynamics and missing depth. Neural Head Avatars combine a morphable head model with neural geometry refinement and dynamic texture networks, producing controllable, photorealistic avatars that support large pose and viewpoint changes. The method is subject-specific and has author-reported limitations for unconstrained mouth-cavity regions and physical hair effects.

  • Problem

    Monocular head reconstruction is difficult because facial dynamics are complex and single-view input lacks 3D information, while some prior methods lack consistent full-head geometry and standard rasterization compatibility.

  • Method

    The method combines FLAME for coarse face shape and expression with coordinate-based neural networks that refine explicit geometry and synthesize dynamic textures, optimized with color-dependent and color-independent energy terms.

  • Results

    The resulting controllable 4D avatar reconstructs human-head geometry and appearance from monocular RGB sequences, remains robust to large pose, viewpoint, and expression changes, and outperforms state-of-the-art head-avatar methods qualitatively and quantitatively.

  • Takeaways & Limitations

    Explicit geometry combined with a deep appearance network provides photorealism and generalizability for dynamic head avatars compatible with standard graphics pipelines.

  • Takeaways & Limitations

    The method is person-specific, requires optimization for every new actor, and can lose visual quality in mouth-cavity regions when poses and expressions lie far outside the training corpus; it also does not synthesize physical hair effects.

Abstract

from arXiv · show

We present Neural Head Avatars, a novel neural representation that explicitly models the surface geometry and appearance of an animatable human avatar that can be used for teleconferencing in AR/VR or other applications in the movie or games industry that rely on a digital human. Our representation can be learned from a monocular RGB portrait video that features a range of different expressions and views. Specifically, we propose a hybrid representation consisting of a morphable model for the coarse shape and expressions of the face, and two feed-forward networks, predicting vertex offsets of the underlying mesh as well as a view- and expression-dependent texture. We demonstrate that this representation is able to accurately extrapolate to unseen poses and view points, and generates natural expressions while providing sharp texture details. Compared to previous works on head avatars, our method provides a disentangled shape and appearance model of the complete human head (including hair) that is compatible with the standard graphics pipeline. Moreover, it quantitatively and qualitatively outperforms current state of the art in terms of reconstruction quality and novel-view synthesis.

1. Introduction

Neural Head Avatars reconstruct a complete, subject-specific human head from short monocular RGB video using explicit geometry and appearance modeling. The representation supports controllable animation, novel viewpoints, poses, and expressions while preserving photorealism.

  • Motivation: Monocular head-avatar reconstruction is difficult because facial dynamics are complex and single-view input lacks 3D information.Existing talking-head methods can avoid explicit surface priors but still lack consistent full-head shape reconstruction and compatibility with standard rasterization pipelines.
  • Contribution: Neural Head Avatars explicitly represent complete human-head shape and appearance, including hair, for triangular-mesh graphics pipelines.The representation is subject-specific and uses an explicit surface model rather than relying only on image synthesis.
  • Method: The method combines FLAME as a coarse shape-and-expression proxy with coordinate-based MLPs that refine geometry and synthesize dynamic textures.Color-dependent and color-independent energy terms disentangle surface shape from color detail during optimization.
  • Results: The resulting controllable 4D avatar generates novel poses and expressions while preserving high photorealism and visual quality under large viewpoint changes.The avatar is learned from a short monocular RGB video sequence.

2. Related Work

Prior head-avatar methods use image-based, implicit, or explicit representations with different strengths and limitations. Neural Head Avatars jointly model subject-specific geometry and dynamic texture in an explicit mesh-based pipeline to support consistent novel-view synthesis.

  • Image-based models: Image-based models synthesize new poses and expressions without explicit 3D representations but lack geometric and temporal consistency.They learn three-dimensional deformations through two-dimensional image warping or encoder-decoder synthesis.
  • Implicit models: Implicit models represent geometry with implicit surfaces or volumetric representations, including deformable feature grids and neural radiance fields.These approaches replace explicit meshes with volumetric or function-based scene representations.
  • Explicit models: Explicit models commonly use triangular meshes and morphable models as priors for reconstructing plausible head shapes and facial movements from incomplete or noisy data.Some methods add mesh displacements for fine wrinkles and hair, but the cited related work does not produce photorealistic outputs under novel views.
  • Neural Head Avatars: Neural Head Avatars jointly optimize subject-specific geometry with facial detail and hair structure alongside a photorealistic dynamic texture in a coordinate-based neural network.The texture-based representation links appearance to geometry and supports spatial consistency, unseen poses and expressions, and intuitive jaw, neck, and eye control.

3. Method

The method reconstructs an explicit 4D head avatar by refining a FLAME mesh with pose-dependent geometry offsets and synthesizing view-, pose-, and expression-dependent texture. It jointly optimizes geometry and appearance from monocular RGB video using landmark, normal, semantic, photometric, perceptual, and regularization terms.

  • Explicit Neural Head Representation: The avatar outputs a classical triangle mesh and a texture function, enabling rendering through the standard graphics pipeline.The mesh is defined by vertices V and faces F, while the texture function assigns RGB values to surface points.
  • Explicit Neural Head Representation: A FLAME template provides the explicit surface topology, while Geometry MLP G predicts pose-dependent offsets and Texture MLP T predicts surface colors.T conditions appearance on canonical surface coordinates, expression, pose, and local rendered normal information.
  • Explicit Neural Head Representation: The template topology is subdivided, lower-neck faces are removed, and mouth-cavity faces are added, increasing the mesh from 5023 to 16227 vertices.
  • Optimization based on Monocular RGB Data: Geometry uses appearance-independent landmark, normal, semantic, and regularization terms to disentangle surface shape from color detail.Landmark matching includes eyelid-relative distances, normal matching targets image-space Laplacians for high-frequency detail, and semantic matching covers skin, neck, eyes, ears, and hair.
  • Optimization based on Monocular RGB Data: Appearance optimization combines dense photometric and perceptual energies, with perceptual distance computed from face-detector image features.The joint objective separates geometry terms from appearance terms while regularization promotes smooth surfaces and view-consistent textures.
  • Optimization based on Monocular RGB Data: Initialization first estimates camera, shape, expression, and pose with FLAME tracking, then optimizes geometry, texture, and finally their joint objective.Geometry is initialized using Egeom before texture parameters are optimized with Eapp.

4. Results

The evaluation uses synthetic and real datasets to assess geometry, appearance, and novel pose, expression, and viewpoint synthesis. Results show accurate reconstruction, improved synthesis quality, robustness to large rotations, and limitations in unconstrained regions and physical hair effects.

  • Datasets: The evaluation combines synthetic sequences with ground-truth geometry and real conversational sequences covering varied hairstyles and head visibility.The synthetic dataset contains 200 training and 210 validation frames; the real dataset contains 750 training and 750 validation frames per sequence.
  • Geometry Reconstruction Quality: Geometry quality is measured using normal angular error and single-sided Hausdorff distance for full-head and facial-region meshes against FLAME.The synthetic evaluation uses 210 validation frames and compares predicted geometry with ground-truth meshes and normals.
  • Novel Pose and Expression Synthesis: The animatable geometric backbone supports novel expression and pose synthesis by optimizing validation-frame expression and pose parameters while keeping the avatar fixed.This analysis-by-synthesis evaluation compares the resulting reconstructions with recent talking-head methods.
  • Novel Pose and Expression Synthesis: The method consistently outperforms related approaches, producing higher detail and better expression conservation in qualitative comparisons.The appearance reconstruction is quantitatively evaluated over validation frames on all subjects in the real dataset.
  • Novel Viewpoint Synthesis: Explicit geometry reconstructs realistic full-head shapes, including longer hair and fine facial details, enabling robust novel-view synthesis under large rotations.Related methods exhibit significant artifacts during novel viewpoint synthesis, whereas the proposed method preserves texture detail.
  • Discussion: The method does not model physical effects such as floating or deforming hair and can lose inner-mouth and teeth quality for poses or expressions far outside training.It is also person-specific, requiring approximately 7 hours of optimization on two Nvidia A100 GPUs at 512×512 resolution.

5. Conclusion

Neural Head Avatars reconstruct complete human heads with explicit geometry and photorealistic appearance, remaining robust across large pose, view, and expression changes. The approach combines explicit geometry with a deep appearance network and outperforms prior head-avatar methods qualitatively and quantitatively.

  • The method accurately reconstructs human-head geometry and appearance from monocular RGB sequences.
  • The resulting 4D avatars remain robust under large pose, viewpoint, and expression changes.
  • Static vertex offsets can produce bulky necks, while dynamic pose-conditioned geometry improves neck reconstruction.
  • Rare artifacts occur with misaligned eyelids, extreme mouth poses, or ear shapes that strongly deviate from the statistical average.
  • Explicit geometry combined with a deep appearance network improves photorealism and generalizability over implicit spatial representations.

A. Implementation Details

The method uses an updated FLAME head model as its geometric backbone and modifies its topology to better represent the complete head.

  • The updated FLAME model is subdivided, trimmed at the lower neck, and closed around the mouth cavity.These changes increase the mesh from 5,023 to 16,227 vertices.

A.2. Network Architectures

Two SIREN-based MLPs refine FLAME geometry and synthesize dynamic textures, conditioning their outputs on surface information and facial pose or expression.

  • Geometry MLP G: The Geometry MLP G refines FLAME meshes with facial detail and hair geometry, using vertex coordinates, embeddings, and neck-pose parameters.
  • Geometry MLP G: G predicts three-dimensional vertex offsets through FiLM-conditioned sinusoidal layers whose frequencies and phase shifts are pose-dependent.
  • Geometry MLP G: Dynamic geometry is restricted to the neck region and smoothly blended with an unconditioned prediction to preserve spatial consistency.
  • Texture MLP T: The Texture MLP T uses surface coordinates, interpolated UV embeddings, and FLAME pose and expression features to synthesize view- and expression-dependent texture.
  • Network design: The architecture uses SIREN MLPs with FiLM conditioning and mapping networks that generate dynamic activation frequencies and phase shifts.

A.3. Optimization from Monocular RGB Data

The optimization pipeline initializes geometry and texture before jointly fitting the avatar to monocular RGB sequences with geometric, photometric, perceptual, semantic, and regularization terms. It supports detailed reconstruction, novel-view synthesis, and pose-expression reenactment, while comparisons and examples document both strengths and failure cases.

  • Optimization objectives: The objective combines landmark, regularization, photometric, perceptual, semantic, and geometric terms during optimization.
  • Geometry regularization: Region-specific regularization weights remain unchanged across subjects while supporting varied geometries such as different hairstyles.
  • Geometry regularization: Geometry regularization includes FLAME-prior, Laplacian, surface-consistency, and edge-length terms to control smoothness and pose-dependent variation.
  • Evaluation: The method reports novel-view synthesis comparisons across rotations and identities, emphasizing spatial consistency, texture detail, and face-recognition similarity.
  • Optimization procedure: The pipeline uses staged initialization followed by joint optimization, with Adam for frame-agnostic parameters and SGD for frame-specific components.
  • Initialization and optimization: Geometry initialization aligns FLAME to the training sequence, then optimizes geometry and frame-specific pose and expression parameters before texture training.
  • Avatar reenactment: Driving-sequence poses and expressions can be transferred to optimized avatars, enabling faithful reenactment across subjects.

B.2. Novel View Synthesis

The method is qualitatively compared with related approaches for novel viewpoints across subjects and quantitatively evaluated using face-recognition feature similarity. It outperforms related approaches for pitch angles within ±15°, while all models degrade at large yaw angles.

  • Qualitative comparisons: The method is qualitatively compared with related approaches for novel viewpoints across different subjects.The comparisons are reported in Figures 14 and 12, with synthesized facial regions considered when comparing against paGAN overlays.
  • Quantitative evaluation: CSIM scores are averaged over validation frames from two real-data sequences to evaluate identity preservation without novel-view ground truth.The metric compares pretrained face-recognition feature vectors between front-facing ground truth and novel-view predictions.
  • Quantitative evaluation: Our method outperforms related approaches consistently for pitch angles in a range of ±15°.The evaluation uses cosine similarity of latent feature vectors from a pretrained face-recognition network.
  • Viewpoint limits: At large yaw angles (≥±30°), scores for all considered models decrease rapidly, but qualitative comparisons show high identity preservation in that range.The large-yaw evaluation exposes a shared degradation while retaining a qualitative advantage for the proposed method.

B.4. Synthesis Results for Non-Caucasian Subjects

The method also produces faithful geometry and visually plausible synthesis for a non-caucasian subject, indicating that the reported synthesis quality extends to this example.

  • Subject diversity: For a person of color, the method reconstructs head geometry faithfully and renders visually plausible images.The result is presented as a validation example for non-caucasian subjects.

B.5. Geometry Evaluation

Geometry is evaluated against separately captured multi-view stereo recordings, but the comparison is restricted to the neutral face region because hair and facial dynamics are not reliably reproduced across recordings.

  • Evaluation setup: The predicted geometry is compared with multi-view stereo recordings of the corresponding real identities.The multi-view stereo data was captured separately using a handheld DSLR camera.
  • Evaluation scope: The comparison covers only the neutral pose of the face region because separately recorded hairstyles and face dynamics are unreliable matches.This restriction excludes hair and dynamic facial behavior from the geometry comparison.

B.6. Energy Ablation

The ablation shows that additional energy terms improve different aspects of synthesis, including color stability, semantic alignment, and texture detail.

  • Energy-term effects: Ephot prevents color shifts in the synthesis results.This identifies the color-dependent photometric term’s observed effect in the ablation.
  • Energy-term effects: Esemantic improves alignment of overlapping semantic regions within the avatar silhouette.The improvement concerns alignment between semantic regions during synthesis.
  • Energy-term effects: Eperc produces textures with more detail.The perceptual term’s effect is observed in the ablated synthesis results.
  • Ablation setup: The ablation evaluates how further energy terms affect synthesis results using reference normals as optimization inputs.The setup is identified in Figure 19.
Loading 2112.01554v2…