Source-linked AI summary

FaceScape: a Large-scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction

Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, Xun Cao

arXiv:2003.13989v3cs.CV

TL;DR

Existing 3D face data and single-image methods provide insufficient detail or static, non-riggable geometry. FaceScape introduces a large high-fidelity dataset and learns expression-dependent details for single-image prediction, producing detailed riggable 3D faces.

  • Problem

    Existing 3D face datasets lack detail and scale, while prior detailed single-image reconstructions are not riggable across expressions.

  • Method

    FaceScape captures 18,760 scans and combines bilinear identity-expression modeling with neural prediction of dynamic details in a three-stage pipeline.

  • Results

    The system predicts detailed riggable 3D face models from a single image with plausible detailed geometry across various expressions.

  • Takeaways & Limitations

    FaceScape provides high-quality large-scale facial data and supports recovery of expression-dependent detailed geometry from a single image.

Abstract

from arXiv · show

In this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniformed. These fine 3D facial models can be represented as a 3D morphable model for rough shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different than the previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. The unprecedented dataset and code will be released to public for research purpose.

1. Introduction

FaceScape addresses the limited detail and scale of existing 3D face datasets with a large, high-fidelity dataset and a method for predicting detailed riggable faces from single images.

  • Motivation: Existing 3D face datasets lack sufficient 3D detail and scale, limiting learning-based methods that rely on 3D information.The paper motivates improved data for face tracking, recognition, reconstruction, and synthesis.
  • FaceScape Dataset: 18,760 high-quality 3D face models were captured from 938 subjects performing 20 specified expressions.A dense 68-camera array under controlled illumination recovers wrinkle- and pore-level geometry.
  • Representation: The scans are converted into topologically uniform base models for rough shape and displacement maps for detailed geometry.These representations are used to build bilinear models across identity and expression dimensions.
  • Evaluation: The generated bilinear model exceeds previous methods in representative ability.
  • Prediction System: A three-stage pipeline predicts detailed riggable face models from a single image: base fitting, displacement prediction, and dynamic-detail synthesis.The predicted model can be rigged to various expressions with plausible detailed geometry.
  • Dynamic Details: FaceScape models expression-dependent geometric variation as dynamic details learned by a deep neural network.The learned relationship targets small-scale changes such as wrinkles caused by expression changes.

2. Related Work

Prior 3D face datasets and single-view reconstruction methods are constrained by acquisition quality, model detail, or static expression-specific geometry. FaceScape combines high-quality, large-scale data with an animable detailed-face prediction objective.

  • 3D Face Datasets: Model-fitting datasets scale well but suffer from uncertain accuracy and limited detailed shape.
  • 3D Face Datasets: Depth sensors and scanners have limited spatial resolution, while sparse multi-view systems produce unstable and inaccurate reconstructions.
  • 3D Morphable Models: 3D morphable models represent face shape and texture in a vector space with explicit correspondences across models.They support model fitting, synthesis, image manipulation, and separate identity or expression control.
  • Single-view Shape Prediction: Single-view methods fit or regress 3D morphable models, but their limited representation power makes small-detail recovery difficult.Multi-layer refinement, bump maps, displacement maps, and conditional GANs have been used to enhance details.
  • FaceScape Advances: FaceScape is presented as larger and higher quality than prior public large-scale 3D face datasets.The paper reports a quantitative comparison in Table 1.
  • FaceScape Advances: Unlike prior work focused on static detailed shape, FaceScape targets detailed rigged 3D face recovery from a single image with expression-dependent wrinkles.

3. Dataset

FaceScape constructs a large, high-fidelity 3D face dataset from controlled multi-view capture and converts scans into a compact, topology-uniform representation. Its bilinear model preserves detailed geometry while achieving lower fitting error than prior parametric models.

  • 3.1. 3D Face Capture: 18,760 reconstructed meshes cover 938 mostly Asian subjects aged 16–70, each performing 20 expressions, with roughly 2 million vertices and 4 million triangle faces.The dataset records subject metadata including age, gender, and job information.
  • 3.1. 3D Face Capture: The capture system uses 68 synchronized DSLR cameras, including 30 cameras recording 8K front-facing images and others recording 4K side images.Camera shutters are triggered within 5ms under controlled multi-view acquisition.
  • 3.2. Topologically Uniformed Model: Scans are registered to a template using facial landmarks and NICP, producing deformed meshes with shared uniform topology and minor accuracy loss.This topology-uniform registration enables consistent representation across the captured faces.
  • 3.2. Topologically Uniformed Model: FaceScape represents rough geometry with base shapes and middle- and fine-scale geometry with UV-space displacement maps.The displacement values encode signed distances between corresponding points on the base and raw meshes.
  • 3.2. Topologically Uniformed Model: The two-layer representation uses roughly 2% of the original mesh data size while maintaining mean absolute error below 0.3mm.This compact representation retains the dataset’s detailed facial shape information.
  • 3.3. Bilinear Model: The bilinear model decomposes 52-expression blendshape tensors across 938 identities into a core tensor and low-dimensional identity and expression components.It generates vertex positions from identity and expression parameters, then achieves much lower fitting error than FaceWarehouse and lower error than FLAME with fewer expression parameters.

4. Detailed Riggable Model Prediction

The method addresses the gap between detailed but expression-specific reconstructions and riggable but coarse models by predicting expression-dependent dynamic details from a single image. It combines bilinear base-model fitting, displacement-map prediction, and dynamic-detail synthesis.

  • Prior detailed reconstructions are not expression-riggable, whereas parametric riggable models recover only rough geometry.
  • The pipeline has three stages: base model fitting, displacement map prediction, and dynamic details synthesis.
  • 4.1. Base Model Fitting: The bilinear model separates identity and expression parameters, enabling a rough riggable model from fitted identity parameters.
  • 4.2. Displacement Map Prediction: Detailed geometry is represented with multiple displacement maps for 20 basic expressions because one map cannot capture expression-dependent dynamic details.
  • 4.2. Displacement Map Prediction: Static details are linked mainly to texture, while dynamic details vary with expression and are modeled using deforming maps and textures as CNN inputs.
  • 4.3. Dynamic Detail Synthesis: For arbitrary expressions, displacement maps are synthesized by weighted combinations of neutral and key-expression maps using activation-based masks.

5. Experiments

Experiments evaluate training, rigging, ablations, and comparisons with prior methods. The predicted models retain detailed wrinkles across expressions and achieve visually and quantitatively better static reconstruction results than previous methods.

  • The network trains on 17,760 displacement maps from 888 subjects, with 50 subjects reserved for testing.
  • 5.3. Ablation Study: The predicted faces are rigged to five expressions and show photo-realistic detailed wrinkles in the rigged outputs.
  • 5.3. Ablation Study: Dynamic-detail synthesis adds expression-caused wrinkles that are absent when only the source-image displacement map is used.
  • 5.3. Ablation Study: Removing the deforming map produces few expression-caused details compared with the full method.
  • 5.4. Comparisons to Prior Works: Compared with three previous methods, the proposed method predicts the lowest static reconstruction error in the Figure 7 comparison.
  • 5.4. Comparisons to Prior Works: The method is visually and quantitatively better than prior methods, attributed to the bilinear model’s representation power and predicted details.

6. Conclusion

FaceScape combines a detailed, large-scale 3D facial dataset with single-image prediction of riggable faces and high-fidelity dynamic detail synthesis.

  • FaceScape provides the highest geometry quality and largest model amount among previous public large-scale 3D face datasets.
  • The dataset supports detailed face representation through a uniformed base model and displacement maps, enabling comparisons of predicted details on common base meshes.
  • The final model recovers wrinkles in rigged expressions, unlike variants without deforming maps or dynamic details.
  • The method predicts a detailed riggable 3D face from a single image and achieves high fidelity in dynamic detail synthesis.

A. Animation

The animation pipeline fits a base face model from an image, changes its expression, and synthesizes expression-dependent details for the resulting rigged face.

  • Animation: Displacement maps represent middle- and fine-scale details omitted by the topology-uniformed base model.
  • Animation: Base-model fitting includes landmark alignment under a weak-perspective camera, with scale, rotation, and translation parameters.
  • Animation: Pixel-level consistency matches synthetic and input images over pixels corresponding to frontal fitted vertices.
  • Animation: Identity, expression, and albedo parameters are regularized with multivariate Gaussian priors and optimized alternatively.

D. Facial Capture System

FaceScape uses a dense multi-view capture system with two acquisition frameworks, while supplementary figures show synthesis, displacement prediction, and failure cases.

  • D. Facial Capture System: Image synthesis in another expression is demonstrated with detailed shading.
  • D. Facial Capture System: A failure case shows poor aquiline-nose recovery because the feature is uncommon in the dataset.
  • D. Facial Capture System: Another failure case attributes an incorrect displacement map to occlusion.
  • D. Facial Capture System: The capture system uses a 68-camera DSLR array, controlled lighting, and centralized control across two acquisition frameworks.
  • D. Facial Capture System: Predicted displacement maps are compared with ground truth in the supplementary results.

E. More Results

Additional results show photorealistic detailed reconstructions, plausible expression rigging with synthesized details, and broad variation across identities and expressions.

  • E. More Results: Supplementary results recover 3D faces with photo-realistic details and synthesize details when the faces are rigged to other expressions.
  • E. More Results: The predicted and ground-truth displacement maps are compared in an extended figure.
  • E. More Results: Each subject has 20 captured expressions, while additional neutral-expression subjects illustrate identity diversity.
  • E. More Results: Expression transfer uses base-model fitting, target-expression generation, vertex-guided pixel warping, and shading adjustment for expression-caused details.

H. Failure Case

The paper presents failure cases alongside supplementary visualizations of expressions, identities, and topologically uniformed face models.

  • Failure cases of the proposed method are shown in Figure 12.
  • Figure 14 provides additional results extending Figure 6 in the main paper.
  • Figure 15 depicts the 20 specified expressions performed by the subjects.
  • Figure 16 shows different identities, pairing subject images with processed topologically uniformed models.
Loading 2003.13989v3…