Source-linked AI summary

FaceVerse: a Fine-grained and Detail-controllable 3D Face Morphable Model from a Hybrid Dataset

Lizhen Wang, Zhiyuan Chen, Tao Yu, Chenguang Ma, Liang Li, Yebin Liu

arXiv:2203.14057v3cs.CV

TL;DR

FaceVerse addresses the limited scale or fidelity of face datasets and the lack of controllable facial details in prior 3DMMs. It combines hybrid East Asian datasets with a coarse-to-fine model, conditional StyleGAN detail generation, and differentiable-rendering fitting. Experiments report superiority over state-of-the-art methods in 3D face fitting and monocular reconstruction, while the method remains limited by East Asian data and sparse scans of older people.

  • Problem

    Existing 3DMMs are constrained by dataset scale or fidelity, and prior methods generally cannot parameterize detailed facial geometry and texture.

  • Method

    FaceVerse builds a coarse base model from large-scale RGB-D data and a conditional StyleGAN fine model from high-fidelity scans, then fits images through differentiable rendering.

  • Results

    FaceVerse outperforms state-of-the-art methods in 3D face model fitting and monocular face reconstruction.

  • Takeaways & Limitations

    The model supports adjustment of both basic facial attributes and detailed geometry and texture while preserving base facial attributes.

  • Takeaways & Limitations

    The dataset contains only East Asian faces, and limited detailed scans of old people restrict generation of thick beards and deep wrinkles.

Abstract

from arXiv · show

We present FaceVerse, a fine-grained 3D Neural Face Model, which is built from hybrid East Asian face datasets containing 60K fused RGB-D images and 2K high-fidelity 3D head scan models. A novel coarse-to-fine structure is proposed to take better advantage of our hybrid dataset. In the coarse module, we generate a base parametric model from large-scale RGB-D images, which is able to predict accurate rough 3D face models in different genders, ages, etc. Then in the fine module, a conditional StyleGAN architecture trained with high-fidelity scan models is introduced to enrich elaborate facial geometric and texture details. Note that different from previous methods, our base and detailed modules are both changeable, which enables an innovative application of adjusting both the basic attributes and the facial details of 3D face models. Furthermore, we propose a single-image fitting framework based on differentiable rendering. Rich experiments show that our method outperforms the state-of-the-art methods.

1. Introduction

FaceVerse addresses limited dataset scale or fidelity and non-changeable facial details with a hybrid dataset and coarse-to-fine, detail-controllable model.

  • 1. Introduction: A coarse-to-fine structure first builds a base parametric model, then generates detailed geometry and texture using conditional StyleGAN.The detailed module uses UV maps from the base model and additional latent code and noise.
  • 1. Introduction: The conditional StyleGAN uses multi-scale input features and a normal discriminator to constrain outputs and enrich geometric details.These additions distinguish the generator from the original StyleGAN design.
  • 1. Introduction: FaceVerse combines a large-scale RGB-D dataset with high-fidelity scans to improve both generalization ability and facial fidelity.The coarse dataset supports the base model, while scans enrich detailed geometry and texture.
  • 1. Introduction: Both basic facial attributes and detailed facial features are parameter-changeable, unlike previous methods that leave detailed geometry and texture fixed.The design preserves base attributes while allowing facial-detail adjustment.
  • 1. Introduction: FaceVerse provides released East Asian face-modeling resources, including pretrained models and a detailed dataset for research.The paper positions these resources as a tool for East Asian face modeling.

2. Related Work

Prior 3D face models improved dataset diversity, representation flexibility, and reconstruction, but detailed facial features generally remained non-parameterizable; FaceVerse targets this gap.

  • 2. Related Work: Existing 3DMMs use PCA, multilinear, or nonlinear representations to model shape, expression, identity, and texture with increasing flexibility.These approaches build on progressively richer datasets and parameter spaces.
  • 2. Related Work: Monocular reconstruction evolved from landmark-based fitting to CNN parameter regression and differentiable-rendering methods for single-image fitting.These methods use 3DMMs to formulate reconstruction as a model-fitting problem.
  • 2. Related Work: Layered refinement methods improve facial detail after rough reconstruction, but their detailed facial features are still not parameter-changeable.Refinement may adjust depth, displacement, albedo, or normal maps without exposing detail controls.
  • 2. Related Work: FaceVerse differs by combining a hybrid dataset, a PCA-based coarse model, a conditional StyleGAN fine model, and differentiable-rendering fitting.Its pipeline reconstructs high-fidelity faces and allows adjustment through detailed parameters.

3. Hybrid Dataset

The hybrid dataset pairs efficient structured-light RGB-D capture for many identities with a 128-camera system for high-fidelity scans and registration.

  • 3.1. Coarse Dataset: Structured-light sensors collect coarse RGB-D data efficiently, enabling approximately five fused frames per volunteer and large-scale identity coverage.Frames are fused by ICP registration into a smooth facial point cloud.
  • 3.1. Coarse Dataset: A template mesh is aligned to fused point clouds using projected landmarks followed by non-rigid ICP to produce topologically uniform models.The resulting dataset includes documented age and gender distributions.
  • 3.2. Detailed Dataset: The detailed capture system uses 128 DSLR cameras arranged on 16 pillars, synchronously recording 128 images at 6000 × 4000 resolution.The multi-view setup is designed for high-fidelity 3D scan collection.
  • 3.2. Detailed Dataset: Detailed scans are registered to a uniform template using landmarks, the coarse base model, UV up-sampling, and subsequent detailed deformation.The UV resolution increases from 200 × 200 to 1024 × 1024 before detailed registration.

4. FaceVerse Model

FaceVerse combines a PCA-based base model with conditional StyleGAN refinement to generate controllable facial geometry and texture details, then fits the model to a single image through differentiable rendering.

  • FaceVerse Model: The coarse-to-fine model builds a PCA base from large-scale data and a detailed model from high-fidelity data, followed by differentiable-rendering-based single-image fitting.The fitting pipeline has base-model, detailed-model, and expression-refinement phases.
  • Base Model Generation: The base model preserves 100 shape and 200 texture principal components, adds 20 detailed-dataset shape components for cheeks, and uses 64 expression components.Its shape and texture dimensions are m = 120 and k = 200, while the expression model uses l = 64.
  • Base Model Generation: The base model fits faces across ages and genders but lacks fine geometry and texture, which are supplied by the subsequent detailed model.This division combines the base model's fitting generalization with the detailed model's refinement role.
  • Detailed Model Generation: Conditional StyleGAN refines base UV maps with geometry-texture conditioning, latent codes, noise, and discriminators while retaining basic facial attributes.The detail generator jointly processes geometry and texture, and incomplete-supervision training includes detailed pairs and coarse-derived conditional UV maps.
  • Detailed Model Generation: A second conditional generator refines expression-related geometry from detailed geometry and base expression offsets while preserving basic shape and expression.Its output is controlled by zexp and injected noise, enabling detailed changes such as a smiling mouth.
  • Coarse-to-Fine Single-Image Fitting: The three-phase fitting procedure optimizes base parameters, detail latent variables, and expression latent variables using differentiable rendering and image-based losses.Base fitting optimizes shape, texture, expression, pose, and lighting; later phases optimize detail and expression generator variables while fixing earlier outputs as specified.

5. Experiments

FaceVerse is evaluated for monocular reconstruction, 3D model fitting, detail control, and module effectiveness. It is compared with prior methods and tested through ablations of the coarse dataset and detailed generators.

  • 5.1. Qualitative Results: FaceVerse supports detail transfer by combining base parameters from one fitted image with detail parameters from another.This produces faces retaining the source’s basic shape while adopting target details such as bigger eyes, thinner lips, or a broader nose.
  • 5.2. Comparisons to Prior Works: FaceVerse shows better qualitative fitting of rough facial shape and facial details than FaceScape, Hifi3DFace, DECA, and 3DDFAv2.The comparison attributes this performance to the large-scale base model and GAN-based detail generator.
  • 5.2. Comparisons to Prior Works: The base model achieves the best quantitative 3D model-fitting performance against FaceScape, Hifi3DFace, and BFM on 357 test scans from 17 people.The scans are fixed at a length of 200mm, and fitting uses back-propagation through ICP; the detailed model is excluded because it requires additional texture input.
  • 5.3. Ablation Study: The detail generator adds reasonable facial details, while the expression refinement generator addresses its limited expression description power.The ablation also compares a detailed model trained without the normal discriminator.
  • 5.3. Ablation Study: Introducing the coarse dataset significantly improves fitting ability compared with a base model trained only on the detailed dataset.The ablation model uses 50 shape principal components and the same expression principal components as the full base model.

6. Discussion and Conclusion

FaceVerse provides a fine-grained, detail-changeable 3D face morphable model built from a hybrid dataset and coarse-to-fine architecture. The discussion identifies regional and age-related limitations, while the conclusion reports superiority in 3D fitting and monocular reconstruction.

  • Limitations: Performance declines when fitting faces from regions other than East Asia because the dataset contains only East Asian faces.This is an explicit geographic scope limitation of the method.
  • Limitations: The detailed model lacks sufficient scans of older people and cannot generate extreme textures such as thick beards or deep wrinkles.These limitations are illustrated in Figure 14.
  • Potential Social Impact: Single-image reconstruction can be used to generate a 3D fake model of a person, requiring careful consideration before deployment.The paper identifies this as a potential social impact.
  • Conclusion: FaceVerse combines a large-scale coarse dataset with a high-fidelity detailed dataset and uses conditional StyleGAN to control facial geometry and texture details.The detailed model preserves basic facial attributes from the base model.
  • Conclusion: Experiments demonstrate superiority over state-of-the-art methods in 3D face model fitting and monocular face reconstruction.The conclusion presents FaceVerse as a potential tool for face-related research.
Loading 2203.14057v3…