Source-linked AI summary
GNM Head: A Generative aNthropometric Model of the human head
Stylianos Ploumpis, Jan Bednarik, Gaspard Zoss, Ruslan Guseinov, Luca Prasso, Prashanth Chandran, Oliver Boyne, Vasileios Choutas, Timo Bolkart, Daoye Wang, Menglei Chai, Di Qiu, Sebastian Winberg, Gilles Rainer, Lewis Bridgeman, Delio Vicini, Jérémy Riviere, Yannick Boetzel, Alexander Koumis, Jay Busch, Cynthia Herrera, Jacob Still, Scott Ysebert, Peter Lincoln, Sergio Orts Escolano, Christoph Rhemann, Erroll Wood, Thabo Beeler, Stefanos Zafeiriou
TL;DR
Existing parametric head models omit important oral and ocular anatomy, limiting their anatomical scope. GNM unifies these structures in a high-fidelity statistical model and outperforms FLAME in reconstruction and generalization evaluations.
Problem
Existing parametric head models largely treat the head as a hollow shell, omitting oral anatomy and fine ocular structures.
Method
GNM combines high-quality real-world scans with artist-created synthetic assets to model the head, face, neck, eyes, teeth, tongue, and region-specific expressions.
Results
GNM outperforms FLAME in generalization and reconstruction accuracy across demographic groups and facial expressions.
Takeaways & Limitations
GNM provides a publicly available holistic parametric head framework with detailed internal anatomy and high-fidelity shape representation.
Takeaways & Limitations
Identity conditioning uses binary gender and four broad ethnic categories that do not fully represent human gender identities or global population diversity.
Abstract
from arXiv · showhide
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeling only outer geometry while ignoring intra-oral and ocular structures, and frequently suffer from reduced geometric quality stemming from low-fidelity input datasets. In this report we introduce a new parametric model dubbed Generative aNthropometric Model (GNM), named as a homophone of the human genome. GNM encompasses the head, face, neck, eyeballs, teeth, and tongue, and it is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy specific artist-made samples. This report details the data provenance, the model architecture including the specialized sub-models for the ocular and intra-oral structures, and shows its SotA performance on fitting target 3D face scans. To foster community innovation, the complete GNM framework is made publicly available.
1. Introduction
GNM is introduced as a holistic parametric head model addressing existing 3DMMs’ omission of oral and ocular anatomy. It unifies external and internal structures in a controllable statistical space and adds specialized variation and sampling mechanisms for high-fidelity digital humans.
- Motivation: 3D Morphable Models provide compact, expressive, and controllable low-dimensional representations of human head geometry and appearance.They support research across computer vision and graphics.
- Limitations: Existing models such as FLAME, BFM, and LSFM largely treat the head as a hollow shell, omitting teeth, tongue, and fine ocular structures.This structural omission limits downstream applications requiring anatomically complete representations.
- GNM: GNM unifies facial skin, eyes, teeth, and tongue within a single statistical space as a holistic parametric head framework.The model is designed to address physical-plausibility demands in AI, AR, and telepresence.
- Model innovations: GNM embeds teeth variation in the global identity shape space, adds tongue and localized facial-expression blendshapes, and models identity-dependent pupil, sclera, and cornea variation.It is built on high-resolution mesh topology and combines high-fidelity 3D facial data with artist-made samples.
- Accessibility and control: GNM is publicly available for academic and commercial use and includes a dual-CVAE Semantic Sampler mapping demographic and expression attributes onto a smooth parametric manifold.The introduction also presents a fitting pipeline for the model.
2. Related Models
Related head models progress from PCA-based linear 3D morphable models toward neural and photorealistic representations, while specialized ocular and intraoral models remain fragmented. GNM bridges these approaches with a controllable linear model covering the head and dedicated teeth, tongue, and eye components.
- Linear 3D morphable models: PCA-based 3D morphable models established linear representations of facial shape and texture, followed by BFM and FaceWarehouse for identity and expression variation.FLAME later introduced a more anatomically flexible model.
- Detailed capture models: DECA and EMOCA capture high-frequency details and emotional nuances, but still largely rely on the underlying FLAME topology.DECA uses regressed displacement maps for fine-scale wrinkles.
- Neural and rendering models: Neural parametric models replace explicit linear bases with implicit representations, while neural rendering frameworks target photorealistic results but face topology, identity, compute, or geometry-prior limitations.These approaches include NPHM, imHead, Shape Transformer, AIM, StyleMorpheus, and Gaussian Head Avatar.
- Region-specific anatomical models: Specialized models represent ocular, tongue, and ear anatomy with high regional fidelity, but intraoral assets lack a unified multi-part registration manifold and coordinated lips-teeth-tongue boundaries.Tongue models include linear blendshapes and neural paradigms for extreme intraoral expressions.
- GNM positioning: GNM combines linear-model simplicity and controllability with holistic anatomical scope, using a large-scale 3D facial database and dedicated teeth, tongue, and eye components.The design targets compatibility with standard graphics workflows and high-fidelity perioral and ocular detail.
3. GNM · 3.1. Data Foundation & Acquisition
GNM combines a statistical head model with articulated neck and eyeball structures, built on high-quality real-world facial scans and artist-created samples. Its data foundation uses synchronized multi-view capture to produce high-fidelity geometry across diverse subjects and expressions.
- 3. GNM: GNM comprises linear identity and expression bases, a skeletal rig for neck and eyeball articulation, joint-location identity bases, and standard linear blend skinning.The model conceptually follows the standard 3D morphable model definition while adding specialized components for articulation and posing.
- 3. GNM: The model is built from a large high-quality real-world dataset of registered expressive human faces combined with artist-created samples.
- 3.1. Data Foundation & Acquisition: The empirical foundation is a large-scale, high-resolution 3D facial database curated to maximize morphological and demographic diversity.
- 3.1. Data Foundation & Acquisition: 22 high-resolution 6144 x 4096 ZCam E2 S6G cameras and 14 controllable lights provide synchronized capture with uniform, diffused illumination.The lighting setup is designed to maximize data quality and subject comfort.
- 3.1. Data Foundation & Acquisition: The synchronized multi-view streams span roughly 150 degrees horizontally and 60 degrees vertically for high-fidelity facial reconstruction using multi-view depth refinement.The protocol captures a neutral relaxed expression and a set of static facial expressions.
- 3.1. Data Foundation & Acquisition: Over ∼5 000 individuals perform static facial expressions, producing ∼150 000 samples across flexing, 10 standard viseme categories, lips motion, and global emotions.The subjects cover diverse demographic backgrounds, supporting construction of a highly generalizable 3D morphable model.
3.2. Head Registration
GNM head registration iteratively alternates face registration and statistical model building, then refines scan fits with non-rigid offsets, image-based supervision, and geometric regularization. A cross-domain regression step reconstructs anatomically plausible cranial geometry while preserving internal structures under occlusion.
- Registration strategy: GNM uses iterative coregistration that alternates face registration with statistical model building, beginning from a curated 3D head-mesh model and continually refining registrations.The process follows FLAME and registers a large multi-view dataset.
- Scan fitting: The fitting pipeline first aligns GNM to each scan, then jointly optimizes per-vertex offsets, RGB texture, and GNM parameters using rendered image losses and geometric priors.It uses unshaded inverse rendering of the template mesh and texture in Mitsuba with edge sampling for visibility gradients.
- Scan fitting: Image supervision combines RGB, dense face landmarks, reconstructed normals, and semantic segmentation, with L1 losses for rendered modalities and SSIM for RGB.The renderer produces corresponding normal and semantic Arbitrary Output Variables.
- Geometric regularization: Geometry is stabilized with a gradient descent preconditioner, L2 regularization on vertex offsets, graph-Laplacian smoothing penalties, edge-deviation penalties, and a differentiable self-intersection loss.The self-intersection constraint targets high-curvature regions including the ears, tongue, and lips.
- Cranial reconstruction: A cross-domain latent-space regression model built from 200 artist-sculpted head meshes reconstructs an anatomically plausible cranium despite hair occlusion and limited direct measurement.Projecting registered faces onto the auxiliary model retains the captured facial information while supplying plausible cranial shape.
- Internal structures: The final registration meshes reconstruct external skin together with eyeballs, teeth, and tongue under severe occlusion, using GNM as a prior to deform internal components from sparse signals.Artist-sculpted internal parts were integrated into the initial template to bootstrap the model.
3.3. Model Formulation
GNM formulates human-head generation as a parametric function producing a 3D mesh from identity, expression, joint rotation, and translation parameters. Its bind-pose geometry is deformed with linear blend skinning using learned model bases and manually designed articulation components.
- Model formulation: GNM maps model parameters Θ and fixed model data Ψ to a human-head mesh with N_V 3D vertices.The function is defined as M(Θ; Ψ): ℝ^|Θ| → ℝ^(N_V×3).
- Model parameters: The parameter vector Θ comprises identity β, expression ϕ, angle-axis rotations θ for K = 4 joints, and global translation τ.The four articulated joints are the right eye, left eye, neck, and head.
- Model data: The fixed data Ψ includes a template mesh, template joints, identity and expression bases, joint-location identity basis, LBS weights, and a kinematic chain.These components define the model geometry, articulation, and skinning structure.
- Deformation: GNM computes bind-pose vertices from identity and expression offsets, then applies standard linear blend skinning using joint transformations and blendweights.The skinning function rotates and blends bind-pose vertices according to the articulated joint transformations.
- Parameter sources: The template mesh and geometric bases are learned from data, while skinning blendweights and the kinematic chain are manually designed by an artist.Template joint locations combine an artist-defined joint regressor with the data-driven identity basis.
3.4. Composite Linear Bases
GNM uses composite linear bases that partition identity and expression components by anatomical region, improving shape-space fidelity and enabling specialized datasets. Identity covers the head, teeth, and eyeballs, while expression covers periocular, lower-face, tongue, and pupil regions.
- Composite basis design: Identity and expression bases are split into anatomical portions to improve shape-space fidelity and support highly specialized datasets.The portions correspond to anatomical regions of the human head.
- Identity basis: The identity basis is partitioned into head, teeth, and eyeballs portions stacked along the first tensor dimension.These portions are represented as I(head), I(eye), and I(teeth) in ℝ^|𝜷(r)|×N_V×3.
- Expression basis: The expression basis comprises left-eye, right-eye, lower-face, tongue, and pupil portions.These portions correspond to the left and right periocular regions, lower face, tongue, and eyeball pupil.
- Expression basis: Each expression portion has 𝝓(r) components, where r denotes one of the five anatomical regions.The component count is defined separately for the left eye, right eye, lower face, tongue, and pupil portions.
3.5. Head Identity
GNM defines head identity as a PCA basis over Procrustes-aligned neutral face meshes, retaining dominant variance while separately preserving plausible eyeball–eyelid contact. Artist-placed template joints and a learned linear regressor provide identity-dependent joint locations.
- Identity basis: Head identity is computed by applying PCA to neutral face meshes after rigid Procrustes alignment removes head movement from identity variation.The basis is obtained from the covariance matrix of the centered, flattened mesh dataset.
- Ocular deformation: Eyeball vertices are excluded during basis construction and manually reintroduced to improve eyeball–eyelid contact fidelity.This modified dataset zeroes eyeball vertices before computing the identity basis.
- Identity basis: The retained identity basis uses orthogonal, variance-scaled vectors, with the first β(head) ≪ N_I components jointly explaining approximately 99% of dataset variance.Variance scaling unifies the effective range of identity coefficients β, while the basis vectors remain non-unit-length.
- Ocular deformation: For each identity component, eyeballs receive an optimized uniform rigid translation so they remain aligned with eyelids under positive and negative identity deformations.The translation minimizes mismatch between displaced eyeballs and corresponding eyelid vertices at fixed deformation magnitude.
- Joint locations: Template joints are manually placed, while a linear joint regressor automatically computes identity-dependent joint basis components as Q_i = ℜI_i.The regressor is optimized to map the template mesh T to the artist-placed joint locations J, with regularization.
3.6. Head Expression
GNM models facial expression through three spatially partitioned regions after two-stage mesh stabilization removes rigid misalignment. The regional basis enables localized control, mirrors periocular motion, and explains approximately 99% of dataset variance while separately handling ocular and tongue deformation.
- Regional expression basis: Expression bases are computed separately for the left periocular, right periocular, and lower-face regions using uncentered PCA of stabilized expressive-face vertex displacements.The basis is intended to capture expression-induced shape variation without rigid face motion.
- Expression components: Tongue and eyeball deformations are zeroed in the general expression basis, while jaw motion and lower-teeth movement remain fully modeled within it.Ocular and intra-oral expression bases are composed individually.
- Regional expression basis: Regional partitioning provides localized expression control, preventing semantic leakage such as a jaw movement triggering an eye wink.The three continuous vertex masks spatially blend neighboring regional components.
- Regional expression basis: Approximately 99% of dataset variance is explained by retained regional expression components, which are orthogonal within each region but not across regions.Only the first components are retained, with component contributions scaled by explained variance.
- Regional expression basis: The right periocular basis is obtained by mirroring the left-eye displacement vectors across the head’s vertical symmetry plane.This design makes left- and right-periocular expressions perfectly mirrored.
- Mesh stabilization: A two-stage stabilization process first applies automatic confidence-map alignment, then semi-automatic PCA-based stabilization across five granular facial regions.Stabilization estimates a rigid 6-DOF transform so the subject’s underlying skull aligns in space.
3.7. Internal Anatomy
GNM models internal anatomy through dedicated parametric subsystems for teeth, tongue, and eyeballs. These components combine artist-guided synthetic data with real-world scans or physiological modeling to capture anatomical variation and expressive geometry.
- Dental subsystem: GNM’s dental subsystem uses identity basis I(teeth) and parameters 𝜷(teeth) to represent diverse dental arches, tooth shapes, sizes, and alignments.The subsystem replaces static, generic mouth templates while remaining compatible with the overall head shape.
- Dental subsystem: 5 000 synthetic dental shapes were procedurally generated from an artist-rigged generic template, covering upper and lower teeth with associated gums.The resulting dataset XT is used to compute I(teeth) through PCA.
- Tongue model: GNM represents the tongue with E(tongue), a PCA-based expression subspace designed to increase inner-cavity expressiveness.The tongue dataset combines artist-made rig fitting with expressive face data and contains approximately 2.5K samples, with retained components explaining ∼99% of variance.
- Tongue model: The tongue mean is absorbed as the first E(tongue) component so the default coefficient produces a retracted tongue tucked inside the mouth cavity.This avoids treating the independently computed tongue mean as a global template transformation.
- Ocular model: GNM defines eyeballs with a linear parametric model based on a two-sphere geometry, bridging simplified spherical representations and person-specific ocular geometry.The model includes I(eye) for identity variation and E(pupil) for pupil expression, while corneal and limbal parameters support biologically plausible ocular features.
3.8. Initial Placement of Teeth and Tongue
GNM initializes artist-designed teeth and tongue within its identity and expression bases by approximating their rigid motion from surrounding lip, chin, and jaw deformations.
- Initialization: The initial GNM model manually injected artist-designed teeth, tongue, and eyeballs into the computed identity and expression bases.This procedure backfills structures after the bases are computed from registered data.
- Identity basis: Identity changes backfill teeth and tongue using a rigid 3-DOF translation estimated from mean inner upper- and lower-lip displacements for each identity component.The estimated displacement is stored in the previously zeroed teeth-and-tongue vertices of each identity basis component.
- Expression basis: Expression changes backfill lower teeth and tongue using a rigid 6-DOF transform estimated by Procrustes alignment of lower-lip and chin vertices for each expression component.The recovered transform is applied to all lower-teeth and tongue vertices in the corresponding expression component.
3.9. Model Implementation Details
GNM represents anatomical regions with disjoint PCA-based identity and expression bases, retaining components that explain sufficient dataset variance. Its topology includes 17 821 vertices and 4 skeletal joints, with the template oriented in a defined global coordinate system.
- Basis construction: Identity and expression bases use disjoint anatomical regions, each represented by PCA over vertex displacements, retaining components that explain sufficient dataset variance.The final component counts for individual regions are summarized in Table 1.
- Topology: 17 821 vertices span the outer skin, teeth, tongue, and eyeballs, while the skeletal structure contains 4 joints.These counts describe the final GNM topology.
- Coordinate system: The template head mesh is placed globally with the neck upright along positive Y and the face looking along positive Z.This establishes the model’s coordinate convention.
4. GNM Functionality & Experiments
GNM is evaluated for statistical validity, generative capability, and downstream face reconstruction, with experiments covering generalization, specificity, in-the-wild reconstruction, and semantic sampling. The results indicate improved reconstruction performance over FLAME and support detailed intra-oral recovery, while the semantic sampler enables interpretable generation and attribute editing.
- Intrinsic Evaluation: GNM’s intrinsic evaluation measures generalization and specificity against FLAME, assessing representation of diverse head shapes and restriction to plausible human-face geometries.The evaluation uses 15 000 held-out high-resolution 3D scans spanning identities, expressions, and demographic categories.
- Intrinsic Evaluation: GNM shows meaningful improvements over FLAME across reconstruction metrics and categories, with error distributions further analyzed by expression, gender, age, and ethnicity.Generalization curves also examine how scan-to-mesh distance changes with the number of identity and expression components.
- Specificity: 2 000 randomly sampled identity coefficient vectors are used to evaluate whether GNM generates realistic head geometries within the human-shape manifold.Specificity is assessed through surface-to-surface distances between sampled meshes and scans in an evaluation database, with FLAME evaluated likewise.
- In-the-Wild Reconstruction: GNM’s explicit modeling of the mouth cavity, teeth, tongue, and gums enables in-the-wild reconstruction of peri-oral deformations including extreme jaw openings and tongue expressions.The method anchors optimization to detected 2D landmarks for these regions and recovers fine geometric details in close-up reconstructions.
- Semantic Sampling: The Semantic Sampler generates plausible, human-interpretable expressions and identities for precise generation and targeted attribute editing.It is introduced to address the coupled deformations often produced by conventional parameter changes and to support animation rigs and conditioning interfaces.
- Limitations: The identity sampler’s demographic categories are limited by binary gender classification and four broad ethnic groupings inherited from source literature and scan data.The paper states that these categories do not represent the full spectrum of human gender identities or all individuals.
5. Discussion and Future Work
GNM is presented as a holistic parametric framework for digital human-head representation, unifying external facial characteristics with oral and ocular anatomy. By addressing the “hollow shell” limitation of traditional 3DMMs, it provides a rich latent space for identity and expression.
- Contributions: GNM unifies external facial characteristics with the oral cavity, teeth, tongue, and ocular anatomy in one parametric framework.This holistic scope is designed to address the “hollow shell” limitation of traditional 3DMMs.
- Contributions: GNM addresses the “hollow shell” limitations inherent in traditional 3DMMs by representing internal head structures alongside external facial characteristics.The cited internal structures include the oral cavity, teeth, tongue, and ocular anatomy.
- Contributions: GNM establishes a rich latent space for both identity and expression domains, built upon an extensive database.The passage presents the database and latent-space design as foundations of the framework.