Source-linked AI summary
Detailed, accurate, human shape estimation from clothed 3D scan sequences
Chao Zhang, Sergi Pujades, Michael Black, Gerard Pons-Moll
TL;DR
The paper estimates personalized minimally-clothed body shape and pose from clothed 3D scan sequences, addressing the challenge that clothing obscures the target shape and prior model-based methods lose identity details. It fuses temporal registrations into a single shape estimate and introduces BUFF for quantitative evaluation; the method improves pose and shape estimation over prior work.
Problem
Estimating minimally-clothed body shape from clothed 3D scans is difficult because clothing occludes the body, while statistical body models can produce overly smooth shapes lacking identity details.
Method
The method registers scans into a common unposed space, fuses the registrations into a fusion scan, and estimates a single detailed body shape before recovering sequence pose and shape.
Results
The method produces numerically and visually accurate body-shape and pose estimates fitting clothed scans and outperforms previous state-of-the-art methods qualitatively and quantitatively.
Takeaways & Limitations
BUFF provides high-resolution clothed scan sequences with ground-truth minimally-clothed shapes, enabling quantitative evaluation of body-shape estimation.
Takeaways & Limitations
Visual hulls can overestimate or underestimate true shape, limiting their suitability for quantitative evaluation.
Abstract
from arXiv · showhide
We address the problem of estimating human pose and body shape from 3D scans over time. Reliable estimation of 3D body shape is necessary for many applications including virtual try-on, health monitoring, and avatar creation for virtual reality. Scanning bodies in minimal clothing, however, presents a practical barrier to these applications. We address this problem by estimating body shape under clothing from a sequence of 3D scans. Previous methods that have exploited body models produce smooth shapes lacking personalized details. We contribute a new approach to recover a personalized shape of the person. The estimated shape deviates from a parametric model to fit the 3D scans. We demonstrate the method using high quality 4D data as well as sequences of visual hulls extracted from multi-view images. We also make available BUFF, a new 4D dataset that enables quantitative evaluation (http://buff.is.tue.mpg.de). Our method outperforms the state of the art in both pose estimation and shape estimation, qualitatively and quantitatively.
1. Introduction
The paper estimates personalized minimally-clothed body shape and pose from clothed 3D scan sequences by combining free-form shape optimization with temporal fusion. It introduces BUFF and reports accurate results that outperform prior methods.
- Clothing occludes minimally-clothed shape, complicating body estimation despite its relevance to virtual try-on, biometrics, fitness, and cloth simulation.
- Statistical body models constrain estimates to smooth model spaces, losing identity details and making local clothing-surface constraints difficult to satisfy.
- A single-frame objective keeps cloth outside the body, fits visible skin, and robustly attracts the body surface to nearby cloth vertices.
- The method directly optimizes 6,890 template vertices while regularizing them toward SMPL, allowing local details beyond the parametric model.
- Temporal processing registers scans into a common unposed space, fuses them into a fusion scan, estimates a fusion shape, then uses it to recover pose and time-varying details.
- BUFF contains 11,054 high-resolution clothed scans from six subjects with ground-truth naked shapes, and evaluations report improvements over prior state-of-the-art methods.
2. Related Work
Prior work uses statistical models, image or depth cues, and model-free tracking, but commonly ignores clothing, requires manual initialization, or lacks personalized shape detail. This method combines body-model constraints with free-form subject-specific deformation.
- Body Models: Modern body models separate pose and shape deformations and are learned from thousands of real-human scans.
- Pose and Shape Estimation: Depth-based methods estimate pose and shape from silhouettes, depth, color, or correspondences, but several focus on minimal clothing or do not explicitly model clothing.
- Pose and Shape Estimation: Image and multi-view approaches often ignore clothing or treat it as noise, while joint-based fitting simplifies shape because bone lengths cannot determine the full body.
- Shape Under Clothing: Shape-under-clothing methods use layered clothing models, body-model fitting, or temporal averaging, but commonly require manual pose initialization.
- Shape Under Clothing: Other approaches model clothing statistics in 2D or through simulation, but show limited body regions or clothing variety.
- Body Models: Model-based prior methods optimize only model parameters, restricting results to smooth model spaces and omitting subject-specific details.
- Shape Under Clothing: The proposed method combines the strong constraints of a body model with the deformation freedom of model-free approaches.
3. Body Model
The paper uses SMPL, a learned rigged body model that represents identity-dependent shape and pose-dependent deformation over a 6,890-vertex template and 24-joint kinematic skeleton.
- SMPL uses a learned rigged template with N = 6890 vertices whose positions depend on identity shape parameters and skeletal pose.
- The human skeleton is represented as a kinematic chain of n = 24 rigidly linked ball joints, each parameterized by three rotational degrees of freedom.
- SMPL applies additive shape- and pose-dependent offsets to the template and predicts joint locations from the deformed template.
- A linear blend skinning function transforms rest-pose vertices using joint locations, pose, and blend weights into posed vertices.
4. Method
The method estimates personalized naked shape and pose from clothed scan sequences by combining scan-fitting objectives with temporal fusion and model-based regularization.
- The outputs are a personalized static template shape, per-frame poses, and per-frame detailed template shapes that may vary slightly over time.Allowing temporal variation captures details that the pose deformation model cannot represent.
- Single-Frame Objective: The single-frame objective enforces cloth outside the body, fits visible skin, and robustly ignores distant cloth points.Its energy combines skin, cloth, model-coupling, and prior terms.
- Single-Frame Objective: Skin residuals are smoothly weighted by geodesic distance to cloth boundaries, making the objective more robust to inaccurate segmentation.A logistic mapping converts geodesic distances into weights that decrease near skin-cloth boundaries.
- Single-Frame Objective: Cloth penetration is penalized, while a robust Geman-McClure fit term prevents wide garments from forcing the estimated body excessively thin.Points far from the body receive a small nearly constant penalty.
- Single-Frame Objective: The template’s 6,890 vertices are optimized directly and regularized toward SMPL, preserving local details while maintaining anthropometric constraints.Joint optimization of the template and model coefficients couples the detailed estimate with its statistical representation.
- Fusion Shape Estimation: Because single-frame estimates vary with pose, aligned scans are brought into a canonical pose, united into a fusion scan, and used to estimate one fusion shape.Different poses constrain different cloth regions; the fusion scan therefore carves the volume in which the naked shape should lie.
5. Datasets
The paper evaluates its method on the existing INRIA dataset and introduces BUFF, a higher-resolution dataset designed to preserve shape details for quantitative evaluation.
- Existing Dataset: INRIA contains visual-hull mesh sequences from six subjects performing three motions in three clothing styles.It was captured with 68 color cameras at 30fps and includes sparse motion-capture data, but no scan texture.
- Existing Dataset: INRIA’s S-SCAPE-based ground-truth shape does not capture individual human-shape details and biases quantitative evaluation toward the model space.Visual hulls offer an alternative qualitative reference but can over- or under-estimate the true shape.
- BUFF Dataset: BUFF contains high-resolution clothed scan sequences of three males and three females in different clothing styles, with ground-truth naked shape for each subject.The dataset is publicly available for research and supports qualitative and quantitative evaluation.
- BUFF Dataset: BUFF was captured with a multi-camera active-stereo system at 60 frames per second, producing full-body meshes with approximately 150K vertices on average.The system uses 22 stereo-camera pairs, 22 color cameras, 34 speckle projectors, and LED panels.
- BUFF Dataset: BUFF includes six subjects wearing two clothing styles, with sequences lasting 4–9 seconds and totaling 13,632 3D scans.Subjects perform varied activities, and the dataset also includes texture data.
- Ground-Truth Shapes: Ground-truth minimally clothed shapes are computed from an “A-T-U-Squat” sequence by fitting all frames and averaging the resulting template meshes.More than half of scan points lie within 1.5mm of the estimated shape, and 80% lie within 3mm.
6. Experiments
Experiments evaluate pose and shape estimation on previous datasets and BUFF, using qualitative comparisons and quantitative registration error. The method achieves state-of-the-art pose results and systematically outperforms prior work on BUFF shape estimation.
- Evaluation on Previous Datasets: Pose estimation on INRIA achieves state-of-the-art accuracy across average landmark errors and error-versus-distance evaluations.The evaluation uses 10 frames sampled from the first 50 frames of each sequence to obtain 10 correspondence sets.
- Evaluation on Previous Datasets: Qualitative INRIA and Dancer comparisons show plausible minimally-clothed shape estimates that visually outperform previous state-of-the-art methods.The Dancer comparison includes Wuhrer et al. [44] and Yang et al..
- Evaluation on BUFF: BUFF numerical results systematically outperform the best state-of-the-art method for the fusion and detailed meshes.Table 1 reports root mean squared point-to-surface distance in millimeters between the posed ground-truth mesh and each method result.
- Evaluation on BUFF: BUFF qualitative results recover scan pose effectively, especially at elbows and shoulders, while fusion and detailed shapes accurately capture body shape.The detailed shape captures missing details and is visually closer to ground truth, although both shape variants are quantitatively very similar.
- Evaluation on BUFF: Using all scan vertices as cloth increases mean error from ≈2.5mm for the detailed method to ≈3mm, while retaining accurate shapes.The all-cloth setting evaluates robustness when skin/cloth segmentation is unavailable; the resulting shapes are less detailed.
- Computation time and parameters: Computation takes ∼10 seconds per frame for the single-frame objective, ∼200 seconds for fusion, and ∼40 seconds per frame for detail refinement.Sequence computations run in parallel on a 3GHz 8-Core Intel Xeon E5.
7. Conclusion
The paper estimates detailed body shape under clothing from scan sequences by fusing clothed registrations, and introduces BUFF for quantitative evaluation. Results improve over the state of the art, but SMPL limits female breast-shape estimation.
- Conclusion: The method fuses clothed registrations into a single frame to estimate detailed body shape under clothing from 3D scan sequences.The fused approach produces accurate shape estimates while using sequence information.
- Conclusion: BUFF provides high-resolution clothed scan sequences and ground-truth minimally-clothed shapes for quantitative body-shape evaluation.The paper describes BUFF as the first dataset of high-quality 4D scans of clothed people.
- Conclusion: Results on BUFF reveal a clear improvement over the state of the art.
- Conclusion: Underestimation of female breast shape remains a limitation attributed to SMPL's omission of soft-tissue deformations.The authors identify soft-tissue deformation models as future work for improving accuracy.