Source-linked AI summary

3D Menagerie: Modeling the 3D shape and pose of animals

Silvia Zuffi, Angjoo Kanazawa, David Jacobs, Michael J. Black

arXiv:1611.07700v2cs.CV

TL;DR

Animal 3D modeling lacks the extensive scan data available for humans because live animals are difficult to scan. The paper learns an articulated statistical model from 41 toy scans, aligns and normalizes them, and demonstrates fitting to real animals, including unseen species.

  • Problem

    Realistic articulated animal models are scarce, while human models rely on thousands of scans that are infeasible to collect from live animals.

  • Method

    The authors use a part-based GLoSS registration, pose normalization, PCA shape modeling, and iterative co-registration to learn SMAL from toy scans.

  • Results

    SMAL generates and reposes realistic quadruped shapes and fits 2D images of real animals, including species absent from training.

  • Takeaways & Limitations

    Toy scans provide a starting point for animal shape and motion capture models that generalize across quadruped families and to real-animal images.

  • Takeaways & Limitations

    The model covers a limited set of quadrupeds and does not yet handle varying numbers or radically different types of body parts.

Abstract

from arXiv · show

There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousands of 3D scans of people in specific poses, which is infeasible with live animals. Consequently, we learn our model from a small set of 3D scans of toy figurines in arbitrary poses. We employ a novel part-based shape model to compute an initial registration to the scans. We then normalize their pose, learn a statistical shape model, and refine the registrations and the model together. In this way, we accurately align animal scans from different quadruped families with very different shapes and poses. With the registration to a common template we learn a shape space representing animals including lions, cats, dogs, horses, cows and hippos. Animal shapes can be sampled from the model, posed, animated, and fit to data. We demonstrate generalization by fitting it to images of real animals including species not seen in training.

1. Introduction

The paper extends human 3D pose-and-shape modeling to animals using toy scans, a part-based registration method, and an articulated statistical model that fits 2D data.

  • Animal modeling supports applications across biology, neuroscience, ecology, farming, and entertainment, but computer vision has focused more heavily on humans.
  • Animals are harder to model because species vary greatly in shape, tails are highly deformable, and collecting live-animal scans is impractical.
  • The authors learn from 41 toy-animal scans and show that the resulting model generalizes to real animals.
  • GLoSS provides coarse template registrations across very different animals, which ARAP-constrained deformation refines toward scan surfaces.
  • Pose-normalized registrations support PCA learning, while co-registration jointly refines the registrations and shape space until convergence.
  • SMAL represents quadruped shape variation from 41 scans, generates and reposes realistic animals, and fits them to keypoints and segmentations in 2D data.

2. Related Work

Prior work spans 2D animal analysis, part-based 3D representations, image- and video-based reconstruction, and human scan models; this paper addresses multi-animal articulated shape spaces.

  • Earlier 3D animal work includes part-based models that connect shape primitives in a kinematic tree.
  • Existing animal scan datasets have limited coverage or realism, while live-animal scanning is difficult because of varied size, shape, and movement.
  • Image-based methods deform templates using keypoints, segmentations, and silhouettes, but prior models often target individual species or produce low-resolution complex shapes.
  • The paper is complementary to image-only approaches and identifies combining scans with image data as a route to richer models.
  • Video methods track parts, appearance, or surfaces, but do not model the detail of 3D scans or a multi-animal articulated shape space.
  • SMPL provides the human-modeling foundation, but it uses far more scans while animals require representing greater shape variability with less data.

3. Dataset

The dataset contains 41 toy-animal scans spanning multiple quadruped species, with scale normalized across manufacturers.

  • The authors scanned 41 toy figurines across cats, big cats, dogs, foxes, wolves, hyenas, deer, horses, zebras, cows, and hippos.
  • They estimated a scaling factor so animals from different manufacturers were comparable in size.

4. Global/Local Stitched Shape Model

GLoSS is a globally differentiable, part-based articulated model that represents local shape deformations and assembles parts through interface stitching. Its analytic deformations allow fitting novel animal shapes without prior training data.

  • GLoSS defines local shape deformations for each articulated part and assembles parts by minimizing stitching costs at their interfaces.The model is globally differentiable, enabling gradient-based fitting to data.
  • GLoSS uses analytic rather than learned shape deformations, making it more approximate but applicable to novel animal shapes without a priori training data.
  • The model requires a template mesh, part segmentation, skinning weights, and an animation sequence; the template is manually divided into N = 33 parts.
  • Each part is parameterized by location, absolute rotation, intrinsic shape variables, and pose deformation variables.Vertex coordinates are computed in a global frame from these part-specific variables.
  • Pose-dependent deformation bases are computed from animated template examples by applying PCA to each part’s local vertices.

5. Initial Registration

Initial registration first fits GLoSS to each scan, then refines mesh vertices with model-free ARAP optimization. The objective combines model, stitching, scan-distance, keypoint, and curvature terms to obtain tighter registrations.

  • Initial registration uses gradient-based GLoSS fitting followed by model-free As-Rigid-As-Possible refinement to capture fine scan details.The GLoSS stage provides a coarse alignment, while ARAP moves vertices closer to the scan.
  • The registration objective includes model terms based on Mahalanobis shape distance and L2 pose deformation penalties.
  • The stitching term sums squared distances between corresponding vertices at interfaces, favoring connected parts.
  • Keypoint matching helps align scans with extremely different animal shapes, while curvature preserves pairwise relationships between part surfaces and the template.
  • Figure 4 progresses from the initial template and scan to the GLoSS fit, part visualization, and a merged mesh with global topology.
  • Figure 5 compares coarse GLoSS registrations with tighter ARAP-refined fits, and Figure 6 shows toy scans registered in neutral pose.

6. Skinned Multi-Animal Linear Model

SMAL learns a low-dimensional animal shape space from registered toy scans, combining shape variation with articulated pose. The model captures family structure and can generate and repose animal meshes.

  • Animal shape space: Pose-normalized registrations support learning mean shapes and principal components that capture differences between animals.Normalization uses linear blend skinning before shape-space estimation.
  • Model representation: SMAL represents animal meshes as a function of shape, pose, and global translation.Shape uses PCA coefficients, pose specifies relative joint rotations, and translation shifts the root joint.
  • Model refinement: The registrations and model are refined jointly through four co-registration iterations using a 30-dimensional shape space.Each iteration alternates SMAL fitting with model-free registration regularized toward the current model.
  • Animal shape space: The final PCA space captures shape variability across animal families and separates family-specific characteristics.The first component captures scale differences, while t-SNE visualizes the first eight PCA dimensions and family means.

7. Fitting Animals to Images

The paper fits SMAL to images by optimizing pose, shape, translation, and camera parameters against manually extracted keypoints and silhouettes. Priors, bounds, multiscale optimization, and staged weighting help constrain the fit.

  • Image fitting: SMAL is fitted to manually extracted 2D keypoints and silhouettes by optimizing its shape and pose parameters.The image-fitting objective also estimates translation and focal length.
  • Objective function: The objective combines keypoint and silhouette reprojection errors with shape and pose priors.Silhouette matching uses a bidirectional distance, while shape and pose priors are Mahalanobis penalties.
  • Pose constraints: Pose fitting uses reflected training poses to create a symmetric prior and hand-defined limits because the training set is small.Global rotation is not limited by these bounds.
  • Optimization: Optimization proceeds from torso-based depth and rotation initialization through staged prior relaxation before adding the silhouette term.The staged procedure is intended to reduce trapping in local optima.

8. Experiments

Experiments test whether a toy-trained SMAL model captures real-animal shape variation. Fits generalize to unseen animal families, while the main failures arise from depth ambiguity.

  • Experimental setup: The experiments fit generic and family-specific SMAL models to annotated images of real animals.The fitting data contains 19 semantic keypoints plus an additional tail-tip point.
  • Generalization: SMAL generalizes to boars, donkeys, sheep, and pigs, which were not present in training.The fits capture broad quadruped shape variation but cannot exactly represent characteristic properties such as a pig snout.
  • Failure cases: The main fitting failures result from inherent depth ambiguity in global rotation and pose.These failures are illustrated in Figure 10.

9. Conclusions

The work shows that toy figurines can seed an animal model that generalizes to real images and unseen animal types. Extending beyond the studied quadrupeds requires richer evidence and mechanisms for variable animal parts.

  • Conclusions: A model learned from toy figurines generalizes to images of real animals and animal types absent from training.The authors present this as a procedure for building richer models from more animals and scans.
  • Future directions: The authors identify image and video evidence as needed for a much richer animal model.Image fitting is described as a starting point for learning richer deformations from 2D evidence.
  • Limitations: The demonstrated model is limited to a restricted set of four-legged mammals with shared part counts.Broader animal classes introduce varying numbers and substantially different types of parts, such as horns, tusks, trunks, and elephant ears.
Loading 1611.07700v2…