Source-linked AI summary
HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling
Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin, Tao Yu, Wenjia Wang, Xiangyu Fan, Yang Gao, Yifan Yu, Liang Pan, Fangzhou Hong, Mingyuan Zhang, Chen Change Loy, Lei Yang, Ziwei Liu
TL;DR
HuMMan addresses the need for versatile datasets for 4D human sensing and modeling. It contributes a large-scale multimodal dataset with mobile-device data, a comprehensive action set, and benchmarks across human-vision tasks, while experiments identify fine-grained recognition and cross-device transfer as continuing challenges.
Problem
Existing human-sensing research needs more versatile datasets spanning modalities, mobile devices, comprehensive actions, and multiple tasks.
Method
HuMMan constructs a large-scale multimodal 4D dataset with synchronized annotations, mobile-device sensing, 500 anatomically designed movements, and benchmarks for multiple tasks.
Results
Experiments support HuMMan’s usefulness across multiple fields and expose challenges in fine-grained action recognition, dynamic mesh reconstruction, point-cloud recovery, and cross-device transfer.
Takeaways & Limitations
HuMMan provides a benchmark for studying multimodal 4D human sensing and modeling across actions, representations, reconstruction, and devices.
Takeaways & Limitations
Knowledge transfer across devices remains challenging, especially for point-cloud methods because mobile devices produce sparser point clouds from lower-resolution depth maps.
Abstract
from arXiv · showhide
4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this work, we contribute HuMMan, a large-scale multi-modal 4D human dataset with 1000 human subjects, 400k sequences and 60M frames. HuMMan has several appealing properties: 1) multi-modal data and annotations including color images, point clouds, keypoints, SMPL parameters, and textured meshes; 2) popular mobile device is included in the sensor suite; 3) a set of 500 actions, designed to cover fundamental movements; 4) multiple tasks such as action recognition, pose estimation, parametric human recovery, and textured mesh reconstruction are supported and evaluated. Extensive experiments on HuMMan voice the need for further study on challenges such as fine-grained action recognition, dynamic human mesh reconstruction, point cloud-based parametric human recovery, and cross-device domain gaps.
1 Introduction
HuMMan is a large-scale 4D human dataset designed for versatile sensing and modeling. It combines multiple synchronized modalities, mobile-device data, a 500-action set, and benchmarks across several human-vision tasks.
- HuMMan contains 1,000 subjects, 400k sequences, and 60M frames for 4D human sensing and modeling.
- Multiple Modalities: It provides synchronized color, point-cloud, keypoint, SMPL-parameter, and textured-mesh data, with spatial alignment for 3D annotations.
- Mobile Device: A mobile phone with built-in LiDAR is included to address insufficient mobile-device data for emerging applications.
- Action Set: HuMMan designs 500 movements by categorizing major muscle groups for a more complete and fundamental representation of human actions.
- Multiple Tasks: The dataset supports benchmarks for action recognition, pose estimation, parametric human recovery, and textured mesh reconstruction.
2 Related Works
Prior datasets support human action, pose, parametric-recovery, and textured-mesh research, but limitations remain in scale, 3D annotations, action coverage, and sequential textured data. HuMMan is positioned to address these gaps with a broader action set and multimodal sequential data.
- Action Recognition: RGB action-recognition datasets often lack 3D annotations, while earlier RGB-D datasets were small and NTU RGB-D emphasizes upper-body actions.
- HuMMan differs by developing a larger, more complete action set while providing sequential multimodal data for related sensing and modeling tasks.
- 2D and 3D Keypoint Detection: Pose-estimation datasets span 2D and 3D keypoints and single- or multi-view settings, using images or videos with differing annotation formats.
- 3D Parametric Human Recovery: Parametric human recovery methods use keypoints, images, videos, or point clouds to estimate models such as SMPL, SMPL-X, STAR, and GHUM.
- Textured Mesh Reconstruction: Textured human-mesh reconstruction uses methods including multi-view stereo, volumetric fusion, neural surfaces, texture mapping, and neural rendering.
3 Hardware Setup
HuMMan uses a calibrated, synchronized octagonal sensor framework for multiview RGB-D capture, mobile-device sensing, and high-resolution static scans. The setup addresses coverage, calibration, and synchronization requirements for 4D data collection.
- The collection system uses a 1.7 m-high octagonal framework with a 3.4 m side length to accommodate calibrated, synchronized sensors.
- RGB-D Sensors: Ten Azure Kinects capture multiview RGB-D sequences with 1920×1080 color and 640×576 depth resolution.
- RGB-D Sensors: The Kinects are spaced for wide coverage so expressive poses remain visible to at least two sensors.
- Calibration: Image-based calibration uses a light-absorbent chessboard modification to mitigate infrared overexposure and calibrate Kinects and iPhones.
- Synchronization: Synchronization cables distribute a unified clock among Kinects, with images rejected when synchronization error reaches 33 ms or above.
4 Toolchain
HuMMan’s toolchain converts synchronized multi-view RGB-D data into several aligned human annotations, including keypoints, SMPL parameters, and textured meshes. It combines automated processing with geometric constraints, scan-based registration, denoising, and depth-aware texture reconstruction.
- Annotation pipeline: The automatic toolchain produces keypoint annotations, SMPL parameters, and dynamic textured meshes from captured sequences.Human inspection rejects low-quality data with erroneous annotations.
- Keypoint annotation: Keypoint annotation uses virtual cameras on static scans and multi-view RGB-D color images across two stages.Stage I renders multi-view images from minimally clothed scans; stage II uses captured multi-view RGB-D imagery.
- Keypoint annotation: 3D keypoints are triangulated with calibrated camera parameters and constrained by temporal smoothness and connected-bone consistency.The resulting 2D and 3D keypoints support separate pose-estimation tasks.
- Parametric registration: SMPL registration estimates pose, shape, and translation parameters, using 3D keypoint energy, scan-surface alignment, shape consistency, and joint-angle constraints.Stage I estimates shape from static high-resolution scans, while stage II estimates pose for dynamic sequences using the stage-I shape.
- Textured mesh reconstruction: Point-cloud denoising removes boundary noise, while depth-aware texture reconstruction prevents texture misprojection caused by geometry–subject misalignment.Overlapping camera views allow more aggressive removal of noisy boundary pixels.
5 Action Set
HuMMan designs 500 actions around completeness and unambiguity. Its hierarchy covers upper extremity, lower-limb, and whole-body movements, while muscle-based definitions make action classes more specific.
- Design principles: HuMMan’s 500-action set is designed around completeness and unambiguity rather than only empirically selected daily activities.The design takes an anatomical perspective based on body movements and driving muscles.
- Completeness: The hierarchy divides actions into upper-extremity, lower-limb, and whole-body movements to balance coverage across body parts.Whole-body movements require multiple body parts to collaborate, including lying down and sprawling.
- Unambiguity: Muscle-based action definitions aim to make classes clear and unambiguous.The paper contrasts this with general motion descriptions used in earlier datasets.
- Coverage: The action design seeks broader coverage than datasets whose actions mainly emphasize upper-body movements.HuMMan cross-checks action definitions against existing datasets to support wide coverage.
6 Subjects
HuMMan includes 1000 subjects spanning diverse genders, ages, body shapes, and ethnicities. Subjects wear their personal daily clothes to provide naturally varied appearances.
- Subject diversity: HuMMan contains 1000 subjects with broad coverage of genders, ages, body shapes, and ethnicities.Body-shape diversity includes variation in height and weight.
- Subject appearance: Subjects wear personal daily clothes, creating a collection of natural appearances.The dataset also provides high-resolution scans of the subjects.
7 Experiments
Experiments evaluate HuMMan across action recognition, 3D keypoint detection, parametric human recovery, textured mesh reconstruction, and cross-device transfer. The results show both strong support for diverse sensing tasks and persistent challenges in fine-grained actions, genuine point clouds, mesh reconstruction, and device transfer.
- Action Recognition: 74.1% Top-1 accuracy is achieved by 2s-AGCN on HuMMan, compared with 88.9% on NTU RGB+D 60 and 82.9% on NTU RGB+D 120.The lower accuracy indicates that HuMMan is a more challenging action-recognition benchmark than these NTU RGB+D settings.
- Action Recognition: A roughly 30% Top-1–Top-5 accuracy gap reflects fine-grained intra-action variants such as quadruped, kneeling, and leg push-ups.HuMMan’s whole-body action coverage requires models to attend to subtle differences between similar movements.
- 3D Keypoint Detection: HuMMan-trained 3D keypoint methods generalize better than Human3.6M-trained methods, while in-domain values are slightly higher than Human3.6M’s FCN MPJPE of 53.4 mm.The paper attributes this pattern to HuMMan’s diverse subjects and actions.
- 3D Parametric Human Recovery: HMR achieves low MPJPE and PA-MPJPE, whereas VoteHMR performs poorly on genuine RGB-D point clouds from commercial sensors.The paper argues that existing point-cloud methods rely heavily on synthetic data, which differs from HuMMan’s real sensor point clouds.
- Textured Mesh Reconstruction: HuMMan supports reconstruction methods using RGB, single-view RGB-D video, and multi-view depth point clouds through its multimodal signals.The evaluated method families include PIFu, PIFuHD, Function4D, 3D Self-Portrait, and CON.
- Mobile Device: Cross-device transfer remains challenging: image-based methods show a considerable domain gap, while point-cloud methods face a much more significant gap from sparser mobile-device clouds.The mobile device’s lower depth-map resolution is identified as the source of its sparser point clouds.
8 Discussion
HuMMan is presented as a large-scale 4D human dataset combining multimodal sensing, mobile-device capture, a comprehensive action set, and benchmarks for multiple vision tasks. The authors identify several directions for future research in human sensing and modeling.
- 8 Discussion: HuMMan combines multimodal data and annotations, mobile-device capture, atomic-motion actions, and standard benchmarks for multiple vision tasks.The dataset is intended to support broader sensing and modeling studies.
- 8 Discussion: The authors highlight fine-grained action recognition, point-cloud parametric human estimation, dynamic mesh reconstruction, cross-device transfer, and potentially multi-task joint training as future directions.These directions are drawn from the experiments and the dataset’s supported tasks.
- 8 Discussion: Additional dataset-collection, hardware, toolchain, action-set, subject, experiment, and dataset-comparison details are provided in the supplementary material.
B Additional Details of Data Collection
HuMMan combines synchronized multi-view RGB-D capture, mobile sensing, calibration, annotation, and scan-based quality evaluation. Its collection spans 500 actions and supports temporally synchronized, spatially aligned multimodal data.
- Data collection: Each subject receives natural-clothing and tight-fitting high-resolution scans, then performs 40–60 randomly sampled actions from the 500-action set.Each action contains ten Kinect RGB-D sequences and one iPhone RGB-D sequence, lasting 5–15 seconds at 30 FPS.
- Sensor suite: HuMMan uses ten Azure Kinects and an iPhone 12 Pro Max to capture synchronized multi-view RGB-D sequences.The Kinect arrangement provides broad body coverage, while Fig. 9 shows synchronized RGB frames and device IDs.
- Synchronization: The synchronization method estimates the Kinect-to-iPhone ARKit clock offset by combining workstation-to-Kinect, iPhone-to-workstation, and ARKit-to-iPhone offsets.The workstation communicates with the iPhone app through TCP and uses a round-trip message to measure clock relationships.
- Point clouds: Raw Kinect and iPhone point clouds differ substantially: iPhone clouds are sparser and empirically noisier, especially near object boundaries.The Fig. 10 visualization uses the same 10× downsampling factor for both raw clouds.
- Annotations: The keypoint annotation pipeline filters detections, triangulates 3D points, reprojections them, iteratively selects cameras, and returns failure if convergence is not reached.Its inputs include detected 2D keypoints, camera parameters, confidence thresholds, and camera-selection settings.
- Body-shape and pose annotation: The full-body angle prior bounds three-degree-of-freedom movement using per-axis maximum ranges, with empirically inspected values and Euler angles used only for the loss.Joint rotations remain represented in axis-angle form, and the coordinate design aims to reduce gimbal lock.
E Additional Details of Action Set
HuMMan constructs its action set hierarchically from body parts and muscle groups, targeting systematic coverage of fundamental movements. Its motion statistics indicate greater joint-angle diversity than two established datasets.
- Design process: HuMMan organizes actions hierarchically from the body to major body parts and then to anatomically defined muscle groups.This structure is intended to make action coverage complete and unambiguous.
- Motion diversity: HuMMan’s mean joint-angle standard deviation is 0.269, compared with 0.208 for AMASS and 0.159 for 3DPW.The paper uses higher mean standard deviation as an indicator of greater motion diversity.
- Subjects and ethics: The dataset contains 1000 subjects, with gender, age, height, and weight reported as subject-diversity statistics.Participants were recruited voluntarily and signed agreements permitting public research use of their data.
G.1 Splits and Protocols
HuMMan defines subject-, action-, and view-based evaluation protocols while sampling 10% of its large corpus for computational feasibility. Supplementary benchmarks show substantial difficulty under unseen-action and cross-view shifts.
- Protocols: Only 10% of HuMMan is sampled for training and testing because the full dataset contains 1000 subjects, 500 actions, 400k sequences, and 60M frames.Protocol 1 splits subjects into mutually exclusive 70% training and 30% test sets and is used throughout the main paper.
- Cross-action evaluation: Cross-action 3D keypoint precision degrades significantly when models train on fewer actions and test on unseen actions.The whole-body category shows especially pronounced cross-evaluation degradation under action-distribution misalignment.
- Cross-view evaluation: Cross-view 3D keypoint errors increase as the test view deviates from the training view, demonstrating a considerable view-domain gap.The experiment uses a model trained only on View 0 and evaluated on all views.
- Parametric recovery: Testing 3D parametric human recovery on unseen poses is challenging, and whole-body actions are more distributionally distant from limb-specific actions.These findings come from Protocol 2 evaluation of the HMR baseline.
- Mesh reconstruction: HuMMan supports textured-mesh reconstruction comparisons using Function4D with four selected views.The broader evaluation also compares reconstruction methods and datasets across multiple vision tasks.
H A More Complete Dataset Comparison
HuMMan is compared with published datasets across action recognition, keypoint detection, parametric human recovery, and mesh reconstruction. The comparison tracks scale, modalities, annotations, mobile sensing, and video support.
- Task coverage: The dataset comparison covers four task families: action recognition, 2D and 3D keypoint detection, 3D parametric human recovery, and mesh reconstruction.The comparison is presented as a more complete evaluation against similar published datasets.
- Comparison dimensions: The comparison table records subjects, actions, sequences, video, mobile sensing, depth or point clouds, action labels, keypoints, parametric models, and texture.It distinguishes genuine depth-sensor point clouds from other depth representations and marks unavailable entries as not applicable or not reported.