Source-linked AI summary

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han

arXiv:2608.16222v1cs.ROcs.AI

TL;DR

Existing humanoid datasets trade behavioral coverage for physical fidelity, limiting reference motion for learning whole-body skills and interactions. HiPHI addresses this with a FrameNet-guided, high-precision motion-and-object-interaction dataset and benchmark, demonstrating broader coverage, higher motion quality, stronger interaction consistency, and improved humanoid tracking.

  • Problem

    Existing humanoid-learning data trade broad behavioral coverage for precisely measured physical states and interaction grounding, or provide high fidelity across only narrow behaviors.

  • Method

    HiPHI uses FrameNet-guided capture scripts expanded across motion and interaction dimensions to construct a 617.5-hour high-precision human-motion and object-interaction benchmark.

  • Results

    HiPHI demonstrates broader motion-space coverage, higher motion quality, stronger human-object geometric consistency, and improved downstream humanoid tracking performance.

  • Takeaways & Limitations

    HiPHI provides executable reference data that transfers to real humanoid hardware across diverse whole-body and object-interaction behaviors.

  • Takeaways & Limitations

    HiPHI captures only single-person studio motion and focuses on kinematics rather than directly measuring contact forces or tactile signals.

Abstract

from arXiv · show

Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets remain fundamentally limited: internet-scale video data lack precise physical states and interaction grounding, while laboratory motion datasets provide high fidelity but only narrow behavioral coverage. This mismatch creates a critical bottleneck for scalable humanoid policy learning. We present HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold. HiPHI is theoretically guided by FrameNet, a linguistic framework organizing human primitives. Created using an optical motion capture pipeline, HiPHI provides sub-millimeter spatial marker tracking accuracy for full-body human motion and mesh-level object trajectories. We further introduce a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications. Our analyses demonstrate that HiPHI significantly expands motion coverage compared to existing motion datasets while maintaining high-fidelity interaction quality, and establishes a scalable data foundation for training, evaluating, and generalizing humanoid policies in real-world embodied tasks, where similar extensions are also applicable to motion prior models in computer graphics.

1 Introduction

HiPHI addresses the need for broad, physically faithful reference motion by using FrameNet-guided motion-space design to build a large-scale, high-precision human motion and object-interaction benchmark. The dataset and evaluation protocol target motion coverage, data quality, interaction consistency, and humanoid policy learning.

  • Motivation: Existing humanoid data sources lack high-precision behavioral coverage, while robot demonstrations and teleoperation are costly, small-scale, and embodiment-specific.Humanoid policies require balance, posture transitions, limb coordination, contact, and load-dependent strategies.
  • Method: HiPHI uses FrameNet frames and lexical units as semantic seeds for performer-facing capture scripts, framing data collection as motion-space design.The pipeline selects embodied-intelligence-relevant frames and lexical units, then expands each seed across multiple dimensions.
  • Dataset: 617.5-hour HiPHI was captured from 132 performers with sub-millimeter optical motion-capture accuracy, combining 371.8 hours of whole-body motion and 245.7 hours of object interaction.The dataset spans 214 Frame–LU motion units across 22 frames and includes synchronized object trajectories and meshes, indexing, and natural-language descriptions.
  • Evaluation: HiPHI is evaluated for motion-space coverage, data quality, interaction consistency, and humanoid policy learning through kinematic, physical, geometric, tracking, and deployment analyses.The analyses assess label-free kinematic embeddings, smoothness, ground contact, interaction geometry, humanoid motion tracking, and real-world deployment.

2 Related Work

Existing embodied datasets trade off embodiment, scale, behavioral diversity, physical-state precision, or interaction fidelity. HiPHI addresses these limitations with 245.7 hours of contact-rich human-object interaction data emphasizing physical accuracy.

  • Robot Teleoperation and Egocentric Vision Data: Real-robot datasets avoid embodiment gaps but remain constrained by collection cost, scale, behavioral diversity, and dependence on specific hardware, sensors, and viewpoints.These sources include teleoperated robot data, manipulation benchmarks, and low-cost interfaces such as UMI.
  • Robot Teleoperation and Egocentric Vision Data: Egocentric visual data capture long-tail real-world behaviors but make precise human-motion and physical-interaction states difficult to recover.Their primary limitation is that they are typically restricted to visual modalities.
  • Human Motion Capture Data: Motion-capture datasets accurately record human motion but were generally designed for synthesis, annotation, generation, or animation rather than systematically for embodied intelligence.Representative datasets include AMASS, LAFAN1, BABEL/KIT/HumanML3D, Motion-X/Motion-X++, and BONES-SEED.
  • Human-Object Interaction: 245.7 hours of HOI data distinguish HiPHI from typically small datasets by emphasizing physical accuracy in contact-rich interactions.Existing HOI datasets are often only a few hours long, with limited contact-level accuracy and interaction diversity.

3 FrameNet-Guided Motion-Space Construction

HiPHI constructs humanoid-relevant motion coverage with FrameNet rather than ad hoc semantic scripts, using Frame–LU seeds and systematic expansion across physical and interaction factors. This approach targets control-centric, reachable human motion while preserving the object states needed for physically complete interactions.

  • Motivation: Humanoid-relevant data should emphasize control-centric patterns such as locomotion, posture transitions, limb coordination, dynamic balance, and responses to external constraints.The paper contrasts these patterns with behaviors such as sleeping, eating, or watching television, which robots are not expected to over-learn.
  • Limitations of Manual Scripting: Manually designed scripts provide natural, interpretable actions but offer no principled way to identify covered or missing regions of the motion space.A single action such as walking can yield diverse whole-body motions through changes in direction, speed, stride length, turning, or posture.
  • FrameNet Foundation: FrameNet supplies structured motion meanings through frames and lexical units, with Frame–LU pairs distinguishing specific meanings such as walk, jog, and run.HiPHI uses this linguistic organization analogously to ImageNet’s use of WordNet for structured visual concepts.
  • Systematic Expansion: HiPHI selects humanoid-relevant FrameNet units and expands them across path, direction, speed, rhythm, amplitude, posture, body-part involvement, support, and object/contact conditions.The selected units cover body motion, posture change, directional movement, body-part motion, object actuation, and human-object interaction.
  • Construction Framework: HiPHI organizes collection around Frame–LU motion seeds and shared factors, replacing case-by-case scripts with structured, scalable motion-space expansion.The construction also captures relevant object states alongside human motion to avoid physically incomplete interactions such as a person appearing to sit in the air.
  • Novelty: HiPHI is, to the authors’ knowledge, the first MoCap dataset for robot learning to prospectively adopt a linguistic scaffold for data construction.The scaffold is intended to support systematic enumeration and expansion rather than relying on manually designed task scripts.

4 Dataset Composition and Statistics

HiPHI comprises 617.5 hours of high-fidelity optical motion capture data, combining body-only and object-interaction sequences. Its Frame-LU indexing organizes the dataset across 214 labels and 22 FrameNet frames, while the interaction subset captures motions grounded in real objects and physical contact.

  • Scale and capture: 617.5 hours of optical motion capture data comprise approximately 200.1 million frames, captured at 90 Hz from 132 distinct performers.The dataset applies left-right mirroring to 308.7 hours of original motion capture data.
  • Frame-LU composition: Each clip has a Frame-LU label and natural-language description, with 214 Frame-LU labels distributed across 22 FrameNet frames.Frame-LU labels support retrieval, sampling, and analysis while situating motion meanings within event contexts.
  • Comparison with existing datasets: 245.7 hours of object-state-aligned motion give HiPHI greater scale and interaction coverage than representative human motion and interaction datasets.Table 1 identifies 617.5 hours as the total MoCap volume.
  • Object-interaction subset: 371.8 hours are body-only motion and 245.7 hours are human-object interaction motion, covering self-motion, posture, dynamic movement, and whole-body coordination.The object-interaction subset includes motions shaped by real geometry, contact, load, and object motion.
  • Object-interaction subset: The object-interaction subset contains 40 objects across 12 categories, spanning examples such as sitting, leaning, supporting, pushing, pulling, and carrying.Categories range from furniture and containers to additional object types described in the dataset composition.

5 Experiments

HiPHI is evaluated for dataset coverage and quality, then for downstream humanoid tracking in simulation and sim-to-real deployment. It spans the broadest reported kinematic region, achieves strong motion and object-interaction fidelity, and supports effective physical execution on simulated and real humanoids.

  • Dataset coverage: HiPHI spans the broadest kinematic region, covering most regions occupied by other datasets while extending into additional motion-space areas.The comparison uses a shared 16-D latent space and occupancy analysis with at most 5,000 windows per dataset.
  • Body-motion precision: HiPHI achieves the best value on every reported body-motion metric among datasets sharing the same floor-plane convention.The metrics capture jerk, acceleration, ground penetration, unsupported floating, and support-point drift; Motion-X++ is excluded from the three floor-related metrics.
  • Object-interaction consistency: 245.7 h vs. 21.6 h for HIMO and 9.8 h for OMOMO, HiPHI provides substantially longer object-interaction duration while achieving 98.1% non-conflict and 95.7% near-surface grounding.These grounding results improve over HIMO’s 79.0% and OMOMO’s 50.4% interaction-grounding values.
  • Humanoid tracking: HiPHI achieves the highest whole-body tracking success rates and fastest convergence in both matched 3-hour and 20-hour settings.Across five independent runs, it maintains lower failure rates and smaller variance than the compared datasets; it also achieves the best overall body tracking performance across most categories and best object-tracking results for several actions.
  • Sim-to-real deployment: On Unitree G1 hardware, HiPHI-trained policies successfully perform diverse locomotion, whole-body, and object-interaction behaviors despite actuation limits, sensing noise, and sim-to-real discrepancies.Demonstrated behaviors include sitting, crawling, carrying, flipping, and pulling.

6 Conclusion

HiPHI is a 617.5-hour high-precision optical MoCap dataset and benchmark, including 245.7 hours of synchronized human-object interaction, built through a Frame-LU-guided construction pipeline. Evaluations show broad motion coverage, high-quality and geometrically consistent interactions, improved humanoid tracking, scalable policy gains, and transfer to real humanoid hardware.

  • Dataset and construction: 617.5 hours comprise HiPHI’s high-precision optical MoCap dataset and benchmark, including 245.7 hours of synchronized human-object interaction with object trajectories and meshes.The dataset is built upon a Frame-LU-guided motion-space construction pipeline.
  • Dataset and construction: HiPHI provides a systematic and scalable approach for collecting broad whole-body motion and physically grounded interactions.
  • Evaluation and transfer: Evaluations demonstrate broader motion-space coverage, higher motion quality, stronger human-object geometric consistency, and improved downstream humanoid tracking performance.
  • Evaluation and transfer: Policies trained with HiPHI achieve strong matched-budget tracking, improve with increasing data scale, and transfer successfully to real humanoid hardware across diverse whole-body and object-interaction behaviors.
  • Conclusion: Together, these results establish HiPHI as a large-scale data foundation for physically grounded humanoid skills, bridging precision motion capture, object-aware interaction, and scalable policy learning.

7 Limitations · Appendix · A Additional Details for Dataset Construction

HiPHI is currently limited to single-person, studio-based kinematic capture without direct force or tactile measurements. Its appendices document dataset construction, release details, evaluation setups, training procedures, and FrameNet-guided realization protocols.

  • 7 Limitations: HiPHI captures only single-person motion, leaving multi-person interaction and human-human contact outside its scope.Complementary efforts are required for these interaction settings.
  • 7 Limitations: HiPHI provides precise motion and object trajectories but focuses on kinematics rather than directly measuring contact forces or tactile signals.The dataset is also collected in a studio, motivating future ego-vision capture in the wild.
  • Appendix: Appendix A-B describe FrameNet-guided motion-unit construction, release composition, metadata, and demographics.Appendix C covers licensing and ethics, while Appendix D presents representative sample sequences.
  • Appendix: Appendix E-F provide unsupervised motion-space embedding setup and full-run quality metrics, while Appendix G-H document whole-body and motion-with-object tracking training details.These materials complement the main text with evaluation and training specifications.
  • A Additional Details for Dataset Construction: HiPHI uses FrameNet as a semantic scaffold, seeding motion units with motion-relevant Frame-LU pairs that can expand into multiple concrete realizations.This keeps the public index compact while preserving semantic variation.
  • A Additional Details for Dataset Construction: Each Frame-LU seed becomes repeatable performer-facing capture instructions that vary observable factors including direction, path, speed, intensity, amplitude, height, body parts, objects, contact, and load.Examples include push variations by object, load, and direction; lean variations by support surface and height; and crawl variations by path shape.

B Additional Details for Dataset Composition and Metadata · C Release Terms, License, and Ethics · D Representative Sample Sequences

HiPHI’s release provides 617.5 hours of mirrored-release motion with FrameNet-based semantic metadata, performer and object-interaction annotations, and synchronized human-object assets. It will be publicly released under a non-commercial research license with ethics protections, while representative sequences illustrate the release format and quality.

  • B Additional Details for Dataset Composition and Metadata: 617.5 hours of motion correspond to approximately 200.1M frames at 90 Hz under the mirrored-release convention.The release includes 371.8 hours of body-only motion and 245.7 hours of strict object-interaction motion.
  • B Additional Details for Dataset Composition and Metadata: The semantic layer contains 22 FrameNet frames and 214 Frame-LU motion-unit labels, with the top 10, 20, and 50 labels covering 20.4%, 31.8%, and 53.7% of duration.Nearly half of the release duration remains outside the top 50 Frame-LUs.
  • B Additional Details for Dataset Composition and Metadata: HiPHI includes 132 anonymized performer profiles, comprising 76 male and 56 female profiles, with heights of 155-185 cm and weights of 40-82 kg.Figure 9 summarizes the gender, height, and weight distributions.
  • B Additional Details for Dataset Composition and Metadata: The object-interaction subset contains 40 real-world objects from 12 categories, spanning 0.45-6.25 kg, 15 FrameNet frames, and 90 Frame-LU labels.Strict object-interaction sequences provide synchronized object trajectories and object meshes with human BVH motion.
  • C Release Terms, License, and Ethics: HiPHI will be released under the ModalityNet Open Research License v1.0, a custom non-commercial license for scientific research, education, and evaluation.Complete license terms will accompany the public release.
  • C Release Terms, License, and Ethics: The full dataset, including BVH motions, object trajectories, object meshes, and the Frame-LU index, will be released publicly soon.The supplementary material contains four representative sample sequences described as three body-only and three object-interaction sequences.
  • C Release Terms, License, and Ethics: Data collection received institutional ethics review-board approval, and all adult performers provided written informed consent for capture, research use, and planned public release.Performers were compensated for each capture session and may request withdrawal of their data at any time.
  • D Representative Sample Sequences: Six representative HiPHI sequences are shown through per-sequence renderings in Figure 10, with full motion playback provided in the supplementary video.The public release contains BVH motion, object trajectories, and object meshes, without RGB video, audio, facial imagery, or voice recordings.

E Additional Details for Motion-Space Diversity

The motion-space analysis standardizes all datasets into a shared representation, uses balanced sampling and a joint unsupervised encoder, and evaluates coverage with grid-based statistics. HiPHI’s occupied-region advantage remains positive across tested random seeds, t-SNE perplexities, and grid resolutions.

  • Unified motion representation: All datasets are mapped to 23 keypoints in meters, resampled to 30 FPS, and represented in a z-up coordinate frame.Sequences are divided into 30-frame windows with a stride of 25 frames.
  • Unified motion representation: Each frame uses normalized heading-local, root-relative joint positions and velocities, producing a 23 × 3 × 2 = 138-dimensional feature.Body-scale normalization and per-window median-pose subtraction reduce scale and skeleton-layout differences across datasets.
  • Dataset-balanced sampling: Balanced sampling draws 5,000 motion clips per dataset when available and one valid 30-frame window from each selected clip.LAFAN1 uses all available clips because it contains fewer than 5,000 samples.
  • Shared unsupervised encoder: A shared temporal convolutional autoencoder encodes 30 × 138 inputs into a 16-D latent code using reconstruction loss alone.The model is trained jointly across balanced samples, without dataset identity, text labels, Frame-LU labels, or object labels.
  • Coverage robustness analysis: HiPHI maintains a positive occupied-cell margin in every tested configuration, so its coverage advantage persists across resampling, projection, and discretization choices.Tests vary random seeds 55, 73, and 112; perplexities 40, 55, and 65; and grid resolutions of 55 × 55, 88 × 88, and 111 × 111.

F Additional Details for Data Quality and Precision

The evaluation aggregates quality statistics by recorded duration and defines body-motion, ground-contact, and object-interaction metrics with explicit geometric and temporal conventions. Metric applicability depends on available timing, ground references, synchronized object data, and public dataset subsets.

  • Aggregation across sequences: Quality statistics weight observations by duration, so every recorded second contributes equally without allowing dataset size to bias results.The weighting uses source duration for frames, window length for windows, and sequence duration for sequence aggregates.
  • Aggregation across sequences: Ratio metrics use sequence-duration weighting to recover frame-level dataset aggregation, while distribution metrics aggregate per-frame or per-pair observations.Upper-tail statistics use the time-weighted quantile with τ = 0.95.
  • Body-motion metric sets and smoothing: Body-motion metrics use core joints, support points, and ground-contact frame-point pairs, with contact defined within ϵ = 30 mm of the ground.Finite differences are divided by each source’s frame interval ∆t.
  • Body-motion metric sets and smoothing: Jerk trajectories are smoothed over a fixed physical window of 1/6 s, and Motion-X++ is excluded from floor-related metrics because its absolute ground convention is not comparable.The moving average has odd length and at least three frames; raw positions are used for acceleration.
  • Object-interaction geometry: Object-interaction geometry samples canonical skeleton-segment centerlines at 32 points per segment and measures distances to the posed object mesh boundary in shared coordinates.The body proxy covers hands, arms, legs, torso, pelvis, and head/neck regions.
  • Applicability: Body-motion metrics require aligned joint trajectories and native timing, ground metrics require comparable up-axis and ground conventions, and interaction metrics require synchronized human, object-pose, and mesh data.HUMOTO is evaluated on its publicly released GLB subset rather than the full licensed dataset.

G Training Details for Whole-Body Motion Tracking

HiPHI motions are evaluated through a standardized physics-based imitation-learning pipeline after retargeting to the Unitree G1. The resulting policy produces stable, physically plausible tracking across diverse behaviors, supporting HiPHI as executable references for humanoid control.

  • Evaluation Protocol: All datasets use the same physics-based imitation-learning pipeline, observation design, reward formulation, optimization budget, and evaluation protocol.This standardization makes performance differences mainly reflect the physical usability of the underlying motion data.
  • Evaluation Protocol: Source motions are retargeted to the Unitree G1 humanoid using PyRoki before serving as reference trajectories for policy learning.The retargeted robot motions provide the references used by the learned tracking policy.
  • Tracking Results: HiPHI yields stable and physically plausible whole-body tracking across diverse behaviors on the Unitree G1 humanoid.These qualitative results demonstrate that HiPHI motions can function as effective executable references for humanoid control.

H Training Details for Motion-with-Object Tracking

HiPHI’s motion-with-object tracking pipeline retargets synchronized human-object demonstrations to the Unitree G1 while preserving object trajectories and uses an object-aware BeyondMimic-style policy. Object-state observations, tracking rewards, and domain randomization support robust reproduction of whole-body motion and interactions, with qualitative comparisons showing stronger stability and interaction consistency for HiPHI.

  • Data Processing: Omni-Retarget retargets synchronized human-object demonstrations to the Unitree G1 humanoid while preserving the corresponding object trajectories.
  • Data Processing: A unified object configuration is used across datasets because existing open-source human-object motion datasets usually omit object physical parameters.
  • Tracking Framework: The BeyondMimic-style setup extends observation and reward design to explicitly model object interactions alongside whole-body motion.
  • Tracking Framework: Object-state observations, position-and-orientation tracking rewards, and domain randomization over humanoid and object properties improve the interaction-aware tracking setup’s robustness.
  • Qualitative Results: HiPHI demonstrates more stable whole-body tracking and better human-object interaction consistency than OMOMO and HUMOTO in qualitative carry and kick comparisons.
Loading 2608.16222v1…