Source-linked AI summary

HumanNet: Scaling Human-centric Video Learning to One Million Hours

Yufan Deng, Daquan Zhou

arXiv:2605.06747v1cs.CVcs.RO

TL;DR

Embodied learning remains constrained by smaller, narrower, robot-specific datasets than those enabling progress in language and vision-language modeling. HumanNet introduces a curated one-million-hour corpus of human-centric video, and controlled validation shows 1,000 hours of egocentric HumanNet video matches or modestly surpasses 100 hours of real-robot data while narrowing the gap to a 20,000-hour baseline.

  • Problem

    Embodied learning remains data-limited because physical-interaction datasets are orders of magnitude smaller, narrower, and more robot-specific than web-scale language and vision-language corpora.

  • Method

    HumanNet constructs a one-million-hour corpus of curated first- and third-person human-centric video with interaction-focused annotations and a multi-axis taxonomy.

  • Results

    1,000 hours of egocentric HumanNet pretraining matches or modestly surpasses 100 hours of real-robot pretraining and substantially closes the gap to a 20,000-hour real-robot baseline.

  • Takeaways & Limitations

    Egocentric human video can serve as a scalable and cost-effective substitute when robot data is limited.

  • Takeaways & Limitations

    HumanNet does not eliminate the embodiment gap between human behavior and robot control spaces, so its value is expected in transferable priors rather than direct replacement of robot data.

Abstract

from arXiv · show

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly annotated human activity data. We present HumanNet, a one-million-hour human-centric video corpus that captures how humans interact with the physical world at scale. HumanNet spans both first-person and third-person perspectives and covers fine-grained activities, human-object interactions, tool use, and long-horizon behaviors across diverse real-world environments. Beyond raw video, the dataset provides interaction-centric annotations, including captions, motion descriptions, and hand and body-related signals, enabling motion-aware and interaction-aware learning. Beyond scale, HumanNet introduces a systematic data curation paradigm for embodied learning, where human-centric filtering, temporal structuring, viewpoint diversity, and annotation enrichment are treated as first-class design principles. This design transforms unstructured internet video into a scalable substrate for representation learning, activity understanding, motion generation, and human-to-robot transfer. We conduct a first-step validation on the value of this design through controlled vision-language-action ablation: under a fixed set of validation data, continued training from the Qwen VLM model with 1000 hours of egocentric video drawn from HumanNet surpasses the continued training with 100 hours of real-robot data from Magic Cobot, indicating that egocentric human video could be a scalable and cost-effective substitute for robot data. By building this project, we aim to explore the opportunity to scale embodied foundation models using human-centric videos, rather than relying solely on robot-specific data.

1 Introduction

HumanNet addresses the data bottleneck in embodied learning with a one-million-hour, human-centric video corpus spanning first-person and third-person views. Its curation and annotation design supports physical activity understanding, motion-aware learning, and embodied pretraining, while controlled validation shows egocentric HumanNet data can match or surpass substantially less robot data under a fixed downstream regime.

  • Corpus and motivation: Its design prioritizes breadth across activities, environments, objects, body motions, interaction styles, and camera viewpoints while preserving physical structure for embodied learning.The dataset is intended to support fine-grained activity understanding, motion-aware representation learning, procedural reasoning, and human-to-robot transfer.
  • Corpus and motivation: HumanNet introduces one million hours of human-centric video spanning first-person and third-person views of fine-grained physical activities.The corpus is organized by source type, viewpoint, task structure, environment, interaction style, motion category, and metadata availability.
  • Data curation: The curation pipeline converts heterogeneous web video into pretraining infrastructure through human-centric filtering, viewpoint characterization, segmentation, deduplication, quality control, privacy review, and caption and motion annotation.These steps support representation learning, motion-aware video modeling, and embodied pretraining.
  • Controlled validation: 1,000 hours of egocentric HumanNet pretraining matches or modestly surpasses 100 hours of real-robot from Magic Cobot pretraining under an identical downstream regime.The controlled vision-language-action study holds the policy architecture and downstream corpus fixed, and the HumanNet model substantially closes the gap to a 20,000-hour real-robot baseline.

2 Related Work

Prior work has established human-centric video as a foundation for learning visual, temporal, and physical structure, while emerging approaches use human demonstrations and video supervision to support robot learning. These efforts span broad third-person activity datasets, actor-centered first-person datasets, and transfer-oriented representation learning.

  • Human-centric activity datasets: Human activity datasets support learning visual, temporal, and physical structure from naturally occurring behavior.Third-person resources cover broad actions, household activities, localized behavior, and object-centric temporal reasoning, while first-person datasets expose actor-centered intent and hand-object interactions.
  • Human-centric activity datasets: Third-person datasets include ActivityNet, Kinetics, Charades, AVA, and Something-Something, covering broad actions and object-centric temporal reasoning.These datasets also represent household activities and localized human behavior.
  • Human-centric activity datasets: First-person datasets such as EPIC-KITCHENS and Ego4D expose actor-centered intent and hand-object interactions.The supplied passage identifies these datasets as complementary first-person resources within human-centric activity learning.
  • Robot learning from human data: Human data complement robot supervision by providing diverse demonstrations of manipulation, tool use, locomotion, and procedural behavior at difficult-to-match scale.Prior work has used passive human video and broad visual pretraining to learn representations that transfer to downstream control.

3 The 1M-Hour Human-Centric Video Dataset

HumanNet is a one-million-hour human-centric video corpus designed around physically meaningful human activity across diverse viewpoints, environments, interactions, and task horizons. Its staged collection, processing, and annotation pipeline converts heterogeneous internet and real-world video into quality-controlled clips with motion, geometric, semantic, and robot-relevant supervision.

  • Dataset scope: The corpus spans first-person and third-person views, diverse objects, environments, body configurations, task variations, motion styles, social contexts, and long-horizon behaviors.First-person recordings capture actor-centered intent and hand-object contact, while third-person recordings broaden coverage across settings and viewpoints.
  • Dataset scope: Human-centric clips organize around physically meaningful behavior, including object manipulation, tool use, navigation, assembly, transport, coordination, and multi-step procedures.The organizing signal is human activity with visible physical interaction or environment state changes, rather than passive observation.
  • Construction pipeline: The construction pipeline separates data collection, clip-level processing, and annotation so each stage can be audited, extended, or rerun independently.Collection combines keyword discovery, crawling, search, and existing sources; processing performs deduplication, normalization, content filtering, and quality filtering.
  • Annotation: Annotations add 3D hand and body pose, camera trajectory, humanoid-skeleton retargeting, captions, motion categories, contact estimates, scene tags, and procedural boundaries.The supervision connects video to motion geometry, robot-relevant kinematics, and activity semantics; clips meeting retargeting conditions can be designated robot-ready.
  • Corpus characterization: Quality-filtered clips concentrate high-confidence pose supervision while retaining heavy-tailed motion scores and lengths, combining dense supervision with coverage of long-tail behaviors.The resulting structure supports mixed-supervision training by matching downstream tasks to appropriate corpus subsets.

4 Downstream Relevance

HumanNet is designed as a versatile substrate for video/VLM pretraining, world-action modeling, motion-aware representation learning, multimodal physical-AI objectives, and human-to-robot transfer. Its central advantage is combining first- and third-person viewpoints with annotations that preserve human activity, motion, contact, and interaction structure.

  • Video and VLM pretraining: First- and third-person clips support video encoders and VLMs by exposing complementary object engagement, body pose, spatial context, and social interactions.First-person footage emphasizes actor-object engagement, while third-person footage reveals pose, scene context, and interactions among people.
  • World-action model training: The corpus supports world-action models that jointly learn environment dynamics and the actions driving hand-object contact, tool use, body motion, and scene changes.Caption labels and motion annotations enable action-conditioned forward-dynamics learning and prediction of future visual states from past observations.
  • Multimodal objectives for physical AI: Where metadata permits, the data supports masked or predictive video modeling, language-video alignment, boundary prediction, hand-object learning, pose prediction, and caption-conditioned activity modeling.These objectives share a requirement for scale paired with annotations preserving physically meaningful interaction structure.

5 Conclusion

HumanNet is a one-million-hour human-centric video corpus combining diverse viewpoints with interaction-focused annotations and systematic curation. In controlled vision-language-action post-training, 1,000 hours of HumanNet egocentric video matches or modestly surpasses initialization from 100 hours of real-robot data.

  • Corpus and curation: HumanNet contributes a one-million-hour human-centric video corpus pairing first-person and third-person footage with caption labels, motion annotations, and hand and body signals.The corpus is organized by a multi-axis taxonomy and curated with filtering, viewpoint characterization, quality control, and privacy review as first-class design choices.
  • Corpus and curation: HumanNet treats filtering, viewpoint characterization, quality control, and privacy review as first-class design choices in its curation pipeline.This pipeline organizes the annotated footage through a multi-axis taxonomy.
  • Validation: 1,000 hours of egocentric video drawn from HumanNet matches or modestly surpasses initialization from 100 hours of real-robot data under controlled vision-language-action post-training.The comparison uses the same controlled post-training protocol while varying the initialization data source and amount.

6 Limitations, Ethics, and Broader Impact

HumanNet offers transferable representation-learning value but does not eliminate the embodiment gap or the need for robot data. Its scale also brings annotation noise, uneven coverage, privacy and safety risks, and dual-use consequences.

  • Limitations: Human-centered video cannot directly replace robot data because human bodies, hands, tools, mobility, and control spaces differ from robots.The dataset’s expected value is representation learning and transferable priors.
  • Limitations: One-million-hour scale introduces ambiguous labels, inconsistent task boundaries, missing metadata, viewpoint imbalance, and variable visual quality.Captions, pose estimates, and motion annotations expand coverage but introduce additional errors, motivating transparent confidence and subset-quality reporting.
  • Limitations: Large-scale coverage can remain biased across geographies, socioeconomic contexts, occupations, viewpoints, body types, routines, and public activities.Without careful analysis, scale may create an illusion of universality while leaving significant blind spots.
  • Ethics: Human-centric video raises privacy and safety risks by capturing bystanders, sensitive interiors, private documents, screens, identifiable people, and proprietary workflows.Public release requires license review, redaction, restricted-content filtering, and access controls where necessary.
  • Broader Impact: The dataset may accelerate assistive systems, robotic manipulation, procedural understanding, motion modeling, and physical AI while also enabling surveillance-adjacent systems and bias inheritance.Its broader impact is therefore dual-use, with potential benefits and harms arising from the same data.
Loading 2605.06747v1…