Source-linked AI summary

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks

Yihang Li, Xuelong Wei, Jingzhou Luo, Yingjing Xiao, Yibo Bai, Guangyuan Zhou, Teng Zou, Chenguang Gui, Jiajun Wen, He Zhang, Kangliang Chen, Xing Pan, Shuaiyan Liu, Daming Wang, Tao An, Jiayi Li, Shibo Jin, Wanwan Zhang, Tianyu Wang, Boren Wei, Zhixuan Huang, Fangsheng Liu, Ruodai Li, Hui Zhang, Anson Li, Yicheng Gong, Peng Cao, Jiaming Liang, Liang Lin

arXiv:2604.23570v1cs.RO

TL;DR

Robot manipulation learning lacks large-scale, diverse, and extensible datasets, while existing collection paradigms face scalability or deployment constraints. EgoLive introduces a large-scale egocentric dataset built with custom hardware, automated multimodal annotation, and unconstrained real-world collection, positioning it as an open-source resource for generalizable robot learning.

  • Problem

    Robot learning is hindered by scarce large-scale datasets, while existing datasets have restricted environmental diversity and poor extensibility.

  • Method

    EgoLive combines custom head-mounted stereo capture, a scalable collection pipeline, and automated high-accuracy multimodal annotation for real-world human task demonstrations.

  • Results

    EgoLive is presented as the largest open-source annotated egocentric dataset for real-world human tasks, collected entirely in unconstrained environments with improved diversity and ecological validity.

  • Takeaways & Limitations

    The dataset provides an expanding knowledge base of natural human demonstrations intended to support generalizable robotic models and real-world robot deployment.

  • Takeaways & Limitations

    Real-robot teleoperation remains costly and difficult to scale because it requires specialized hardware and intensive human involvement.

Abstract

from arXiv · show

The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and universal manipulation interfaces dominate current datasets, they suffer from inherent limitations in scalability and real-world deployability. Human egocentric video collection, by contrast, has emerged as a promising approach to enable scalable, natural and in-the-wild data collection. As such, we present EgoLive, a large-scale, high-quality egocentric dataset designed explicitly for robot manipulation learning. EgoLive establishes three distinctive technical advantages over existing egocentric datasets: first, it represents the largest open-source annotated egocentric dataset focused on real-world task-oriented human routines to date; second, it delivers leading data quality via a customized head-mounted capture device and comprehensive high-precision multi-modal annotations; third, all data is collected exclusively in unconstrained real-world scenarios and encompasses vertical field human working data, including home service, retail, and other practical work scenarios, providing superior diversity and ecological validity. With the introduction of EgoLive, we aim to provide the research community with a scalable, high-quality dataset that accelerates breakthroughs in generalizable robotic models and facilitates the real-world deployment of robot systems.

1 Introduction

Robot manipulation learning is limited by datasets that lack environmental diversity and scalable collection. EgoLive addresses these gaps with a large, annotated egocentric dataset collected from unconstrained real-world human tasks.

  • Existing robot datasets are constrained by restricted environmental diversity and poor extensibility, limiting large-scale collection for generalizable manipulation learning.
  • Egocentric collection supports natural first-person perception, unrestricted movement, and scalable recording across diverse environments and users.
  • EgoLive contains 1,680 hours of stereo video at 60 frames per second, covering 65,866 episodes across 346 real-world tasks.
  • EgoLive is presented as the largest open-source annotated egocentric dataset focused on real-world human tasks.
  • Its capture system provides 130° × 130° field of view, 60 frames per second, and 2160×2160 resolution per camera, alongside high-accuracy multimodal annotations.
  • All data comes from unconstrained real-world environments, improving scene diversity and ecological validity for robot learning.

2 Related Work

Prior manipulation datasets trade off fidelity, scalability, embodiment compatibility, annotation richness, and operational coverage. EgoLive combines real-world service scenarios with broad multimodal annotations and a capture system designed for scalable human-to-robot learning.

  • Manipulation Data Collection: Real-robot teleoperation provides high-fidelity trajectories and physical priors but requires specialized hardware, intensive human involvement, and high collection costs.
  • Manipulation Data Collection: Universal Manipulation Interfaces enable natural demonstrations and reduce visual domain gaps, but are often tailored to specific robot embodiments and have limited cross-platform compatibility.
  • Manipulation Data Collection: Human egocentric video collection removes many hardware and spatial constraints, accelerating real-world data acquisition through wearable cameras.
  • Human Egocentric Datasets: Existing egocentric datasets prioritize semantic breadth, interaction fidelity, or deployment-scale operational coverage as distinct design goals.
  • Human Egocentric Datasets: EgoLive is compared with representative datasets using collection duration, spatiotemporal resolution, and annotation comprehensiveness, targeting real-world scenes with the second-longest collection duration.
  • Human Egocentric Datasets: EgoLive targets underexplored service-oriented settings such as home services, retail, and pharmacy, with annotations for camera pose, 3D hand keypoints, depth, masks, and sub-task segmentation.
  • Learning from Egocentric Data: Learning from egocentric data spans capability learning, human-to-robot transfer, and whole-body control.
  • Learning from Egocentric Data: JoyEgoCam uses stereo RGB cameras with a wide field of view and an integrated IMU to support human motion and camera-pose estimation.

3 Dataset

EgoLive combines lightweight stereo egocentric capture with an automated pipeline for geometric, kinematic, and semantic annotation. Its analyses indicate broad semantic coverage, long-tailed distributions, and diverse yet locally coherent interaction patterns.

  • Dataset overview: EgoLive is a large-scale in-the-wild manipulation dataset with rich, high-quality annotations.The dataset is designed for downstream egocentric perception, hand-object interaction understanding, and manipulation policy learning.
  • Data collection: JoyEgoCam uses stereo RGB cameras, a wide human-like field of view, high resolution, high frame rate, and an IMU for camera-pose estimation.The device is designed for long-term, minimally intrusive collection of natural human actions.
  • Automated annotation: The annotation pipeline processes binocular video and sensor data into motion tracking, semantic understanding, depth reconstruction, camera localization, sub-task segmentation, and instruction captions.Its components jointly provide geometric, kinematic, and semantic supervision.
  • Semantic composition: EgoLive spans manipulation-intensive scenarios and uses instruction captions to represent actions, objects, and object attributes.Task categories include household services, organization, cleaning, logistics, and related activities.
  • Dataset distribution: Compared with EgoDex and Xperience-10M, EgoLive shows broader semantic coverage and longer tails across objects, actions, and attributes.Its continuous embedding analysis also places EgoLive across a broader representation manifold with more locally coherent clusters.

4 Accuracy Evaluation

EgoLive’s evaluations assess hand reconstruction, depth reconstruction, and instruction captioning through qualitative alignment, calibration-based measurement, and consistency judgments. The reported examples show accurate hand and scene geometry alongside structured captions for atomic manipulation actions.

  • Hand reconstruction: EgoLive’s 2D hand keypoints align closely with observed hands, unlike the localization errors and spatial misalignments shown for EgoDex.The comparison is presented as qualitative evidence of more accurate and robust hand annotations.
  • Hand reconstruction: Stereo-space optimization produces 3D hand keypoints that align with reconstructed hand point clouds and maintain physical scale under dynamic occlusions.The framework is reported to address depth drift while preserving consistent absolute scale.
  • Depth reconstruction: Depth reconstruction is evaluated against checkerboard references captured across operating distances from 500mm to 3500mm.The protocol compares predicted depths with geometry-derived reference depths after stereo rectification and pose optimization.
  • Depth reconstruction: Qualitative real-world examples show that JoyEgoCam and the depth reconstruction method recover scene structure and complex 3D geometry.The examples cover multiple representative collection scenes.
  • Instruction captioning: Instruction captions are generated for atomic sub-task clips and evaluated for hand, object, action, and global consistency.The evaluation targets whether captions faithfully describe interaction elements and the complete manipulation behavior.

5 Conclusion

EgoLive is presented as a large, high-quality egocentric dataset for realistic human operational activities. Its scalable hardware and production pipeline are intended to support continued growth and robotics research connecting human behavior with robot action.

  • Dataset scope: EgoLive targets realistic, utility-driven human operational activities with large scale, high quality, and diverse coverage.The conclusion characterizes it as the world’s largest open-source egocentric dataset for these activities.
  • Scalability: Purpose-built hardware and a streamlined data production pipeline are designed to sustain increasing dataset scale and coverage.The conclusion describes the resulting resource as an ever-growing knowledge base of natural human behavioral priors.
  • Research implications: The dataset is positioned to support human-to-robot alignment, humanoid robot policy learning, and research bridging human behavior with robot action.These are the research directions explicitly identified in the conclusion.

6 Contributions

The supplied passage identifies the paper’s author team and marks Yihang Li and Liang Lin with dagger symbols. No contribution-specific claims are provided in this passage.

  • Authorship: The paper lists Yihang Li as an author.Yihang Li is the first listed author.
  • Authorship: The paper lists Liang Lin as an author.Liang Lin appears near the end of the author list and is marked with a dagger symbol.
Loading 2604.23570v1…