Source-linked AI summary

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

arXiv:2607.28625v1cs.CV

TL;DR

Existing embodied-AI datasets incompletely capture synchronized perception, motion, manipulation, and physical sensing in everyday environments. ACE addresses this gap with two-scale ambient capture and ACE-Data-0, where existing methods degrade under contact, occlusion, and long-duration activity.

  • Problem

    Existing datasets fragment synchronized perception, whole-body motion, dexterous manipulation, and physical sensing across viewpoints, modalities, and spatial scales.

  • Method

    ACE combines table- and room-scale configurations to synchronously capture egocentric and exocentric video, human and object motion, audio, and tactile signals.

  • Results

    Existing methods degrade across touch prediction, full-body motion recovery, and hand-motion estimation under contact, occlusion, and long-duration activity.

  • Takeaways & Limitations

    ACE-Data-0 provides synchronized demonstrations with contact and motion supervision for studying perception, action, and physical state in real homes.

  • Takeaways & Limitations

    ACE covers only two sites, limiting variation in layouts, furnishings, and lighting, and ground truth is restricted to instrumented entities.

Abstract

from arXiv · show

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

1 Introduction

ACE addresses the embodied-AI data bottleneck by turning real homes into synchronized recording studios for multimodal human interaction. It introduces ACE-Data-0 and a hierarchical benchmark spanning signals, scene components, and interactions.

  • Motivation: Everyday household activity requires coordinated perception, whole-body movement, dexterous manipulation, and physical sensing that existing models have yet to attain.
  • Ambient Capture Engine: ACE captures human actions, body and hand motion, object-state evolution, audio, tactile signals, and synchronized first- and third-person scene views.Its two complementary configurations target table-scale manipulation and room-scale whole-body activity.
  • ACE-Data-0: 150 hours, 200 task categories, 50 participants, 2 environments, 17M video frames, and 75,000 interaction episodes comprise ACE-Data-0.Participants follow goal-level instructions, preserving planning, hesitation, improvisation, and natural behavioral variation.
  • ACE-Data-0: ACE-Data-0 provides synchronized timelines, camera calibrations, human and hand poses, object meshes and 6-DoF poses, motion trails, and event descriptions.A large fraction of these annotations is derived automatically from tracked physical states.
  • Hierarchical benchmark: The benchmark advances from cross-modal signals, to body-and-hand scene components, to interaction estimation, with evaluations of more than 30 state-of-the-art methods.These tracks expose open challenges for embodied perception and robot learning.

2 Related Work

Prior datasets capture complementary aspects of embodied interaction, but typically emphasize isolated modalities, viewpoints, physical supervision, or task scales. ACE-Data-0 unifies egocentric and exocentric observation with synchronized multisensory physical grounding across manipulation and household activity.

  • Multimodal grounding: Everyday interaction couples visual appearance, body and hand motion, object state, audio, and touch, whose signals provide complementary evidence rather than interchangeable observations.Vision captures externally observable changes, kinematics tracks human and object evolution, audio reveals impacts and transitions, and touch directly indicates contact and force.
  • Physical interaction datasets: Earlier physically grounded datasets primarily target local grasping and manipulation, progressing from hand-object contact and pose to bimanual interaction, object tracking, and dynamic 4D understanding.Examples include ContactPose, DexYCB, H2O, H2O-3D, and HOI4D, with supervision centered on RGB-D observations, hand states, object poses, and contact.
  • Dexterity and task structure: Recent resources expand dexterity, compositionality, and task structure, using articulated objects, tool-action-object relationships, affordance hierarchies, and larger multi-camera collections.ARCTIC, TACO, OakInk2, and GigaHands exemplify this progression toward richer bimanual activity and language annotations.
  • ACE-Data-0 distinction: ACE-Data-0 registers wearable egocentric views with synchronized exocentric context, measured headset, human, and object motion, and shared multisensory physical ground truth.Its table-scale setup resolves fine hand-object manipulation, while its room-scale setup captures full-body activity across furnished domestic environments.
  • Egocentric and robot data: Egocentric datasets provide action-relevant first-person context, while manipulation-oriented and robot-demonstration datasets add physical supervision or executable trajectories but remain viewpoint- or embodiment-specific.Human-to-robot methods increasingly align human demonstrations with robot embodiments or learn transferable representations from mixed human and robot data.
  • Long-horizon household activity: ACE-Data-0 exposes long-horizon household structure through goal-level instructions, allowing natural object choices, task ordering, movement, hesitation, interruption, and recovery while synchronizing visual, kinematic, audio, and tactile signals.This design bridges geometric human-object interaction benchmarks, long-form egocentric video, and robot trajectory collections.

3 Ambient Capture Engine

ACE is a synchronized multimodal capture engine built around two complementary spatial scales to record holistic everyday interaction in realistic domestic environments. Its configurations combine close-range manipulation sensing with wide-area whole-room coverage and dedicated sensors for visual, kinematic, object, audio, and tactile signals.

  • Architecture: ACE captures interaction signals across modalities while synchronizing and registering them in a common frame within believably domestic environments.The system is designed to capture every important signal an interaction may produce while preserving shared spatial and temporal alignment.
  • Architecture: ACE uses table-scale and room-scale configurations because fine finger-object contact and whole-scene locomotion require different spatial coverage.Table-scale targets dexterous hand-object manipulation with close-range cameras, optical motion capture, and tactile gloves; room-scale covers a furnished apartment.
  • Capture environments: The environments preserve clutter, furniture, and spatial constraints while supporting dense sensing, including a 30-square-meter table-scale workspace with over 25 interactable objects.The table-scale workspace includes objects from more than 8 categories, while room-scale sensing is deployed across a furnished apartment.
  • Sensor suite: The shared sensor suite observes what participants see, how bodies and objects move, and what interactions sound and feel like through dedicated devices.Sensors include egocentric and exocentric cameras, motion-capture suits and gloves, tracked object geometry and 6-DoF trajectories, audio capture, and tactile pressure maps.

4 ACE-Data-0

Section 4 presents ACE-Data-0 through its data acquisition pipeline, collection process, and annotation pipeline. It covers multimodal recording, synchronization and calibration, dataset design and statistics, and the annotations provided with the dataset.

  • 4.1 Data acquisition pipeline: The data acquisition pipeline covers capture workflow, multimodal recording, sensor synchronization, and calibration.
  • 4.2 Data collection: The data collection process describes task design, capture settings, dataset statistics, and recorded modalities.
  • 4.3 Annotations: The dataset includes annotations produced by a dedicated annotation pipeline.

4.1 Data Acquisition Pipeline

ACE’s data acquisition pipeline synchronizes heterogeneous sensors to a common 60 Hz motion-capture clock, calibrates their spatial relationships, and standardizes capture sessions across both configurations. The resulting protocol jointly records video, motion, object poses, audio, and tactile signals as timestamped multisensory data.

  • Temporal alignment: ACE aligns cameras, motion capture, object trackers, and tactile sensors by referencing every device to the OptiTrack clock and its strictly simultaneous 60 Hz frames.The pipeline addresses independent device clocks and differing sampling rates to preserve contact-level temporal correspondence.
  • Temporal alignment: Per-camera time offsets reestimated during calibration are under 4 ms, and a per-take table maps every camera frame to its corresponding 60 Hz motion-capture frame.Downstream processing consumes this timing table for both exocentric and egocentric cameras.
  • Spatial calibration: Exocentric cameras are calibrated through an ArUco board visible to RGB and OptiTrack systems, bridging camera pairs that lack shared views.Retroreflective markers at the board’s corners allow the motion-capture system to connect the cameras spatially.
  • Spatial calibration: Egocentric camera poses are measured from the ACE-Ego-Head’s rigid marker body rather than estimated from vision, achieving a median reprojection error of about 2 px.OptiTrack tracks the rig at 60 Hz, and poses for 20 FPS egocentric frames are obtained by timestamp interpolation and hand-eye transformation.
  • Capture protocol: Each capture session uses five standardized steps: randomized scene preparation, participant setup, goal-level task briefing, simultaneous acquisition, and session completion.The protocol is applied at both sites, with participants receiving verbal goals rather than prescribed action sequences.
  • Capture protocol: During acquisition, ACE records exocentric RGB, 4 egocentric fisheye streams with IMU, full-body and hand motion at 60 Hz, tracked-object 6-DoF poses at 60 Hz, and tactile signals.All streams share the common clock, and a one-hour session produces approximately 1 TB of raw data.

4.2 Data Collection

ACE-Data-0 uses goal-level household tasks to capture natural variation across atomic manipulation, long-horizon activity chains, and human-scene interaction. The collection comprises more than 150 hours of synchronized multimodal data, over 17M frames, and more than 75,000 episodes from 50 participants.

  • Task design: Goal-level instructions let participants choose different sub-task orders, grasps, routes, objects, and timing, introducing natural behavioral variability.The task design leaves how to reach each household goal open rather than prescribing fixed action sequences.
  • Task design: Atomic HOI takes last roughly three minutes and provide self-contained everyday manipulation instances drawn from more than 15 household activity types.Examples include pouring water, drinking, making tea, watering plants, chopping vegetables, cooking, and tidying up.
  • Task design: Chains of HOI takes last roughly twenty to thirty minutes, interleave sub-tasks, and exercise long-horizon planning, state tracking, and memory.Each take ends with the scene tidied back into order, completing a full household-activity cycle.
  • Task design: HSI takes last about five minutes and record whole-body motion and human-scene contact involving tables, chairs, and sofas.Examples include walking, exercising, sitting, lying, and leaning.
  • Dataset statistics: More than 150 hours of synchronized multimodal capture comprise over 17M frames across more than 75,000 episodes.Episodes are contiguous, semantically self-contained sub-goal segments counted within takes, preserving surrounding context.

4.3 Annotation

ACE-Data-0 enriches every take with five synchronized annotation types spanning objects, body and hand configuration, contact, and activity descriptions. Most annotations derive directly from measured states, calibration, and tactile sensing, enabling metrically aligned multimodal supervision without extensive manual labeling.

  • Annotation types: Five annotation types jointly characterize each take: objects, body configuration, hand articulation, hand-object contact, and acoustic and semantic activity descriptions.Object annotations include category labels, bounding boxes, and per-frame 6-DoF poses for tracked objects.
  • Calibration and synchronization: Every take includes camera intrinsics, world-frame camera poses, and a shared timeline aligning all sensor streams.These calibration outputs allow tracked 3D points to project into any view and arbitrary modalities to be paired at any instant.
  • Human poses: Measured 3D body and hand states generate pixel-aligned 2D poses in every camera view and metric 3D poses in the world frame.Because they are reprojected from measured states rather than image detectors, the poses remain reliable under furniture occlusion, extreme viewpoints, and motion blur.
  • Tactile labels: Directly sensed tactile readings identify when and with what an interaction occurs, rather than inferring contact from appearance or geometric proximity.Each frame pairs the tactile glove signal with the visual streams and the currently used object label.
  • Annotation production: All annotations except textual descriptions are measured rather than estimated, so projections, boxes, motion trails, and contact events follow from recorded states and tactile sensing.This design avoids most of the otherwise prohibitive manual effort required to produce annotations at this breadth.

5 Benchmark

The benchmark evaluates existing methods on ACE-Data-0 with a diagnostic focus on failures in long-horizon home-scene data. It is organized hierarchically into three levels: signals, components, and interactions.

  • Benchmark structure: Three benchmark levels progress from raw sensory-signal inference to scene components and then interactions.Each level builds upon the one beneath it. Unless otherwise noted, baselines use officially released pre-trained checkpoints for fair comparison.

5.1 Low-level Signals

This section benchmarks ego-view full-hand grasp-pressure prediction from video against synchronized tactile-glove measurements. Results show that contact timing is easier to infer than fine-grained pressure location, making robust pressure reconstruction challenging under visual ambiguity and occlusion.

  • Task: The task is to predict full-hand grasp pressure at each moment from egocentric human-object interaction video.Predictions are evaluated against synchronized tactile-glove measurements.
  • Metrics: Four metrics assess temporal contact accuracy, spatial contact overlap, pressure-aware volumetric overlap, and fingertip pressure-centroid error.The metrics are temporal accuracy, Contact IoU (C-IoU), volumetric IoU (V-IoU), and Center-of-Pressure (CoP) error.
  • Baselines: The benchmark compares PressureVision, EgoPressureDiff, and TouchAnything using their released pre-trained weights.The methods use convolutional regression, conditional diffusion, and cross-view vision-to-touch representation learning, respectively.
  • Results: PressureVision produces nearly zero C-IoU and V-IoU with the largest CoP error, while EgoPressureDiff improves temporal recognition and pressure localization but retains limited spatial overlap.These results come from close-range table-scale recordings.
  • Results: Existing tactile-estimation models exhibit a substantial generalization gap, requiring both contact-event recognition and fine-grained pressure-distribution recovery.The challenge is heightened by visually ambiguous and frequently occluded interactions in ego-view recordings.

5.2 Scene Components

ACE-Data-0 evaluates scene-component estimation in furnished domestic scenes using metric ground truth, covering body motion across viewpoints and method families. Results reveal strong local pose recovery but substantially weaker global trajectory estimation, with scene context helping localization and egocentric views remaining most difficult.

  • Human motion estimation: Furniture occlusion, uncommon household postures, and multi-minute sequences make human motion estimation difficult and can cause global trajectory drift.Examples include crouching at a cabinet, reaching overhead, and lying on a sofa.
  • Evaluation setup: The evaluation compares multi-view exocentric, single-view exocentric, and egocentric video across per-frame, temporal, and scene-aware method families against motion-captured ground truth.Using shared ground truth allows direct comparison of viewpoint, temporal information, and scene context.
  • Results: Local pose metrics remain strong under household occlusion and uncommon postures, whereas world-frame trajectory errors are substantially higher and produce different method rankings.This exposes a clear gap between articulated pose recovery and global trajectory recovery.
  • Results: Scene-aware methods achieve lower trajectory errors while retaining similar Procrustes-aligned results to temporal methods, and egocentric methods perform worse across most metrics.Scene geometry mainly constrains where a person stands or moves, while egocentric views omit much of the body and require inference from head motion.
  • Results: Multi-view methods do not outperform the strongest single-view methods in this evaluation, likely reflecting strong recent single-view methods and limited multi-view baselines.The comparison is not interpreted as evidence that multiple views are less useful.

5.3 Embodied Interaction

This section evaluates hand-motion recovery from synchronized egocentric and exocentric video under identical interactions and ground truth. Exocentric methods achieve stronger articulation and global trajectory results, while the viewpoints remain complementary because each has distinct failure modes.

  • Evaluation setup: ACE-Data-0 compares egocentric and exocentric hand-motion estimation using the same takes, interactions, and ground truth.The task estimates articulated MANO hand sequences and, when supported, world-coordinate trajectories.
  • Complementary viewpoints: Egocentric video is affected by truncation, distortion, and motion blur, whereas exocentric video is affected by occlusion from bodies and manipulated objects.Combining both streams could reduce their individual failure cases; measured headset motion could also help separate reconstruction from egomotion errors.
  • Exocentric results: 9.1–10.8 mm PA-MPJPE: exocentric methods achieve similar articulation accuracy, with WiLoR best at 9.1 mm, F@5 0.313, and AUCJ 0.819.HaMeR follows with 9.6 mm PA-MPJPE, while HORT reaches 10.8 mm on manipulation frames.
  • Global trajectory: 63 mm trajectory error: HaPTIC substantially outperforms egocentric world-space methods at 98–102 mm while maintaining 10.0 mm PA-MPJPE.The fixed exocentric camera provides a stable coordinate system that reduces uncertainty from egomotion.
  • Viewpoint comparison: Exocentric methods outperform egocentric methods for both articulation and global trajectory estimation, with the largest gap in trajectory error.The exocentric result is 63 mm, whereas egocentric methods remain close to 100 mm.

6 Conclusion

ACE records synchronized multisensory household behavior at table and room scales, producing ACE-Data-0 for evaluating embodied perception and motion methods. The dataset exposes degradation under contact, occlusion, and long-duration activity while supporting research connecting perception, action, and physical state.

  • ACE methodology: ACE combines table-scale hand-object capture with room-scale whole-body activity recording in furnished homes.Both configurations synchronize egocentric and exocentric video, body, hand, and object motion, audio, and tactile signals using optical-clock alignment.
  • ACE-Data-0: Over 150 hours, 75,000 interaction episodes, and 17M frames comprise ACE-Data-0 with per-frame annotations.The dataset supports benchmarks for touch prediction from video, full-body motion recovery, and hand-motion estimation across egocentric and exocentric views.
  • Benchmark findings: Existing methods degrade under contact, occlusion, and long-duration activity, conditions distinguishing real homes from controlled laboratories.These benchmark results motivate models that address the challenges of everyday household behavior.
  • Research applications: ACE-Data-0 combines egocentric observations, demonstration trajectories, and contact-level supervision in one time-aligned stream for manipulation policies, world models, and vision-language-action systems.Its design targets systems connecting perception, action, and physical state in real homes.
  • Limitations: ACE covers only two sites, limiting variation in layouts, furnishings, and lighting.This limitation constrains the environmental diversity represented by the dataset.
Loading 2607.28625v1…