Source-linked AI summary

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zach Chavis, Joya Chen, Feng Cheng, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Maria Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang, Md Mohaiminul Islam, Suyog Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Miguel Martin, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh Kumar Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Edward Zhang, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina Gonzalez, Prince Gupta, Jiabo Hu, Yifei Huang, Yiming Huang, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu, Mi Luo, Zhengyi Luo, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Wei Shan, Kiran Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Huiyu Wang, Yuchen Wang, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao, Ziwei Zhao, Zhifan Zhu, Jeff Zhuo, Pablo Arbelaez, Gedas Bertasius, David Crandall, Dima Damen, Jakob Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato, Manolis Savva, Jianbo Shi, Mike Zheng Shou, Michael Wray

arXiv:2311.18259v4cs.CVcs.AI

TL;DR

Existing ego-exo datasets are few, small, unsynchronized, staged, or curated, limiting fluid movement between first- and third-person perspectives. Ego-Exo4D addresses this gap with a large-scale multimodal dataset and benchmark suite for skilled activities, with temporal cues improving object correspondence in one evaluated setting.

  • Problem

    Existing ego-exo datasets are few, small, unsynchronized, staged, or curated, leaving fluid first- and third-person activity understanding out of reach.

  • Method

    Ego-Exo4D collects time-synchronized egocentric and multiple exocentric videos of skilled activities in diverse authentic settings, using domain-specific camera placement and multimodal annotations.

  • Results

    IoU improves from 13.88% to 22.14% in the Ego→Exo setting when spatio-temporal baselines replace spatial baselines for object correspondence.

  • Takeaways & Limitations

    The dataset provides a foundation for research on ego-exo video learning and multimodal perception across skilled human activities.

  • Takeaways & Limitations

    Exocentric-camera placement is difficult across settings, causing occlusion, annotation impacts, and manual selection of suitable cameras; activity durations also create a long-tailed distribution.

Abstract

from arXiv · show

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions -- including a novel "expert commentary" done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Project page: http://ego-exo4d-data.org/

1 Introduction

Ego-Exo4D addresses the lack of large-scale, synchronized first- and third-person resources for understanding skilled human activity by providing a diverse multimodal dataset and benchmark suite. It combines complementary viewpoints, extensive annotations, and tasks spanning activity understanding, proficiency, cross-view translation, and 3D pose.

  • Motivation and gap: Existing ego-exo datasets are few, small, poorly synchronized, or staged, while instructional datasets generally provide only a single viewpoint.These limitations leave fluid movement between first- and third-person perspectives out of reach.
  • Dataset: Ego-Exo4D captures 1,286 hours from 740 participants across 123 natural scenes, pairing egocentric video with 4-5 synchronized exocentric streams.Sequences span 1 to 42 minutes and are localized in a metric, gravity-aligned reference frame.
  • Dataset: The dataset covers unscripted skilled physical and procedural activities, including dance, sports, music, cooking, bike repair, and health care, across varied proficiency levels.Captures occur in natural settings rather than scripted laboratory environments.
  • Multimodal resources: Multimodal recordings include 7-channel audio, IMU, eye gaze, SLAM cameras, 3D point clouds, and time-indexed first-person, third-person, and expert language descriptions.Expert commentary focuses on how an activity is executed and critiques performance-specific subtleties.
  • Benchmarks: The benchmark formalizes four task families: ego-exo relation, fine-grained recognition, proficiency estimation, and ego 3D pose.The tasks address viewpoint translation, keysteps and structure, execution quality, and skilled body and hand movements.
  • Impact: The authors open source the data, annotations, camera-rig protocol, and benchmarks to support research in ego-exo, multimodal activity understanding, and broader 3D vision.The exocentric component also supports traditional exocentric activity recognition and body-pose estimation.

2 Related work

Prior datasets rarely combine synchronized egocentric and exocentric video at real-world scale, while existing work often emphasizes single viewpoints, static scenes, or limited skill analysis. Ego-Exo4D addresses these gaps with broader multimodal resources, annotations, and benchmarks for skilled activity understanding.

  • Related research has addressed egocentric video, exocentric video, instructional activity, skill quality, and cross-view modeling, but these areas remain only partially integrated.
  • The dataset provides more modalities, synchronized language resources, key-step segments, object masks, and extensive 3D pose annotations than comparable ego-exo datasets.
  • 740 participants, 123 scenes, 13 cities, and 1,286 hours make Ego-Exo4D an order of magnitude larger and more diverse than prior efforts.
  • Ego-Exo4D targets skilled activities in authentic settings, extending beyond tabletop or household tasks to varied full-body movements and object interactions.

3 Ego-Exo4D dataset

Ego-Exo4D was collected through a coordinated, portable ego-exo rig built around Aria glasses and synchronized GoPros. The dataset spans skilled physical and procedural activities across authentic, geographically varied settings with rich sensor and spatial outputs.

  • A coordinated effort across 12 research labs used common guidelines, scenarios, and camera rigs to create a cohesive but diverse dataset.
  • 3.1 Ego-exo camera rig: The low-cost rig is portable, auto-synchronized, internationally attainable, and combines one Aria device with four GoPros and supporting equipment.The common rig costs under $3,000 excluding the Aria, phone, and laptop.
  • 3.1.1 Aria device and sensors: Aria records RGB, monochrome and eye cameras, seven-channel audio, IMUs, gaze-related signals, and other calibrated sensor streams with metadata.
  • 3.2 Domains and activities: The collection covers eight physical and procedural domains, including soccer, basketball, dance, music, cooking, bike repair, and health care, in authentic varied sites.Cooking alone includes more than 650 takes by over 170 chefs across 60 environments.

Appendix B describes the data collection

Data collection emphasized expert participants, diverse recruitment and demographics, authentic multi-site scenarios, and participant-provided information for skill, actions, and objects. Privacy review and informed consent governed collection.

  • 740 participants with credentials, training, or expertise performed the recorded skills, while varied skill levels supported proficiency estimation.
  • Demographics: Participants ranged from 18 to 74 years old, included varied gender identities, and reported more than 24 ethnicities.Ethnicity reporting was optional and was not collected at several listed Pennsylvania, California, and New York sites.
  • Collection scope: Ego and paired exo views span eight domains and 123 scenes, with odd and even figure columns representing paired first- and third-person views.
  • Participants: Recruitment used locally selected channels including campus lists, flyers, referrals, online advertisements, social media, agencies, schools, gyms, and athletic organizations.
  • Participant annotations: Pre-task and post-task surveys captured perceived skill, performance, mistakes, timing, and experience, while narrate-and-act recordings captured actions and objects.
  • Privacy and ethics: Collection followed independent institutional review, informed consent, and Project Aria responsible-research guidelines.

4 Natural language descriptions

Ego-Exo4D provides three time-indexed paired language resources spanning expert critique, first-person narration, and atomic descriptions. Together, they capture what skilled participants do, why and how they do it, and how experts assess execution.

  • Three paired language datasets provide expert commentary, narrate-and-act descriptions, and atomic action descriptions aligned with video.
  • The language resources are intended as general-purpose data for action and object grounding, multimodal learning, skill assessment, and video-language models.
  • Expert commentary captures strengths, weaknesses, execution quality, and skill ratings through speech, spatial drawings, and numeric proficiency scores.Fifty-two experts critique participant performance and explain how hand/body pose and object use affect execution.
  • Expert commentaries focus on how activities are executed, revealing subtle differences in skilled performance that may be invisible to non-experts.
  • The 52 experts were selected for technical expertise, communication ability, and live commentating performance, with coaching or teaching experience common among them.On average, 90% had more than 10 years of professional experience, and all had served as coaches, instructors, or mentors.

Expert commentary guidelines

Expert commentary is collected as retrospective, time-anchored critique across synchronized ego and exo views. The annotation suite also distinguishes first-person explanations from atomic third-party descriptions and records visibility information for multiview learning.

  • Expert commentary guidelines: Experts first review synchronized ego and exo videos fully, then pause to record comments focused on critique, teaching advice, execution quality, and mistakes.They typically make about 7 comments per minute, and commentary is retrospective and time-anchored.
  • Expert commentary guidelines: Experts can augment spoken commentary with telestrator sketches and assign an overall proficiency score from 1 to 10 with written justification.
  • Expert commentary guidelines: Each video receives commentary from 2-5 distinct experts, yielding 117,812 time-stamped, video-aligned comments from more than 6,000 hours of expert work.
  • Other language annotations: Narrate-and-act descriptions are first-person tutorial-style explanations, whereas atomic descriptions are short third-party statements about individual visible actions.Narrate-and-act data covers about 10% of takes, and participants may perform more slowly while explaining their actions.
  • Other language annotations: Atomic narrations are designed around one verb and one timestamp, while annotators also identify whether actions are visible in ego and which exo camera offers the best view.Visibility tags support exocentric view selection for correspondence and expert commentary, and frame selection for hand and body pose.
  • Language statistics: Expert commentary uses the largest vocabulary and longest statements, while atomic descriptions have the greatest temporal density and narrate-and-act falls between them.

5 Ego-Exo4D benchmark tasks

Ego-Exo4D defines benchmark tasks for learning across ego and exo viewpoints, organized into relation, recognition, proficiency, and ego-pose families. The suite provides multimodal data, annotations, evaluation protocols, and baselines.

  • The benchmark suite comprises four task families: relation, recognition, proficiency, and ego-pose.
  • Each task includes multimodal data, high-quality annotations, evaluation support, and baseline models intended to enable comparable research.
  • Relation tasks: Ego-exo relation tasks address object-level correspondence and synthesis of one viewpoint from the other under extreme viewpoint changes.
  • Correspondence: Correspondence predicts the same object’s mask across synchronized ego and exo frames, without semantic labels, camera pose, IMU, or active range-sensor inputs.The benchmark covers selected objects from six scenarios and excludes Bouldering and Dance because of limited object diversity.
  • Correspondence: Correspondence baselines compare spatial prediction at each time point with spatio-temporal prediction that uses correspondence history.

Results

Temporal modeling substantially improves ego-exo correspondence, but cross-view mask prediction remains difficult, especially from ego queries to exo views. Translation is framed as separate track and clip-generation problems with restricted inputs to preserve applicability to arbitrary third-person video.

  • Correspondence results: 13.88% to 22.14% IoU: spatio-temporal baselines substantially outperform spatial baselines in the Ego→Exo correspondence setting.
  • Correspondence results: Correspondence models perform worse when querying ego masks and predicting exo masks, consistent with exocentric object occlusion and small object size.
  • Correspondence results: Below 23% IoU in Ego→Exo and below 24% IoU in Exo→Ego: all baselines show the task remains challenging.The dataset includes substantial object-shape variation and many very small objects.
  • Correspondence results: Spatio-temporal XView-XMem reliably tracks a single object across a sequence where spatial XSegTx alternates between one and two predicted masks.
  • Translation results: Ego-exo translation separates ego-track prediction from ego-clip generation, with variants that optionally provide relative ego-camera pose at inference.
  • Translation results: Translation inputs are restricted primarily to exo views and object masks, excluding depth, point clouds, IMU, and SLAM except for an ego-pose variant.The restriction is intended to support translation from arbitrary third-person video, where such signals are typically unavailable.

Results

Ego-Exo4D supports cross-view translation and fine-grained keystep recognition using paired ego-exo video, annotations, and view-invariant training. Translation remains challenging, while recognition is formulated around subtle, temporally variable procedural steps.

  • Ego-exo translation: DiT-pix outperforms pix2pix-pix across all reported Ego Clip Generation metrics.Its generations are usually photorealistic and close to ground truth, although some textures deviate despite accurate object shapes.
  • Ego-exo translation: Removing the exo object crop disrupts target color and texture, while removing the ego crop mask causes incorrect object orientation.These ablations show that both cropped inputs contribute distinct information to translation.
  • Ego-exo translation: Multi-frame prediction offers no quantitative advantage over frame-to-frame prediction but often produces more temporally consistent generations under heavy exo-view occlusion.Using multiple frames can provide information missing from individual exocentric frames.
  • Keystep recognition: Fine-grained keystep recognition distinguishes subtle hand-object interactions across steps with different temporal spans, using ego video at inference.The task can exploit exocentric context during training while testing with only the ego view.
  • Keystep recognition: View-invariant learning pretrains on synchronized ego-exo pairs with clip-level contrastive loss before classification training.The approach uses paired views to learn representations intended for ego-view keystep classification.

Results

The energy-efficient recognition benchmark treats online keystep prediction as a joint sensing and inference problem under explicit power budgets. It models compute, memory transfer, and sensor costs while evaluating accuracy-efficiency trade-offs.

  • Task and energy model: Energy-efficient recognition maximizes full-video keystep performance while keeping sensing and inference within a specified energy budget.A sensor-triggering policy selects active modalities, and a predictor estimates the current keystep from the resulting observation.
  • Task and energy model: Total energy accounts for model forward passes, memory transfers, and sensor operation, with modality-specific costs.Camera sensors are described as more expensive to operate than audio and IMU sensors.
  • Motivation: The benchmark is designed around real-world on-device power consumption rather than isolated FLOPs, parameter counts, or throughput.It targets constraints relevant to devices such as mobile phones and AR glasses.
  • Task and energy model: The task uses RGB video with audio and potentially other sensors, while online inference may use only current and past observations.The same egocentric videos and keystep annotations support this benchmark.
  • Evaluation: 20 mW and 2.8W define the high-efficiency and high-performance evaluation tiers, respectively.Recognition is measured with per-frame calibrated mean average precision and energy in mW.

Results

Energy-efficient recognition exhibits modality- and policy-dependent trade-offs, while procedure understanding models keystep dependencies, optionality, mistakes, and missing steps. Task graphs provide the procedural structure, but training and test annotations remain restricted.

  • Energy-efficient recognition: Combining vision and audio improves high-efficiency-tier recognition over either modality alone, while vision-only models generally outperform audio-only models.Raw backbones usually perform better than sampling-policy models but consume more energy.
  • Energy-efficient recognition: High-performance-tier audio-visual models generally underperform vision-only models, possibly because late fusion reduces the expressivity of strong visual features.The reported pattern differs from the high-efficiency tier.
  • Energy-efficient recognition: Audio-only models show affinity for sounding actions such as stirring egg mixture and cutting butter.Performance varies across keystep labels and modalities.
  • Procedure understanding: Procedure understanding infers keystep orderings, optional steps, procedural mistakes, and missing keysteps from segment histories.The task is evaluated under instance-level or procedure-level weak supervision with causal processing.
  • Procedure understanding: Task graphs represent keysteps as nodes and dependencies as directed edges, including constraints for multiple prerequisites and optional or repeatable steps.For example, all required dependencies must be satisfied before a dependent keystep can occur.
  • Annotations and evaluation: Training keystep annotations and labeled task graphs are withheld, validation annotations are limited, and test evaluation requires server submissions.The restrictions reflect the weakly supervised benchmark setup and avoid label leakage.

Results

Procedure understanding baselines show that graph structure and keystep co-occurrences provide useful information, while future-keystep prediction remains difficult. Proficiency benchmarks define both person-level skill classification and timestamped localization of good or improvable executions.

  • Procedure understanding: Graph-based baselines significantly outperform a uniform predictor on most procedure-understanding tasks, except future-keystep prediction.The result indicates that simple keystep co-occurrences capture part of a procedure’s overall structure.
  • Procedure understanding: End-to-end models trade accuracy for test-time efficiency because they lack an explicit procedure graph.Procedure-level baselines perform worse without ground-truth labels, while keystep prediction outperforms keystep assignment in the reported comparisons.
  • Proficiency estimation: Proficiency estimation includes four-class demonstrator skill classification and temporal localization of good executions or areas needing improvement.Parts of a video that do not reveal proficiency remain unlabeled in the demonstration-level task.
  • Proficiency estimation: The benchmarks use top-1 accuracy for demonstrator proficiency and L1-distance-based mAP for timestamp localization.Localization mAP is defined using thresholds on prediction-to-ground-truth timestamp distance in seconds rather than temporal IoU.

Results

Proficiency experiments show that learned video models outperform naïve baselines, with ego video often sufficient and exo views useful for some activities. The pose benchmarks target 3D body and hand recovery from first-person video or IMU inputs, supported by extensive annotations.

  • Proficiency estimation: TimeSFormer trained from random initialization significantly outperforms random and majority-class baselines for demonstrator proficiency estimation.The models quantify skill levels from video using top-1 classification accuracy.
  • Proficiency estimation: Ego video performs well in most proficiency cases, while exo video benefits tasks such as bouldering; simple late fusion does not improve performance.Pre-trained initialization particularly improves ego-video results.
  • Proficiency estimation: ActionFormer outperforms naïve baselines for demonstration proficiency localization, but absolute mAP scores remain fairly low.The task predicts timestamped good executions and improvement needs from ego, exo, or combined views.
  • Ego pose: The ego-pose benchmark separates body-pose and hand-pose estimation from egocentric video or IMU data.Body pose predicts 17 three-dimensional MS COCO joints over time, while hand pose predicts 21 3D joints per partially visible hand.
  • Ego pose: Ego-Exo4D provides approximately 14M combined body and hand 3D ground-truth and pseudo-ground-truth frames, including large manually annotated sets.The dataset reports 376K body and 68K hand manually annotated 3D poses, alongside automatically generated annotations.

Results

Body- and hand-pose baselines improve over static or prior approaches, while joint visibility affects estimation difficulty. The dataset combines realistic multimodal collection, broad skilled-activity coverage, benchmark resources, and explicit limitations from camera placement, synchronization, and long-tailed activity duration.

  • Pose results: The static body-pose baseline has substantially higher MPJPE than other approaches, indicating that poses vary widely across scenarios.The authors conclude that one fixed pose is infeasible for the dataset’s diverse test cases.
  • Pose results: Thumbs and fingertips have larger hand-pose errors, likely because they are more frequently occluded or invisible.The analysis reports MPJPE and PA-MPJPE alongside error distributions across hand joints.
  • Pose results: PA-MPJPE decreases as joints become visible in more views, linking multiview visibility with lower pose-estimation difficulty.Visibility also serves as an indicator of ground-truth uncertainty because fewer observations often reflect entanglement with objects or other hand parts.
  • Dataset and resources: Ego-Exo4D offers synchronized ego and exo recordings across physical and procedural skills, with benchmarks, models, scripts, visualization, and baselines.Its collection pipeline captures activities in natural indoor and outdoor settings rather than mocap suits or laboratories.
  • Limitations: Camera occlusion, imperfect synchronization, long-tailed durations, and unequal skill difficulty constrain dataset coverage and model training.Static exo-camera placement can obstruct actions, while cooking contributes nine times more hours than soccer.

A.4 Aria post-processing

Ego-Exo4D’s post-processing combines Aria and GoPro localization, synchronization, and coordinated multimodal data collection across diverse skilled-activity settings.

  • A.4 Aria post-processing: 783 Aria recordings containing 5,035 takes were processed through MPS before GoPro localization, time synchronization, and take separation.95.9% had successful Aria localization throughout, while 3.5% had partial tracking failures and 0.6% failed completely.
  • A.4 Aria post-processing: 91.4% of 3,724 attempted GoPro recordings were successfully localized on a recording level.Recording-level localization helps with short physical-activity takes, while textureless views are the dominant failure case.
  • Data collection: Twelve research labs coordinated common guidelines, scenarios, and camera rigs while contributing site-specific participants, locations, and modalities.The coordinated process aimed to make the dataset cohesive while retaining geographic and environmental diversity.
  • Data collection: Collections covered skilled activities including soccer, bike repair, cooking, music, and other procedural or physical domains in natural settings.Examples include professional soccer drills, mechanics working in their own shops, and chefs recording in professional kitchens.
  • Camera configuration: Additional hand-focused cameras were found crucial for guitar, violin, and cooking because they improve coverage of hand-object interactions with less self-occlusion.The head-mounted camera follows body motion while focusing on the hand/body region.
  • Language resources: Ego-Exo4D provides parallel expert commentary, narrate-and-act descriptions, and atomic action descriptions as time-indexed video-language resources.The annotation types differ in perspective and granularity, supporting both action description and skilled-performance analysis.

E Benchmarks: annotations and baselines

The benchmarks use unified, leakage-resistant splits and provide annotations and baselines for ego-exo object correspondence across time-synchronized views.

  • Dataset splits: Take-level splits assign each take to training, validation, or testing, with derivatives inheriting the original assignment.Splits are stratified by activity and proficiency, and participants remain exclusively within one split to prevent leakage.
  • Relation annotations: A multistage relation-annotation pipeline enumerates active objects, segments them throughout egocentric videos, and transfers the task to visible exocentric frames.Annotators receive object descriptions, sample ego masks, and synchronized ego-exo videos when marking exocentric masks.
  • Relation annotations: Mask coverage is exhaustive for few-object scenarios but sampled by frequency and size for cooking, health, and bike repair.Objects too small for reliable matching or visible in fewer than 10 ego frames are excluded from exocentric mask annotation.
  • Relation annotations: 5,566 objects across 1,335 ego-exo video pairs yielded 2.2M segmentation masks, including 742K ego and 1.1M exo paired masks.The process also produced approximately 4M annotated frames and 367K ego-only masks.
  • Baselines: The spatial baseline predicts correspondence from egocentric and exocentric frames plus a query mask using cross-image features and mask losses.Training uses binary cross-entropy and Dice losses, with visibility classification trained across sequence frames.
  • Baselines: The spatio-temporal XView-XMem baseline tracks objects through interleaved egocentric and exocentric frames and stores fused features in working memory.XSegTx embeddings are added to XMem memory to help mitigate track drift within and across views.

Results

Results vary substantially by activity and object scale: basketball and soccer are easier, while cooking and bike repair remain more challenging because their objects are more diverse.

  • Correspondence results: Basketball and soccer are generally easier to model than cooking and bike repair in the correspondence benchmarks.The easier scenarios have less variation in object shape and appearance, whereas cooking and bike repair contain more diverse objects.
  • Correspondence results: All baselines struggle on very small target-view objects.The evaluation explicitly stratifies validation examples by predicted object size measured as image-pixel proportion.
  • Translation results: Ego-exo translation shows the same scenario trend for ego track prediction and ego clip generation.Both subtasks perform better in basketball and soccer than in bike and cooking scenarios.
  • Keystep annotations: Keystep annotations mark procedural actions with timestamps, category labels, descriptions, and essential-or-optional status.The interface presents synchronized ego and exo videos for annotation.
  • Keystep taxonomy: The keystep taxonomy is developed iteratively because unscripted activities cannot be fully specified before annotation.Annotators add missing actions, which are reviewed and incorporated across three taxonomy-development iterations.

Dataset splits

The dataset’s keystep-recognition split contains 130,979 segments, while benchmark analyses compare views, modalities, and pose-estimation approaches across scenarios.

  • Keystep recognition: 278 unique keysteps were retained after applying a 20-sample cutoff to the long-tailed keystep distribution.The analysis considers only leaf-node keysteps.
  • Keystep recognition: 130,979 keystep segments average 11.34 seconds, with 74,342 training, 23,636 validation, and 33,001 test segments.The training, validation, and test sets contain 14,326, 4,517, and 6,373 ego-view segments, respectively.
  • View-specific performance: Exocentric views outperform egocentric views on several keysteps, while egocentric views are more effective for manipulating small objects.Exocentric views benefit most on “have a conversation asking different questions,” whereas ego views help on “cut carrots” and “unpack the new tube.”
  • Modality-specific performance: Vision-only models improve most on steps without distinctive sounds, whereas audio-only models help most when task sounds are strongly indicative.Examples include adding green chillies or getting celeries for vision, and stir-frying egg mixture or cutting butter for audio.
  • Proficiency estimation: Scenario-specific proficiency estimation favors ego views in cooking and exo views in bouldering, but usually fails to beat the majority-class baseline on test splits.The reported pattern highlights a distribution shift between validation and test splits.
  • Pose estimation: Three of four hand-pose challenge submissions outperform the baseline, with the best participant improving MPJPE by 11.85% and PA-MPJPE by 23.31%.The challenge compares submissions against POTTER.
Loading 2311.18259v4…