Source-linked AI summary
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina Gonzalez, James Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jachym Kolar, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Ziwei Zhao, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fuegen, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
TL;DR
Existing datasets and models capture only a limited definition of visual perception. Ego4D addresses this gap with a large egocentric dataset and benchmark suite spanning daily-life activity, and makes it available for multimodal perception research.
Problem
Current datasets and models represent only a limited definition of visual perception.
Method
Ego4D combines a massive publicly available egocentric dataset with benchmarks for querying past video, recognizing present object-state changes, and forecasting activities.
Results
Ego4D provides a first-of-its-kind multimodal egocentric dataset and benchmark suite with 931 camera wearers and wide scenario coverage.
Takeaways & Limitations
The dataset and benchmarks provide footing for video-understanding research relevant to augmented reality, robotics, and other domains.
Takeaways & Limitations
Coverage remains incomplete: participants are generally from urban or college-town areas, with pandemic-related and battery-life biases in recorded activities.
Abstract
from arXiv · showhide
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/
1. Introduction
Ego4D addresses the limited definition of visual perception in existing datasets by introducing a massive, diverse egocentric video dataset and benchmark suite for first-person perception.
- Existing Internet datasets capture brief, isolated moments from a third-person view, unlike the long, fluid first-person streams needed in robotics and augmented reality.
- First-person perception requires persistent 3D scene understanding and interpretation of human-object interactions and high-level social behaviors.
- Ego4D introduces a dataset and benchmark suite centered on egocentric visual perception, combining 3D spatial and temporal information.
- 3,670 hours of video from 931 participants across 74 locations and 9 countries provide large-scale, geographically diverse daily-life activity footage.
- The dataset includes long-form interactions across home, workplace, leisure, social, and commuting settings, with portions providing audio, 3D meshes, gaze, stereo, and synchronized multi-camera views.
- Five benchmark tasks span indexing past experiences, analyzing present interactions, and anticipating future activity, supported by millions of annotations from over 250,000 hours of annotator effort.
2. Related Work
Prior large-scale vision datasets mainly use third-person web data, while existing egocentric datasets are smaller and often narrower in demographic, geographic, or task coverage. Ego4D expands egocentric video scale, diversity, and benchmark scope.
- Collections such as Kinetics, AVA, UCF, ActivityNet, HowTo100M, ImageNet, and COCO focus on third-person Web data shaped by human photographers.
- Egocentric video poses challenges including human-object interaction, activity recognition, anticipation, summarization, hand detection, social interaction parsing, and body-pose inference.
- Existing egocentric datasets include unscripted daily-life collections, but other datasets are semi-scripted or use narrower participant populations and settings.
- Ego4D contains 3,670 hours from 931 camera wearers, compared with 100 hours and 71 people in cited prior datasets, while spanning 74 locations and 9 countries.
- Ego4D provides millions of annotations supporting multiple complex tasks, representing a step change in dataset scale and diversity.
3. Ego4D Dataset
Ego4D was collected as diverse, largely unscripted daily-life video from worldwide participants and enriched with multiple modalities, annotations, and privacy safeguards. Its scope remains bounded by geographic, demographic, pandemic, recording-duration, language, and resource-access biases.
- Collection strategy: The collection sought diversity of people, places, objects, and activities through unscripted footage recorded for long periods.
- Collection strategy: 931 participants across 74 cities contributed 3,670 hours of video through distributed teams spanning 9 countries and 5 continents.
- Camera wearers: Participants covered varied occupations and ages, including 96 people over 50, with 45% identifying as female.
- Scenarios: Participants freely selected activities and wore cameras for extended periods, producing typical raw clips lasting 8 minutes rather than 10 seconds.
- Privacy and ethics: Collection followed informed-consent, privacy, institutional-review, sensitive-area avoidance, and de-identification requirements, with some social footage allowing unblurred faces and audio.
- Scenarios: The dataset spans diverse scenarios and activities, with the 931 camera wearers providing a glimpse of daily life around the world.
- Cameras and modalities: Ego4D combines video with 3D scans, audio, gaze, stereo, synchronized multi-camera streams, and textual narrations for multimodal research.
- Possible sources of bias: Dataset biases include incomplete global coverage, urban and college-town concentration, pandemic-related scenario imbalance, active-day sampling, and localized narration vocabulary.
4. Narrations of Camera Wearer Activity
Ego4D uses dense, independently recorded narrations to describe camera-wearer activity and support dataset organization and annotation. These narrations also form a large aligned language-video resource.
- 13.2 sentences per minute produced 3.85M temporally dense narration sentences across the video.Each five-minute clip receives two independent narrations, including summaries and timestamped descriptions of camera-wearer actions.
- Narrations support taxonomy construction for actions and objects through data-driven text mining.
- Narrations map videos to relevant benchmarks and identify temporal windows for seeding additional annotations.
- The narration collection is itself a potentially valuable repository of weakly aligned natural language and video.
5. Ego4D Benchmark Suite
The Ego4D benchmark suite targets fundamental first-person perception challenges across past, present, and future activity. Its tasks cover episodic memory, hand-object state changes, and multimodal social interaction, with task-specific annotations and metrics.
- Benchmark suite overview: Five benchmarks span indexing past experiences, analyzing present interactions, and anticipating future activity.The first release provides annotations for 48–1,000 hours per benchmark on top of 3,670 narrated hours.
- Episodic Memory: Episodic Memory localizes answers in past video for natural-language, visual, and moment queries.The task couples temporal localization with spatial reasoning, including object location and activity occurrence.
- Episodic Memory: Episodic Memory annotations include 13 language-query templates, a 110-activity taxonomy, and approximately 74K queries across 800 hours.
- Hands and Objects: Hands and Objects models object state changes across temporal, spatial, and semantic axes, including point-of-no-return localization.Annotations mark pre, PNR, and post frames with boxes for hands, tools, and objects, plus state-change types, verbs, and nouns.
- Hands and Objects: Hands and Objects is presented as the first video benchmark dedicated to understanding object state changes.Its design seeks generalization across different object-tool-action combinations rather than simple overfitting to actions or objects.
- Audio-Visual and Social Interaction: The benchmark suite adds audio-visual and egocentric social-interaction challenges to account for multimodal present activity.Audio-Visual Diarization is composed of four tasks, while egocentric perception also requires long-form video, attention cues, and person-object interactions.
6. Conclusion
Ego4D is introduced as a large-scale dataset and benchmark suite for multimodal egocentric video perception. Its data and benchmarks are intended to support innovations relevant to video understanding, augmented reality, robotics, and related domains.
- Ego4D is a first-of-its-kind dataset and benchmark suite for multimodal perception of egocentric video.
- The dataset enables learning from daily-life experiences around the world through visual and auditory egocentric data.
- The benchmark suite provides footing for video-understanding innovations relevant to augmented reality, robotics, and other domains.
Contribution statement
The project was led and initiated by Kristen Grauman, with program management and operations led by Andrew Westbury and scientific advising by Jitendra Malik. Video collection and annotation involved Facebook Reality Labs, university partners, and Facebook AI.
- Kristen Grauman led and initiated the project, while Andrew Westbury led program management and operations.
- Jitendra Malik provided scientific advising, and starred authors drove implementation, collection, or annotation development.
- Facebook Reality Labs collected some consented video in a closed Facebook-building environment, while university partners managed other collection and recruitment.
- Facebook AI led the annotation effort.
A. Data Collection
Ego4D used distributed, naturalistic collection protocols across multiple sites, with participants recording diverse daily activities under consent and review procedures. The supplied site records illustrate broad geographic, demographic, and scenario coverage.
- 660.5 hours were recorded from 138 subjects across five Indian states and 36 scenarios, including brickmaking, knitting, egg-carton making, and hairstyling.
- The dataset includes focused Japanese cooking and handcraft recordings from 81 participants, totaling 141 hours.
- Participants were encouraged to record activities they naturally performed, including paired scenarios, while signing consent forms before cameras were posted.
- Collection procedures included participant training, recording windows, footage upload, and researcher review for lighting, setting, and viewpoint.
- 262 hours were recorded by 82 UK participants, who averaged 3.0 hours each (σ = 0.7 hours) across varied commuting, entertainment, work, sports, and household activities.
- Site protocols included screening and censoring identifying information, with participants sometimes able to request additional censoring before sharing.
B.2 Sample pipeline
The sample CMU pipeline combines automated sensitive-object detection with sequential human review, correction, and final image blurring. Its labor cost is substantial relative to the original video duration.
- The four stages are automatic face and license plate detection, false positive removal, negative detection handling, and image blurring.
- Reviewers first identify videos containing sensitive objects such as faces, license plates, and credit cards before software detects them automatically.
- Reviewers manually reject bounding boxes lacking sensitive information and annotate missed detections, using online tracking to propose boxes across frames.
- After detections are modified and corrected, bounding-box regions are de-identified through robust image blurring.
- 780 hours of manual labor were required to de-identify roughly 500 hours of video, with work distributed across multiple review and correction steps.
C. Demographics
The supplied passages describe optional demographic reporting alongside dense, time-aligned narration collection for activities and object interactions. Narrations use independent annotators and explicit tags for the camera wearer and other participants.
- C. Demographics: Demographic information was reported separately by state or country because ethnic-group categories differed in granularity, and it was unavailable for Minnesota participants.
- C. Demographics: UK demographic reporting was optional, with 63% of participants self-reporting ethnic-group membership.
- D.1 Narration instructions and content: Narrations were collected on clips of up to 5 minutes, with two independent annotators providing both clip summaries and dense action descriptions.
- D.1 Narration instructions and content: Each narration marks a timepoint and describes an atomic action or object interaction, including interactions between the camera wearer and other tagged people.
D.2 Narration analysis
Ego4D’s narration analysis quantifies dense textual coverage and lexical diversity, then converts narration language into compact verb and noun taxonomies. These resources also support targeted selection of benchmark annotations.
- 3.85M sentences were collected across the 3,670 hours of video, with narration frequency varying substantially by depicted activity.
- Scenario word distributions highlight both characteristic objects, such as bowls in cooking and cards in board games, and objects common across scenarios.
- The raw narrations contain 1,772 unique verbs and 4,336 unique nouns, reflecting broad lexical coverage of the videos.
- The taxonomy construction parses verbs and nouns, separates verb senses, decomposes compound nouns, maps collective expressions, and manually clusters redundancies.
- The resulting taxonomy contains 115 verbs and 478 nouns.
- Narrations and summaries are used to automatically target benchmark-relevant videos, such as conversation clips for AV Diarization and hand-object clips for Hands and Objects.
E. Benchmark Data Splits
The benchmark suite defines several episodic-memory query types and aligns data splits within related task families. Its tasks retrieve temporally localized evidence for objects, actions, and natural-language questions from egocentric video.
- Data coverage: 764 hours are relevant to the AVD and Social tasks, while memory queries can apply across all 3,670 hours.The annotated subset for AVD and Social totals 47.7 hours; other benchmarks depend more flexibly on video content.
- Data splits: Forecasting and Hands+Objects share splits, preventing videos used for training in one task from appearing in validation for the other.Episodic Memory tasks likewise share splits, but consistency across very different task families is harder because their annotated videos differ.
- Episodic Memory tasks: Episodic Memory includes visual, natural-language, and moments queries, each requiring localization of the response in video.The tasks retrieve an object’s most recent occurrence, answerable temporal evidence for a language query, or all instances of an action category.
- Episodic Memory tasks: Visual queries identify the last occurrence of an object from a static image crop and can additionally report its 3D displacement when an environment scan exists.The response is a temporally contiguous track of bounding boxes; repeated appearances are resolved by selecting the most recent one before the query frame.
- Episodic Memory tasks: Natural-language queries return a contiguous response track sufficient to answer an episodic question without external knowledge.The task requires multimodal reasoning over objects, activities, and events in the video.
- Episodic Memory tasks: Moments queries retrieve all temporal instances of a predefined action category, allowing multiple response windows for one query.Unlike natural-language grounding, the query uses action categories rather than sentences and can correspond to multiple instances.
F.4 Data Analysis
The analysis characterizes Ego4D’s query annotations, baseline performance, and remaining challenges across visual, natural-language, and moment localization tasks. Results show strong task difficulty, especially for efficient object search, camera-pose availability, and short-moment detection.
- Visual queries: 433 hours contain 22,602 visual queries spanning 10 universities and 54 scenarios, with disjoint video splits.
- Visual queries: Visual-query bounding boxes are largely position-unbiased, while query-response distances and response-track sizes show potential distributional bias.
- Natural language queries: 19.2K natural-language queries over 227 video hours form a short-window “needle in the haystack” retrieval problem.
- Visual queries 3D localization: Camera pose estimates are available for only 15% of queries, limiting VQ3D success despite qualitatively good pose estimates under scene changes.
- Moment queries: Moment-query errors arise from both localization and classification, with short moments performing worse and motivating improved classifiers and short-moment handling.
- Visual queries: 42.9% success, 0.13 tAP, and 0.06 stAP define the visual-query baseline, while improving search efficiency causes drastic performance reductions.
G.6 Data Analysis
The analysis characterizes Ego4D’s annotations, interaction diversity, baseline performance, and limitations across object-state-change and audio-visual tasks.
- Critical frames: PNR frames cluster near the center of the 8-second snippets, while pre and post frames usually occur shortly before and after them.The PNR distribution aligns closely with the 4-second narration point, and the nearby frames reflect rapid state changes.
- Hands and objects: The dataset contains ∼825K hand, object, and tool bounding boxes distributed across varied sizes and image locations.This includes ∼245K left-hand, ∼260K right-hand, ∼280K object, and ∼40K tool boxes.
- Actions: Annotated actions emphasize low-level manipulation verbs and nouns, with common actions and objects forming a natural long tail.Most object categories are uncommon in standard detection datasets: 442 of 478 extend beyond COCO’s 80 categories.
- Baseline results: Learnable state-change classifiers exceed 60% accuracy, outperforming the close-to-50% always-positive baseline, with Bidirectional LSTM performing best.The remaining difficulty reflects substantial variation in object types and state changes.
- Baseline results: PNR localization improves from around 1.1 seconds for center-frame prediction to 0.425 seconds on validation and 0.489 seconds on test with SlowFast + Perceiver.These results indicate that baseline models learn meaningful temporal cues, although some changes produce little visible appearance change.
- Baseline results: Single-frame state-change object detection remains difficult, with all baselines achieving only 8–14% AP.Object-size variation and the absence of cross-frame appearance information constrain single-frame detection.
- Discussion: The benchmark defines complementary when, where, and what aspects of hand-induced object changes, while encouraging future methods to model their dependencies jointly.The discussion connects semantic transformations such as splitting with PNR localization and post-change box structure.
- Audio-visual data: The audio-visual benchmark targets speaker localization, voice activity, and speech content from the egocentric viewpoint, while Ego4D adds casual, mobile, multi-speaker settings missing from several prior datasets.Compared with existing datasets, Ego4D emphasizes first-person video and in-the-wild daily-life conversations.
H.6 Baseline Modeling Framework
The baseline framework addresses the four linked Audio-Visual Diarization tasks sequentially, combining trajectory grouping, active-speaker classification, audio association, wearer voice detection, and transcription. Results show that in-the-wild egocentric diarization and transcription remain difficult, especially with overlapping speech and noisy cross-modal information.
- Framework: The framework treats the four linked tasks sequentially, using representations from one task to support the others.The pipeline proceeds through localization and tracking, active-speaker classification, audio association, wearer voice detection, and ASR.
- Framework: A greedy trajectory-grouping procedure trades global optimality guarantees for lower computational complexity and strong empirical results.The method can produce an overall cost of 11 when the optimal solution costs 7, although accurate person embeddings may reduce such corner cases.
- Framework: Audio-visual baselines combine cropped face video and corresponding audio to classify whether each tracked person is speaking.Using multiple cropped mouth images did not significantly improve active-speaker classification, partly because camera motion and difficult face angles produce inaccurate crops.
- Results: Voice activity detection substantially improves active-speaker results, while imperfect active-speaker predictions can contaminate speaker signature libraries.The reported baselines compare AVA-pretrained models with models trained only on Ego4D data, and video-only approaches can be combined with VAD to remove false alarms.
- Results: Joint audio-visual diarization and transcription remain challenging, with baseline DER above 80% and WER above 60% amid overlapping speech and in-the-wild variation.The discussion identifies overlapping speakers, interruptions, noise, accents, vocabulary differences, and head-motion blur as unresolved challenges for modeling and annotation.
I.3 Social Baseline Models and Results
The Social benchmarks evaluate first-person attention and speaking behaviors using multimodal annotations, synchronized recordings, and baseline models. Baselines show that LAM is substantially stronger than TTM, while both tasks leave room for improvement.
- LAM: LAM uses cropped face sequences processed by ResNet-18 and a Bidirectional LSTM to predict whether a person looks at the camera wearer.The model predicts the binary label for the center frame and addresses class imbalance with weighted cross-entropy.
- LAM: 78.07% mAP is achieved by LAM with Gaze360 initialization, compared with 66.07% mAP when initialized randomly.The results indicate a close relationship between LAM and gaze estimation.
- TTM: TTM combines face-crop video features with MFCC-based audio features to determine whether an utterance is directed toward visible faces.Audio segments are truncated to at most 1.5 seconds, and visual and audio embeddings are fused before classification.
- TTM: TTM improves mAP by only 9.77% over random guessing, making it more challenging than LAM.The task requires both audio-content analysis and fusion of audio and video cues.
- Data and future tasks: The Social benchmark includes naturalistic multimodal recordings from five sites, while its annotations support extensions such as social gaze and utterance target prediction.The dataset contains 764 hours of video and audio, along with synchronized videos from multiple social members and gaze measurements for subsets of participants.
J. Forecasting Benchmark
The Forecasting benchmark covers four future-prediction tasks, with annotations and baseline evaluations spanning locomotion, hand movement, object interactions, and long-term action sequences. Results show useful but limited anticipation, with performance affected by prediction horizon, visual complexity, and dataset annotation challenges.
- Task definitions: Four tasks predict future locomotion trajectories, hand positions, short-term object interactions, and long-term action sequences.Long-term anticipation targets sequences of future actions for ordered, long-horizon planning.
- Long-term action anticipation: SlowFast-to-MViT changes improve verb forecasting but hurt noun forecasting, while multiple-clip aggregation gives the best reported performance.Explicitly training multiple heads also improves verb, noun, and action prediction relative to the No Change baseline.
- Annotation analysis: Forecasting annotations were difficult because diverse taxonomies, “stuff” bounding boxes, and large-tool interactions created ambiguity.Annotators often used OTHER labels, struggled to bound objects such as grass, and faced uncertainty about which object constituted the interaction target.
- Locomotion prediction: Global visual features provide useful locomotion cues but do not model complex movement such as avoiding pedestrians.The result motivates finer-grained visual representations for future movement prediction.
- Short-term object interaction anticipation: 2.07% validation and 2.45% test Top-5 mAP were achieved for short-term object interaction anticipation.The baseline substantially outperformed random prediction, but complex scenarios and missed active objects remained limiting factors.
- Long-term action anticipation: More temporal context improves transformer-based anticipation, while longer prediction horizons become progressively harder and eventually plateau.Using more context increases memory consumption, and immediate next actions are easier to predict than actions farther ahead.
K. Societal Impact
Ego4D pairs a large-scale egocentric video resource with privacy and ethics safeguards, while acknowledging risks of misuse, privacy breaches, careless future collections, and incomplete demographic coverage.
- Ego4D offers a large-scale video resource with rigorous privacy and ethics standards, diverse subjects, and benchmarks intended to support reproducible technical advances.The authors identify potential applications including assistive technology, education, fitness, entertainment, gaming, and eldercare.
- Ego4D could be misused, creating potential negative societal impact as egocentric vision technology develops.The authors explicitly call for future research to guard against misuse.
- Wearable-camera deployments create privacy risks requiring notice, consent, and user controls over data use, storage, and sharing.Models transcribing speech or performing related tasks should also include robust user controls.
- Audio-visual and social benchmarks use fully consented participant data, including unblurred faces and conversation audio, addressing a previously unmet need for large-scale study.The resource is designed to respect privacy protocols across different countries.
- Large-scale data releases may encourage future collections that lack comparable privacy and ethical care, so the authors document procedures and plan to disseminate best-practice recommendations.The mitigation strategy is intended to reduce the risk of less careful follow-on efforts.
- Dataset coverage remains imbalanced: Rwanda has relatively little data, 74 cities do not represent all demographics, and complete global coverage is elusive.The authors propose expanding collaborations with researchers and participants in underrepresented areas as a mitigation.