← 论文海报合集← Paper Notes|
CVPR 2022 · 大规模数据集CVPR 2022 · Large-Scale Dataset

Ego4D: Around the World in 3,000 Hours of Egocentric Video

史上最大规模第一视角视频数据集与基准套件,覆盖全球 74 地、9 国、931 名拍摄者The largest egocentric video dataset and benchmark suite to date, spanning 74 locations, 9 countries and 931 camera wearers worldwide
Kristen Grauman, Andrew Westbury, et al. (85 位共同作者) · Meta AI, CMU, 及全球 13 所合作高校Kristen Grauman, Andrew Westbury, et al. (85 co-authors) · Meta AI, CMU, and 13 partner universities worldwide

Ego4D 收录 3,670 小时日常生活第一视角视频,横跨家务、户外、职场、休闲等数百种场景,是此前最大 egocentric 数据集的 20 倍以上。配套提供五大基准任务(情景记忆、手-物交互、音视频说话人分析、社交互动、行为预测),旨在催生增强现实与机器人感知领域的新一轮研究浪潮。Ego4D collects 3,670 hours of egocentric video of daily life, spanning hundreds of settings across household chores, the outdoors, the workplace and leisure — more than 20× the size of the largest previous egocentric dataset. It ships with five benchmark tasks (episodic memory, hand-object interaction, audio-visual diarization, social interaction and action forecasting), aiming to spark a new wave of research in augmented reality and robot perception.

CVPR 2022 3,670 小时视频3,670 hours of video 931 名拍摄者 · 74 地 · 9 国931 camera wearers · 74 locations · 9 countries arXiv:2110.07058 ego4d-data.org
egocentric video first-person perception episodic memory hand-object interaction action forecasting large-scale dataset 第一视角视频first-person video 增强现实augmented reality 机器人感知robot perception

01 动机Motivation

当前计算机视觉系统高度依赖「第三人称」互联网图像数据集,善于识别孤立的短片段对象与动作。然而,增强现实(AR)与机器人技术中的核心输入是第一视角(egocentric)的长流式视频——摄像头戴在人身上,实时记录日常互动、物体操作与社交行为。现有数据集在规模、多样性与真实性上均无法满足这一需求。Current computer vision systems rely heavily on "third-person" Internet image datasets and are good at recognizing objects and actions in isolated short clips. In augmented reality (AR) and robotics, however, the core input is a long, streaming video from the first-person (egocentric) viewpoint — a camera worn on the body, recording everyday interactions, object manipulation and social behavior as they happen. Existing datasets meet this need neither in scale, nor in diversity, nor in realism.

"Today's influential Internet datasets capture brief, isolated moments in time from a third-person 'spectator' view. However, in both robotics and augmented reality, the input is a long, fluid video stream from the first-person or 'egocentric' point of view — where we see the world through the eyes of an agent actively engaged with its environment."
Ego4D dataset overview: global locations and diverse activities
Ego4D 数据集全貌:随机采样 5% 视频片段,展示跨越 74 个全球地点、丰富多样的活动场景与拍摄模态。地图颜色区分各合作机构。Overview of the Ego4D dataset: a random 5% sample of video clips, showing the richly varied activities and capture modalities spanning 74 locations worldwide. Map colors distinguish the partner institutions.
3,670小时视频(hours of video)hours of video
931独立拍摄者(unique camera wearers)unique camera wearers
74全球地点(worldwide locations)worldwide locations
20×超越此前最大 egocentric 数据集the size of the largest previous egocentric dataset

Ego4D 名称来源:Ego 代表 egocentric(第一视角),4D 代表三维空间加时间维度。数据集由来自 9 个国家、5 大洲的 14 个机构历时两年联合采集,历经逾 250,000 小时的标注工作,产出数百万条时序、空间与语义标注。Origin of the name Ego4D: Ego stands for egocentric (first-person view), and 4D for three-dimensional space plus the time dimension. The dataset was collected jointly over two years by 14 institutions from 9 countries on 5 continents, went through more than 250,000 hours of annotation work, and yielded millions of temporal, spatial and semantic annotations.

Ego4D scenario distribution
Ego4D 场景分布:外圈为最常见的 14 个场景(占数据的 70%),词云展示其余 30% 场景;内圈颜色对应各合作机构。数据覆盖家务、室外、职场、休闲等数百种日常场景。Scenario distribution of Ego4D: the outer ring shows the 14 most common scenarios (70% of the data) and the word cloud the remaining 30%; inner-ring colors correspond to the partner institutions. The data covers hundreds of everyday settings including household chores, the outdoors, the workplace and leisure.

02 数据集构建与基准套件Dataset Construction and Benchmark Suite

Ego4D 的核心贡献分为两部分:(1)大规模、多样化的第一视角视频数据集;(2)覆盖「过去·现在·未来」三个时间维度的五大基准任务。The core contribution of Ego4D has two parts: (1) a large-scale, diverse egocentric video dataset; (2) five benchmark tasks covering the three temporal dimensions of past, present and future.

数据采集策略Data Collection Strategy

采用分布式采集策略:14 个合作团队分布于全球 9 个国家、5 大洲,各自招募志愿者佩戴摄像头持续拍摄 1–10 小时。绝大多数视频为非脚本、自然状态下的日常活动,拍摄者涵盖不同职业(面包师、木匠、园丁、机械师等)、年龄段(96 人超过 50 岁)与性别(45% 为女性)。平均每段原始视频约 8 分钟,远长于第三人称视频研究中常见的 10 秒片段。A distributed collection strategy was used: 14 partner teams spread over 9 countries and 5 continents each recruited volunteers to wear a camera and record continuously for 1–10 hours. The vast majority of the videos are unscripted, natural everyday activity, and the camera wearers span a range of occupations (baker, carpenter, gardener, mechanic and so on), age groups (96 people over 50) and genders (45% women). Each raw video averages about 8 minutes, far longer than the 10-second clips common in third-person video research.

采集设备使用七种不同头戴摄像头(GoPro、Vuzix Blade、Pupil Labs、ZShades、ORDRO EP6、iVue Rincon 1080、Weeview),避免模型过度拟合单一设备。除 RGB 视频外,数据还包含多种模态:Seven different head-mounted cameras were used for capture (GoPro, Vuzix Blade, Pupil Labs, ZShades, ORDRO EP6, iVue Rincon 1080, Weeview), so that models do not overfit to a single device. Besides RGB video, the data also contains several modalities:

模态Modality时长(小时)Duration (hours)
RGB video3,670
Text narrations(文本叙述)Text narrations3,670
Audio(音频)Audio2,535
IMU(惯性测量)IMU (inertial measurement)836
Faces(人脸,已授权)Faces (consented)612
3D scans(Matterport3D 扫描)3D scans (Matterport3D)491
Multi-cam(多视角同步)Multi-cam (synchronized multi-view)224
Stereo(立体视频)Stereo (stereo video)80
Gaze(眼动追踪)Gaze (eye tracking)45

叙事标注(Narrations)Narrations

所有视频在进入基准标注前,先经过叙事(narration)流程:标注人员每看完 5 分钟视频片段,以密集的时间戳句子描述拍摄者的每个动作。平均密度为 13.2 句/分钟,共产出 3.85M 条叙事句子,覆盖 1,772 个唯一动词与 4,336 个唯一名词。这些叙事既用于后续标注任务的引导,本身也是一项多模态自然语言研究资源。Before entering benchmark annotation, every video first goes through a narration pass: after watching each 5-minute video segment, annotators describe every action of the camera wearer with dense time-stamped sentences. The average density is 13.2 sentences per minute, producing 3.85M narration sentences in total and covering 1,772 unique verbs and 4,336 unique nouns. These narrations guide the later annotation tasks and are themselves a multimodal natural language research resource.

五大基准任务:过去 · 现在 · 未来Five Benchmark Tasks: Past · Present · Future

Episodic Memory benchmark
过去(Past)— Episodic Memory:给定 egocentric 视频和查询,在用户过去的视频中定位答案所在时刻或区域。包含三种查询类型:自然语言查询(NLQ)、视觉查询(VQ)、时刻查询(MQ),共约 74K 条查询,覆盖 800 小时视频。Past — Episodic Memory: given an egocentric video and a query, localize the moment or region holding the answer inside the user’s past video. It has three query types: natural language queries (NLQ), visual queries (VQ) and moment queries (MQ), about 74K queries in total covering 800 hours of video.
Present benchmarks: Hands&Objects and Audio-Visual
现在(Present)— Hands & Objects + Audio-Visual Diarization + Social Interactions: (1) Hands and Objects:识别物体状态变化的时序定位(Point-of-No-Return)、检测与分类; (2) Audio-Visual Diarization (AVD):说话人定位、追踪、身份识别、说话活动检测与语音转录; (3) Social Interactions:判断对话者是否在看向(Looking at Me, LAM)或对着摄像头拍摄者说话(Talking to Me, TTM)。Present — Hands & Objects + Audio-Visual Diarization + Social Interactions: (1) Hands and Objects: temporal localization of object state changes (Point-of-No-Return), plus detection and classification; (2) Audio-Visual Diarization (AVD): speaker localization, tracking, identification, speech activity detection and speech transcription; (3) Social Interactions: deciding whether an interlocutor is looking at (Looking at Me, LAM) or speaking to (Talking to Me, TTM) the camera wearer.
Forecasting benchmark
未来(Future)— Forecasting:包含四个子任务:(1) 运动轨迹预测(Locomotion prediction);(2) 手部运动预测(Hand movement prediction);(3) 短期物体交互预测(Short-term object interaction anticipation);(4) 长期动作序列预测(Long-term action anticipation)。Future — Forecasting: comprises four subtasks: (1) locomotion prediction; (2) hand movement prediction; (3) short-term object interaction anticipation; (4) long-term action anticipation.
Episodic Memory query types
Episodic Memory 的三种查询类型:自然语言查询(如"我把什么放进抽屉了?")、视觉查询(给定物体图片定位其最后出现位置)、时刻查询("我什么时候给孩子读书了?")。The three query types of Episodic Memory: natural language queries (e.g. "What did I put in the drawer?"), visual queries (given a picture of an object, localize where it last appeared), and moment queries ("When did I read to my child?").

03 基线实验结果Baseline Results

论文为所有五大基准设计并评测了基线模型,使用 SlowFast(ResNet-101,Kinetics-400 预训练)作为视频特征主干,BERT 作为语言特征编码器。以下为各任务关键基线结果(verbatim from the paper)。The paper designs and evaluates baseline models for all five benchmarks, using SlowFast (ResNet-101, pretrained on Kinetics-400) as the video feature backbone and BERT as the language feature encoder. The key baseline results for each task follow (verbatim from the paper).

Episodic Memory — Natural Language Query (NLQ)

基线模型BaselineR@1, IoU=0.3 (%)R@5, IoU=0.3 (%)R@1, IoU=0.5 (%)R@5, IoU=0.5 (%)
2D-TAN (test)5.8013.902.345.96
VSLNet (test)5.4711.212.806.57
2D-TAN — visual2.296.771.323.46
2D-TAN — text3.4610.131.784.38

消融实验表明,视觉特征与文本特征对 NLQ 任务均有显著贡献(去除任一特征均导致明显性能下降)。The ablations show that visual and text features both contribute substantially to the NLQ task (removing either one causes a clear drop in performance).

Forecasting — 各子任务基线Forecasting — Baselines per Subtask

子任务Subtask基线模型Baseline关键指标Key metric
Short-term object interaction anticipationFaster RCNN + SlowFastTop-5 mAP: 1.75%
Long-term action anticipation (verbs)SlowFast + TransformerED@20: 0.741
Long-term action anticipation (nouns)SlowFast + TransformerED@20: 0.784

Locomotion prediction 基线(AlexNet + KNN):1 秒预测 1-MTE = 0.73m,5 秒预测 1-MTE = 2.73m。手部运动预测(I3D encoder):左/右手 Mean Key Frame Displacement Error 分别为 64.28 / 61.18。Locomotion prediction baseline (AlexNet + KNN): 1-MTE = 0.73m at a 1-second horizon and 1-MTE = 2.73m at a 5-second horizon. Hand movement prediction (I3D encoder): Mean Key Frame Displacement Error of 64.28 / 61.18 for the left / right hand.

Social Interactions 基线Social Interactions Baselines

任务Task基线模型BaselinemAPTop-1 Acc
Looking at Me (LAM)BiLSTM + ResNet-180.690.87
Talking to Me (TTM)Video + Audio (MFCC + ResNet-18)0.540.58
Ego4D camera wearer demographics
Ego4D 拍摄者人口统计信息(自报告数据,覆盖 64% 参与者):涵盖年龄、性别、居住国与职业分布。字体大小反映职业出现频率。45% 为女性,96 人年龄超过 50 岁。Demographics of the Ego4D camera wearers (self-reported, covering 64% of the participants): age, gender, country of residence and occupation distributions. Font size reflects how often an occupation occurs. 45% are women and 96 people are over 50.

消融与分析Ablations and Analysis

通过对 NLQ 任务的视觉/文本特征消融,论文证明两类特征缺一不可:仅去除视觉特征使 R@1(IoU=0.3) 从 5.80% 降至 2.29%;仅去除文本特征降至 3.46%。这表明 egocentric NLQ 需要真正的多模态理解,而非单靠语言偏见即可解决。Through visual / text feature ablations on the NLQ task, the paper shows that neither kind of feature can be dispensed with: removing only the visual features drops R@1(IoU=0.3) from 5.80% to 2.29%, and removing only the text features drops it to 3.46%. This indicates that egocentric NLQ requires genuine multimodal understanding and cannot be solved by language priors alone.

叙事数据分析:13.2 句/分钟的标注密度、1,772 个唯一动词、4,336 个唯一名词,体现了 Ego4D 词汇的丰富性与真实日常活动的多样性。Analysis of the narration data: an annotation density of 13.2 sentences per minute, 1,772 unique verbs and 4,336 unique nouns reflect the lexical richness of Ego4D and the diversity of genuine everyday activity.

04 局限性Limitations

Note:以下局限性部分由作者在论文中明确陈述(标注为「stated」),部分为根据数据集设计推断(标注为「inferred」)。Note: some of the limitations below are explicitly stated by the authors in the paper (marked "stated"), while others are inferred from the dataset design (marked "inferred").
地理与人口覆盖不完整 [stated]Incomplete geographic and demographic coverage [stated]

论文明确指出:"74 locations is still a long way from complete coverage of the globe. In addition, the camera wearers are generally located in urban or college town areas."——农村与欠发达地区代表性不足,全球日常生活活动的完整覆盖仍十分困难。The paper states explicitly: "74 locations is still a long way from complete coverage of the globe. In addition, the camera wearers are generally located in urban or college town areas." — rural and less developed regions are under-represented, and complete coverage of everyday activity worldwide remains very hard.

COVID-19 对采集场景的影响 [stated]Impact of COVID-19 on the collected scenarios [stated]

"The COVID-19 pandemic led to ample footage in stay-at-home scenarios such as cooking, cleaning, crafts, etc. and more limited opportunities to collect video at major social public events."——疫情造成社交公共场景数据偏少,采集时间也因设备电池寿命限制而集中在一天中较为活跃的时段。"The COVID-19 pandemic led to ample footage in stay-at-home scenarios such as cooking, cleaning, crafts, etc. and more limited opportunities to collect video at major social public events." — the pandemic left social public settings under-represented, and battery life also concentrated capture in the more active hours of the day.

标注语言偏差 [stated]Language bias in the annotations [stated]

"Ego4D annotations are done by crowdsourced workers in two sites in Africa. This means that there will be at least subtle ways in which the language-based narrations are biased towards their local word choices."——叙事标注集中于少数标注站点,语言表达存在地域性偏差。"Ego4D annotations are done by crowdsourced workers in two sites in Africa. This means that there will be at least subtle ways in which the language-based narrations are biased towards their local word choices." — narration annotation is concentrated in a few annotation sites, so the wording carries a regional bias.

数据集规模与可及性挑战 [stated + inferred]Dataset scale and accessibility challenges [stated + inferred]

作者明确提供了缓解措施(预计算 SlowFast 特征、按基准子集下载),但 3,670 小时的原始视频对计算资源有限的研究者仍构成较高门槛。各基准标注仅覆盖部分小时数(48–1,000 小时不等),与全量数据之间存在差距(inferred)。The authors do provide mitigations (precomputed SlowFast features, per-benchmark subset downloads), but 3,670 hours of raw video is still a high barrier for researchers with limited compute. Each benchmark’s annotations cover only part of the hours (ranging from 48 to 1,000), leaving a gap with respect to the full data (inferred).

隐私与伦理风险 [stated]Privacy and ethical risks [stated]

论文专设附录讨论潜在社会影响:可穿戴摄像头在公共场所的普及带来隐私隐患;egocentric 感知技术若被滥用可能用于监控;未来采集工作可能缺乏同等严格的知情同意与去标识化程序。Ego4D 通过许可协议限制数据用途,但无法完全规避风险。The paper devotes an appendix to potential societal impact: the spread of wearable cameras in public places raises privacy concerns; egocentric perception technology could be abused for surveillance; and future collection efforts may lack equally rigorous informed-consent and de-identification procedures. Ego4D restricts the use of the data through a license agreement, but cannot fully avoid the risks.