Ego4D: Around the World in 3,000 Hours of Egocentric Video
1. Paper overview
In one sentence: Ego4D supplies diverse first-person experience and task-specific supervision, while uneven sensor coverage and separately documented baselines limit what this paper alone can establish. e-motivatione-modalitiese-suite-scopee-memorye-state-changee-forecast
| At a glance | What to know |
|---|---|
| Research problem | Source description Wearable cameras produce moving, first-person streams without the temporal selection typical of Internet clips. Useful perception must connect actions, objects, people and persistent surroundings over time. Ego4D addresses the shortage of broad recordings and task definitions for that setting. Robotics and augmented reality motivate benchmarks of recorded human experience. e-motivation |
| Core mechanism | Source description A distributed collection combines 14 teams, 74 cities and 9 countries. Participants record daily activities for 1–10 hours at a time, broadening occupations and scenarios beyond a single kitchen or laboratory. e-collectione-scenarios |
| A key reported result | Ego4D-3K collection coverage: 3,670 hours; 931 wearers; 74 cities across 9 countries. Hours, unique wearers and collection locations. First release described by the CVPR 2022 paper; dataset inventory, without a train/test performance comparison. No matched model baseline accompanies these counts. Collection scale does not establish predictive accuracy or representative global sampling. e-collection |
| Reading caution | Source description The authors identify urban/college-town concentration, pandemic-era home activity, limited public events, battery-driven sampling of active periods and vocabulary influenced by two annotation sites in Africa. Demographic information covers only 64% of participants. e-collectione-ethics-bias |
Core contributions
- Source description
A distributed collection combines 14 teams, 74 cities and 9 countries. Participants record daily activities for 1–10 hours at a time, broadening occupations and scenarios beyond a single kitchen or laboratory. e-collectione-scenarios
- Source description
Dense narrations and five benchmark families connect language, spatial localization, object transformations, conversation and future activity, providing supervision at several temporal and semantic scales. e-narrationse-suite-scopee-memorye-state-changee-diarizatione-sociale-forecast
Figure 6. Time relative to the wearer organizes five distinct benchmark families. Original paper, p. 6 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Start with the three colored headers. Past contains Episodic Memory, whose query refers back to recorded experience. Present contains three different forms of interpretation: changing objects, identifying who spoke when, and recognizing attention toward the wearer. Future contains Forecasting. The small questions beneath each panel motivate the tasks; their exact prediction formats are specified in Sections 5.1–5.5. In particular, the memory panel's spatial illustration does not mean every memory task returns a 3D answer. Read the grouping as a conceptual organization of benchmarks, not as connected neural modules or a sequence of executed actions. e-suite-scopee-memorye-state-changee-diarizatione-sociale-forecast
What it supports. The contribution separates several capabilities that a wearable assistant might need. Retrieving evidence, recognizing a transformation, understanding a conversation and anticipating an action have different targets and metrics. Reader interpretation: success on one panel cannot stand in for an evaluation of the complete collection of capabilities.
Where the evidence stops. This is the suite's task overview, not a model architecture or measured result. Section 5 refers baseline models and scores to appendices that are absent from the supplied PDF.
2. Motivation
2.1 The problem and the proposed response
Wearable cameras produce moving, first-person streams without the temporal selection typical of Internet clips. Useful perception must connect actions, objects, people and persistent surroundings over time. Ego4D addresses the shortage of broad recordings and task definitions for that setting. Robotics and augmented reality motivate benchmarks of recorded human experience. e-motivation
2.2 What this reading follows
A camera wearer sees actions unfold from inside an activity: objects move in and out of view, conversations involve an unseen participant, and a useful answer may lie minutes earlier in the recording. Ego4D organizes these challenges into a dataset and five benchmark families. Read the illustrations as a progression from collection coverage to the meanings of the prediction targets. The modality table establishes available data; the memory, state-change and forecasting figures explain what a system must return. This supplied CVPR 2022 paper delegates baseline models and scores to separate appendices, so the illustrated evidence supports task design and resource scope, not comparative model performance. e-motivatione-modalitiese-suite-scopee-memorye-state-changee-forecast
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Datasets |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
3.1 Evidence-based assessment
Supports the recorded classification
The Datasets classification is supported by collection, multisensor subsets and benchmark annotations. Architecture, prediction paradigm and quadrant are appropriately Not applicable: the suite supplies no single world-action architecture. Forecasting human activity does not itself establish joint future/action generation, inverse dynamics or an executable policy. e-modalitiese-suite-scopee-forecast
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
5. Method in detail
5.1 Begin with the annotation pipeline and the eligible data
The collection process explains what learning signals become available. Participants record lengthy daily activities; annotators then summarize and densely narrate clips before specialized annotations are created. Those narrations help derive vocabularies, select relevant videos and seed temporal windows for closer labeling. This ordering connects broad coverage to task-specific supervision: a clip about carpentry can lead to a state-change annotation, while an object interaction can seed a forecasting example. The source also reports modality subsets rather than uniform sensor coverage. Reader interpretation: defining the eligible subset is part of experimental design, since a model requiring gaze and scene meshes cannot be assumed to use every RGB clip. Likewise, benchmark-specific annotation hours should not be confused with the complete narrated collection. e-collectione-narrationse-suite-scopee-modalitiese-state-changee-forecast
Figure 7. Three query types require different kinds of retrieved evidence. Original paper, p. 6 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow the red, blue and purple markings separately. A visual query supplies an object example, illustrated by keys, and asks when and where it was last observed. Its red arrows associate the query with a response in video and spatial context; they do not show a navigation command. The blue language query points to an interval containing the evidence needed to answer a question. The purple moments query points to several intervals for a repeated high-level activity. Section 5.1 confirms the distinction between a last occurrence, an answer-bearing window and all matching activity windows. e-memorye-modalities
What it supports. The benchmark evaluates grounding in personal video history. A natural language query can be answered by returning relevant footage, without generating a verbal answer. Visual-query localization adds object position, while moments queries require finding repeated events. The supplied examples clarify these output contracts rather than measuring retrieval quality.
Where the evidence stops. Optional 3D displacement belongs to VQ and should not be assumed for all clips or query types. Search speed and localization metrics are named, but complete baselines and results are assigned to the absent Appendix F.
5.2 Translate a useful question into the exact supervised target
The benchmark suite becomes clearer when each task is read as an input/output contract. For memory, a language question yields an interval with evidence, a picture of an object yields its last localized occurrence, and an activity name yields all matching intervals. For manipulation, a short clip supports onset localization or change classification, whereas the detection task receives the designated pre-condition, PNR and post-condition frames. In conversation, identifying who is speaking and identifying whom they address are separate targets: the social labels build on face tracks and active-speaker annotations. Reader interpretation: a reproduction should preserve these supplied inputs and dependencies. Requiring a detector to infer inputs that the benchmark provides would change the problem, as would scoring a fluent verbal answer when the required output is a temporal window. e-memorye-state-changee-diarizatione-social
Figure 8. Object change is anchored in before, onset and after frames. Original paper, p. 7 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read each row from left to right, keeping the state-change description attached to that row. The plant example and the wood example illustrate different transformations using the same three temporal anchors. Section 5.2 defines PNR, point of no return, as when the change begins; it should not be read as the time the activity finishes. The figure shows these anchors, while the accompanying task definition adds bounding boxes for objects, hands and tools. Distinguish the task that predicts the PNR keyframe from the task that receives the three frames and detects the object undergoing change. e-state-change
What it supports. The annotation scheme separates when a change begins, which object changes and whether a change occurred. The authors motivate this representation by the possibility that different tools or motions produce the same transformation. The examples explain that intended abstraction, but do not demonstrate invariance to tool choice.
Where the evidence stops. These are annotation examples, not predictions with measured errors. The main paper gives temporal error, detection AP and classification accuracy as metrics; model scores and implementation details are referred to Appendix G.
5.3 Keep future prediction separate from physical control evidence
Forecasting connects the dataset most directly to world-model research. The annotations relate an observed prefix to later movement, object contact and actions, with structure from motion providing wearer trajectories and annotated contact frames organizing interactions. These are concrete ways to study regularities in human behavior. The supplied paper, however, defines prediction tasks and names their metrics; it does not present a learned simulator that accepts candidate control commands, a policy that chooses among them or a feedback loop that executes them. Reader interpretation: improved anticipation could be useful to an embodied system, but that transfer would require another experiment. Even the offline baseline comparison is incomplete here because Appendix J contains details and scores beyond the supplied artifact. Dataset scale and illustrative future labels cannot supply that missing evidence. e-forecaste-forecast-metricse-suite-scope
Figure 10. Forecasting spans motion, object contact and sequences of action labels. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Across the upper row, distinguish the wearer's movement from hand movement and object interaction. The short-term panel attaches a verb and time-to-contact to candidate objects in the most recent observed frame. Its displayed 0.8 seconds is an illustrative prediction, not a reported error or universal horizon. In the lower row, the bracket marks input video and the dashed divider separates it from the illustrated future. Arrows between verb–noun labels order future actions. Section 5.5 defines these outputs as prediction targets; neither the dashed boundary nor the arrows specify an internal network or an action execution loop. e-forecaste-forecast-metrics
What it supports. The same recorded activity can motivate several evaluation problems. Locomotion and hand prediction use L2 distance, short-term object anticipation uses Top-5 mAP, and long-term anticipation uses edit distance. The figure's future frames provide context for the action-label sequence; they do not establish a benchmark result for generated video.
Where the evidence stops. The main text describes a Top-4 false-negative discount within Top-5 mAP without defining its implementation. Appendix J is absent, leaving that operational rule and baseline performance unresolved. No robot action execution is demonstrated here.
5.4 Training and inference
During training
Narrations, benchmark labels and precomputed SlowFast features provide learning resources. The main paper specifies no unified learned representation, objective, optimization schedule, frozen modules or compute budget. Baseline architectures, training details and scores are assigned to separate appendices, so a complete training recipe cannot be recovered here. e-modalitiese-suite-scope
During inference
Tasks return predictions for offline evaluation. Forecasting uses L2 distance for movement, Top-5 mAP for short-term anticipation and edit distance for action sequences. Reader assessment: these outputs do not establish a demonstrated robot controller, executed action policy or closed-loop feedback system. e-memorye-state-changee-diarizatione-sociale-forecaste-forecast-metrics
5.5 Implementation flow
- Collect varied activity under explicit consent
Recruitment seeks everyday scenarios while allowing participants to choose activities. A typical raw clip lasts 8 minutes. Some social sessions are arranged with consent for conversation and unblurred faces. Site policies allow review, redaction and withdrawal. e-collectione-scenariose-ethics-bias
- Preserve modality and viewpoint distinctions
Seven camera types provide different fields of view and sensor combinations. Head-down cameras emphasize manipulation; heads-up views can miss near-body interactions. Some videos have static Matterport3D meshes. Modality availability applies to subsets, so an all-sensor input cannot be assumed. e-modalities
- Narrate before specializing annotations
Annotators summarize five-minute clips, rewatch and pause to write timestamped descriptions. Each video receives two independent narrations. These seed action/object taxonomies, select benchmark videos and locate windows for precise annotations. e-narrations
- Make episodic memory queryable
NLQ returns an answer-bearing window; VQ locates the last occurrence of an example object; MQ returns every matching activity interval. VQ includes 2D boxes and optional 3D displacement. NLQ uses overlap-threshold recall, MQ uses mAP and recall, and VQ also measures search timeliness. e-memory
- Annotate object transformations
Hands and Objects anchors changes with pre-condition, PNR and post-condition frames, plus boxes for hands, tools and objects. Tasks predict change onset, the changed object and change/no-change classification, evaluated with temporal error in seconds, detection AP and accuracy. e-state-change
- Separate speaking from social address
Audio-Visual Diarization tracks faces, detects speakers including the unseen wearer, diarizes speech and transcribes English speech. Social tasks ask whether tracked faces look at or talk to the wearer. Metrics include tracking measures, speaker error, DER/WER and frame-level social mAP/Top-1 accuracy. e-diarizatione-social
- Define several futures
Forecasting targets wearer locomotion, hand positions, next interacted objects with verbs/contact times, and action sequences. Contact and pre-condition frames guide labels, with earlier boxes at 0.5, 1 and 1.5 seconds before pre-condition. Structure from motion supplies wearer trajectories. e-forecast
6. Experiments & results
Ego4D-3K turns long recordings of human daily activity into a resource for retrieving past events, interpreting present interactions and anticipating future behavior. Its contribution is collection and task design: 3,670 hours from 931 wearers, with uneven additional modalities. The supplied paper establishes dataset scope and evaluation targets, but delegates model scores and implementation details to absent appendices.
The supplied 18-page dataset paper contains task illustrations and one modality inventory table, but no neural architecture diagram, quantitative model-performance table or controlled ablation. Section 5 and Sections 5.1–5.5 explicitly send baseline models, scores and details to separate appendices that are not supplied. Figure 6 therefore serves as the method overview, Table 1 supplies descriptive quantitative evidence, and Figure 3 supplies a distribution diagnostic under the ablation section without being represented as a controlled experiment. The edition cannot establish comparative accuracy, training cost or mechanism-isolating improvements. e-suite-scopee-modalitiese-scenariose-memorye-state-changee-diarizatione-sociale-forecast-metrics
6.1 Read the original evidence
Table 1. Additional sensing is available on subsets of the RGB collection. Original paper, p. 5 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the second row as video hours associated with each column, not counts of sensors or annotations. RGB, text narrations and Features each cover 3,670 hours. The table caption defines Features as precomputed SlowFast video features, Faces as video whose participants consented to remain unblurred, and Multi-cam as synchronized recordings of an event from several wearers. The 3D scans column refers to video coupled with Matterport3D environment meshes, rather than a count of meshes. These definitions matter when deciding which inputs are actually available for an experiment; the table does not provide the intersections of the modality subsets. e-modalities
What it supports. The inventory reports 2,535 hours with audio, 491 with 3D scans and 45 with gaze, against 3,670 RGB hours. Reader interpretation: a study that requires all of these inputs must determine the eligible common subset before claiming to use the full Ego4D collection.
Where the evidence stops. These are dataset statistics, not accuracy measurements. The columns overlap and must not be summed. Joint sensor coverage, benchmark membership and train/test composition cannot be inferred from the marginal totals.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| Ego4D-3K collection coverage First release described by the CVPR 2022 paper; dataset inventory, without a train/test performance comparison. | 3,670 hours; 931 wearers; 74 cities across 9 countries. Hours, unique wearers and collection locations | No matched model baseline accompanies these counts. Collection scale does not establish predictive accuracy or representative global sampling. e-collection |
| Ego4D-3K modality coverage Table 1, first-release video hours associated with each modality; overlapping totals, not split-specific. | RGB, narrations and features: 3,670 each; audio: 2,535; faces: 612; 3D scans: 491; stereo: 80; gaze: 45; IMU: 836; multi-cam: 224. Hours | Additional modalities cover less video than RGB; no multimodal model comparison is supplied. These marginal counts do not specify joint sensor availability or a shared evaluation subset. e-modalities |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Figure 3. A descriptive diagnostic reveals broad coverage with substantial concentration. Original paper, p. 4 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the labeled outer sectors first: Arts & Crafts, Cleaning & Laundry, Cooking and Construction occupy visibly large shares. The yellow Other sector expands into a word cloud that names additional scenarios; individual word labels do not provide exact category percentages. The inner ring encodes contributing partners, with the color key in the paper's Figure 1 on PDF page 2, rather than another set of scenario classes. Section 3.2 explains that a scenario can contain many distinct actions. Thus a large Cooking sector is a setting-level observation, not a direct count of any particular hand motion or object interaction. e-scenariose-ethics-bias
What it supports. The plotted shares include Arts & Crafts at 15%, Cleaning & Laundry at 13%, Cooking at 11% and Construction at 9%. These observations support a breadth-and-concentration reading of the collection. They do not show balanced coverage or establish that a model generalizes across the less frequent settings.
Where the evidence stops. The plot says Other 31%, but its caption describes the remaining 30%; the source does not reconcile this. This is a dataset diagnostic, not an ablation. Partner-color assignments require Figure 1's separate map key.
7. Analysis & limitations
7.1 What the evidence leaves open
The authors identify urban/college-town concentration, pandemic-era home activity, limited public events, battery-driven sampling of active periods and vocabulary influenced by two annotation sites in Africa. Demographic information covers only 64% of participants. e-collectione-ethics-bias
Baseline development is reported, but no performance table or controlled ablation is supplied. Scenario frequencies are descriptive. Figure 3 labels Other as 31%, while its caption says 30%; this discrepancy is preserved. e-suite-scopee-scenarios
The forecasting metric paragraph describes discounting Top-4 false negative predictions without an operational definition. Its exact Top-5 matching/discount rule remains unresolved without Appendix J; substituting a familiar mAP implementation would be an assumption. e-forecast-metrics
7.2 Questions for discussion
- How much NLQ performance survives query/video shuffling under participant-disjoint evaluation?
- Can state-change models generalize to new object/tool combinations?
- How does requiring several modalities change the geographic and scenario mix of an evaluation subset?
8. Reproducibility audit
8.1 Requirements and known gaps
A faithful baseline comparison requires the release, benchmark annotations, splits, sampling rules, evaluation implementations and appendix-specific training configurations. The paper names an Ego4D license without establishing its complete terms in this text. Modality totals cannot reconstruct benchmark membership. e-collectione-modalitiese-suite-scope
Proposed checks: compare NLQ grounding against shuffled query/video pairings, and test state-change recognition on held-out object/tool combinations against an object-only control. Both require explicit participant-disjoint splits and uncertainty reporting; neither check was performed. e-memorye-state-change
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Does natural language retrieval use the query and its video evidence?
Reader-proposed check, not performed: once the release, split definitions and NLQ evaluator are available, compare a fixed retrieval model on the original query/video pairs against the same videos with queries shuffled within question-template groups. Keep the original target intervals, candidate-window budget and temporal-overlap thresholds fixed; add a query-independent temporal-prior control. Use participant-disjoint data, report per-query recall differences and participant-level bootstrap uncertainty, and inspect ambiguous shuffled examples separately. If correct queries show no reliable advantage, the model's apparent retrieval ability may be dominated by temporal or dataset priors. A paired advantage supports sensitivity to the query, while still falling short of proving unrestricted question answering. e-memorye-suite-scope
Check 2: Does state-change recognition survive new object and tool combinations?
Reader-proposed check, not performed: construct a supplementary participant-disjoint split that holds out object/tool combinations while retaining state-change types in training, using the source's object, tool and change annotations. Compare a fixed video model with an object-only pre-condition-frame classifier. Match training examples and class balance, stratify test results by familiar versus held-out combinations, and report classification accuracy with participant-level uncertainty. Evaluate PNR temporal error separately on change-positive clips to avoid merging incompatible metrics. If both models collapse similarly on new combinations, recognition may rely on object appearance instead of transformation evidence. A smaller video-model drop would motivate further tests of the invariance sought by the benchmark. This would be a new diagnostic split, not an official reproduced score. e-state-change
8.3 Reading coverage
Visual audit: Visually inspected the title/author block and CVF version watermark; Figures 1–10; Table 1; all main-body pages supporting collection, modalities, annotation, task, metric, bias and proposed-check claims; and the first and last reference pages for document boundaries and pagination. All six final crops were opened and checked. Figure 7's query arrows, Figure 8's onset anchors and Figure 10's observed/future boundary were cross-checked with Sections 5.1, 5.2 and 5.5. Figure 3's 31%/30% discrepancy is disclosed. All eight text chunks, including the complete bibliography, were read. Bibliography-only pages 11–17 were read as text, not visually inspected. No supplemental appendices, code, dataset files or linked resources were inspected.
PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 18. Appendix coverage: not read.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Abstract
- 1. Introduction
- 2. Related Work
- 3. Ego4D Dataset, including 3.1–3.5
- 4. Narrations of Camera Wearer Activity
- 5. Ego4D Benchmark Suite, including 5.1–5.5
- 6. Conclusion
- Contribution statement and Acknowledgements
- References, PDF pp. 10–18
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- All eight supplied text chunks were read individually. Full-text status covers this 18-page main paper and references; separately referenced appendices were not supplied or read.
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout. This acquisition limitation was addressed by inspecting original PDF pages and final crops.
- Separate supplemental material availability has not been fully verified.
- Referenced appendices A, D, F–K cover collection, narrations, baselines, evaluation and societal impact. Their detailed sampling/splits, training configurations, scores and ablations cannot be verified here.
- Identity notes: the title and author roster match, with catalog/observed name-form differences: Kiran K. Somasundaram / Kiran Somasundaram, Richard A. Newcombe / Richard Newcombe, and Jáchym Kolár / Jáchym Kolář. Metadata uses the inspected forms (e-identity).
- Edition notes: the supplied CVPR 2022 CVF Open Access artifact states equivalence to the accepted version except for its watermark. No independent IEEE or revision comparison was made. Printed pages are 18995–19012 versus catalog BibTeX pages 18973–18990. No numbered revision is observed (e-identity).
- Code, dataset files, external license terms and linked resources were not inspected; no experiments were reproduced.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e-identityPDF p. 1, title, author block and CVF watermark; pp. 1 and 18, printed page numbers
The title matches. The CVF Open Access watermark states equivalence to the accepted version except for the watermark. Author forms include Kiran Somasundaram, Richard Newcombe and Jáchym Kolář. Printed pagination is 18995–19012.
Go to primary source ↓e-motivationPDF pp. 1–2, Abstract and Section 1
Egocentric daily-life video and benchmarks address past, present and future perception, motivated by long, uncurated first-person streams and persistent spatial understanding.
Go to primary source ↓e-collectionPDF p. 3, Sections 3–3.1, Figure 2 and footnote 1
Ego4D-3K reports 3,670 hours, 931 wearers, 74 cities, 9 countries and 14 collecting teams. Sessions last 1–10 hours. Demographics cover 64% of participants. Availability is under an Ego4D license.
Go to primary source ↓e-scenariosPDF p. 4, Section 3.2 and Figure 3, outer ring and caption
Recruitment seeks varied naturally occurring scenarios; typical raw clips last 8 minutes. Some consented social activities are arranged at five sites. The plot labels Arts & Crafts 15%, Cleaning & Laundry 13%, Cooking 11%, Construction 9%, Other 31%; the caption describes a 70%/30% division.
Go to primary source ↓e-modalitiesPDF pp. 4–5, Section 3.3, Figure 4, Table 1 and caption
Seven camera types have different views and modalities. Hours: RGB, narrations and features 3,670 each; audio 2,535; faces 612; 3D scans 491; stereo 80; gaze 45; IMU 836; multi-cam 224. Features are precomputed SlowFast features; faces means consented unblurred video.
Go to primary source ↓e-ethics-biasPDF pp. 4–5, Sections 3.4–3.5
Policies include informed consent, withdrawal/redaction and de-identification. Biases include urban/college-town geography, pandemic-era home activities, battery-limited active periods and local vocabulary from two annotation sites in Africa.
Go to primary source ↓e-narrationsPDF p. 5, Section 4 and Figure 5
Two annotators independently summarize and repeatedly pause five-minute clips to describe actions with timestamps. Reported totals are 13.2 sentences/minute, 3.85 million sentences, 1,772 verbs and 4,336 nouns. Narrations seed taxonomies, benchmark selection and annotation windows.
Go to primary source ↓e-suite-scopePDF p. 5, Section 5; p. 6, Figure 6; pp. 10 and 18, References boundaries
Five benchmark families organize past, present and future perception. Benchmark annotations span 48–1,000 hours per family. Section 5 places baseline models and quantitative results in appendices; the supplied main paper proceeds to References and ends there.
Go to primary source ↓e-memoryPDF p. 6, Section 5.1 and Figure 7
NLQ returns an answer-bearing interval; VQ returns the last object occurrence with 2D localization and optional 3D displacement; MQ returns all matching activity intervals. Approximately 74K queries span 1,000 hours. Metrics include temporal-overlap recall, mAP and VQ timeliness. Baselines are referred to Appendix F.
Go to primary source ↓e-state-changePDF p. 7, Section 5.2 and Figure 8
Pre-condition, point-of-no-return (PNR) and post-condition frames anchor state-change labels. PNR means onset of change. Metrics are absolute temporal error in seconds, object-detection AP and classification accuracy. Baselines are referred to Appendix G.
Go to primary source ↓e-diarizationPDF pp. 7–8, Section 5.3 and Figure 9
Audio-Visual Diarization covers face localization/tracking, active speaker detection including the unseen wearer, speech diarization and English-only transcription. Metrics are MOT measures, speaker error rate, DER and WER; baselines are referred to Appendix H.
Go to primary source ↓e-socialPDF p. 8, Section 5.4 and Figure 9
Looking at me and Talking to me classify visible tracked faces relative to the wearer, building on face tracks and active-speaker annotations. Evaluation uses mAP and Top-1 accuracy, with precision measured per frame; details are referred to Appendix I.
Go to primary source ↓e-forecastPDF pp. 8–9, Section 5.5 and Figure 10
Targets are ground-plane locomotion, hand positions, next interacted objects with verbs/contact times, and future action sequences. Annotations include contact/pre-condition frames and boxes 0.5, 1 and 1.5 seconds before pre-condition. Wearer trajectories use structure from motion.
Go to primary source ↓e-forecast-metricsPDF p. 9, Section 5.5, Evaluation metrics and baselines
Movement uses L2 distance; short-term anticipation uses Top-5 mAP, described as discounting Top-4 false negative predictions; long-term anticipation uses edit distance. Operational rules and baselines are referred to Appendix J.
Go to primary source ↓8.5 Primary sources
Ego4D: Around the World in 3,000 Hours of Egocentric Video ↗
PDF · 13,559 extracted words
Source fingerprint
e5b723a0531314b530e0275ec6a92e0379e32410631ec888f3fc76da4f5b8666