paper-with-me

Papers

Ego4D: Around the World in 3,000 Hours of Egocentric Video

2021-10-13 · CVPR 2022 1 · Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina Gonzalez, James Hillis, Xuhua Huang, Yifei HUANG, Wenqi Jia, Weslie Khoo, Jachym Kolar, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Ziwei Zhao, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fuegen, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, Jitendra Malik

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/

📄 PDF Abstract BibTeX arXiv:2110.07058

Code (8)

facebookresearch/Ego4d 공식 구현 pytorch
ego4d/episodic-memory pytorch
ego4d/forecasting pytorch
ego4d/hands-and-objects
frenchkrab/is2023-powerset-diarization
nnnnai/ego4d_nlq_2022_1st_place_solution pytorch
pyannote/pyannote-audio pytorch
yeliudev/R2-Tuning pytorch

Tasks

De-identificationEthics

Methods 이 논문이 사용한 방법론

Uphold 설명 없음

Similar Papers 제목 키워드 기반

RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D

2023-08-23 · ICCV 2023 1 · Shuhei Kurita, Naoki Katsura, Eri Onami

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capa…

ObjectObject TrackingReferring ExpressionReferring Expression Comprehension

LookOut: Real-World Humanoid Egocentric Navigation

2025-08-20 · Boxiao Pan, Adam W. Harley, C. Karen Liu, Leonidas J. Guibas arxiv

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging p…

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

2026-06-15 · Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou 외 arxiv

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos…

Wanderlust: Online Continual Object Detection in the Real World

2021-08-25 · ICCV 2021 10 · Jianren Wang, Xin Wang, Yue Shang-Guan, Abhinav Gupta

Online continual learning from data streams in dynamic environments is a critical direction in the computer vision field. However, realistic benchmarks and fundamental studies in this line are still missing. To bridge th…

Continual LearningObjectobject-detectionObject Detection

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

2025-01-12 · Wenqi Zhou, Kai Cao, Hao Zheng, Xinyi Zheng 외

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activit…

Video Understanding