paper-with-me

Papers

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

2024-03-24 · CVPR 2024 1 · Yifei HUANG, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, LiMin Wang, Yu Qiao

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that emulates the human demonstration following process, in which individuals record egocentric videos as they execute tasks guided by demonstration videos. Focusing on the potential applications in daily assistance and professional support, EgoExoLearn contains egocentric and demonstration video data spanning 120 hours captured in daily life scenarios and specialized laboratories. Along with the videos we record high-quality gaze data and provide detailed multimodal annotations, formulating a playground for modeling the human ability to bridge asynchronous procedural actions from different viewpoints. To this end, we present benchmarks such as cross-view association, cross-view action planning, and cross-view referenced skill assessment, along with detailed analysis. We expect EgoExoLearn can serve as an important resource for bridging the actions across views, thus paving the way for creating AI agents capable of seamlessly learning by observing humans in the real world. Code and data can be found at: https://github.com/OpenGVLab/EgoExoLearn

📄 PDF Abstract BibTeX arXiv:2403.16182

Code (1)

opengvlab/egoexolearn 공식 구현 pytorch

Tasks

Action AnticipationAction Quality AssessmentLong Term AnticipationVideo Retrieval

Similar Papers 제목 키워드 기반

Towards Generalizing Temporal Action Segmentation to Unseen Views

2025-04-03 · Emad Bahrami, Olga Zatsarynna, Gianpiero Francesca, Juergen Gall

While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera vie…

Action SegmentationSegmentationTemporal Action Segmentation

Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency

2026-03-10 · Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 외 arxiv

Efficient adaptation between Egocentric (Ego) and Exocentric (Exo) views is crucial for applications such as human-robot cooperation. However, the success of most existing Ego-Exo adaptation methods relies heavily on tar…

Test-time AdaptationAction Anticipation

PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

2026-08-31 · Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim 외 arxiv

Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as…

Representation Learning

Egocentric Activity Prediction via Event Modulated Attention

2018-09-01 · ECCV 2018 9 · Yang Shen, Bingbing Ni, Zefan Li, Ning Zhuang

Predicting future activities from an egocentric viewpoint is of particular interest in assisted living. However, state-of-the-art egocentric activity understanding techniques are mostly NOT capable of predictive tasks, a…

Activity PredictionEvent ExtractionPrediction

Results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in Conjunction with CVPR 2026

2026-08-24 · Luca Rossetto, Werner Bailer, Cathal Gurrin, Graham Healy 외 arxiv

This report summarizes the contributions and results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in conjunction with CVPR 2026.