paper-with-me

홈 › Papers

A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives

2024-03-05 · CVPR 2024 1 · Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Giuseppe Averta

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once. We believe that - to effectively transfer such an holistic perception to intelligent machines - an important role is played by learning to correlate concepts and to abstract knowledge coming from different tasks, to synergistically exploit them when learning novel skills. To accomplish this, we seek for a unified approach to video understanding which combines shared temporal modelling of human actions with minimal overhead, to support multiple downstream tasks and enable cooperation when learning novel skills. We then propose EgoPack, a solution that creates a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed. We demonstrate the effectiveness and efficiency of our approach on four Ego4D benchmarks, outperforming current state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2403.03037

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

Learning reusable concepts across different egocentric video understanding tasks

2025-05-30 · Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Tatiana Tommasi 외

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecas…

Video Understanding

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…

SkillSight: Efficient First-Person Skill Assessment with Gaze

2025-11-24 · Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman arxiv

Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We introduce SkillSight for power-efficient …

MM-Ego: Towards Building Egocentric Multimodal LLMs

2024-10-09 · Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen 외

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric …

Video Understanding

EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding

2024-06-13 · Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng 외

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-bo…

Action ClassificationAction LocalizationAction Understanding