paper-with-me

홈 › Papers

Learning reusable concepts across different egocentric video understanding tasks

2025-05-30 · Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Tatiana Tommasi, Giuseppe Averta

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen soon, everything all at once. To endow autonomous systems with such holistic perception, learning how to correlate concepts, abstract knowledge across diverse tasks, and leverage tasks synergies when learning novel skills is essential. In this paper, we introduce Hier-EgoPack, a unified framework able to create a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed.

📄 PDF Abstract BibTeX arXiv:2505.24690

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

EgoNCE++: Do Egocentric Video-Language Models Really Understand Hand-Object Interactions?

2024-05-28 · Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song 외

Egocentric video-language pretraining is a crucial paradigm to advance the learning of egocentric hand-object interactions (EgoHOI). Despite the great success on existing testbeds, these benchmarks focus more on closed-s…

Action RecognitionAttributeIn-Context LearningMulti-Instance Retrieval

EgoReID: Cross-view Self-Identification and Human Re-identification in Egocentric and Surveillance Videos

2016-12-24 · Shervin Ardeshir, Sandesh Sharma, Ali Broji

Human identification remains to be one of the challenging tasks in computer vision community due to drastic changes in visual features across different viewpoints, lighting conditions, occlusion, etc. Most of the literat…

Person Re-IdentificationVisual Reasoning

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

2024-08-07 · Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang 외

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…

Multi-Instance RetrievalRepresentation LearningStyle Transfer

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

2025-12-30 · Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 외 arxiv

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for u…

Referring Video Object Segmentation

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

2026-05-13 · Qinchuan Cheng, Zhantao Gong, Pengzhan Sun, Angela Yao 외 arxiv

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egoc…