paper-with-me

Papers

HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model

2024-06-01 · Khoa Vo, Thinh Phan, Kashu Yamazaki, Minh Tran, Ngan Le

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretation; and (2) learning is limited in capturing inherent fine-grained relationships between two modalities. In this paper, we take an inspiration from human perception and explore a compositional approach for egocentric video representation. We introduce HENASY (Hierarchical ENtities ASsemblY), which includes a spatiotemporal token grouping mechanism to explicitly assemble dynamically evolving scene entities through time and model their relationship for video representation. By leveraging compositional structure understanding, HENASY possesses strong interpretability via visual grounding with free-form text queries. We further explore a suite of multi-grained contrastive losses to facilitate entity-centric understandings. This comprises three alignment types: video-narration, noun-entity, verb-entities alignments. Our method demonstrates strong interpretability in both quantitative and qualitative experiments; while maintaining competitive performances on five downstream tasks via zero-shot transfer or as video/text representation, including video/text retrieval, action recognition, multi-choice query, natural language query, and moments query.

📄 PDF Abstract BibTeX arXiv:2406.00307

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionActivity RecognitionDecoderLanguage ModelingLanguage ModellingText RetrievalVideo-Text RetrievalVideo UnderstandingVisual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

2026-08-20 · Marko Haralović, Akash Ramakrishnan, Estefania Talavera Martinez arxiv

Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving o…

Action UnderstandingAction Recognition

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

2023-12-06 · CVPR 2024 1 · Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi 외

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun a…

Action AnticipationFormVideo Understanding

MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction

2025-12-12 · Bate Li, Houqiang Zhong, Zhengxue Cheng, Qiang Hu 외 arxiv

Multi-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-vi…

PlayerOne: Egocentric World Simulator

2025-06-11 · Yuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai 외

We introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric scene image from the user, PlayerOne ca…

Video Generation

EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding

2026-02-27 · Shitong Sun, Ke Han, Yukai Huang, Weitong Cai 외 arxiv

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their a…

Video Question Answering