paper-with-me

홈 › Papers

Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation

2024-07-28 · Tz-Ying Wu, Kyle Min, Subarna Tripathi, Nuno Vasconcelos

Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that Ego-VPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning.

📄 PDF Abstract BibTeX arXiv:2407.19520

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

2026-07-01 · Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri 외 arxiv

Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We s…

EgoX: Egocentric Video Generation from a Single Exocentric Video

2025-12-09 · Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park 외 arxiv

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibili…

Video Generation

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

2024-06-19 · Alessandro Suglia, Claudio Greco, Katie Baker, Jose L. Part 외

AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, n…

Question AnsweringSpatial ReasoningVideo CaptioningVideo Question Answering+1

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

2023-12-06 · CVPR 2024 1 · Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi 외

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun a…

Action AnticipationFormVideo Understanding

Unsupervised Segmentation of Action Segments in Egocentric Videos using Gaze

2017-09-30 · I. Hipiny, H. Ujir, J. L. Minoi, S. F. Samson Juan 외

Unsupervised segmentation of action segments in egocentric videos is a desirable feature in tasks such as activity recognition and content-based video retrieval. Reducing the search space into a finite set of action segm…

Activity RecognitionRetrievalVideo Retrieval