paper-with-me

Papers

Egocentric Video Description based on Temporally-Linked Sequences

2017-04-07 · Marc Bolaños, Álvaro Peris, Francisco Casacuberta, Sergi Soler, Petia Radeva

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the user. A natural topic that arises in egocentric vision is storytelling, that is, how to understand and tell the story relying behind the pictures. In this paper, we tackle storytelling as an egocentric sequences description problem. We propose a novel methodology that exploits information from temporally neighboring events, matching precisely the nature of egocentric sequences. Furthermore, we present a new method for multimodal data fusion consisting on a multi-input attention recurrent network. We also publish the first dataset for egocentric image sequences description, consisting of 1,339 events with 3,991 descriptions, from 55 days acquired by 11 people. Furthermore, we prove that our proposal outperforms classical attentional encoder-decoder methods for video description.

📄 PDF Abstract BibTeX arXiv:1704.02163

Code (1)

MarcBS/TMA 공식 구현

Tasks

DecoderVideo Description

Similar Papers 제목 키워드 기반

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

2026-06-24 · Agnese Taluzzi, Riccardo Santambrogio, Simone Mentasti, Chiara Plizzari 외 arxiv

Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video understanding approaches, especially in …

Video Question Answering

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

2023-12-06 · CVPR 2024 1 · Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi 외

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun a…

Action AnticipationFormVideo Understanding

You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions

2019-04-22 · CVPR 2020 6 · Evonne Ng, Donglai Xiang, Hanbyul Joo, Kristen Grauman

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a le…

Pose Estimation

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

2025-09-05 · Ahad Jawaid, Yu Xiang arxiv

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce Op…

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

2025-11-22 · Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 외 arxiv

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controll…

Video GenerationVideo Prediction