paper-with-me

홈 › Papers

In the Eye of the Beholder: Gaze and Actions in First Person Video

2020-05-31 · Yin Li, Miao Liu, James M. Rehg

We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. To facilitate our research, we first introduce the EGTEA Gaze+ dataset. Our dataset comes with videos, gaze tracking data, hand masks and action annotations, thereby providing the most comprehensive benchmark for First Person Vision (FPV). Moving beyond the dataset, we propose a novel deep model for joint gaze estimation and action recognition in FPV. Our method describes the participant's gaze as a probabilistic variable and models its distribution using stochastic units in a deep network. We further sample from these stochastic units, generating an attention map to guide the aggregation of visual features for action recognition. Our method is evaluated on our EGTEA Gaze+ dataset and achieves a performance level that exceeds the state-of-the-art by a significant margin. More importantly, we demonstrate that our model can be applied to larger scale FPV dataset---EPIC-Kitchens even without using gaze, offering new state-of-the-art results on FPV action recognition.

📄 PDF Abstract BibTeX arXiv:2006.00626

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionGaze Estimation

Similar Papers 제목 키워드 기반

In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video

2018-09-01 · ECCV 2018 9 · Yin Li, Miao Liu, James M. Rehg

We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. We propose a novel deep model for joint gaze estimation and actio…

Action RecognitionGaze EstimationTemporal Action Localization

Action is in the Eye of the Beholder: Eye-gaze Driven Model for Spatio-Temporal Action Localization

2013-12-01 · NeurIPS 2013 12 · Nataliya Shapovalova, Michalis Raptis, Leonid Sigal, Greg Mori

We propose a new weakly-supervised structured learning approach for recognition and spatio-temporal localization of actions in video. As part of the proposed approach we develop a generalization of the Max-Path search al…

Action LocalizationClassificationGeneral ClassificationSpatio-Temporal Action Localization+2

Following Gaze in Video

2017-10-01 · ICCV 2017 10 · Adria Recasens, Carl Vondrick, Aditya Khosla, Antonio Torralba

Following the gaze of people inside videos is an important signal for understanding people and their actions. In this paper, we present an approach for following gaze in video by predicting where a person (in the video) …

Mutual Context Network for Jointly Estimating Egocentric Gaze and Actions

2019-01-07 · Yifei Huang, Zhenqiang Li, Minjie Cai, Yoichi Sato

In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what…

Action RecognitionGaze PredictionPredictionTemporal Action Localization

Social Behavior Prediction from First Person Videos

2016-11-29 · Shan Su, Jung Pyo Hong, Jianbo Shi, Hyun Soo Park

This paper presents a method to predict the future movements (location and gaze direction) of basketball players as a whole from their first person videos. The predicted behaviors reflect an individual physical space tha…

3D ReconstructionPrediction