Action is in the Eye of the Beholder: Eye-gaze Driven Model for Spatio-Temporal Action Localization
We propose a new weakly-supervised structured learning approach for recognition and spatio-temporal localization of actions in video. As part of the proposed approach we develop a generalization of the Max-Path search algorithm, which allows us to efficiently search over a structured space of multiple spatio-temporal paths, while also allowing to incorporate context information into the model. Instead of using spatial annotations, in the form of bounding boxes, to guide the latent model during training, we utilize human gaze data in the form of a weak supervisory signal. This is achieved by incorporating gaze, along with the classification, into the structured loss within the latent SVM learning framework. Experiments on a challenging benchmark dataset, UCF-Sports, show that our model is more accurate, in terms of classification, and achieves state-of-the-art results in localization. In addition, we show how our model can produce top-down saliency maps conditioned on the classification label and localized latent paths.
Code (0)
등록된 구현이 없습니다.
Tasks
Action LocalizationClassificationGeneral ClassificationSpatio-Temporal Action LocalizationTemporal Action LocalizationTemporal LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video
We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. We propose a novel deep model for joint gaze estimation and actio…
Action RecognitionGaze EstimationTemporal Action LocalizationIn the Eye of the Beholder: Gaze and Actions in First Person Video
We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. To facilitate our research, we first introduce the EGTEA Gaze+ da…
Action RecognitionGaze EstimationUnderstanding Human Gaze Communication by Spatio-Temporal Graph Reasoning
This paper addresses a new problem of understanding human gaze communication in social videos from both atomic-level and event-level, which is significant for studying human social interactions. To tackle this novel and …
DecoderGraph Neural NetworkLearning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
Video-based gaze estimation methods aim to capture the inherently temporal dynamics of human eye gaze from multiple image frames. However, since models must capture both spatial and temporal relationships, performance is…
Gaze EstimationObject Referring in Videos with Language and Human Gaze
We investigate the problem of object referring (OR) i.e. to localize a target object in a visual scene coming with a language description. Humans perceive the world more as continued video snippets than as static images,…
ObjectReferring Expression