paper-with-me

홈 › Papers

Where and when to look? Spatial-temporal attention for action recognition in videos

2019-05-01 · ICLR 2019 5 · Lili Meng, Bo Zhao, Bo Chang, Gao Huang, Frederick Tung, Leonid Sigal

Inspired by the observation that humans are able to process videos efficiently by only paying attention when and where it is needed, we propose a novel spatial-temporal attention mechanism for video-based action recognition. For spatial attention, we learn a saliency mask to allow the model to focus on the most salient parts of the feature maps. For temporal attention, we employ a soft temporal attention mechanism to identify the most relevant frames from an input video. Further, we propose a set of regularizers that ensure that our attention mechanism attends to coherent regions in space and time. Our model is efficient, as it proposes a separable spatio-temporal mechanism for video attention, while being able to identify important parts of the video both spatially and temporally. We demonstrate the efficacy of our approach on three public video action recognition datasets. The proposed approach leads to state-of-the-art performance on all of them, including the new large-scale Moments in Time dataset. Furthermore, we quantitatively and qualitatively evaluate our model's ability to accurately localize discriminative regions spatially and critical frames temporally. This is despite our model only being trained with per video classification labels.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAction Recognition In VideosTemporal Action LocalizationVideo Classification

Similar Papers 제목 키워드 기반

Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention

2020-04-02 · Juan-Manuel Perez-Rua, Brais Martinez, Xiatian Zhu, Antoine Toisoul 외

Attentive video modeling is essential for action recognition in unconstrained videos due to their rich yet redundant information over space and time. However, introducing attention in a deep neural network for action rec…

Action Recognition

Toward Improving the Evaluation of Visual Attention Models: a Crowdsourcing Approach

2020-02-11 · Dario Zanca, Stefano Melacci, Marco Gori

Human visual attention is a complex phenomenon. A computational modeling of this phenomenon must take into account where people look in order to evaluate which are the salient locations (spatial distribution of the fixat…

Saliency Prediction

Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification

2018-08-03 · Lin Wu, Yang Wang, Junbin Gao, Xue Li

Video-based person re-identification (re-id) is a central application in surveillance systems with significant concern in security. Matching persons across disjoint camera views in their video fragments is inherently cha…

Metric LearningPerson Re-IdentificationVideo-Based Person Re-Identification

Brain Effective Connectivity Estimation via Fourier Spatiotemporal Attention

2025-03-14 · Wen Xiong, Jinduo Liu, Junzhong Ji, Fenglong Ma

Estimating brain effective connectivity (EC) from functional magnetic resonance imaging (fMRI) data can aid in comprehending the neural mechanisms underlying human behavior and cognition, providing a foundation for disea…

Connectivity Estimation

ST-GRAT: A Novel Spatio-temporal Graph Attention Network for Accurately Forecasting Dynamically Changing Road Speed

2019-11-29 · Cheonbok Park, Chunggi Lee, Hyojin Bahng, Yunwon Tae 외

Predicting road traffic speed is a challenging task due to different types of roads, abrupt speed change and spatial dependencies between roads; it requires the modeling of dynamically changing spatial dependencies among…

Graph Attention