Human Action Recognition: Pose-based Attention draws focus to Hands
We propose a new spatio-temporal attention based mechanism for human action recognition able to automatically attend to the hands most involved into the studied action and detect the most discriminative moments in an action. Attention is handled in a recurrent manner employing Recurrent Neural Network (RNN) and is fully-differentiable. In contrast to standard soft-attention based mechanisms, our approach does not use the hidden RNN state as input to the attention model. Instead, attention distributions are extracted using external information: human articulated pose. We performed an extensive ablation study to show the strengths of this approach and we particularly studied the conditioning aspect of the attention mechanism. We evaluate the method on the largest currently available human action recognition dataset, NTU-RGB+D, and report state-of-the-art results. Other advantages of our model are certain aspects of explanability, as the spatial and temporal attention distributions at test time allow to study and verify on which parts of the input data the method focuses.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Do We Train on Test Data? The Impact of Near-Duplicates on License Plate Recognition
This work draws attention to the large fraction of near-duplicates in the training and test sets of datasets widely adopted in License Plate Recognition (LPR) research. These duplicates refer to images that, although dif…
License Plate RecognitionAttention-Oriented Action Recognition for Real-Time Human-Robot Interaction
Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the actio…
Action RecognitionPose EstimationHuman Action Recognition Based on Spatial-Temporal Attention
Many state-of-the-art methods of recognizing human action are based on attention mechanism, which shows the importance of attention mechanism in action recognition. With the rapid development of neural networks, human ac…
Action RecognitionTemporal Action LocalizationSimultaneous Implementation Features Extraction and Recognition Using C3D Network for WiFi-based Human Activity Recognition
Human actions recognition has attracted more and more people's attention. Many technology have been developed to express human action's features, such as image, skeleton-based, and channel state information(CSI). Among t…
Activity RecognitionHuman Activity RecognitionTime Series AnalysisWeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increas…
Managementspeech-recognitionSpeech RecognitionTarget Speaker Extraction