Spatial-Temporal Alignment Network for Action Recognition and Detection
This paper studies how to introduce viewpoint-invariant feature representations that can help action recognition and detection. Although we have witnessed great progress of action recognition in the past decade, it remains challenging yet interesting how to efficiently model the geometric variations in large scale datasets. This paper proposes a novel Spatial-Temporal Alignment Network (STAN) that aims to learn geometric invariant representations for action recognition and action detection. The STAN model is very light-weighted and generic, which could be plugged into existing action recognition models like ResNet3D and the SlowFast with a very low extra computational cost. We test our STAN model extensively on AVA, Kinetics-400, AVA-Kinetics, Charades, and Charades-Ego datasets. The experimental results show that the STAN model can consistently improve the state of the arts in both action detection and action recognition tasks. We will release our data, models and code.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionAction RecognitionSimilar Papers 제목 키워드 기반
On the Importance of Spatial Relations for Few-shot Action Recognition
Deep learning has achieved great success in video recognition, yet still struggles to recognize novel actions when faced with only a few examples. To tackle this challenge, few-shot action recognition methods have been p…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionVideo RecognitionSpatial-Temporal Alignment Network for Action Recognition
This paper studies introducing viewpoint invariant feature representations in existing action recognition architecture. Despite significant progress in action recognition, efficiently handling geometric variations in lar…
Action RecognitionTA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarit…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMetric Learning+1Parallel Attention Interaction Network for Few-Shot Skeleton-Based Action Recognition
Learning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings …
Action RecognitionFew-Shot Skeleton-Based Action RecognitionSkeleton Based Action RecognitionAction recognition in real-world videos
The goal of human action recognition is to temporally or spatially localize the human action of interest in video sequences. Temporal localization (i.e. indicating the start and end frames of the action in a video) is re…
Action RecognitionTemporal Action LocalizationTemporal Localization