Spatial-Temporal Alignment Network for Action Recognition
This paper studies introducing viewpoint invariant feature representations in existing action recognition architecture. Despite significant progress in action recognition, efficiently handling geometric variations in large-scale datasets remains challenging. To tackle this problem, we propose a novel Spatial-Temporal Alignment Network (STAN), which explicitly learns geometric invariant representations for action recognition. Notably, the STAN model is light-weighted and generic, which could be plugged into existing action recognition models (e.g., MViTv2) with a low extra computational cost. We test our STAN model on widely-used datasets like UCF101 and HMDB51. The experimental results show that the STAN model can consistently improve the state-of-the-art models in action recognition tasks in trained-from-scratch settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionSimilar Papers 제목 키워드 기반
On the Importance of Spatial Relations for Few-shot Action Recognition
Deep learning has achieved great success in video recognition, yet still struggles to recognize novel actions when faced with only a few examples. To tackle this challenge, few-shot action recognition methods have been p…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionVideo RecognitionTA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarit…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMetric Learning+1Spatial-Temporal Alignment Network for Action Recognition and Detection
This paper studies how to introduce viewpoint-invariant feature representations that can help action recognition and detection. Although we have witnessed great progress of action recognition in the past decade, it remai…
Action DetectionAction RecognitionParallel Attention Interaction Network for Few-Shot Skeleton-Based Action Recognition
Learning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings …
Action RecognitionFew-Shot Skeleton-Based Action RecognitionSkeleton Based Action RecognitionTwo Stream Self-Supervised Learning for Action Recognition
We present a self-supervised approach using spatio-temporal signals between video frames for action recognition. A two-stream architecture is leveraged to tangle spatial and temporal representation learning. Our task is …
Action RecognitionRepresentation LearningSelf-Supervised LearningTemporal Action Localization+1