Few-shot Action Recognition with Implicit Temporal Alignment and Pair Similarity Optimization
Few-shot learning aims to recognize instances from novel classes with few labeled samples, which has great value in research and application. Although there has been a lot of work in this area recently, most of the existing work is based on image classification tasks. Video-based few-shot action recognition has not been explored well and remains challenging: 1) the differences of implementation details among different papers make a fair comparison difficult; 2) the wide variations and misalignment of temporal sequences make the video-level similarity comparison difficult; 3) the scarcity of labeled data makes the optimization difficult. To solve these problems, this paper presents 1) a specific setting to evaluate the performance of few-shot action recognition algorithms; 2) an implicit sequence-alignment algorithm for better video-level similarity comparison; 3) an advanced loss for few-shot learning to optimize pair similarity with limited data. Specifically, we propose a novel few-shot action recognition framework that uses long short-term memory following 3D convolutional layers for sequence modeling and alignment. Circle loss is introduced to maximize the within-class similarity and minimize the between-class similarity flexibly towards a more definite convergence target. Instead of using random or ambiguous experimental settings, we set a concrete criterion analogous to the standard image-based few-shot learning setting for few-shot action recognition evaluation. Extensive experiments on two datasets demonstrate the effectiveness of our proposed method.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionFew-Shot Learningimage-classificationImage ClassificationTemporal SequencesSimilar Papers 제목 키워드 기반
TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarit…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMetric Learning+1TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
In this paper we propose a novel Temporal Attentive Relation Network (TARN) for the problems of few-shot and zero-shot action recognition. At the heart of our network is a meta-learning approach that learns to compare re…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMeta-Learning+4Temporal-Viewpoint Transportation Plan for Skeletal Few-shot Action Recognition
We propose a Few-shot Learning pipeline for 3D skeleton-based action recognition by Joint tEmporal and cAmera viewpoiNt alIgnmEnt (JEANIE). To factor out misalignment between query and support sequences of 3D body joints…
Action RecognitionDynamic Time WarpingFew-Shot action recognitionFew Shot Action Recognition+3On the Importance of Spatial Relations for Few-shot Action Recognition
Deep learning has achieved great success in video recognition, yet still struggles to recognize novel actions when faced with only a few examples. To tackle this challenge, few-shot action recognition methods have been p…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionVideo RecognitionMotion-Modulated Temporal Fragment Alignment Network for Few-Shot Action Recognition
While the majority of FSL models focus on image classification, the extension to action recognition is rather challenging due to the additional temporal dimension in videos. To address this issue, we propose an end-t…
Action RecognitionFew-Shot action recognitionFew Shot Action Recognitionimage-classification+1