CLTA: Contents and Length-based Temporal Attention for Few-shot Action Recognition
Few-shot action recognition has attracted increasing attention due to the difficulty in acquiring the properly labelled training samples. Current works have shown that preserving spatial information and comparing video descriptors are crucial for few-shot action recognition. However, the importance of preserving temporal information is not well discussed. In this paper, we propose a Contents and Length-based Temporal Attention (CLTA) model, which learns customized temporal attention for the individual video to tackle the few-shot action recognition problem. CLTA utilizes the Gaussian likelihood function as the template to generate temporal attention and trains the learning matrices to study the mean and standard deviation based on both frame contents and length. We show that even a not fine-tuned backbone with an ordinary softmax classifier can still achieve similar or better results compared to the state-of-the-art few-shot action recognition with precisely captured temporal attention.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Focused Large Language Models are Stable Many-Shot Learners
In-Context Learning (ICL) enables large language models (LLMs) to achieve rapid task adaptation by learning from demonstrations. With the increase in available context length of LLMs, recent experiments have shown that t…
In-Context LearningTARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
In this paper we propose a novel Temporal Attentive Relation Network (TARN) for the problems of few-shot and zero-shot action recognition. At the heart of our network is a meta-learning approach that learns to compare re…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMeta-Learning+4SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
Video captioning works on the two fundamental concepts, feature detection and feature composition. While modern day transformers are beneficial in composing features, they lack the fundamental problems of selecting and u…
Dense Video CaptioningVideo CaptioningFew-shot Action Recognition with Permutation-invariant Attention
Many few-shot learning models focus on recognising images. In contrast, we tackle a challenging task of few-shot action recognition from videos. We build on a C3D encoder for spatio-temporal video blocks to capture short…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionFew-Shot Learning+2Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA
Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal st…
audio-visual learningRetrievalVideo Retrieval