Two-stream joint matching method based on contrastive learning for few-shot action recognition
Although few-shot action recognition based on metric learning paradigm has achieved significant success, it fails to address the following issues: (1) inadequate action relation modeling and underutilization of multi-modal information; (2) challenges in handling video matching problems with different lengths and speeds, and video matching problems with misalignment of video sub-actions. To address these issues, we propose a Two-Stream Joint Matching method based on contrastive learning (TSJM), which consists of two modules: Multi-modal Contrastive Learning Module (MCL) and Joint Matching Module (JMM). The objective of the MCL is to extensively investigate the inter-modal mutual information relationships, thereby thoroughly extracting modal information to enhance the modeling of action relationships. The JMM aims to simultaneously address the aforementioned video matching problems. The effectiveness of the proposed method is evaluated on two widely used few shot action recognition datasets, namely, SSv2 and Kinetics. Comprehensive ablation experiments are also conducted to substantiate the efficacy of our proposed approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action RecognitionMetric LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MoLo: Motion-augmented Long-short Contrastive Learning for Few-shot Action Recognition
Current state-of-the-art approaches for few-shot action recognition achieve promising performance by conducting frame-level matching on learned visual features. However, they generally suffer from two limitations: i) the…
Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action RecognitionContrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation Learning and Retrieval
Recently, the cross-modal pre-training task has been a hotspot because of its wide application in various down-streaming researches including retrieval, captioning, question answering and so on. However, exiting methods …
Contrastive LearningCross-Modal RetrievalImage RetrievalQuestion Answering+3MEDBind: Unifying Language and Multimodal Medical Data Embeddings
Medical vision-language pretraining models (VLPM) have achieved remarkable progress in fusing chest X-rays (CXR) with clinical texts, introducing image-text data binding approaches that enable zero-shot learning and down…
Language ModelingLanguage ModellingLarge Language ModelRetrieval+2MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning
Existing contrastive language-image pre-training aims to learn a joint representation by matching abundant image-text pairs. However, the number of image-text pairs in medical datasets is usually orders of magnitude smal…
Contrastive LearningRepresentation LearningSentenceCOOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-Training for Vision-Language Representation
There has been a recent surge of interest in cross-modal pre-training. However, existed approaches pre-train a one-stream model to learn joint vision-language representation, which suffers from calculation explosion …
Contrastive LearningCross-Modal RetrievalImage RetrievalRetrieval+1