Video Contrastive Learning with Global Context
Contrastive learning has revolutionized self-supervised image representation learning field, and recently been adapted to video domain. One of the greatest advantages of contrastive learning is that it allows us to flexibly define powerful loss objectives as long as we can find a reasonable way to formulate positive and negative samples to contrast. However, existing approaches rely heavily on the short-range spatiotemporal salience to form clip-level contrastive signals, thus limit themselves from using global context. In this paper, we propose a new video-level contrastive learning method based on segments to formulate positive pairs. Our formulation is able to capture global context in a video, thus robust to temporal content change. We also incorporate a temporal order regularization term to enforce the inherent sequential structure of videos. Extensive experiments show that our video-level contrastive learning framework (VCLR) is able to outperform previous state-of-the-arts on five video datasets for downstream action classification, action localization and video retrieval. Code is available at https://github.com/amazon-research/video-contrastive-learning.
Code (1)
Tasks
Action ClassificationAction LocalizationContrastive LearningRepresentation LearningRetrievalVideo RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Denoising-Contrastive Alignment for Continuous Sign Language Recognition
Continuous sign language recognition (CSLR) aims to recognize signs in untrimmed sign language videos to textual glosses. A key challenge of CSLR is achieving effective cross-modality alignment between video and gloss se…
DenoisingRepresentation LearningSign Language RecognitionContextual Augmented Global Contrast for Multimodal Intent Recognition
Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language visual and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Exi…
Contrastive LearningIntent RecognitionMultimodal Intent RecognitionMultimodal Sentiment Analysis+2Supervised Contrastive Frame Aggregation for Video Representation Learning
We propose a supervised contrastive learning framework for video representation learning that leverages temporally global context. We introduce a video to image aggregation strategy that spatially arranges multiple frame…
Representation LearningContrastive LearningData AugmentationContext-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
Most research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification ta…
Anomaly DetectionContrastive LearningDeepFake DetectionFace Swapping+1Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval
Attempt to fully discover the temporal diversity and chronological characteristics for self-supervised video representation learning, this work takes advantage of the temporal dependencies within videos and further propo…
Action RecognitionContrastive LearningGraph LearningRepresentation Learning+3