Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision
Modern self-supervised learning algorithms typically enforce persistency of instance representations across views. While being very effective on learning holistic image and video representations, such an objective becomes sub-optimal for learning spatio-temporally fine-grained features in videos, where scenes and instances evolve through space and time. In this paper, we present Contextualized Spatio-Temporal Contrastive Learning (ConST-CL) to effectively learn spatio-temporally fine-grained video representations via self-supervision. We first design a region-based pretext task which requires the model to transform in-stance representations from one view to another, guided by context features. Further, we introduce a simple network design that successfully reconciles the simultaneous learning process of both holistic and local representations. We evaluate our learned representations on a variety of downstream tasks and show that ConST-CL achieves competitive results on 6 datasets, including Kinetics, UCF, HMDB, AVA-Kinetics, AVA and OTB.
Code (2)
Tasks
Action LocalizationAction RecognitionContrastive LearningObject TrackingSelf-Supervised LearningSpatio-Temporal Action LocalizationTemporal Action LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Spatio-temporal Contrastive Domain Adaptation for Action Recognition
Unsupervised domain adaptation (UDA) for human action recognition is a practical and challenging problem. Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial rep…
Action RecognitionContrastive LearningDomain AdaptationSelf-Supervised Learning+2Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal Regions
We introduce the task of weakly supervised learning for detecting human and object interactions in videos. Our task poses unique challenges as a system does not know what types of human-object interactions are present in…
Human-Object Interaction DetectionObjectSentenceWeakly-supervised LearningSpatio-Temporal Pixel-Level Contrastive Learning-based Source-Free Domain Adaptation for Video Semantic Segmentation
Unsupervised Domain Adaptation (UDA) of semantic segmentation transfers labeled source knowledge to an unlabeled target domain by relying on accessing both the source and target data. However, the access to source data i…
Contrastive LearningDomain AdaptationSemantic SegmentationSource-Free Domain Adaptation+2Contrastive Spatio-Temporal Pretext Learning for Self-supervised Video Representation
Spatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discrim…
Contrastive LearningRepresentation LearningVideo UnderstandingWhen Do Contrastive Learning Signals Help Spatio-Temporal Graph Forecasting?
Deep learning models are modern tools for spatio-temporal graph (STG) forecasting. Though successful, we argue that data scarcity is a key factor limiting their recent improvements. Meanwhile, contrastive learning has be…
Contrastive LearningData AugmentationSemantic SimilaritySemantic Textual Similarity