Time-Equivariant Contrastive Video Representation Learning
We introduce a novel self-supervised contrastive learning method to learn representations from unlabelled videos. Existing approaches ignore the specifics of input distortions, e.g., by learning invariance to temporal transformations. Instead, we argue that video representation should preserve video dynamics and reflect temporal manipulations of the input. Therefore, we exploit novel constraints to build representations that are equivariant to temporal transformations and better capture video dynamics. In our method, relative temporal transformations between augmented clips of a video are encoded in a vector and contrasted with other transformation vectors. To support temporal equivariance learning, we additionally propose the self-supervised classification of two clips of a video into 1. overlapping 2. ordered, or 3. unordered. Our experiments show that time-equivariant representations achieve state-of-the-art results in video retrieval and action recognition benchmarks on UCF101, HMDB51, and Diving48.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionContrastive LearningRepresentation LearningRetrievalVideo RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CIPER: Combining Invariant and Equivariant Representations Using Contrastive and Predictive Learning
Self-supervised representation learning (SSRL) methods have shown great success in computer vision. In recent studies, augmentation-based contrastive learning methods have been proposed for learning representations that …
Contrastive LearningData AugmentationRepresentation LearningLearning Temporally Equivariance for Degenerative Disease Progression in OCT by Predicting Future Representations
Contrastive pretraining provides robust representations by ensuring their invariance to different image transformations while simultaneously preventing representational collapse. Equivariant contrastive learning, on the …
Contrastive LearningESCL: Equivariant Self-Contrastive Learning for Sentence Representations
Previous contrastive learning methods for sentence representations often focus on insensitive transformations to produce positive pairs, but neglect the role of sensitive transformations that are harmful to semantic repr…
Contrastive LearningMulti-Task LearningSemantic Textual SimilaritySentenceContrastive Learning Via Equivariant Representation
Invariant Contrastive Learning (ICL) methods have achieved impressive performance across various domains. However, the absence of latent space representation for distortion (augmentation)-related information in the laten…
Contrastive LearningPeCLR: Self-Supervised 3D Hand Pose Estimation from monocular RGB via Equivariant Contrastive Learning
Encouraged by the success of contrastive learning on image classification tasks, we propose a new self-supervised method for the structured regression task of 3D hand pose estimation. Contrastive learning makes use of un…
3D Hand Pose EstimationContrastive LearningHand Pose Estimationimage-classification+4