Self-supervised Temporal Learning
Self-supervised learning (SSL) has shown its powerful ability in discriminative representations for various visual, audio, and video applications. However, most recent works still focus on the different paradigms of spatial-level SSL on video representations. How to self-supervised learn the inherent representation on the temporal dimension is still unrevealed. In this work we propose self-supervised temporal learning (SSTL), aiming at learning spatial-temporal-invariance. Inspired by spatial-based contrastive SSL, we show that significant improvement can be achieved by a proposed temporal-based contrastive learning approach, which includes three novel and efficient modules: temporal augmentations, temporal memory bank and SSTL loss. The temporal augmentations include three operators -- temporal crop, temporal dropout, and temporal jitter. Besides the contrastive paradigm, we observe the temporal contents vary between each layer of the temporal pyramid. The SSTL extends the upper-bound of the current SSL approaches by $\sim$6% on the famous video classification tasks and surprisingly improves the current state-of-the-art approaches by $\sim$100% on some famous video retrieval tasks. The code of SSTL is released with this draft, hoping to nourish the progress of the booming self-supervised learning community.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningRetrievalSelf-Supervised LearningVideo ClassificationVideo RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Self-Supervised Learning for Semi-Supervised Temporal Action Proposal
Self-supervised learning presents a remarkable performance to utilize unlabeled data for various video tasks. In this paper, we focus on applying the power of self-supervised methods to improve semi-supervised action pro…
RelationSelf-Supervised LearningSemi-Supervised Action DetectionTemporal Action LocalizationTwo Stream Self-Supervised Learning for Action Recognition
We present a self-supervised approach using spatio-temporal signals between video frames for action recognition. A two-stream architecture is leveraged to tangle spatial and temporal representation learning. Our task is …
Action RecognitionRepresentation LearningSelf-Supervised LearningTemporal Action Localization+1Can Temporal Information Help with Contrastive Self-Supervised Learning?
Leveraging temporal information has been regarded as essential for developing video understanding models. However, how to properly incorporate temporal information into the recent successful instance discrimination based…
Data AugmentationRepresentation LearningSelf-Supervised LearningVideo UnderstandingPhySU-Net: Long Temporal Context Transformer for rPPG with Self-Supervised Pre-training
Remote photoplethysmography (rPPG) is a promising technology that consists of contactless measuring of cardiac activity from facial videos. Most recent approaches utilize convolutional networks with limited temporal mode…
Self-supervised Learning for Semi-supervised Temporal Language Grounding
Given a text description, Temporal Language Grounding (TLG) aims to localize temporal boundaries of the segments that contain the specified semantics in an untrimmed video. TLG is inherently a challenging task, as it req…
Contrastive LearningPseudo LabelSelf-Supervised LearningSentence