FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language supervision has contributed to developing transferable visual pretraining. Our goal is to achieve a more semantic video representation by leveraging the text related to the video content during the pretraining in a fully self-supervised manner. To this end, we present FILS, a novel self-supervised video Feature prediction In semantic Language Space (FILS). The vision model can capture valuable structured information by correctly predicting masked feature semantics in language space. It is learned using a patch-wise video-text contrastive strategy, in which the text representations act as prototypes for transforming vision features into a language space, which are then used as targets for semantically meaningful feature prediction using our masked encoder-decoder structure. FILS demonstrates remarkable transferability on downstream action recognition tasks, achieving state-of-the-art on challenging egocentric datasets, like Epic-Kitchens, Something-SomethingV2, Charades-Ego, and EGTEA, using ViT-Base. Our efficient method requires less computation and smaller batches compared to previous works.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionDecoderSimilar Papers 제목 키워드 기반
Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
The success of deep neural networks generally requires a vast amount of training data to be labeled, which is expensive and unfeasible in scale, especially for video collections. To alleviate this problem, in this paper,…
Action RecognitionPredictionSelf-Supervised Action RecognitionTemporal Action Localization+1Hierarchical Self-supervised Representation Learning for Movie Understanding
Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understanding and propose a novel hierarchical se…
Action RecognitionContrastive LearningRepresentation LearningRepresentation Learning with Video Deep InfoMax
Self-supervised learning has made unsupervised pretraining relevant again for difficult computer vision tasks. The most effective self-supervised methods involve prediction tasks based on features extracted from diverse …
Action RecognitionData AugmentationRepresentation LearningSelf-Supervised LearningCollaboratively Self-supervised Video Representation Learning for Action Recognition
Considering the close connection between action recognition and human pose estimation, we design a Collaboratively Self-supervised Video Representation (CSVR) learning framework specific to action recognition by jointly …
Action RecognitionPose EstimationPose PredictionRepresentation LearningPixel-level Correspondence for Self-Supervised Learning from Video
While self-supervised learning has enabled effective representation learning in the absence of labels, for vision, video remains a relatively untapped source of supervision. To address this, we propose Pixel-level Corres…
Contrastive Learningimage-classificationImage ClassificationOptical Flow Estimation+3