Motion-Augmented Self-Training for Video Recognition at Smaller Scale
The goal of this paper is to self-train a 3D convolutional neural network on an unlabeled video collection for deployment on small-scale video collections. As smaller video datasets benefit more from motion than appearance, we strive to train our network using optical flow, but avoid its computation during inference. We propose the first motion-augmented self-training regime, we call MotionFit. We start with supervised training of a motion model on a small, and labeled, video collection. With the motion model we generate pseudo-labels for a large unlabeled video collection, which enables us to transfer knowledge by learning to predict these pseudo-labels with an appearance model. Moreover, we introduce a multi-clip loss as a simple yet efficient way to improve the quality of the pseudo-labeling, even without additional auxiliary tasks. We also take into consideration the temporal granularity of videos during self-training of the appearance model, which was missed in previous works. As a result we obtain a strong motion-augmented representation model suited for video downstream tasks like action recognition and clip retrieval. On small-scale video datasets, MotionFit outperforms alternatives for knowledge transfer by 5%-8%, video-only self-supervision by 1%-7% and semi-supervised learning by 9%-18% using the same amount of class labels.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionOptical Flow EstimationRetrievalTransfer LearningVideo RecognitionSimilar Papers 제목 키워드 기반
Self-Supervised Learning of Motion-Informed Latents
Siamese network architectures trained for self-supervised instance recognition can learn powerful visual representations that are useful in various tasks. Many such approaches work by simply maximizing the similarity bet…
Action RecognitionData AugmentationPose EstimationSelf-Supervised LearningSelf-supervised Group Meiosis Contrastive Learning for EEG-Based Emotion Recognition
The progress of EEG-based emotion recognition has received widespread attention from the fields of human-machine interactions and cognitive science in recent years. However, how to recognize emotions with limited labels …
Contrastive LearningData AugmentationEEGElectroencephalogram (EEG)+1Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization
We propose a self-supervised method for learning motion-focused video representations. Existing approaches minimize distances between temporally augmented videos, which maintain high spatial similarity. We instead propos…
Self-supervised Temporal Discriminative Learning for Video Representation Learning
Temporal cues in videos provide important information for recognizing actions accurately. However, temporal-discriminative features can hardly be extracted without using an annotated large-scale video action dataset for …
Action RecognitionRepresentation LearningTemporal Action LocalizationTripletMOFO: MOtion FOcused Self-Supervision for Video Understanding
Self-supervised learning (SSL) techniques have recently produced outstanding results in learning visual representations from unlabeled videos. Despite the importance of motion in supervised learning techniques for action…
Action ClassificationAction RecognitionRepresentation LearningSelf-Supervised Learning+1