paper-with-me

홈 › Papers

Motion-Augmented Self-Training for Video Recognition at Smaller Scale

2021-05-04 · ICCV 2021 10 · Kirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. Snoek

The goal of this paper is to self-train a 3D convolutional neural network on an unlabeled video collection for deployment on small-scale video collections. As smaller video datasets benefit more from motion than appearance, we strive to train our network using optical flow, but avoid its computation during inference. We propose the first motion-augmented self-training regime, we call MotionFit. We start with supervised training of a motion model on a small, and labeled, video collection. With the motion model we generate pseudo-labels for a large unlabeled video collection, which enables us to transfer knowledge by learning to predict these pseudo-labels with an appearance model. Moreover, we introduce a multi-clip loss as a simple yet efficient way to improve the quality of the pseudo-labeling, even without additional auxiliary tasks. We also take into consideration the temporal granularity of videos during self-training of the appearance model, which was missed in previous works. As a result we obtain a strong motion-augmented representation model suited for video downstream tasks like action recognition and clip retrieval. On small-scale video datasets, MotionFit outperforms alternatives for knowledge transfer by 5%-8%, video-only self-supervision by 1%-7% and semi-supervised learning by 9%-18% using the same amount of class labels.

📄 PDF Abstract BibTeX arXiv:2105.01646

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionOptical Flow EstimationRetrievalTransfer LearningVideo Recognition

Similar Papers 제목 키워드 기반

Self-Supervised Learning of Motion-Informed Latents

2021-09-29 · Raphaël Jean, Pierre-Luc St-Charles, Soren Pirk, Simon Brodeur

Siamese network architectures trained for self-supervised instance recognition can learn powerful visual representations that are useful in various tasks. Many such approaches work by simply maximizing the similarity bet…

Action RecognitionData AugmentationPose EstimationSelf-Supervised Learning

Self-supervised Group Meiosis Contrastive Learning for EEG-Based Emotion Recognition

2022-07-12 · Haoning Kan, Jiale Yu, Jiajin Huang, Zihe Liu 외

The progress of EEG-based emotion recognition has received widespread attention from the fields of human-machine interactions and cognitive science in recent years. However, how to recognize emotions with limited labels …

Contrastive LearningData AugmentationEEGElectroencephalogram (EEG)+1

Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization

2023-03-20 · ICCV 2023 1 · Fida Mohammad Thoker, Hazel Doughty, Cees Snoek

We propose a self-supervised method for learning motion-focused video representations. Existing approaches minimize distances between temporally augmented videos, which maintain high spatial similarity. We instead propos…

Self-supervised Temporal Discriminative Learning for Video Representation Learning

2020-08-05 · Jinpeng Wang, Yiqi Lin, Andy J. Ma, Pong C. Yuen

Temporal cues in videos provide important information for recognizing actions accurately. However, temporal-discriminative features can hardly be extracted without using an annotated large-scale video action dataset for …

Action RecognitionRepresentation LearningTemporal Action LocalizationTriplet

MOFO: MOtion FOcused Self-Supervision for Video Understanding

2023-08-23 · Mona Ahmadian, Frank Guerin, Andrew Gilbert

Self-supervised learning (SSL) techniques have recently produced outstanding results in learning visual representations from unlabeled videos. Despite the importance of motion in supervised learning techniques for action…

Action ClassificationAction RecognitionRepresentation LearningSelf-Supervised Learning+1