Evolving Losses for Unlabeled Video Representation Learning
We present a new method to learn video representations from unlabeled data. Given large-scale unlabeled video data, the objective is to benefit from such data by learning a generic and transferable representation space that can be directly used for a new task such as zero/few-shot learning. We formulate our unsupervised representation learning as a multi-modal, multi-task learning problem, where the representations are also shared across different modalities via distillation. Further, we also introduce the concept of finding a better loss function to train such multi-task multi-modal representation space using an evolutionary algorithm; our method automatically searches over different combinations of loss functions capturing multiple (self-supervised) tasks and modalities. Our formulation allows for the distillation of audio, optical flow and temporal information into a single, RGB-based convolutional neural network. We also compare the effects of using additional unlabeled video data and evaluate our representation learning on standard public video datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningMulti-Task LearningOptical Flow EstimationRepresentation LearningSimilar Papers 제목 키워드 기반
Evolving Losses for Unsupervised Video Representation Learning
We present a new method to learn video representations from large-scale unlabeled video data. Ideally, this representation will be generic and transferable, directly usable for new tasks such as action recognition and ze…
Action RecognitionFew-Shot LearningMulti-Task LearningRepresentation Learning+1ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations
We propose a novel strategy ES3 for self-supervised learning of robust audio-visual speech representations from unlabeled talking face videos. While many recent approaches for this task primarily rely on guiding the …
Audio-Visual Speech RecognitionLipreadingSelf-Supervised LearningSpeech RecognitionSpatial-then-Temporal Self-Supervised Learning for Video Correspondence
In low-level video analyses, effective representations are important to derive the correspondences between video frames. These representations have been learned in a self-supervised fashion from unlabeled images or video…
Contrastive LearningSelf-Supervised LearningExploiting Temporal Coherence for Self-Supervised One-shot Video Re-identification
While supervised techniques in re-identification are extremely effective, the need for large amounts of annotations makes them impractical for large camera networks. One-shot re-identification, which uses a singular labe…
One-Shot LearningSemi-Supervised Learning for Video Captioning
Deep neural networks have made great success on video captioning in supervised learning setting. However, annotating videos with descriptions is very expensive and time-consuming. If the video captioning algorithm can be…
Video Captioning