Representation Learning with Video Deep InfoMax
Self-supervised learning has made unsupervised pretraining relevant again for difficult computer vision tasks. The most effective self-supervised methods involve prediction tasks based on features extracted from diverse views of the data. DeepInfoMax (DIM) is a self-supervised method which leverages the internal structure of deep networks to construct such views, forming prediction tasks between local features which depend on small patches in an image and global features which depend on the whole image. In this paper, we extend DIM to the video domain by leveraging similar structure in spatio-temporal networks, producing a method we call Video Deep InfoMax(VDIM). We find that drawing views from both natural-rate sequences and temporally-downsampled sequences yields results on Kinetics-pretrained action recognition tasks which match or outperform prior state-of-the-art methods that use more costly large-time-scale transformer models. We also examine the effects of data augmentation and fine-tuning methods, accomplishingSoTA by a large margin when training only on the UCF-101 dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionData AugmentationRepresentation LearningSelf-Supervised LearningSimilar Papers 제목 키워드 기반
The Variational InfoMax Learning Objective
Bayesian Inference and Information Bottleneck are the two most popular objectives for neural networks, but they can be optimised only via a variational lower bound: the Variational Information Bottleneck (VIB). In this m…
Bayesian InferenceSmooth InfoMax -- Towards easier Post-Hoc interpretability
We introduce Smooth InfoMax (SIM), a novel method for self-supervised representation learning that incorporates an interpretability constraint into the learned representations at various depths of the neural network. SIM…
DecoderRepresentation LearningEstablishing Deep InfoMax as an effective self-supervised learning methodology in materials informatics
The scarcity of property labels remains a key challenge in materials informatics, whereas materials data without property labels are abundant in comparison. By pretraining supervised property prediction models on self-su…
Band GapFormation EnergyPredictionProperty Prediction+3Efficient Distribution Matching of Representations via Noise-Injected Deep InfoMax
Deep InfoMax (DIM) is a well-established method for self-supervised representation learning (SSRL) based on maximization of the mutual information between the input and the output of a deep neural network encoder. Despit…
DisentanglementRepresentation LearningDiscrete Infomax Codes for Supervised Representation Learning
Learning compact discrete representations of data is a key task on its own or for facilitating subsequent processing of data. In this paper we present a model that produces Discrete InfoMax Codes (DIMCO); we learn a prob…
Meta-LearningMetric LearningRepresentation LearningRetrieval