Divide and Contrast: Self-supervised Learning from Uncurated Data
Self-supervised learning holds promise in leveraging large amounts of unlabeled data, however much of its progress has thus far been limited to highly curated pre-training data such as ImageNet. We explore the effects of contrastive learning from larger, less-curated image datasets such as YFCC, and find there is indeed a large difference in the resulting representation quality. We hypothesize that this curation gap is due to a shift in the distribution of image classes -- which is more diverse and heavy-tailed -- resulting in less relevant negative samples to learn from. We test this hypothesis with a new approach, Divide and Contrast (DnC), which alternates between contrastive learning and clustering-based hard negative mining. When pretrained on less curated datasets, DnC greatly improves the performance of self-supervised learning on downstream tasks, while remaining competitive with the current state-of-the-art on curated datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringContrastive LearningSelf-Supervised Image ClassificationSelf-Supervised LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
The abundance and ease of utilizing sound, along with the fact that auditory clues reveal so much about what happens in the scene, make the audio-visual space a perfectly intuitive choice for self-supervised representati…
Contrastive LearningRepresentation LearningSelf-Supervised LearningVariational Self-Supervised Contrastive Learning Using Beta Divergence
Learning a discriminative semantic space using unlabelled and noisy data remains unaddressed in a multi-label setting. We present a contrastive self-supervised learning method which is robust to data noise, grounded in t…
Face RecognitionLinear evaluationSelf-Supervised LearningPretrained Encoders are All You Need
Data-efficiency and generalization are key challenges in deep learning and deep reinforcement learning as many models are trained on large-scale, domain-specific, and expensive-to-label datasets. Self-supervised models t…
AllContrastive LearningDeep Reinforcement Learningreinforcement-learning+2Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data
Remote sensing and automatic earth monitoring are key to solve global-scale challenges such as disaster prevention, land use monitoring, or tackling climate change. Although there exist vast amounts of remote sensing dat…
Change DetectionSelf-Supervised LearningTransfer LearningUnsupervised Pre-trainingDeploying self-supervised learning in the wild for hybrid automatic speech recognition
Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Event DetectionScheduling+3