Coarse Is Better? A New Pipeline Towards Self-Supervised Learning with Uncurated Images
Most self-supervised learning (SSL) methods often work on curated datasets where the object-centric assumption holds. This assumption breaks down in uncurated images. Existing scene image SSL methods try to find the two views from original scene images that are well matched or dense, which is both complex and computationally heavy. This paper proposes a conceptually different pipeline: first find regions that are coarse objects (with adequate objectness), crop them out as pseudo object-centric images, then any SSL method can be directly applied as in a real object-centric dataset. That is, coarse crops benefits scene images SSL. A novel cropping strategy that produces coarse object box is proposed. The new pipeline and cropping strategy successfully learn quality features from uncurated datasets without ImageNet. Experiments show that our pipeline outperforms existing SSL methods (MoCo-v2, DenseCL and MAE) on classification, detection and segmentation tasks. We further conduct extensively ablations to verify that: 1) the pipeline do not rely on pretrained models; 2) the cropping strategy is better than existing object discovery methods; 3) our method is not sensitive to hyperparameters and data augmentations.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectObject DiscoverySelf-Supervised LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data
Remote sensing and automatic earth monitoring are key to solve global-scale challenges such as disaster prevention, land use monitoring, or tackling climate change. Although there exist vast amounts of remote sensing dat…
Change DetectionSelf-Supervised LearningTransfer LearningUnsupervised Pre-trainingDeploying self-supervised learning in the wild for hybrid automatic speech recognition
Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Event DetectionScheduling+3RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data
Semi-supervised learning aims to train a model using limited labels. State-of-the-art semi-supervised methods for image classification such as PAWS rely on self-supervised representations learned with large-scale unlabel…
Density Estimationimage-classificationImage ClassificationRepresentation LearningFine-grained Multi-Modal Self-Supervised Learning
Multi-Modal Self-Supervised Learning from videos has been shown to improve model's performance on various downstream tasks. However, such Self-Supervised pre-training requires large batch sizes and a large amount of comp…
Action RecognitionSelf-Supervised LearningSEPT: Towards Scalable and Efficient Visual Pre-Training
Recently, the self-supervised pre-training paradigm has shown great potential in leveraging large-scale unlabeled data to improve downstream task performance. However, increasing the scale of unlabeled pre-training data …
Retrieval