Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations
Recent progress in self-supervised learning has demonstrated promising results in multiple visual tasks. An important ingredient in high-performing self-supervised methods is the use of data augmentation by training models to place different augmented views of the same image nearby in embedding space. However, commonly used augmentation pipelines treat images holistically, ignoring the semantic relevance of parts of an image-e.g. a subject vs. a background-which can lead to the learning of spurious correlations. Our work addresses this problem by investigating a class of simple, yet highly effective "background augmentations", which encourage models to focus on semantically-relevant content by discouraging them from focusing on image backgrounds. Through a systematic investigation, we show that background augmentations lead to substantial improvements in performance across a spectrum of state-of-the-art self-supervised methods (MoCo-v2, BYOL, SwAV) on a variety of tasks, e.g. $\sim$+1-2% gains on ImageNet, enabling performance on par with the supervised baseline. Further, we find the improvement in limited-labels settings is even larger (up to 4.2%). Background augmentations also improve robustness to a number of distribution shifts, including natural adversarial examples, ImageNet-9, adversarial attacks, ImageNet-Renditions. We also make progress in completely unsupervised saliency detection, in the process of generating saliency masks used for background augmentations.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningData AugmentationImage ClassificationRepresentation LearningSaliency DetectionSelf-Supervised LearningUnsupervised Saliency DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Evaluating The Robustness of Self-Supervised Representations to Background/Foreground Removal
Despite impressive empirical advances of SSL in solving various tasks, the problem of understanding and characterizing SSL representations learned from input data remains relatively under-explored. We provide a comparati…
image-classificationImage ClassificationLearning Background Invariance Improves Generalization and Robustness in Self-Supervised Learning on ImageNet and Beyond
Recent progress in self-supervised learning has demonstrated promising results in multiple visual tasks. An important ingredient in high-performing self-supervised methods is the use of data augmentation by training mode…
Data AugmentationSaliency DetectionSelf-Supervised LearningUnsupervised Saliency DetectionCharacterizing the adversarial vulnerability of speech self-supervised learning
A leaderboard named Speech processing Universal PERformance Benchmark (SUPERB), which aims at benchmarking the performance of a shared self-supervised learning (SSL) speech model across various downstream speech tasks wi…
Adversarial RobustnessBenchmarkingRepresentation LearningSelf-Supervised Learning+1Self-Supervised Dynamic Networks for Covariate Shift Robustness
As supervised learning still dominates most AI applications, test-time performance is often unexpected. Specifically, a shift of the input covariates, caused by typical nuisances like background-noise, illumination varia…
image-classificationImage ClassificationJoint Self-Supervised Learning for Vision-based Reinforcement Learning
Vision-based reinforcement learning requires efficient and robust representations of image-based observations, especially when the images contain distracting (task-irrelevant) elements such as shadows, clouds, and light.…
Autonomous Drivingcontinuous-controlContinuous Controlreinforcement-learning+3