Exploring Localization for Self-supervised Fine-grained Contrastive Learning
Self-supervised contrastive learning has demonstrated great potential in learning visual representations. Despite their success in various downstream tasks such as image classification and object detection, self-supervised pre-training for fine-grained scenarios is not fully explored. We point out that current contrastive methods are prone to memorizing background/foreground texture and therefore have a limitation in localizing the foreground object. Analysis suggests that learning to extract discriminative texture information and localization are equally crucial for fine-grained self-supervised pre-training. Based on our findings, we introduce cross-view saliency alignment (CVSA), a contrastive learning framework that first crops and swaps saliency regions of images as a novel view generation and then guides the model to localize on foreground objects via a cross-view alignment loss. Extensive experiments on both small- and large-scale fine-grained classification benchmarks show that CVSA significantly improves the learned representation.
Code (1)
Tasks
Contrastive LearningFine-Grained Image ClassificationFine-Grained Image Recognitionimage-classificationImage ClassificationObject DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Joint-task Self-supervised Learning for Temporal Correspondence
This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grai…
Object TrackingSelf-Supervised LearningSemi-Supervised Video Object SegmentationUnsupervised Video Object SegmentationFine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual Localization
Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice, for example, encountered in autonom…
Autonomous DrivingSegmentationVisual LocalizationFine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual Localization
Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice that is, for example, encountered in…
Autonomous DrivingSegmentationVisual LocalizationSelf-supervised Spatiotemporal Representation Learning by Exploiting Video Continuity
Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored…
Action LocalizationAction RecognitionRepresentation LearningRetrieval+1LayerCAM: Exploring Hierarchical Class Activation Maps for Localization
The class activation maps are generated from the final convolutional layer of CNN. They can highlight discriminative object regions for the class of interest. These discovered object regions have been widely used for wea…
ObjectObject LocalizationSemantic SegmentationWeakly-Supervised Object Localization