Soft Contrastive Learning for Visual Localization
Localization by image retrieval is inexpensive and scalable due to simple mapping and matching techniques. Such localization, however, depends upon the quality of image features often obtained using Contrastive learning frameworks. Most contrastive learning strategies opt for features to distinguish different classes. In the context of localization, however, there is no natural definition of classes. Therefore, images are usually artificially separated into positive and negative classes, with respect to the chosen anchor images, based on some geometric proximity measure. In this paper, we show why such divisions are problematic for learning localization features. We argue that any artificial division based on some proximity measure is undesirable, due to the inherently ambiguous supervision for images near proximity threshold. To this end, we propose a novel technique that uses soft positive/negative assignments of images for contrastive learning, avoiding the aforementioned problem. Our soft assignment makes a gradual distinction between close and far images in both geometric and feature spaces. Experiments on four large-scale benchmark datasets demonstrate the superiority of the proposed soft contrastive learning over the state-of-the-art method for retrieval-based visual localization.
Code (1)
Tasks
Contrastive LearningImage RetrievalRetrievalVisual LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Semantic Pose Verification for Outdoor Visual Localization with Self-supervised Contrastive Learning
Any city-scale visual localization system has to overcome long-term appearance changes, such as varying illumination conditions or seasonal changes between query and database images. Since semantic content is more robust…
Contrastive LearningSemantic SimilaritySemantic Textual SimilarityVisual LocalizationProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization
Visual Geo-localization (VG) refers to the process to identify the location described in query images, which is widely applied in robotics field and computer vision tasks, such as autonomous driving, metaverse, augmented…
geo-localizationVisual Place RecognitionMarginNCE: Robust Sound Localization with a Negative Margin
The goal of this work is to localize sound sources in visual scenes with a self-supervised approach. Contrastive learning in the context of sound source localization leverages the natural correspondence between audio and…
Contrastive LearningSound Source LocalizationContrastive Self-Supervised Learning of Global-Local Audio-Visual Representations
Contrastive self-supervised learning has delivered impressive results in many audio-visual recognition tasks. However, existing approaches optimize for learning either global representations useful for high-level underst…
ClassificationDeepFake DetectionFace SwappingGeneral Classification+4GLIPv2: Unifying Localization and Vision-Language Understanding
We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2…
2D Object DetectionContrastive LearningImage CaptioningInstance Segmentation+10