Rethinking Visual Geo-localization for Large-Scale Applications
Visual Geo-localization (VG) is the task of estimating the position where a given photo was taken by comparing it with a large database of images of known locations. To investigate how existing techniques would perform on a real-world city-wide VG application, we build San Francisco eXtra Large, a new dataset covering a whole city and providing a wide range of challenging cases, with a size 30x bigger than the previous largest dataset for visual geo-localization. We find that current methods fail to scale to such large datasets, therefore we design a new highly scalable training technique, called CosPlace, which casts the training as a classification problem avoiding the expensive mining needed by the commonly used contrastive learning. We achieve state-of-the-art performance on a wide range of datasets and find that CosPlace is robust to heavy domain changes. Moreover, we show that, compared to the previous state-of-the-art, CosPlace requires roughly 80% less GPU memory at train time, and it achieves better results with 8x smaller descriptors, paving the way for city-wide real-world visual geo-localization. Dataset, code and trained models are available for research purposes at https://github.com/gmberton/CosPlace.
Code (2)
Tasks
Contrastive Learninggeo-localizationGPUImage ClassificationImage RetrievalVisual Place RecognitionSimilar Papers 제목 키워드 기반
Accurate Visual Localization for Automotive Applications
Accurate vehicle localization is a crucial step towards building effective Vehicle-to-Vehicle networks and automotive applications. Yet standard grade GPS data, such as that provided by mobile phones, is often noisy and …
RetrievalVisual LocalizationRSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and high semantic ambiguity of remote sensing s…
Spatial ReasoningVisual GroundingRenderNet: Visual Relocalization Using Virtual Viewpoints in Large-Scale Indoor Environments
Visual relocalization has been a widely discussed problem in 3D vision: given a pre-constructed 3D visual map, the 6 DoF (Degrees-of-Freedom) pose of a query image is estimated. Relocalization in large-scale indoor envir…
Image RetrievalRetrievalRobot NavigationFrom Coarse to Fine: Robust Hierarchical Localization at Large Scale
Robust and accurate visual localization is a fundamental capability for numerous applications, such as autonomous driving, mobile robotics, or augmented reality. It remains, however, a challenging task, particularly for …
Autonomous DrivingRetrievalVisual LocalizationVisual Place RecognitionRethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
3D Visual Grounding (3DVG) and 3D Dense Captioning (3DDC) are two crucial tasks in various 3D applications, which require both shared and complementary information in localization and visual-language relationships. There…
3D dense captioning3D visual groundingDense CaptioningVisual Grounding