VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps
This paper presents Vision-Language Global Localization (VLG-Loc), a novel global localization method that uses human-readable labeled footprint maps containing only names and areas of distinctive visual landmarks in an environment. While humans naturally localize themselves using such maps, translating this capability to robotic systems remains highly challenging due to the difficulty of establishing correspondences between observed landmarks and those in the map without geometric and appearance details. To address this challenge, VLG-Loc leverages a vision-language model (VLM) to search the robot's multi-directional image observations for the landmarks noted in the map. The method then identifies robot poses within a Monte Carlo localization framework, where the found landmarks are used to evaluate the likelihood of each pose hypothesis. Experimental validation in simulated and real-world retail environments demonstrates superior robustness compared to existing scan-based methods, particularly under environmental changes. Further improvements are achieved through the probabilistic fusion of visual and scan-based localization.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Exploring Representation Learning for Small-Footprint Keyword Spotting
In this paper, we investigate representation learning for low-resource keyword spotting (KWS). The main challenges of KWS are limited labeled data and limited available device resources. To address those challenges, we e…
Contrastive LearningKeyword SpottingRepresentation LearningSmall-Footprint Keyword SpottingConstructing Indoor Region-based Radio Map without Location Labels
Radio map construction requires a large amount of radio measurement data with location labels, which imposes a high deployment cost. This paper develops a region-based radio map from received signal strength (RSS) measur…
ClusteringCross-Domain Generalization of Multimodal LLMs for Global Photovoltaic Assessment
The rapid expansion of distributed photovoltaic (PV) systems poses challenges for power grid management, as many installations remain undocumented. While satellite imagery provides global coverage, traditional computer v…
Domain GeneralizationImproving Small Footprint Few-shot Keyword Spotting with Supervision on Auxiliary Data
Few-shot keyword spotting (FS-KWS) models usually require large-scale annotated datasets to generalize to unseen target keywords. However, existing KWS datasets are limited in scale and gathering keyword-like labeled dat…
Keyword SpottingMulti-Task LearningSelf-Supervised LearningCLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps
This paper describes a multi-modal data association method for global localization using object-based maps and camera images. In global localization, or relocalization, using object-based maps, existing methods typically…
Language ModelingLanguage ModellingObject