VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text-point cloud correspondence without effective spatial reasoning, limiting their accuracy in complex environments. To address this limitation, we propose VLM-Loc, a framework that leverages the spatial reasoning capability of large vision-language models (VLMs) for T2P localization. Specifically, we transform point clouds into bird's-eye-view (BEV) images and scene graphs that jointly encode geometric and semantic context, providing structured inputs for the VLM to learn cross-modal representations bridging linguistic and spatial semantics. On top of these representations, we introduce a partial node assignment mechanism that explicitly associates textual cues with scene graph nodes, enabling interpretable spatial reasoning for accurate localization. To facilitate systematic evaluation across diverse scenes, we present CityLoc, a benchmark built from multi-source point clouds for fine-grained T2P localization. Experiments on CityLoc demonstrate VLM-Loc achieves superior accuracy and robustness compared to state-of-the-art methods. Our code, model, and dataset are available at \href{https://github.com/MCG-NKU/nku-3d-vision}{repository}.
Code (0)
등록된 구현이 없습니다.
Tasks
Spatial ReasoningPoint CloudsSimilar Papers 제목 키워드 기반
FlexCloud: Direct, Modular Georeferencing and Drift-Correction of Point Cloud Maps
Current software stacks for real-world applications of autonomous driving leverage map information to ensure reliable localization, path planning, and motion prediction. An important field of research is the generation o…
Autonomous Drivingmotion predictionSimultaneous Localization and MappingIncorporating GNSS Information with LIDAR-Inertial Odometry for Accurate Land-Vehicle Localization
Currently, visual odometry and LIDAR odometry are performing well in pose estimation in some typical environments, but they still cannot recover the localization state at high speed or reduce accumulated drifts. In order…
Pose EstimationVisual OdometryMonocular Visual Place Recognition in LiDAR Maps via Cross-Modal State Space Model and Multi-View Matching
Achieving monocular camera localization within pre-built LiDAR maps can bypass the simultaneous mapping process of visual SLAM systems, potentially reducing the computational overhead of autonomous localization. To this …
Camera LocalizationContrastive LearningCross-modal place recognitionVisual Place RecognitionSparse 3D Point-cloud Map Upsampling and Noise Removal as a vSLAM Post-processing Step: Experimental Evaluation
The monocular vision-based simultaneous localization and mapping (vSLAM) is one of the most challenging problem in mobile robotics and computer vision. In this work we study the post-processing techniques applied to spar…
Simultaneous Localization and MappingDepth-Guided Privacy-Preserving Visual Localization Using 3D Sphere Clouds
The emergence of deep neural networks capable of revealing high-fidelity scene details from sparse 3D point clouds has raised significant privacy concerns in visual localization involving private maps. Lifting map points…
Camera Pose EstimationVisual LocalizationPoint Clouds