Visual Localization Using Semantic Segmentation and Depth Prediction
In this paper, we propose a monocular visual localization pipeline leveraging semantic and depth cues. We apply semantic consistency evaluation to rank the image retrieval results and a practical clustering technique to reject estimation outliers. In addition, we demonstrate a substantial performance boost achieved with a combination of multiple feature extractors. Furthermore, by using depth prediction with a deep neural network, we show that a significant amount of falsely matched keypoints are identified and eliminated. The proposed pipeline outperforms most of the existing approaches at the Long-Term Visual Localization benchmark 2020.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDepth EstimationDepth PredictionImage RetrievalPredictionReal-Time Semantic SegmentationRetrievalSemantic SegmentationVisual LocalizationSimilar Papers 제목 키워드 기반
Domain Adaptive Semantic Segmentation with Self-Supervised Depth Estimation
Domain adaptation for semantic segmentation aims to improve the model performance in the presence of a distribution shift between source and target domain. Leveraging the supervision from auxiliary tasks~(such as depth e…
Depth EstimationDepth PredictionDomain AdaptationFeature Correlation+3Learning Geocentric Object Pose in Oblique Monocular Images
An object's geocentric pose, defined as the height above ground and orientation with respect to gravity, is a powerful representation of real-world structure for object detection, segmentation, and localization tasks usi…
Earth ObservationObjectobject-detectionObject Detection+2Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization
Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing m…
3D Semantic SegmentationScene UnderstandingAG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
Large multimodal models (LMMs) exhibit strong task generalization capabilities, offering new opportunities for zero-shot visual anomaly segmentation (ZSAS). However, existing LMM-based segmentation approaches still face …
How semantic and geometric information mutually reinforce each other in ToF object localization
We propose a novel approach to localize a 3D object from the intensity and depth information images provided by a Time-of-Flight (ToF) sensor. Our method uses two CNNs. The first one uses raw depth and intensity images a…
ObjectObject Localization