paper-with-me

홈 › Papers

Pix2Map: Cross-modal Retrieval for Inferring Street Maps from Images

2023-01-10 · CVPR 2023 1 · Xindi Wu, KwunFung Lau, Francesco Ferroni, Aljoša Ošep, Deva Ramanan

Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a complex urban road topology directly from raw image data. The main insight of this paper is that this problem can be posed as cross-modal retrieval by learning a joint, cross-modal embedding space for images and existing maps, represented as discrete graphs that encode the topological layout of the visual surroundings. We conduct our experimental evaluation using the Argoverse dataset and show that it is indeed possible to accurately retrieve street maps corresponding to both seen and unseen roads solely from image data. Moreover, we show that our retrieved maps can be used to update or expand existing maps and even show proof-of-concept results for visual localization and image retrieval from spatial graphs.

📄 PDF Abstract BibTeX arXiv:2301.04224

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous NavigationCross-Modal RetrievalImage RetrievalRetrievalVisual Localization

Similar Papers 제목 키워드 기반

Inferring and Improving Street Maps with Data-Driven Automation

2019-10-02 · Favyen Bastani, Songtao He, Satvat Jagwani, Edward Park 외

Street maps are a crucial data source that help to inform a wide range of decisions, from navigating a city to disaster relief and urban planning. However, in many parts of the world, street maps are incomplete or lag be…

A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

2023-05-05 · Weijia Wu, Yuzhong Zhao, Zhuang Li, Jiahong Li 외

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to un…

Reading ComprehensionRetrievalSentenceVideo Retrieval

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

2026-06-08 · Quang Long Ho Ngo, Zimin Xia, Alexandre Alahi arxiv

Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (e.g., OpenStreetMap tiles), these approaches often lag in performanc…

Cross-View Image Retrieval -- Ground to Aerial Image Retrieval through Deep Learning

2020-05-02 · Numan Khurshid, Talha Hanif, Mohbat Tharani, Murtaza Taj

Cross-modal retrieval aims to measure the content similarity between different types of data. The idea has been previously applied to visual, text, and speech data. In this paper, we present a novel cross-modal retrieval…

Cross-Modal RetrievalImage RetrievalMetric LearningRetrieval

Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming

2026-03-26 · Yunus Talha Erzurumlu, Jiyong Kwag, Alper Yilmaz arxiv

Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation. Existing methods almost universally fo…

Spatial Reasoning