paper-with-me

홈 › Papers

Cross-View Image Retrieval -- Ground to Aerial Image Retrieval through Deep Learning

2020-05-02 · Numan Khurshid, Talha Hanif, Mohbat Tharani, Murtaza Taj

Cross-modal retrieval aims to measure the content similarity between different types of data. The idea has been previously applied to visual, text, and speech data. In this paper, we present a novel cross-modal retrieval method specifically for multi-view images, called Cross-view Image Retrieval CVIR. Our approach aims to find a feature space as well as an embedding space in which samples from street-view images are compared directly to satellite-view images (and vice-versa). For this comparison, a novel deep metric learning based solution "DeepCVIR" has been proposed. Previous cross-view image datasets are deficient in that they (1) lack class information; (2) were originally collected for cross-view image geolocalization task with coupled images; (3) do not include any images from off-street locations. To train, compare, and evaluate the performance of cross-view image retrieval, we present a new 6 class cross-view image dataset termed as CrossViewRet which comprises of images including freeway, mountain, palace, river, ship, and stadium with 700 high-resolution dual-view images for each class. Results show that the proposed DeepCVIR outperforms conventional matching approaches on the CVIR task for the given dataset and would also serve as the baseline for future research.

📄 PDF Abstract BibTeX arXiv:2005.00725

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage RetrievalMetric LearningRetrieval

Similar Papers 제목 키워드 기반

Text-based Aerial-Ground Person Retrieval

2025-11-11 · Xinyu Zhou, Yu Wu, Jiayao Ma, Wenhao Wang 외 arxiv

This work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Ret…

Person RetrievalText Generation

GeoCapsNet: Aerial to Ground view Image Geo-localization using Capsule Network

2019-04-12 · Bin Sun, Chen Chen, Yingying Zhu, Jianmin Jiang

The task of cross-view image geo-localization aims to determine the geo-location (GPS coordinates) of a query ground-view image by matching it with the GPS-tagged aerial (satellite) images in a reference dataset. Due to …

geo-localizationImage RetrievalRetrievalTriplet

C-BEV: Contrastive Bird's Eye View Training for Cross-View Image Retrieval and 3-DoF Pose Estimation

2023-12-13 · Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens 외

To find the geolocation of a street-view image, cross-view geolocalization (CVGL) methods typically perform image retrieval on a database of georeferenced aerial images and determine the location from the visually most s…

Image RetrievalPose EstimationRetrieval

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

2026-03-21 · Yifei Deng, Chenglong Li, Yuyang Zhang, Guyue Hu 외 arxiv

Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. Compared to ground-view text--image perso…

Person RetrievalText Generation

CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation

2026-06-03 · Yurim Jeon, Dongseong Seo, Seung-Woo Seo arxiv

Cross-view geo-localization estimates the geographic location of a ground image by matching it against an aerial image database. Existing methods tackle this through either large-scale retrieval or precise pose estimatio…

Pose Estimation