paper-with-me

홈 › Papers

C-BEV: Contrastive Bird's Eye View Training for Cross-View Image Retrieval and 3-DoF Pose Estimation

2023-12-13 · Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen

To find the geolocation of a street-view image, cross-view geolocalization (CVGL) methods typically perform image retrieval on a database of georeferenced aerial images and determine the location from the visually most similar match. Recent approaches focus mainly on settings where street-view and aerial images are preselected to align w.r.t. translation or orientation, but struggle in challenging real-world scenarios where varying camera poses have to be matched to the same aerial image. We propose a novel trainable retrieval architecture that uses bird's eye view (BEV) maps rather than vectors as embedding representation, and explicitly addresses the many-to-one ambiguity that arises in real-world scenarios. The BEV-based retrieval is trained using the same contrastive setting and loss as classical retrieval. Our method C-BEV surpasses the state-of-the-art on the retrieval task on multiple datasets by a large margin. It is particularly effective in challenging many-to-one scenarios, e.g. increasing the top-1 recall on VIGOR's cross-area split with unknown orientation from 31.1% to 65.0%. Although the model is supervised only through a contrastive objective applied on image pairings, it additionally learns to infer the 3-DoF camera pose on the matching aerial image, and even yields a lower mean pose error than recent methods that are explicitly trained with metric groundtruth.

📄 PDF Abstract BibTeX arXiv:2312.08060

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalPose EstimationRetrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and Mapping

2023-10-29 · Srikumar Sastry, Subash Khanal, Aayush Dhakal, Di Huang 외

We propose a metadata-aware self-supervised learning~(SSL)~framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework unifies two SSL strategies: Contrastive L…

Contrastive LearningCross-Modal RetrievalFine-Grained Image ClassificationRetrieval+2

BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation

2023-08-12 · ICCV 2023 1 · Miaoyu Li, Yachao Zhang, Xu Ma, Yanyun Qu 외

Cross-modal Unsupervised Domain Adaptation (UDA) aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in a new domain. However, UDA methods rely on access to the target domain during train…

3D Semantic SegmentationContrastive LearningDomain AdaptationDomain Generalization+2

BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning

2025-08-06 · Ziyang Leng, Jiawei Yang, Zhicheng Ren, Bolei Zhou arxiv

We present BEVCon, a simple yet effective contrastive learning framework designed to improve Bird's Eye View (BEV) perception in autonomous driving. BEV perception offers a top-down-view representation of the surrounding…

Representation LearningTrajectory PredictionContrastive Learning3D Object Detection

Generative Adversarial Frontal View to Bird View Synthesis

2018-08-01 · Xinge Zhu, Zhichao Yin, Jianping Shi, Hongsheng Li 외

Environment perception is an important task with great practical value and bird view is an essential part for creating panoramas of surrounding environment. Due to the large gap and severe deformation between the frontal…

Bird View SynthesisHomography EstimationTranslation

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

2025-12-11 · Phu Pham, Damon Conover, Aniket Bera arxiv

Aerial-ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with a…

Contrastive Learning