paper-with-me

홈 › Papers

Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEs

2025-01-01 · CVPR 2025 1 · Sijie Wang, Rui She, Qiyu Kang, Siqi Li, Disheng Li, Tianyu Geng, Shangshu Yu, Wee Peng Tay

Place recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal PR works have mainly focused on the same-view scenario in which ground-view descriptors are matched with a database of ground-view descriptors during inference, the multi-modal cross-view scenario, in which ground-view descriptors are matched with aerial-view descriptors in a database, remains under-explored. We propose AGPlace, a model that effectively integrates information from multi-modal ground sensors (cameras and LiDARs) to achieve accurate aerial-ground PR. AGPlace achieves effective aerial-ground cross-view PR by leveraging a manifold-based neural ordinary differential equation (ODE) framework with a multi-domain alignment loss. It outperforms existing state-of-the-art cross-view PR models on large-scale datasets. As most existing PR models are designed for ground-ground PR, we adapt these baselines into our cross-view pipeline. Experiments demonstrate that this direct adaptation performs worse than our overall model architecture AGPlace. AGPlace represents a significant advancement in multi-modal aerial-ground PR, with promising implications for real-world applications.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Cross-View Policy Learning for Street Navigation

2019-06-13 · ICCV 2019 10 · Ang Li, Huiyi Hu, Piotr Mirowski, Mehrdad Farajtabar

The ability to navigate from visual observations in unfamiliar environments is a core component of intelligent agents and an ongoing challenge for Deep Reinforcement Learning (RL). Street View can be a sensible testbed f…

Deep Reinforcement LearningNavigateReinforcement LearningReinforcement Learning (RL)+1

MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition

2026-05-10 · Zhengyi Xu, Yuhang Ming, Zhihao Zhan, Hanyu Zhu 외 arxiv

Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spatial-structure discrepancies between ground observations and aerial r…

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

2025-12-11 · Phu Pham, Damon Conover, Aniket Bera arxiv

Aerial-ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with a…

Contrastive Learning

Cross-View Meets Diffusion: Aerial Image Synthesis with Geometry and Text Guidance

2024-08-08 · Ahmad Arrabi, Xiaohan Zhang, Waqas Sultani, Chen Chen 외

Aerial imagery analysis is critical for many research fields. However, obtaining frequent high-quality aerial images is not always accessible due to its high effort and cost requirements. One solution is to use the Groun…

BEV SegmentationData Augmentationgeo-localizationImage Generation

AG-VPReID.VIR: Bridging Aerial and Ground Platforms for Video-based Visible-Infrared Person Re-ID

2025-07-24 · Huy Nguyen, Kien Nguyen, Akila Pemasiri, Akmal Jahan 외 arxiv

Person re-identification (Re-ID) across visible and infrared modalities is crucial for 24-hour surveillance systems, but existing datasets primarily focus on ground-level perspectives. While ground-based IR systems offer…

Person Re-Identification