paper-with-me

홈 › Papers

MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition

2026-05-10 · Zhengyi Xu, Yuhang Ming, Zhihao Zhan, Hanyu Zhu, Javier Civera, Wanzeng Kong arxiv

Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spatial-structure discrepancies between ground observations and aerial references. To address this challenge, we present MAG-VLAQ, a foundation-model-enhanced query aggregation framework for multi-modal aerial-ground cross-view place recognition. Specifically, our approach leverages pre-trained foundation models to extract dense visual tokens from both ground and aerial images, as well as expressive geometric tokens from ground LiDAR observations. These heterogeneous tokens are then projected into a shared embedding space for cross-modal alignment and fusion. As our main contribution, we propose ODE-conditioned VLAQ, which tightly couples neural ordinary differential equations (ODE)-based RGB-LiDAR fusion with vectors of locally aggregated queries (VLAQ). In this design, the VLAQ query centers are dynamically adapted according to the fused multi-modal state. This mechanism allows the final global descriptor to preserve globally learned retrieval prototypes while remaining responsive to scene-specific visual and geometric evidence, significantly improving aerial-ground matching. Extensive experiments on KITTI360-AG and nuScenes-AG validate the effectiveness of our proposed MAG-VLAQ. Notably, on KITTI360-AG, our MAG-VLAQ nearly doubles the state-of-the-art performance, achieving 61.1 Recall@1 in the satellite setting, compared with 34.5 from the closest competing approach.

📄 PDF Abstract BibTeX arXiv:2605.09418

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition

2026-01-19 · Hanyu Zhu, Zhihao Zhan, Yuhang Ming, Liang Li 외 arxiv

One of the central challenges in visual place recognition (VPR) is learning a robust global representation that remains discriminative under large viewpoint changes, illumination variations, and severe domain shifts. Whi…

Visual Place Recognition

Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEs

2025-01-01 · CVPR 2025 1 · Sijie Wang, Rui She, Qiyu Kang, Siqi Li 외

Place recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal …

Autonomous Driving

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

2026-07-09 · Haibin Tian, Huichao Xie, Xuelin Qian, Ruitao Lu 외 arxiv

Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) b…

Referring ExpressionVisual Grounding

Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark

2025-03-10 · Jiahao Wang, Xiangyu Cao, Jiaru Zhong, Yuner Zhang 외

Despite significant advancements, autonomous driving systems continue to struggle with occluded objects and long-range detection due to the inherent limitations of single-perspective sensing. Aerial-ground cooperation of…

Autonomous DrivingBenchmarking

Evaluation of Cross-View Matching to Improve Ground Vehicle Localization with Aerial Perception

2020-03-13 · Deeksha Dixit, Surabhi Verma, Pratap Tokekar

Cross-view matching refers to the problem of finding the closest match for a given query ground view image to one from a database of aerial images. If the aerial images are geotagged, then the closest matching aerial ima…