paper-with-me

Papers

Satellite-Free Training for Drone-View Geo-Localization

2026-04-02 · Tao Liu, Yingzhi Zhang, Kan Ren, Xiaoqi Zhao arxiv

Drone-view geo-localization (DVGL) aims to determine the location of drones in GPS-denied environments by retrieving the corresponding geotagged satellite tile from a reference gallery given UAV observations of a location. In many existing formulations, these observations are represented by a single oblique UAV image. In contrast, our satellite-free setting is designed for multi-view UAV sequences, which are used to construct a geometry-normalized UAV-side location representation before cross-view retrieval. Existing approaches rely on satellite imagery during training, either through paired supervision or unsupervised alignment, which limits practical deployment when satellite data are unavailable or restricted. In this paper, we propose a satellite-free training (SFT) framework that converts drone imagery into cross-view compatible representations through three main stages: drone-side 3D scene reconstruction, geometry-based pseudo-orthophoto generation, and satellite-free feature aggregation for retrieval. Specifically, we first reconstruct dense 3D scenes from multi-view drone images using 3D Gaussian splatting and project the reconstructed geometry into pseudo-orthophotos via PCA-guided orthographic projection. This rendering stage operates directly on reconstructed scene geometry without requiring camera parameters at rendering time. Next, we refine these orthophotos with lightweight geometry-guided inpainting to obtain texture-complete drone-side views. Finally, we extract DINOv3 patch features from the generated orthophotos, learn a Fisher vector aggregation model solely from drone data, and reuse it at test time to encode satellite tiles for cross-view retrieval. Experimental results on University-1652 and SUES-200 show that our SFT framework substantially outperforms satellite-free generalization baselines and narrows the gap to methods trained with satellite imagery.

📄 PDF Abstract BibTeX arXiv:2604.01581

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization

2020-02-27 · Zhedong Zheng, Yunchao Wei, Yi Yang

We consider the problem of cross-view geo-localization. The primary challenge of this task is to learn the robust feature against large viewpoint changes. Existing benchmarks can help, but are limited in the number of vi…

Drone navigationDrone-view target localizationgeo-localizationImage-Based Localization+1

DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization

2025-11-09 · Tao Liu, Kan Ren, Qian Chen arxiv

With the rapid growth of the low-altitude economy, unmanned aerial vehicles (UAVs) have become key platforms for measurement and tracking in intelligent patrol systems. However, in GNSS-denied environments, localization …

Image Retrieval

Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval

2022-05-22 · Zelong Zeng, Zheng Wang, Fan Yang, Shin'ichi Satoh

The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view …

geo-localizationImage RetrievalRepresentation LearningRetrieval

SMGeo: Cross-View Object Geo-Localization with Grid-Level Mixture-of-Experts

2025-11-18 · Fan Zhang, Haoyuan Ren, Fei Ma, Qiang Yin 외 arxiv

Cross-view object Geo-localization aims to precisely pinpoint the same object across large-scale satellite imagery based on drone images. Due to significant differences in viewpoint and scale, coupled with complex backgr…

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

2026-06-29 · Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi 외 arxiv

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance…