paper-with-me

Papers

Geo$^\textbf{2}$: Geometry-Guided Cross-view Geo-Localization and Image Synthesis

2026-03-26 · Yancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu, Safwan Wshah, Chen Chen arxiv

Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and aerial views. Recent Geometric Foundation Models (GFMs) have demonstrated strong capabilities in extracting generalizable 3D geometric features from images, but their potential in cross-view geo-spatial tasks remains underexplored. In this work, we present Geo^2, a unified framework that leverages Geometric priors from GFMs (e.g., VGGT) to jointly perform geo-spatial tasks, CVGL and bidirectional CVIS. Despite the 3D reconstruction ability of GFMs, directly applying them to CVGL and CVIS remains challenging due to the large viewpoint gap between ground and aerial imagery. We propose GeoMap, which embeds ground and aerial features into a shared 3D-aware latent space, effectively reducing cross-view discrepancies for localization. This shared latent space naturally bridges cross-view image synthesis in both directions. To exploit this, we propose GeoFlow, a flow-matching model conditioned on geometry-aware latent embeddings. We further introduce a consistency loss to enforce latent alignment between the two synthesis directions, ensuring bidirectional coherence. Extensive experiments on standard benchmarks, including CVUSA, CVACT, and VIGOR, demonstrate that Geo^2 achieves state-of-the-art performance in both localization and synthesis, highlighting the effectiveness of 3D geometric priors for cross-view geo-spatial learning.

📄 PDF Abstract BibTeX arXiv:2603.25819

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction

Similar Papers 제목 키워드 기반

Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View Transformer

2023-07-16 · ICCV 2023 1 · Yujiao Shi, Fei Wu, Akhil Perincherry, Ankit Vora 외

Image retrieval-based cross-view localization methods often lead to very coarse camera pose estimation, due to the limited sampling density of the database satellite images. In this paper, we propose a method to increase…

Camera LocalizationCamera Pose EstimationImage RetrievalPose Estimation+1

GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization

2025-07-15 · Shaowen Tong, Zimin Xia, Alexandre Alahi, Xuming He 외

Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with satellite images, is crucial for large-scale outdoor applications like autonomous navigati…

Autonomous Navigation

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

2026-08-04 · Wanhao Liu, Jinsong Lin, Rulin Zhou, Chi Kit Ng 외 arxiv

Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mot…

Video PredictionVideo Generation

Satellite-Free Training for Drone-View Geo-Localization

2026-04-02 · Tao Liu, Yingzhi Zhang, Kan Ren, Xiaoqi Zhao arxiv

Drone-view geo-localization (DVGL) aims to determine the location of drones in GPS-denied environments by retrieving the corresponding geotagged satellite tile from a reference gallery given UAV observations of a locatio…

SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

2026-09-04 · Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao 외 arxiv

Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scan…

Unsupervised Domain Adaptation3D Object Detection