paper-with-me

홈 › Papers

AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis

2025-04-17 · CVPR 2025 1 · Khiem Vuong, Anurag Ghosh, Deva Ramanan, Srinivasa Narasimhan, Shubham Tulsiani

We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-registered aerial-ground datasets for training is a key reason for this failure. Such data is difficult to assemble precisely because it is difficult to reconstruct in a scalable way. To overcome this challenge, we propose a scalable framework combining pseudo-synthetic renderings from 3D city-wide meshes (e.g., Google Earth) with real, ground-level crowd-sourced images (e.g., MegaDepth). The pseudo-synthetic data simulates a wide range of aerial viewpoints, while the real, crowd-sourced images help improve visual fidelity for ground-level images where mesh-based renderings lack sufficient detail, effectively bridging the domain gap between real images and pseudo-synthetic renderings. Using this hybrid dataset, we fine-tune several state-of-the-art algorithms and achieve significant improvements on real-world, zero-shot aerial-ground tasks. For example, we observe that baseline DUSt3R localizes fewer than 5% of aerial-ground pairs within 5 degrees of camera rotation error, while fine-tuning with our data raises accuracy to nearly 56%, addressing a major failure point in handling large viewpoint changes. Beyond camera estimation and scene reconstruction, our dataset also improves performance on downstream tasks like novel-view synthesis in challenging aerial-ground scenarios, demonstrating the practical value of our approach in real-world applications.

📄 PDF Abstract BibTeX arXiv:2504.13157

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View Synthesis

Similar Papers 제목 키워드 기반

Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty

2024-08-27 · Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen 외

Robust and realistic rendering for large-scale road scenes is essential in autonomous driving simulation. Recently, 3D Gaussian Splatting (3D-GS) has made groundbreaking progress in neural rendering, but the general fide…

Autonomous DrivingNeural RenderingNovel View Synthesis

Geo$^\textbf{2}$: Geometry-Guided Cross-view Geo-Localization and Image Synthesis

2026-03-26 · Yancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu 외 arxiv

Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and a…

3D Reconstruction

SkyDiffusion: Ground-to-Aerial Image Synthesis with Diffusion Models and BEV Paradigm

2024-08-03 · Junyan Ye, Jun He, Weijia Li, Zhutao Lv 외

Ground-to-aerial image synthesis focuses on generating realistic aerial images from corresponding ground street view images while maintaining consistent content layout, simulating a top-down view. The significant viewpoi…

Image GenerationSSIM

3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification

2026-04-29 · William Grolleau, Astrid Sabourin, Guillaume Lapouge, Catherine Achard arxiv

Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminative features, making cross-viewpoint image retrieval challenging. While…

Representation LearningNovel View SynthesisImage Retrieval

Cross-View Meets Diffusion: Aerial Image Synthesis with Geometry and Text Guidance

2024-08-08 · Ahmad Arrabi, Xiaohan Zhang, Waqas Sultani, Chen Chen 외

Aerial imagery analysis is critical for many research fields. However, obtaining frequent high-quality aerial images is not always accessible due to its high effort and cost requirements. One solution is to use the Groun…

BEV SegmentationData Augmentationgeo-localizationImage Generation