paper-with-me

홈 › Papers

Top2Ground: A Height-Aware Dual Conditioning Diffusion Model for Robust Aerial-to-Ground View Generation

2025-11-11 · Jae Joong Lee, Bedrich Benes arxiv

Generating ground-level images from aerial views is a challenging task due to extreme viewpoint disparity, occlusions, and a limited field of view. We introduce Top2Ground, a novel diffusion-based method that directly generates photorealistic ground-view images from aerial input images without relying on intermediate representations such as depth maps or 3D voxels. Specifically, we condition the denoising process on a joint representation of VAE-encoded spatial features (derived from aerial RGB images and an estimated height map) and CLIP-based semantic embeddings. This design ensures the generation is both geometrically constrained by the scene's 3D structure and semantically consistent with its content. We evaluate Top2Ground on three diverse datasets: CVUSA, CVACT, and the Auto Arborist. Our approach shows 7.3% average improvement in SSIM across three benchmark datasets, showing Top2Ground can robustly handle both wide and narrow fields of view, highlighting its strong generalization capabilities.

📄 PDF Abstract BibTeX arXiv:2511.08258

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diffusion-Based, Data-Assimilation-Enabled Super-Resolution of Hub-height Winds

2025-10-03 · Xiaolong Ma, Xu Dong, Ashley Tarrant, Lei Yang 외 arxiv

High-quality observations of hub-height winds are valuable but sparse in space and time. Simulations are widely available on regular grids but are generally biased and too coarse to inform wind-farm siting or to assess e…

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy

2026-05-06 · Yuan Wu, Zhiqiang Yan, Jiawei Lian, Zhengxue Wang 외 arxiv

3D occupancy prediction aims to infer dense, voxel-wise scene semantics from sensor observations, where the 2D-to-3D view transformation serves as a crucial step in bridging image features and volumetric representations.…

DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion

2025-05-03 · Haoteng Li, Zhao Yang, Zezhong Qian, Gongpeng Zhao 외

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and…

3D Object DetectionAutonomous DrivingBEV Segmentationobject-detection+2

Conditioning Matters: Stabilizing Inversion and Attention in Diffusion Image Editing

2026-06-12 · Zheyuan Zhan, Hongchen Li, Can Wang, Yinfei Ma 외 arxiv

Inversion-based image editing offers flexible and training-free control but still struggles with inversion accuracy and the trade-off between editing fidelity and background preservation. While recent methods improve inv…

Image Editing

DEMIST: Decoupled Multi-stream latent diffusion for Quantitative Myelin Map Synthesis

2025-11-16 · Jiacheng Wang, Hao Li, Xing Yao, Ahmad Toubasi 외 arxiv

Quantitative magnetization transfer (qMT) imaging provides myelin-sensitive biomarkers, such as the pool size ratio (PSR), which is valuable for multiple sclerosis (MS) assessment. However, qMT requires specialized 20-30…