Retrieval-guided Cross-view Image Synthesis
Information retrieval techniques have demonstrated exceptional capabilities in identifying semantic similarities across diverse domains through robust feature representations. However, their potential in guiding synthesis tasks, particularly cross-view image synthesis, remains underexplored. Cross-view image synthesis presents significant challenges in establishing reliable correspondences between drastically different viewpoints. To address this, we propose a novel retrieval-guided framework that reimagines how retrieval techniques can facilitate effective cross-view image synthesis. Unlike existing methods that rely on auxiliary information, such as semantic segmentation maps or preprocessing modules, our retrieval-guided framework captures semantic similarities across different viewpoints, trained through contrastive learning to create a smooth embedding space. Furthermore, a novel fusion mechanism leverages these embeddings to guide image synthesis while learning and encoding both view-invariant and view-specific features. To further advance this area, we introduce VIGOR-GEN, a new urban-focused dataset with complex viewpoint variations in real-world scenarios. Extensive experiments demonstrate that our retrieval-guided approach significantly outperforms existing methods on the CVUSA, CVACT and VIGOR-GEN datasets, particularly in retrieval accuracy (R@1) and synthesis quality (FID). Our work bridges information retrieval and synthesis tasks, offering insights into how retrieval techniques can address complex cross-domain synthesis challenges.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningDiversityImage GenerationInformation RetrievalRetrievalSemantic SegmentationSSIMMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Geometry-guided Cross-view Diffusion for One-to-many Cross-view Image Synthesis
This paper presents a novel approach for cross-view synthesis aimed at generating plausible ground-level images from corresponding satellite imagery or vice versa. We refer to these tasks as satellite-to-ground (Sat2Grd)…
Image GenerationGeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction
3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to severe artifacts when trained under sparse-view constraints. While recent…
Novel View Synthesis3D ReconstructionTexture Synthesis Guided Deep Hashing for Texture Image Retrieval
With the large-scale explosion of images and videos over the internet, efficient hashing methods have been developed to facilitate memory and time efficient retrieval of similar images. However, none of the existing work…
Data AugmentationDeep HashingImage RetrievalRetrieval+2Geometry-guided Online 3D Video Synthesis with Multi-View Temporal Consistency
We introduce a novel geometry-guided online video view synthesis method with enhanced view and temporal consistency. Traditional approaches achieve high-quality synthesis from dense multi-view camera setups but require s…
Novel View SynthesisCross-view image synthesis using geometry-guided conditional GANs
We address the problem of generating images across two drastically different views, namely ground (street) and aerial (overhead) views. Image synthesis by itself is a very challenging computer vision task and is even mor…
Cross-View Image-to-Image TranslationImage Generation