Multi-View Unsupervised Image Generation with Cross Attention Guidance
The growing interest in novel view synthesis, driven by Neural Radiance Field (NeRF) models, is hindered by scalability issues due to their reliance on precisely annotated multi-view images. Recent models address this by fine-tuning large text2image diffusion models on synthetic multi-view data. Despite robust zero-shot generalization, they may need post-processing and can face quality issues due to the synthetic-real domain gap. This paper introduces a novel pipeline for unsupervised training of a pose-conditioned diffusion model on single-category datasets. With the help of pretrained self-supervised Vision Transformers (DINOv2), we identify object poses by clustering the dataset through comparing visibility and locations of specific object parts. The pose-conditioned diffusion model, trained on pose labels, and equipped with cross-frame attention at inference time ensures cross-view consistency, that is further aided by our novel hard-attention guidance. Our model, MIRAGE, surpasses prior work in novel view synthesis on real images. Furthermore, MIRAGE is robust to diverse textures and geometries, as demonstrated with our experiments on synthetic images generated with pretrained Stable Diffusion.
Code (0)
등록된 구현이 없습니다.
Tasks
Hard AttentionImage GenerationNeRFNovel View SynthesisZero-shot GeneralizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Multi-view UAV Image Geo-localization via Iterative Rendering
Unmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) presents significant challenges due to the view discrepancy between oblique UAV images and overhead satellite images. Existing methods heavily rely on the …
geo-localizationImage GenerationUCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space Translation
Although existing image inpainting approaches have been able to produce visually realistic and semantically correct results, they produce only one result for each masked input. In order to produce multiple and diverse re…
DiversityGenerative Adversarial NetworkImage InpaintingTranslationUNeR3D: Versatile and Scalable 3D RGB Point Cloud Generation from 2D Images in Unsupervised Reconstruction
In the realm of 3D reconstruction from 2D images, a persisting challenge is to achieve high-precision reconstructions devoid of 3D Ground Truth data reliance. We present UNeR3D, a pioneering unsupervised methodology that…
3D ReconstructionPoint Cloud GenerationPAC-GAN: An Effective Pose Augmentation Scheme for Unsupervised Cross-View Person Re-identification
Person re-identification (person Re-Id) aims to retrieve the pedestrian images of a same person that captured by disjoint and non-overlapping cameras. Lots of researchers recently focuse on this hot issue and propose dee…
Cross-Modal Person Re-IdentificationGenerative Adversarial NetworkImage RetrievalPerson Re-Identification+1GANSeg: Learning to Segment by Unsupervised Hierarchical Image Generation
Segmenting an image into its parts is a frequent preprocess for high-level vision tasks such as image editing. However, annotating masks for supervised training is expensive. Weakly-supervised and unsupervised methods ex…
Image AugmentationImage GenerationSegmentationUnsupervised Facial Landmark Detection+2