SinDiffusion: Learning a Diffusion Model from a Single Natural Image
We present SinDiffusion, leveraging denoising diffusion models to capture internal distribution of patches from a single natural image. SinDiffusion significantly improves the quality and diversity of generated samples compared with existing GAN-based approaches. It is based on two core designs. First, SinDiffusion is trained with a single model at a single scale instead of multiple models with progressive growing of scales which serves as the default setting in prior work. This avoids the accumulation of errors, which cause characteristic artifacts in generated results. Second, we identify that a patch-level receptive field of the diffusion network is crucial and effective for capturing the image's patch statistics, therefore we redesign the network structure of the diffusion model. Coupling these two designs enables us to generate photorealistic and diverse images from a single image. Furthermore, SinDiffusion can be applied to various applications, i.e., text-guided image generation, and image outpainting, due to the inherent capability of diffusion models. Extensive experiments on a wide range of images demonstrate the superiority of our proposed method for modeling the patch distribution.
Code (1)
Tasks
DenoisingDiversityImage GenerationImage OutpaintingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SinFusion: Training Diffusion Models on a Single Image or Video
Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate …
DiversityImage ManipulationVideo GenerationUniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation
Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single architecture. While autoregressive VLMs can reason across modalities, the…
multimodal generationImage GenerationText GenerationPC2: Projection-Conditioned Point Cloud Diffusion for Single-Image 3D Reconstruction
Reconstructing the 3D shape of an object from a single RGB image is a long-standing problem in computer vision. In this paper, we propose a novel method for single-image 3D reconstruction which generates a sparse poi…
3D ReconstructionDenoising$PC^2$: Projection-Conditioned Point Cloud Diffusion for Single-Image 3D Reconstruction
Reconstructing the 3D shape of an object from a single RGB image is a long-standing and highly challenging problem in computer vision. In this paper, we propose a novel method for single-image 3D reconstruction which gen…
3D ReconstructionDenoisingZero-1-to-3: Zero-shot One Image to 3D Object
We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors…
3D ReconstructionImage to 3DNovel View SynthesisSingle-View 3D Reconstruction+1