paper-with-me

홈 › Papers

MultiDiff: Consistent Novel View Synthesis from a Single Image

2024-06-26 · CVPR 2024 1 · Norman Müller, Katja Schwarz, Barbara Roessle, Lorenzo Porzi, Samuel Rota Bulò, Matthias Nießner, Peter Kontschieder

We introduce MultiDiff, a novel approach for consistent novel view synthesis of scenes from a single RGB image. The task of synthesizing novel views from a single reference image is highly ill-posed by nature, as there exist multiple, plausible explanations for unobserved areas. To address this issue, we incorporate strong priors in form of monocular depth predictors and video-diffusion models. Monocular depth enables us to condition our model on warped reference images for the target views, increasing geometric stability. The video-diffusion prior provides a strong proxy for 3D scenes, allowing the model to learn continuous and pixel-accurate correspondences across generated images. In contrast to approaches relying on autoregressive image generation that are prone to drifts and error accumulation, MultiDiff jointly synthesizes a sequence of frames yielding high-quality and multi-view consistent results -- even for long-term scene generation with large camera movements, while reducing inference time by an order of magnitude. For additional consistency and image quality improvements, we introduce a novel, structured noise distribution. Our experimental results demonstrate that MultiDiff outperforms state-of-the-art methods on the challenging, real-world datasets RealEstate10K and ScanNet. Finally, our model naturally supports multi-view consistent editing without the need for further tuning.

📄 PDF Abstract BibTeX arXiv:2406.18524

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationNovel View SynthesisScene Generation

Similar Papers 제목 키워드 기반

MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose

2026-02-22 · Sirine Bhouri, Lan Wei, Jian-Qing Zheng, Dandan Zhang arxiv

Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cr…

Image GenerationPose Estimation

Consistent Mesh Diffusion

2023-12-01 · Julian Knodt, Xifeng Gao

Given a 3D mesh with a UV parameterization, we introduce a novel approach to generating textures from text prompts. While prior work uses optimization from Text-to-Image Diffusion models to generate textures and geometry…

Spherical Dense Text-to-Image Synthesis

2025-02-18 · Timon Winter, Stanislav Frolov, Brian Bernhard Moser, Andreas Dengel

Recent advancements in text-to-image (T2I) have improved synthesis results, but challenges remain in layout control and generating omnidirectional panoramic images. Dense T2I (DT2I) and spherical T2I (ST2I) models addres…

Image Generation

MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

2023-02-16 · Omer Bar-Tal, Lior Yariv, Yaron Lipman, Tali Dekel

Recent advances in text-to-image generation with diffusion models present transformative capabilities in image quality. However, user controllability of the generated image, and fast adaptation to new tasks still remains…

Image GenerationText to Image GenerationText-to-Image Generation

HumanGif: Single-View Human Diffusion with Generative Prior

2025-02-17 · Shoukang Hu, Takuya Narihira, Kazumi Fukuda, Ryosuke Sawata 외

While previous single-view-based 3D human reconstruction methods made significant progress in novel view synthesis, it remains a challenge to synthesize both view-consistent and pose-consistent results for animatable hum…

3D Human ReconstructionNeRFNovel View Synthesis