paper-with-me

홈 › Papers

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

2025-01-30 · CVPR 2025 1 · Vitor Guizilini, Muhammad Zubair Irshad, Dian Chen, Greg Shakhnarovich, Rares Ambrus

Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.

📄 PDF Abstract BibTeX arXiv:2501.18804

Code (0)

등록된 구현이 없습니다.

Tasks

3D Scene ReconstructionDepth EstimationNovel View Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

iNVS: Repurposing Diffusion Inpainters for Novel View Synthesis

2023-10-24 · Yash Kant, Aliaksandr Siarohin, Michael Vasilkovsky, Riza Alp Guler 외

We present a method for generating consistent novel views from a single source image. Our approach focuses on maximizing the reuse of visible pixels from the source image. To achieve this, we use a monocular depth estima…

Novel View Synthesis

ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

2023-10-27 · CVPR 2024 1 · Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann 외

We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to…

DiversityNeRFNovel View Synthesis

Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

2024-03-21 · Francesco Di Felice, Alberto Remus, Stefano Gasperini, Benjamin Busam 외

Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state…

6D Pose EstimationNovel View SynthesisPose Estimation

StereoCrafter-Zero: Zero-Shot Stereo Video Generation with Noisy Restart

2024-11-21 · Jian Shi, Qian Wang, Zhenyu Li, Peter Wonka

Generating high-quality stereo videos that mimic human binocular vision requires maintaining consistent depth perception and temporal coherence across frames. While diffusion models have advanced image and video synthesi…

Video Generation

Learning A Zero-shot Occupancy Network from Vision Foundation Models via Self-supervised Adaptation

2025-03-10 · Sihao Lin, Daqi Liu, Ruochong Fu, Dongrui Liu 외

Estimating the 3D world from 2D monocular images is a fundamental yet challenging task due to the labour-intensive nature of 3D annotations. To simplify label acquisition, this work proposes a novel approach that bridges…

Novel View Synthesis