paper-with-me

홈 › Papers

ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion

2023-10-16 · CVPR 2024 1 · Jiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji, Hongdong Li

Given a single image of a 3D object, this paper proposes a novel method (named ConsistNet) that is able to generate multiple images of the same object, as if seen they are captured from different viewpoints, while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Central to our method is a multi-view consistency block which enables information exchange across multiple single-view diffusion processes based on the underlying multi-view geometry principles. ConsistNet is an extension to the standard latent diffusion model, and consists of two sub-modules: (a) a view aggregation module that unprojects multi-view features into global 3D volumes and infer consistency, and (b) a ray aggregation module that samples and aggregate 3D consistent features back to each view to enforce consistency. Our approach departs from previous methods in multi-view image generation, in that it can be easily dropped-in pre-trained LDMs without requiring explicit pixel correspondences or depth prediction. Experiments show that our method effectively learns 3D consistency over a frozen Zero123 backbone and can generate 16 surrounding views of the object within 40 seconds on a single A100 GPU. Our code will be made available on https://github.com/JiayuYANG/ConsistNet

📄 PDF Abstract BibTeX arXiv:2310.10343

Code (1)

jiayuyang/consistnet 공식 구현

Tasks

Depth EstimationDepth PredictionGPUImage GenerationObject

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Consolidating Attention Features for Multi-view Image Editing

2024-02-22 · Or Patashnik, Rinon Gal, Daniel Cohen-Or, Jun-Yan Zhu 외

Large-scale text-to-image models enable a wide range of image editing techniques, using text prompts or even spatial controls. However, applying these editing methods to multi-view images depicting a single scene leads t…

MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation

2024-04-04 · CVPR 2024 1 · Hanzhe Hu, Zhizhuo Zhou, Varun Jampani, Shubham Tulsiani

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models, these…

DenoisingDepth EstimationDepth Prediction

Deep Facial Non-Rigid Multi-View Stereo

2020-06-01 · CVPR 2020 6 · Ziqian Bai, Zhaopeng Cui, Jamal Ahmed Rahim, Xiaoming Liu 외

We present a method for 3D face reconstruction from multi-view images with different expressions. We formulate this problem from the perspective of non-rigid multi-view stereo (NRMVS). Unlike previous learning-based meth…

3D Face ReconstructionFace Reconstruction

Share With Thy Neighbors: Single-View Reconstruction by Cross-Instance Consistency

2022-04-21 · Tom Monnier, Matthew Fisher, Alexei A. Efros, Mathieu Aubry

Approaches for single-view reconstruction typically rely on viewpoint annotations, silhouettes, the absence of background, multiple views of the same instance, a template shape, or symmetry. We avoid all such supervision…

3D Object Reconstruction3D Object Reconstruction From A Single Image3D ReconstructionObject+1

Neural Radiance Flow for 4D View Synthesis and Video Processing

2020-12-17 · ICCV 2021 10 · Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenenbaum 외

We present a method, Neural Radiance Flow (NeRFlow),to learn a 4D spatial-temporal representation of a dynamic scene from a set of RGB images. Key to our approach is the use of a neural implicit representation that learn…

Image Super-ResolutionSuper-ResolutionTemporal View Synthesis