paper-with-me

Papers

4Diffusion: Multi-view Video Diffusion Model for 4D Generation

2024-05-31 · Haiyu Zhang, Xinyuan Chen, Yaohui Wang, Xihui Liu, Yunhong Wang, Yu Qiao

Current 4D generation methods have achieved noteworthy efficacy with the aid of advanced diffusion generative models. However, these methods lack multi-view spatial-temporal modeling and encounter challenges in integrating diverse prior knowledge from multiple diffusion models, resulting in inconsistent temporal appearance and flickers. In this paper, we propose a novel 4D generation pipeline, namely 4Diffusion, aimed at generating spatial-temporally consistent 4D content from a monocular video. We first design a unified diffusion model tailored for multi-view video generation by incorporating a learnable motion module into a frozen 3D-aware diffusion model to capture multi-view spatial-temporal correlations. After training on a curated dataset, our diffusion model acquires reasonable temporal consistency and inherently preserves the generalizability and spatial consistency of the 3D-aware diffusion model. Subsequently, we propose 4D-aware Score Distillation Sampling loss, which is based on our multi-view video diffusion model, to optimize 4D representation parameterized by dynamic NeRF. This aims to eliminate discrepancies arising from multiple diffusion models, allowing for generating spatial-temporally consistent 4D content. Moreover, we devise an anchor loss to enhance the appearance details and facilitate the learning of dynamic NeRF. Extensive qualitative and quantitative experiments demonstrate that our method achieves superior performance compared to previous methods.

📄 PDF Abstract BibTeX arXiv:2405.20674

Code (0)

등록된 구현이 없습니다.

Tasks

NeRFVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

2024-06-12 · Bing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai 외

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie…

Video Generation

Zero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion Model

2025-03-28 · Jangho Park, Taesung Kwon, Jong Chul Ye

Recently, multi-view or 4D video generation has emerged as a significant research topic. Nonetheless, recent approaches to 4D generation still struggle with fundamental limitations, as they primarily rely on harnessing m…

Video Generation

View-Consistent Diffusion Representations for 3D-Consistent Video Generation

2025-11-24 · Duolikun Danier, Ge Gao, Steven McDonagh, Changjian Li 외 arxiv

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arisi…

Video Generation

Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models

2024-09-11 · Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao 외

Despite having tremendous progress in image-to-3D generation, existing methods still struggle to produce multi-view consistent images with high-resolution textures in detail, especially in the paradigm of 2D diffusion th…

3D Generation3D ReconstructionImage GenerationImage to 3D+2

DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

2023-10-11 · Xiaofan Li, Yifu Zhang, Xiaoqing Ye

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urg…

Autonomous DrivingImage GenerationVideo Generation