paper-with-me

Papers

World-consistent Video Diffusion with Explicit 3D Modeling

2024-12-02 · CVPR 2025 1 · Qihang Zhang, Shuangfei Zhai, Miguel Angel Bautista, Kevin Miao, Alexander Toshev, Joshua Susskind, Jiatao Gu

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like single-image-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model.

📄 PDF Abstract BibTeX arXiv:2412.01821

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationImage GenerationImage to 3DVideo Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

2026-07-23 · Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh 외 hf

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines …

Video Generation

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

2026-03-03 · Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang 외 arxiv

Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible d…

Reinforcement Learning

Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models

2025-01-01 · CVPR 2025 1 · Ruofan Liang, Zan Gojcic, Huan Ling, Jacob Munkberg 외

Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene repres…

3D geometryInverse Rendering

DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models

2025-01-30 · Ruofan Liang, Zan Gojcic, Huan Ling, Jacob Munkberg 외

Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene representa…

3D geometryInverse Rendering

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

2026-08-05 · Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou 외 hf

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance use…

Novel View Synthesis