paper-with-me

홈 › Papers

Controlling Space and Time with Diffusion Models

2024-07-10 · Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet

We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works that focus on limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes. For an overview see https://4d-diffusion.github.io

📄 PDF Abstract BibTeX arXiv:2407.07860

Code (0)

등록된 구현이 없습니다.

Tasks

Image to 3DNovel View Synthesis

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Diffusion Models already have a Semantic Latent Space

2022-10-20 · Mingi Kwon, Jaeseok Jeong, Youngjung Uh

Diffusion models achieve outstanding generative performance in various domains. Despite their great success, they lack semantic latent space which is essential for controlling the generative process. To address the probl…

Image Manipulation

Exploring the Design Space of Diffusion Bridge Models via Stochasticity Control

2024-10-28 · Shaorong Zhang, Yuanbin Cheng, Xianghao Kong, Greg Ver Steeg

Diffusion bridge models effectively facilitate image-to-image (I2I) translation by connecting two distributions. However, existing methods overlook the impact of noise in sampling SDEs, transition kernel, and the base di…

Diversity

MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model

2024-04-30 · Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu 외

This work introduces MotionLCM, extending controllable motion generation to a real-time level. Existing methods for spatial-temporal control in text-conditioned motion generation suffer from significant runtime inefficie…

Motion GenerationMotion Synthesis

Unifying Diffusion Models' Latent Space, with Applications to CycleDiffusion and Guidance

2022-10-11 · Chen Henry Wu, Fernando de la Torre

Diffusion models have achieved unprecedented performance in generative modeling. The commonly-adopted formulation of the latent code of diffusion models is a sequence of gradually denoised samples, as opposed to the simp…

Image GenerationImage-to-Image Translation

Diffusion-LM Improves Controllable Text Generation

2022-05-27 · Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 외

Controlling the behavior of language models (LMs) without re-training is a major open problem in natural language generation. While recent works have demonstrated successes on controlling simple sentence attributes (e.g.…

Language ModelingLanguage ModellingSentenceText Generation