paper-with-me

홈 › Papers

A Unified Approach for Text- and Image-guided 4D Scene Generation

2023-11-28 · CVPR 2024 1 · Yufeng Zheng, Xueting Li, Koki Nagano, Sifei Liu, Karsten Kreis, Otmar Hilliges, Shalini De Mello

Large-scale diffusion generative models are greatly simplifying image, video and 3D asset creation from user-provided text prompts and images. However, the challenging problem of text-to-4D dynamic 3D scene generation with diffusion guidance remains largely unexplored. We propose Dream-in-4D, which features a novel two-stage approach for text-to-4D synthesis, leveraging (1) 3D and 2D diffusion guidance to effectively learn a high-quality static 3D asset in the first stage; (2) a deformable neural radiance field that explicitly disentangles the learned static asset from its deformation, preserving quality during motion learning; and (3) a multi-resolution feature grid for the deformation field with a displacement total variation loss to effectively learn motion with video diffusion guidance in the second stage. Through a user preference study, we demonstrate that our approach significantly advances image and motion quality, 3D consistency and text fidelity for text-to-4D generation compared to baseline approaches. Thanks to its motion-disentangled representation, Dream-in-4D can also be easily adapted for controllable generation where appearance is defined by one or multiple images, without the need to modify the motion learning stage. Thus, our method offers, for the first time, a unified approach for text-to-4D, image-to-4D and personalized 4D generation tasks.

📄 PDF Abstract BibTeX arXiv:2311.16854

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Align, Adapt and Inject: Sound-guided Unified Image Generation

2023-06-20 · Yue Yang, Kaipeng Zhang, Yuying Ge, Wenqi Shao 외

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid represen…

Image GenerationRetrievalText Retrieval

UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance

2022-10-28 · Wei Li, Xue Xu, Xinyan Xiao, Jiachen Liu 외

Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusion model and cross-modal guided diffusion…

Image GenerationImage-text matchingLanguage ModelingLanguage Modelling+1

Panoptic Diffusion Models: co-generation of images and segmentation maps

2024-12-04 · Yinghan Long, Kaushik Roy

Recently, diffusion models have demonstrated impressive capabilities in text-guided and image-conditioned image generation. However, existing diffusion models cannot simultaneously generate a segmentation map of objects …

Image GenerationPanoptic SegmentationSegmentation

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

2025-12-29 · Tianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen 외 arxiv

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, wi…

Scene UnderstandingScene Generation

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

2026-02-09 · Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu 외 arxiv

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are aligned. To investigate thi…

Image Generation