paper-with-me

홈 › Papers

4Dynamic: Text-to-4D Generation with Hybrid Priors

2024-07-17 · Yu-Jie Yuan, Leif Kobbelt, Jiwen Liu, Yuan Zhang, Pengfei Wan, Yu-Kun Lai, Lin Gao

Due to the fascinating generative performance of text-to-image diffusion models, growing text-to-3D generation works explore distilling the 2D generative priors into 3D, using the score distillation sampling (SDS) loss, to bypass the data scarcity problem. The existing text-to-3D methods have achieved promising results in realism and 3D consistency, but text-to-4D generation still faces challenges, including lack of realism and insufficient dynamic motions. In this paper, we propose a novel method for text-to-4D generation, which ensures the dynamic amplitude and authenticity through direct supervision provided by a video prior. Specifically, we adopt a text-to-video diffusion model to generate a reference video and divide 4D generation into two stages: static generation and dynamic generation. The static 3D generation is achieved under the guidance of the input text and the first frame of the reference video, while in the dynamic generation stage, we introduce a customized SDS loss to ensure multi-view consistency, a video-based SDS loss to improve temporal consistency, and most importantly, direct priors from the reference video to ensure the quality of geometry and texture. Moreover, we design a prior-switching training strategy to avoid conflicts between different priors and fully leverage the benefits of each prior. In addition, to enrich the generated motion, we further introduce a dynamic modeling representation composed of a deformation network and a topology network, which ensures dynamic continuity while modeling topological changes. Our method not only supports text-to-4D generation but also enables 4D generation from monocular videos. The comparison experiments demonstrate the superiority of our method compared to existing methods.

📄 PDF Abstract BibTeX arXiv:2407.12684

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

2025-12-05 · Su Sun, Cheng Zhao, Himangi Mittal, Gaurav Mittal 외 arxiv

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize …

Video Generation

3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion Priors

2024-03-04 · Fangzhou Hong, Jiaxiang Tang, Ziang Cao, Min Shi 외

We present a two-stage text-to-3D generation system, namely 3DTopia, which generates high-quality general 3D assets within 5 minutes using hybrid diffusion priors. The first stage samples from a 3D diffusion prior direct…

3D GenerationText to 3DTexture Synthesis

Hybrid Fourier Score Distillation for Efficient One Image to 3D Object Generation

2024-05-31 · Shuzhou Yang, Yu Wang, Haijie Li, Jiarui Meng 외

Single image-to-3D generation is pivotal for crafting controllable 3D assets. Given its under-constrained nature, we attempt to leverage 3D geometric priors from a novel view diffusion model and 2D appearance priors from…

3D GenerationImage GenerationImage to 3D

HumanGen: Generating Human Radiance Fields with Explicit Priors

2022-12-10 · CVPR 2023 1 · Suyi Jiang, Haoran Jiang, Ziyu Wang, Haimin Luo 외

Recent years have witnessed the tremendous progress of 3D GANs for generating view-consistent radiance fields with photo-realism. Yet, high-quality generation of human radiance fields remains challenging, partially due t…

AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving

2025-09-24 · Jinhao Chai, Anqing Jiang, Hao Jiang, Shiyi Mu 외 arxiv

End-to-end multi-modal planning has become a transformative paradigm in autonomous driving, effectively addressing behavioral multi-modality and the generalization challenge in long-tail scenarios. We propose AnchDrive, …

Autonomous Driving