paper-with-me

Papers

4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency

2023-12-28 · Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao, Yunchao Wei

Aided by text-to-image and text-to-video diffusion models, existing 4D content creation pipelines utilize score distillation sampling to optimize the entire dynamic 3D scene. However, as these pipelines generate 4D content from text or image inputs directly, they are constrained by limited motion capabilities and depend on unreliable prompt engineering for desired results. To address these problems, this work introduces \textbf{4DGen}, a novel framework for grounded 4D content creation. We identify monocular video sequences as a key component in constructing the 4D content. Our pipeline facilitates controllable 4D generation, enabling users to specify the motion via monocular video or adopt image-to-video generations, thus offering superior control over content creation. Furthermore, we construct our 4D representation using dynamic 3D Gaussians, which permits efficient, high-resolution supervision through rendering during training, thereby facilitating high-quality 4D generation. Additionally, we employ spatial-temporal pseudo labels on anchor frames, along with seamless consistency priors implemented through 3D-aware score distillation sampling and smoothness regularizations. Compared to existing video-to-4D baselines, our approach yields superior results in faithfully reconstructing input signals and realistically inferring renderings from novel viewpoints and timesteps. More importantly, compared to previous image-to-4D and text-to-4D works, 4DGen supports grounded generation, offering users enhanced control and improved motion generation capabilities, a feature difficult to achieve with previous methods. Project page: https://vita-group.github.io/4DGen/

📄 PDF Abstract BibTeX arXiv:2312.17225

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationPrompt Engineering

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Turbo4DGen: Ultra-Fast Acceleration for 4D Generation

2026-01-24 · Yuanbin Man, Ying Huang, Zhile Ren, Miao Yin arxiv

4D generation, or dynamic 3D content generation, integrates spatial, temporal, and view dimensions to model realistic dynamic scenes, playing a foundational role in advancing world models and physical AI. However, mainta…

Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization

2025-04-05 · Yikai Wang, Guangce Liu, Xinzhou Wang, Zilong Chen 외

The advancement of 4D (i.e., sequential 3D) generation opens up new possibilities for lifelike experiences in various applications, where users can explore dynamic objects or characters from any viewpoint. Meanwhile, vid…

3D GenerationVideo AlignmentVideo Generation

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

2025-06-10 · Agnese Taluzzi, Davide Gesualdi, Riccardo Santambrogio, Chiara Plizzari 외

This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large language model (MLLM) to capture fine-graine…

Knowledge GraphsLanguage ModelingLanguage ModellingLarge Language Model+1

Phys4DGen: A Physics-Driven Framework for Controllable and Efficient 4D Content Generation from a Single Image

2024-11-25 · Jiajing Lin, Zhenzhong Wang, Shu Jiang, Yongjie Hou 외

The task of 4D content generation involves creating dynamic 3D models that evolve over time in response to specific input conditions, such as images. Existing methods rely heavily on pre-trained video diffusion models to…

Physical Simulations

PBR3DGen: A VLM-guided Mesh Generation with High-quality PBR Texture

2025-03-14 · Xiaokang Wei, BoWen Zhang, Xianghui Yang, Yuxuan Wang 외

Generating high-quality physically based rendering (PBR) materials is important to achieve realistic rendering in the downstream tasks, yet it remains challenging due to the intertwined effects of materials and lighting.…

3D Generation