paper-with-me

홈 › Papers

Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization

2025-04-05 · Yikai Wang, Guangce Liu, Xinzhou Wang, Zilong Chen, Jiafang Li, Xin Liang, Fuchun Sun, Jun Zhu

The advancement of 4D (i.e., sequential 3D) generation opens up new possibilities for lifelike experiences in various applications, where users can explore dynamic objects or characters from any viewpoint. Meanwhile, video generative models are receiving particular attention given their ability to produce realistic and imaginative frames. These models are also observed to exhibit strong 3D consistency, indicating the potential to act as world simulators. In this work, we present Video4DGen, a novel framework that excels in generating 4D representations from single or multiple generated videos as well as generating 4D-guided videos. This framework is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. The 4D outputs generated by Video4DGen are represented using our proposed Dynamic Gaussian Surfels (DGS), which optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. We design warped-state geometric regularization and refinements on Gaussian surfels, to preserve the structural integrity and fine-grained appearance details. To perform 4D generation from multiple videos and capture representation across spatial, temporal, and pose dimensions, we design multi-video alignment, root pose optimization, and pose-guided frame sampling strategies. The leveraging of continuous warping fields also enables a precise depiction of pose, motion, and deformation over per-video frames. Further, to improve the overall fidelity from the observation of all camera poses, Video4DGen performs novel-view video generation guided by the 4D content, with the proposed confidence-filtered DGS to enhance the quality of generated sequences. With the ability of 4D and video generation, Video4DGen offers a powerful tool for applications in virtual reality, animation, and beyond.

📄 PDF Abstract BibTeX arXiv:2504.04153

Code (1)

yikaiw/vidu4d 공식 구현 pytorch

Tasks

3D GenerationVideo AlignmentVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency

2023-12-28 · Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao 외

Aided by text-to-image and text-to-video diffusion models, existing 4D content creation pipelines utilize score distillation sampling to optimize the entire dynamic 3D scene. However, as these pipelines generate 4D conte…

Motion GenerationPrompt Engineering

VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

2024-08-05 · Zhiyu Tan, Xiaomeng Yang, Luozheng Qin, Hao Li

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consis…

Text-to-Video GenerationVideo Generation

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

2025-07-08 · Rongsheng Wang, Junying Chen, Ke Ji, Zhenyang Cai 외 arxiv

Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical traini…

Video Generation

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

2025-09-24 · Jianxuan Yang, Xiaoran Yang, Lipan Zhang, Xinyue Guo 외 arxiv

Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods fac…

Audio Generation

Robust Multi-Object 4D Generation for In-the-wild Videos

2025-01-01 · CVPR 2025 1 · Wen-Hsuan Chu, Lei Ke, Jianmeng Liu, Mingxiao Huo 외

We address the challenge of generating dynamic 4D scenes from monocular multi-object videos with heavy occlusions and introduce Robust4DGen, a novel approach that integrates rendering-based deformable 3D Gaussian opt…

ObjectScene Generation