paper-with-me

Papers

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

2023-12-21 · CVPR 2024 1 · Huan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler, Karsten Kreis

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here, we instead focus on the underexplored text-to-4D setting and synthesize dynamic, animated 3D objects using score distillation methods with an additional temporal dimension. Compared to previous work, we pursue a novel compositional generation-based approach, and combine text-to-image, text-to-video, and 3D-aware multiview diffusion models to provide feedback during 4D object optimization, thereby simultaneously enforcing temporal consistency, high-quality visual appearance and realistic geometry. Our method, called Align Your Gaussians (AYG), leverages dynamic 3D Gaussian Splatting with deformation fields as 4D representation. Crucial to AYG is a novel method to regularize the distribution of the moving 3D Gaussians and thereby stabilize the optimization and induce motion. We also propose a motion amplification mechanism as well as a new autoregressive synthesis scheme to generate and combine multiple 4D sequences for longer generation. These techniques allow us to synthesize vivid dynamic scenes, outperform previous work qualitatively and quantitatively and achieve state-of-the-art text-to-4D performance. Due to the Gaussian 4D representation, different 4D animations can be seamlessly combined, as we demonstrate. AYG opens up promising avenues for animation, simulation and digital content creation as well as synthetic data generation.

📄 PDF Abstract BibTeX arXiv:2312.13763

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

SAGA: Surface-Aligned Gaussian Avatar

2024-12-01 · Ronghan Chen, Yang Cong, Jiayue Liu

This paper presents a Surface-Aligned Gaussian representation for creating animatable human avatars from monocular videos,aiming at improving the novel view and pose synthesis performance while ensuring fast training and…

3DGSDynamic ReconstructionNeRF

1000+ FPS 4D Gaussian Splatting for Dynamic Scene Rendering

2025-03-20 · Yuheng Yuan, Qiuhong Shen, Xingyi Yang, Xinchao Wang

4D Gaussian Splatting (4DGS) has recently gained considerable attention as a method for reconstructing dynamic scenes. Despite achieving superior quality, 4DGS typically requires substantial storage and suffers from slow…

Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction

2026-04-06 · Ahan Shabanov, Peter Hedman, Ethan Weber, Zhengqin Li 외 arxiv

We present Free-Range Gaussians, a multi-view reconstruction method that predicts non-pixel, non-voxel-aligned 3D Gaussians from as few as four images. This is done through flow matching over Gaussian parameters. Our gen…

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

2025-01-01 · CVPR 2025 1 · Yuheng Jiang, Zhehao Shen, Chengcheng Guo, Yu Hong 외

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform gener…

AttributePosition

Neural Signed Distance Function Inference through Splatting 3D Gaussians Pulled on Zero-Level Set

2024-10-18 · Wenyuan Zhang, Yu-Shen Liu, Zhizhong Han

It is vital to infer a signed distance function (SDF) in multi-view based surface reconstruction. 3D Gaussian splatting (3DGS) provides a novel perspective for volume rendering, and shows advantages in rendering efficien…

3DGSNeural RenderingSurface Reconstruction