paper-with-me

홈 › Papers

DreamMapping: High-Fidelity Text-to-3D Generation via Variational Distribution Mapping

2024-09-08 · Zeyu Cai, Duotun Wang, Yixun Liang, Zhijing Shao, Ying-Cong Chen, Xiaohang Zhan, Zeyu Wang

Score Distillation Sampling (SDS) has emerged as a prevalent technique for text-to-3D generation, enabling 3D content creation by distilling view-dependent information from text-to-2D guidance. However, they frequently exhibit shortcomings such as over-saturated color and excess smoothness. In this paper, we conduct a thorough analysis of SDS and refine its formulation, finding that the core design is to model the distribution of rendered images. Following this insight, we introduce a novel strategy called Variational Distribution Mapping (VDM), which expedites the distribution modeling process by regarding the rendered images as instances of degradation from diffusion-based generation. This special design enables the efficient training of variational distribution by skipping the calculations of the Jacobians in the diffusion U-Net. We also introduce timestep-dependent Distribution Coefficient Annealing (DCA) to further improve distilling precision. Leveraging VDM and DCA, we use Gaussian Splatting as the 3D representation and build a text-to-3D generation framework. Extensive experiments and evaluations demonstrate the capability of VDM and DCA to generate high-fidelity and realistic assets with optimization efficiency.

📄 PDF Abstract BibTeX arXiv:2409.05099

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationText to 3D

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

2023-05-25 · NeurIPS 2023 11 · Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 외

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity …

3D GenerationDiversityNeRFText to 3D

Single Image to High-Quality 3D Object via Latent Features

2025-11-24 · Huanning Dong, Yinuo Huang, Fan Li, Ping Kuang arxiv

3D assets are essential in the digital age. While automatic 3D generation, such as image-to-3d, has made significant strides in recent years, it often struggles to achieve fast, detailed, and high-fidelity generation sim…

3D Generation

Efficient Generative Modeling with Residual Vector Quantization-Based Tokens

2024-12-13 · Jaehyeon Kim, Taehong Moon, Keon Lee, Jaewoong Cho

We explore the use of Residual Vector Quantization (RVQ) for high-fidelity generation in vector-quantized generative models. This quantization technique maintains higher data fidelity by employing more in-depth tokens. H…

Conditional Image GenerationImage GenerationQuantizationSpeech Synthesis+4

Enhancing Variational Autoencoders with Smooth Robust Latent Encoding

2025-04-24 · Hyomin Lee, Minseon Kim, Sangwon Jang, Jongheon Jeong 외

Variational Autoencoders (VAEs) have played a key role in scaling up diffusion-based generative models, as in Stable Diffusion, yet questions regarding their robustness remain largely underexplored. Although adversarial …

Image Reconstructiontext-guided-image-editing

High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

2024-07-04 · Gael Le Lan, Bowen Shi, Zhaoheng Ni, Sidd Srinivasan 외

We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48 kHz stereo variational auto encoder code…

DenoisingMusic Generation