paper-with-me

홈 › Papers

HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation

2024-01-15 · Antoine Mercier, Ramin Nakhli, Mahesh Reddy, Rajeev Yasarla, Hong Cai, Fatih Porikli, Guillaume Berger

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in data scarcity: the most extensive 3D datasets encompass merely millions of assets, while their 2D counterparts contain billions of text-image pairs. To address this, we propose a novel approach which harnesses the power of large, pretrained 2D diffusion models. More specifically, our approach, HexaGen3D, fine-tunes a pretrained text-to-image model to jointly predict 6 orthographic projections and the corresponding latent triplane. We then decode these latents to generate a textured mesh. HexaGen3D does not require per-sample optimization, and can infer high-quality and diverse objects from textual prompts in 7 seconds, offering significantly better quality-to-latency trade-offs when comparing to existing approaches. Furthermore, HexaGen3D demonstrates strong generalization to new objects or compositions.

📄 PDF Abstract BibTeX arXiv:2401.07727

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation

2024-01-22 · Zhihong Chen, Maya Varma, Justin Xu, Magdalini Paschali 외

Over 1.4 billion chest X-rays (CXRs) are performed annually due to their cost-effectiveness as an initial diagnostic test. This scale of radiological studies provides a significant opportunity to streamline CXR interpret…

BenchmarkingDiagnosticFairnessLanguage Modelling+1

Designing a Better Asymmetric VQGAN for StableDiffusion

2023-06-07 · Zixin Zhu, Xuelu Feng, Dongdong Chen, Jianmin Bao 외

StableDiffusion is a revolutionary text-to-image generator that is causing a stir in the world of image generation and editing. Unlike traditional methods that learn a diffusion model in pixel space, StableDiffusion lear…

DecoderImage GenerationImage Inpainting

Parallel Sampling of Diffusion Models

2023-05-25 · NeurIPS 2023 11 · Andy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh 외

Diffusion models are powerful generative models but suffer from slow sampling, often taking 1000 sequential denoising steps for one sample. As a result, considerable efforts have been directed toward reducing the number …

DenoisingImage Generation

Predict then Propagate: Graph Neural Networks meet Personalized PageRank

2018-10-14 · ICLR 2019 5 · Johannes Gasteiger, Aleksandar Bojchevski, Stephan Günnemann

Neural message passing algorithms for semi-supervised classification on graphs have recently achieved great success. However, for classifying a node these methods only consider nodes that are a few propagation steps away…

General ClassificationNode ClassificationNode Classification on Non-Homophilic (Heterophilic) Graphs

Exploring the Capability of Text-to-Image Diffusion Models with Structural Edge Guidance for Multi-Spectral Satellite Image Inpainting

2023-11-06 · Mikolaj Czerkawski, Christos Tachtatzis

The letter investigates the utility of text-to-image inpainting models for satellite image data. Two technical challenges of injecting structural guiding signals into the generative process as well as translating the inp…

Image InpaintingTranslation