paper-with-me

Papers

ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model

2024-10-12 · Hongbin Xu, Weitao Chen, Zhipeng Zhou, Feng Xiao, Baigui Sun, Mike Zheng Shou, Wenxiong Kang

Despite recent advancements in 3D generation methods, achieving controllability still remains a challenging issue. Current approaches utilizing score-distillation sampling are hindered by laborious procedures that consume a significant amount of time. Furthermore, the process of first generating 2D representations and then mapping them to 3D lacks internal alignment between the two forms of representation. To address these challenges, we introduce ControLRM, an end-to-end feed-forward model designed for rapid and controllable 3D generation using a large reconstruction model (LRM). ControLRM comprises a 2D condition generator, a condition encoding transformer, and a triplane decoder transformer. Instead of training our model from scratch, we advocate for a joint training framework. In the condition training branch, we lock the triplane decoder and reuses the deep and robust encoding layers pretrained with millions of 3D data in LRM. In the image training branch, we unlock the triplane decoder to establish an implicit alignment between the 2D and 3D representations. To ensure unbiased evaluation, we curate evaluation samples from three distinct datasets (G-OBJ, GSO, ABO) rather than relying on cherry-picking manual generation. The comprehensive experiments conducted on quantitative and qualitative comparisons of 3D controllability and generation quality demonstrate the strong generalization capacity of our proposed approach.

📄 PDF Abstract BibTeX arXiv:2410.09592

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationDecoder

Similar Papers 제목 키워드 기반

Fast Controllable Diffusion Models for Undersampled MRI Reconstruction

2023-11-20 · Wei Jiang, Zhuang Xiong, Feng Liu, Nan Ye 외

Supervised deep learning methods have shown promise in undersampled Magnetic Resonance Imaging (MRI) reconstruction, but their requirement for paired data limits their generalizability to the diverse MRI acquisition para…

MRI Reconstruction

3D MedDiffusion: A 3D Medical Diffusion Model for Controllable and High-quality Medical Image Generation

2024-12-17 · Haoshen Wang, Zhentao Liu, Kaicong Sun, Xiaodong Wang 외

The generation of medical images presents significant challenges due to their high-resolution and three-dimensional nature. Existing methods often yield suboptimal performance in generating high-quality 3D medical images…

CT ReconstructionData AugmentationDenoisingImage Generation+2

DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

2025-10-17 · Weijie Wang, Jiagang Zhu, Zeyu Zhang, Xiaofeng Wang 외 arxiv

We present DriveGen3D, a novel framework for generating high-quality and highly controllable dynamic 3D driving scenes that addresses critical limitations in existing methodologies. Current approaches to driving scene sy…

Scene GenerationVideo Generation

A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion

2026-01-29 · Pu Cao, Yiyang Ma, Feng Zhou, Xuedan Yin 외 arxiv

In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID). In recent ImageNet-scale AE studies, we…

MetaHead: An Engine to Create Realistic Digital Head

2023-04-03 · Dingyun Zhang, Chenglai Zhong, Yudong Guo, Yang Hong 외

Collecting and labeling training data is one important step for learning-based methods because the process is time-consuming and biased. For face analysis tasks, although some generative models can be used to generate fa…

DiversityImage Generation