paper-with-me

홈 › Papers

From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy

2026-03-23 · Bi'an Du, Daizong Liu, Pufan Li, Wei Hu arxiv

Single-image 3D generation lies at the core of vision-to-graphics models in the real world. However, it remains a fundamental challenge to achieve reliable generalization across diverse semantic categories and highly variable structural complexity under sparse supervision. Existing approaches typically model objects in a monolithic manner or rely on a fixed number of parts, including recent part-aware models such as PartCrafter, which still require a labor-intensive user-specified part count. Such designs easily lead to overfitting, fragmented or missing structural components, and limited compositional generalization when encountering novel object layouts. To this end, this paper rethinks single-image 3D generation as learning an adaptive part-whole hierarchy in the flexible 3D latent space. We present a novel part-to-whole 3D generative world model that autonomously discovers latent structural slots by inferring soft and compositional masks directly from image tokens. Specifically, an adaptive slot-gating mechanism dynamically determines the slot-wise activation probabilities and smoothly consolidates redundant slots within different objects, ensuring that the emergent structure remains compact yet expressive across categories. Each distilled slot is then aligned to a learnable, class-agnostic prototype bank, enabling powerful cross-category shape sharing and denoising through universal geometric prototypes in the real world. Furthermore, a lightweight 3D denoiser is introduced to reconstruct geometry and appearance via unified diffusion objectives. Experiments show consistent gains in cross-category transfer and part-count extrapolation, and ablations confirm complementary benefits of the prototype bank for shape-prior sharing as well as slot-gating for structural adaptation.

📄 PDF Abstract BibTeX arXiv:2603.21557

Code (0)

등록된 구현이 없습니다.

Tasks

3D Generation

Similar Papers 제목 키워드 기반

BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics

2026-06-16 · Junfeng Xia, Wenhao Ye, Junxiang Zhang, Xuanye Pan 외 arxiv

Whole-brain 4D fMRI generation is valuable for modeling functional brain dynamics, yet existing fMRI foundation models mainly target representation learning and downstream prediction rather than conditional predictive ge…

Representation Learning

Conformal Prediction for Generative Models via Adaptive Cluster-Based Density Estimation

2026-01-29 · Qidong Yang, Qianyu Julie Zhu, Jonathan Giezendanner, Youssef Marzouk 외 arxiv

Conditional generative models map input variables to complex, high-dimensional distributions, enabling realistic sample generation in a diverse set of domains. A critical challenge with these models is the absence of cal…

Density Estimation

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

2025-04-09 · Wei Huang, Meiyu Liang, Peining Li, Xu Hou 외

Most current MKGC approaches are predominantly based on discriminative models that maximize conditional likelihood. These approaches struggle to efficiently capture the complex connections in real-world knowledge graphs,…

Graph AttentionKnowledge Graph CompletionKnowledge GraphsRepresentation Learning+1

3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising

2026-01-11 · Peiyuan Jing, Yue Yang, Chun-Wun Cheng, Zhenxuan Zhang 외 arxiv

Low-dose Positron Emission Tomography (PET) imaging reduces patient radiation exposure but suffers from increased noise that degrades image quality and diagnostic reliability. Although diffusion models have demonstrated …

Compositional Generative Modeling from Decentralized Data

2026-06-08 · Mashrur M. Morshed, Vishnu Naresh Boddeti arxiv

Learning the compositional nature of the physical world requires joint observation of interacting factors. However, because practical data is often decentralized, these factors are fragmented across isolated silos. Exist…

Conditional Image GenerationFederated Learning