paper-with-me

홈 › Papers

MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation

2025-02-09 · Zhifei Yang, Keyang Lu, Chao Zhang, Jiaxing Qi, Hanqi Jiang, Ruifei Ma, Shenglin Yin, Yifan Xu, Mingzhe Xing, Zhen Xiao, Jieyi Long, Guangyao Zhai

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. However, current graph-based methods for scene generation are constrained to text-based inputs and exhibit insufficient adaptability to flexible user inputs, hindering the ability to precisely control object geometry. To address this issue, we propose MMGDreamer, a dual-branch diffusion model for scene generation that incorporates a novel Mixed-Modality Graph, visual enhancement module, and relation predictor. The mixed-modality graph allows object nodes to integrate textual and visual modalities, with optional relationships between nodes. It enhances adaptability to flexible user inputs and enables meticulous control over the geometry of objects in the generated scenes. The visual enhancement module enriches the visual fidelity of text-only nodes by constructing visual representations using text embeddings. Furthermore, our relation predictor leverages node representations to infer absent relationships between nodes, resulting in more coherent scene layouts. Extensive experimental results demonstrate that MMGDreamer exhibits superior control of object geometry, achieving state-of-the-art scene generation performance. Project page: https://yangzhifeio.github.io/project/MMGDreamer.

📄 PDF Abstract BibTeX arXiv:2502.05874

Code (1)

yangzhifeio/MMGDreamer 공식 구현 pytorch

Tasks

Scene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Mixed-Modality Dual Face-Hair Retrieval

2026-06-02 · Quoc-Anh Bui-Huynh, Mai-Tuyen Lam, Dai-Anh-Tuan Nguyen, Thanh Duc Ngo arxiv

We introduce Dual Face-Hair Retrieval (DFHR), a new mixed-modality dual-reference task in image retrieval where a query consists of a face image specifying identity and a hairstyle reference expressed as either an image …

Image Retrieval

Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets

2025-09-25 · Team Hunyuan3D, :, Bowen Zhang, Chunchao Guo 외 arxiv

Recent advances in 3D-native generative models have accelerated asset creation for games, film, and design. However, most methods still rely primarily on image or text conditioning and lack fine-grained, cross-modal cont…

Point Clouds

e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings

2026-01-07 · Haonan Chen, Sicheng Gao, Radu Timofte, Tetsuya Sakai 외 arxiv

Modern information systems often involve different types of items, e.g., a text query, an image, a video clip, or an audio segment. This motivates omni-modal embedding models that map heterogeneous modalities into a shar…

GraphVF: Controllable Protein-Specific 3D Molecule Generation with Variational Flow

2023-02-23 · Fang Sun, Zhihao Zhan, Hongyu Guo, Ming Zhang 외

Designing molecules that bind to specific target proteins is a fundamental task in drug discovery. Recent models leverage geometric constraints to generate ligand molecules that bind cohesively with specific protein pock…

3D geometry3D Molecule GenerationDrug Discovery

TriHuman : A Real-time and Controllable Tri-plane Representation for Detailed Human Geometry and Appearance Synthesis

2023-12-08 · Heming Zhu, Fangneng Zhan, Christian Theobalt, Marc Habermann

Creating controllable, photorealistic, and geometrically detailed digital doubles of real humans solely from video data is a key challenge in Computer Graphics and Vision, especially when real-time performance is require…

NeRF