paper-with-me

Papers

WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion

2026-03-24 · Manuel-Andreas Schneider, Angela Dai arxiv

Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -video approaches struggle to maintain scene- and object-level consistency beyond a limited environment scale without a persistent, explicit geometric representation. We thus present a geometry-first approach that decouples this complex problem of large-scale 3D scene synthesis into its structural composition, represented as a mesh scaffold, and realistic appearance synthesis, which leverages powerful image synthesis models conditioned on the mesh scaffold. From an input text description, we first construct a mesh capturing the environment's geometry (walls, floors, etc.), and then use image synthesis, segmentation and object reconstruction to populate the mesh structure with objects in realistic layouts. This mesh scaffold is then rendered to condition image synthesis, providing a structural backbone for consistent appearance generation. This enables scalable, arbitrarily-sized 3D scenes of high object richness and diversity, combining robust 3D consistency with photorealistic detail. We believe this marks a significant step toward generating truly environment-scale, immersive 3D worlds.

📄 PDF Abstract BibTeX arXiv:2603.22972

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Similar Papers 제목 키워드 기반

RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation

2024-12-11 · CVPR 2025 1 · Mingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova 외

Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address this, we introduce RoomTour3D, a video-i…

3D ReconstructionDiversityVision and Language Navigation

A Database-Driven Framework for 3D Level Generation with LLMs

2025-08-25 · Kaijie Xu, Clark Verbrugge arxiv

Procedural Content Generation for 3D game levels faces challenges in balancing spatial coherence, navigational functionality, and adaptable gameplay progression across multi-floor environments. This paper introduces a no…

MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments

2025-03-18 · Zhixuan Liu, Haokun Zhu, Rui Chen, Jonathan Francis 외

We introduce a novel diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel Multi-view Overlapped Scene Align…

DenoisingPrivacy Preserving

Frankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-Plane

2024-03-24 · Han Yan, Yang Li, Zhennan Wu, Shenzhou Chen 외

We present Frankenstein, a diffusion-based framework that can generate semantic-compositional 3D scenes in a single pass. Unlike existing methods that output a single, unified 3D shape, Frankenstein simultaneously genera…

DenoisingObject Rearrangement

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

2025-12-12 · Wentang Chen, Shougao Zhang, Yiman Zhang, Tianhao Zhou 외 arxiv

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on im…

Indoor Scene SynthesisScene GenerationSemantic Parsing