paper-with-me

홈 › Papers

Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation

2026-03-14 · Stefan Ainetter, Thomas Deixelberger, Edoardo A. Dominici, Philipp Drescher, Konstantinos Vardis, Markus Steinberger arxiv

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or scale ambiguity, our approach maintains an absolute world coordinate frame throughout the entire generation process. Starting from a textual scene description, we predict a global 3D layout encoding both semantic and geometric structure, which serves as a guiding proxy for downstream stages. A semantics- and depth-conditioned panoramic diffusion model then synthesizes 360° imagery aligned with the global layout, substantially improving spatial coherence. To explore unobserved regions, we employ a video diffusion model guided by optimized camera trajectories that balances coverage and collision avoidance, achieving up to 10x faster sampling compared to exhaustive path exploration. The generated views are fused using 3D Gaussian Splatting, yielding a consistent and fully navigable 3D scene in absolute scale. GuidedSceneGen enables accurate transfer of object poses and semantic labels from layout to reconstruction, and supports progressive scene expansion without re-alignment. Quantitative results and a user study demonstrate greater 3D consistency and layout plausibility compared to recent panoramic text-to-3D baselines.

📄 PDF Abstract BibTeX arXiv:2603.13910

Code (0)

등록된 구현이 없습니다.

Tasks

Collision AvoidanceScene Generation3D Generation

Similar Papers 제목 키워드 기반

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

2020-06-30 · Fei Yu, Jiji Tang, Weichong Yin, Yu Sun 외

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic co…

AttributePredictionReferring Expression ComprehensionSentence+1

Augmenting Crowd-Sourced 3D Reconstructions Using Semantic Detections

2018-06-01 · CVPR 2018 6 · True Price, Johannes L. Schönberger, Zhen Wei, Marc Pollefeys 외

Image-based 3D reconstruction for Internet photo collections has become a robust technology to produce impressive virtual representations of real-world scenes. However, several fundamental challenges remain for Structure…

3D Reconstruction

VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery

2025-05-05 · Bojin Wu, Jing Chen

We propose a robust method for monocular depth scale recovery. Monocular depth estimation can be divided into two main directions: (1) relative depth estimation, which provides normalized or inverse depth without scale i…

Depth EstimationMonocular Depth Estimation

SGFormer: Semantic Graph Transformer for Point Cloud-based 3D Scene Graph Generation

2023-03-20 · Changsheng Lv, Mengshi Qi, Xia Li, Zhengyuan Yang 외

In this paper, we propose a novel model called SGFormer, Semantic Graph TransFormer for point cloud-based 3D scene graph generation. The task aims to parse a point cloud-based scene into a semantic structural graph, with…

3d scene graph generationGraph EmbeddingGraph GenerationLanguage Modelling+1

Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes

2020-03-15 · ECCV 2020 8 · Liang Liao, Jing Xiao, Zheng Wang, Chia-Wen Lin 외

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic in…

Image InpaintingSemantic SegmentationTexture Synthesis