paper-with-me

홈 › Papers

NavCrafter: Exploring 3D Scenes from a Single Image

2026-04-03 · Hongbo Duan, Peiyu Zhuang, Yi Liu, Zhengyang Zhang, Yuxin Zhang, Pengting Luo, Fangming Liu, Xueqian Wang arxiv

Creating flexible 3D scenes from a single image is vital when direct 3D data acquisition is costly or impractical. We introduce NavCrafter, a novel framework that explores 3D scenes from a single image by synthesizing novel-view video sequences with camera controllability and temporal-spatial consistency. NavCrafter leverages video diffusion models to capture rich 3D priors and adopts a geometry-aware expansion strategy to progressively extend scene coverage. To enable controllable multi-view synthesis, we introduce a multi-stage camera control mechanism that conditions diffusion models with diverse trajectories via dual-branch camera injection and attention modulation. We further propose a collision-aware camera trajectory planner and an enhanced 3D Gaussian Splatting (3DGS) pipeline with depth-aligned supervision, structural regularization and refinement. Extensive experiments demonstrate that NavCrafter achieves state-of-the-art novel-view synthesis under large viewpoint shifts and substantially improves 3D reconstruction fidelity.

📄 PDF Abstract BibTeX arXiv:2604.02828

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction

Similar Papers 제목 키워드 기반

Structured Prediction of Unobserved Voxels From a Single Depth Image

2016-06-01 · CVPR 2016 6 · Michael Firman, Oisin Mac Aodha, Simon Julier, Gabriel J. Brostow

Building a complete 3D model of a scene, given only a single depth image, is underconstrained. To gain a full volumetric model, one needs either multiple views, or a single view together with a library of unambiguous 3D …

Structured Prediction

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

2024-03-09 · Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad 외

Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose t…

Image GenerationRetrieval

IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes

2021-12-10 · ICLR 2022 4 · Qi Li, Kaichun Mo, Yanchao Yang, Hang Zhao 외

Building embodied intelligent agents that can interact with 3D indoor environments has received increasing research attention in recent years. While most works focus on single-object or agent-object visual functionality …

Object

SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model

2025-03-18 · Xinqing Li, Ruiqi Song, Qingyu Xie, Ye Wu 외

With the rapid advancement of autonomous driving technology, a lack of data has become a major obstacle to enhancing perception model accuracy. Researchers are now exploring controllable data generation using world model…

Autonomous DrivingImage GenerationScene Generation

Joint Learning of Intrinsic Images and Semantic Segmentation

2018-07-31 · ECCV 2018 9 · Anil S. Baslamisli, Thomas T. Groenestege, Partha Das, Hoang-An Le 외

Semantic segmentation of outdoor scenes is problematic when there are variations in imaging conditions. It is known that albedo (reflectance) is invariant to all kinds of illumination effects. Thus, using reflectance ima…

Intrinsic Image DecompositionSegmentationSemantic Segmentation