paper-with-me

Papers

Unleashing Semantic and Geometric Priors for 3D Scene Completion

2025-08-19 · Shiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers, John See, Cong Yang arxiv

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric priors, which forces the model to make a trade-off between conflicting demands and limits its overall performance. To tackle these challenges, we propose FoundationSSC, a novel framework that performs dual decoupling at both the source and pathway levels. At the source level, we introduce a foundation encoder that provides rich semantic feature priors for the semantic branch and high-fidelity stereo cost volumes for the geometric branch. At the pathway level, these priors are refined through specialised, decoupled pathways, yielding superior semantic context and depth distributions. Our dual-decoupling design produces disentangled and refined inputs, which are then utilised by a hybrid view transformation to generate complementary 3D features. Additionally, we introduce a novel Axis-Aware Fusion (AAF) module that addresses the often-overlooked challenge of fusing these features by anisotropically merging them into a unified representation. Extensive experiments demonstrate the advantages of FoundationSSC, achieving simultaneous improvements in both semantic and geometric metrics, surpassing prior bests by +0.23 mIoU and +2.03 IoU on SemanticKITTI. Additionally, we achieve state-of-the-art performance on SSCBench-KITTI-360, with 21.78 mIoU and 48.61 IoU.

📄 PDF Abstract BibTeX arXiv:2508.13601

Code (0)

등록된 구현이 없습니다.

Tasks

3D Semantic Scene CompletionAutonomous Driving

Results from the Paper

RankTaskDatasetModelMetrics
#62 3D Semantic Scene Completion SemanticKITTI FoundationSSC mIoU: 2.03

Similar Papers 제목 키워드 기반

VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion

2025-03-08 · Meng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin 외

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity cau…

3D Semantic Scene CompletionAutonomous DrivingLanguage ModelingLanguage Modelling+1

VOIC: Visible-Occluded Integrated Guidance for 3D Semantic Scene Completion

2025-12-22 · Zaidao Han, Risa Higashita, Jiang Liu arxiv

Camera-based 3D Semantic Scene Completion (SSC) is a critical task for autonomous driving and robotic scene understanding. It aims to infer a complete 3D volumetric representation of both semantics and geometry from a si…

3D Semantic Scene CompletionSemantic SegmentationScene UnderstandingAutonomous Driving

2D Semantic-Guided Semantic Scene Completion

2024-09-12 · International Journal of Computer Vision (IJCV) 2024 9 · Xianzhu Liu, Haozhe Xie, Shengping Zhang, Hongxun Yao 외

Semantic scene completion (SSC) aims to simultaneously perform scene completion (SC) and predict semantic categories of a 3D scene from a single depth and/or RGB image. Most existing SSC methods struggle to handle compl…

2D Semantic Segmentation3D Semantic Scene CompletionMissing ValuesSemantic Segmentation

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

2026-03-19 · Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia 외 arxiv

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions ty…

Scene UnderstandingSpatial ReasoningVideo Generation

Geospatial-Prior Guidance for 3D Semantic Scene Completion

2026-08-04 · Meng Wang, Shougao Zhang, Wenzhe He, Ruihui Li 외 arxiv

Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-a…

3D Semantic Scene Completion