paper-with-me

홈 › Papers

Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings

2025-09-16 · Abdalla Arafa, Didier Stricker arxiv

Novel view synthesis has seen significant advancements with 3D Gaussian Splatting (3DGS), enabling real-time photorealistic rendering. However, the inherent fuzziness of Gaussian Splatting presents challenges for 3D scene understanding, restricting its broader applications in AR/VR and robotics. While recent works attempt to learn semantics via 2D foundation model distillation, they inherit fundamental limitations: alpha blending averages semantics across objects, making 3D-level understanding impossible. We propose a paradigm-shifting alternative that bypasses differentiable rendering for semantics entirely. Our key insight is to leverage predecomposed object-level Gaussians and represent each object through multiview CLIP feature aggregation, creating comprehensive "bags of embeddings" that holistically describe objects. This allows: (1) accurate open-vocabulary object retrieval by comparing text queries to object-level (not Gaussian-level) embeddings, and (2) seamless task adaptation: propagating object IDs to pixels for 2D segmentation or to Gaussians for 3D extraction. Experiments demonstrate that our method effectively overcomes the challenges of 3D open-vocabulary object extraction while remaining comparable to state-of-the-art performance in 2D open-vocabulary segmentation, ensuring minimal compromise.

📄 PDF Abstract BibTeX arXiv:2509.12938

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View SynthesisScene Understanding

Similar Papers 제목 키워드 기반

OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding

2024-08-20 · Youjun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang 외

Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem with…

ObjectScene Understanding

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

2026-07-06 · Xianhao Chen, Jiarui Hu, Yuanbo Yang, Xiyu Zhang 외 arxiv

Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have advanced this task by lifting language-a…

Scene Understanding

UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

2024-01-21 · Qingdong He, Jinlong Peng, Zhengkai Jiang, Kai Wu 외

3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain…

Instance SegmentationScene UnderstandingSemantic Segmentation

PLA: Language-Driven Open-Vocabulary 3D Scene Understanding

2022-11-29 · CVPR 2023 1 · Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang 외

Open-vocabulary scene understanding aims to localize and recognize unseen categories beyond the annotated label space. The recent breakthrough of 2D open-vocabulary perception is largely driven by Internet-scale paired i…

3D Open-Vocabulary Instance SegmentationContrastive LearningInstance SegmentationRepresentation Learning+2

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation

2025-05-26 · Zuyao Chen, Jinlin Wu, Zhen Lei, Chang Wen Chen

We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventional methods restrict both object and relat…

Graph GenerationKnowledge DistillationNovel ConceptsOpen-vocabulary object detection+3