paper-with-me

홈 › Papers

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

2026-05-26 · Jianzhe Gao, Rui Liu, Wenguan Wang arxiv

Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene representations to enhance spatial awareness, they often neglect the complex 3D geometry and rich semantics in VLN scenarios, limiting the ability to generalize across diverse and unseen environments. To address these challenges, this work proposes a 3D Gaussian Map that represents the environment as a set of differentiable 3D Gaussians and accordingly develops a navigation strategy for VLN. Specifically, Egocentric Scene Map is constructed online by initializing 3D Gaussians from sparse pseudo-lidar point clouds, providing informative geometric priors for scene understanding. Each Gaussian primitive is further enriched through Open-Set Semantic Grouping operation, which groups 3D Gaussians based on their membership in object instances or stuff categories within the open world, resulting in a unified 3D Gaussian Map. Building on this map, Multi-Level Action Prediction strategy, which combines spatial-semantic cues at multiple granularities, is designed to assist agents in decision-making. Extensive experiments conducted on three public benchmarks (i.e., R2R, R4R, and REVERIE) validate the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2605.26500

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationScene UnderstandingPoint Clouds

Similar Papers 제목 키워드 기반

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

2026-06-29 · Jameel Hassan, Yasiru Ranasinghe, Vishal Patel arxiv

3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such a…

Referring ExpressionSpatial Reasoning

ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning

2025-03-30 · CVPR 2025 1 · Zhenyang Liu, Yikai Wang, Sixiao Zheng, Tongying Pan 외

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigat…

3D visual groundingFeature SplattingVision-Language NavigationVisual Grounding

GroupViT: Semantic Segmentation Emerges from Text Supervision

2022-02-22 · CVPR 2022 1 · Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon 외

Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end-to-end deep learning systems, grouping of image regions usually happens impl…

Object DetectionScene UnderstandingSemantic SegmentationTransfer Learning+1

Bayesian Fields: Task-driven Open-Set Semantic Gaussian Splatting

2025-03-07 · Dominic Maggio, Luca Carlone

Open-set semantic mapping requires (i) determining the correct granularity to represent the scene (e.g., how should objects be defined), and (ii) fusing semantic knowledge across multiple 2D observations into an overall …

3D Reconstruction

Open-Vocabulary Semantic Part Segmentation of 3D Human

2025-02-27 · Keito Suzuki, Bang Du, Girish Krishnan, Kunyao Chen 외

3D part segmentation is still an open problem in the field of 3D vision and AR/VR. Due to limited 3D labeled data, traditional supervised segmentation methods fall short in generalizing to unseen shapes and categories. R…

3D Part SegmentationSegmentation