paper-with-me

홈 › Papers

ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding

2025-12-03 · Lingjun Zhao, Yandong Luo, James Hays, Lu Gan arxiv

We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision foundation models (VFMs). Gaussian-based methods have demonstrated superior performance and computational efficiency across a wide range of scene understanding tasks. However, existing methods either model objects as closed-set semantic Gaussians supervised by annotated 3D labels, neglecting their rendering ability, or learn open-set Gaussian representations via purely 2D self-supervision, leading to degraded geometry and limited to camera-only settings. To fully exploit the potential of Gaussians, we propose a Multi-Modal Gaussian Transformer that enables Gaussians to query features from diverse sensor modalities, and a Shelf-Supervised Learning Paradigm that efficiently optimizes Gaussians with VFM features jointly at 2D image and 3D scene levels. We evaluate ShelfGaussian on various perception and planning tasks. Experiments on Occ3D-nuScenes demonstrate its state-of-the-art zero-shot semantic occupancy prediction performance. ShelfGaussian is further evaluated on an unmanned ground vehicle (UGV) to assess its in the-wild performance across diverse urban scenarios. Project website: https://lunarlab-gatech.github.io/ShelfGaussian/.

📄 PDF Abstract BibTeX arXiv:2512.03370

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyScene Understanding

Similar Papers 제목 키워드 기반

FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

2026-04-30 · Zeyu Jiang, Changqing Zhou, Xingxing Zuo, Changhao Chen arxiv

Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction fr…

Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features

2022-12-20 · Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan, Austin Myers 외

Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detection utilizing pretrained image-text co-…

Action DetectionOptical Flow Estimation

SuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians

2024-12-13 · Siyun Liang, Sen Wang, Kunyi Li, Michael Niemeyer 외

3D Gaussian Splatting has recently gained traction for its efficient training and real-time rendering. While the vanilla Gaussian Splatting representation is mainly designed for view synthesis, more recent works investig…

GPUObject LocalizationScene UnderstandingSegmentation+1

GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane

2024-05-27 · Yansong Qu, Shaohui Dai, Xinyang Li, Jianghang Lin 외

3D open-vocabulary scene understanding, crucial for advancing augmented reality and robotic applications, involves interpreting and locating specific regions within a 3D space as directed by natural language instructions…

3DGSfeature selectionReferring ExpressionReferring Expression Segmentation+1

GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting

2025-08-19 · Elena Alegret, Kunyi Li, Sen Wang, Siyun Liang 외 arxiv

3D scene reconstruction and understanding have gained increasing popularity, yet existing methods still struggle to capture fine-grained, language-aware 3D representations from 2D images. In this paper, we present GALA, …

Contrastive LearningScene Understanding