paper-with-me

홈 › Papers

GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic Fields

2024-04-01 · CVPR 2024 1 · Yunsong Wang, Hanlin Chen, Gim Hee Lee

Recent advancements in vision-language foundation models have significantly enhanced open-vocabulary 3D scene understanding. However, the generalizability of existing methods is constrained due to their framework designs and their reliance on 3D data. We address this limitation by introducing Generalizable Open-Vocabulary Neural Semantic Fields (GOV-NeSF), a novel approach offering a generalizable implicit representation of 3D scenes with open-vocabulary semantics. We aggregate the geometry-aware features using a cost volume, and propose a Multi-view Joint Fusion module to aggregate multi-view features through a cross-view attention mechanism, which effectively predicts view-specific blending weights for both colors and open-vocabulary features. Remarkably, our GOV-NeSF exhibits state-of-the-art performance in both 2D and 3D open-vocabulary semantic segmentation, eliminating the need for ground truth semantic labels or depth priors, and effectively generalize across scenes and datasets without fine-tuning.

📄 PDF Abstract BibTeX arXiv:2404.00931

Code (1)

wangys16/gov-nesf 공식 구현 pytorch

Tasks

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationScene UnderstandingSemantic Segmentation

Similar Papers 제목 키워드 기반

NeSF: Neural Semantic Fields for Generalizable Semantic Segmentation of 3D Scenes

2021-11-25 · Suhani Vora, Noha Radwan, Klaus Greff, Henning Meyer 외

We present NeSF, a method for producing 3D semantic fields from posed RGB images alone. In place of classical 3D representations, our method builds on recent work in implicit neural scene representations wherein 3D struc…

3D Semantic SegmentationSegmentationSemantic Segmentation

GNeSF: Generalizable Neural Semantic Fields

2023-10-24 · NeurIPS 2023 11

3D scene segmentation based on neural implicit representation has emerged recently with the advantage of training only on 2D supervision. However, existing approaches still requires expensive per-scene optimization that …

3D Semantic SegmentationScene SegmentationSegmentationSemantic Segmentation

Neural Structure Fields with Application to Crystal Structure Autoencoders

2022-12-08 · Naoya Chiba, Yuta Suzuki, Tatsunori Taniai, Ryo Igarashi 외

Representing crystal structures of materials to facilitate determining them via neural networks is crucial for enabling machine-learning applications involving crystal structure estimation. Among these applications, the …

Learning Generalizable Feature Fields for Mobile Manipulation

2024-03-12 · Ri-Zhao Qiu, Yafei Hu, Yuchen Song, Ge Yang 외

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while unders…

Novel View Synthesis

Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images

2025-08-05 · Xiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam 외 arxiv

Reconstructing and semantically interpreting 3D scenes from sparse 2D views remains a fundamental challenge in computer vision. Conventional methods often decouple semantic understanding from reconstruction or necessitat…

3D Semantic SegmentationNovel View Synthesis3D Reconstruction