paper-with-me

홈 › Papers

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

2025-04-28 · CVPR 2025 1 · Yan Wang, Baoxiong Jia, Ziyu Zhu, Siyuan Huang

Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked Point-Entity Contrastive learning method for open-vocabulary 3D semantic segmentation that leverages both 3D entity-language alignment and point-entity consistency across different point cloud views to foster entity-specific feature representations. Our method improves semantic discrimination and enhances the differentiation of unique instances, achieving state-of-the-art results on ScanNet for open-vocabulary 3D semantic segmentation and demonstrating superior zero-shot scene understanding capabilities. Extensive fine-tuning experiments on 8 datasets, spanning from low-level perception to high-level reasoning tasks, showcase the potential of learned 3D features, driving consistent performance gains across varied 3D scene understanding tasks. Project website: https://mpec-3d.github.io/

📄 PDF Abstract BibTeX arXiv:2504.19500

Code (0)

등록된 구현이 없습니다.

Tasks

3D Semantic SegmentationContrastive LearningScene UnderstandingSemantic Segmentation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Contrastive Feature Masking Open-Vocabulary Vision Transformer

2023-09-02 · ICCV 2023 1 · Dahun Kim, Anelia Angelova, Weicheng Kuo

We present Contrastive Feature Masking Vision Transformer (CFM-ViT) - an image-text pretraining methodology that achieves simultaneous learning of image- and region-level representation for open-vocabulary object detecti…

Contrastive LearningImage-text Retrievalobject-detectionObject Detection+4

Open-Vocabulary 3D Semantic Segmentation with Foundation Models

2024-01-01 · CVPR 2024 1 · Li Jiang, Shaoshuai Shi, Bernt Schiele

In dynamic 3D environments the ability to recognize a diverse range of objects without the constraints of predefined categories is indispensable for real-world applications. In response to this need we introduce OV3D…

3D Semantic SegmentationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentation+2

Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs

2023-09-27 · Haonan Chang, Kowndinya Boyalakuntla, Shiyang Lu, Siwei Cai 외

We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries. Unlike conventional semantic-…

FormNavigateObjectObject Localization+1

Open-Vocabulary 3D Detection via Image-level Class and Debiased Cross-modal Contrastive Learning

2022-07-05 · Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie 외

Current point-cloud detection methods have difficulty detecting the open-vocabulary objects in the real world, due to their limited generalization capability. Moreover, it is extremely laborious and expensive to collect …

Cloud DetectionContrastive Learning

HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos

2026-03-06 · Tingting Han, Xinsong Tao, Yufei Yin, Min Tan 외 arxiv

Temporal Sentence Grounding in Videos (TSGV) aims to temporally localize segments of a video that correspond to a given natural language query. Despite recent progress, most existing TSGV approaches operate under closed-…

Temporal Sentence Grounding