paper-with-me

홈 › Papers

Learning Open-vocabulary Semantic Segmentation Models From Natural Language Supervision

2023-01-22 · CVPR 2023 1 · Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, Weidi Xie

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled image-text pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slot-attention based binding module, and aligns the group tokens to the corresponding caption embedding. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given the group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, which encourages the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on three benchmark datasets, PASCAL VOC 2012, PASCAL Context, and COCO Object. Our model achieves superior segmentation results over the state-of-the-art method by using only 3\% data (4M vs 134M) for pre-training. Code and pre-trained models will be released for future research.

📄 PDF Abstract BibTeX arXiv:2301.09121

Code (1)

Jazzcharles/OVSegmentor 공식 구현 pytorch

Tasks

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

2026-03-06 · Sandeep Inuganti, Hideaki Kanayama, Kanta Shimizu, Mahdi Chamseddine 외 arxiv

Semantic segmentation across visual modalities such as 3D point clouds and panoramic images remains a challenging task, primarily due to the scarcity of annotated data and the limited adaptability of fixed-label models. …

Open Vocabulary Semantic Segmentation3D Semantic SegmentationScene UnderstandingPoint Clouds

SENSE: Stereo OpEN Vocabulary SEmantic Segmentation

2026-04-17 · Thomas Campagnolo, Ezio Malis, Philippe Martinet, Gaétan Bahl arxiv

Open-vocabulary semantic segmentation enables models to segment objects or image regions beyond fixed class sets, offering flexibility in dynamic environments. However, existing methods often rely on single-view images a…

Open Vocabulary Semantic SegmentationScene UnderstandingSpatial Reasoning

OpenLex3D: A New Evaluation Benchmark for Open-Vocabulary 3D Scene Representations

2025-03-25 · Christina Kassab, Sacha Morin, Martin Büchner, Matías Mattamala 외

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, the evaluation of these representations is limited to closed-set semantics that do not…

3D Semantic SegmentationScene UnderstandingSemantic Segmentation

DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes

2025-06-02 · Jiajun Jiang, Yiming Zhu, Zirui Wu, Jie Song

We introduce DualMap, an online open-vocabulary mapping system that enables robots to understand and navigate dynamically changing environments through natural language queries. Designed for efficient semantic mapping an…

Natural Language QueriesNavigateRobot Navigation

RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation

2025-05-21 · Naman Patel, Prashanth Krishnamurthy, Farshad Khorrami

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. W…

GPUNatural Language QueriesObjectobject-detection+3