paper-with-me

홈 › Papers

A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future

2023-07-18 · Chaoyang Zhu, Long Chen

As the most fundamental scene understanding tasks, object detection and segmentation have made tremendous progress in deep learning era. Due to the expensive manual labeling cost, the annotated categories in existing datasets are often small-scale and pre-defined, i.e., state-of-the-art fully-supervised detectors and segmentors fail to generalize beyond the closed vocabulary. To resolve this limitation, in the last few years, the community has witnessed an increasing attention toward Open-Vocabulary Detection (OVD) and Segmentation (OVS). By ``open-vocabulary'', we mean that the models can classify objects beyond pre-defined categories. In this survey, we provide a comprehensive review on recent developments of OVD and OVS. A taxonomy is first developed to organize different tasks and methodologies. We find that the permission and usage of weak supervision signals can well discriminate different methodologies, including: visual-semantic space mapping, novel visual feature synthesis, region-aware training, pseudo-labeling, knowledge distillation, and transfer learning. The proposed taxonomy is universal across different tasks, covering object detection, semantic/instance/panoptic segmentation, 3D and video understanding. The main design principles, key challenges, development routes, methodology strengths, and weaknesses are thoroughly analyzed. In addition, we benchmark each task along with the vital components of each method in appendix and updated online at https://github.com/seanzhuh/awesome-open-vocabulary-detection-and-segmentation. Finally, several promising directions are provided and discussed to stimulate future research.

📄 PDF Abstract BibTeX arXiv:2307.09220

Code (1)

seanzhuh/awesome-open-vocabulary-detection-and-segmentation 공식 구현 pytorch

Tasks

Knowledge Distillationobject-detectionObject DetectionPanoptic SegmentationScene UnderstandingSegmentationTransfer LearningVideo Understanding

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Towards Open Vocabulary Learning: A Survey

2023-06-28 · Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan 외

In the field of visual scene understanding, deep neural networks have made impressive advancements in various core tasks like segmentation, tracking, and detection. However, most approaches operate on the close-set assum…

Open Set LearningOut-of-Distribution DetectionScene UnderstandingSegmentation+2

OpenSD: Unified Open-Vocabulary Segmentation and Detection

2023-12-10 · Shuai Li, Minghan Li, Pengfei Wang, Lei Zhang

Recently, a few open-vocabulary methods have been proposed by employing a unified architecture to tackle generic segmentation and detection tasks. However, their performance still lags behind the task-specific models due…

DecoderPrompt LearningSegmentationZero Shot Segmentation

Rethinking Evaluation Metrics of Open-Vocabulary Segmentaion

2023-11-06 · Hao Zhou, Tiancheng Shen, Xu Yang, Hai Huang 외

In this paper, we highlight a problem of evaluation metrics adopted in the open-vocabulary segmentation. That is, the evaluation process still heavily relies on closed-set metrics on zero-shot or cross-dataset pipelines …

Segmentation

A Survey on Training-free Open-Vocabulary Semantic Segmentation

2025-05-28 · Naomi Kombol, Ivan Martinović, Siniša Šegvić

Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional methods strive to train models up from scr…

Multi-modal ClassificationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation+1

A Simple Framework for Open-Vocabulary Segmentation and Detection

2023-03-14 · ICCV 2023 1 · Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu 외

We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we fi…

Instance SegmentationPanoptic SegmentationSegmentationSemantic Segmentation+1