MSPL: Multi-Step Pseudo-Labeling for Open-Vocabulary Object Detection
Open-vocabulary object detection (OVD) aims to recognize and localize object categories beyond the training set. Recent approaches leverage vision-language models to generate pseudo-labels using image-text alignment, allowing detectors to generalize to unseen classes without explicit supervision. However, these methods depend heavily on single-step image-text matching, neglecting the intermediate reasoning steps crucial for interpreting semantically complex visual contexts, such as crowding or occlusion. In this paper, we introduce MSPL, a framework that incorporates multi-step visual reasoning into the pseudo-labeling process for OVD. It decomposes complex scene understanding into three interpretable steps-object localization, category recognition, and background grounding-where these intermediate reasoning states serve as rich supervision sources. Extensive experiments on standard OVD evaluation protocols demonstrate that MSPL achieves state-of-the-art performance with superior pseudo-labeling efficiency, outperforming the strong baseline by 9.4 AP50 for novel classes on OV-COCO and improving box and mask APr by 3.2 and 2.2, respectively, on OV-LVIS. Code and models are available at https://github.com/hchoi256/mspl.
Code (0)
등록된 구현이 없습니다.
Tasks
Object LocalizationImage-text matchingScene UnderstandingObject DetectionSimilar Papers 제목 키워드 기반
One-Shot Federated Unsupervised Domain Adaptation with Scaled Entropy Attention and Multi-Source Smoothed Pseudo Labeling
Federated Learning (FL) is a promising approach for privacy-preserving collaborative learning. However, it faces significant challenges when dealing with domain shifts, especially when each client has access only to its …
Domain AdaptationFederated LearningPrivacy PreservingUnsupervised Domain AdaptationOpen-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling
Open-vocabulary instance segmentation aims at segmenting novel classes without mask annotations. It is an important step toward reducing laborious human supervision. Most existing works first pretrain a model on captione…
Instance SegmentationSemantic SegmentationLearning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection
Open-vocabulary object detection (OVOD) has recently gained significant attention as a crucial step toward achieving human-like visual intelligence. Existing OVOD methods extend target vocabulary from pre-defined categor…
Image to textobject-detectionObject DetectionOpen-vocabulary object detection+3Integrating splice-isoform expression into genome-scale models characterizes breast cancer metabolism
Motivation: Despite being often perceived as the main contributors to cell fate and physiology, genes alone cannot predict cellular phenotype. During the process of gene expression, 95% of human genes can code for multip…
PseudoAugment: Learning to Use Unlabeled Data for Data Augmentation in Point Clouds
Data augmentation is an important technique to improve data efficiency and save labeling cost for 3D detection in point clouds. Yet, existing augmentation policies have so far been designed to only utilize labeled data, …
Data AugmentationPseudo Labelvehicle detection