Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection
Open-vocabulary object detection (OVOD) has recently gained significant attention as a crucial step toward achieving human-like visual intelligence. Existing OVOD methods extend target vocabulary from pre-defined categories to open-world by transferring knowledge of arbitrary concepts from vision-language pre-training models to the detectors. While previous methods have shown remarkable successes, they suffer from indirect supervision or limited transferable concepts. In this paper, we propose a simple yet effective method to directly learn region-text alignment for arbitrary concepts. Specifically, the proposed method aims to learn arbitrary image-to-text mapping for pseudo-labeling of arbitrary concepts, named Pseudo-Labeling for Arbitrary Concepts (PLAC). The proposed method shows competitive performance on the standard OVOD benchmark for noun concepts and a large improvement on referring expression comprehension benchmark for arbitrary concepts.
Code (0)
등록된 구현이 없습니다.
Tasks
Image to textobject-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object DetectionReferring ExpressionReferring Expression ComprehensionSimilar Papers 제목 키워드 기반
BSNet: Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation
3D instance segmentation (3DIS) is a crucial task, but point-level annotations are tedious in fully supervised settings. Thus, using bounding boxes (bboxes) as annotations has shown great potential. The current mainstrea…
3D Instance SegmentationDecoderInstance SegmentationSemantic SegmentationLego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and Dr…
Language ModellingLarge Language ModelQuestion AnsweringVisual Question Answering+1Mapping and Generating Classifiers using an Open Chinese Ontology
In languages such as Chinese, classifiers (CLs) play a central role in the quantification of noun-phrases. This can be a problem when generating text from input that does not specify the classifier, as in machine transla…
Machine TranslationTranslationProgressive Representative Labeling for Deep Semi-Supervised Learning
Deep semi-supervised learning (SSL) has experienced significant attention in recent years, to leverage a huge amount of unlabeled data to improve the performance of deep learning with limited labeled data. Pseudo-labelin…
Graph Neural NetworkGoing Beyond Nouns With Vision & Language Models Using Synthetic Data
Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbi…
SentenceVisual Reasoning