paper-with-me

홈 › Papers

Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection

2023-12-04 · Sunghun Kang, Junbum Cha, Jonghwan Mun, Byungseok Roh, Chang D. Yoo

Open-vocabulary object detection (OVOD) has recently gained significant attention as a crucial step toward achieving human-like visual intelligence. Existing OVOD methods extend target vocabulary from pre-defined categories to open-world by transferring knowledge of arbitrary concepts from vision-language pre-training models to the detectors. While previous methods have shown remarkable successes, they suffer from indirect supervision or limited transferable concepts. In this paper, we propose a simple yet effective method to directly learn region-text alignment for arbitrary concepts. Specifically, the proposed method aims to learn arbitrary image-to-text mapping for pseudo-labeling of arbitrary concepts, named Pseudo-Labeling for Arbitrary Concepts (PLAC). The proposed method shows competitive performance on the standard OVOD benchmark for noun concepts and a large improvement on referring expression comprehension benchmark for arbitrary concepts.

📄 PDF Abstract BibTeX arXiv:2312.02103

Code (0)

등록된 구현이 없습니다.

Tasks

Image to textobject-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object DetectionReferring ExpressionReferring Expression Comprehension

Similar Papers 제목 키워드 기반

BSNet: Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation

2024-03-22 · CVPR 2024 1 · Jiahao Lu, Jiacheng Deng, Tianzhu Zhang

3D instance segmentation (3DIS) is a crucial task, but point-level annotations are tedious in fully supervised settings. Thus, using bounding boxes (bboxes) as annotations has shown great potential. The current mainstrea…

3D Instance SegmentationDecoderInstance SegmentationSemantic Segmentation

Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models

2023-11-23 · Saman Motamed, Danda Pani Paudel, Luc van Gool

Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and Dr…

Language ModellingLarge Language ModelQuestion AnsweringVisual Question Answering+1

Mapping and Generating Classifiers using an Open Chinese Ontology

2016-01-01 · GWC 2016 1 · Luis Morgado Da Costa, Francis Bond, Helena Gao

In languages such as Chinese, classifiers (CLs) play a central role in the quantification of noun-phrases. This can be a problem when generating text from input that does not specify the classifier, as in machine transla…

Machine TranslationTranslation

Progressive Representative Labeling for Deep Semi-Supervised Learning

2021-08-13 · Xiaopeng Yan, Riquan Chen, Litong Feng, Jingkang Yang 외

Deep semi-supervised learning (SSL) has experienced significant attention in recent years, to leverage a huge amount of unlabeled data to improve the performance of deep learning with limited labeled data. Pseudo-labelin…

Graph Neural Network

Going Beyond Nouns With Vision & Language Models Using Synthetic Data

2023-03-30 · ICCV 2023 1 · Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh 외

Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbi…

SentenceVisual Reasoning