paper-with-me

홈 › Papers

Open Vocabulary Multi-Label Classification with Dual-Modal Decoder on Aligned Visual-Textual Features

2022-08-19 · Shichao Xu, Yikang Li, Jenhao Hsiao, Chiuman Ho, Zhu Qi

In computer vision, multi-label recognition are important tasks with many real-world applications, but classifying previously unseen labels remains a significant challenge. In this paper, we propose a novel algorithm, Aligned Dual moDality ClaSsifier (ADDS), which includes a Dual-Modal decoder (DM-decoder) with alignment between visual and textual features, for open-vocabulary multi-label classification tasks. Then we design a simple and yet effective method called Pyramid-Forwarding to enhance the performance for inputs with high resolutions. Moreover, the Selective Language Supervision is applied to further enhance the model performance. Extensive experiments conducted on several standard benchmarks, NUS-WIDE, ImageNet-1k, ImageNet-21k, and MS-COCO, demonstrate that our approach significantly outperforms previous methods and provides state-of-the-art performance for open-vocabulary multi-label classification, conventional multi-label classification and an extreme case called single-to-multi label classification where models trained on single-label datasets (ImageNet-1k, ImageNet-21k) are tested on multi-label ones (MS-COCO and NUS-WIDE).

📄 PDF Abstract BibTeX arXiv:2208.09562

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDecoderMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-label zero-shot learning

Similar Papers 제목 키워드 기반

Open Vocabulary Multi-Label Video Classification

2024-07-12 · Rohit Gupta, Mamshad Nayeem Rizve, Jayakrishnan Unnikrishnan, Ashish Tawari 외

Pre-trained vision-language models (VLMs) have enabled significant progress in open vocabulary computer vision tasks such as image classification, object detection and image segmentation. Some recent works have focused o…

Action ClassificationClassificationimage-classificationImage Classification+6

Open Vocabulary Extreme Classification Using Generative Models

2022-05-12 · Findings (ACL) 2022 5 · Daniel Simig, Fabio Petroni, Pouya Yanki, Kashyap Popat 외

The extreme multi-label classification (XMC) task aims at tagging content with a subset of labels from an extremely large label set. The label vocabulary is typically defined in advance by domain experts and assumed to c…

ClassificationExtreme Multi-Label ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+3

Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification

2024-01-02 · Xuelin Zhu, Jian Liu, Dongqi Tang, Jiawei Ge 외

Identifying labels that did not appear during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. To this end, recent studies have attempted to explore the multi-modal knowledge o…

Knowledge DistillationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-label zero-shot learning+1

TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training

2023-12-20 · Yuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li 외

Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish dif…

ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONSemantic Segmentation+3

OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer

2025-03-13 · Jinyang Li, En Yu, Sijia Chen, Wenbing Tao

Open-vocabulary multiple object tracking aims to generalize trackers to unseen categories during training, enabling their application across a variety of real-world scenarios. However, the existing open-vocabulary tracke…

Decodermultimodal interactionMultiple Object TrackingMultiple Object Tracking with Transformer+1