paper-with-me

Papers

LLM meets Vision-Language Models for Zero-Shot One-Class Classification

2024-03-31 · Yassir Bendou, Giulia Lioi, Bastien Pasdeloup, Lukas Mauch, Ghouthi Boukli Hacene, Fabien Cardinaux, Vincent Gripon

We consider the problem of zero-shot one-class visual classification, extending traditional one-class classification to scenarios where only the label of the target class is available. This method aims to discriminate between positive and negative query samples without requiring examples from the target class. We propose a two-step solution that first queries large language models for visually confusing objects and then relies on vision-language pre-trained models (e.g., CLIP) to perform classification. By adapting large-scale vision benchmarks, we demonstrate the ability of the proposed method to outperform adapted off-the-shelf alternatives in this setting. Namely, we propose a realistic benchmark where negative query samples are drawn from the same original dataset as positive ones, including a granularity-controlled version of iNaturalist, where negative samples are at a fixed distance in the taxonomy tree from the positive ones. To our knowledge, we are the first to demonstrate the ability to discriminate a single category from other semantically related ones using only its label.

📄 PDF Abstract BibTeX arXiv:2404.00675

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationOne-Class Classification

Similar Papers 제목 키워드 기반

SLIP: Self-supervision meets Language-Image Pre-training

2021-12-23 · Norman Mu, Alexander Kirillov, David Wagner, Saining Xie

Recent work has shown that self-supervised pre-training leads to improvements over supervised learning on challenging visual recognition tasks. CLIP, an exciting new approach to learning with language supervision, demons…

Multi-Task LearningRepresentation LearningSelf-Supervised Learning

CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement

2023-10-21 · Mohammadreza Salehi, Mehrdad Farajtabar, Maxwell Horton, Fartash Faghri 외

Contrastive language image pretraining (CLIP) is a standard method for training vision-language models. While CLIP is scalable, promptable, and robust to distribution shifts on image classification tasks, it lacks object…

Depth Estimationimage-classificationImage ClassificationObject Localization+3

Cross-Modal Retrieval Meets Inference:Improving Zero-Shot Classification with Cross-Modal Retrieval

2023-08-29 · Seongha Eom, Namgyu Ho, Jaehoon Oh, Se-Young Yun

Contrastive language-image pre-training (CLIP) has demonstrated remarkable zero-shot classification ability, namely image classification using novel text labels. Existing works have attempted to enhance CLIP by fine-tuni…

Cross-Modal Retrievalimage-classificationImage ClassificationRetrieval+3

DiffCLIP: Differential Attention Meets CLIP

2025-03-09 · Hasan Abed Al Kader Hammoud, Bernard Ghanem

We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for large language models to amplify relevant co…

Language ModelingLanguage ModellingRetrievalzero-shot-classification+1

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

2024-12-20 · CVPR 2025 1 · Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre 외

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual fe…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentationzero-shot-classification+1