paper-with-me

Papers

Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity

2023-06-28 · Zhenlin Xu, Yi Zhu, Tiffany Deng, Abhay Mittal, Yanbei Chen, Manchen Wang, Paolo Favaro, Joseph Tighe, Davide Modolo

This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although VLMs excel in tasks like image captioning, they face challenges in open-world settings. Our benchmarks test VLMs' consistency in understanding concepts across semantic granularity levels and their response to varying text specificity. Findings show that VLMs favor moderately fine-grained concepts and struggle with specificity, often misjudging texts that differ from their training data. Extensive evaluations reveal limitations in current VLMs, particularly in distinguishing between correct and subtly incorrect descriptions. While fine-tuning offers some improvements, it doesn't fully address these issues, highlighting the need for VLMs with enhanced generalization capabilities for real-world applications. This study provides insights into VLM limitations and suggests directions for developing more robust models.

📄 PDF Abstract BibTeX arXiv:2306.16048

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingImage CaptioningSpecificityZero-Shot Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Prompting Scientific Names for Zero-Shot Species Recognition

2023-10-15 · Shubham Parashar, Zhiqiu Lin, Yanan Li, Shu Kong

Trained on web-scale image-text pairs, Vision-Language Models (VLMs) such as CLIP can recognize images of common objects in a zero-shot fashion. However, it is underexplored how to use CLIP for zero-shot recognition of h…

BenchmarkingZero-Shot Learning

PEVA-Net: Prompt-Enhanced View Aggregation Network for Zero/Few-Shot Multi-View 3D Shape Recognition

2024-04-30 · Dongyun Lin, Yi Cheng, Shangbo Mao, Aiyuan Guo 외

Large vision-language models have impressively promote the performance of 2D visual recognition under zero/few-shot scenarios. In this paper, we focus on exploiting the large vision-language model, i.e., CLIP, to address…

3D Shape RecognitionFew-Shot LearningLanguage ModellingZero-Shot Learning

Benchmarking Multimodal Large Language Models for Face Recognition

2025-10-16 · Hatef Otroshi Shahreza, Sébastien Marcel arxiv

Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance o…

Face Recognition

Benchmarking Foundation Models for Zero-Shot Biometric Tasks

2025-05-30 · Redwan Sony, Parisa Farmanifard, Hamzeh Alzwairy, Nitish Shukla 외

The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across…

AttributeBenchmarkingDeepFake DetectionFace Swapping+2

Vision-Language Models for Vision Tasks: A Survey

2023-04-03 · Jingyi Zhang, Jiaxing Huang, Sheng Jin, Shijian Lu

Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming…

BenchmarkingKnowledge DistillationSurveyTransfer Learning