Probing the Prompting of CLIP on Human Faces
Large-scale multimodal models such as CLIP have caught great attention due to their generalization capability. CLIP can take free-form text prompts, but the performance varies with different text prompt manipulations, which is considered unpredictable. In this paper, we conduct a controlled study to understand how CLIP perceives images with different forms of text prompts, particularly on human facial attributes. We find that (1) using the prompt starter "a photo of" can guide the model to allocate higher attention weights to human faces, leading to better classification performance; (2) CLIP model is better at aligning information from shorter text prompts, as additional textual details shift away the attention from key words; (3) properly adding punctuation or removing stop words in the text prompt can shift attention to target information. Our practice on facial attributes shed light on the design of reliable text prompts for CLIP in other tasks.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting
In Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these V-FER models cannot deal with unknown classes that are preval…
Facial Expression RecognitionMulti-Task LearningOpen Set LearningPrompt Learning+2Language-Guided Invariance Probing of Vision-Language Models
Recent vision-language models (VLMs) such as CLIP, OpenCLIP, EVA02-CLIP and SigLIP achieve strong zero-shot performance, but it is unclear how reliably they respond to controlled linguistic perturbations. We introduce La…
Image-text matchingUnleashing the Power of Visual Prompting At the Pixel Level
This paper presents a simple and effective visual prompting method for adapting pre-trained models to downstream recognition tasks. Our method includes two key designs. First, rather than directly adding together the pro…
DiversityVisual PromptingProbe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection
Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but existing approaches impose restrictive inter…
Video Anomaly DetectionProbing via Prompting
Probing is a popular method to discern what linguistic information is contained in the representations of pre-trained language models. However, the mechanism of selecting the probe model has recently been subject to inte…
DiagnosticLanguage ModelingLanguage Modelling