LoGoPrompt: Synthetic Text Images Can Be Good Visual Prompts for Vision-Language Models
Prompt engineering is a powerful tool used to enhance the performance of pre-trained models on downstream tasks. For example, providing the prompt "Let's think step by step" improved GPT-3's reasoning accuracy to 63% on MutiArith while prompting "a photo of" filled with a class name enables CLIP to achieve $80$\% zero-shot accuracy on ImageNet. While previous research has explored prompt learning for the visual modality, analyzing what constitutes a good visual prompt specifically for image recognition is limited. In addition, existing visual prompt tuning methods' generalization ability is worse than text-only prompting tuning. This paper explores our key insight: synthetic text images are good visual prompts for vision-language models! To achieve that, we propose our LoGoPrompt, which reformulates the classification objective to the visual prompt selection and addresses the chicken-and-egg challenge of first adding synthetic text images as class-wise visual prompts or predicting the class first. Without any trainable visual prompt parameters, experimental results on 16 datasets demonstrate that our method consistently outperforms state-of-the-art methods in few-shot learning, base-to-new generalization, and domain generalization.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationFew-Shot LearningPrompt EngineeringPrompt LearningVisual Prompt TuningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Estimates of maize plant density from UAV RGB images using Faster-RCNN detection model: impact of the spatial resolution
Early-stage plant density is an essential trait that determines the fate of a genotype under given environmental conditions and management practices. The use of RGB images taken from UAVs may replace traditional visual c…
Generative Adversarial NetworkManagementSuper-ResolutionA Novel Visual Representation on Text Using Diverse Conditional GAN for Visual Recognition
Abstract— Automatic image visual recognition can make full use of largely available images with text descriptions on social media platforms to build large-scale image labeled datasets. In this paper, we propose a nove…
Generative Adversarial NetworkSemantic SegmentationVITAL: A Visual Interpretation on Text with Adversarial Learning for Image Labeling
In this paper, we propose a novel way to interpret text information by extracting visual feature presentation from multiple high-resolution and photo-realistic synthetic images generated by Text-to-image Generative Adver…
Generative Adversarial NetworkHyCIR: Boosting Zero-Shot Composed Image Retrieval with Synthetic Labels
Composed Image Retrieval (CIR) aims to retrieve images based on a query image with text. Current Zero-Shot CIR (ZS-CIR) methods try to solve CIR tasks without using expensive triplet-labeled training datasets. However, t…
Contrastive LearningImage RetrievalImage to textLanguage Modelling+4As Good As A Coin Toss: Human detection of AI-generated images, videos, audio, and audiovisual stimuli
One of the current principal defenses against weaponized synthetic media continues to be the ability of the targeted individual to visually or auditorily recognize AI-generated content when they encounter it. However, as…
Human Detection