paper-with-me

홈 › Papers

What does CLIP know about a red circle? Visual prompt engineering for VLMs

2023-04-13 · ICCV 2023 1 · Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi

Large-scale Vision-Language Models, such as CLIP, learn powerful image-text representations that have found numerous applications, from zero-shot classification to text-to-image generation. Despite that, their capabilities for solving novel discriminative tasks via prompting fall behind those of large language models, such as GPT-3. Here we explore the idea of visual prompt engineering for solving computer vision tasks beyond classification by editing in image space instead of text. In particular, we discover an emergent ability of CLIP, where, by simply drawing a red circle around an object, we can direct the model's attention to that region, while also maintaining global information. We show the power of this simple approach by achieving state-of-the-art in zero-shot referring expressions comprehension and strong performance in keypoint localization tasks. Finally, we draw attention to some potential ethical concerns of large language-vision models.

📄 PDF Abstract BibTeX arXiv:2304.06712

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationPrompt EngineeringText to Image GenerationText-to-Image Generationzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
Multi-Head Attention 설명 없음
Attention 설명 없음
Weight Decay 설명 없음
Adam 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.

Similar Papers 제목 키워드 기반

Sequence and Circle: Exploring the Relationship Between Patches

2022-10-18 · Zhengyang Yu, Jochen Triesch

The vision transformer (ViT) has achieved state-of-the-art results in various vision tasks. It utilizes a learnable position embedding (PE) mechanism to encode the location of each image patch. However, it is presently u…

Tune-An-Ellipse: CLIP Has Potential to Find What You Want

2024-01-01 · CVPR 2024 1 · Jinheng Xie, Songhe Deng, Bing Li, Haozhe Liu 외

Visual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle commonly used for highlighting can guide CLIP's attention to the surrounding regio…

ObjectReferring ExpressionReferring Expression ComprehensionVisual Prompting

What does CLIP know about peeling a banana?

2024-04-18 · Claudia Cuttano, Gabriele Rosi, Gabriele Trivigno, Giuseppe Averta

Humans show an innate capability to identify tools to support specific actions. The association between objects parts and the actions they facilitate is usually named affordance. Being able to segment objects parts depen…

Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo

2026-04-13 · Artem Gadzhiev, Andrew Kislov arxiv

Providing AI agents with reliable long-term memory that does not hallucinate remains an open problem. Current approaches to memory for LLM agents -- sliding windows, summarization, embedding-based RAG, and flat fact extr…

Adversarial Robustness

CLIPDrawX: Primitive-based Explanations for Text Guided Sketch Synthesis

2023-12-04 · Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri, Anjan Dutta

With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transformations on simple geometric primitives …