paper-with-me

홈 › Papers

Unleashing the Power of Visual Prompting At the Pixel Level

2022-12-20 · Junyang Wu, Xianhang Li, Chen Wei, Huiyu Wang, Alan Yuille, Yuyin Zhou, Cihang Xie

This paper presents a simple and effective visual prompting method for adapting pre-trained models to downstream recognition tasks. Our method includes two key designs. First, rather than directly adding together the prompt and the image, we treat the prompt as an extra and independent learnable component. We show that the strategy of reconciling the prompt and the image matters, and find that warping the prompt around a properly shrinked image empirically works the best. Second, we re-introduce two "old tricks" commonly used in building transferable adversarial examples, i.e., input diversity and gradient normalization, into visual prompting. These techniques improve optimization and enable the prompt to generalize better. We provide extensive experimental results to demonstrate the effectiveness of our method. Using a CLIP model, our prompting method sets a new record of 82.8% average accuracy across 12 popular classification datasets, substantially surpassing the prior art by +5.6%. It is worth noting that this prompting performance already outperforms linear probing by +2.1% and can even match fully fine-tuning in certain datasets. In addition, our prompting method shows competitive performance across different data scales and against distribution shifts. The code is publicly available at https://github.com/UCSC-VLAA/EVP.

📄 PDF Abstract BibTeX arXiv:2212.10556

Code (1)

ucsc-vlaa/evp 공식 구현 pytorch

Tasks

DiversityVisual Prompting

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

3DAxiesPrompts: Unleashing the 3D Spatial Task Capabilities of GPT-4V

2023-12-15 · Dingning Liu, Xiaomeng Dong, Renrui Zhang, Xu Luo 외

In this work, we present a new visual prompting method called 3DAxiesPrompts (3DAP) to unleash the capabilities of GPT-4V in performing 3D spatial tasks. Our investigation reveals that while GPT-4V exhibits proficiency i…

3D Object Detectionobject-detectionObject DetectionVisual Prompting

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

2025-11-03 · Wenqi Liang, Gan Sun, Yao He, Jiahua Dong 외 arxiv

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image-text-action data and remain limite…

Scene Understanding

Fine-Grained Visual Prompting

2023-06-07 · NeurIPS 2023 11 · Lingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang 외

Vision-Language Models (VLMs), such as CLIP, have demonstrated impressive zero-shot transfer capabilities in image-level visual perception. However, these models have shown limited performance in instance-level tasks tha…

Visual Prompting

LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation

2024-11-30 · Huadong Tang, Youpeng Zhao, Yan Huang, Min Xu 외

It is widely agreed that open-vocabulary-based approaches outperform classical closed-set training solutions for recognizing unseen objects in images for semantic segmentation. Existing open-vocabulary approaches leverag…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation

Explicit Visual Prompting for Low-Level Structure Segmentations

2023-03-20 · CVPR 2023 1 · Weihuang Liu, Xi Shen, Chi-Man Pun, Xiaodong Cun

We consider the generic problem of detecting low-level structures in images, which includes segmenting the manipulated parts, identifying out-of-focus pixels, separating shadow regions, and detecting concealed objects. W…

Camouflaged Object SegmentationDefocus Blur DetectionForeground SegmentationImage Manipulation Detection+3