paper-with-me

홈 › Papers

AutoVP: An Automated Visual Prompting Framework and Benchmark

2023-10-12 · Hsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Sijia Liu, Tsung-Yi Ho

Visual prompting (VP) is an emerging parameter-efficient fine-tuning approach to adapting pre-trained vision models to solve various downstream image-classification tasks. However, there has hitherto been little systematic study of the design space of VP and no clear benchmark for evaluating its performance. To bridge this gap, we propose AutoVP, an end-to-end expandable framework for automating VP design choices, along with 12 downstream image-classification tasks that can serve as a holistic VP-performance benchmark. Our design space covers 1) the joint optimization of the prompts; 2) the selection of pre-trained models, including image classifiers and text-image encoders; and 3) model output mapping strategies, including nonparametric and trainable label mapping. Our extensive experimental results show that AutoVP outperforms the best-known current VP methods by a substantial margin, having up to 6.7% improvement in accuracy; and attains a maximum performance increase of 27.5% compared to linear-probing (LP) baseline. AutoVP thus makes a two-fold contribution: serving both as an efficient tool for hyperparameter tuning on VP design choices, and as a comprehensive benchmark that can reasonably be expected to accelerate VP's development. The source code is available at https://github.com/IBM/AutoVP.

📄 PDF Abstract BibTeX arXiv:2310.08381

Code (1)

IBM/AutoVP 공식 구현 pytorch

Tasks

image-classificationImage Classificationparameter-efficient fine-tuningVisual Prompting

Similar Papers 제목 키워드 기반

Benchmarking Human and Automated Prompting in the Segment Anything Model

2024-10-29 · Jorge Quesada, Zoe Fowler, Mohammad Alotaibi, Mohit Prabhushankar 외

The remarkable capabilities of the Segment Anything Model (SAM) for tackling image segmentation tasks in an intuitive and interactive manner has sparked interest in the design of effective visual prompts. Such interest h…

BenchmarkingImage SegmentationSemantic SegmentationVisual Prompting

Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization

2025-02-16 · Wonduk Seo, Seungyong Lee, Daye Kang, Hyunjin An 외

Rapid advancements in Large Language Models (LLMs) have accelerated their integration into automated visualization code generation applications. Despite advancements through few-shot prompting and query expansion, existi…

Code GenerationData VisualizationNatural Language Queries

VRPTEST: Evaluating Visual Referring Prompting in Large Multimodal Models

2023-12-07 · Zongjie Li, Chaozheng Wang, Chaowei Liu, Pingchuan Ma 외

With recent advancements in Large Multimodal Models (LMMs) across various domains, a novel prompting method called visual referring prompting has emerged, showing significant potential in enhancing human-computer interac…

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

2026-04-24 · Advait Tilak, Jiwon Choi, Nazifa Mouli, Wei Le arxiv

The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing benchmarks largely overlooked implausible scenarios and do not measure audio-vis…

Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection

2025-11-16 · Xi Xiao, Zhuxuanzi Wang, Mingqiao Mo, Chen Liu 외 arxiv

The deployment of automated pavement defect detection is often hindered by poor cross-domain generalization. Supervised detectors achieve strong in-domain accuracy but require costly re-annotation for new environments, w…

Domain GeneralizationRoad Damage Detection