paper-with-me

홈 › Papers

Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation

2025-10-09 · Shohei Enomoto arxiv

Visual prompting (VP) has emerged as a promising parameter-efficient fine-tuning approach for adapting pre-trained vision models to downstream tasks without modifying model parameters. Despite offering advantages like negligible computational overhead and compatibility with black-box models, conventional VP methods typically achieve lower accuracy than other adaptation approaches. Our analysis reveals two critical limitations: the restricted expressivity of simple additive transformation and a tendency toward overfitting when the parameter count increases. To address these challenges, we propose ACAVP (Affine, Color, and Additive Visual Prompting), which enhances VP's expressive power by introducing complementary transformation operations: affine transformation for creating task-specific prompt regions while preserving original image information, and color transformation for emphasizing task-relevant visual features. Additionally, we identify that overfitting is a critical issue in VP training and introduce TrivialAugment as an effective data augmentation, which not only benefits our approach but also significantly improves existing VP methods, with performance gains of up to 12 percentage points on certain datasets. This demonstrates that appropriate data augmentation is universally beneficial for VP training. Extensive experiments across twelve diverse image classification datasets with two different model architectures demonstrate that ACAVP achieves state-of-the-art accuracy among VP methods, surpasses linear probing in average accuracy, and exhibits superior robustness to distribution shifts, all while maintaining minimal computational overhead during inference. Our code is available at https://github.com/s-enmt/ACAVP.

📄 PDF Abstract BibTeX arXiv:2510.07823

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningImage ClassificationData Augmentation

Similar Papers 제목 키워드 기반

Exploring Visual Prompts for Adapting Large-Scale Models

2022-03-31 · Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, Phillip Isola

We investigate the efficacy of visual prompting to adapt large-scale models in vision. Following the recent approach from prompt tuning and adversarial reprogramming, we learn a single image perturbation such that a froz…

Visual Prompting

Context-Aware Visual Prompting: Automating Geospatial Web Dashboards with Large Language Models and Agent Self-Validation for Decision Support

2025-10-10 · Haowen Xu, Jose Tupayachi, Xiao-Ying Yu arxiv

The development of web-based geospatial dashboards for risk analysis and decision support is often challenged by the difficulty in visualization of big, multi-dimensional environmental data, implementation complexity, an…

Code GenerationDecision Making

I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction

2024-07-19 · Zaiqiao Meng, Hao Zhou, Yifang Chen

Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. Howeve…

3D ReconstructionSpatial ReasoningVisual Reasoning

Interpretable dimensionality reduction using weighted linear transformation

2025-03-26 · Adv. Artif. Intell. Mach. Learn. 2025 3 · Erik Bergh

Dimensionality reduction techniques are fundamental for analyzing and visualizing high-dimensional data. With established methods like t-SNE and PCA presenting a trade-off between representational power and interpretabil…

Dimensionality Reduction

Interpretable non-linear dimensionality reduction using gaussian weighted linear transformation

2025-04-24 · Erik Bergh

Dimensionality reduction techniques are fundamental for analyzing and visualizing high-dimensional data. With established methods like t-SNE and PCA presenting a trade-off between representational power and interpretabil…

Dimensionality Reduction