paper-with-me

홈 › Papers

(PASS) Visual Prompt Locates Good Structure Sparsity through a Recurrent HyperNetwork

2024-07-24 · Tianjin Huang, Fang Meng, Li Shen, Fan Liu, Yulong Pei, Mykola Pechenizkiy, Shiwei Liu, Tianlong Chen

Large-scale neural networks have demonstrated remarkable performance in different domains like vision and language processing, although at the cost of massive computation resources. As illustrated by compression literature, structural model pruning is a prominent algorithm to encourage model efficiency, thanks to its acceleration-friendly sparsity patterns. One of the key questions of structural pruning is how to estimate the channel significance. In parallel, work on data-centric AI has shown that prompting-based techniques enable impressive generalization of large language models across diverse downstream tasks. In this paper, we investigate a charming possibility - \textit{leveraging visual prompts to capture the channel importance and derive high-quality structural sparsity}. To this end, we propose a novel algorithmic framework, namely \texttt{PASS}. It is a tailored hyper-network to take both visual prompts and network weight statistics as input, and output layer-wise channel sparsity in a recurrent manner. Such designs consider the intrinsic channel dependency between layers. Comprehensive experiments across multiple network architectures and six datasets demonstrate the superiority of \texttt{PASS} in locating good structural sparsity. For example, at the same FLOPs level, \texttt{PASS} subnetworks achieve $1\%\sim 3\%$ better accuracy on Food101 dataset; or with a similar performance of $80\%$ accuracy, \texttt{PASS} subnetworks obtain $0.35\times$ more speedup than the baselines.

📄 PDF Abstract BibTeX arXiv:2407.17412

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

2026-05-08 · Tao Wang, Shuo Li, Yan Sun, Dongsheng Ding 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimization methods, such as GRPO, typically …

Reinforcement LearningMathematical Reasoning

EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling

2025-10-13 · Daniel Scalena, Leonidas Zotos, Elisabetta Fersini, Malvina Nissim 외 arxiv

With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computation is often required to generate multiple candidate sequences from the same pro…

Visual prompt engineering for video models

2026-07-28 · Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat 외 arxiv

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently bec…

Prompt EngineeringVisual ReasoningImage Editing

How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A

2026-05-09 · YiJie Huang, Yiqun Zhang, Zhuoyue Jia, Xiaocui Yang 외 arxiv

Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling question: as multimodal models grow, how many…

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

2025-06-19 · Yuan Zhang, Chun-Kai Fan, Sicheng Yu, Junwen Pan 외 arxiv

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, performance tends to saturate under single vis…