paper-with-me

홈 › Papers

All You Need is One: Capsule Prompt Tuning with a Single Vector

2025-10-19 · Yiyang Liu, James C. Liang, Heng Fan, Wenhao Yang, Yiming Cui, Xiaotian Han, Lifu Huang, Dongfang Liu, Qifan Wang, Cheng Han arxiv

Prompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning generation with task-aware guidance. Despite its successes, current prompt-based learning methods heavily rely on laborious grid searching for optimal prompt length and typically require considerable number of prompts, introducing additional computational burden. Worse yet, our pioneer findings indicate that the task-aware prompt design is inherently limited by its absence of instance-aware information, leading to a subtle attention interplay with the input sequence. In contrast, simply incorporating instance-aware information as a part of the guidance can enhance the prompt-tuned model performance without additional fine-tuning. Moreover, we find an interesting phenomenon, namely "attention anchor", that incorporating instance-aware tokens at the earliest position of the sequence can successfully preserve strong attention to critical structural information and exhibit more active attention interaction with all input tokens. In light of our observation, we introduce Capsule Prompt-Tuning (CaPT), an efficient and effective solution that leverages off-the-shelf, informative instance semantics into prompt-based learning. Our approach innovatively integrates both instance-aware and task-aware information in a nearly parameter-free manner (i.e., one single capsule prompt). Empirical results demonstrate that our method can exhibit superior performance across various language tasks (e.g., 84.03\% average accuracy on T5-Large), serving as an "attention anchor," while enjoying high parameter efficiency (e.g., 0.003\% of model parameters on Llama3.2-1B).

📄 PDF Abstract BibTeX arXiv:2510.16670

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

No Routing Needed Between Capsules

2020-01-24 · Adam Byerly, Tatiana Kalganova, Ian Dear

Most capsule network designs rely on traditional matrix multiplication between capsule layers and computationally expensive routing mechanisms to deal with the capsule dimensional entanglement that the matrix multiplicat…

Image Classification

InceptionCapsule: Inception-Resnet and CapsuleNet with self-attention for medical image Classification

2024-02-03 · Elham Sadeghnezhad, Sajjad Salem

Initial weighting is significant in deep neural networks because the random selection of weights produces different outputs and increases the probability of overfitting and underfitting. On the other hand, vector-based a…

Classificationimage-classificationImage ClassificationMedical Image Classification+1

Homogeneous Vector Capsules Enable Adaptive Gradient Descent in Convolutional Neural Networks

2019-06-20 · Adam Byerly, Tatiana Kalganova

Capsules are the name given by Geoffrey Hinton to vector-valued neurons. Neural networks traditionally produce a scalar value for an activated neuron. Capsules, on the other hand, produce a vector of values, which Hinton…

ClassificationGeneral Classification

Fast Dynamic Routing Based on Weighted Kernel Density Estimation

2018-05-28 · Suofei Zhang, Wei Zhao, Xiaofu Wu, Quan Zhou

Capsules as well as dynamic routing between them are most recently proposed structures for deep neural networks. A capsule groups data into vectors or matrices as poses rather than conventional scalars to represent speci…

Density EstimationImage Classification

Dynamic Routing Between Capsules

2017-10-26 · NeurIPS 2017 12 · Sara Sabour, Nicholas Frosst, Geoffrey E. Hinton

A capsule is a group of neurons whose activity vector represents the instantiation parameters of a specific type of entity such as an object or an object part. We use the length of the activity vector to represent the pr…

Image Classification