paper-with-me

홈 › Papers

Understanding Prompt Tuning for V-L Models Through the Lens of Neural Collapse

2023-06-28 · Didi Zhu, Zexi Li, Min Zhang, Junkun Yuan, Yunfeng Shao, Jiashuo Liu, Kun Kuang, Yinchuan Li, Chao Wu

Large-scale vision-language (V-L) models have demonstrated remarkable generalization capabilities for downstream tasks through prompt tuning. However, the mechanisms behind the learned text representations are unknown, limiting further generalization gains, especially under class imbalance scenarios. Recent advances in the neural collapse (NC) phenomenon of vision-only models suggest that the optimal representation structure is the simplex ETF, which paves the way to study representations in V-L models. In this paper, we make the first attempt to use NC for examining the representations in V-L models via prompt tuning. It is found that NC optimality of text-to-image representations shows a positive correlation with downstream generalizability, which is more severe under class imbalance settings. To improve the representations, we propose Neural-collapse-anchored Prompt Tuning (NPT), a novel method that learns prompts with text and image representations that satisfy the same simplex ETF. NPT incorporates two regularization terms: language-modality collapse and multi-modality isomorphism; and it is compatible with other prompt tuning methods. Extensive experiments show that NPT can consistently help to improve existing prompt tuning techniques across 11 datasets for both balanced and imbalanced settings.

📄 PDF Abstract BibTeX arXiv:2306.15955

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization

2026-06-29 · Kunal Samanta, Ari Holtzman, Peter West arxiv

The alignment of language models is typically studied through the lens of capability benchmarks, but the dynamics of how models change during post-training remain poorly understood. We argue that the physical sciences, a…

Reinforcement Learning

Collapsed Language Models Promote Fairness

2024-10-06 · Jingxuan Xu, Wuyang Chen, Linyi Li, Yao Zhao 외

To mitigate societal biases implicitly encoded in recent successful pretrained language models, a diverse array of approaches have been proposed to encourage model fairness, focusing on prompting, data augmentation, regu…

Data AugmentationFairnessNatural Language UnderstandingWord Embeddings

Inducer-tuning: Connecting Prefix-tuning and Adapter-tuning

2022-10-26 · Yifan Chen, Devamanyu Hazarika, Mahdi Namazifar, Yang Liu 외

Prefix-tuning, or more generally continuous prompt tuning, has become an essential paradigm of parameter-efficient transfer learning. Using a large pre-trained language model (PLM), prefix-tuning can obtain strong perfor…

Language ModelingLanguage ModellingNatural Language UnderstandingTransfer Learning

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

2025-08-19 · Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng 외 arxiv

Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically igno…

Image Segmentation

Understanding BLOOM: An empirical study on diverse NLP tasks

2022-11-27 · Parag Pravin Dakle, SaiKrishna Rallabandi, Preethi Raghavan

We view the landscape of large language models (LLMs) through the lens of the recently released BLOOM model to understand the performance of BLOOM and other decoder-only LLMs compared to BERT-style encoder-only models. W…

DecoderFew-Shot Text ClassificationQuestion Answeringtext-classification+3