paper-with-me

홈 › Papers

Task Addition and Weight Disentanglement in Closed-Vocabulary Models

2025-11-18 · Adam Hazimeh, Alessandro Favero, Pascal Frossard arxiv

Task arithmetic has recently emerged as a promising method for editing pre-trained \textit{open-vocabulary} models, offering a cost-effective alternative to standard multi-task fine-tuning. However, despite the abundance of \textit{closed-vocabulary} models that are not pre-trained with language supervision, applying task arithmetic to these models remains unexplored. In this paper, we deploy and study task addition in closed-vocabulary image classification models. We consider different pre-training schemes and find that \textit{weight disentanglement} -- the property enabling task arithmetic -- is a general consequence of pre-training, as it appears in different pre-trained closed-vocabulary models. In fact, we find that pre-trained closed-vocabulary vision transformers can also be edited with task arithmetic, achieving high task addition performance and enabling the efficient deployment of multi-task models. Finally, we demonstrate that simple linear probing is a competitive baseline to task addition. Overall, our findings expand the applicability of task arithmetic to a broader class of pre-trained models and open the way for more efficient use of pre-trained models in diverse settings.

📄 PDF Abstract BibTeX arXiv:2511.14569

Code (0)

등록된 구현이 없습니다.

Tasks

Image Classification

Similar Papers 제목 키워드 기반

Set-Based Face Recognition Beyond Disentanglement: Burstiness Suppression With Variance Vocabulary

2023-04-13 · Jiong Wang, Zhou Zhao, Fei Wu

Set-based face recognition (SFR) aims to recognize the face sets in the unconstrained scenario, where the appearance of same identity may change dramatically with extreme variances (e.g., illumination, pose, expression).…

DisentanglementFace Recognition

Opening the Vocabulary of Egocentric Actions

2023-08-22 · NeurIPS 2023 11 · Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations …

Action RecognitionObjectOpen Vocabulary Action Recognition

Scaling Open-Vocabulary Action Detection

2025-04-04 · Zhen Hao Sia, Yogesh Singh Rawat

In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending…

Action DetectionMultiple Action DetectionOpen Vocabulary Action DetectionSpatio-Temporal Action Localization+2

Open-vocabulary Panoptic Segmentation with Embedding Modulation

2023-03-20 · ICCV 2023 1 · Xi Chen, Shuang Li, Ser-Nam Lim, Antonio Torralba 외

Open-vocabulary image segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whe…

Image SegmentationOpen Vocabulary Panoptic SegmentationPanoptic SegmentationSegmentation+1

Text Attribute Control via Closed-Loop Disentanglement

2023-12-01 · Lei Sha, Thomas Lukasiewicz

Changing an attribute of a text without changing the content usually requires to first disentangle the text into irrelevant attributes and content representations. After that, in the inference phase, the representation o…

AttributeContrastive LearningDisentanglementSentence