paper-with-me

홈 › Papers

Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning

2022-11-06 · Zafir Stojanovski, Karsten Roth, Zeynep Akata

Large pre-trained, zero-shot capable models have shown considerable success both for standard transfer and adaptation tasks, with particular robustness towards distribution shifts. In addition, subsequent fine-tuning can considerably improve performance on a selected downstream task. However, through naive fine-tuning, these zero-shot models lose their generalizability and robustness towards distribution shifts. This is a particular problem for tasks such as Continual Learning (CL), where continuous adaptation has to be performed as new task distributions are introduced sequentially. In this work, we showcase that where fine-tuning falls short to adapt such zero-shot capable models, simple momentum-based weight interpolation can provide consistent improvements for CL tasks in both memory-free and memory-based settings. In particular, we find improvements of over $+4\%$ on standard CL benchmarks, while reducing the error to the upper limit of jointly training on all tasks at once in parts by more than half, allowing the continual learner to inch closer to the joint training limits.

📄 PDF Abstract BibTeX arXiv:2211.03186

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight Optimization

2023-02-01 · Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 외

Contrastive Language-Image Pretraining (CLIP) has demonstrated impressive zero-shot learning abilities for image understanding, yet limited effort has been made to investigate CLIP for zero-shot video recognition. We int…

Action RecognitionContinual LearningVideo RecognitionZero-Shot Learning

Convergence rates of stochastic gradient method with independent sequences of step-size and momentum weight

2024-07-31 · Wen-Liang Hwang

In large-scale learning algorithms, the momentum term is usually included in the stochastic sub-gradient method to improve the learning speed because it can navigate ravines efficiently to reach a local minimum. However,…

Navigate

Zero-Shot Dense Retrieval with Momentum Adversarial Domain Invariant Representation

2021-09-29 · Ji Xin, Chenyan Xiong, Ashwin Srinivasan, Ankita Sharma 외

Dense retrieval (DR) methods conduct text retrieval by first encoding texts in the embedding space and then matching them by nearest neighbor search. This requires strong locality properties from the representation space…

Representation LearningRetrievalText Retrieval

Zero-Shot Dense Retrieval with Momentum Adversarial Domain Invariant Representations

2021-10-14 · Findings (ACL) 2022 5 · Ji Xin, Chenyan Xiong, Ashwin Srinivasan, Ankita Sharma 외

Dense retrieval (DR) methods conduct text retrieval by first encoding texts in the embedding space and then matching them by nearest neighbor search. This requires strong locality properties from the representation space…

Representation LearningRetrievalText Retrieval

mcBERT: Momentum Contrastive Learning with BERT for Zero-Shot Slot Filling

2022-03-24 · Seong-Hwan Heo, WonKee Lee, Jong-Hyeok Lee

Zero-shot slot filling has received considerable attention to cope with the problem of limited available data for the target domain. One of the important factors in zero-shot learning is to make the model learn generaliz…

Contrastive Learningslot-fillingSlot FillingZero-Shot Learning+1