paper-with-me

홈 › Papers

Parameter-Level Soft-Masking for Continual Learning

2023-06-26 · Tatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke, Gyuhak Kim, Bing Liu

Existing research on task incremental learning in continual learning has primarily focused on preventing catastrophic forgetting (CF). Although several techniques have achieved learning with no CF, they attain it by letting each task monopolize a sub-network in a shared network, which seriously limits knowledge transfer (KT) and causes over-consumption of the network capacity, i.e., as more tasks are learned, the performance deteriorates. The goal of this paper is threefold: (1) overcoming CF, (2) encouraging KT, and (3) tackling the capacity problem. A novel technique (called SPG) is proposed that soft-masks (partially blocks) parameter updating in training based on the importance of each parameter to old tasks. Each task still uses the full network, i.e., no monopoly of any part of the network by any task, which enables maximum KT and reduction in capacity usage. To our knowledge, this is the first work that soft-masks a model at the parameter-level for continual learning. Extensive experiments demonstrate the effectiveness of SPG in achieving all three objectives. More notably, it attains significant transfer of knowledge not only among similar tasks (with shared knowledge) but also among dissimilar tasks (with little shared knowledge) while mitigating CF.

📄 PDF Abstract BibTeX arXiv:2306.14775

Code (1)

uic-liu-lab/spg 공식 구현 pytorch

Tasks

Continual LearningIncremental LearningTransfer Learning

Similar Papers 제목 키워드 기반

Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning

2023-09-26 · Hoyong Kim, Minchan Kwon, Kangil Kim

In replay-based methods for continual learning, replaying input samples in episodic memory has shown its effectiveness in alleviating catastrophic forgetting. However, the potential key factor of cross-entropy loss with …

Continual LearningIncremental Learning

Soft-TransFormers for Continual Learning

2024-11-25 · Haeyong Kang, Chang D. Yoo

Inspired by Well-initialized Lottery Ticket Hypothesis (WLTH), which provides suboptimal fine-tuning solutions, we propose a novel fully fine-tuned continual learning (CL) method referred to as Soft-TransFormers (Soft-TF…

class-incremental learningClass Incremental LearningContinual LearningIncremental Learning

Sub-network Discovery and Soft-masking for Continual Learning of Mixed Tasks

2023-10-13 · Zixuan Ke, Bing Liu, Wenhan Xiong, Asli Celikyilmaz 외

Continual learning (CL) has two main objectives: preventing catastrophic forgetting (CF) and encouraging knowledge transfer (KT). The existing literature mainly focused on overcoming CF. Some work has also been done on K…

Continual LearningTransfer Learning

Continual Pre-training of Language Models

2023-02-07 · Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi 외

Language models (LMs) have been instrumental for the rapid advance of natural language processing. This paper studies continual pre-training of LMs, in particular, continual domain-adaptive pre-training (or continual DAP…

Continual LearningContinual PretrainingGeneral KnowledgeTransfer Learning

Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages

2025-01-10 · Aloka Fernando, Surangika Ranathunga

Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM) objective are commonly being used for cross-lingual tasks such as bitext mining. However, the performance of these mode…

Language ModellingSentiment Analysis