paper-with-me

홈 › Papers

Neural Networks Remember More: The Power of Parameter Isolation and Combination

2025-02-16 · Biqing Zeng, Zehan Li, Aladdin Ayesh

Catastrophic forgetting is a pervasive issue for pre-trained language models (PLMs) during continual learning, where models lose previously acquired knowledge when sequentially trained on a series of tasks. The model's ability to retain old tasks is referred to as stability, while its adaptability to new tasks is called plasticity. Therefore, the key to solving this problem is to find a trade-off between the plasticity and stability of the model. To address this issue, in this paper, we propose a novel method to achieve a balance between model stability and plasticity, thereby mitigating catastrophic forgetting. More specifically, our proposed approach leverages parameter isolation and a subsequent combination strategy. Initially, in the training stage, the model adapts to each downstream task via a parameter isolation method to prevent potential interference among different tasks. We then combine all trained parameters, which contain acquired knowledge, using the task arithmetic method and finally apply them to the backbone model. Empirical evaluations on continual language learning benchmarks substantiate the effectiveness of our approach, revealing a marked enhancement over existing state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:2502.10966

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningTask Arithmetic

Similar Papers 제목 키워드 기반

Forgetting to Remember: A Scalable Incremental Learning Framework for Cross-Task Blind Image Quality Assessment

2022-09-15 · Rui Ma, Qingbo Wu, King Ngi Ngan, Hongliang Li 외

Recent years have witnessed the great success of blind image quality assessment (BIQA) in various task-specific scenarios, which present invariable distortion types and evaluation criteria. However, due to the rigid stru…

Image Quality AssessmentIncremental LearningNo-Reference Image Quality Assessment

Those Aren't Your Memories, They're Somebody Else's: Seeding Misinformation in Chat Bot Memories

2023-04-06 · Conor Atkins, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Ian Wood 외

One of the new developments in chit-chat bots is a long-term memory mechanism that remembers information from past conversations for increasing engagement and consistency of responses. The bot is designed to extract know…

Misinformation

ResRep: Lossless CNN Pruning via Decoupling Remembering and Forgetting

2020-07-07 · ICCV 2021 10 · Xiaohan Ding, Tianxiang Hao, Jianchao Tan, Ji Liu 외

We propose ResRep, a novel method for lossless channel pruning (a.k.a. filter pruning), which slims down a CNN by reducing the width (number of output channels) of convolutional layers. Inspired by the neurobiology resea…

Learning to Remember Translation History with a Continuous Cache

2017-11-26 · TACL 2018 1 · Zhaopeng Tu, Yang Liu, Shuming Shi, Tong Zhang

Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a …

Machine TranslationNMTTranslation

ImpressLearn: Continual Learning via Combined Task Impressions

2022-10-05 · Dhrupad Bhardwaj, Julia Kempe, Artem Vysogorets, Angela M. Teng 외

This work proposes a new method to sequentially train deep neural networks on multiple tasks without suffering catastrophic forgetting, while endowing it with the capability to quickly adapt to unseen tasks. Starting fro…

Continual Learningimage-classificationImage ClassificationTransfer Learning