Meta-learnt priors slow down catastrophic forgetting in neural networks
Current training regimes for deep learning usually involve exposure to a single task / dataset at a time. Here we start from the observation that in this context the trained model is not given any knowledge of anything outside its (single-task) training distribution, and has thus no way to learn parameters (i.e., feature detectors or policies) that could be helpful to solve other tasks, and to limit future interference with the acquired knowledge, and thus catastrophic forgetting. Here we show that catastrophic forgetting can be mitigated in a meta-learning context, by exposing a neural network to multiple tasks in a sequential manner during training. Finally, we present SeqFOMAML, a meta-learning algorithm that implements these principles, and we evaluate it on sequential learning problems composed by Omniglot and MiniImageNet classification tasks.
Code (1)
Tasks
Meta-LearningSimilar Papers 제목 키워드 기반
Fast Training of Neural Lumigraph Representations using Meta Learning
Novel view synthesis is a long-standing problem in machine learning and computer vision. Significant progress has recently been made in developing neural scene representations and rendering techniques that synthesize pho…
Meta-LearningNeural RenderingNovel View SynthesisSlowing Down Forgetting in Continual Learning
A common challenge in continual learning (CL) is catastrophic forgetting, where the performance on old tasks drops after new, additional tasks are learned. In this paper, we propose a novel framework called ReCL to slow …
Continual LearningIncremental LearningIncremental Meta-Learning via Indirect Discriminant Alignment
Majority of the modern meta-learning methods for few-shot classification tasks operate in two phases: a meta-training phase where the meta-learner learns a generic representation by solving multiple few-shot tasks sample…
Incremental LearningMeta-LearningStochastic Gradient Descent: Going As Fast As Possible But Not Faster
When applied to training deep neural networks, stochastic gradient descent (SGD) often incurs steady progression phases, interrupted by catastrophic episodes in which loss and gradient norm explode. A possible mitigation…
Change Point DetectionPerformance analysis of Zero Black-Derman-Toy interest rate model in catastrophic events: COVID-19 case study
In this paper we continue the research of our recent interest rate tree model called Zero Black-Derman-Toy (ZBDT) model, which includes the possibility of a jump at each step to a practically zero interest rate. This app…