paper-with-me

홈 › Papers

Nearly Optimal Bounds for Cyclic Forgetting

2023-09-21 · NeurIPS 2023 11

We provide theoretical bounds on the forgetting quantity in the continual learning setting for linear tasks, where each round of learning corresponds to projecting onto a linear subspace. For a cyclic task ordering on $T$ tasks repeated $m$ times each, we prove the best known upper bound of $O(T^2/m)$ on the forgetting. Notably, our bound holds uniformly over all choices of tasks and is independent of the ambient dimension. Our main technical contribution is a characterization of the union of all numerical ranges of products of $T$ (real or complex) projections as a sinusoidal spiral, which may be of independent interest.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification

2025-04-17 · Hyunji Jung, Hanseul Cho, Chulhee Yun

We study continual learning on multiple linear classification tasks by sequentially running gradient descent (GD) for a fixed budget of iterations per task. When all tasks are jointly linearly separable and are presented…

Continual LearningTransfer Learning

How catastrophic can catastrophic forgetting be in linear regression?

2022-05-19 · Itay Evron, Edward Moroshko, Rachel Ward, Nati Srebro 외

To better understand catastrophic forgetting, we study fitting an overparameterized linear model to a sequence of tasks with different input distributions. We analyze how much the model forgets the true labels of earlier…

Continual Learningregression

Distributional Equivalence and Structure Learning for Bow-free Acyclic Path Diagrams

2015-08-07 · Christopher Nowzohour, Marloes H. Maathuis, Robin J. Evans, Peter Bühlmann

We consider the problem of structure learning for bow-free acyclic path diagrams (BAPs). BAPs can be viewed as a generalization of linear Gaussian DAG models that allow for certain hidden variables. We present a first me…

Cyclical Learning Rates for Training Neural Networks

2015-06-03 · Leslie N. Smith

It is known that the learning rate is the most important hyper-parameter to tune for training deep neural networks. This paper describes a new method for setting the learning rate, named cyclical learning rates, which pr…

Nearly Optimal Algorithms with Sublinear Computational Complexity for Online Kernel Regression

2023-06-14 · Junfan Li, Shizhong Liao

The trade-off between regret and computational cost is a fundamental problem for online kernel regression, and previous algorithms worked on the trade-off can not keep optimal regret bounds at a sublinear computational c…

regression