Tilting the Odds at the Lottery: the Interplay of Overparameterisation and Curricula in Neural Networks
A wide range of empirical and theoretical works have shown that overparameterisation can amplify the performance of neural networks. According to the lottery ticket hypothesis, overparameterised networks have an increased chance of containing a sub-network that is well-initialised to solve the task at hand. A more parsimonious approach, inspired by animal learning, consists in guiding the learner towards solving the task by curating the order of the examples, i.e. providing a curriculum. However, this learning strategy seems to be hardly beneficial in deep learning applications. In this work, we undertake an analytical study that connects curriculum learning and overparameterisation. In particular, we investigate their interplay in the online learning setting for a 2-layer network in the XOR-like Gaussian Mixture problem. Our results show that a high degree of overparameterisation -- while simplifying the problem -- can limit the benefit from curricula, providing a theoretical account of the ineffectiveness of curricula in deep learning.
Code (1)
Similar Papers 제목 키워드 기반
Optimal Lottery Tickets via Subset Sum: Logarithmic Over-Parameterization is Sufficient
The strong lottery ticket hypothesis (LTH) postulates that one can approximate any target neural network by only pruning the weights of a sufficiently over-parameterized random network. A recent work by Malach et al. [M…
Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
The strong {\it lottery ticket hypothesis} (LTH) postulates that one can approximate any target neural network by only pruning the weights of a sufficiently over-parameterized random network. A recent work by Malach et a…
Overparameterisation and worst-case generalisation: friend or foe?
Overparameterised neural networks have demonstrated the remarkable ability to perfectly fit training samples, while still generalising to unseen test samples. However, several recent works have revealed that such models'…
Structured PredictionThe dynamic interplay between in-context and in-weight learning in humans and neural networks
Human learning embodies a striking duality: sometimes, we appear capable of following logical, compositional rules and benefit from structured curricula (e.g., in formal education), while other times, we rely on an incre…
BlockingIn-Context LearningLow-level Pose Control of Tilting Multirotor for Wall Perching Tasks Using Reinforcement Learning
Recently, needs for unmanned aerial vehicles (UAVs) that are attachable to the wall have been highlighted. As one of the ways to address the need, researches on various tilting multirotors that can increase maneuverabili…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)