paper-with-me

홈 › Papers

The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes

2022-12-23 · Alexander Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz Pehlevan

For small training set sizes $P$, the generalization error of wide neural networks is well-approximated by the error of an infinite width neural network (NN), either in the kernel or mean-field/feature-learning regime. However, after a critical sample size $P^*$, we empirically find the finite-width network generalization becomes worse than that of the infinite width network. In this work, we empirically study the transition from infinite-width behavior to this variance limited regime as a function of sample size $P$ and network width $N$. We find that finite-size effects can become relevant for very small dataset sizes on the order of $P^* \sim \sqrt{N}$ for polynomial regression with ReLU networks. We discuss the source of these effects using an argument based on the variance of the NN's final neural tangent kernel (NTK). This transition can be pushed to larger $P$ by enhancing feature learning or by ensemble averaging the networks. We find that the learning curve for regression with the final NTK is an accurate approximation of the NN learning curve. Using this, we provide a toy model which also exhibits $P^* \sim \sqrt{N}$ scaling and has $P$-dependent benefits from feature learning.

📄 PDF Abstract BibTeX arXiv:2212.12147

Code (1)

Pehlevan-Group/onset-of-variance 공식 구현 jax

Tasks

regression

Methods 이 논문이 사용한 방법론

NTK 설명 없음

Similar Papers 제목 키워드 기반

Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes

2024-05-27 · Zhenfeng Tu, Santiago Aranguri, Arthur Jacot

The training dynamics of linear networks are well studied in two distinct setups: the lazy regime and balanced/active regime, depending on the initialization and width of the network. We provide a surprisingly simple uni…

Lazy-MDPs: Towards Interpretable Reinforcement Learning by Learning When to Act

2022-03-16 · Alexis Jacq, Johan Ferret, Olivier Pietquin, Matthieu Geist

Traditionally, Reinforcement Learning (RL) aims at deciding how to act optimally for an artificial agent. We argue that deciding when to act is equally important. As humans, we drift from default, instinctive or memorize…

Atari GamesDecision Makingreinforcement-learningReinforcement Learning (RL)

The Importance of Being Lazy: Scaling Limits of Continual Learning

2025-06-20 · Jacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann 외

Despite recent efforts, neural networks still struggle to learn in non-stationary environments, and our understanding of catastrophic forgetting (CF) is far from complete. In this work, we perform a systematic study on t…

Continual Learning

How Transformers Get Rich: Approximation and Dynamics Analysis

2024-10-15 · Mingze Wang, Ruoxi Yu, Weinan E, Lei Wu

Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remains limited. A recent work (Elhage et al., 2021) identified a ``rich'' in-co…

In-Context Learning

How connectivity structure shapes rich and lazy learning in neural circuits

2023-10-12 · Yuhan Helena Liu, Aristide Baratin, Jonathan Cornford, Stefan Mihalas 외

In theoretical neuroscience, recent work leverages deep learning tools to explore how some network attributes critically influence its learning dynamics. Notably, initial weight distributions with small (resp. large) var…