paper-with-me

Papers

Multi-scale Feature Learning Dynamics: Insights for Double Descent

2021-12-06 · Mohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume Lajoie

A key challenge in building theoretical foundations for deep learning is the complex optimization dynamics of neural networks, resulting from the high-dimensional interactions between the large number of network parameters. Such non-trivial dynamics lead to intriguing behaviors such as the phenomenon of "double descent" of the generalization error. The more commonly studied aspect of this phenomenon corresponds to model-wise double descent where the test error exhibits a second descent with increasing model complexity, beyond the classical U-shaped error curve. In this work, we investigate the origins of the less studied epoch-wise double descent in which the test error undergoes two non-monotonous transitions, or descents as the training time increases. By leveraging tools from statistical physics, we study a linear teacher-student setup exhibiting epoch-wise double descent similar to that in deep neural networks. In this setting, we derive closed-form analytical expressions for the evolution of generalization error over training. We find that double descent can be attributed to distinct features being learned at different scales: as fast-learning features overfit, slower-learning features start to fit, resulting in a second descent in test error. We validate our findings through numerical experiments where our theory accurately predicts empirical findings and remains consistent with observations in deep neural networks.

📄 PDF Abstract BibTeX arXiv:2112.03215

Code (1)

nndoubledescent/doubledescent 공식 구현 pytorch

Similar Papers 제목 키워드 기반

When and how epochwise double descent happens

2021-08-26 · Cory Stephenson, Tyler Lee

Deep neural networks are known to exhibit a `double descent' behavior as the number of parameters increases. Recently, it has also been shown that an `epochwise double descent' effect exists in which the generalization e…

ModLaNets: Learning Generalisable Dynamics via Modularity and Physical Inductive Bias

2022-06-24 · Yupu Lu, ShiJie Lin, Guanqi Chen, Jia Pan

Deep learning models are able to approximate one specific dynamical system but struggle at learning generalisable dynamics, where dynamical systems obey the same laws of physics but contain different numbers of elements …

Inductive Bias

Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network

2025-10-21 · Zhanhong He, Hanyu Meng, David Huang, Roberto Togneri arxiv

Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, be…

Beat Tracking

A dynamic view of the double descent

2025-05-03 · Vivek Shripad Borkar

It has been observed by Belkin et al.\ that overparametrized neural networks exhibit a `double descent' phenomenon. That is, as the model complexity, as reflected in the number of features, increases, the training error …

Learning Networked Dynamical System Models with Weak Form and Graph Neural Networks

2024-07-23 · Yin Yu, Daning Huang, Seho Park, Herschel C. Pangborn

This paper presents a sequence of two approaches for the data-driven control-oriented modeling of networked systems, i.e., the systems that involve many interacting dynamical components. First, a novel deep learning appr…

Form