paper-with-me

홈 › Papers

Generalization of Overparametrized Deep Neural Network Under Noisy Observations

2021-09-29 · ICLR 2022 4 · Namjoon Suh, Hyunouk Ko, Xiaoming Huo

We study the generalization properties of the overparameterized deep neural network (DNN) with Rectified Linear Unit (ReLU) activations. Under the non-parametric regression framework, it is assumed that the ground-truth function is from a reproducing kernel Hilbert space (RKHS) induced by a neural tangent kernel (NTK) of ReLU DNN, and a dataset is given with the noises. Without a delicate adoption of early stopping, we prove that the overparametrized DNN trained by vanilla gradient descent does not recover the ground-truth function. It turns out that the estimated DNN's $L_{2}$ prediction error is bounded away from $0$. As a complement of the above result, we show that the $\ell_{2}$-regularized gradient descent enables the overparametrized DNN achieve the minimax optimal convergence rate of the $L_{2}$ prediction error, without early stopping. Notably, the rate we obtained is faster than $\mathcal{O}(n^{-1/2})$ known in the literature.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Regularization Matters: A Nonparametric Perspective on Overparametrized Neural Network

2020-07-06 · Tianyang Hu, Wenjia Wang, Cong Lin, Guang Cheng

Overparametrized neural networks trained by gradient descent (GD) can provably overfit any training data. However, the generalization guarantee may not hold for noisy data. From a nonparametric perspective, this paper st…

Architecture independent generalization bounds for overparametrized deep ReLU networks

2025-04-08 · Thomas Chen, Chun-Kai Kevin Chien, Patricia Muñoz Ewald, Andrew G. Moore

We prove that overparametrized neural networks are able to generalize with a test error that is independent of the level of overparametrization, and independent of the Vapnik-Chervonenkis (VC) dimension. We prove explici…

Generalization Bounds

Nesterov acceleration despite very noisy gradients

2023-02-10 · Kanan Gupta, Jonathan W. Siegel, Stephan Wojtowytsch

We present a generalization of Nesterov's accelerated gradient descent algorithm. Our algorithm (AGNES) provably achieves acceleration for smooth convex and strongly convex minimization tasks with noisy gradient estimate…

Small Data, Big Decisions: Model Selection in the Small-Data Regime

2020-09-26 · ICML 2020 1 · Jorg Bornschein, Francesco Visin, Simon Osindero

Highly overparametrized neural networks can display curiously strong generalization performance - a phenomenon that has recently garnered a wealth of theoretical and empirical research in order to better understand it. I…

Model Selection

Nonasymptotic theory for two-layer neural networks: Beyond the bias-variance trade-off

2021-06-09 · Huiyuan Wang, Wei Lin

Large neural networks have proved remarkably effective in modern deep learning practice, even in the overparametrized regime where the number of active parameters is large relative to the sample size. This contradicts th…

Vocal Bursts Valence Prediction