paper-with-me

홈 › Papers

Nonasymptotic theory for two-layer neural networks: Beyond the bias-variance trade-off

2021-06-09 · Huiyuan Wang, Wei Lin

Large neural networks have proved remarkably effective in modern deep learning practice, even in the overparametrized regime where the number of active parameters is large relative to the sample size. This contradicts the classical perspective that a machine learning model must trade off bias and variance for optimal generalization. To resolve this conflict, we present a nonasymptotic generalization theory for two-layer neural networks with ReLU activation function by incorporating scaled variation regularization. Interestingly, the regularizer is equivalent to ridge regression from the angle of gradient-based optimization, but plays a similar role to the group lasso in controlling the model complexity. By exploiting this "ridge-lasso duality," we obtain new prediction bounds for all network widths, which reproduce the double descent phenomenon. Moreover, the overparametrized minimum risk is lower than its underparametrized counterpart when the signal is strong, and is nearly minimax optimal over a suitable class of functions. By contrast, we show that overparametrized random feature models suffer from the curse of dimensionality and thus are suboptimal.

📄 PDF Abstract BibTeX arXiv:2106.04795

Code (0)

등록된 구현이 없습니다.

Tasks

Vocal Bursts Valence Prediction

Similar Papers 제목 키워드 기반

A Near Complete Nonasymptotic Generalization Theory For Multilayer Neural Networks: Beyond the Bias-Variance Tradeoff

2025-03-03 · Hao Yu, Xiangyang Ji

We propose a first near complete (that will make explicit sense in the main text) nonasymptotic generalization theory for multilayer neural networks with arbitrary Lipschitz activations and general Lipschitz loss functio…

Beyond Consistency: Inference for the Relative risk functional in Deep Nonparametric Cox Models

2026-03-25 · Sattwik Ghosal, Xuran Meng, Yi Li arxiv

There remain theoretical gaps in deep neural network estimators for the nonparametric Cox proportional hazards model. In particular, it is unclear how gradient-based optimization error propagates to population risk under…

Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models

2020-10-26 · Jason W. Rocks, Pankaj Mehta

The bias-variance trade-off is a central concept in supervised learning. In classical statistics, increasing the complexity of a model (e.g., number of parameters) reduces bias but also increases variance. Until recently…

Improving Kernel-Based Nonasymptotic Simultaneous Confidence Bands

2024-01-28 · Balázs Csanád Csáji, Bálint Horváth

The paper studies the problem of constructing nonparametric simultaneous confidence bands with nonasymptotic and distribition-free guarantees. The target function is assumed to be band-limited and the approach is based o…

On Nonasymptotic Confidence Intervals for Treatment Effects in Randomized Experiments

2026-01-16 · Ricardo J. Sandoval, Sivaraman Balakrishnan, Avi Feller, Michael I. Jordan 외 arxiv

We study nonasymptotic (finite-sample) confidence intervals for treatment effects in randomized experiments. In the existing literature, the effective sample sizes of nonasymptotic confidence intervals tend to be looser …