paper-with-me

홈 › Papers

Generalization performance of narrow one-hidden layer networks in the teacher-student setting

2025-07-01 · Rodrigo Pérez Ortiz, Gibbs Nwemadji, Jean Barbier, Federica Gerace, Alessandro Ingrosso, Clarissa Lauditi, Enrico M. Malatesta arxiv

Understanding the generalization properties of neural networks on simple input-output distributions is key to explaining their performance on real datasets. The classical teacher-student setting, where a network is trained on data generated by a teacher model, provides a canonical theoretical test bed. In this context, a complete theoretical characterization of fully connected one-hidden-layer networks with generic activation functions remains missing. In this work, we develop a general framework for such networks with large width, yet much smaller than the input dimension. Using methods from statistical physics, we derive closed-form expressions for the typical performance of both finite-temperature (Bayesian) and empirical risk minimization estimators in terms of a small number of order parameters. We uncover a transition to a specialization phase, where hidden neurons align with teacher features once the number of samples becomes sufficiently large and proportional to the number of network parameters. Our theory accurately predicts the generalization error of networks trained on regression and classification tasks using either noisy full-batch gradient descent (Langevin dynamics) or deterministic full-batch gradient descent.

📄 PDF Abstract BibTeX arXiv:2507.00629

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Online Learning for the Random Feature Model in the Student-Teacher Framework

2023-03-24 · Roman Worschech, Bernd Rosenow

Deep neural networks are widely used prediction algorithms whose performance often improves as the number of weights increases, leading to over-parametrization. We consider a two-layered neural network whose first layer …

Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions

2020-06-27 · NeurIPS 2020 12 · Stefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka Zdeborová

We study the dynamics of optimization and the generalization properties of one-hidden layer neural networks with quadratic activation function in the over-parametrized regime where the layer width $m$ is larger than the …

Learning and generalization of one-hidden-layer neural networks, going beyond standard Gaussian data

2022-07-07 · Hongkang Li, Shuai Zhang, Meng Wang

This paper analyzes the convergence and generalization of training a one-hidden-layer neural network when the input features follow the Gaussian mixture model consisting of a finite number of Gaussian distributions. Assu…

Continual Learning in the Teacher-Student Setup: Impact of Task Similarity

2021-07-09 · Sebastian Lee, Sebastian Goldt, Andrew Saxe

Continual learning-the ability to learn many tasks in sequence-is critical for artificial learning systems. Yet standard training methods for deep networks often suffer from catastrophic forgetting, where learning new ta…

Continual Learning

FitNets: Hints for Thin Deep Nets

2014-12-19 · Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang 외

While depth tends to improve network performances, it also makes gradient-based training more difficult since deeper networks tend to be more non-linear. The recently proposed knowledge distillation approach is aimed at …

Knowledge Distillation