paper-with-me

홈 › Papers

Learning and generalization of one-hidden-layer neural networks, going beyond standard Gaussian data

2022-07-07 · Hongkang Li, Shuai Zhang, Meng Wang

This paper analyzes the convergence and generalization of training a one-hidden-layer neural network when the input features follow the Gaussian mixture model consisting of a finite number of Gaussian distributions. Assuming the labels are generated from a teacher model with an unknown ground truth weight, the learning problem is to estimate the underlying teacher model by minimizing a non-convex risk function over a student neural network. With a finite number of training samples, referred to the sample complexity, the iterations are proved to converge linearly to a critical point with guaranteed generalization error. In addition, for the first time, this paper characterizes the impact of the input distributions on the sample complexity and the learning rate.

📄 PDF Abstract BibTeX arXiv:2207.03615

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Exponential Families

2014-11-10 · Rajesh Ranganath, Linpeng Tang, Laurent Charlin, David M. Blei

We describe \textit{deep exponential families} (DEFs), a class of latent variable models that are inspired by the hidden structures used in deep neural networks. DEFs capture a hierarchy of dependencies between latent va…

Variational Inference

A Recipe for Global Convergence Guarantee in Deep Neural Networks

2021-04-12 · Kenji Kawaguchi, Qingyun Sun

Existing global convergence guarantees of (stochastic) gradient descent do not apply to practical deep networks in the practical regime of deep learning beyond the neural tangent kernel (NTK) regime. This paper proposes …

On the Generalization Power of the Overfitted Three-Layer Neural Tangent Kernel Model

2022-06-04 · Peizhong Ju, Xiaojun Lin, Ness B. Shroff

In this paper, we study the generalization performance of overparameterized 3-layer NTK models. We show that, for a specific set of ground-truth functions (which we refer to as the "learnable set"), the test error of the…

Depth-Attention: Cross-Layer Value Mixing for Language Models

2026-06-03 · Boyi Zeng, Yiqin Hao, Zitong Wang, Shixiang Song 외 arxiv

Self-attention selects information freely across the sequence, but across depth, Transformers merely add each layer's output to the residual stream, so later layers cannot selectively reuse earlier-layer representations.…

Automatic Node Selection for Deep Neural Networks using Group Lasso Regularization

2016-11-17 · Tsubasa Ochiai, Shigeki Matsuda, Hideyuki Watanabe, Shigeru Katagiri

We examine the effect of the Group Lasso (gLasso) regularizer in selecting the salient nodes of Deep Neural Network (DNN) hidden layers by applying a DNN-HMM hybrid speech recognizer to TED Talks speech data. We test two…

General ClassificationPlaying the Game of 2048