paper-with-me

홈 › Papers

Learning One-hidden-layer Neural Networks on Gaussian Mixture Models with Guaranteed Generalizability

2021-01-01 · Hongkang Li, Shuai Zhang, Meng Wang

We analyze the learning problem of fully connected neural networks with the sigmoid activation function for binary classification from the setup of model estimation. The outputs are assumed to be generated by a ground-truth neural network with the unknown parameters, and the learning objective is to estimate the ground-truth model parameters by minimizing a non-convex cross-entropy loss function of the training data. Instead of following the conventional and restrictive assumption in the literature that the input features follow the standard Gaussian distribution, this paper, for the first time, analyzes a more general and practical scenario that the input features follow a Gaussian mixture model of a finite number of Gaussian distributions of various mean and variance. We propose a gradient descent algorithm with a tensor initialization approach and show that our algorithm converges linearly to a critical point that has a diminishing distance to the ground-truth model with guaranteed generalizability. We characterize the required number of samples for successful convergence, referred to as the sample complexity, as a function of the parameters of the Gaussian mixture model. We prove analytically that when any mean or variance in the mixture model is large, or when all variances are close to zero, the sample complexity increases, and the convergence slows down, indicating a more challenging learning problem. Although focusing on one-hidden-layer neural networks, this paper provides the first theoretical analyses of the impact of the parameters of the input distributions on the learning performance.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

Learning and generalization of one-hidden-layer neural networks, going beyond standard Gaussian data

2022-07-07 · Hongkang Li, Shuai Zhang, Meng Wang

This paper analyzes the convergence and generalization of training a one-hidden-layer neural network when the input features follow the Gaussian mixture model consisting of a finite number of Gaussian distributions. Assu…

Convergence of neural networks to Gaussian mixture distribution

2022-04-26 · Yasuhiko Asao, Ryotaro Sakamoto, Shiro Takagi

We give a proof that, under relatively mild conditions, fully-connected feed-forward deep random neural networks converge to a Gaussian mixture distribution as only the width of the last hidden layer goes to infinity. We…

Efficient Learning of Convolution Weights as Gaussian Mixture Model Posteriors

2024-01-30 · Lifan Liang

In this paper, we showed that the feature map of a convolution layer is equivalent to the unnormalized log posterior of a special kind of Gaussian mixture for image modeling. Then we expanded the model to drive diverse f…

FiMReSt: Finite Mixture of Multivariate Regulated Skew-t Kernels -- A Flexible Probabilistic Model for Multi-Clustered Data with Asymmetrically-Scattered Non-Gaussian Kernels

2023-05-15 · Sarmad Mehrdad, S. Farokh Atashzar

Recently skew-t mixture models have been introduced as a flexible probabilistic modeling technique taking into account both skewness in data clusters and the statistical degree of freedom (S-DoF) to improve modeling gene…

Deep neural networks with dependent weights: Gaussian Process mixture limit, heavy tails, sparsity and compressibility

2022-05-17 · Hoil Lee, Fadhel Ayed, Paul Jung, Juho Lee 외

This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegati…

Gaussian ProcessesRepresentation Learning