Learning One-hidden-layer Neural Networks on Gaussian Mixture Models with Guaranteed Generalizability
We analyze the learning problem of fully connected neural networks with the sigmoid activation function for binary classification from the setup of model estimation. The outputs are assumed to be generated by a ground-truth neural network with the unknown parameters, and the learning objective is to estimate the ground-truth model parameters by minimizing a non-convex cross-entropy loss function of the training data. Instead of following the conventional and restrictive assumption in the literature that the input features follow the standard Gaussian distribution, this paper, for the first time, analyzes a more general and practical scenario that the input features follow a Gaussian mixture model of a finite number of Gaussian distributions of various mean and variance. We propose a gradient descent algorithm with a tensor initialization approach and show that our algorithm converges linearly to a critical point that has a diminishing distance to the ground-truth model with guaranteed generalizability. We characterize the required number of samples for successful convergence, referred to as the sample complexity, as a function of the parameters of the Gaussian mixture model. We prove analytically that when any mean or variance in the mixture model is large, or when all variances are close to zero, the sample complexity increases, and the convergence slows down, indicating a more challenging learning problem. Although focusing on one-hidden-layer neural networks, this paper provides the first theoretical analyses of the impact of the parameters of the input distributions on the learning performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Binary ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning and generalization of one-hidden-layer neural networks, going beyond standard Gaussian data
This paper analyzes the convergence and generalization of training a one-hidden-layer neural network when the input features follow the Gaussian mixture model consisting of a finite number of Gaussian distributions. Assu…
Convergence of neural networks to Gaussian mixture distribution
We give a proof that, under relatively mild conditions, fully-connected feed-forward deep random neural networks converge to a Gaussian mixture distribution as only the width of the last hidden layer goes to infinity. We…
Efficient Learning of Convolution Weights as Gaussian Mixture Model Posteriors
In this paper, we showed that the feature map of a convolution layer is equivalent to the unnormalized log posterior of a special kind of Gaussian mixture for image modeling. Then we expanded the model to drive diverse f…
FiMReSt: Finite Mixture of Multivariate Regulated Skew-t Kernels -- A Flexible Probabilistic Model for Multi-Clustered Data with Asymmetrically-Scattered Non-Gaussian Kernels
Recently skew-t mixture models have been introduced as a flexible probabilistic modeling technique taking into account both skewness in data clusters and the statistical degree of freedom (S-DoF) to improve modeling gene…
Deep neural networks with dependent weights: Gaussian Process mixture limit, heavy tails, sparsity and compressibility
This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegati…
Gaussian ProcessesRepresentation Learning