paper-with-me

홈 › Papers

On Energy-Based Models with Overparametrized Shallow Neural Networks

2021-04-15 · Carles Domingo-Enrich, Alberto Bietti, Eric Vanden-Eijnden, Joan Bruna

Energy-based models (EBMs) are a simple yet powerful framework for generative modeling. They are based on a trainable energy function which defines an associated Gibbs measure, and they can be trained and sampled from via well-established statistical tools, such as MCMC. Neural networks may be used as energy function approximators, providing both a rich class of expressive models as well as a flexible device to incorporate data structure. In this work we focus on shallow neural networks. Building from the incipient theory of overparametrized neural networks, we show that models trained in the so-called "active" regime provide a statistical advantage over their associated "lazy" or kernel regime, leading to improved adaptivity to hidden low-dimensional structure in the data distribution, as already observed in supervised learning. Our study covers both maximum likelihood and Stein Discrepancy estimators, and we validate our theoretical results with numerical experiments on synthetic data.

📄 PDF Abstract BibTeX arXiv:2104.07531

Code (1)

CDEnrich/ebms_shallow_nn 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Dual Training of Energy-Based Models with Overparametrized Shallow Neural Networks

2021-07-11 · Carles Domingo-Enrich, Alberto Bietti, Marylou Gabrié, Joan Bruna 외

Energy-based models (EBMs) are generative models that are usually trained via maximum likelihood estimation. This approach becomes challenging in generic situations where the trained energy is non-convex, due to the need…

Deep orthogonal linear networks are shallow

2020-11-27 · Pierre Ablin

We consider the problem of training a deep orthogonal linear network, which consists of a product of orthogonal matrices, with no non-linearity in-between. We show that training the weights with Riemannian gradient desce…

Optimisation of Overparametrized Sum-Product Networks

2019-05-20 · Martin Trapp, Robert Peharz, Franz Pernkopf

It seems to be a pearl of conventional wisdom that parameter learning in deep sum-product networks is surprisingly fast compared to shallow mixture models. This paper examines the effects of overparameterization in sum-p…

Nearly Minimal Over-Parametrization of Shallow Neural Networks

2019-10-09 · Armin Eftekhari, ChaeHwan Song, Volkan Cevher

A recent line of work has shown that an overparametrized neural network can perfectly fit the training data, an otherwise often intractable nonconvex optimization problem. For (fully-connected) shallow networks, in the b…

Stationary Points of Shallow Neural Networks with Quadratic Activation Function

2019-12-03 · David Gamarnik, Eren C. Kızıldağ, Ilias Zadik

We consider the teacher-student setting of learning shallow neural networks with quadratic activations and planted weight matrix $W^*\in\mathbb{R}^{m\times d}$, where $m$ is the width of the hidden layer and $d\le m$ is …