paper-with-me

홈 › Papers

Demystifying overcomplete nonlinear auto-encoders: fast SGD convergence towards sparse representation from random initialization

2018-01-01 · ICLR 2018 1 · Cheng Tang, Claire Monteleoni

Auto-encoders are commonly used for unsupervised representation learning and for pre-training deeper neural networks. When its activation function is linear and the encoding dimension (width of hidden layer) is smaller than the input dimension, it is well known that auto-encoder is optimized to learn the principal components of the data distribution (Oja1982). However, when the activation is nonlinear and when the width is larger than the input dimension (overcomplete), auto-encoder behaves differently from PCA, and in fact is known to perform well empirically for sparse coding problems. We provide a theoretical explanation for this empirically observed phenomenon, when rectified-linear unit (ReLu) is adopted as the activation function and the hidden-layer width is set to be large. In this case, we show that, with significant probability, initializing the weight matrix of an auto-encoder by sampling from a spherical Gaussian distribution followed by stochastic gradient descent (SGD) training converges towards the ground-truth representation for a class of sparse dictionary learning models. In addition, we can show that, conditioning on convergence, the expected convergence rate is O(1/t), where t is the number of updates. Our analysis quantifies how increasing hidden layer width helps the training performance when random initialization is used, and how the norm of network weights influence the speed of SGD convergence.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dictionary LearningRepresentation Learning

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
PCA Principle Components Analysis (PCA) is an unsupervised method primary used for dimensionality reduction within machine learning. PCA is calculated via a singular value…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Beyond Linear and Overcomplete Regimes: A Mean-Field Analysis of Bottleneck Autoencoders

2026-06-05 · Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee arxiv

Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error. Despite their empirical success, theoretical understanding remains limited and largely r…

ICA with Reconstruction Cost for Efficient Overcomplete Feature Learning

2011-12-01 · NeurIPS 2011 12 · Quoc V. Le, Alexandre Karpenko, Jiquan Ngiam, Andrew Y. Ng

Independent Components Analysis (ICA) and its variants have been successfully used for unsupervised feature learning. However, standard ICA requires an orthonoramlity constraint to be enforced, which makes it difficult to…

Image ClassificationObject Recognition

Deep Convolutional Autoencoders as Generic Feature Extractors in Seismological Applications

2021-10-22 · Qingkai Kong, Andrea Chiang, Ana C. Aguiar, M. Giselle Fernández-Godino 외

The idea of using a deep autoencoder to encode seismic waveform features and then use them in different seismological applications is appealing. In this paper, we designed tests to evaluate this idea of using autoencoder…

MIDA: Multiple Imputation using Denoising Autoencoders

2017-05-08 · Lovedeep Gondara, Ke Wang

Missing data is a significant problem impacting all domains. State-of-the-art framework for minimizing missing data bias is multiple imputation, for which the choice of an imputation model remains nontrivial. We propose …

DenoisingImputation

Learning overcomplete, low coherence dictionaries with linear inference

2016-06-10 · Jesse A. Livezey, Alejandro F. Bujan, Friedrich T. Sommer

Finding overcomplete latent representations of data has applications in data analysis, signal processing, machine learning, theoretical neuroscience and many other fields. In an overcomplete representation, the number of…

compressed sensing