Parametric Information Bottleneck to Optimize Stochastic Neural Networks
In this paper, we present a layer-wise learning of stochastic neural networks (SNNs) in an information-theoretic perspective. In each layer of an SNN, the compression and the relevance are defined to quantify the amount of information that the layer contains about the input space and the target space, respectively. We jointly optimize the compression and the relevance of all parameters in an SNN to better exploit the neural network's representation. Previously, the Information Bottleneck (IB) framework (\cite{Tishby99}) extracts relevant information for a target variable. Here, we propose Parametric Information Bottleneck (PIB) for a neural network by utilizing (only) its model parameters explicitly to approximate the compression and the relevance. We show that, as compared to the maximum likelihood estimate (MLE) principle, PIBs : (i) improve the generalization of neural networks in classification tasks, (ii) push the representation of neural networks closer to the optimal information-theoretical representation in a faster manner.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Variational AutoEncoder for Transformers with Nonparametric Variational Information Bottleneck
We propose a VAE for Transformers by developing a variational information bottleneck regulariser for Transformer embeddings. We formalise the embedding space of Transformer encoders as mixture probability distributions, …
DecoderVariational Information Bottleneck for Unsupervised Clustering: Deep Gaussian Mixture Embedding
In this paper, we develop an unsupervised generative clustering framework that combines the Variational Information Bottleneck and the Gaussian Mixture Model. Specifically, in our approach, we use the Variational Informa…
ClusteringVariational InferenceImproving Generalization of Deep Networks for Inverse Reconstruction of Image Sequences
Deep learning networks have shown state-of-the-art performance in many image reconstruction problems. However, it is not well understood what properties of representation and learning may improve the generalization abili…
DecoderImage ReconstructionLearning TheoryInformation Theoretic Meta Learning with Gaussian Processes
We formulate meta learning using information theoretic concepts; namely, mutual information and the information bottleneck. The idea is to learn a stochastic representation or encoding of the task description, given by a…
Gaussian ProcessesMeta-LearningLayer-wise Learning of Stochastic Neural Networks with Information Bottleneck
Information Bottleneck (IB) is a generalization of rate-distortion theory that naturally incorporates compression and relevance trade-offs for learning. Though the original IB has been extensively studied, there has not …
Adversarial Robustness