A Probabilistic Representation for Deep Learning: Delving into The Information Bottleneck Principle
The Information Bottleneck (IB) principle has recently attracted great attention to explaining Deep Neural Networks (DNNs), and the key is to accurately estimate the mutual information between a hidden layer and dataset. However, some unsettled limitations weaken the validity of the IB explanation for DNNs. To address these limitations and fully explain deep learning in an information theoretic fashion, we propose a probabilistic representation for deep learning that allows the framework to estimate the mutual information, more accurately than existing non-parametric models, and also quantify how the components of a hidden layer affect the mutual information. Leveraging the probabilistic representation, we take into account the back-propagation training and derive two novel Markov chains to characterize the information flow in DNNs. We show that different hidden layers achieve different IB trade-offs depending on the architecture and the position of the layers in DNNs, whereas a DNN satisfies the IB principle no matter the architecture of the DNN.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Rewiring Techniques to Mitigate Oversquashing and Oversmoothing in GNNs: A Survey
Graph Neural Networks (GNNs) are powerful tools for learning from graph-structured data, but their effectiveness is often constrained by two critical challenges: oversquashing, where the excessive compression of informat…
Delving Globally into Texture and Structure for Image Inpainting
Image inpainting has achieved remarkable progress and inspired abundant methods, where the critical bottleneck is identified as how to fulfill the high-frequency structure and low-frequency texture information on the mas…
DecoderImage InpaintingUnveiling the Potential of Probabilistic Embeddings in Self-Supervised Learning
In recent years, self-supervised learning has played a pivotal role in advancing machine learning by allowing models to acquire meaningful representations from unlabeled data. An intriguing research avenue involves devel…
Out-of-Distribution DetectionSelf-Supervised LearningA Probabilistic Representation of Deep Learning for Improving The Information Theoretic Interpretability
In this paper, we propose a probabilistic representation of MultiLayer Perceptrons (MLPs) to improve the information-theoretic interpretability. Above all, we demonstrate that the activations being i.i.d. is not valid fo…
Structural Entropy Guided Probabilistic Coding
Probabilistic embeddings have several advantages over deterministic embeddings as they map each data point to a distribution, which better describes the uncertainty and complexity of data. Many works focus on adjusting t…
Natural Language UnderstandingregressionRepresentation Learning