paper-with-me

Papers

Understanding Autoencoders with Information Theoretic Concepts

2018-03-30 · Shujian Yu, Jose C. Principe

Despite their great success in practical applications, there is still a lack of theoretical and systematic methods to analyze deep neural networks. In this paper, we illustrate an advanced information theoretic methodology to understand the dynamics of learning and the design of autoencoders, a special type of deep learning architectures that resembles a communication channel. By generalizing the information plane to any cost function, and inspecting the roles and dynamics of different layers using layer-wise information quantities, we emphasize the role that mutual information plays in quantifying learning from data. We further suggest and also experimentally validate, for mean square error training, three fundamental properties regarding the layer-wise flow of information and intrinsic dimensionality of the bottleneck layer, using respectively the data processing inequality and the identification of a bifurcation point in the information plane that is controlled by the given data. Our observations have a direct impact on the optimal design of autoencoders, the design of alternative feedforward training methods, and even in the problem of generalization.

📄 PDF Abstract BibTeX arXiv:1804.00057

Code (0)

등록된 구현이 없습니다.

Tasks

Information Plane

Similar Papers 제목 키워드 기반

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders

2026-06-05 · Chenhao Zhang, Chris Lin, Su-In Lee arxiv

We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpretability of neural networks by learning sp…

A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima

2025-12-05 · Yiming Tang, Harshvardhan Saini, Zhaoqian Yao, Zheng Lin 외 arxiv

As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important for both scientific progress and trustwort…

Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

2025-02-21 · Xuansheng Wu, Jiayi Yuan, Wenlin Yao, Xiaoming Zhai 외

Large language models (LLMs) excel at handling human queries, but they can occasionally generate flawed or unexpected responses. Understanding their internal states is crucial for understanding their successes, diagnosin…

Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models

2024-11-01 · Aashiq Muhamed, Mona Diab, Virginia Smith

Understanding and mitigating the potential risks associated with foundation models (FMs) hinges on developing effective interpretability methods. Sparse Autoencoders (SAEs) have emerged as a promising tool for disentangl…

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

2026-06-19 · Sergio Lanza, Jae Hee Lee, Stefan Wermter arxiv

Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answering (VQA), but our understanding of the…

Visual Question AnsweringImage Captioning