paper-with-me

홈 › Papers

Understanding Deep Learning Generalization by Maximum Entropy

2017-11-21 · ICLR 2018 1 · Guanhua Zheng, Jitao Sang, Changsheng Xu

Deep learning achieves remarkable generalization capability with overwhelming number of model parameters. Theoretical understanding of deep learning generalization receives recent attention yet remains not fully explored. This paper attempts to provide an alternative understanding from the perspective of maximum entropy. We first derive two feature conditions that softmax regression strictly apply maximum entropy principle. DNN is then regarded as approximating the feature conditions with multilayer feature learning, and proved to be a recursive solution towards maximum entropy principle. The connection between DNN and maximum entropy well explains why typical designs such as shortcut and regularization improves model generalization, and provides instructions for future model development.

📄 PDF Abstract BibTeX arXiv:1711.07758

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learningregression

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Limit of the Maximum Random Permutation Set Entropy

2024-03-10 · Jiefeng Zhou, Zhen Li, Kang Hao Cheong, Yong Deng

The Random Permutation Set (RPS) is a new type of set proposed recently, which can be regarded as the generalization of evidence theory. To measure the uncertainty of RPS, the entropy of RPS and its corresponding maximum…

A Minimax Approach to Supervised Learning

2016-06-07 · NeurIPS 2016 12 · Farzan Farnia, David Tse

Given a task of predicting $Y$ from $X$, a loss function $L$, and a set of probability distributions $\Gamma$ on $(X,Y)$, what is the optimal decision rule minimizing the worst-case expected loss over $\Gamma$? In this p…

A Scheme for Molecular Computation of Maximum Likelihood Estimators for Log-Linear Models

2015-06-10 · Manoj Gopalkrishnan

We propose a novel molecular computing scheme for statistical inference. We focus on the much-studied statistical inference problem of computing maximum likelihood estimators for log-linear models. Our scheme takes log-l…

Maximum Entropy Auto-Encoding

2021-04-13 · Paul M Baggenstoss

In this paper, it is shown that an auto-encoder using optimal reconstruction significantly outperforms a conventional auto-encoder. Optimal reconstruction uses the conditional mean of the input given the features, under …

Image Reconstruction

Towards Understanding Distributional Reinforcement Learning: Regularization, Optimization, Acceleration and Sinkhorn Algorithm

2021-09-29 · Ke Sun, Yingnan Zhao, Yi Liu, Enze Shi 외

Distributional reinforcement learning~(RL) is a class of state-of-the-art algorithms that estimate the whole distribution of the total return rather than only its expectation. Despite the remarkable performance of distri…

Atari GamesDistributional Reinforcement Learningreinforcement-learningReinforcement Learning (RL)