paper-with-me

홈 › Papers

Designing a Linearized Potential Function in Neural Network Optimization Using Csiszár Type of Tsallis Entropy

2024-11-06 · Keito Akiyama

In recent years, learning for neural networks can be viewed as optimization in the space of probability measures. To obtain the exponential convergence to the optimizer, the regularizing term based on Shannon entropy plays an important role. Even though an entropy function heavily affects convergence results, there is almost no result on its generalization, because of the following two technical difficulties: one is the lack of sufficient condition for generalized logarithmic Sobolev inequality, and the other is the distributional dependence of the potential function within the gradient flow equation. In this paper, we establish a framework that utilizes a linearized potential function via Csisz\'{a}r type of Tsallis entropy, which is one of the generalized entropies. We also show that our new framework enable us to derive an exponential convergence result.

📄 PDF Abstract BibTeX arXiv:2411.03611

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Promises and Pitfalls of the Linearized Laplace in Bayesian Optimization

2023-04-17 · Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Vincent Fortuin

The linearized-Laplace approximation (LLA) has been shown to be effective and efficient in constructing Bayesian neural networks. It is theoretically compelling since it can be seen as a Gaussian process posterior with t…

Bayesian OptimizationDecision MakingGaussian Processesimage-classification+2

Cramér-Rao Lower Bounds Arising from Generalized Csiszár Divergences

2020-01-14 · M. Ashok Kumar, Kumar Vijay Mishra

We study the geometry of probability distributions with respect to a generalized family of Csisz\'ar $f$-divergences. A member of this family is the relative $\alpha$-entropy which is also a R\'enyi analog of relative en…

Linearized Relative Positional Encoding

2023-07-18 · Zhen Qin, Weixuan Sun, Kaiyue Lu, Hui Deng 외

Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are not always directly applicable to a line…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+2

Controllable Neural Architectures for Multi-Task Control

2025-01-31 · Umberto Casti, Giacomo Baggio, Sandro Zampieri, Fabio Pasqualetti

This paper studies a multi-task control problem where multiple linear systems are to be regulated by a single non-linear controller. In particular, motivated by recent advances in multi-task learning and the design of br…

Multi-Task Learning

Differentiable Linearized ADMM

2019-05-15 · Xingyu Xie, Jianlong Wu, Zhisheng Zhong, Guangcan Liu 외

Recently, a number of learning-based optimization methods that combine data-driven architectures with the classical optimization algorithms have been proposed and explored, showing superior empirical performance in solvi…