Nonextensive information theoretical machine
In this paper, we propose a new discriminative model named \emph{nonextensive information theoretical machine (NITM)} based on nonextensive generalization of Shannon information theory. In NITM, weight parameters are treated as random variables. Tsallis divergence is used to regularize the distribution of weight parameters and maximum unnormalized Tsallis entropy distribution is used to evaluate fitting effect. On the one hand, it is showed that some well-known margin-based loss functions such as $\ell_{0/1}$ loss, hinge loss, squared hinge loss and exponential loss can be unified by unnormalized Tsallis entropy. On the other hand, Gaussian prior regularization is generalized to Student-t prior regularization with similar computational complexity. The model can be solved efficiently by gradient-based convex optimization and its performance is illustrated on standard datasets.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Anomaly Detection in Global Financial Markets with Graph Neural Networks and Nonextensive Entropy
Anomaly detection is a challenging task, particularly in systems with many variables. Anomalies are outliers that statistically differ from the analyzed data and can arise from rare events, malfunctions, or system misuse…
Anomaly DetectionStatistical crossover and nonextensive behavior of the neuronal short-term depression
The theoretical basis of neuronal coding, associated with short term degradation in synaptic transmission, is a matter of debate in the literature. In fact, electrophysiological signals are commonly characterized as inve…
Fourier Analysis and q-Gaussian Functions: Analytical and Numerical Results
It is a consensus in signal processing that the Gaussian kernel and its partial derivatives enable the development of robust algorithms for feature detection. Fourier analysis and convolution theory have central role in …
On Power-law Kernels, corresponding Reproducing Kernel Hilbert Space and Applications
The role of kernels is central to machine learning. Motivated by the importance of power-law distributions in statistical modeling, in this paper, we propose the notion of power-law kernels to investigate power-laws in l…
General Classificationregressionq-Paths: Generalizing the Geometric Annealing Path using Power Means
Many common machine learning methods involve the geometric annealing path, a sequence of intermediate densities between two distributions of interest constructed using the geometric average. While alternatives such as th…
Bayesian Inference