paper-with-me

홈 › Papers

Minimum information divergence of Q-functions for dynamic treatment resumes

2022-11-16 · Shinto Eguchi

This paper aims at presenting a new application of information geometry to reinforcement learning focusing on dynamic treatment resumes. In a standard framework of reinforcement learning, a Q-function is defined as the conditional expectation of a reward given a state and an action for a single-stage situation. We introduce an equivalence relation, called the policy equivalence, in the space of all the Q-functions. A class of information divergence is defined in the Q-function space for every stage. The main objective is to propose an estimator of the optimal policy function by a method of minimum information divergence based on a dataset of trajectories. In particular, we discuss the $\gamma$-power divergence that is shown to have an advantageous property such that the $\gamma$-power divergence between policy-equivalent Q-functions vanishes. This property essentially works to seek the optimal policy, which is discussed in a framework of a semiparametric model for the Q-function. The specific choices of power index $\gamma$ give interesting relationships of the value function, and the geometric and harmonic means of the Q-function. A numerical experiment demonstrates the performance of the minimum $\gamma$-power divergence method in the context of dynamic treatment regimes.

📄 PDF Abstract BibTeX arXiv:2211.08741

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Regularizing cross entropy loss via minimum entropy and K-L divergence

2025-01-23 · Abdulrahman Oladipupo Ibraheem

I introduce two novel loss functions for classification in deep learning. The two loss functions extend standard cross entropy loss by regularizing it with minimum entropy and Kullback-Leibler (K-L) divergence terms. The…

Image ClassificationPosition

Empirically Estimable Classification Bounds Based on a New Divergence Measure

2014-12-19 · Visar Berisha, Alan Wisler, Alfred O. Hero, Andreas Spanias

Information divergence functions play a critical role in statistics and information theory. In this paper we show that a non-parametric f-divergence measure can be used to provide improved bounds on the minimum binary cl…

Binary ClassificationClassificationfeature selectionGeneral Classification

Combination of Evidence Using the Principle of Minimum Information Gain

2013-03-27 · Michael S. K. M. Wong, P. Lingras

One of the most important aspects in any treatment of uncertain information is the rule of combination for updating the degrees of uncertainty. The theory of belief functions uses the Dempster rule to combine two belief …

Unifying Non-Maximum Likelihood Learning Objectives with Minimum KL Contraction

2011-12-01 · NeurIPS 2011 12 · Siwei Lyu

When used to learn high dimensional parametric probabilistic models, the clas- sical maximum likelihood (ML) learning often suffers from computational in- tractability, which motivates the active developments of non-ML l…

Bounds on the Excess Minimum Risk via Generalized Information Divergence Measures

2025-05-30 · Ananya Omanwar, Fady Alajaji, Tamás Linder

Given finite-dimensional random vectors $Y$, $X$, and $Z$ that form a Markov chain in that order (i.e., $Y \to X \to Z$), we derive upper bounds on the excess minimum risk using generalized information divergence measure…