paper-with-me

Papers

Disentangling Options with Hellinger Distance Regularizer

2019-04-15 · Minsung Hyun, Junyoung Choi, Nojun Kwak

In reinforcement learning (RL), temporal abstraction still remains as an important and unsolved problem. The options framework provided clues to temporal abstraction in the RL, and the option-critic architecture elegantly solved the two problems of finding options and learning RL agents in an end-to-end manner. However, it is necessary to examine whether the options learned through this method play a mutually exclusive role. In this paper, we propose a Hellinger distance regularizer, a method for disentangling options. In addition, we will shed light on various indicators from the statistical point of view to compare with the options learned through the existing option-critic architecture.

📄 PDF Abstract BibTeX arXiv:1904.06887

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Private Minimum Hellinger Distance Estimation via Hellinger Distance Differential Privacy

2025-01-24 · Fengnan Deng, Anand N. Vidyashankar

Objective functions based on Hellinger distance yield robust and efficient estimators of model parameters. Motivated by privacy and regulatory requirements encountered in contemporary applications, we derive in this pape…

Robust hypothesis testing and distribution estimation in Hellinger distance

2020-11-03 · Ananda Theertha Suresh

We propose a simple robust hypothesis test that has the same sample complexity as that of the optimal Neyman-Pearson test up to constants, but robust to distribution perturbations under Hellinger distance. We discuss the…

Two-sample testing

Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures

2026-02-03 · Joonhyuk Jung, Chao Gao arxiv

We study the relation between the total variation (TV) and Hellinger distances between two Gaussian location mixtures. Our first result establishes a general upper bound: for any two mixing distributions supported on a c…

Hellinger Distance Constrained Regression

2021-01-01 · Egor Rotinov

This paper introduces the off-policy reinforcement learning method that uses the Hellinger distance between sampling policy and current policy as a constraint. Hellinger distance squared multiplied by two is greater than…

MuJoCoregressionReinforcement Learning (RL)

HELLINGER-UCB: A novel algorithm for stochastic multi-armed bandit problem and cold start problem in recommender system

2024-04-16 · Ruibo Yang, Jiazhou Wang, Andrew Mullhaupt

In this paper, we study the stochastic multi-armed bandit problem, where the reward is driven by an unknown random variable. We propose a new variant of the Upper Confidence Bound (UCB) algorithm called Hellinger-UCB, wh…

Recommendation Systems