paper-with-me

Papers

Gamma-Nets: Generalizing Value Estimation over Timescale

2019-11-18 · Craig Sherstan, Shibhansh Dohare, James Macglashan, Johannes Günther, Patrick M. Pilarski

We present $\Gamma$-nets, a method for generalizing value function estimation over timescale. By using the timescale as one of the estimator's inputs we can estimate value for arbitrary timescales. As a result, the prediction target for any timescale is available and we are free to train on multiple timescales at each timestep. Here we empirically evaluate $\Gamma$-nets in the policy evaluation setting. We first demonstrate the approach on a square wave and then on a robot arm using linear function approximation. Next, we consider the deep reinforcement learning setting using several Atari video games. Our results show that $\Gamma$-nets can be effective for predicting arbitrary timescales, with only a small cost in accuracy as compared to learning estimators for fixed timescales. $\Gamma$-nets provide a method for compactly making predictions at many timescales without requiring a priori knowledge of the task, making it a valuable contribution to ongoing work on model-based planning, representation learning, and lifelong learning algorithms.

📄 PDF Abstract BibTeX arXiv:1911.07794

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningLifelong learningReinforcement LearningRepresentation Learning

Similar Papers 제목 키워드 기반

Contraction of $E_γ$-Divergence and Its Applications to Privacy

2020-12-20 · Shahab Asoodeh, Mario Diaz, Flavio P. Calmon

We investigate the contraction coefficients derived from strong data processing inequalities for the $E_\gamma$-divergence. By generalizing the celebrated Dobrushin's coefficient from total variation distance to $E_\gamm…

Machine-Learned Phase Diagrams of Generalized Kitaev Honeycomb Magnets

2021-02-01 · Nihal Rao, Ke Liu, Marc Machaczek, Lode Pollet

We use a recently developed interpretable and unsupervised machine-learning method, the tensorial kernel support vector machine (TK-SVM), to investigate the low-temperature classical phase diagram of a generalized Heisen…

A Semiparametric Bayesian Extreme Value Model Using a Dirichlet Process Mixture of Gamma Densities

2013-04-02 · Jairo Fuquene

In this paper we propose a model with a Dirichlet process mixture of gamma densities in the bulk part below threshold and a generalized Pareto density in the tail for extreme value estimation. The proposed model is simpl…

Density Estimation

Factors of Influence of the Overestimation Bias of Q-Learning

2022-10-11 · Julius Wagenbach, Matthia Sabatelli

We study whether the learning rate $\alpha$, the discount factor $\gamma$ and the reward signal $r$ have an influence on the overestimation bias of the Q-Learning algorithm. Our preliminary results in environments which …

Q-Learning

Gamma-Models: Generative Temporal Difference Learning for Infinite-Horizon Prediction

2020-12-01 · NeurIPS 2020 12 · Michael Janner, Igor Mordatch, Sergey Levine

We introduce the gamma-model, a predictive model of environment dynamics with an infinite, probabilistic horizon. Replacing standard single-step models with gamma-models leads to generalizations of the procedures that fo…

Generative Adversarial Network