paper-with-me

Papers

Non-decreasing Quantile Function Network with Efficient Exploration for Distributional Reinforcement Learning

2021-05-14 · Fan Zhou, Zhoufan Zhu, Qi Kuang, Liwen Zhang

Although distributional reinforcement learning (DRL) has been widely examined in the past few years, there are two open questions people are still trying to address. One is how to ensure the validity of the learned quantile function, the other is how to efficiently utilize the distribution information. This paper attempts to provide some new perspectives to encourage the future in-depth studies in these two fields. We first propose a non-decreasing quantile function network (NDQFN) to guarantee the monotonicity of the obtained quantile estimates and then design a general exploration framework called distributional prediction error (DPE) for DRL which utilizes the entire distribution of the quantile function. In this paper, we not only discuss the theoretical necessity of our method but also show the performance gain it achieves in practice by comparing with some competitors on Atari 2600 Games especially in some hard-explored games.

📄 PDF Abstract BibTeX arXiv:2105.06696

Code (0)

등록된 구현이 없습니다.

Tasks

Atari GamesDistributional Reinforcement LearningEfficient Explorationreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Non-Crossing Quantile Regression for Distributional Reinforcement Learning

2020-12-01 · NeurIPS 2020 12 · Fan Zhou, Jianing Wang, Xingdong Feng

Distributional reinforcement learning (DRL) estimates the distribution over future returns instead of the mean to more efficiently capture the intrinsic uncertainty of MDPs. However, batch-based DRL algorithms cannot gua…

Atari GamesDistributional Reinforcement Learningquantile regressionregression+3

QUOTA: The Quantile Option Architecture for Reinforcement Learning

2018-11-05 · Shangtong Zhang, Borislav Mavrin, Linglong Kong, Bo Liu 외

In this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL). In QUOTA, decision making is based on quantiles of a value distri…

Decision MakingDistributional Reinforcement Learningreinforcement-learningReinforcement Learning+1

Toward Risk-based Optimistic Exploration for Cooperative Multi-Agent Reinforcement Learning

2023-03-03 · Jihwan Oh, Joonkee Kim, Minchan Jeong, Se-Young Yun

The multi-agent setting is intricate and unpredictable since the behaviors of multiple agents influence one another. To address this environmental uncertainty, distributional reinforcement learning algorithms that incorp…

Distributional Reinforcement LearningMulti-agent Reinforcement Learningquantile regressionreinforcement-learning+2

Distributional Reinforcement Learning with Monotonic Splines

2021-09-29 · ICLR 2022 4 · Yudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte 외

Distributional Reinforcement Learning (RL) differs from traditional RL by estimating the distribution over returns to capture the intrinsic uncertainty of MDPs. One key challenge in distributional RL lies in how to param…

Distributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Distributional Reinforcement Learning for Efficient Exploration

2019-05-13 · Borislav Mavrin, Shangtong Zhang, Hengshuai Yao, Linglong Kong 외

In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient exploration method for deep RL that has…

Atari GamesDistributional Reinforcement LearningEfficient Explorationreinforcement-learning+2