paper-with-me

홈 › Papers

Why is Posterior Sampling Better than Optimism for Reinforcement Learning?

2016-07-01 · ICML 2017 8 · Ian Osband, Benjamin Van Roy

Computational results demonstrate that posterior sampling for reinforcement learning (PSRL) dramatically outperforms algorithms driven by optimism, such as UCRL2. We provide insight into the extent of this performance boost and the phenomenon that drives it. We leverage this insight to establish an $\tilde{O}(H\sqrt{SAT})$ Bayesian expected regret bound for PSRL in finite-horizon episodic Markov decision processes, where $H$ is the horizon, $S$ is the number of states, $A$ is the number of actions and $T$ is the time elapsed. This improves upon the best previous bound of $\tilde{O}(H S \sqrt{AT})$ for any reinforcement learning algorithm.

📄 PDF Abstract BibTeX arXiv:1607.00215

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

(More) Efficient Reinforcement Learning via Posterior Sampling

2013-06-04 · NeurIPS 2013 12 · Ian Osband, Daniel Russo, Benjamin Van Roy

Most provably-efficient learning algorithms introduce optimism about poorly-understood states and actions to encourage exploration. We study an alternative approach for efficient exploration, posterior sampling for reinf…

Efficient Explorationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Posterior Sampling-based Online Learning for Episodic POMDPs

2023-10-16 · Dengwang Tang, Dongze Ye, Rahul Jain, Ashutosh Nayyar 외

Learning in POMDPs is known to be significantly harder than in MDPs. In this paper, we consider the online learning problem for episodic POMDPs with unknown transition and observation models. We propose a Posterior Sampl…

A Self-Play Posterior Sampling Algorithm for Zero-Sum Markov Games

2022-10-04 · Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen 외

Existing studies on provably efficient algorithms for Markov games (MGs) almost exclusively build on the "optimism in the face of uncertainty" (OFU) principle. This work focuses on a different approach of posterior sampl…

Better Optimism By Bayes: Adaptive Planning with Rich Models

2014-02-09 · Arthur Guez, David Silver, Peter Dayan

The computational costs of inference and planning have confined Bayesian model-based reinforcement learning to one of two dismal fates: powerful Bayes-adaptive planning but only for simplistic models, or powerful, Bayesi…

Model-based Reinforcement LearningReinforcement LearningReinforcement Learning (RL)Thompson Sampling

Online Learning for Stochastic Shortest Path Model via Posterior Sampling

2021-06-09 · Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain, Haipeng Luo

We consider the problem of online reinforcement learning for the Stochastic Shortest Path (SSP) problem modeled as an unknown MDP with an absorbing state. We propose PSRL-SSP, a simple posterior sampling-based reinforcem…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)