paper-with-me

Papers

Model-based RL with Optimistic Posterior Sampling: Structural Conditions and Sample Complexity

2022-06-15 · Alekh Agarwal, Tong Zhang

We propose a general framework to design posterior sampling methods for model-based RL. We show that the proposed algorithms can be analyzed by reducing regret to Hellinger distance in conditional probability estimation. We further show that optimistic posterior sampling can control this Hellinger distance, when we measure model error via data likelihood. This technique allows us to design and analyze unified posterior sampling algorithms with state-of-the-art sample complexity guarantees for many model-based RL settings. We illustrate our general result in many special cases, demonstrating the versatility of our framework.

📄 PDF Abstract BibTeX arXiv:2206.07659

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling

2022-03-15 · Alekh Agarwal, Tong Zhang

Provably sample-efficient Reinforcement Learning (RL) with rich observations and function approximation has witnessed tremendous recent progress, particularly when the underlying function approximators are linear. In thi…

Reinforcement Learning (RL)

Partially Observable RL with B-Stability: Unified Structural Condition and Sharp Sample-Efficient Algorithms

2022-09-29 · Fan Chen, Yu Bai, Song Mei

Partial Observability -- where agents can only observe partial information about the true underlying state of the system -- is ubiquitous in real-world applications of Reinforcement Learning (RL). Theoretically, learning…

Reinforcement Learning (RL)

Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees

2022-09-28 · Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 외

We consider reinforcement learning in an environment modeled by an episodic, finite, stage-dependent Markov decision process of horizon $H$ with $S$ states, and $A$ actions. The performance of an agent is measured by the…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Variational Bayesian Optimistic Sampling

2021-10-29 · NeurIPS 2021 12 · Brendan O'Donoghue, Tor Lattimore

We consider online sequential decision problems where an agent must balance exploration and exploitation. We derive a set of Bayesian `optimistic' policies which, in the stochastic multi-armed bandit case, includes the T…

Thompson Sampling

Eluder Dimension and the Sample Complexity of Optimistic Exploration

2013-12-01 · NeurIPS 2013 12 · Daniel Russo, Benjamin Van Roy

This paper considers the sample complexity of the multi-armed bandit with dependencies among the arms. Some of the most successful algorithms for this problem use the principle of optimism in the face of uncertainty to g…

Thompson Sampling