paper-with-me

Papers

Tight Bayesian Ambiguity Sets for Robust MDPs

2018-11-15 · Reazul Hasan Russel, Marek Petrik

Robustness is important for sequential decision making in a stochastic dynamic environment with uncertain probabilistic parameters. We address the problem of using robust MDPs (RMDPs) to compute policies with provable worst-case guarantees in reinforcement learning. The quality and robustness of an RMDP solution is determined by its ambiguity set. Existing methods construct ambiguity sets that lead to impractically conservative solutions. In this paper, we propose RSVF, which achieves less conservative solutions with the same worst-case guarantees by 1) leveraging a Bayesian prior, 2) optimizing the size and location of the ambiguity set, and, most importantly, 3) relaxing the requirement that the set is a confidence interval. Our theoretical analysis shows the safety of RSVF, and the empirical results demonstrate its practical promise.

📄 PDF Abstract BibTeX arXiv:1811.06512

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingReinforcement LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs

2019-02-20 · NeurIPS 2019 12 · Marek Petrik, Reazul Hasan Russell

Robust MDPs (RMDPs) can be used to compute policies with provable worst-case guarantees in reinforcement learning. The quality and robustness of an RMDP solution are determined by the ambiguity set---the set of plausible…

Bayesian InferencePositionreinforcement-learningReinforcement Learning+1

Optimizing Percentile Criterion Using Robust MDPs

2019-10-23 · Bahram Behzadian, Reazul Hasan Russel, Marek Petrik, Chin Pang Ho

We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high confidence when deployed. This objective,…

Reinforcement LearningReinforcement Learning (RL)

Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets

2026-02-05 · Chin Pang Ho, Marek Petrik, Wolfram Wiesemann arxiv

Robust Markov decision processes (MDPs) have attracted significant interest due to their ability to protect MDPs from poor out-of-sample performance in the presence of ambiguity. In contrast to classical MDPs, which acco…

Reinforcement Learning in Factored MDPs: Oracle-Efficient Algorithms and Tighter Regret Bounds for the Non-Episodic Setting

2020-02-06 · NeurIPS 2020 12 · Ziping Xu, Ambuj Tewari

We study reinforcement learning in non-episodic factored Markov decision processes (FMDPs). We propose two near-optimal and oracle-efficient algorithms for FMDPs. Assuming oracle access to an FMDP planner, they enjoy a B…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Robust Phi-Divergence MDPs

2022-05-27 · Chin Pang Ho, Marek Petrik, Wolfram Wiesemann

In recent years, robust Markov decision processes (MDPs) have emerged as a prominent modeling framework for dynamic decision problems affected by uncertainty. In contrast to classical MDPs, which only account for stochas…