paper-with-me

Papers

Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization

2023-09-03 · Uri Gadot, Esther Derman, Navdeep Kumar, Maxence Mohamed Elfatihi, Kfir Levy, Shie Mannor

In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance sensitivity to misspecified environments. Yet, to preserve computational tractability, the uncertainty set is traditionally independently structured for each state. This so-called rectangularity condition is solely motivated by computational concerns. As a result, it lacks a practical incentive and may lead to overly conservative behavior. In this work, we study coupled reward RMDPs where the transition kernel is fixed, but the reward function lies within an $\alpha$-radius from a nominal one. We draw a direct connection between this type of non-rectangular reward-RMDPs and applying policy visitation frequency regularization. We introduce a policy-gradient method and prove its convergence. Numerical experiments illustrate the learned policy's robustness and its less conservative behavior when compared to rectangular uncertainty.

📄 PDF Abstract BibTeX arXiv:2309.01107

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient Policy Iteration for Robust Markov Decision Processes via Regularization

2022-05-28 · Navdeep Kumar, Kfir Levy, Kaixin Wang, Shie Mannor

Robust Markov decision processes (MDPs) provide a general framework to model decision problems where the system dynamics are changing or only partially known. Efficient methods for some \texttt{sa}-rectangular robust MDP…

Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization

2026-02-11 · Anirudh Satheesh, Ziyi Chen, Furong Huang, Heng Huang arxiv

We study robust Markov decision processes (RMDPs) with general policy parameterization under s-rectangular and non-rectangular uncertainty sets. Prior work is largely limited to tabular policies, and hence either lacks s…

On the convex formulations of robust Markov decision processes

2022-09-21 · Julien Grand-Clément, Marek Petrik

Robust Markov decision processes (MDPs) are used for applications of dynamic optimization in uncertain environments and have been studied extensively. Many of the main properties and algorithms of MDPs, such as value ite…

Partial Policy Iteration for L1-Robust Markov Decision Processes

2020-06-16 · Chin Pang Ho, Marek Petrik, Wolfram Wiesemann

Robust Markov decision processes (MDPs) allow to compute reliable solutions for dynamic decision problems whose evolution is modeled by rewards and partially-known transition probabilities. Unfortunately, accounting for …

Strongly Polynomial Time Complexity of Policy Iteration for $L_\infty$ Robust MDPs

2026-01-30 · Ali Asadi, Krishnendu Chatterjee, Ehsan Goharshady, Mehrdad Karrabi 외 arxiv

Markov decision processes (MDPs) are a fundamental model in sequential decision making. Robust MDPs (RMDPs) extend this framework by allowing uncertainty in transition probabilities and optimizing against the worst-case …

Decision Making