paper-with-me

홈 › Papers

Analysis of Lower Bounds for Simple Policy Iteration

2019-11-28 · Sarthak Consul, Bhishma Dedhia, Kumar Ashutosh, Parthasarathi Khirwadkar

Policy iteration is a family of algorithms that are used to find an optimal policy for a given Markov Decision Problem (MDP). Simple Policy iteration (SPI) is a type of policy iteration where the strategy is to change the policy at exactly one improvable state at every step. Melekopoglou and Condon [1990] showed an exponential lower bound on the number of iterations taken by SPI for a 2 action MDP. The results have not been generalized to $k-$action MDP since. In this paper, we revisit the algorithm and the analysis done by Melekopoglou and Condon. We generalize the previous result and prove a novel exponential lower bound on the number of iterations taken by policy iteration for $N-$state, $k-$action MDPs. We construct a family of MDPs and give an index-based switching rule that yields a strong lower bound of $\mathcal{O}\big((3+k)2^{N/2-3}\big)$.

📄 PDF Abstract BibTeX arXiv:1911.12842

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lower Bounds for Policy Iteration on Multi-action MDPs

2020-09-16 · Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia, Parthasarathi Khirwadkar 외

Policy Iteration (PI) is a classical family of algorithms to compute an optimal policy for any given Markov Decision Problem (MDP). The basic idea in PI is to begin with some initial policy and to repeatedly update the p…

Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model

2025-05-30 · Oliver Mortensen, Mohammad Sadegh Talebi

In this paper we analyze the sample complexities of learning the optimal state-action value function $Q^*$ and an optimal policy $\pi^*$ in a discounted Markov decision process (MDP) where the agent has recursive entropi…

Q-Learning

Off-Policy Interval Estimation with Lipschitz Value Iteration

2020-10-29 · NeurIPS 2020 12 · Ziyang Tang, Yihao Feng, Na Zhang, Jian Peng 외

Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such as medical diagnosis or financial decis…

Decision MakingMedical DiagnosisOff-policy evaluation

Easy Monotonic Policy Iteration

2016-02-29 · Joshua Achiam

A key problem in reinforcement learning for control with general function approximators (such as deep neural networks and other nonlinear functions) is that, for many algorithms employed in practice, updates to the polic…

Reinforcement Learning

Improved and Generalized Upper Bounds on the Complexity of Policy Iteration

2013-06-03 · NeurIPS 2013 12 · Bruno Scherrer

Given a Markov Decision Process (MDP) with $n$ states and a totalnumber $m$ of actions, we study the number of iterations needed byPolicy Iteration (PI) algorithms to converge to the optimal$\gamma$-discounted policy. We…