paper-with-me

홈 › Papers

Refined Policy Improvement Bounds for MDPs

2021-07-16 · J. G. Dai, Mark Gluzman

The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algorithm. The existing bound leads to a degenerate bound when the discount factor approaches one, making the applicability of TRPO and related algorithms questionable when the discount factor is close to one. We refine the results in \cite{Schulman2015, Achiam2017} and propose a novel bound that is "continuous" in the discount factor. In particular, our bound is applicable for MDPs with the long-run average rewards as well.

📄 PDF Abstract BibTeX arXiv:2107.08068

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

TRPO Trust Region Policy Optimization, or TRPO, is a policy gradient method in reinforcement learning that avoids parameter updates that change the policy too much with a KL…

Similar Papers 제목 키워드 기반

Performance Improvement Bounds for Lipschitz Configurable Markov Decision Processes

2024-02-21 · Alberto Maria Metelli

Configurable Markov Decision Processes (Conf-MDPs) have recently been introduced as an extension of the traditional Markov Decision Processes (MDPs) to model the real-world scenarios in which there is the possibility to …

Gap-Dependent Bounds for Federated $Q$-learning

2025-02-05 · Haochen Zhang, Zhong Zheng, Lingzhou Xue

We present the first gap-dependent analysis of regret and communication cost for on-policy federated $Q$-Learning in tabular episodic finite-horizon Markov decision processes (MDPs). Existing FRL methods focus on worst-c…

Q-Learning

Lower Bounds for Policy Iteration on Multi-action MDPs

2020-09-16 · Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia, Parthasarathi Khirwadkar 외

Policy Iteration (PI) is a classical family of algorithms to compute an optimal policy for any given Markov Decision Problem (MDP). The basic idea in PI is to begin with some initial policy and to repeatedly update the p…

Processing Network Controls via Deep Reinforcement Learning

2022-05-01 · Mark Gluzman

Novel advanced policy gradient (APG) algorithms, such as proximal policy optimization (PPO), trust region policy optimization, and their variations, have become the dominant reinforcement learning (RL) algorithms because…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes

2024-03-11 · Navdeep Kumar, Yashaswini Murthy, Itai Shufaro, Kfir Y. Levy 외

We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we focus on ergodic tabular MDPs with fini…