paper-with-me

홈 › Papers

Concentration of Cumulative Reward in Markov Decision Processes

2024-11-27 · Borna Sayedana, Peter E. Caines, Aditya Mahajan

In this paper, we investigate the concentration properties of cumulative rewards in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward concentration in MDPs, covering both infinite-horizon settings (i.e., average and discounted reward frameworks) and finite-horizon setting. Our asymptotic results include the law of large numbers, the central limit theorem, and the law of iterated logarithms, while our non-asymptotic bounds include Azuma-Hoeffding-type inequalities and a non-asymptotic version of the law of iterated logarithms. Additionally, we explore two key implications of our results. First, we analyze the sample path behavior of the difference in rewards between any two stationary policies. Second, we show that two alternative definitions of regret for learning policies proposed in the literature are rate-equivalent. Our proof techniques rely on a novel martingale decomposition of cumulative rewards, properties of the solution to the policy evaluation fixed-point equation, and both asymptotic and non-asymptotic concentration results for martingale difference sequences.

📄 PDF Abstract BibTeX arXiv:2411.18551

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal Nudging: Solving Average-Reward Semi-Markov Decision Processes as a Minimal Sequence of Cumulative Tasks

2015-04-20 · Reinaldo Uribe Muriel, Fernando Lozando, Charles Anderson

This paper describes a novel method to solve average-reward semi-Markov decision processes, by reducing them to a minimal sequence of cumulative reward problems. The usual solution methods for this type of problems updat…

Reinforcement Learning

Safe Reinforcement Learning in Constrained Markov Decision Processes

2020-08-15 · ICML 2020 1 · Akifumi Wachi, Yanan Sui

Safe reinforcement learning has been a promising approach for optimizing the policy of an agent that operates in safety-critical applications. In this paper, we propose an algorithm, SNO-MDP, that explores and optimizes …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning

Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning

2024-05-22 · Maximilian Nägele, Jan Olle, Thomas Fösel, Remmy Zen 외

Markov decision processes (MDPs) are used to model a wide variety of applications ranging from game playing over robotics to finance. Their optimal policy typically maximizes the expected sum of rewards given at each ste…

Portfolio Optimizationreinforcement-learningReinforcement Learning

Optimizing Quantiles in Preference-based Markov Decision Processes

2016-12-01 · Hugo Gilbert, Paul Weng, Yan Xu

In the Markov decision process model, policies are usually evaluated by expected cumulative rewards. As this decision criterion is not always suitable, we propose in this paper an algorithm for computing a policy optimal…

Learning in Markov Decision Processes under Constraints

2020-02-27 · Rahul Singh, Abhishek Gupta, Ness B. Shroff

We consider reinforcement learning (RL) in Markov Decision Processes in which an agent repeatedly interacts with an environment that is modeled by a controlled Markov process. At each time step $t$, it earns a reward, an…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)