paper-with-me

Papers

Model approximation in MDPs with unbounded per-step cost

2024-02-13 · Berk Bozkurt, Aditya Mahajan, Ashutosh Nayyar, Yi Ouyang

We consider the problem of designing a control policy for an infinite-horizon discounted cost Markov decision process $\mathcal{M}$ when we only have access to an approximate model $\hat{\mathcal{M}}$. How well does an optimal policy $\hat{\pi}^{\star}$ of the approximate model perform when used in the original model $\mathcal{M}$? We answer this question by bounding a weighted norm of the difference between the value function of $\hat{\pi}^\star $ when used in $\mathcal{M}$ and the optimal value function of $\mathcal{M}$. We then extend our results and obtain potentially tighter upper bounds by considering affine transformations of the per-step cost. We further provide upper bounds that explicitly depend on the weighted distance between cost functions and weighted distance between transition kernels of the original and approximate models. We present examples to illustrate our results.

📄 PDF Abstract BibTeX arXiv:2402.08813

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces

2025-10-05 · Osman Bicer, Ali D. Kara, Serdar Yuksel arxiv

In this paper, for Markov decision processes (MDPs) with unbounded state spaces we present refined upper bounds presented in [Kara et. al. JMLR'23] on finite model approximation errors via optimizing the quantizers used …

Learning Infinite-Horizon Average-Reward Markov Decision Processes with Constraints

2022-01-31 · Liyu Chen, Rahul Jain, Haipeng Luo

We study regret minimization for infinite-horizon average-reward Markov Decision Processes (MDPs) under cost constraints. We start by designing a policy optimization algorithm with carefully designed action-value estimat…

Operator-Theoretic Foundations and Policy Gradient Methods for General MDPs with Unbounded Costs

2026-03-18 · Abhishek Gupta, Aditya Mahajan arxiv

Markov decision processes (MDPs) is viewed as an optimization of an objective function over certain linear operators over general function spaces. A new existence result is established for the existence of optimal polici…

Reinforcement Learning

Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation

2024-04-19 · Jianliang He, Han Zhong, Zhuoran Yang

We study infinite-horizon average-reward Markov decision processes (AMDPs) in the context of general function approximation. Specifically, we propose a novel algorithmic framework named Local-fitted Optimization with OPt…

Anytime-Constrained Reinforcement Learning

2023-11-09 · Jeremy McMahan, Xiaojin Zhu

We introduce and study constrained Markov Decision Processes (cMDPs) with anytime constraints. An anytime constraint requires the agent to never violate its budget at any point in time, almost surely. Although Markovian …

reinforcement-learningReinforcement Learning