paper-with-me

Papers

The Cesàro Value Iteration

2025-04-07 · Jonas Mair, Lukas Schwenkel, Matthias A. Müller, Frank Allgöwer

In this paper, we address the problem of undiscouted infinite-horizon optimal control for deterministic systems where the classic value iteration does not converge. For such systems, we propose to use the Ces\aro mean to define the infinite-horizon optimal control problem and the corresponding infinite-horizon value function. Moreover, for this value function, we introduce the Ces\aro value iteration and prove its convergence for the special case of systems with periodic optimal operating behavior.

📄 PDF Abstract BibTeX arXiv:2504.04889

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generalized Second Order Value Iteration in Markov Decision Processes

2019-05-10 · Chandramouli Kamanchi, Raghuram Bharadwaj Diddigi, Shalabh Bhatnagar

Value iteration is a fixed point iteration technique utilized to obtain the optimal value function and policy in a discounted reward Markov Decision Process (MDP). Here, a contraction operator is constructed and applied …

Reinforcement Learning

The Value Iteration Algorithm is Not Strongly Polynomial for Discounted Dynamic Programming

2013-12-19 · Eugene A. Feinberg, Jefferson Huang

This note provides a simple example demonstrating that, if exact computations are allowed, the number of iterations required for the value iteration algorithm to find an optimal policy for discounted dynamic programming …

Approximate Modified Policy Iteration

2012-05-14 · Bruno Scherrer, Victor Gabillon, Mohammad Ghavamzadeh, Matthieu Geist

Modified policy iteration (MPI) is a dynamic programming (DP) algorithm that contains the two celebrated policy and value iteration methods. Despite its generality, MPI has not been thoroughly studied, especially its app…

General Classification

On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes

2023-02-08 · Yashaswini Murthy, Mehrdad Moharrami, R. Srikant

Modified policy iteration (MPI) is a dynamic programming algorithm that combines elements of policy iteration and value iteration. The convergence of MPI has been well studied in the context of discounted and average-cos…

Computational Efficiency

A Contracting Dynamical System Perspective toward Interval Markov Decision Processes

2023-09-17 · Saber Jafarpour, Samuel Coogan

Interval Markov decision processes are a class of Markov models where the transition probabilities between the states belong to intervals. In this paper, we study the problem of efficient estimation of the optimal polici…