paper-with-me

홈 › Papers

Hitting time for Markov decision process

2022-05-06 · Ruichao Jiang, Javad Tavakoli, Yiqinag Zhao

We define the hitting time for a Markov decision process (MDP). We do not use the hitting time of the Markov process induced by the MDP because the induced chain may not have a stationary distribution. Even it has a stationary distribution, the stationary distribution may not coincide with the (normalized) occupancy measure of the MDP. We observe a relationship between the MDP and the PageRank. Using this observation, we construct an MP whose stationary distribution coincides with the normalized occupancy measure of the MDP and we define the hitting time of the MDP as the hitting time of the associated MP.

📄 PDF Abstract BibTeX arXiv:2205.03476

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Similar Papers 제목 키워드 기반

Maximum Expected Hitting Cost of a Markov Decision Process and Informativeness of Rewards

2019-12-01 · NeurIPS 2019 12 · Falcon Dai, Matthew Walter

We propose a new complexity measure for Markov decision processes (MDPs), the maximum expected hitting cost (MEHC). This measure tightens the closely related notion of diameter [JOA10] by accounting for the reward struct…

Informativeness

Maximum Expected Hitting Cost of a Markov Decision Process and Informativeness of Rewards

2019-07-03 · Falcon Z. Dai, Matthew R. Walter

We propose a new complexity measure for Markov decision processes (MDPs), the maximum expected hitting cost (MEHC). This measure tightens the closely related notion of diameter [JOA10] by accounting for the reward struct…

Informativeness

ULTRA-MC: A Unified Approach to Learning Mixtures of Markov Chains via Hitting Times

2024-05-23 · Fabian Spaeh, Konstantinos Sotiropoulos, Charalampos E. Tsourakakis

This study introduces a novel approach for learning mixtures of Markov chains, a critical process applicable to various fields, including healthcare and the analysis of web users. Existing research has identified a clear…

On some dynamical features of the complete Moran model for neutral evolution in the presence of mutations

2024-02-20 · Giuseppe Gaeta

We present a version of the classical Moran model, in which mutations are taken into account; the possibility of mutations was introduced by Moran in his seminal paper, but it is more often overlooked in discussing the M…

Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies

2026-05-07 · Magnus Victor Boock, Abdullah Akgül, Mustafa Mert Çelikok, Melih Kandemir arxiv

We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process from hitting time observations. While p…

Representation LearningReinforcement Learning