paper-with-me

Papers

Approximation Methods for Partially Observed Markov Decision Processes (POMDPs)

2021-08-31 · Caleb M. Bowyer

POMDPs are useful models for systems where the true underlying state is not known completely to an outside observer; the outside observer incompletely knows the true state of the system, and observes a noisy version of the true system state. When the number of system states is large in a POMDP that often necessitates the use of approximation methods to obtain near optimal solutions for control. This survey is centered around the origins, theory, and approximations of finite-state POMDPs. In order to understand POMDPs, it is required to have an understanding of finite-state Markov Decision Processes (MDPs) in \autoref{mdp} and Hidden Markov Models (HMMs) in \autoref{hmm}. For this background theory, I provide only essential details on MDPs and HMMs and leave longer expositions to textbook treatments before diving into the main topics of POMDPs. Once the required background is covered, the POMDP is introduced in \autoref{pomdp}. The origins of the POMDP are explained in the classical papers section \autoref{classical}. Once the high computational requirements are understood from the exact methodological point of view, the main approximation methods are surveyed in \autoref{approximations}. Then, I end the survey with some new research directions in \autoref{conclusion}.

📄 PDF Abstract BibTeX arXiv:2108.13965

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

Reinforcement Learning with Function Approximation for Non-Markov Processes

2026-01-01 · Ali Devran Kara arxiv

We study reinforcement learning methods with linear function approximation under non-Markov state and cost processes. We first consider the policy evaluation method and show that the algorithm converges under suitable er…

Reinforcement Learning

Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes

2020-10-15 · Ali Devran Kara, Serdar Yuksel

In the theory of Partially Observed Markov Decision Processes (POMDPs), existence of optimal policies have in general been established via converting the original partially observed stochastic control problem to a fully …

Finite-Time Analysis of Natural Actor-Critic for POMDPs

2022-02-20 · Semih Cayci, Niao He, R. Srikant

We consider the reinforcement learning problem for partially observed Markov decision processes (POMDPs) with large or even countably infinite state spaces, where the controller has access to only noisy observations of t…

Active Trajectory Estimation for Partially Observed Markov Decision Processes via Conditional Entropy

2021-04-04 · Timothy L. Molloy, Girish N. Nair

In this paper, we consider the problem of controlling a partially observed Markov decision process (POMDP) in order to actively estimate its state trajectory over a fixed horizon with minimal uncertainty. We pose a novel…

A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes

2021-11-12 · Chengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan Jiang

We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs), where the evaluation policy depends only on observable variables and the behavior policy depends on unobservable latent …

Off-policy evaluation