paper-with-me

Papers

Markov Decision Processes with Continuous Side Information

2017-11-15 · Aditya Modi, Nan Jiang, Satinder Singh, Ambuj Tewari

We consider a reinforcement learning (RL) setting in which the agent interacts with a sequence of episodic MDPs. At the start of each episode the agent has access to some side-information or context that determines the dynamics of the MDP for that episode. Our setting is motivated by applications in healthcare where baseline measurements of a patient at the start of a treatment episode form the context that may provide information about how the patient might respond to treatment decisions. We propose algorithms for learning in such Contextual Markov Decision Processes (CMDPs) under an assumption that the unobserved MDP parameters vary smoothly with the observed context. We also give lower and upper PAC bounds under the smoothness assumption. Because our lower bound has an exponential dependence on the dimension, we consider a tractable linear setting where the context is used to create linear combinations of a finite set of MDPs. For the linear setting, we give a PAC learning algorithm based on KWIK learning techniques.

📄 PDF Abstract BibTeX arXiv:1711.05726

Code (0)

등록된 구현이 없습니다.

Tasks

PAC learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Logarithmic regret bounds for continuous-time average-reward Markov decision processes

2022-05-23 · Xuefeng Gao, Xun Yu Zhou

We consider reinforcement learning for continuous-time Markov decision processes (MDPs) in the infinite-horizon, average-reward setting. In contrast to discrete-time MDPs, a continuous-time process moves to a state and s…

Point Processesreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Approximating Euclidean by Imprecise Markov Decision Processes

2020-06-26 · Manfred Jaeger, Giorgio Bacci, Giovanni Bacci, Kim Guldstrand Larsen 외

Euclidean Markov decision processes are a powerful tool for modeling control problems under uncertainty over continuous domains. Finite state imprecise, Markov decision processes can be used to approximate the behavior o…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

ChronosPerseus: Randomized Point-based Value Iteration with Importance Sampling for POSMDPs

2022-07-16 · Richard Kohar, François Rivest, Alain Gosselin

In reinforcement learning, agents have successfully used environments modeled with Markov decision processes (MDPs). However, in many problem domains, an agent may suffer from noisy observations or random times until its…

Decision Making

Policy Gradient for Continuous-Time Robust Markov Decision Processes

2026-06-03 · Tanya Veeravalli, David M. Bossens, Atsushi Nitanda arxiv

The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under worst-case transition dynamics. Traditional RMDPs consider discrete-t…

Reinforcement Learning

Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs

2020-06-29 · NeurIPS 2020 12 · Jianzhun Du, Joseph Futoma, Finale Doshi-Velez

We present two elegant solutions for modeling continuous-time dynamics, in a novel model-based reinforcement learning (RL) framework for semi-Markov decision processes (SMDPs), using neural ordinary differential equation…

Model-based Reinforcement Learningreinforcement-learningReinforcement Learning (RL)