paper-with-me

Papers

POPCORN: Partially Observed Prediction COnstrained ReiNforcement Learning

2020-01-13 · Joseph Futoma, Michael C. Hughes, Finale Doshi-Velez

Many medical decision-making tasks can be framed as partially observed Markov decision processes (POMDPs). However, prevailing two-stage approaches that first learn a POMDP and then solve it often fail because the model that best fits the data may not be well suited for planning. We introduce a new optimization objective that (a) produces both high-performing policies and high-quality generative models, even when some observations are irrelevant for planning, and (b) does so in batch off-policy settings that are typical in healthcare, when only retrospective data is available. We demonstrate our approach on synthetic examples and a challenging medical decision-making problem.

📄 PDF Abstract BibTeX arXiv:2001.04032

Code (1)

dtak/POPCORN-POMDP 공식 구현

Tasks

Decision MakingPredictionreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

POPCORN: Progressive Pseudo-labeling with Consistency Regularization and Neighboring

2021-09-13 · Reda Abdellah Kamraoui, Vinh-Thong Ta, Nicolas Papadakis, Fanny Compaire 외

Semi-supervised learning (SSL) uses unlabeled data to compensate for the scarcity of annotated images and the lack of method generalization to unseen domains, two usual problems in medical segmentation tasks. In this wor…

Image SegmentationLesion SegmentationSegmentationSemantic Segmentation

Regret Minimization for Partially Observable Deep Reinforcement Learning

2017-10-31 · ICML 2018 7 · Peter Jin, Kurt Keutzer, Sergey Levine

Deep reinforcement learning algorithms that estimate state and state-action value functions have been shown to be effective in a variety of challenging domains, including learning control strategies from raw image pixels…

counterfactualDeep Reinforcement LearningMinecraftreinforcement-learning+2

Neural Algorithms for Graph Navigation

2020-10-17 · NeurIPS Workshop LMCA 2020 12 · Aaron Zweig, Nesreen Ahmed, Theodore L. Willke, Guixiang Ma

The application of deep reinforcement learning (RL) to graph learning and meta-learning admits challenges from both topics. We consider the task of one-shot, partially observed graph navigation, acknowledging and addres…

Deep Reinforcement LearningGraph LearningMeta-Learningreinforcement-learning+1

Generative Temporal Models with Spatial Memory for Partially Observed Environments

2018-04-25 · ICML 2018 7 · Marco Fraccaro, Danilo Jimenez Rezende, Yori Zwols, Alexander Pritzel 외

In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent's representations during training or via use as part of an exp…

Model-based Reinforcement LearningReinforcement Learning

Learning over All Stabilizing Nonlinear Controllers for a Partially-Observed Linear System

2021-12-08 · Ruigang Wang, Nicholas H. Barbara, Max Revay, Ian R. Manchester

This paper proposes a nonlinear policy architecture for control of partially-observed linear dynamical systems providing built-in closed-loop stability guarantees. The policy is based on a nonlinear version of the Youla …

AllReinforcement Learning (RL)