POPCORN: Partially Observed Prediction COnstrained ReiNforcement Learning
Many medical decision-making tasks can be framed as partially observed Markov decision processes (POMDPs). However, prevailing two-stage approaches that first learn a POMDP and then solve it often fail because the model that best fits the data may not be well suited for planning. We introduce a new optimization objective that (a) produces both high-performing policies and high-quality generative models, even when some observations are irrelevant for planning, and (b) does so in batch off-policy settings that are typical in healthcare, when only retrospective data is available. We demonstrate our approach on synthetic examples and a challenging medical decision-making problem.
Code (1)
Tasks
Decision MakingPredictionreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
POPCORN: Progressive Pseudo-labeling with Consistency Regularization and Neighboring
Semi-supervised learning (SSL) uses unlabeled data to compensate for the scarcity of annotated images and the lack of method generalization to unseen domains, two usual problems in medical segmentation tasks. In this wor…
Image SegmentationLesion SegmentationSegmentationSemantic SegmentationRegret Minimization for Partially Observable Deep Reinforcement Learning
Deep reinforcement learning algorithms that estimate state and state-action value functions have been shown to be effective in a variety of challenging domains, including learning control strategies from raw image pixels…
counterfactualDeep Reinforcement LearningMinecraftreinforcement-learning+2Neural Algorithms for Graph Navigation
The application of deep reinforcement learning (RL) to graph learning and meta-learning admits challenges from both topics. We consider the task of one-shot, partially observed graph navigation, acknowledging and addres…
Deep Reinforcement LearningGraph LearningMeta-Learningreinforcement-learning+1Generative Temporal Models with Spatial Memory for Partially Observed Environments
In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent's representations during training or via use as part of an exp…
Model-based Reinforcement LearningReinforcement LearningLearning over All Stabilizing Nonlinear Controllers for a Partially-Observed Linear System
This paper proposes a nonlinear policy architecture for control of partially-observed linear dynamical systems providing built-in closed-loop stability guarantees. The policy is based on a nonlinear version of the Youla …
AllReinforcement Learning (RL)