Plan To Predict: Learning an Uncertainty-Foreseeing Model for Model-Based Reinforcement Learning
In Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an accurate model can be difficult since the policy is continually updated and the induced distribution over visited states used for model learning shifts accordingly. Prior methods alleviate this issue by quantifying the uncertainty of model-generated samples. However, these methods only quantify the uncertainty passively after the samples were generated, rather than foreseeing the uncertainty before model trajectories fall into those highly uncertain regions. The resulting low-quality samples can induce unstable learning targets and hinder the optimization of the policy. Moreover, while being learned to minimize one-step prediction errors, the model is generally used to predict for multiple steps, leading to a mismatch between the objectives of model learning and model usage. To this end, we propose \emph{Plan To Predict} (P2P), an MBRL framework that treats the model rollout process as a sequential decision making problem by reversely considering the model as a decision maker and the current policy as the dynamics. In this way, the model can quickly adapt to the current policy and foresee the multi-step future uncertainty when generating trajectories. Theoretically, we show that the performance of P2P can be guaranteed by approximately optimizing a lower bound of the true environment return. Empirical results demonstrate that P2P achieves state-of-the-art performance on several challenging benchmark tasks.
Code (1)
Tasks
Decision MakingmodelModel-based Reinforcement LearningSequential Decision MakingSimilar Papers 제목 키워드 기반
Risk Sensitive Model-Based Reinforcement Learning using Uncertainty Guided Planning
Identifying uncertainty and taking mitigating actions is crucial for safe and trustworthy reinforcement learning agents, especially when deployed in high-risk environments. In this paper, risk sensitivity is promoted in …
Model-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Selective Dyna-style Planning Under Limited Model Capacity
In model-based reinforcement learning, planning with an imperfect model of the environment has the potential to harm learning progress. But even when a model is imperfect, it may still contain information that is useful …
modelModel-based Reinforcement LearningSafe Chance Constrained Reinforcement Learning for Batch Process Control
Reinforcement Learning (RL) controllers have generated excitement within the control community. The primary advantage of RL controllers relative to existing methods is their ability to optimize uncertain systems independ…
Gaussian ProcessesModel Predictive Controlreinforcement-learningReinforcement Learning+1Explainability of Predictive Process Monitoring Results: Can You See My Data Issues?
Predictive business process monitoring (PPM) has been around for several years as a use case of process mining. PPM enables foreseeing the future of a business process through predicting relevant information about how a …
Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Predictive Process MonitoringUMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning
Offline reinforcement learning (RL) provides a framework for learning decision-making from offline data and therefore constitutes a promising approach for real-world applications as automated driving. Self-driving vehicl…
Decision MakingOffline RLreinforcement-learningReinforcement Learning+1