Modeling the Long Term Future in Model-Based Reinforcement Learning
In model-based reinforcement learning, the agent interleaves between model learning and planning. These two components are inextricably intertwined. If the model is not able to provide sensible long-term prediction, the executed planer would exploit model flaws, which can yield catastrophic failures. This paper focuses on building a model that reasons about the long-term future and demonstrates how to use this for efficient planning and exploration. To this end, we build a latent-variable autoregressive model by leveraging recent ideas in variational inference. We argue that forcing latent variables to carry future information through an auxiliary task substantially improves long-term predictions. Moreover, by planning in the latent space, the planner's solution is ensured to be within regions where the model is valid. An exploration strategy can be devised by searching for unlikely trajectories under the model. Our methods achieves higher reward faster compared to baselines on a variety of tasks and environments in both the imitation learning and model-based reinforcement learning settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Imitation LearningModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)validVariational InferenceSimilar Papers 제목 키워드 기반
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently. As opposed to traditional one-step dynamics models, DWM offers long-horizon …
D4RLQ-LearningPRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning
We propose PRISM, a novel framework designed to overcome the limitations of 2D-based Preference-Based Reinforcement Learning (PBRL) by unifying 3D point cloud modeling and future-aware preference refinement. At its core,…
Autonomous NavigationDecision MakingLanguage ModelingLanguage Modelling+2Imitation Learning for Human Pose Prediction
Modeling and prediction of human motion dynamics has long been a challenging problem in computer vision, and most existing methods rely on the end-to-end supervised training of various architectures of recurrent neural n…
Deep Reinforcement LearningHuman Pose ForecastingImitation LearningPose Prediction+4Chat More If You Like: Dynamic Cue Words Planning to Flow Longer Conversations
To build an open-domain multi-turn conversation system is one of the most interesting and challenging tasks in Artificial Intelligence. Many research efforts have been dedicated to building such dialogue systems, yet few…
DiversityReinforcement LearningTransformer-Enhanced Reinforcement Learning: Fundamentals and Applications in Communication Networks
Reinforcement Learning (RL) has long been a powerful solution to various problems in communication networks. However, traditional RL models still face with several limitations. Not only do they rely on large numbers of i…
Reinforcement LearningSemantic Communication