Exploring Model-based Planning with Policy Networks
Model-based reinforcement learning (MBRL) with model-predictive control or online planning has shown great potential for locomotion control tasks in terms of both sample efficiency and asymptotic performance. Despite their initial successes, the existing planning methods search from candidate sequences randomly generated in the action space, which is inefficient in complex high-dimensional environments. In this paper, we propose a novel MBRL algorithm, model-based policy planning (POPLIN), that combines policy networks with online planning. More specifically, we formulate action planning at each time-step as an optimization problem using neural networks. We experiment with both optimization w.r.t. the action sequences initialized from the policy network, and also online optimization directly w.r.t. the parameters of the policy network. We show that POPLIN obtains state-of-the-art performance in the MuJoCo benchmarking environments, being about 3x more sample efficient than the state-of-the-art algorithms, such as PETS, TD3 and SAC. To explain the effectiveness of our algorithm, we show that the optimization surface in parameter space is smoother than in action space. Further more, we found the distilled policy network can be effectively applied without the expansive model predictive control during test time for some environments such as Cheetah. Code is released in https://github.com/WilsonWangTHU/POPLIN.
Code (1)
Tasks
BenchmarkingmodelModel-based Reinforcement LearningModel Predictive ControlMuJoCoReinforcement LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Broadly-Exploring, Local-Policy Trees for Long-Horizon Task Planning
Long-horizon planning in realistic environments requires the ability to reason over sequential tasks in high-dimensional state spaces with complex dynamics. Classical motion planning algorithms, such as rapidly-exploring…
Motion PlanningTask PlanningDeep imagination is a close to optimal policy for planning in large decision trees under limited resources
Many decisions involve choosing an uncertain course of actions in deep and wide decision trees, as when we plan to visit an exotic country for vacation. In these cases, exhaustive search for the best sequence of actions …
validDistributed Safe Learning and Planning for Multi-robot Systems
This paper considers the problem of online multi-robot motion planning with general nonlinear dynamics subject to unknown external disturbances. We propose dSLAP, a distributed safe learning and planning framework that a…
Active LearningCollision AvoidanceModel Predictive ControlMotion Planning+2Dual policy as self-model for planning
Planning is a data efficient decision-making strategy where an agent selects candidate actions by exploring possible future states. To simulate future states when there is a high-dimensional action space, the knowledge o…
Decision MakingmodelRL-RRT: Kinodynamic Motion Planning via Learning Reachability Estimators from RL Policies
This paper addresses two challenges facing sampling-based kinodynamic motion planning: a way to identify good candidate states for local transitions and the subsequent computationally intractable steering between these c…
Deep Reinforcement LearningMotion PlanningReinforcement Learning