Deep Value Model Predictive Control
In this paper, we introduce an actor-critic algorithm called Deep Value Model Predictive Control (DMPC), which combines model-based trajectory optimization with value function estimation. The DMPC actor is a Model Predictive Control (MPC) optimizer with an objective function defined in terms of a value function estimated by the critic. We show that our MPC actor is an importance sampler, which minimizes an upper bound of the cross-entropy to the state distribution of the optimal sampling policy. In our experiments with a Ballbot system, we show that our algorithm can work with sparse and binary reward signals to efficiently solve obstacle avoidance and target reaching tasks. Compared to previous work, we show that including the value function in the running cost of the trajectory optimizer speeds up the convergence. We also discuss the necessary strategies to robustify the algorithm in practice.
Code (0)
등록된 구현이 없습니다.
Tasks
modelModel Predictive ControlSimilar Papers 제목 키워드 기반
Predictive Control with Learning-Based Terminal Costs Using Approximate Value Iteration
Stability under model predictive control (MPC) schemes is frequently ensured by terminal ingredients. Employing a (control) Lyapunov function as the terminal cost constitutes a common choice. Learning-based methods may b…
Model Predictive ControlSoft MPCritic: Amortized Model Predictive Value Iteration
Reinforcement learning (RL) and model predictive control (MPC) offer complementary strengths, yet combining them at scale remains computationally challenging. We propose soft MPCritic, an RL-MPC framework that learns in …
Reinforcement LearningBootstrapped Model Predictive Control
Model Predictive Control (MPC) has been demonstrated to be effective in continuous control tasks. When a world model and a value function are available, planning a sequence of actions ahead of time leads to a better poli…
continuous-controlContinuous ControlImitation Learningmodel+1Lessons from AlphaZero for Optimal, Model Predictive, and Adaptive Control
In this paper we aim to provide analysis and insights (often based on visualization), which explain the beneficial effects of on-line decision making on top of off-line training. In particular, through a unifying abstrac…
Bayesian OptimizationDecision MakingModel Predictive ControlEfficient Recursive Data-enabled Predictive Control (Extended Version)
In the field of model predictive control, Data-enabled Predictive Control (DeePC) offers direct predictive control, bypassing traditional modeling. However, challenges emerge with increased computational demand due to re…
FormModel Predictive Control