Model-Based Policy Search for Automatic Tuning of Multivariate PID Controllers
PID control architectures are widely used in industrial applications. Despite their low number of open parameters, tuning multiple, coupled PID controllers can become tedious in practice. In this paper, we extend PILCO, a model-based policy search framework, to automatically tune multivariate PID controllers purely based on data observed on an otherwise unknown system. The system's state is extended appropriately to frame the PID policy as a static state feedback policy. This renders PID tuning possible as the solution of a finite horizon optimal control problem without further a priori knowledge. The framework is applied to the task of balancing an inverted pendulum on a seven degree-of-freedom robotic arm, thereby demonstrating its capabilities of fast and data-efficient policy learning, even on complex real world problems.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Sample-Efficient Reinforcement Learning of Koopman eNMPC
Reinforcement learning (RL) can be used to tune data-driven (economic) nonlinear model predictive controllers ((e)NMPCs) for optimal performance in a specific control task by optimizing the dynamic model or parameters in…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Learning Robust Controllers Via Probabilistic Model-Based Policy Search
Model-based Reinforcement Learning estimates the true environment through a world model in order to approximate the optimal policy. This family of algorithms usually benefits from better sample efficiency than their mode…
Model-based Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Automatic LQR Tuning Based on Gaussian Process Global Optimization
This paper proposes an automatic controller tuning framework based on linear optimal control combined with Bayesian optimization. With this framework, an initial set of controller gains is automatically improved accordin…
Bayesian Optimizationglobal-optimizationLearning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, s…
Curriculum-based Reinforcement Learning for Distribution System Critical Load Restoration
This paper focuses on the critical load restoration problem in distribution systems following major outages. To provide fast online response and optimal sequential decision-making support, a reinforcement learning (RL) b…
Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1