Learning convex bounds for linear quadratic control policy synthesis
Learning to make decisions from observed data in dynamic environments remains a problem of fundamental importance in a number of fields, from artificial intelligence and robotics, to medicine and finance. This paper concerns the problem of learning control policies for unknown linear dynamical systems so as to maximize a quadratic reward function. We present a method to optimize the expected value of the reward over the posterior distribution of the unknown system parameters, given data. The algorithm involves sequential convex programing, and enjoys reliable local convergence and robust stability guarantees. Numerical simulations and stabilization of a real-world inverted pendulum are used to demonstrate the approach, with strong performance and robustness properties observed in both.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Policy Optimization for Markovian Jump Linear Quadratic Control: Gradient-Based Methods and Global Convergence
Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the global convergence of gradient-based policy op…
Policy Gradient MethodsRegret Analysis of Certainty Equivalence Policies in Continuous-Time Linear-Quadratic Systems
This work theoretically studies a ubiquitous reinforcement learning policy for controlling the canonical model of continuous-time stochastic linear-quadratic systems. We show that randomized certainty equivalent policy a…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
Nonlinear control systems with partial information to the decision maker are prevalent in a variety of applications. As a step toward studying such nonlinear systems, this work explores reinforcement learning methods for…
Online learning with dynamics: A minimax perspective
We study the problem of online learning with dynamics, where a learner interacts with a stateful environment over multiple rounds. In each round of the interaction, the learner selects a policy to deploy and incurs a cos…
counterfactualData-Driven LQR using Reinforcement Learning and Quadratic Neural Networks
This paper introduces a novel data-driven approach to design a linear quadratic regulator (LQR) using a reinforcement learning (RL) algorithm that does not require a system model. The key contribution is to perform polic…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)