Policy Gradient Methods for Discrete Time Linear Quadratic Regulator With Random Parameters
This paper studies an infinite horizon optimal control problem for discrete-time linear system and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. In this general setting, we apply the policy gradient method, a reinforcement learning technique, to search for the optimal control without requiring knowledge of statistical information of the parameters. We investigate the sub-Gaussianity of the state process and establish global linear convergence guarantee for this approach based on assumptions that are weaker and easier to verify compared to existing results. Numerical experiments are presented to illustrate our result.
Code (0)
등록된 구현이 없습니다.
Tasks
Policy Gradient Methodsreinforcement-learningSimilar Papers 제목 키워드 기반
Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems
We study the global linear convergence of policy gradient (PG) methods for finite-horizon continuous-time exploratory linear-quadratic control (LQC) problems. The setting includes stochastic LQC problems with indefinite …
Policy Gradient MethodsOptimization Landscape of Policy Gradient Methods for Discrete-time Static Output Feedback
In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-fee…
Policy Gradient MethodsGlobal Convergence Using Policy Gradient Methods for Model-free Markovian Jump Linear Quadratic Control
Owing to the growth of interest in Reinforcement Learning in the last few years, gradient based policy control methods have been gaining popularity for Control problems as well. And rightly so, since gradient policy meth…
Policy Gradient MethodsPolicy Optimization for Markovian Jump Linear Quadratic Control: Gradient-Based Methods and Global Convergence
Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the global convergence of gradient-based policy op…
Policy Gradient MethodsGlobal Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time
Reinforcement learning is a powerful tool to learn the optimal policy of possibly multiple agents by interacting with the environment. As the number of agents grow to be very large, the system can be approximated by a me…