paper-with-me

홈 › Papers

Linear Convergence of Data-Enabled Policy Optimization for Linear Quadratic Tracking

2024-10-08 · Shubo Kang, Feiran Zhao, Keyou You

Data-enabled policy optimization (DeePO) is a newly proposed method to attack the open problem of direct adaptive LQR. In this work, we extend the DeePO framework to the linear quadratic tracking (LQT) with offline data. By introducing a covariance parameterization of the LQT policy, we derive a direct data-driven formulation of the LQT problem. Then, we use gradient descent method to iteratively update the parameterized policy to find an optimal LQT policy. Moreover, by revealing the connection between DeePO and model-based policy optimization, we prove the linear convergence of the DeePO iteration. Finally, a numerical experiment is given to validate the convergence results. We hope our work paves the way to direct adaptive LQT with online closed-loop data.

📄 PDF Abstract BibTeX arXiv:2410.05596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neural Temporal-Difference Learning Converges to Global Optima

2019-12-01 · NeurIPS 2019 12 · Qi Cai, Zhuoran Yang, Jason D. Lee, Zhaoran Wang

Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coup…

Deep Reinforcement LearningQ-LearningReinforcement LearningReinforcement Learning (RL)

Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima

2019-05-24 · NeurIPS 2019 12 · Qi Cai, Zhuoran Yang, Jason D. Lee, Zhaoran Wang

Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coup…

Deep Reinforcement LearningQ-LearningReinforcement Learning

Policy Optimization for Markovian Jump Linear Quadratic Control: Gradient-Based Methods and Global Convergence

2020-11-24 · Joao Paulo Jansch-Porto, Bin Hu, Geir Dullerud

Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the global convergence of gradient-based policy op…

Policy Gradient Methods

Convergence Guarantees of Policy Optimization Methods for Markovian Jump Linear Systems

2020-02-10 · Joao Paulo Jansch-Porto, Bin Hu, Geir Dullerud

Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the convergence of policy optimization for quadrat…

Reinforcement Learning

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

2019-05-31 · NeurIPS 2019 12 · Kaiqing Zhang, Zhuoran Yang, Tamer Başar

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-…

Reinforcement Learning