paper-with-me

홈 › Papers

On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost

2019-07-14 · Zhuoran Yang, Yongxin Chen, Mingyi Hong, Zhaoran Wang

Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence is known to be fragile. To understand the instability of actor-critic, we focus on its application to linear quadratic regulators, a simple yet fundamental setting of reinforcement learning. We establish a nonasymptotic convergence analysis of actor-critic in this setting. In particular, we prove that actor-critic finds a globally optimal pair of actor (policy) and critic (action-value function) at a linear rate of convergence. Our analysis may serve as a preliminary step towards a complete theoretical understanding of bilevel optimization with nonconvex subproblems, which is NP-hard in the worst case and is often solved using heuristics.

📄 PDF Abstract BibTeX arXiv:1907.06246

Code (0)

등록된 구현이 없습니다.

Tasks

Bilevel OptimizationReinforcement Learning

Similar Papers 제목 키워드 기반

Provably Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost

2019-12-01 · NeurIPS 2019 12 · Zhuoran Yang, Yongxin Chen, Mingyi Hong, Zhaoran Wang

Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization,…

Bilevel OptimizationReinforcement Learning

Single-Timescale Actor-Critic Provably Finds Globally Optimal Policy

2020-08-02 · ICLR 2021 1 · Zuyue Fu, Zhuoran Yang, Zhaoran Wang

We study the global convergence and global optimality of actor-critic, one of the most popular families of reinforcement learning algorithms. While most existing works on actor-critic employ bi-level or two-timescale upd…

Global Convergence of Two-timescale Actor-Critic for Solving Linear Quadratic Regulator

2022-08-18 · Xuyang Chen, Jingliang Duan, Yingbin Liang, Lin Zhao

The actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly …

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning

2025-05-24 · Zhiyao Zhang, Myeung Suk Oh, FNU Hairi, Ziyue Luo 외

Actor-critic methods for decentralized multi-agent reinforcement learning (MARL) facilitate collaborative optimal decision making without centralized coordination, thus enabling a wide range of applications in practice. …

Decision MakingMulti-agent Reinforcement Learning

Convergence of gradient descent for learning linear neural networks

2021-08-04 · Gabin Maxime Nguegnang, Holger Rauhut, Ulrich Terstiege

We study the convergence properties of gradient descent for training deep linear neural networks, i.e., deep matrix factorizations, by extending a previous analysis for the related gradient flow. We show that under suita…