paper-with-me

홈 › Papers

On the Global Optimum Convergence of Momentum-based Policy Gradient

2021-10-19 · Yuhao Ding, Junzi Zhang, Javad Lavaei

Policy gradient (PG) methods are popular and efficient for large-scale reinforcement learning due to their relative stability and incremental nature. In recent years, the empirical success of PG methods has led to the development of a theoretical foundation for these methods. In this work, we generalize this line of research by studying the global convergence of stochastic PG methods with momentum terms, which have been demonstrated to be efficient recipes for improving PG methods. We study both the soft-max and the Fisher-non-degenerate policy parametrizations, and show that adding a momentum improves the global optimality sample complexity of vanilla PG methods by $\tilde{\mathcal{O}}(\epsilon^{-1.5})$ and $\tilde{\mathcal{O}}(\epsilon^{-1})$, respectively, where $\epsilon>0$ is the target tolerance. Our work is the first one that obtains global convergence results for the momentum-based PG methods. For the generic Fisher-non-degenerate policy parametrizations, our result is the first single-loop and finite-batch PG algorithm achieving $\tilde{O}(\epsilon^{-3})$ global optimality sample complexity. Finally, as a by-product, our methods also provide general framework for analyzing the global convergence rates of stochastic PG methods, which can be easily applied and extended to different PG estimators.

📄 PDF Abstract BibTeX arXiv:2110.10116

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decentralized Federated Learning with Gradient Tracking over Time-Varying Directed Networks

2024-09-25 · Duong Thuy Anh Nguyen, Su Wang, Duong Tung Nguyen, Angelia Nedich 외

We investigate the problem of agent-to-agent interaction in decentralized (federated) learning over time-varying directed graphs, and, in doing so, propose a consensus-based algorithm called DSGTm-TV. The proposed algori…

Federated Learningimage-classificationImage Classification

Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization

2022-11-21 · Jiafeng Wang, Zhaoyu Chen, Kaixun Jiang, Dingkang Yang 외

Deep Neural Networks (DNNs) are vulnerable to adversarial examples, which are crafted by adding human-imperceptible perturbations to the benign inputs. Simultaneously, adversarial examples exhibit transferability across …

Global Convergence of Natural Policy Gradient with Hessian-aided Momentum Variance Reduction

2024-01-02 · Jie Feng, Ke Wei, Jinchi Chen

Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessi…

MuJoCoPolicy Gradient Methods

Mean-field analysis for heavy ball methods: Dropout-stability, connectivity, and global convergence

2022-10-13 · Diyuan Wu, Vyacheslav Kungurtsev, Marco Mondelli

The stochastic heavy ball method (SHB), also known as stochastic gradient descent (SGD) with Polyak's momentum, is widely used in training neural networks. However, despite the remarkable success of such algorithm in pra…

Policy Gradient Method For Robust Reinforcement Learning

2022-05-15 · Yue Wang, Shaofeng Zou

This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch. Robust reinforcement learning is to learn a policy rob…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)