paper-with-me

Papers

Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization

2024-05-03 · Mudit Gaur, Amrit Singh Bedi, Di Wang, Vaneet Aggarwal

The current state-of-the-art theoretical analysis of Actor-Critic (AC) algorithms significantly lags in addressing the practical aspects of AC implementations. This crucial gap needs bridging to bring the analysis in line with practical implementations of AC. To address this, we advocate for considering the MMCLG criteria: \textbf{M}ulti-layer neural network parametrization for actor/critic, \textbf{M}arkovian sampling, \textbf{C}ontinuous state-action spaces, the performance of the \textbf{L}ast iterate, and \textbf{G}lobal optimality. These aspects are practically significant and have been largely overlooked in existing theoretical analyses of AC algorithms. In this work, we address these gaps by providing the first comprehensive theoretical analysis of AC algorithms that encompasses all five crucial practical aspects (covers MMCLG criteria). We establish global convergence sample complexity bounds of $\tilde{\mathcal{O}}\left({\epsilon^{-3}}\right)$. We achieve this result through our novel use of the weak gradient domination property of MDP's and our unique analysis of the error in critic estimation.

📄 PDF Abstract BibTeX arXiv:2405.01843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Average-Iterate to Last-Iterate Convergence in Games: A Reduction and Its Applications

2025-06-04 · Yang Cai, Haipeng Luo, Chen-Yu Wei, Weiqiang Zheng

The convergence of online learning algorithms in games under self-play is a fundamental question in game theory and machine learning. Among various notions of convergence, last-iterate convergence is particularly desirab…

Last-iterate convergence rates for min-max optimization

2019-06-05 · ICLR 2020 1 · Jacob Abernethy, Kevin A. Lai, Andre Wibisono

While classic work in convex-concave min-max optimization relies on average-iterate convergence results, the emergence of nonconvex applications such as training Generative Adversarial Networks has led to renewed interes…

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs

2026-05-12 · Michael Lu, Max Qiushi Lin, Mo Chen, Sharan Vaswani arxiv

We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture policy, deploying such a policy is computa…

Continuous Control

Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual

2026-02-25 · Yining Li, Peizhong Ju, Ness Shroff arxiv

Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constraints can be formulated as a primal-dual o…

Reinforcement Learning

Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games

2023-11-01 · Yang Cai, Gabriele Farina, Julien Grand-Clément, Christian Kroer 외

We study last-iterate convergence properties of algorithms for solving two-player zero-sum games based on Regret Matching$^+$ (RM$^+$). Despite their widespread use for solving real games, virtually nothing is known abou…