paper-with-me

홈 › Papers

Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization

2026-01-18 · Safwan Labbi, Daniil Tiapkin, Paul Mangold, Eric Moulines arxiv

Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditioned optimization landscapes and lead to exponentially slow convergence. Although this can be mitigated by preconditioning, this solution is often computationally expensive. Instead, we propose replacing the softmax with an alternative family of policy parameterizations based on the generalized f-softargmax. We further advocate coupling this parameterization with a regularizer induced by the same f-divergence, which improves the optimization landscape and ensures that the resulting regularized objective satisfies a Polyak-Lojasiewicz inequality. Leveraging this structure, we establish the first explicit non-asymptotic last-iterate convergence guarantees for stochastic policy gradient methods for finite MDPs without any form of preconditioning. We also derive sample-complexity bounds for the unregularized problem and show that f-PG, with Tsallis divergences achieves polynomial sample complexity in contrast to the exponential complexity incurred by the standard softmax parameterization.

📄 PDF Abstract BibTeX arXiv:2601.12604

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

2026-05-24 · Ziyue Chen, David Šiška, Lukasz Szpruch arxiv

We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider log-linear softmax policies with linear fu…

On the Global Convergence Rates of Softmax Policy Gradient Methods

2020-05-13 · ICML 2020 1 · Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, Dale Schuurmans

We make three contributions toward better understanding policy gradient methods in the tabular setting. First, we show that with the true gradient, policy gradient with a softmax parametrization converges at a $O(1/t)$ r…

Open-Ended Question AnsweringPolicy Gradient Methods

Elementary Analysis of Policy Gradient Methods

2024-04-04 · Jiacai Liu, Wenye Li, Ke Wei

Projected policy gradient under the simplex parameterization, policy gradient and natural policy gradient under the softmax parameterization, are fundamental algorithms in reinforcement learning. There have been a flurry…

Policy Gradient Methods

Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization

2020-07-13 · Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei 외

Natural policy gradient (NPG) methods are among the most widely used policy optimization algorithms in contemporary reinforcement learning. This class of methods is often applied in conjunction with entropy regularizatio…

Policy Gradient Methods

Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality

2023-03-22 · François Ged, Maria Han Veiga

A novel Policy Gradient (PG) algorithm, called $\textit{Matryoshka Policy Gradient}$ (MPG), is introduced and studied, in the context of fixed-horizon max-entropy reinforcement learning, where an agent aims at maximizing…