paper-with-me

홈 › Papers

Correcting discount-factor mismatch in on-policy policy gradient methods

2023-06-23 · Fengdi Che, Gautham Vasan, A. Rupam Mahmood

The policy gradient theorem gives a convenient form of the policy gradient in terms of three factors: an action value, a gradient of the action likelihood, and a state distribution involving discounting called the \emph{discounted stationary distribution}. But commonly used on-policy methods based on the policy gradient theorem ignores the discount factor in the state distribution, which is technically incorrect and may even cause degenerate learning behavior in some environments. An existing solution corrects this discrepancy by using $\gamma^t$ as a factor in the gradient estimate. However, this solution is not widely adopted and does not work well in tasks where the later states are similar to earlier states. We introduce a novel distribution correction to account for the discounted stationary distribution that can be plugged into many existing gradient estimators. Our correction circumvents the performance degradation associated with the $\gamma^t$ correction with a lower variance. Importantly, compared to the uncorrected estimators, our algorithm provides improved state emphasis to evade suboptimal policies in certain environments and consistently matches or exceeds the original performance on several OpenAI gym and DeepMind suite benchmarks.

📄 PDF Abstract BibTeX arXiv:2306.13284

Code (0)

등록된 구현이 없습니다.

Tasks

OpenAI GymPolicy Gradient Methods

Similar Papers 제목 키워드 기반

Coordinate Ascent for Off-Policy RL with Global Convergence Guarantees

2022-12-10 · Hsin-En Su, Yen-ju Chen, Ping-Chun Hsieh, Xi Liu

We revisit the domain of off-policy policy optimization in RL from the perspective of coordinate ascent. One commonly-used approach is to leverage the off-policy policy gradient to optimize a surrogate objective -- the t…

counterfactual

Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement Learning

2022-06-14 · Shentao Yang, Yihao Feng, Shujian Zhang, Mingyuan Zhou

Offline reinforcement learning (RL) extends the paradigm of classical RL algorithms to purely learning from static datasets, without interacting with the underlying environment during the learning process. A key challeng…

continuous-controlContinuous ControlOffline RLreinforcement-learning+1

A Tale of Sampling and Estimation in Discounted Reinforcement Learning

2023-04-11 · Alberto Maria Metelli, Mirco Mutti, Marcello Restelli

The most relevant problems in discounted reinforcement learning involve estimating the mean of a function under the stationary distribution of a Markov reward process, such as the expected return in policy evaluation, or…

reinforcement-learningReinforcement Learning

Analysis of On-policy Policy Gradient Methods under the Distribution Mismatch

2025-03-28 · Weizhen Wang, Jianping He, Xiaoming Duan

Policy gradient methods are one of the most successful methods for solving challenging reinforcement learning problems. However, despite their empirical successes, many SOTA policy gradient algorithms for discounted prob…

Policy Gradient Methods

Refined Policy Improvement Bounds for MDPs

2021-07-16 · J. G. Dai, Mark Gluzman

The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algorithm. The existing bound leads to a degen…