paper-with-me

Papers

Revisiting a Design Choice in Gradient Temporal Difference Learning

2023-08-02 · Xiaochi Qian, Shangtong Zhang

Off-policy learning enables a reinforcement learning (RL) agent to reason counterfactually about policies that are not executed and is one of the most important ideas in RL. It, however, can lead to instability when combined with function approximation and bootstrapping, two arguably indispensable ingredients for large-scale reinforcement learning. This is the notorious deadly triad. The seminal work Sutton et al. (2008) pioneers Gradient Temporal Difference learning (GTD) as the first solution to the deadly triad, which has enjoyed massive success thereafter. During the derivation of GTD, some intermediate algorithm, called $A^\top$TD, was invented but soon deemed inferior. In this paper, we revisit this $A^\top$TD and prove that a variant of $A^\top$TD, called $A_t^\top$TD, is also an effective solution to the deadly triad. Furthermore, this $A_t^\top$TD only needs one set of parameters and one learning rate. By contrast, GTD has two sets of parameters and two learning rates, making it hard to tune in practice. We provide asymptotic analysis for $A^\top_t$TD and finite sample analysis for a variant of $A^\top_t$TD that additionally involves a projection operator. The convergence rate of this variant is on par with the canonical on-policy temporal difference learning.

📄 PDF Abstract BibTeX arXiv:2308.01170

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Revisiting Design Choices in Proximal Policy Optimization

2020-09-23 · Chloe Ching-Yun Hsu, Celestine Mendler-Dünner, Moritz Hardt

Proximal Policy Optimization (PPO) is a popular deep policy gradient algorithm. In standard implementations, PPO regularizes policy updates with clipped probability ratios, and parameterizes policies with either continuo…

MuJoCo

Gradient Temporal Difference with Momentum: Stability and Convergence

2021-11-22 · Rohan Deb, Shalabh Bhatnagar

Gradient temporal difference (Gradient TD) algorithms are a popular class of stochastic approximation (SA) algorithms used for policy evaluation in reinforcement learning. Here, we consider Gradient TD algorithms with an…

Temporal-Difference estimation of dynamic discrete choice models

2019-12-19 · Karun Adusumilli, Dita Eckardt

We study the use of Temporal-Difference learning for estimating the structural parameters in dynamic discrete choice models. Our algorithms are based on the conditional choice probability approach but use functional appr…

Discrete Choice Models

On the Sample Complexity of Actor-Critic Method for Reinforcement Learning with Function Approximation

2019-10-18 · Harshat Kumar, Alec Koppel, Alejandro Ribeiro

Reinforcement learning, mathematically described by Markov Decision Problems, may be approached either through dynamic programming or policy search. Actor-critic algorithms combine the merits of both approaches by altern…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Temporal Difference Learning with Linear Function Approximation

2020-02-20 · Tao Sun, Han shen, Tianyi Chen, Dongsheng Li

This paper revisits the temporal difference (TD) learning algorithm for the policy evaluation tasks in reinforcement learning. Typically, the performance of TD(0) and TD($\lambda$) is very sensitive to the choice of step…

OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)