paper-with-me

홈 › Papers

Deep Reinforcement Learning and the Deadly Triad

2018-12-06 · Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, Joseph Modayil

We know from reinforcement learning theory that temporal difference learning can fail in certain cases. Sutton and Barto (2018) identify a deadly triad of function approximation, bootstrapping, and off-policy learning. When these three properties are combined, learning can diverge with the value estimates becoming unbounded. However, several algorithms successfully combine these three properties, which indicates that there is at least a partial gap in our understanding. In this work, we investigate the impact of the deadly triad in practice, in the context of a family of popular deep reinforcement learning models - deep Q-networks trained with experience replay - analysing how the components of this system play a role in the emergence of the deadly triad, and in the agent's performance

📄 PDF Abstract BibTeX arXiv:1812.02648

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningLearning Theoryreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Experience Replay Experience Replay is a replay memory technique used in reinforcement learning where we store the agent’s experiences at each time-step, $e\_{t} = \left(s\_{t}, a\_{t}, r\_{t},…

Similar Papers 제목 키워드 기반

Breaking the Deadly Triad with a Target Network

2021-01-21 · Shangtong Zhang, Hengshuai Yao, Shimon Whiteson

The deadly triad refers to the instability of a reinforcement learning algorithm when it employs off-policy learning, function approximation, and bootstrapping simultaneously. In this paper, we investigate the target net…

Q-Learning

Revisiting a Design Choice in Gradient Temporal Difference Learning

2023-08-02 · Xiaochi Qian, Shangtong Zhang

Off-policy learning enables a reinforcement learning (RL) agent to reason counterfactually about policies that are not executed and is one of the most important ideas in RL. It, however, can lead to instability when comb…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning

2019-09-09 · Kristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton 외

We explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a $\textit{fixed}$ number of future time steps. To learn…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Towards Characterizing Divergence in Deep Q-Learning

2019-03-21 · Joshua Achiam, Ethan Knight, Pieter Abbeel

Deep Q-Learning (DQL), a family of temporal difference algorithms for control, employs three techniques collectively known as the `deadly triad' in reinforcement learning: bootstrapping, off-policy learning, and function…

continuous-controlContinuous ControlMuJoCoOpenAI Gym+2

Average-Reward Off-Policy Policy Evaluation with Function Approximation

2021-01-08 · Shangtong Zhang, Yi Wan, Richard S. Sutton, Shimon Whiteson

We consider off-policy policy evaluation with function approximation (FA) in average-reward MDPs, where the goal is to estimate both the reward rate and the differential value function. For this problem, bootstrapping is…