Deep Reinforcement Learning and the Deadly Triad
We know from reinforcement learning theory that temporal difference learning can fail in certain cases. Sutton and Barto (2018) identify a deadly triad of function approximation, bootstrapping, and off-policy learning. When these three properties are combined, learning can diverge with the value estimates becoming unbounded. However, several algorithms successfully combine these three properties, which indicates that there is at least a partial gap in our understanding. In this work, we investigate the impact of the deadly triad in practice, in the context of a family of popular deep reinforcement learning models - deep Q-networks trained with experience replay - analysing how the components of this system play a role in the emergence of the deadly triad, and in the agent's performance
Code (0)
등록된 구현이 없습니다.
Tasks
Deep Reinforcement LearningLearning Theoryreinforcement-learningReinforcement LearningReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Breaking the Deadly Triad with a Target Network
The deadly triad refers to the instability of a reinforcement learning algorithm when it employs off-policy learning, function approximation, and bootstrapping simultaneously. In this paper, we investigate the target net…
Q-LearningRevisiting a Design Choice in Gradient Temporal Difference Learning
Off-policy learning enables a reinforcement learning (RL) agent to reason counterfactually about policies that are not executed and is one of the most important ideas in RL. It, however, can lead to instability when comb…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning
We explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a $\textit{fixed}$ number of future time steps. To learn…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Towards Characterizing Divergence in Deep Q-Learning
Deep Q-Learning (DQL), a family of temporal difference algorithms for control, employs three techniques collectively known as the `deadly triad' in reinforcement learning: bootstrapping, off-policy learning, and function…
continuous-controlContinuous ControlMuJoCoOpenAI Gym+2Average-Reward Off-Policy Policy Evaluation with Function Approximation
We consider off-policy policy evaluation with function approximation (FA) in average-reward MDPs, where the goal is to estimate both the reward rate and the differential value function. For this problem, bootstrapping is…