Adapting Double Q-Learning for Continuous Reinforcement Learning
Majority of off-policy reinforcement learning algorithms use overestimation bias control techniques. Most of these techniques rooted in heuristics, primarily addressing the consequences of overestimation rather than its fundamental origins. In this work we present a novel approach to the bias correction, similar in spirit to Double Q-Learning. We propose using a policy in form of a mixture with two components. Each policy component is maximized and assessed by separate networks, which removes any basis for the overestimation bias. Our approach shows promising near-SOTA results on a small set of MuJoCo environments.
Code (0)
등록된 구현이 없습니다.
Tasks
MuJoCoQ-Learningreinforcement-learningReinforcement LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Investigating Reinforcement Learning Agents for Continuous State Space Environments
Given an environment with continuous state spaces and discrete actions, we investigate using a Double Deep Q-learning Reinforcement Agent to find optimal policies using the LunarLander-v2 OpenAI gym environment.
OpenAI GymQ-Learningreinforcement-learningReinforcement Learning+1Efficient Continuous Control with Double Actors and Regularized Critics
How to obtain good value estimation is one of the key problems in Reinforcement Learning (RL). Current value estimation methods, such as DDPG and TD3, suffer from unnecessary over- or underestimation bias. In this paper,…
continuous-controlContinuous ControlReinforcement Learning (RL)Using Reinforcement Learning to Validate Empirical Game-Theoretic Analysis: A Continuous Double Auction Study
Empirical game-theoretic analysis (EGTA) has recently been applied successfully to analyze the behavior of large numbers of competing traders in a continuous double auction market. Multiagent simulation methods like EGTA…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning
To obtain better value estimation in reinforcement learning, we propose a novel algorithm based on the double actor-critic framework with temporal difference error-driven regularization, abbreviated as TDDR. TDDR employs…
continuous-controlContinuous ControlAction Candidate Driven Clipped Double Q-learning for Discrete and Continuous Action Tasks
Double Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped Double Q-learning, as an effective variant of Double Q-learning, employs the clipped double estimator to …
Q-Learning