paper-with-me

홈 › Papers

Adapting Double Q-Learning for Continuous Reinforcement Learning

2023-09-25 · Arsenii Kuznetsov

Majority of off-policy reinforcement learning algorithms use overestimation bias control techniques. Most of these techniques rooted in heuristics, primarily addressing the consequences of overestimation rather than its fundamental origins. In this work we present a novel approach to the bias correction, similar in spirit to Double Q-Learning. We propose using a policy in form of a mixture with two components. Each policy component is maximized and assessed by separate networks, which removes any basis for the overestimation bias. Our approach shows promising near-SOTA results on a small set of MuJoCo environments.

📄 PDF Abstract BibTeX arXiv:2309.14471

Code (0)

등록된 구현이 없습니다.

Tasks

MuJoCoQ-Learningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Investigating Reinforcement Learning Agents for Continuous State Space Environments

2017-08-08 · David Von Dollen

Given an environment with continuous state spaces and discrete actions, we investigate using a Double Deep Q-learning Reinforcement Agent to find optimal policies using the LunarLander-v2 OpenAI gym environment.

OpenAI GymQ-Learningreinforcement-learningReinforcement Learning+1

Efficient Continuous Control with Double Actors and Regularized Critics

2021-06-06 · Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu Li

How to obtain good value estimation is one of the key problems in Reinforcement Learning (RL). Current value estimation methods, such as DDPG and TD3, suffer from unnecessary over- or underestimation bias. In this paper,…

continuous-controlContinuous ControlReinforcement Learning (RL)

Using Reinforcement Learning to Validate Empirical Game-Theoretic Analysis: A Continuous Double Auction Study

2016-04-22 · Mason Wright

Empirical game-theoretic analysis (EGTA) has recently been applied successfully to analyze the behavior of large numbers of competing traders in a continuous double auction market. Multiagent simulation methods like EGTA…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning

2024-09-28 · Haohui Chen, Zhiyong Chen, Aoxiang Liu, Wentuo Fang

To obtain better value estimation in reinforcement learning, we propose a novel algorithm based on the double actor-critic framework with temporal difference error-driven regularization, abbreviated as TDDR. TDDR employs…

continuous-controlContinuous Control

Action Candidate Driven Clipped Double Q-learning for Discrete and Continuous Action Tasks

2022-03-22 · Haobo Jiang, Jin Xie, Jian Yang

Double Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped Double Q-learning, as an effective variant of Double Q-learning, employs the clipped double estimator to …

Q-Learning