paper-with-me

홈 › Papers

Hybrid TD3: Overestimation Bias Analysis and Stable Policy Optimization for Hybrid Action Space

2026-03-01 · Thanh-Tuan Tran, Thanh Nguyen Canh, Nak Young Chong, Xiem HoangVan arxiv

Reinforcement learning in discrete-continuous hybrid action spaces presents fundamental challenges for robotic manipulation, where high-level task decisions and low-level joint-space execution must be jointly optimized. Existing approaches either discretize continuous components or relax discrete choices into continuous approximations, which suffer from scalability limitations and training instability in high-dimensional action spaces and under domain randomization. In this paper, we propose Hybrid TD3, an extension of Twin Delayed Deep Deterministic Policy Gradient (TD3) that natively handles parameterized hybrid action spaces in a principled manner. We conduct a rigorous theoretical analysis of overestimation bias in hybrid action settings, deriving formal bounds under twin-critic architectures and establishing a complete bias ordering across five algorithmic variants under synchronized Gaussian error assumptions. Building on this analysis, we introduce a weighted clipped Q-learning target that marginalizes over the discrete action distribution, achieving equivalent bias reduction to standard clipped minimization while improving policy smoothness. Experimental results demonstrate that Hybrid TD3 achieves superior training stability and competitive performance against state-of-the-art hybrid action baselines.

📄 PDF Abstract BibTeX arXiv:2603.01302

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Deep Reinforcement Learning With Adaptive Combined Critics

2021-01-01 · Huihui Zhang, Wu Huang

The overestimation problem has long been popular in deep value learning, because function approximation errors may lead to amplified value estimates and suboptimal policies. There have been several methods to deal with t…

continuous-controlContinuous ControlDeep Reinforcement Learningreinforcement-learning+2

Adapting Double Q-Learning for Continuous Reinforcement Learning

2023-09-25 · Arsenii Kuznetsov

Majority of off-policy reinforcement learning algorithms use overestimation bias control techniques. Most of these techniques rooted in heuristics, primarily addressing the consequences of overestimation rather than its …

MuJoCoQ-Learningreinforcement-learningReinforcement Learning

Elastic Step DQN: A novel multi-step algorithm to alleviate overestimation in Deep QNetworks

2022-10-07 · Adrian Ly, Richard Dazeley, Peter Vamplew, Francisco Cruz 외

Deep Q-Networks algorithm (DQN) was the first reinforcement learning algorithm using deep neural network to successfully surpass human level performance in a number of Atari learning environments. However, divergent and …

OpenAI Gym

Sign-Separated Asymmetric Finite-Time Error Analysis of Q-Learning

2026-05-15 · Donghwan Lee arxiv

Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, positive errors can be selected and propagated, causing learned values to exceed the…

Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics

2020-05-08 · ICML 2020 1 · Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry Vetrov

The overestimation bias is one of the major impediments to accurate off-policy learning. This paper investigates a novel way to alleviate the overestimation bias in a continuous control setting. Our method---Truncated Qu…

continuous-controlContinuous Control