paper-with-me

홈 › Papers

Heavy-Ball Q-Learning with Residual Weighting Correction

2026-06-25 · Donghwan Lee arxiv

This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes convergence of its deterministic mean dynamics. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-learning with linear function approximation, where analogous convergence and acceleration statements are derived for the corresponding corrected fixed point. The sampled stochastic versions are treated through conditional-mean recursions and, in the stated linear-function-approximation setting, finite-time bounds. The analysis is based on a switched linear system (SLS) representation of Q-learning algorithms and on the joint spectral radius (JSR) of the associated switching families. This SLS viewpoint is not commonly used in standard analyses of Q-learning, and it provides a complementary framework and new insight into how heavy-ball momentum can accelerate Q-learning.

📄 PDF Abstract BibTeX arXiv:2606.27112

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Activation-Weighted Seeded Residual Coding for Low-Bit LLM Weight Repair

2026-08-24 · Zehao Liu, Chuangchuang Fang, Yang Ren arxiv

Low-bit weight quantization saves storage but leaves errors that degrade LLM quality. We introduce activation-weighted seeded residual coding (AWSRC), a compact repair codec for an existing quantization backbone. Given a…

Residual Physics Learning and System Identification for Sim-to-real Transfer of Policies on Buoyancy Assisted Legged Robots

2023-03-16 · Nitish Sontakke, Hosik Chae, Sangjoon Lee, Tianle Huang 외

The light and soft characteristics of Buoyancy Assisted Lightweight Legged Unit (BALLU) robots have a great potential to provide intrinsically safe interactions in environments involving humans, unlike many heavy and rig…

Deep Reinforcement Learning

Understanding the Acceleration Phenomenon via High-Resolution Differential Equations

2018-10-21 · Bin Shi, Simon S. Du, Michael. I. Jordan, Weijie J. Su

Gradient-based optimization algorithms can be studied from the perspective of limiting ordinary differential equations (ODEs). Motivated by the fact that existing ODEs do not distinguish between two fundamentally differe…

Vocal Bursts Intensity Prediction

Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling

2025-09-16 · Zhefei Gong, Shangke Lyu, Pengxiang Ding, Wei Xiao 외 arxiv

Imitation learning (IL) enables efficient skill acquisition from demonstrations but often struggles with long-horizon tasks and high-precision control due to compounding errors. Residual policy learning offers a promisin…

Heavy-ball Algorithms Always Escape Saddle Points

2019-07-23 · Tao Sun, Dongsheng Li, Zhe Quan, Hao Jiang 외

Nonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this …