Target Updates May Stabilize Linear Q-Learning: Periodic and Soft Dynamics
Periodic target updates in Q-learning and soft target updates in actor-critic methods are empirically well established stabilization mechanisms, but their precise theoretical explanation is still incomplete. This paper gives a rigorous and exact analysis of these mechanisms for Q-learning with linear function approximation (linear Q-learning) using the exact switched linear system (SLS) dynamics induced by the Bellman maximum and the joint spectral radius (JSR) of the resulting switching matrix families. Although linear Q-learning can fail to converge in general, we prove that, under explicit spectral and step-size conditions, periodic hard target updates and soft target updates can guarantee convergence to the exact projected Q-Bellman solution. The main analysis is carried out for deterministic linear Q-learning, where the target-update mechanism is most transparent. Once the corresponding JSR certificate is established for the mean recursion, the stochastic reinforcement-learning setting can be treated by replacing deterministic modes with sampled stochastic modes and adding the corresponding stochastic-noise analysis.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Geometrically Averaged Hard Target Updates for Linear Q-Learning
Periodic hard target updates are among the most common stabilization devices in modern deep Q-learning. Recent studies suggest that target updates can improve stability in Q-learning with function approximation, includin…
Deep Q-Learning with Gradient Target Tracking
This paper introduces Q-learning with gradient target tracking, a novel reinforcement learning framework that provides a learned continuous target update mechanism as an alternative to the conventional hard update paradi…
Q-LearningLearning Control as Enabling Layer for Embodied Intelligence Research explored with Soft Robotic Swimming in diverse Flow Speeds
Soft robots are valuable robophysical platforms for studying body-caudal undulatory locomotion, but their compliant bodies are difficult to control precisely under changing hydrodynamic loading. Conventional proportional…
Communication-Efficient Neural Tangent Kernels for Heterogeneous Decentralized Federated Learning
Decentralized federated learning (DFL) enables collaborative model training without a central server, but converges slowly under statistical heterogeneity. Recent work has shown that neural tangent kernel (NTK) methods a…
Federated LearningOrbital Stabilization and Time Synchronization of Unstable Periodic Motions in Underactuated Robots
This paper presents a control methodology for achieving orbital stabilization with simultaneous time synchronization of periodic trajectories in underactuated robotic systems. The proposed approach extends the classical …