paper-with-me

홈 › Papers

Target Updates May Stabilize Linear Q-Learning: Periodic and Soft Dynamics

2026-05-31 · Donghwan Lee arxiv

Periodic target updates in Q-learning and soft target updates in actor-critic methods are empirically well established stabilization mechanisms, but their precise theoretical explanation is still incomplete. This paper gives a rigorous and exact analysis of these mechanisms for Q-learning with linear function approximation (linear Q-learning) using the exact switched linear system (SLS) dynamics induced by the Bellman maximum and the joint spectral radius (JSR) of the resulting switching matrix families. Although linear Q-learning can fail to converge in general, we prove that, under explicit spectral and step-size conditions, periodic hard target updates and soft target updates can guarantee convergence to the exact projected Q-Bellman solution. The main analysis is carried out for deterministic linear Q-learning, where the target-update mechanism is most transparent. Once the corresponding JSR certificate is established for the mean recursion, the stochastic reinforcement-learning setting can be treated by replacing deterministic modes with sampled stochastic modes and adding the corresponding stochastic-noise analysis.

📄 PDF Abstract BibTeX arXiv:2606.02645

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometrically Averaged Hard Target Updates for Linear Q-Learning

2026-06-09 · Donghwan Lee arxiv

Periodic hard target updates are among the most common stabilization devices in modern deep Q-learning. Recent studies suggest that target updates can improve stability in Q-learning with function approximation, includin…

Deep Q-Learning with Gradient Target Tracking

2025-03-20 · Donghwan Lee, Bum Geun Park, Taeho Lee

This paper introduces Q-learning with gradient target tracking, a novel reinforcement learning framework that provides a learned continuous target update mechanism as an alternative to the conventional hard update paradi…

Q-Learning

Learning Control as Enabling Layer for Embodied Intelligence Research explored with Soft Robotic Swimming in diverse Flow Speeds

2026-06-09 · Fabian Schwab, Federico Allione, Bingcheng Wang, Mohamed El Arayshi 외 arxiv

Soft robots are valuable robophysical platforms for studying body-caudal undulatory locomotion, but their compliant bodies are difficult to control precisely under changing hydrodynamic loading. Conventional proportional…

Communication-Efficient Neural Tangent Kernels for Heterogeneous Decentralized Federated Learning

2025-12-14 · Li Xia arxiv

Decentralized federated learning (DFL) enables collaborative model training without a central server, but converges slowly under statistical heterogeneity. Recent work has shown that neural tangent kernel (NTK) methods a…

Federated Learning

Orbital Stabilization and Time Synchronization of Unstable Periodic Motions in Underactuated Robots

2025-09-24 · Surov Maksim arxiv

This paper presents a control methodology for achieving orbital stabilization with simultaneous time synchronization of periodic trajectories in underactuated robotic systems. The proposed approach extends the classical …