paper-with-me

Papers

Asynchronous Stochastic Approximation and Average-Reward Reinforcement Learning

2024-09-05 · Huizhen Yu, Yi Wan, Richard S. Sutton

This paper studies asynchronous stochastic approximation (SA) algorithms and their theoretical application to reinforcement learning in semi-Markov decision processes (SMDPs) with an average-reward criterion. We first extend Borkar and Meyn's stability proof method to accommodate more general noise conditions, yielding broader convergence guarantees for asynchronous SA. To sharpen the convergence analysis, we further examine shadowing properties in the asynchronous setting, building on a dynamical-systems approach of Hirsch and Bena\"{i}m. Leveraging these SA results, we establish the convergence of an asynchronous SA analogue of Schweitzer's classical relative value iteration algorithm, RVI Q-learning, for finite-space, weakly communicating SMDPs. Moreover, to make full use of these SA results in this application, we introduce new monotonicity conditions for estimating the optimal reward rate in RVI Q-learning. These conditions substantially expand the previously considered algorithmic framework, and we address them with novel arguments in the stability and convergence analysis of RVI Q-learning.

📄 PDF Abstract BibTeX arXiv:2409.03915

Code (0)

등록된 구현이 없습니다.

Tasks

Q-Learningreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays

2023-12-22 · Huizhen Yu, Yi Wan, Richard S. Sutton

In this paper, we study asynchronous stochastic approximation algorithms without communication delays. Our main contribution is a stability proof for these algorithms that extends a method of Borkar and Meyn by accommoda…

reinforcement-learningReinforcement Learning

Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration

2025-12-05 · Huizhen Yu, Yi Wan, Richard S. Sutton arxiv

This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov decision processes (SMDPs). We establish t…

Reinforcement Learning

Finite-Time Bounds for Two-Time-Scale Stochastic Approximation with Arbitrary Norm Contractions and Markovian Noise

2025-03-24 · Siddharth Chandak, Shaan ul Haque, Nicholas Bambos

Two-time-scale Stochastic Approximation (SA) is an iterative algorithm with applications in reinforcement learning and optimization. Prior finite time analysis of such algorithms has focused on fixed point iterations wit…

Q-Learningreinforcement-learningReinforcement Learning

Concentration of Contractive Stochastic Approximation and Reinforcement Learning

2021-06-27 · Siddharth Chandak, Vivek S. Borkar, Parth Dodhia

Using a martingale concentration inequality, concentration bounds `from time $n_0$ on' are derived for stochastic approximation algorithms with contractive maps and both martingale difference and Markov noises. These are…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Average-Reward Reinforcement Learning with Entropy Regularization

2025-01-15 · Jacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. Kulkarni

The average-reward formulation of reinforcement learning (RL) has drawn increased interest in recent years due to its ability to solve temporally-extended problems without discounting. Independently, RL algorithms have b…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)