paper-with-me

Papers

Two Timescale Stochastic Approximation with Controlled Markov noise and Off-policy temporal difference learning

2015-03-31 · Prasenjit Karmakar, Shalabh Bhatnagar

We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise. In particular, both the faster and slower recursions have non-additive controlled Markov noise components in addition to martingale difference noise. We analyze the asymptotic behavior of our framework by relating it to limiting differential inclusions in both time-scales that are defined in terms of the ergodic occupation measures associated with the controlled Markov processes. Finally, we present a solution to the off-policy convergence problem for temporal difference learning with linear function approximation, using our results.

📄 PDF Abstract BibTeX arXiv:1503.09105

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Central Limit Theorem for Two-Timescale Stochastic Approximation with Markovian Noise: Theory and Applications

2024-01-17 · Jie Hu, Vishwaraj Doshi, Do Young Eun

Two-timescale stochastic approximation (TTSA) is among the most general frameworks for iterative stochastic algorithms. This includes well-known stochastic optimization methods such as SGD variants and those designed for…

Stochastic Optimization

Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning

2026-05-29 · Vagul Mahadevan, Claire Chen, Shuze Daniel Liu, Shangtong Zhang arxiv

This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in fast and slow timescales respectively. Notable examples of two-time…

Reinforcement Learning

Finite Time Analysis of Linear Two-timescale Stochastic Approximation with Markovian Noise

2020-02-04 · Maxim Kaledin, Eric Moulines, Alexey Naumov, Vladislav Tadic 외

Linear two-timescale stochastic approximation (SA) scheme is an important class of algorithms which has become popular in reinforcement learning (RL), particularly for the policy evaluation problem. Recently, a number of…

Reinforcement LearningReinforcement Learning (RL)

Gaussian Approximation for Two-Timescale Linear Stochastic Approximation

2025-08-11 · Bogdan Butyrin, Artemy Rubtsov, Alexey Naumov, Vladimir Ulyanov 외 arxiv

In this paper, we establish non-asymptotic bounds for accuracy of normal approximation for linear two-timescale stochastic approximation (TTSA) algorithms driven by martingale difference or Markov noise. Focusing on both…

Finite-sample Analysis of Greedy-GQ with Linear Function Approximation under Markovian Noise

2020-05-20 · Yue Wang, Shaofeng Zou

Greedy-GQ is an off-policy two timescale algorithm for optimal control in reinforcement learning. This paper develops the first finite-sample analysis for the Greedy-GQ algorithm with linear function approximation under …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)