Finite-Time Analysis of Temporal Difference Learning: Discrete-Time Linear System Perspective
TD-learning is a fundamental algorithm in the field of reinforcement learning (RL), that is employed to evaluate a given policy by estimating the corresponding value function for a Markov decision process. While significant progress has been made in the theoretical analysis of TD-learning, recent research has uncovered guarantees concerning its statistical efficiency by developing finite-time error bounds. This paper aims to contribute to the existing body of knowledge by presenting a novel finite-time analysis of tabular temporal difference (TD) learning, which makes direct and effective use of discrete-time stochastic linear system models and leverages Schur matrix properties. The proposed analysis can cover both on-policy and off-policy settings in a unified manner. By adopting this approach, we hope to offer new and straightforward templates that not only shed further light on the analysis of TD-learning and related RL algorithms but also provide valuable insights for future research in this domain.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement Learning (RL)Similar Papers 제목 키워드 기반
Numerical analysis for a unified 2 factor model of structural and reduced form types for corporate bonds with fixed discrete coupon
Conditions of Stability for explicit finite difference scheme and some results of numerical analysis for a unified 2 factor model of structural and reduced form types for corporate bonds with fixed discrete coupon are pr…
FormApplications of Tao General Difference in Discrete Domain
Numerical difference computation is one of the cores and indispensable in the modern digital era. Tao general difference (TGD) is a novel theory and approach to difference computation for discrete sequences and arrays in…
Edge DetectionA Theory of General Difference in Continuous and Discrete Domain
Though a core element of the digital age, numerical difference algorithms struggle with noise susceptibility. This stems from a key disconnect between the infinitesimal quantities in continuous differentiation and the fi…
A Finite Time Analysis of Temporal Difference Learning With Linear Function Approximation
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in…
Q-LearningReinforcement LearningTime-adaptive high-order compact finite difference schemes for option pricing in a family of stochastic volatility models
We propose a time-adaptive, high-order compact finite difference scheme for option pricing in a family of stochastic volatility models. We employ a semi-discrete high-order compact finite difference method for the spatia…