paper-with-me

Papers

Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces

2024-03-09 · Yang Peng, Liangyu Zhang, Zhihua Zhang

Distributional reinforcement learning (DRL) has achieved empirical success in various domains. One core task in DRL is distributional policy evaluation, which involves estimating the return distribution $\eta^\pi$ for a given policy $\pi$. Distributional temporal difference learning has been accordingly proposed, which extends the classic temporal difference learning (TD) in RL. In this paper, we focus on the non-asymptotic statistical rates of distributional TD. To facilitate theoretical analysis, we propose non-parametric distributional TD (NTD). For a $\gamma$-discounted infinite-horizon tabular Markov decision process, we show that for NTD with a generative model, we need $\tilde{O}(\varepsilon^{-2}\mu_{\min}^{-1}(1-\gamma)^{-3})$ interactions with the environment to achieve an $\varepsilon$-optimal estimator with high probability, when the estimation error is measured by the $1$-Wasserstein. This sample complexity bound is minimax optimal up to logarithmic factors. In addition, we revisit categorical distributional TD (CTD), showing that the same non-asymptotic convergence bounds hold for CTD in the case of the $1$-Wasserstein distance. We also extend our analysis to the more general setting where the data generating process is Markovian. In the Markovian setting, we propose variance-reduced variants of NTD and CTD, and show that both can achieve a $\tilde{O}(\varepsilon^{-2} \mu_{\pi,\min}^{-1}(1-\gamma)^{-3}+t_{mix}\mu_{\pi,\min}^{-1}(1-\gamma)^{-1})$ sample complexity bounds in the case of the $1$-Wasserstein distance, which matches the state-of-the-art statistical results for classic policy evaluation. To achieve the sharp statistical rates, we establish a novel Freedman's inequality in Hilbert spaces. This new Freedman's inequality would be of independent interest for statistical analysis of various infinite-dimensional online learning problems.

📄 PDF Abstract BibTeX arXiv:2403.05811

Code (0)

등록된 구현이 없습니다.

Tasks

Distributional Reinforcement Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Accelerated Distributional Temporal Difference Learning with Linear Function Approximation

2025-11-16 · Kaicheng Jin, Yang Peng, Jiansheng Yang, Zhihua Zhang arxiv

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD learning is to estimate the return dist…

Reinforcement Learning

A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation

2025-02-20 · Yang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua Zhang

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD learning is to estimate the return distribu…

Distributional Reinforcement Learning

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

2026-08-13 · Zijie Cheng, Yang Peng, Zhihua Zhang arxiv

In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional …

Reinforcement Learning

Online Inference in Distributional Temporal-Difference Learning

2026-08-14 · Yang Peng, Liangyu Zhang arxiv

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Marko…

The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation

2023-05-28 · Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos 외

We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorithm, quantile temporal-difference learning…

Distributional Reinforcement Learningreinforcement-learningReinforcement Learning