paper-with-me

Papers

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

2026-08-13 · Zijie Cheng, Yang Peng, Zhihua Zhang arxiv

In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.

📄 PDF Abstract BibTeX arXiv:2608.12973

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Online Inference in Distributional Temporal-Difference Learning

2026-08-14 · Yang Peng, Liangyu Zhang arxiv

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Marko…

Distributional Reinforcement Learning with Monotonic Splines

2021-09-29 · ICLR 2022 4 · Yudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte 외

Distributional Reinforcement Learning (RL) differs from traditional RL by estimating the distribution over returns to capture the intrinsic uncertainty of MDPs. One key challenge in distributional RL lies in how to param…

Distributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation

2023-05-28 · Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos 외

We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorithm, quantile temporal-difference learning…

Distributional Reinforcement Learningreinforcement-learningReinforcement Learning

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

2026-08-27 · Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang arxiv

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global compariso…

Reinforcement Learning

Distributional Reinforcement Learning with Dual Expectile-Quantile Regression

2023-05-26 · Sami Jullien, Romain Deffayet, Jean-Michel Renders, Paul Groth 외

Distributional reinforcement learning (RL) has proven useful in multiple benchmarks as it enables approximating the full distribution of returns and makes a better use of environment samples. The commonly used quantile r…

Continuous ControlDistributional Reinforcement Learningquantile regressionregression+3