paper-with-me

홈 › Papers

On the Statistical Benefits of Temporal Difference Learning

2023-01-30 · David Cheikhi, Daniel Russo

Given a dataset on actions and resulting long-term rewards, a direct estimation approach fits value functions that minimize prediction error on the training data. Temporal difference learning (TD) methods instead fit value functions by minimizing the degree of temporal inconsistency between estimates made at successive time-steps. Focusing on finite state Markov chains, we provide a crisp asymptotic theory of the statistical advantages of this approach. First, we show that an intuitive inverse trajectory pooling coefficient completely characterizes the percent reduction in mean-squared error of value estimates. Depending on problem structure, the reduction could be enormous or nonexistent. Next, we prove that there can be dramatic improvements in estimates of the difference in value-to-go for two states: TD's errors are bounded in terms of a novel measure - the problem's trajectory crossing time - which can be much smaller than the problem's time horizon.

📄 PDF Abstract BibTeX arXiv:2301.13289

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation

2023-05-28 · Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos 외

We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorithm, quantile temporal-difference learning…

Distributional Reinforcement Learningreinforcement-learningReinforcement Learning

Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status

2022-06-01 · LREC 2022 6 · Yves Bestgen

This paper argues for the widest possible use of bootstrap confidence intervals for comparing NLP system performances instead of the state-of-the-art status (SOTA) and statistical significance testing. Their main benefit…

Online Inference in Distributional Temporal-Difference Learning

2026-08-14 · Yang Peng, Liangyu Zhang arxiv

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Marko…

SpatioTemporal Difference Network for Video Depth Super-Resolution

2025-08-02 · Zhengxue Wang, Yuan Wu, Xiang Li, Zhiqiang Yan 외 arxiv

Depth super-resolution has achieved impressive performance, and the incorporation of multi-frame information further enhances reconstruction quality. Nevertheless, statistical analyses reveal that video depth super-resol…

Please, Don't Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status

2022-05-23 · Yves Bestgen

This paper argues for the widest possible use of bootstrap confidence intervals for comparing NLP system performances instead of the state-of-the-art status (SOTA) and statistical significance testing. Their main benefit…