paper-with-me

홈 › Papers

Multi-step Off-policy Learning Without Importance Sampling Ratios

2017-02-09 · Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton

To estimate the value functions of policies from exploratory data, most model-free off-policy algorithms rely on importance sampling, where the use of importance sampling ratios often leads to estimates with severe variance. It is thus desirable to learn off-policy without using the ratios. However, such an algorithm does not exist for multi-step learning with function approximation. In this paper, we introduce the first such algorithm based on temporal-difference (TD) learning updates. We show that an explicit use of importance sampling ratios can be eliminated by varying the amount of bootstrapping in TD updates in an action-dependent manner. Our new algorithm achieves stability using a two-timescale gradient-based TD update. A prior algorithm based on lookup table representation called Tree Backup can also be retrieved using action-dependent bootstrapping, becoming a special case of our algorithm. In two challenging off-policy tasks, we demonstrate that our algorithm is stable, effectively avoids the large variance issue, and can perform substantially better than its state-of-the-art counterpart.

📄 PDF Abstract BibTeX arXiv:1702.03006

Code (1)

sinaghiassian/OffpolicyAlgorithms

Similar Papers 제목 키워드 기반

Asymptotic optimality of adaptive importance sampling

2018-06-04 · NeurIPS 2018 12 · Bernard Delyon, François Portier

Adaptive importance sampling (AIS) uses past samples to update the \textit{sampling policy} $q_t$ at each stage $t$. Each stage $t$ is formed with two steps : (i) to explore the space with $n_t$ points according to $q_t$…

Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning

2023-01-26 · Brett Daley, Martha White, Christopher Amato, Marlos C. Machado

Off-policy learning from multistep returns is crucial for sample-efficient reinforcement learning, but counteracting off-policy bias without exacerbating variance is challenging. Classically, off-policy bias is corrected…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Weighted importance sampling for off-policy learning with linear function approximation

2014-12-01 · NeurIPS 2014 12 · A. Rupam Mahmood, Hado P. Van Hasselt, Richard S. Sutton

Importance sampling is an essential component of off-policy model-free reinforcement learning algorithms. However, its most effective variant, \emph{weighted} importance sampling, does not carry over easily to function a…

Reinforcement Learning

An Approximate Policy Iteration Viewpoint of Actor-Critic Algorithms

2022-08-05 · Zaiwei Chen, Siva Theja Maguluri

In this work, we consider policy-based methods for solving the reinforcement learning problem, and establish the sample complexity guarantees. A policy-based algorithm typically consists of an actor and a critic. We cons…

Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators

2021-06-24 · NeurIPS 2021 12 · Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan Shanmugam

In temporal difference (TD) learning, off-policy sampling is known to be more practical than on-policy sampling, and by decoupling learning from data collection, it enables data reuse. It is known that policy evaluation …