paper-with-me

홈 › Papers

Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning

2020-03-24 · ICLR 2020 1 · Ali Mousavi, Lihong Li, Qiang Liu, Denny Zhou

Off-policy estimation for long-horizon problems is important in many real-life applications such as healthcare and robotics, where high-fidelity simulators may not be available and on-policy evaluation is expensive or impossible. Recently, \cite{liu18breaking} proposed an approach that avoids the \emph{curse of horizon} suffered by typical importance-sampling-based methods. While showing promising results, this approach is limited in practice as it requires data be drawn from the \emph{stationary distribution} of a \emph{known} behavior policy. In this work, we propose a novel approach that eliminates such limitations. In particular, we formulate the problem as solving for the fixed point of a certain operator. Using tools from Reproducing Kernel Hilbert Spaces (RKHSs), we develop a new estimator that computes importance ratios of stationary distributions, without knowledge of how the off-policy data are collected. We analyze its asymptotic consistency and finite-sample generalization. Experiments on benchmarks verify the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2003.11126

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

2024-12-15 · Juntao Dai, Yaodong Yang, Qian Zheng, Gang Pan

A key aspect of Safe Reinforcement Learning (Safe RL) involves estimating the constraint condition for the next policy, which is crucial for guiding the optimization of safe policy updates. However, the existing Advantag…

reinforcement-learningReinforcement LearningSafe Reinforcement Learning

Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation

2019-10-16 · ICLR 2020 1 · Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou 외

Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that sign…

Density Ratio EstimationOff-policy evaluation

Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation

2018-10-29 · NeurIPS 2018 12 · Qiang Liu, Lihong Li, Ziyang Tang, Dengyong Zhou

We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (near…

On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent Samples

2023-03-07 · Mustafa O. Karabag, Ufuk Topcu

Offline reinforcement learning (offline RL) considers problems where learning is performed using only previously collected samples and is helpful for the settings in which collecting new data is costly or risky. In model…

Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1

Q-Learning with Shift-Aware Upper Confidence Bound in Non-Stationary Reinforcement Learning

2025-10-03 · Ha Manh Bui, Felix Parker, Kimia Ghobadi, Anqi Liu arxiv

We study the Non-Stationary Reinforcement Learning (RL) under distribution shifts in both finite-horizon episodic and infinite-horizon discounted Markov Decision Processes (MDPs). In the finite-horizon case, the transiti…

Computational EfficiencyReinforcement Learning