paper-with-me

홈 › Papers

Distributional Offline Policy Evaluation with Predictive Error Guarantees

2023-02-19 · Runzhe Wu, Masatoshi Uehara, Wen Sun

We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE). We propose an algorithm called Fitted Likelihood Estimation (FLE), which conducts a sequence of Maximum Likelihood Estimation (MLE) and has the flexibility of integrating any state-of-the-art probabilistic generative models as long as it can be trained via MLE. FLE can be used for both finite-horizon and infinite-horizon discounted settings where rewards can be multi-dimensional vectors. Our theoretical results show that for both finite-horizon and infinite-horizon discounted settings, FLE can learn distributions that are close to the ground truth under total variation distance and Wasserstein distance, respectively. Our theoretical results hold under the conditions that the offline data covers the test policy's traces and that the supervised learning MLE procedures succeed. Experimentally, we demonstrate the performance of FLE with two generative models, Gaussian mixture models and diffusion models. For the multi-dimensional reward setting, FLE with diffusion models is capable of estimating the complicated distribution of the return of a test policy.

📄 PDF Abstract BibTeX arXiv:2302.09456

Code (1)

ziqian2000/fitted-likelihood-estimation 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Test 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator

2025-04-23 · Chenhao Li, Andreas Krause, Marco Hutter

Reinforcement Learning (RL) has demonstrated impressive capabilities in robotic control but remains challenging due to high sample complexity, safety concerns, and the sim-to-real gap. While offline RL eliminates the nee…

Offline RLReinforcement Learning (RL)

Distributional Off-policy Evaluation with Bellman Residual Minimization

2024-02-02 · Sungee Hong, Zhengling Qi, Raymond K. W. Wong

We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many…

Distributional Reinforcement LearningOff-policy evaluation

Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

2023-09-23 · Wenzhuo Zhou, Yuhan Li, Ruoqing Zhu, Annie Qu

We study high-confidence off-policy evaluation in the context of infinite-horizon Markov decision processes, where the objective is to establish a confidence interval (CI) for the target policy value using only offline d…

Off-policy evaluation

Policy Constraint by Only Support Constraint for Offline Reinforcement Learning

2025-03-07 · Yunkai Gao, Jiaming Guo, Fan Wu, Rui Zhang

Offline reinforcement learning (RL) aims to optimize a policy by using pre-collected datasets, to maximize cumulative rewards. However, offline reinforcement learning suffers challenges due to the distributional shift be…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows

2024-05-06 · Minjae Cho, Jonathan P. How, Chuangchuang Sun

Despite notable successes of Reinforcement Learning (RL), the prevalent use of an online learning paradigm prevents its widespread adoption, especially in hazardous or costly scenarios. Offline RL has emerged as an alter…

Causal InferencecounterfactualCounterfactual ReasoningOffline RL+2