Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear Models
We consider the problem of offline policy evaluation~(OPE) with Markov decision processes~(MDPs), where the goal is to estimate the utility of given decision-making policies based on static datasets. Recently, theoretical understanding of OPE has been rapidly advanced under (approximate) realizability assumptions, i.e., where the environments of interest are well approximated with the given hypothetical models. On the other hand, the OPE under unrealizability has not been well understood as much as in the realizable setting despite its importance in real-world applications.To address this issue, we study the behavior of a simple existing OPE method called the linear direct method~(DM) under the unrealizability. Consequently, we obtain an asymptotically exact characterization of the OPE error in a doubly robust form. Leveraging this result, we also establish the nonparametric consistency of the tile-coding estimators under quite mild assumptions.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingSimilar Papers 제목 키워드 기반
Optimal Sensor Collaboration for Parameter Tracking Using Energy Harvesting Sensors
In this paper, we design an optimal sensor collaboration strategy among neighboring nodes while tracking a time-varying parameter using wireless sensor networks in the presence of imperfect communication channels. The se…
Asymptotically Efficient Off-Policy Evaluation for Tabular Reinforcement Learning
We consider the problem of off-policy evaluation for reinforcement learning, where the goal is to estimate the expected reward of a target policy $\pi$ using offline data collected by running a logging policy $\mu$. Stan…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Optimal Learning for Sequential Decision Making for Expensive Cost Functions with Stochastic Binary Feedbacks
We consider the problem of sequentially making decisions that are rewarded by "successes" and "failures" which can be predicted through an unknown relationship that depends on a partially controllable vector of attribute…
Decision MakingMulti-Armed BanditsSequential Decision MakingPessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning
Offline Reinforcement Learning (RL) aims to learn policies from previously collected datasets without exploring the environment. Directly applying off-policy algorithms to offline RL usually fails due to the extrapolatio…
D4RLOffline RLreinforcement-learningReinforcement Learning+2On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with respect to KL-regularized performance m…
Multi-Armed Bandits