paper-with-me

홈 › Papers

Doubly Robust Estimator for Off-Policy Evaluation with Large Action Spaces

2023-08-07 · Tatsuhiro Shimizu, Laura Forastiere

We study Off-Policy Evaluation (OPE) in contextual bandit settings with large action spaces. The benchmark estimators suffer from severe bias and variance tradeoffs. Parametric approaches suffer from bias due to difficulty specifying the correct model, whereas ones with importance weight suffer from variance. To overcome these limitations, Marginalized Inverse Propensity Scoring (MIPS) was proposed to mitigate the estimator's variance via embeddings of an action. Nevertheless, MIPS is unbiased under the no direct effect, which assumes that the action embedding completely mediates the effect of an action on a reward. To overcome the dependency on these unrealistic assumptions, we propose a Marginalized Doubly Robust (MDR) estimator. Theoretical analysis shows that the proposed estimator is unbiased under weaker assumptions than MIPS while reducing the variance against MIPS. The empirical experiment verifies the supremacy of MDR against existing estimators with large action spaces.

📄 PDF Abstract BibTeX arXiv:2308.03443

Code (1)

tatsu432/DR-estimator-OPE-large-action 공식 구현

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Doubly robust off-policy evaluation with shrinkage

2019-07-22 · ICML 2020 1 · Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav Dudík

We propose a new framework for designing estimators for off-policy evaluation in contextual bandits. Our approach is based on the asymptotically optimal doubly robust estimator, but we shrink the importance weights to mi…

Model SelectionMulti-Armed BanditsOff-policy evaluation

Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning

2019-12-11 · Riashat Islam, Raihan Seraj, Samin Yeasar Arnob, Doina Precup

We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+2

The Adaptive Doubly Robust Estimator for Policy Evaluation in Adaptive Experiments and a Paradox Concerning Logging Policy

2020-10-08 · Masahiro Kato, Shota Yasui, Kenichiro McAlinn

The doubly robust (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper propose…

Causal InferenceTime Series Analysis

Off-Policy Evaluation Using Information Borrowing and Context-Based Switching

2021-12-18 · Sutanoy Dasgupta, Yabo Niu, Kishan Panaganti, Dileep Kalathil 외

We consider the off-policy evaluation (OPE) problem in contextual bandits, where the goal is to estimate the value of a target policy using the data collected by a logging policy. Most popular approaches to the OPE are v…

Multi-Armed BanditsOff-policy evaluation

Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy

2024-04-02 · Kyungbok Lee, Myunghee Cho Paik

We introduce a novel doubly-robust (DR) off-policy evaluation (OPE) estimator for Markov decision processes, DRUnknown, designed for situations where both the logging policy and the value function are unknown. The propos…

Multi-Armed BanditsOff-policy evaluation