paper-with-me

홈 › Papers

Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling

2023-05-14 · Yuta Saito, Qingyang Ren, Thorsten Joachims

We study off-policy evaluation (OPE) of contextual bandit policies for large discrete action spaces where conventional importance-weighting approaches suffer from excessive variance. To circumvent this variance issue, we propose a new estimator, called OffCEM, that is based on the conjunct effect model (CEM), a novel decomposition of the causal effect into a cluster effect and a residual effect. OffCEM applies importance weighting only to action clusters and addresses the residual causal effect through model-based reward estimation. We show that the proposed estimator is unbiased under a new condition, called local correctness, which only requires that the residual-effect model preserves the relative expected reward differences of the actions within each cluster. To best leverage the CEM and local correctness, we also propose a new two-step procedure for performing model-based estimation that minimizes bias in the first step and variance in the second step. We find that the resulting OffCEM estimator substantially improves bias and variance compared to a range of conventional estimators. Experiments demonstrate that OffCEM provides substantial improvements in OPE especially in the presence of many actions.

📄 PDF Abstract BibTeX arXiv:2305.08062

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Doubly Robust Estimator for Off-Policy Evaluation with Large Action Spaces

2023-08-07 · Tatsuhiro Shimizu, Laura Forastiere

We study Off-Policy Evaluation (OPE) in contextual bandit settings with large action spaces. The benchmark estimators suffer from severe bias and variance tradeoffs. Parametric approaches suffer from bias due to difficul…

Off-policy evaluation

Relational Abstractions for Generalized Reinforcement Learning on Symbolic Problems

2022-04-27 · Rushang Karia, Siddharth Srivastava

Reinforcement learning in problems with symbolic state spaces is challenging due to the need for reasoning over long horizons. This paper presents a new approach that utilizes relational abstractions in conjunction with …

Objectreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Leveraging Factored Action Spaces for Off-Policy Evaluation

2023-07-13 · Aaman Rebello, Shengpu Tang, Jenna Wiens, Sonali Parbhoo

Off-policy evaluation (OPE) aims to estimate the benefit of following a counterfactual sequence of actions, given data collected from executed sequences. However, existing OPE estimators often exhibit high bias and high …

counterfactualOff-policy evaluation

Learning and Planning in Complex Action Spaces

2021-04-13 · Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain 외

Many important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for t…

continuous-controlContinuous ControlGame of Go

Balanced off-policy evaluation in general action spaces

2019-06-09 · Arjun Sondhi, David Arbour, Drew Dimmery

Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In …

Binary ClassificationcounterfactualMulti-Armed BanditsOff-policy evaluation