paper-with-me

Papers

Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal Approach

2023-12-20 · Wen Huang, Xintao Wu

This paper studies bandit problems where an agent has access to offline data that might be utilized to potentially improve the estimation of each arm's reward distribution. A major obstacle in this setting is the existence of compound biases from the observational data. Ignoring these biases and blindly fitting a model with the biased data could even negatively affect the online learning phase. In this work, we formulate this problem from a causal perspective. First, we categorize the biases into confounding bias and selection bias based on the causal structure they imply. Next, we extract the causal bound for each arm that is robust towards compound biases from biased observational data. The derived bounds contain the ground truth mean reward and can effectively guide the bandit agent to learn a nearly-optimal decision policy. We also conduct regret analysis in both contextual and non-contextual bandit settings and show that prior causal bounds could help consistently reduce the asymptotic regret.

📄 PDF Abstract BibTeX arXiv:2312.12731

Code (0)

등록된 구현이 없습니다.

Tasks

Selection bias

Similar Papers 제목 키워드 기반

Adaptive Experimentation When You Can't Experiment

2024-06-15 · Yao Zhao, Kwang-Sung Jun, Tanner Fiez, Lalit Jain

This paper introduces the \emph{confounded pure exploration transductive linear bandit} (\texttt{CPET-LB}) problem. As a motivating example, often online services cannot directly assign users to specific control or treat…

Experimental Design

Bandits with Partially Observable Confounded Data

2020-06-11 · Guy Tennenholtz, Uri Shalit, Shie Mannor, Yonathan Efroni

We study linear contextual bandits with access to a large, confounded, offline dataset that was sampled from some fixed policy. We show that this problem is closely related to a variant of the bandit problem with side in…

Multi-Armed Bandits

Dual Instrumental Method for Confounded Kernelized Bandits

2022-09-07 · Xueping Gong, Jiheng Zhang

The contextual bandit problem is a theoretically justified framework with wide applications in various fields. While the previous study on this problem usually requires independence between noise and contexts, our work c…

Bias-Robust Bayesian Optimization via Dueling Bandits

2021-05-25 · Johannes Kirschner, Andreas Krause

We consider Bayesian optimization in settings where observations can be adversarially biased, for example by an uncontrolled hidden confounder. Our first contribution is a reduction of the confounded setting to the dueli…

Bayesian Optimization

Semiparametric Contextual Bandits

2018-03-12 · ICML 2018 7 · Akshay Krishnamurthy, Zhiwei Steven Wu, Vasilis Syrgkanis

This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of known action features confounded by an non-…

Multi-Armed Bandits