paper-with-me

Papers

A Two-armed Bandit Framework for A/B Testing

2025-07-24 · Jinjuan Wang, Qianglin Wen, Yu Zhang, Xiaodong Yan, Chengchun Shi arxiv

A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods.

📄 PDF Abstract BibTeX arXiv:2507.18118

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningCausal Inference

Similar Papers 제목 키워드 기반

BanditMF: Multi-Armed Bandit Based Matrix Factorization Recommender System

2021-06-21 · Shenghao Xu

Multi-armed bandits (MAB) provide a principled online learning approach to attain the balance between exploration and exploitation. Due to the superior performance and low feedback learning without the learning to act in…

Collaborative FilteringMulti-Armed BanditsRecommendation Systemsvalid

A Batched Multi-Armed Bandit Approach to News Headline Testing

2019-08-17 · Yizhi Mao, Miao Chen, Abhinav Wagle, Junwei Pan 외

Optimizing news headlines is important for publishers and media sites. A compelling headline will increase readership, user engagement and social shares. At Yahoo Front Page, headline testing is carried out using a test-…

ArticlesThompson Sampling

A framework for optimizing COVID-19 testing policy using a Multi Armed Bandit approach

2020-07-28 · Hagit Grushka-Cohen, Raphael Cohen, Bracha Shapira, Jacob Moran-Gilad 외

Testing is an important part of tackling the COVID-19 pandemic. Availability of testing is a bottleneck due to constrained resources and effective prioritization of individuals is necessary. Here, we discuss the impact o…

Decision MakingMulti-Armed Bandits

On the Complexity of Best Arm Identification in Multi-Armed Bandit Models

2014-07-16 · Emilie Kaufmann, Olivier Cappé, Aurélien Garivier

The stochastic multi-armed bandit model is a simple abstraction that has proven useful in many different contexts in statistics and machine learning. Whereas the achievable limit in terms of regret minimization is now we…

LEMMA

Accelerated learning from recommender systems using multi-armed bandit

2019-08-16 · Meisam Hejazinia, Kyler Eastman, Shuqin Ye, Abbas Amirabadi 외

Recommendation systems are a vital component of many online marketplaces, where there are often millions of items to potentially present to users who have a wide variety of wants or needs. Evaluating recommender system a…

Recommendation Systems