paper-with-me

Papers

Online Batch Decision-Making with High-Dimensional Covariates

2020-02-21 · Chi-Hua Wang, Guang Cheng

We propose and investigate a class of new algorithms for sequential decision making that interacts with \textit{a batch of users} simultaneously instead of \textit{a user} at each decision epoch. This type of batch models is motivated by interactive marketing and clinical trial, where a group of people are treated simultaneously and the outcomes of the whole group are collected before the next stage of decision. In such a scenario, our goal is to allocate a batch of treatments to maximize treatment efficacy based on observed high-dimensional user covariates. We deliver a solution, named \textit{Teamwork LASSO Bandit algorithm}, that resolves a batch version of explore-exploit dilemma via switching between teamwork stage and selfish stage during the whole decision process. This is made possible based on statistical properties of LASSO estimate of treatment efficacy that adapts to a sequence of batch observations. In general, a rate of optimal allocation condition is proposed to delineate the exploration and exploitation trade-off on the data collection scheme, which is sufficient for LASSO to identify the optimal treatment for observed user covariates. An upper bound on expected cumulative regret of the proposed algorithm is provided.

📄 PDF Abstract BibTeX arXiv:2002.09438

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingMarketingSequential Decision MakingVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Batched Online Contextual Sparse Bandits with Sequential Inclusion of Features

2024-09-13 · Rowan Swiers, Subash Prabanantham, Andrew Maher

Multi-armed Bandits (MABs) are increasingly employed in online platforms and e-commerce to optimize decision making for personalized user experiences. In this work, we focus on the Contextual Bandit problem with linear r…

Decision MakingFairnessMulti-Armed Bandits

Parallelizing Thompson Sampling

2021-06-02 · NeurIPS 2021 12 · Amin Karbasi, Vahab Mirrokni, Mohammad Shadravan

How can we make use of information parallelism in online decision making problems while efficiently balancing the exploration-exploitation trade-off? In this paper, we introduce a batch Thompson Sampling framework for tw…

Decision MakingThompson Sampling

Dynamic Batch Learning in High-Dimensional Sparse Linear Contextual Bandits

2020-08-27 · Zhimei Ren, Zhengyuan Zhou

We study the problem of dynamic batch learning in high-dimensional sparse linear contextual bandits, where a decision maker, under a given maximum-number-of-batch constraint and only able to observe rewards at the end of…

Decision MakingMarketingMulti-Armed BanditsVocal Bursts Intensity Prediction

Sequential Batch Learning in Finite-Action Linear Contextual Bandits

2020-04-14 · Yanjun Han, Zhengqing Zhou, Zhengyuan Zhou, Jose Blanchet 외

We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can …

Decision MakingMulti-Armed BanditsProduct RecommendationSequential Decision Making

A Reinforcement Learning Approach to Online Learning of Decision Trees

2015-07-24 · Abhinav Garlapati, aditi raghunathan, Vaishnavh Nagarajan, Balaraman Ravindran

Online decision tree learning algorithms typically examine all features of a new data point to update model parameters. We propose a novel alternative, Reinforcement Learning- based Decision Trees (RLDT), that uses Reinf…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)