paper-with-me

홈 › Papers

DTR Bandit: Learning to Make Response-Adaptive Decisions With Low Regret

2020-05-06 · Yichun Hu, Nathan Kallus

Dynamic treatment regimes (DTRs) are personalized, adaptive, multi-stage treatment plans that adapt treatment decisions both to an individual's initial features and to intermediate outcomes and features at each subsequent stage, which are affected by decisions in prior stages. Examples include personalized first- and second-line treatments of chronic conditions like diabetes, cancer, and depression, which adapt to patient response to first-line treatment, disease progression, and individual characteristics. While existing literature mostly focuses on estimating the optimal DTR from offline data such as from sequentially randomized trials, we study the problem of developing the optimal DTR in an online manner, where the interaction with each individual affect both our cumulative reward and our data collection for future learning. We term this the DTR bandit problem. We propose a novel algorithm that, by carefully balancing exploration and exploitation, is guaranteed to achieve rate-optimal regret when the transition and reward models are linear. We demonstrate our algorithm and its benefits both in synthetic experiments and in a case study of adaptive treatment of major depressive disorder using real-world data.

📄 PDF Abstract BibTeX arXiv:2005.02791

Code (1)

CausalML/DTR-Bandit

Similar Papers 제목 키워드 기반

Combining Offline Causal Inference and Online Bandit Learning for Data Driven Decision

2020-01-16 · Li Ye, Yishi Lin, Hong Xie, John C. S. Lui

A fundamental question for companies with large amount of logged data is: How to use such logged data together with incoming streaming data to make good decisions? Many companies currently make decisions via online A/B t…

Causal Inference

Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability

2021-11-24 · Aadirupa Saha, Akshay Krishnamurthy

We study the $K$-armed contextual dueling bandit problem, a sequential decision making setting in which the learner uses contextual information to make two decisions, but only observes \emph{preference-based feedback} su…

Decision MakingSequential Decision Making

A note on continuous-time online learning

2024-05-16 · Lexing Ying

In online learning, the data is provided in a sequential order, and the goal of the learner is to make online decisions to minimize overall regrets. This note is concerned with continuous-time models and algorithms for s…

Bayesian Regret Minimization in Offline Bandits

2023-06-02 · Marek Petrik, Guy Tennenholtz, Mohammad Ghavamzadeh

We study how to make decisions that minimize Bayesian regret in offline linear bandits. Prior work suggests that one must take actions with maximum lower confidence bound (LCB) on their reward. We argue that the reliance…

An Asymptotically Optimal Batched Algorithm for the Dueling Bandit Problem

2022-09-25 · Arpit Agarwal, Rohan Ghuge, Viswanath Nagarajan

We study the $K$-armed dueling bandit problem, a variation of the traditional multi-armed bandit problem in which feedback is obtained in the form of pairwise comparisons. Previous learning algorithms have focused on the…

Recommendation Systems