Estimation of Optimal Dynamic Treatment Assignment Rules under Policy Constraints
Many policies involve dynamics in their treatment assignments, where individuals receive sequential interventions over multiple stages. We study estimation of an optimal dynamic treatment regime that guides the optimal treatment assignment for each individual at each stage based on their history. We propose an empirical welfare maximization approach in this dynamic framework, which estimates the optimal dynamic treatment regime using data from an experimental or quasi-experimental study while satisfying exogenous constraints on policies. The paper proposes two estimation methods: one solves the treatment assignment problem sequentially through backward induction, and the other solves the entire problem simultaneously across all stages. We establish finite-sample upper bounds on worst-case average welfare regrets for these methods and show their optimal $n^{-1/2}$ convergence rates. We also modify the simultaneous estimation method to accommodate intertemporal budget/capacity constraints.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Stochastic Treatment Choice with Empirical Welfare Updating
This paper proposes a novel method to estimate individualised treatment assignment rules. The method is designed to find rules that are stochastic, reflecting uncertainty in estimation of an assignment rule and about its…
Optimal treatment assignment rules under capacity constraints
We study treatment assignment problems under capacity constraints, where a planner aims to maximize social welfare by assigning treatments based on observable covariates. Such constraints are common in practice, as treat…
Gradient Regularized V-Learning for Dynamic Treatment Regimes
Deciding how to optimally treat a patient, including how to select treatments over time among the multiple available treatments, represents one of the most important issues that need to be addressed in medicine today. A …
PAC-Bayesian Treatment Allocation Under Budget Constraints
This paper considers the estimation of treatment assignment rules when the policy maker faces a general budget or resource constraint. Utilizing the PAC-Bayesian framework, we propose new treatment assignment rules that …
Generalization BoundsNear-Optimal Reinforcement Learning in Dynamic Treatment Regimes
A dynamic treatment regime (DTR) consists of a sequence of decision rules, one per stage of intervention, that dictates how to determine the treatment assignment to patients based on evolving treatments and covariates' h…
Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)