Leave No One Undermined: Policy Targeting with Regret Aversion
While the importance of personalized policymaking is widely recognized, fully personalized implementation remains rare in practice. We study the problem of policy targeting for a regret-averse planner when training data gives a rich set of observable characteristics while the assignment rules can only depend on its subset. Grounded in decision theory, our regret-averse criterion reflects a planner's concern about regret inequality across the population, which generally leads to a fractional optimal rule due to treatment effect heterogeneity beyond the average treatment effects conditional on the subset characteristics. We propose a debiased empirical risk minimization approach to learn the optimal rule from data. Viewing our debiased criterion as a weighted least squares problem, we establish new upper and lower bounds for the excess risk, indicating a convergence rate of 1/n and asymptotic efficiency in certain cases. We apply our approach to the National JTPA Study and the International Stroke Trial.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Concurrent bandits and cognitive radio networks
We consider the problem of multiple users targeting the arms of a single multi-armed stochastic bandit. The motivation for this problem comes from cognitive radio networks, where selfish users need to coexist without any…
Collision AvoidancePolicy Targeting under Network Interference
This paper discusses the problem of estimating treatment allocation rules under network interference. I propose a method with several attractive features for applications: (i) it does not rely on the correct specificatio…
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
We study online learning in \emph{constrained MDPs} (CMDPs), focusing on the goal of attaining sublinear strong regret and strong cumulative constraint violation. Differently from their standard (weak) counterparts, thes…
Fair Policy Targeting
One of the major concerns of targeting interventions on individuals in social welfare programs is discrimination: individualized treatments may induce disparities across sensitive attributes such as age, gender, or race.…
FairnessStochastic Multi-armed Bandits: Optimal Trade-off among Optimality, Consistency, and Tail Risk
We consider the stochastic multi-armed bandit problem and fully characterize the interplays among three desired properties for policy design: worst-case optimality, instance-dependent consistency, and light-tailed risk. …