paper-with-me

홈 › Papers

Leave No One Undermined: Policy Targeting with Regret Aversion

2025-06-19 · Toru Kitagawa, Sokbae Lee, Chen Qiu

While the importance of personalized policymaking is widely recognized, fully personalized implementation remains rare in practice. We study the problem of policy targeting for a regret-averse planner when training data gives a rich set of observable characteristics while the assignment rules can only depend on its subset. Grounded in decision theory, our regret-averse criterion reflects a planner's concern about regret inequality across the population, which generally leads to a fractional optimal rule due to treatment effect heterogeneity beyond the average treatment effects conditional on the subset characteristics. We propose a debiased empirical risk minimization approach to learn the optimal rule from data. Viewing our debiased criterion as a weighted least squares problem, we establish new upper and lower bounds for the excess risk, indicating a convergence rate of 1/n and asymptotic efficiency in certain cases. We apply our approach to the National JTPA Study and the International Stroke Trial.

📄 PDF Abstract BibTeX arXiv:2506.16430

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Concurrent bandits and cognitive radio networks

2014-04-22 · Orly Avner, Shie Mannor

We consider the problem of multiple users targeting the arms of a single multi-armed stochastic bandit. The motivation for this problem comes from cognitive radio networks, where selfish users need to coexist without any…

Collision Avoidance

Policy Targeting under Network Interference

2019-06-24 · Davide Viviano

This paper discusses the problem of estimating treatment allocation rules under network interference. I propose a method with several attractive features for applications: (i) it does not rely on the correct specificatio…

Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

2024-10-03 · Francesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi, Nicola Gatti

We study online learning in \emph{constrained MDPs} (CMDPs), focusing on the goal of attaining sublinear strong regret and strong cumulative constraint violation. Differently from their standard (weak) counterparts, thes…

Fair Policy Targeting

2020-05-25 · Davide Viviano, Jelena Bradic

One of the major concerns of targeting interventions on individuals in social welfare programs is discrimination: individualized treatments may induce disparities across sensitive attributes such as age, gender, or race.…

Fairness

Stochastic Multi-armed Bandits: Optimal Trade-off among Optimality, Consistency, and Tail Risk

2023-09-21 · NeurIPS 2023 11

We consider the stochastic multi-armed bandit problem and fully characterize the interplays among three desired properties for policy design: worst-case optimality, instance-dependent consistency, and light-tailed risk. …