paper-with-me

Papers

Policy Learning with $α$-Expected Welfare

2025-05-01 · Yanqin Fan, Yuan Qi, Gaoqian Xu

This paper proposes an optimal policy that targets the average welfare of the worst-off $\alpha$-fraction of the post-treatment outcome distribution. We refer to this policy as the $\alpha$-Expected Welfare Maximization ($\alpha$-EWM) rule, where $\alpha \in (0,1]$ denotes the size of the subpopulation of interest. The $\alpha$-EWM rule interpolates between the expected welfare ($\alpha=1$) and the Rawlsian welfare ($\alpha\rightarrow 0$). For $\alpha\in (0,1)$, an $\alpha$-EWM rule can be interpreted as a distributionally robust EWM rule that allows the target population to have a different distribution than the study population. Using the dual formulation of our $\alpha$-expected welfare function, we propose a debiased estimator for the optimal policy and establish its asymptotic upper regret bounds. In addition, we develop asymptotically valid inference for the optimal welfare based on the proposed debiased estimator. We examine the finite sample performance of the debiased estimator and inference via both real and synthetic data.

📄 PDF Abstract BibTeX arXiv:2505.00256

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

Multi-agent Multi-armed Bandits with Minimum Reward Guarantee Fairness

2025-02-21 · Piyushi Manupriya, Himanshu, SakethaNath Jagarlapudi, Ganesh Ghalme

We investigate the problem of maximizing social welfare while ensuring fairness in a multi-agent multi-armed bandit (MA-MAB) setting. In this problem, a centralized decision-maker takes actions over time, generating rand…

FairnessMulti-Armed Bandits

Welfare and Fairness in Multi-objective Reinforcement Learning

2022-11-30 · Zimeng Fan, Nianli Peng, Muhang Tian, Brandon Fain

We study fair multi-objective reinforcement learning in which an agent must learn a policy that simultaneously achieves high reward on multiple dimensions of a vector-valued reward. Motivated by the fair resource allocat…

FairnessMulti-Objective Reinforcement LearningQ-Learningreinforcement-learning+2

Optimal and Robust Disclosure of Public Information

2022-03-31 · Takashi Ui

A policymaker discloses public information to interacting agents who also acquire costly private information. More precise public information reduces the precision and cost of acquired private information. Considering th…

General Bayesian Policy Learning

2026-02-27 · Masahiro Kato arxiv

This study proposes a General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from a given set to maximize expected welfare. Typical examples include treatme…

Learning under Invariable Bayesian Safety

2020-06-08 · Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe Tennenholtz

A recent body of work addresses safety constraints in explore-and-exploit systems. Such constraints arise where, for example, exploration is carried out by individuals whose welfare should be balanced with overall welfar…