paper-with-me

Papers

Pessimistic Risk-Aware Policy Learning in Contextual Bandits

2026-05-15 · Yilong Wan, Yuqiang Li, Xianyi Wu arxiv

We study risk-aware offline policy learning, aiming to learn a decision rule from logged data that is optimal under general risk criteria. This problem is crucial in high-stakes domains where online interaction is infeasible and adverse outcomes must be carefully controlled. However, existing literature on offline contextual bandits either centers on expected-reward criteria or restricts risk considerations to policy evaluation instead of optimization. In this work, we propose a unified distributional framework for optimizing Lipschitz-continuous risk functionals, a broad class of risk measures encompassing mean-variance, entropic risk, and conditional value-at-risk, among others. By developing novel empirical concentration inequalities for importance sampling-based distributional estimators, our analysis derives data-dependent suboptimality bounds with an $\tilde{\mathcal{O}}(1/\sqrt{n})$ rate, without relying on restrictive uniform overlap assumptions. This rate is minimax optimal and matches that of risk-neutral offline policy optimization, indicating that optimizing general Lipschitz risk criteria incurs no additional statistical cost relative to the expected-reward.

📄 PDF Abstract BibTeX arXiv:2605.15620

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits

2023-06-13 · Lequn Wang, Akshay Krishnamurthy, Aleksandrs Slivkins

We consider offline policy optimization (OPO) in contextual bandits, where one is given a fixed dataset of logged interactions. While pessimistic regularizers are typically used to mitigate distribution shift, prior impl…

Multi-Armed Bandits

Fast Best-in-Class Regret for Contextual Bandits

2025-10-17 · Samuel Girard, Aurelien Bibaut, Arthur Gretton, Nathan Kallus 외 arxiv

We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or imposing model restrictions on losses or…

Off-Policy Risk Assessment in Contextual Bandits

2021-04-18 · NeurIPS 2021 12 · Audrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar Azizzadenesheli

Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set of objectives, most research on off-poli…

Multi-Armed BanditsOff-policy evaluation

Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits

2025-02-09 · Qingyue Zhao, Kaixuan Ji, Heyang Zhao, Tong Zhang 외

Although many popular reinforcement learning algorithms are underpinned by $f$-divergence regularization, their sample complexity with respect to the \emph{regularized objective} still lacks a tight characterization. In …

Multi-Armed Bandits

Mathematics of statistical sequential decision-making: concentration, risk-awareness and modelling in stochastic bandits, with applications to bariatric surgery

2024-05-03 · Patrick Saux

This thesis aims to study some of the mathematical challenges that arise in the analysis of statistical sequential decision-making algorithms for postoperative patients follow-up. Stochastic bandits (multiarmed, contextu…

Decision MakingInterpretable Machine LearningMulti-Armed BanditsSequential Decision Making+1