paper-with-me

Papers

Off-Policy Risk Assessment in Contextual Bandits

2021-04-18 · NeurIPS 2021 12 · Audrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar Azizzadenesheli

Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set of objectives, most research on off-policy evaluation to date focuses on the expected reward. In this paper, we introduce Lipschitz risk functionals, a broad class of objectives that subsumes conditional value-at-risk (CVaR), variance, mean-variance, many distorted risks, and CPT risks, among others. We propose Off-Policy Risk Assessment (OPRA), a framework that first estimates a target policy's CDF and then generates plugin estimates for any collection of Lipschitz risks, providing finite sample guarantees that hold simultaneously over the entire class. We instantiate OPRA with both importance sampling and doubly robust estimators. Our primary theoretical contributions are (i) the first uniform concentration inequalities for both CDF estimators in contextual bandits and (ii) error bounds on our Lipschitz risk estimates, which all converge at a rate of $O(1/\sqrt{n})$.

📄 PDF Abstract BibTeX arXiv:2104.08977

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Armed BanditsOff-policy evaluation

Similar Papers 제목 키워드 기반

Off-Policy Risk Assessment in Markov Decision Processes

2022-09-21 · Audrey Huang, Liu Leqi, Zachary Chase Lipton, Kamyar Azizzadenesheli

Addressing such diverse ends as safety alignment with human preferences, and the efficiency of learning, a growing line of reinforcement learning research focuses on risk functionals that depend on the entire distributio…

Multi-Armed BanditsSafety Alignment

Pessimistic Risk-Aware Policy Learning in Contextual Bandits

2026-05-15 · Yilong Wan, Yuqiang Li, Xianyi Wu arxiv

We study risk-aware offline policy learning, aiming to learn a decision rule from logged data that is optimal under general risk criteria. This problem is crucial in high-stakes domains where online interaction is infeas…

Improving Offline Contextual Bandits with Distributional Robustness

2020-11-13 · Otmane Sakhi, Louis Faury, Flavian vasile

This paper extends the Distributionally Robust Optimization (DRO) approach for offline contextual bandits. Specifically, we leverage this framework to introduce a convex reformulation of the Counterfactual Risk Minimizat…

counterfactualMulti-Armed BanditsStochastic Optimization

On Minimax Optimal Offline Policy Evaluation

2014-09-12 · Lihong Li, Remi Munos, Csaba Szepesvari

This paper studies the off-policy evaluation problem, where one aims to estimate the value of a target policy based on a sample of observations collected by another policy. We first consider the multi-armed bandit case, …

Multi-Armed BanditsOff-policy evaluation

Conditionally Risk-Averse Contextual Bandits

2022-10-24 · Mónika Farsang, Paul Mineiro, Wangda Zhang

Contextual bandits with average-case statistical guarantees are inadequate in risk-averse situations because they might trade off degraded worst-case behaviour for better average performance. Designing a risk-averse cont…

ManagementMulti-Armed Banditsregression