paper-with-me

Papers

Improving Offline Contextual Bandits with Distributional Robustness

2020-11-13 · Otmane Sakhi, Louis Faury, Flavian vasile

This paper extends the Distributionally Robust Optimization (DRO) approach for offline contextual bandits. Specifically, we leverage this framework to introduce a convex reformulation of the Counterfactual Risk Minimization principle. Besides relying on convex programs, our approach is compatible with stochastic optimization, and can therefore be readily adapted tothe large data regime. Our approach relies on the construction of asymptotic confidence intervals for offline contextual bandits through the DRO framework. By leveraging known asymptotic results of robust estimators, we also show how to automatically calibrate such confidence intervals, which in turn removes the burden of hyper-parameter selection for policy optimization. We present preliminary empirical results supporting the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2011.06835

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualMulti-Armed BanditsStochastic Optimization

Similar Papers 제목 키워드 기반

Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization

2021-11-27 · ICLR 2022 4 · Thanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha Venkatesh

Offline policy learning (OPL) leverages existing data collected a priori for policy optimization without any active exploration. Despite the prevalence and recent interest in this problem, its theoretical and algorithmic…

Multi-Armed Bandits

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

2024-02-11 · Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus 외

In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximat…

Distributional Reinforcement LearningMulti-Armed BanditsOffline RL

Pessimistic Risk-Aware Policy Learning in Contextual Bandits

2026-05-15 · Yilong Wan, Yuqiang Li, Xianyi Wu arxiv

We study risk-aware offline policy learning, aiming to learn a decision rule from logged data that is optimal under general risk criteria. This problem is crucial in high-stakes domains where online interaction is infeas…

Distributionally Robust Policy Evaluation and Learning in Offline Contextual Bandits

2020-01-01 · ICML 2020 1 · Nian Si, Fan Zhang, Zhengyuan Zhou, Jose Blanchet

Policy learning using historical observational data is an important problem that has found widespread applications. However, existing literature rests on the crucial assumption that the future environment where the learn…

Multi-Armed Bandits

Online and Distribution-Free Robustness: Regression and Contextual Bandits with Huber Contamination

2020-10-08 · Sitan Chen, Frederic Koehler, Ankur Moitra, Morris Yau

In this work we revisit two classic high-dimensional online learning problems, namely linear regression and contextual bandits, from the perspective of adversarial robustness. Existing works in algorithmic robust statist…

Adversarial RobustnessMulti-Armed Banditsregression