paper-with-me

홈 › Papers

Vector preference-based contextual bandits under distributional shifts

2025-08-21 · Apurv Shukla, P. R. Kumar arxiv

We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and optimistic elimination based policy that self-tunes to the underlying distribution shift. To measure the performance of this policy, we introduce the notion of preference-based regret which measures the performance of a policy in terms of distance between Pareto fronts. We study the performance of this policy by establishing upper bounds on its regret under various assumptions on the nature of distribution shift. Our regret bounds generalize known results for the existing case of no distribution shift and vectorial reward settings, and scale gracefully with problem parameters in presence of distribution shifts.

📄 PDF Abstract BibTeX arXiv:2508.15966

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Linear Bandits with Limited Adaptivity and Learning Distributional Optimal Design

2020-07-04 · Yufei Ruan, Jiaqi Yang, Yuan Zhou

Motivated by practical needs such as large-scale learning, we study the impact of adaptivity constraints to linear contextual bandits, a central problem in online active learning. We consider two popular limited adaptivi…

Active LearningMulti-Armed Bandits

Online and Distribution-Free Robustness: Regression and Contextual Bandits with Huber Contamination

2020-10-08 · Sitan Chen, Frederic Koehler, Ankur Moitra, Morris Yau

In this work we revisit two classic high-dimensional online learning problems, namely linear regression and contextual bandits, from the perspective of adversarial robustness. Existing works in algorithmic robust statist…

Adversarial RobustnessMulti-Armed Banditsregression

Universal and data-adaptive algorithms for model selection in linear contextual bandits

2021-11-08 · Vidya Muthukumar, Akshay Krishnamurthy

Model selection in contextual bandits is an important complementary problem to regret minimization with respect to a fixed model class. We consider the simplest non-trivial instance of model-selection: distinguishing a s…

DiversityModel SelectionMulti-Armed Bandits

Improving Offline Contextual Bandits with Distributional Robustness

2020-11-13 · Otmane Sakhi, Louis Faury, Flavian vasile

This paper extends the Distributionally Robust Optimization (DRO) approach for offline contextual bandits. Specifically, we leverage this framework to introduce a convex reformulation of the Counterfactual Risk Minimizat…

counterfactualMulti-Armed BanditsStochastic Optimization

Online Clustering of Dueling Bandits

2025-02-04 · Zhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie 외

The contextual multi-armed bandit (MAB) is a widely used framework for problems requiring sequential decision-making under uncertainty, such as recommendation systems. In applications involving a large number of users, t…

ClusteringDecision MakingDecision Making Under UncertaintyOnline Clustering+2