paper-with-me

Papers

Position-based Multiple-play Bandit Problem with Unknown Position Bias

2017-12-01 · NeurIPS 2017 12 · Junpei Komiyama, Junya Honda, Akiko Takeda

Motivated by online advertising, we study a multiple-play multi-armed bandit problem with position bias that involves several slots and the latter slots yield fewer rewards. We characterize the hardness of the problem by deriving an asymptotic regret bound. We propose the Permutation Minimum Empirical Divergence (PMED) algorithm and derive its asymptotically optimal regret bound. Because of the uncertainty of the position bias, the optimal algorithm for such a problem requires non-convex optimizations that are different from usual partial monitoring and semi-bandit problems. We propose a cutting-plane method and related bi-convex relaxation for these optimizations by using auxiliary variables.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Similar Papers 제목 키워드 기반

Unknown Delay for Adversarial Bandit Setting with Multiple Play

2020-10-01 · Olusola T. Odeyomi

This paper addresses the problem of unknown delays in adversarial multi-armed bandit (MAB) with multiple play. Existing work on similar game setting focused on only the case where the learner selects an arm in each round…

Adversarial Sleeping Bandit Problems with Multiple Plays: Algorithm and Ranking Application

2023-07-27 · Jianjun Yuan, Wei Lee Woon, Ludovik Coba

This paper presents an efficient algorithm to solve the sleeping bandit with multiple plays problem in the context of an online recommendation system. The problem involves bounded, adversarial loss and unknown i.i.d. dis…

Multitask Bandit Learning Through Heterogeneous Feedback Aggregation

2020-10-29 · Zhi Wang, Chicheng Zhang, Manish Kumar Singh, Laurel D. Riek 외

In many real-world applications, multiple agents seek to learn how to perform highly related yet slightly different tasks in an online bandit learning protocol. We formulate this problem as the $\epsilon$-multi-player mu…

Optimal cross-learning for contextual bandits with unknown context distributions

2024-01-03 · NeurIPS 2023 11 · Jon Schneider, Julian Zimmert

We consider the problem of designing contextual bandit algorithms in the ``cross-learning'' setting of Balseiro et al., where the learner observes the loss for the action they play in all possible contexts, not just the …

Multi-Armed Bandits

Multiple-Play Bandits in the Position-Based Model

2016-06-08 · NeurIPS 2016 12 · Paul Lagrée, Claire Vernade, Olivier Cappé

Sequentially learning to place items in multi-position displays or lists is a task that can be cast into the multiple-play semi-bandit setting. However, a major concern in this context is when the system cannot decide wh…

modelPosition