paper-with-me

홈 › Papers

Generalizing distribution of partial rewards for multi-armed bandits with temporally-partitioned rewards

2022-11-13 · Ronald C. van den Broek, Rik Litjens, Tobias Sagis, Luc Siecker, Nina Verbeeke, Pratik Gajane

We investigate the Multi-Armed Bandit problem with Temporally-Partitioned Rewards (TP-MAB) setting in this paper. In the TP-MAB setting, an agent will receive subsets of the reward over multiple rounds rather than the entire reward for the arm all at once. In this paper, we introduce a general formulation of how an arm's cumulative reward is distributed across several rounds, called Beta-spread property. Such a generalization is needed to be able to handle partitioned rewards in which the maximum reward per round is not distributed uniformly across rounds. We derive a lower bound on the TP-MAB problem under the assumption that Beta-spread holds. Moreover, we provide an algorithm TP-UCB-FR-G, which uses the Beta-spread property to improve the regret upper bound in some scenarios. By generalizing how the cumulative reward is distributed, this setting is applicable in a broader range of applications.

📄 PDF Abstract BibTeX arXiv:2211.06883

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Armed Bandits

Similar Papers 제목 키워드 기반

Randomized Allocation with Nonparametric Estimation for Contextual Multi-Armed Bandits with Delayed Rewards

2019-02-03 · Sakshi Arya, Yuhong Yang

We study a multi-armed bandit problem with covariates in a setting where there is a possible delay in observing the rewards. Under some mild assumptions on the probability distributions for the delays and using an approp…

Multi-Armed Bandits

Regime Switching Bandits

2020-01-26 · NeurIPS 2021 12 · Xiang Zhou, Yi Xiong, Ningyuan Chen, Xuefeng Gao

We study a multi-armed bandit problem where the rewards exhibit regime switching. Specifically, the distributions of the random rewards generated from all arms are modulated by a common underlying state modeled as a fini…

Reinforcement Learning

Stochastic Multi-Armed Bandits with Limited Control Variates

2026-03-02 · Arun Verma, Manjesh Kumar Hanawal, Arun Rajkumar arxiv

Motivated by wireless networks where interference or channel state estimates provide partial insight into throughput, we study a variant of the classical stochastic multi-armed bandit problem in which the learner has lim…

Multi-Armed Bandits

Multi-Armed Bandits with Generalized Temporally-Partitioned Rewards

2023-03-01 · Ronald C. van den Broek, Rik Litjens, Tobias Sagis, Luc Siecker 외

Decision-making problems of sequential nature, where decisions made in the past may have an impact on the future, are used to model many practically important applications. In some real-world applications, feedback about…

Decision MakingMulti-Armed Bandits

Multi-Armed Bandit Problem with Temporally-Partitioned Rewards: When Partial Feedback Counts

2022-06-01 · Giulia Romano, Andrea Agostini, Francesco Trovò, Nicola Gatti 외

There is a rising interest in industrial online applications where data becomes available sequentially. Inspired by the recommendation of playlists to users where their preferences can be collected during the listening o…