paper-with-me

홈 › Papers

Budgeted Multi-Armed Bandits with Asymmetric Confidence Intervals

2023-06-12 · Marco Heyden, Vadim Arzamasov, Edouard Fouché, Klemens Böhm

We study the stochastic Budgeted Multi-Armed Bandit (MAB) problem, where a player chooses from $K$ arms with unknown expected rewards and costs. The goal is to maximize the total reward under a budget constraint. A player thus seeks to choose the arm with the highest reward-cost ratio as often as possible. Current state-of-the-art policies for this problem have several issues, which we illustrate. To overcome them, we propose a new upper confidence bound (UCB) sampling policy, $\omega$-UCB, that uses asymmetric confidence intervals. These intervals scale with the distance between the sample mean and the bounds of a random variable, yielding a more accurate and tight estimation of the reward-cost ratio compared to our competitors. We show that our approach has logarithmic regret and consistently outperforms existing policies in synthetic and real settings.

📄 PDF Abstract BibTeX arXiv:2306.07071

Code (1)

heymarco/omegaucb 공식 구현

Tasks

Multi-Armed Bandits

Similar Papers 제목 키워드 기반

Adaptive Budgeted Multi-Armed Bandits for IoT with Dynamic Resource Constraints

2025-05-05 · Shubham Vaishnav, Praveen Kumar Donta, Sindri Magnússon

Internet of Things (IoT) systems increasingly operate in environments where devices must respond in real time while managing fluctuating resource constraints, including energy and bandwidth. Yet, current approaches often…

Multi-Armed Bandits

Thompson Sampling for Budgeted Multi-armed Bandits

2015-05-01 · Yingce Xia, Haifang Li, Tao Qin, Nenghai Yu 외

Thompson sampling is one of the earliest randomized algorithms for multi-armed bandits (MAB). In this paper, we extend the Thompson sampling to Budgeted MAB, where there is random cost for pulling an arm and the total co…

Multi-Armed BanditsThompson Sampling

Improving Thompson Sampling via Information Relaxation for Budgeted Multi-armed Bandits

2024-08-28 · Woojin Jeong, Seungki Min

We consider a Bayesian budgeted multi-armed bandit problem, in which each arm consumes a different amount of resources when selected and there is a budget constraint on the total amount of resources that can be used. Bud…

Multi-Armed BanditsThompson Sampling

Budgeted Combinatorial Multi-Armed Bandits

2022-02-08 · Debojit Das, Shweta Jain, Sujit Gujar

We consider a budgeted combinatorial multi-armed bandit setting where, in every round, the algorithm selects a super-arm consisting of one or more arms. The goal is to minimize the total expected regret after all rounds …

Multi-Armed Bandits

Confidence-Budget Matching for Sequential Budgeted Learning

2021-02-05 · Yonathan Efroni, Nadav Merlis, Aadirupa Saha, Shie Mannor

A core element in decision-making under uncertainty is the feedback on the quality of the performed actions. However, in many applications, such feedback is restricted. For example, in recommendation systems, repeatedly …

Decision MakingDecision Making Under UncertaintyMulti-Armed BanditsRecommendation Systems