paper-with-me

홈 › Papers

Rate-optimal Bayesian Simple Regret in Best Arm Identification

2021-11-18 · Junpei Komiyama, Kaito Ariu, Masahiro Kato, Chao Qin

We consider best arm identification in the multi-armed bandit problem. Assuming certain continuity conditions of the prior, we characterize the rate of the Bayesian simple regret. Differing from Bayesian regret minimization (Lai, 1987), the leading term in the Bayesian simple regret derives from the region where the gap between optimal and suboptimal arms is smaller than $\sqrt{\frac{\log T}{T}}$. We propose a simple and easy-to-compute algorithm with its leading term matching with the lower bound up to a constant factor; simulation results support our theoretical findings.

📄 PDF Abstract BibTeX arXiv:2111.09885

Code (1)

jkomiyama/bayesbai_paper 공식 구현

Similar Papers 제목 키워드 기반

Suboptimal Performance of the Bayes Optimal Algorithm in Frequentist Best Arm Identification

2022-02-10 · Junpei Komiyama

We consider the fixed-budget best arm identification problem with rewards following normal distributions. In this problem, the forecaster is given $K$ arms (or treatments) and $T$ time steps. The forecaster attempts to f…

UCB Exploration for Fixed-Budget Bayesian Best Arm Identification

2024-08-09 · Rong J. B. Zhu, Yanqi Qiu

We study best-arm identification (BAI) in the fixed-budget setting. Adaptive allocations based on upper confidence bounds (UCBs), such as UCBE, are known to work well in BAI. However, it is well-known that its optimal re…

A Causal Bandit Approach to Learning Good Atomic Interventions in Presence of Unobserved Confounders

2021-07-06 · Aurghya Maiti, Vineet Nair, Gaurav Sinha

We study the problem of determining the best intervention in a Causal Bayesian Network (CBN) specified only by its causal graph. We model this as a stochastic multi-armed bandit (MAB) problem with side-information, where…

qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization

2023-03-28 · Raul Astudillo, Zhiyuan Jerry Lin, Eytan Bakshy, Peter I. Frazier

Preferential Bayesian optimization (PBO) is a framework for optimizing a decision maker's latent utility function using preference feedback. This work introduces the expected utility of the best option (qEUBO) as a novel…

Bayesian Optimization

Bayesian Design Principles for Frequentist Sequential Learning

2023-10-01 · Yunbei Xu, Assaf Zeevi

We develop a general theory to optimize the frequentist regret for sequential learning problems, where efficient bandit and reinforcement learning algorithms can be derived from unified Bayesian principles. We propose a …

Multi-Armed Banditsreinforcement-learningReinforcement Learning