paper-with-me

홈 › Papers

Stochastically Constrained Best Arm Identification with Thompson Sampling

2025-01-07 · Le Yang, Siyang Gao, Cheng Li, Yi Wang

We consider the problem of the best arm identification in the presence of stochastic constraints, where there is a finite number of arms associated with multiple performance measures. The goal is to identify the arm that optimizes the objective measure subject to constraints on the remaining measures. We will explore the popular idea of Thompson sampling (TS) as a means to solve it. To the best of our knowledge, it is the first attempt to extend TS to this problem. We will design a TS-based sampling algorithm, establish its asymptotic optimality in the rate of posterior convergence, and demonstrate its superior performance using numerical examples.

📄 PDF Abstract BibTeX arXiv:2501.03877

Code (0)

등록된 구현이 없습니다.

Tasks

Thompson Sampling

Methods 이 논문이 사용한 방법론

TS Spatio-temporal features extraction that measure the stabilty. The proposed method is based on a compression algorithm named Run Length Encoding. The workflow of the method is…

Similar Papers 제목 키워드 기반

Racing Thompson: an Efficient Algorithm for Thompson Sampling with Non-conjugate Priors

2017-08-16 · ICML 2018 7 · Yichi Zhou, Jun Zhu, Jingwei Zhuo

Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors since these algorithms require to infer…

Thompson Sampling

Top Two Algorithms Revisited

2022-06-13 · Marc Jourdan, Rémy Degenne, Dorian Baudry, Rianne de Heide 외

Top Two algorithms arose as an adaptation of Thompson sampling to best arm identification in multi-armed bandit models (Russo, 2016), for parametric families of arms. They select the next arm to sample from by randomizin…

Thompson SamplingVocal Bursts Valence Prediction

Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits

2018-07-19 · Julian Zimmert, Yevgeny Seldin

We derive an algorithm that achieves the optimal (within constants) pseudo-regret in both adversarial and stochastic multi-armed bandits without prior knowledge of the regime and time horizon. The algorithm is based on o…

Multi-Armed BanditsThompson Sampling

Thompson Exploration with Best Challenger Rule in Best Arm Identification

2023-10-01 · Jongyeong Lee, Junya Honda, Masashi Sugiyama

This paper studies the fixed-confidence best arm identification (BAI) problem in the bandit framework in the canonical single-parameter exponential models. For this problem, many policies have been proposed, but most of …

Thompson Sampling

Bayesian Best-Arm Identification for Selecting Influenza Mitigation Strategies

2017-11-16 · Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Jelena Grujic 외

Pandemic influenza has the epidemic potential to kill millions of people. While various preventive measures exist (i.a., vaccination and school closures), deciding on strategies that lead to their most effective and effi…

Decision MakingThompson Sampling