paper-with-me

홈 › Papers

A study of Thompson Sampling with Parameter h

2017-10-05 · Qiang Ha

Thompson Sampling algorithm is a well known Bayesian algorithm for solving stochastic multi-armed bandit. At each time step the algorithm chooses each arm with probability proportional to it being the current best arm. We modify the strategy by introducing a paramter h which alters the importance of the probability of an arm being the current best arm. We show that the optimality of Thompson sampling is robust to this perturbation within a range of parameter values for two arm bandits.

📄 PDF Abstract BibTeX arXiv:1710.02174

Code (0)

등록된 구현이 없습니다.

Tasks

Thompson Sampling

Similar Papers 제목 키워드 기반

Policy Gradient Optimization of Thompson Sampling Policies

2020-06-30 · Seungki Min, Ciamac C. Moallemi, Daniel J. Russo

We study the use of policy gradient algorithms to optimize over a class of generalized Thompson sampling policies. Our central insight is to view the posterior parameter sampled by Thompson sampling as a kind of pseudo-a…

Policy Gradient MethodsThompson Sampling

Odds-Ratio Thompson Sampling to Control for Time-Varying Effect

2020-03-04 · Sulgi Kim, Kyung-Min Kim

Multi-armed bandit methods have been used for dynamic experiments particularly in online services. Among the methods, thompson sampling is widely used because it is simple but shows desirable performance. Many thompson s…

Thompson Sampling

Diffusion Approximations for Thompson Sampling

2021-05-19 · Lin Fan, Peter W. Glynn

We study the behavior of Thompson sampling from the perspective of weak convergence. In the regime with small $\gamma > 0$, where the gaps between arm means scale as $\sqrt{\gamma}$ and over time horizons that scale as $…

Multi-Armed BanditsThompson Sampling

Efficient and Adaptive Posterior Sampling Algorithms for Bandits

2024-05-02 · Bingshan Hu, Zhiming Huang, Tianyue H. Zhang, Mathias Lécuyer 외

We study Thompson Sampling-based algorithms for stochastic bandits with bounded rewards. As the existing problem-dependent regret bound for Thompson Sampling with Gaussian priors [Agrawal and Goyal, 2017] is vacuous when…

Thompson Sampling

Thompson Sampling for Parameterized Markov Decision Processes with Uninformative Actions

2023-05-13 · Michael Gimelfarb, Michael Jong Kim

We study parameterized MDPs (PMDPs) in which the key parameters of interest are unknown and must be learned using Bayesian inference. One key defining feature of such models is the presence of "uninformative" actions tha…

Bayesian InferenceThompson Sampling