paper-with-me

홈 › Papers

A KL-LUCB algorithm for Large-Scale Crowdsourcing

2017-12-01 · NeurIPS 2017 12 · Ervin Tanczos, Robert Nowak, Bob Mankoff

This paper focuses on best-arm identification in multi-armed bandits with bounded rewards. We develop an algorithm that is a fusion of lil-UCB and KL-LUCB, offering the best qualities of the two algorithms in one method. This is achieved by proving a novel anytime confidence bound for the mean of bounded distributions, which is the analogue of the LIL-type bounds recently developed for sub-Gaussian distributions. We corroborate our theoretical results with numerical experiments based on the New Yorker Cartoon Caption Contest.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Armed Bandits

Similar Papers 제목 키워드 기반

Batched Neural Bandits

2021-02-25 · Quanquan Gu, Amin Karbasi, Khashayar Khosravi, Vahab Mirrokni 외

In many sequential decision-making problems, the individuals are split into several batches and the decision-maker is only allowed to change her policy at the end of batches. These batch problems have a large number of a…

Decision MakingSequential Decision Making

Best Arm Identification with Possibly Biased Offline Data

2025-05-29 · Le Yang, Vincent Y. F. Tan, Wang Chi Cheung

We study the best arm identification (BAI) problem with potentially biased offline data in the fixed confidence setting, which commonly arises in real-world scenarios such as clinical trials. We prove an impossibility re…

On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits

2026-05-25 · Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan arxiv

We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimization or pure exploration paradigms. The…

Multi-Armed Bandits

Almost Optimal Variance-Constrained Best Arm Identification

2022-01-25 · Yunlong Hou, Vincent Y. F. Tan, Zixin Zhong

We design and analyze VA-LUCB, a parameter-free algorithm, for identifying the best arm under the fixed-confidence setup and under a stringent constraint that the variance of the chosen arm is strictly smaller than a giv…

Distributed Contextual Linear Bandits with Minimax Optimal Communication Cost

2022-05-26 · Sanae Amani, Tor Lattimore, András György, Lin F. Yang

We study distributed contextual linear bandits with stochastic contexts, where $N$ agents act cooperatively to solve a linear bandit-optimization problem with $d$-dimensional features over the course of $T$ rounds. For t…