paper-with-me

Papers

Bandit optimisation of functions in the Matérn kernel RKHS

2020-01-28 · David Janz, David R. Burt, Javier González

We consider the problem of optimising functions in the reproducing kernel Hilbert space (RKHS) of a Mat\'ern kernel with smoothness parameter $\nu$ over the domain $[0,1]^d$ under noisy bandit feedback. Our contribution, the $\pi$-GP-UCB algorithm, is the first practical approach with guaranteed sublinear regret for all $\nu>1$ and $d \geq 1$. Empirical validation suggests better performance and drastically improved computational scalablity compared with its predecessor, Improved GP-UCB.

📄 PDF Abstract BibTeX arXiv:2001.10396

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Kernel $ε$-Greedy for Multi-Armed Bandits with Covariates

2023-06-29 · Sakshi Arya, Bharath K. Sriperumbudur

We consider the $\epsilon$-greedy strategy for the multi-arm bandit with covariates (MABC) problem, where the mean reward functions are assumed to lie in a reproducing kernel Hilbert space (RKHS). We propose to estimate …

Multi-Armed Bandits

Approximation Theory Based Methods for RKHS Bandits

2020-10-23 · Sho Takemori, Masahiro Sato

The RKHS bandit problem (also called kernelized multi-armed bandit problem) is an online optimization problem of non-linear functions with noisy feedback. Although the problem has been extensively studied, there are unsa…

Bandit Convex Optimisation Revisited: FTRL Achieves $\tilde{O}(t^{1/2})$ Regret

2023-02-01 · David Young, Douglas Leith, George Iosifidis

We show that a kernel estimator using multiple function evaluations can be easily converted into a sampling-based bandit estimator with expectation equal to the original kernel estimate. Plugging such a bandit estimator …

Adaptation to Misspecified Kernel Regularity in Kernelised Bandits

2023-04-26 · Yusha Liu, Aarti Singh

In continuum-armed bandit problems where the underlying function resides in a reproducing kernel Hilbert space (RKHS), namely, the kernelised bandit problems, an important open problem remains of how well learning algori…

Model Selection

Laplacian Kernelized Bandit

2026-01-01 · Shuang Wu, Arash A. Amini arxiv

We study multi-user contextual bandits where users are related by a graph and their reward functions exhibit both non-linear behavior and graph homophily. We introduce a principled joint penalty for the collection of use…