paper-with-me

홈 › Papers

Lenient Regret and Good-Action Identification in Gaussian Process Bandits

2021-02-11 · Xu Cai, Selwyn Gomes, Jonathan Scarlett

In this paper, we study the problem of Gaussian process (GP) bandits under relaxed optimization criteria stating that any function value above a certain threshold is "good enough". On the theoretical side, we study various {\em lenient regret} notions in which all near-optimal actions incur zero penalty, and provide upper bounds on the lenient regret for GP-UCB and an elimination algorithm, circumventing the usual $O(\sqrt{T})$ term (with time horizon $T$) resulting from zooming extremely close towards the function maximum. In addition, we complement these upper bounds with algorithm-independent lower bounds. On the practical side, we consider the problem of finding a single "good action" according to a known pre-specified threshold, and introduce several good-action identification algorithms that exploit knowledge of the threshold. We experimentally find that such algorithms can often find a good action faster than standard optimization-based approaches.

📄 PDF Abstract BibTeX arXiv:2102.05793

Code (1)

caitree/GoodAction 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

Lenient Regret for Multi-Armed Bandits

2020-08-10 · Nadav Merlis, Shie Mannor

We consider the Multi-Armed Bandit (MAB) problem, where an agent sequentially chooses actions and observes rewards for the actions it took. While the majority of algorithms try to minimize the regret, i.e., the cumulativ…

Multi-Armed BanditsThompson Sampling

On Regret Bounds of Thompson Sampling for Bayesian Optimization

2026-03-10 · Shion Takeno, Shogo Iwazaki arxiv

We study a widely used Bayesian optimization method, Gaussian process Thompson sampling (GP-TS), under the assumption that the objective function is a sample path from a GP. Compared with the GP upper confidence bound (G…

High-dimensional Nonparametric Contextual Bandit Problem

2025-05-20 · Shogo Iwazaki, Junpei Komiyama, Masaaki Imaizumi

We consider the kernelized contextual bandit problem with a large feature space. This problem involves $K$ arms, and the goal of the forecaster is to maximize the cumulative rewards through learning the relationship betw…

Decision MakingMulti-Armed BanditsRecommendation Systems

Quick Best Action Identification in Linear Bandit Problems

2018-12-02 · Jun Geng, Lifeng Lai

In this paper, we consider a best action identification problem in the stochastic linear bandit setup with a fixed confident constraint. In the considered best action identification problem, instead of minimizing the acc…

Decision Theory for Treatment Choice Problems with Partial Identification

2023-12-29 · José Luis Montiel Olea, Chen Qiu, Jörg Stoye

We apply classical statistical decision theory to a large class of treatment choice problems with partial identification. We show that, in a general class of problems with Gaussian likelihood, all decision rules are admi…

All