paper-with-me

Papers

A Scale Free Algorithm for Stochastic Bandits with Bounded Kurtosis

2017-03-27 · NeurIPS 2017 12 · Tor Lattimore

Existing strategies for finite-armed stochastic bandits mostly depend on a parameter of scale that must be known in advance. Sometimes this is in the form of a bound on the payoffs, or the knowledge of a variance or subgaussian parameter. The notable exceptions are the analysis of Gaussian bandits with unknown mean and variance by Cowan and Katehakis [2015] and of uniform distributions with unknown support [Cowan and Katehakis, 2015]. The results derived in these specialised cases are generalised here to the non-parametric setup, where the learner knows only a bound on the kurtosis of the noise, which is a scale free measure of the extremity of outliers.

📄 PDF Abstract BibTeX arXiv:1703.08937

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lipschitz Bandits with Stochastic Delayed Feedback

2025-09-30 · Zhongxuan Liu, Yue Kang, Thomas C. M. Lee arxiv

The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem…

Stochastic Linear Contextual Bandits with Bounded Noise: A Set-Membership Approach

2026-06-18 · Haonan Xu, Yingying Li arxiv

This paper considers stochastic linear contextual bandits (SLCB) with bounded reward noise. Existing works typically assume sub-Gaussian reward noise and bounded expected rewards, under which the optimal regret bound sca…

The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond

2011-02-12 · Aurélien Garivier, Olivier Cappé

This paper presents a finite-time analysis of the KL-UCB algorithm, an online, horizon-free index policy for stochastic bandit problems. We prove two distinct results: first, for arbitrary bounded rewards, the KL-UCB alg…

Improved Algorithms for Adversarial Bandits with Unbounded Losses

2023-10-03 · Mingyu Chen, Xuezhou Zhang

We consider the Adversarial Multi-Armed Bandits (MAB) problem with unbounded losses, where the algorithms have no prior knowledge on the sizes of the losses. We present UMAB-NN and UMAB-G, two algorithms for non-negative…

Multi-Armed Bandits

Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms

2020-09-20 · Mengfan Xu, Diego Klabjan

We study the challenging exploration incentive problem in both bandit and reinforcement learning, where the rewards are scale-free and potentially unbounded, driven by real-world scenarios and differing from existing wor…

Multi-Armed Banditsreinforcement-learningReinforcement LearningReinforcement Learning (RL)