paper-with-me

홈 › Papers

Fighting Bandits with a New Kind of Smoothness

2015-12-14 · NeurIPS 2015 12 · Jacob Abernethy, Chansoo Lee, Ambuj Tewari

We define a novel family of algorithms for the adversarial multi-armed bandit problem, and provide a simple analysis technique based on convex smoothing. We prove two main results. First, we show that regularization via the \emph{Tsallis entropy}, which includes EXP3 as a special case, achieves the $\Theta(\sqrt{TN})$ minimax regret. Second, we show that a wide class of perturbation methods achieve a near-optimal regret as low as $O(\sqrt{TN \log N})$ if the perturbation distribution has a bounded hazard rate. For example, the Gumbel, Weibull, Frechet, Pareto, and Gamma distributions all satisfy this key property.

📄 PDF Abstract BibTeX arXiv:1512.04152

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fighting Contextual Bandits with Stochastic Smoothing

2018-10-11 · Young Hun Jung, Ambuj Tewari

We introduce a new stochastic smoothing perspective to study adversarial contextual bandit problems. We propose a general algorithm template that represents random perturbation based algorithms and identify several pertu…

Multi-Armed Bandits

DareFightingICE Competition: A Fighting Game Sound Design and AI Competition

2022-03-03 · Ibrahim Khan, Thai Van Nguyen, Xincheng Dai, Ruck Thawonmas

This paper presents a new competition -- at the 2022 IEEE Conference on Games (CoG) -- called DareFightingICE Competition. The competition has two tracks: a sound design track and an AI track. The game platform for this …

Enhanced DareFightingICE Competitions: Sound Design and AI Competitions

2024-03-05 · Ibrahim Khan, Chollakorn Nimpattanavong, Thai Van Nguyen, Kantinan Plupattanakit 외

This paper presents a new and improved DareFightingICE platform, a fighting game platform with a focus on visually impaired players (VIPs), in the Unity game engine. It also introduces the separation of the DareFightingI…

Unity

Bandits on graphs and structures

2026-05-05 · Michal Valko arxiv

The goal of this thesis is to investigate the structural properties of certain sequential problems in order to bring the solutions closer to a practical use. In the first part, we put a special emphasis on structures tha…

RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting

2026-04-23 · Yucheng Xin, Jiacheng Bao, Yubo Dong, Xueqian Wang 외 arxiv

Humanoid robots have demonstrated impressive motor skills in a wide range of tasks, yet whole-body control for humanlike long-time, dynamic fighting remains particularly challenging due to the stringent requirements on a…