Multi-Armed Bandits
1개 벤치마크 · 논문 1,377편 · 이 태스크의 논문 보기 →
Benchmarks
Mushroom
Most implemented
Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling
Neural Contextual Bandits with UCB-based Exploration
Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling
Hypothesis Generation with Large Language Models
Off-Policy Evaluation for Large Action Spaces via Embeddings
Online Limited Memory Neural-Linear Bandits with Likelihood Matching
Papers
Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits
Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instab…
Multi-Armed BanditsQuantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms
We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al. [2023], where the learner queries each arm or action through a quantum reward oracle or its inverse. Prior work give…
Multi-Armed BanditsLearning from Local Walks on Dynamic Graphs with Bandit Feedback
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its curr…
Multi-Armed BanditsLeveraging Similarities in Multi-Armed Bandits
In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure. We study online learning with a similari…
Multi-Armed BanditsQuantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning
Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical…
Reinforcement LearningMulti-Armed BanditsBayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits
Identifying Pareto optimal solutions is critical to support multi-objective decision-making. We introduce the first anytime Multi-Objective Multi-Armed Bandit algorithm for the Pareto Set Identification problem, taking a…
Multi-Armed Bandits