paper-with-me

Multi-Armed Bandits

1개 벤치마크 · 논문 1,377편 · 이 태스크의 논문 보기 →

Benchmarks

Mushroom

결과 2개

Most implemented

Papers

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

2026-08-18 · Kaifei Wang, Yinyu Ye, Han Zhong arxiv

Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instab…

Multi-Armed Bandits

Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms

2026-08-14 · Maoli Liu, Zhuohua Li, John C. S. Lui arxiv

We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al. [2023], where the learner queries each arm or action through a quantum reward oracle or its inverse. Prior work give…

Multi-Armed Bandits

Learning from Local Walks on Dynamic Graphs with Bandit Feedback

2026-07-12 · Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen arxiv

We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its curr…

Multi-Armed Bandits

Leveraging Similarities in Multi-Armed Bandits

2026-06-22 · Khaled Eldowa, Thibaud Rahier, Augustin Cablant, Panayotis Mertikopoulos 외 arxiv

In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure. We study online learning with a similari…

Multi-Armed Bandits

Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

2026-06-18 · Asaf Cassel, Aviv Rosenberg arxiv

Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical…

Reinforcement LearningMulti-Armed Bandits

Bayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits

2026-06-17 · Lennert Saerens, Bram Silue, Eleni Litsa, Peter Vrancx 외 arxiv

Identifying Pareto optimal solutions is critical to support multi-objective decision-making. We introduce the first anytime Multi-Objective Multi-Armed Bandit algorithm for the Pareto Set Identification problem, taking a…

Multi-Armed Bandits

전체 1,377편 보기 →