paper-with-me

홈 › Papers

Towards Scalable and Robust Structured Bandits: A Meta-Learning Framework

2022-02-26 · Runzhe Wan, Lin Ge, Rui Song

Online learning in large-scale structured bandits is known to be challenging due to the curse of dimensionality. In this paper, we propose a unified meta-learning framework for a general class of structured bandit problems where the parameter space can be factorized to item-level. The novel bandit algorithm is general to be applied to many popular problems,scalable to the huge parameter and action spaces, and robust to the specification of the generalization model. At the core of this framework is a Bayesian hierarchical model that allows information sharing among items via their features, upon which we design a meta Thompson sampling algorithm. Three representative examples are discussed thoroughly. Both theoretical analysis and numerical results support the usefulness of the proposed method.

📄 PDF Abstract BibTeX arXiv:2202.13227

Code (0)

등록된 구현이 없습니다.

Tasks

Meta-LearningThompson Sampling

Similar Papers 제목 키워드 기반

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

2026-09-09 · Hao Li, Jie Xu, Zheng Xie arxiv

Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-l…

Influence Diagram Bandits

2020-01-01 · ICML 2020 1 · Tong Yu, Branislav Kveton, Zheng Wen, Ruiyi Zhang 외

We propose a novel framework for structured bandits, which we call influence diagram bandit. Our framework captures complicated statistical dependencies between actions, latent variables, and observations; and unifies an…

Learning-To-RankPosition

Metadata-based Multi-Task Bandits with Bayesian Hierarchical Models

2021-08-13 · NeurIPS 2021 12 · Runzhe Wan, Lin Ge, Rui Song

How to explore efficiently is a central problem in multi-armed bandits. In this paper, we introduce the metadata-based multi-task bandit problem, where the agent needs to solve a large number of related multi-armed bandi…

Multi-Armed BanditsThompson Sampling

Influence Diagram Bandits: Variational Thompson Sampling for Structured Bandit Problems

2020-07-09 · Tong Yu, Branislav Kveton, Zheng Wen, Ruiyi Zhang 외

We propose a novel framework for structured bandits, which we call an influence diagram bandit. Our framework captures complex statistical dependencies between actions, latent variables, and observations; and thus unifie…

Thompson Sampling

Modified Meta-Thompson Sampling for Linear Bandits and Its Bayes Regret Analysis

2024-09-10 · Hao Li, Dong Liang, Zheng Xie

Meta-learning is characterized by its ability to learn how to learn, enabling the adaptation of learning strategies across different tasks. Recent research introduced the Meta-Thompson Sampling (Meta-TS), which meta-lear…

Meta-LearningMulti-Armed BanditsThompson Sampling