paper-with-me

Papers

Conservative Contextual Combinatorial Cascading Bandit

2021-04-17 · Kun Wang, Canzhe Zhao, Shuai Li, Shuo Shao

Conservative mechanism is a desirable property in decision-making problems which balance the tradeoff between the exploration and exploitation. We propose the novel \emph{conservative contextual combinatorial cascading bandit ($C^4$-bandit)}, a cascading online learning game which incorporates the conservative mechanism. At each time step, the learning agent is given some contexts and has to recommend a list of items but not worse than the base strategy and then observes the reward by some stopping rules. We design the $C^4$-UCB algorithm to solve the problem and prove its n-step upper regret bound for two situations: known baseline reward and unknown baseline reward. The regret in both situations can be decomposed into two terms: (a) the upper bound for the general contextual combinatorial cascading bandit; and (b) a constant term for the regret from the conservative mechanism. We also improve the bound of the conservative contextual combinatorial bandit as a by-product. Experiments on synthetic data demonstrate its advantages and validate our theoretical analysis.

📄 PDF Abstract BibTeX arXiv:2104.08615

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingRecommendation Systems

Similar Papers 제목 키워드 기반

Cascading Contextual Assortment Bandits

2023-09-21 · NeurIPS 2023 11

We present a new combinatorial bandit model, the \textit{cascading contextual assortment bandit}. This model serves as a generalization of both existing cascading bandits and assortment bandits, broadening their applicab…

Contextual Combinatorial Conservative Bandits

2019-11-26 · Xiaojin Zhang, Shuai Li, Weiwen Liu, Shengyu Zhang

The problem of multi-armed bandits (MAB) asks to make sequential decisions while balancing between exploitation and exploration, and have been successfully applied to a wide range of practical scenarios. Various algorith…

Multi-Armed Bandits

Multi-User Contextual Cascading Bandits for Personalized Recommendation

2025-08-19 · Jiho Park, Huiwen Jia arxiv

We introduce a Multi-User Contextual Cascading Bandit model, a new combinatorial bandit framework that captures realistic online advertising scenarios where multiple users interact with sequentially displayed items simul…

A One-Size-Fits-All Solution to Conservative Bandit Problems

2020-12-14 · Yihan Du, Siwei Wang, Longbo Huang

In this paper, we study a family of conservative bandit problems (CBPs) with sample-path reward constraints, i.e., the learner's reward performance must be at least as well as a given baseline at any time. We propose a O…

AllMulti-Armed Bandits

Contextual Combinatorial Bandits with Probabilistically Triggered Arms

2023-03-30 · Xutong Liu, Jinhang Zuo, Siwei Wang, John C. S. Lui 외

We study contextual combinatorial bandits with probabilistically triggered arms (C$^2$MAB-T) under a variety of smoothness conditions that capture a wide range of applications, such as contextual cascading bandits and co…