paper-with-me

홈 › Papers

Hierarchical Causal Bandit

2021-03-07 · Ruiyang Song, Stefano Rini, Kuang Xu

Causal bandit is a nascent learning model where an agent sequentially experiments in a causal network of variables, in order to identify the reward-maximizing intervention. Despite the model's wide applicability, existing analytical results are largely restricted to a parallel bandit version where all variables are mutually independent. We introduce in this work the hierarchical causal bandit model as a viable path towards understanding general causal bandits with dependent variables. The core idea is to incorporate a contextual variable that captures the interaction among all variables with direct effects. Using this hierarchical framework, we derive sharp insights on algorithmic design in causal bandits with dependent arms and obtain nearly matching regret bounds in the case of a binary context.

📄 PDF Abstract BibTeX arXiv:2103.04215

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

2026-05-28 · Tim Woydt, Paul-David Zuercher arxiv

Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is made; standard bandit and reinforcement-learning theory does not ca…

Causal Bandits with Unknown Graph Structure

2021-06-05 · NeurIPS 2021 12 · Yangyi Lu, Amirhossein Meisami, Ambuj Tewari

In causal bandit problems, the action set consists of interventions on variables of a causal graph. Several researchers have recently studied such bandit problems and pointed out their practical applications. However, al…

Causal Bandits with Propagating Inference

2018-06-06 · ICML 2018 7 · Akihiro Yabe, Daisuke Hatano, Hanna Sumita, Shinji Ito 외

Bandit is a framework for designing sequential experiments. In each experiment, a learner selects an arm $A \in \mathcal{A}$ and obtains an observation corresponding to $A$. Theoretically, the tight regret lower-bound fo…

Causal Bandits: Learning Good Interventions via Causal Inference

2016-06-10 · NeurIPS 2016 12 · Finnian Lattimore, Tor Lattimore, Mark D. Reid

We study the problem of using causal models to improve the rate at which good interventions can be learned online in a stochastic environment. Our formalism combines multi-arm bandits and causal inference to model a nove…

Causal Inference

Chronological Causal Bandits

2021-12-03 · Neil Dhir

This paper studies an instance of the multi-armed bandit (MAB) problem, specifically where several causal MABs operate chronologically in the same dynamical system. Practically the reward distribution of each bandit is g…

Decision Making