paper-with-me

Papers

Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms

2025-08-04 · Katherine Avery, Chinmay Pendse, David Jensen arxiv

Causal graphical models can encode large amounts structural knowledge, both from the background knowledge of domain experts and the structural knowledge discovered from randomized experiments or observational data. However, though we may know the general structure of causal relationships, we often do not know the exact causal mechanisms. In this work, we propose a causal multi-armed bandit evaluation and learning algorithm that can reason effectively despite uncertainty over conditional probability distributions. Further, we show how conditional independence testing can be used to choose variables for modeling. We find that the structural equation model (SEM) approach gives more accurate evaluations compared to traditional approaches, particularly as the range of possible causal mechanisms grows. Further, the SEM approach learns low-variance policies, and it learns an optimal policy, assuming the model is sufficiently well-specified. Traditional approaches can converge to local extrema or fail to converge at all.

📄 PDF Abstract BibTeX arXiv:2508.02812

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Asymmetric Graph Error Control with Low Complexity in Causal Bandits

2024-08-20 · Chen Peng, Di Zhang, Urbashi Mitra

In this paper, the causal bandit problem is investigated, in which the objective is to select an optimal sequence of interventions on nodes in a causal graph. It is assumed that the graph is governed by linear structural…

Change DetectionGraph Learning

Evaluating COVID-19 vaccine allocation policies using Bayesian $m$-top exploration

2023-01-30 · Alexandra Cimpean, Timothy Verstraeten, Lander Willem, Niel Hens 외

Individual-based epidemiological models support the study of fine-grained preventive measures, such as tailored vaccine allocation policies, in silico. As individual-based models are computationally intensive, it is pivo…

Worst-case Performance of Greedy Policies in Bandits with Imperfect Context Observations

2022-04-10 · Hongju Park, Mohamad Kazem Shirani Faradonbeh

Contextual bandits are canonical models for sequential decision-making under uncertainty in environments with time-varying components. In this setting, the expected reward of each bandit arm consists of the inner product…

Decision MakingDecision Making Under UncertaintyMulti-Armed BanditsSequential Decision Making

Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood

2026-02-11 · Jiangrong Ouyang, Mingming Gong, Howard Bondell arxiv

Policy inference plays an essential role in the contextual bandit problem. In this paper, we use empirical likelihood to develop a Bayesian inference method for the joint analysis of multiple contextual bandit policies i…

Bayesian Inference

Collapsing Bandits and Their Application to Public Health Intervention

2020-12-01 · NeurIPS 2020 12 · Aditya Mate, Jackson Killian, Haifeng Xu, Andrew Perrault 외

We propose and study Collapsing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully ob…