paper-with-me

홈 › Papers

Restarted Bayesian Online Change-point Detection for Non-Stationary Markov Decision Processes

2023-04-01 · REDA ALAMI, Mohammed Mahfoud, Eric Moulines

We consider the problem of learning in a non-stationary reinforcement learning (RL) environment, where the setting can be fully described by a piecewise stationary discrete-time Markov decision process (MDP). We introduce a variant of the Restarted Bayesian Online Change-Point Detection algorithm (R-BOCPD) that operates on input streams originating from the more general multinomial distribution and provides near-optimal theoretical guarantees in terms of false-alarm rate and detection delay. Based on this, we propose an improved version of the UCRL2 algorithm for MDPs with state transition kernel sampled from a multinomial distribution, which we call R-BOCPD-UCRL2. We perform a finite-time performance analysis and show that R-BOCPD-UCRL2 enjoys a favorable regret bound of $O\left(D O \sqrt{A T K_T \log\left (\frac{T}{\delta} \right) + \frac{K_T \log \frac{K_T}{\delta}}{\min\limits_\ell \: \mathbf{KL}\left( {\mathbf{\theta}^{(\ell+1)}}\mid\mid{\mathbf{\theta}^{(\ell)}}\right)}}\right)$, where $D$ is the largest MDP diameter from the set of MDPs defining the piecewise stationary MDP setting, $O$ is the finite number of states (constant over all changes), $A$ is the finite number of actions (constant over all changes), $K_T$ is the number of change points up to horizon $T$, and $\mathbf{\theta}^{(\ell)}$ is the transition kernel during the interval $[c_\ell, c_{\ell+1})$, which we assume to be multinomially distributed over the set of states $\mathbb{O}$. Interestingly, the performance bound does not directly scale with the variation in MDP state transition distributions and rewards, ie. can also model abrupt changes. In practice, R-BOCPD-UCRL2 outperforms the state-of-the-art in a variety of scenarios in synthetic environments. We provide a detailed experimental setup along with a code repository (upon publication) that can be used to easily reproduce our experiments.

📄 PDF Abstract BibTeX arXiv:2304.00232

Code (0)

등록된 구현이 없습니다.

Tasks

Change Point DetectionReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Restarted Bayesian Online Change-point Detector achieves Optimal Detection Delay

2020-01-01 · ICML 2020 1 · REDA ALAMI, Odalric-Ambrym Maillard, Raphaël Féraud

In this paper, we consider the problem of sequential change-point detection where both the change-points and the distributions before and after the change are assumed to be unknown. For this key problem in statistical an…

Change Point DetectionLearning Theory

Distributed Consensus Algorithm for Decision-Making in Multi-agent Multi-armed Bandit

2023-06-09 · Xiaotong Cheng, Setareh Maghsudi

We study a structured multi-agent multi-armed bandit (MAMAB) problem in a dynamic environment. A graph reflects the information-sharing structure among agents, and the arms' reward distributions are piecewise-stationary …

Change Point DetectionDecision Making

Bayesian Online Prediction of Change Points

2019-02-12 · Diego Agudelo-España, Sebastian Gomez-Gonzalez, Stefan Bauer, Bernhard Schölkopf 외

Online detection of instantaneous changes in the generative process of a data sequence generally focuses on retrospective inference of such change points without considering their future occurrences. We extend the Bayesi…

Bayesian InferenceChange Point DetectionPrediction

A Risk-Averse Framework for Non-Stationary Stochastic Multi-Armed Bandits

2023-10-24 · REDA ALAMI, Mohammed Mahfoud, Mastane Achab

In a typical stochastic multi-armed bandit problem, the objective is often to maximize the expected sum of rewards over some time horizon $T$. While the choice of a strategy that accomplishes that is optimal with no addi…

Change Point DetectionMulti-Armed Bandits

Lagged Exact Bayesian Online Changepoint Detection with Parameter Estimation

2017-10-09 · Michael Byrd, Linh Nghiem, Jing Cao

Identifying changes in the generative process of sequential data, known as changepoint detection, has become an increasingly important topic for a wide variety of fields. A recently developed approach, which we call EXac…

parameter estimation