paper-with-me

Papers

Risk averse non-stationary multi-armed bandits

2021-09-28 · Leo Benac, Frédéric Godin

This paper tackles the risk averse multi-armed bandits problem when incurred losses are non-stationary. The conditional value-at-risk (CVaR) is used as the objective function. Two estimation methods are proposed for this objective function in the presence of non-stationary losses, one relying on a weighted empirical distribution of losses and another on the dual representation of the CVaR. Such estimates can then be embedded into classic arm selection methods such as epsilon-greedy policies. Simulation experiments assess the performance of the arm selection algorithms based on the two novel estimation approaches, and such policies are shown to outperform naive benchmarks not taking non-stationarity into account.

📄 PDF Abstract BibTeX arXiv:2109.13977

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Armed Bandits

Similar Papers 제목 키워드 기반

A Risk-Averse Framework for Non-Stationary Stochastic Multi-Armed Bandits

2023-10-24 · REDA ALAMI, Mohammed Mahfoud, Mastane Achab

In a typical stochastic multi-armed bandit problem, the objective is often to maximize the expected sum of rewards over some time horizon $T$. While the choice of a strategy that accomplishes that is optimal with no addi…

Change Point DetectionMulti-Armed Bandits

A Central Limit Theorem, Loss Aversion and Multi-Armed Bandits

2021-06-10 · Zengjing Chen, Larry G. Epstein, Guodong Zhang

This paper studies a multi-armed bandit problem where the decision-maker is loss averse, in particular she is risk averse in the domain of gains and risk loving in the domain of losses. The focus is on large horizons. Co…

Multi-Armed Bandits

Risk-Averse Multi-Armed Bandits with Unobserved Confounders: A Case Study in Emotion Regulation in Mobile Health

2022-09-09 · Yi Shen, Jessilyn Dunn, Michael M. Zavlanos

In this paper, we consider a risk-averse multi-armed bandit (MAB) problem where the goal is to learn a policy that minimizes the risk of low expected return, as opposed to maximizing the expected return itself, which is …

Multi-Armed BanditsTransfer Learning

Exploration vs Exploitation vs Safety: Risk-averse Multi-Armed Bandits

2014-01-06 · Nicolas Galichet, Michèle Sebag, Olivier Teytaud

Motivated by applications in energy management, this paper presents the Multi-Armed Risk-Aware Bandit (MARAB) algorithm. With the goal of limiting the exploration of risky arms, MARAB takes as arm quality its conditional…

energy managementManagementMulti-Armed Bandits

A Unifying Theory of Thompson Sampling for Continuous Risk-Averse Bandits

2021-08-25 · Joel Q. L. Chang, Vincent Y. F. Tan

This paper unifies the design and the analysis of risk-averse Thompson sampling algorithms for the multi-armed bandit problem for a class of risk functionals $\rho$ that are continuous and dominant. We prove generalised …

Thompson Sampling