Generalized Risk-Aversion in Stochastic Multi-Armed Bandits
We consider the problem of minimizing the regret in stochastic multi-armed bandit, when the measure of goodness of an arm is not the mean return, but some general function of the mean and the variance.We characterize the conditions under which learning is possible and present examples for which no natural algorithm can achieve sublinear regret.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Armed BanditsSimilar Papers 제목 키워드 기반
Risk-Aversion in Multi-armed Bandits
In stochastic multi--armed bandits the objective is to solve the exploration--exploitation dilemma and ultimately maximize the expected reward. Nonetheless, in many practical problems, maximizing the expected reward is n…
Multi-Armed BanditsA Central Limit Theorem, Loss Aversion and Multi-Armed Bandits
This paper studies a multi-armed bandit problem where the decision-maker is loss averse, in particular she is risk averse in the domain of gains and risk loving in the domain of losses. The focus is on large horizons. Co…
Multi-Armed BanditsProbabilistic risk aversion for generalized rank-dependent functions
Probabilistic risk aversion, defined through quasi-convexity in probabilistic mixtures, is a common useful property in decision analysis. We study a general class of non-monotone mappings, called the generalized rank-dep…
ManagementDiversification for infinite-mean Pareto models without risk aversion
We study stochastic dominance between portfolios of independent and identically distributed (iid) extremely heavy-tailed (i.e., infinite-mean) Pareto random variables. With the notion of majorization order, we show that …
Deep Hedging: Continuous Reinforcement Learning for Hedging of General Portfolios across Multiple Risk Aversions
We present a method for finding optimal hedging policies for arbitrary initial portfolios and market states. We develop a novel actor-critic algorithm for solving general risk-averse stochastic control problems and use i…
Reinforcement Learning (RL)