paper-with-me

Papers

Decision Variance in Online Learning

2018-07-24 · Sattar Vakili, Alexis Boukouvalas, Qing Zhao

Online learning has traditionally focused on the expected rewards. In this paper, a risk-averse online learning problem under the performance measure of the mean-variance of the rewards is studied. Both the bandit and full information settings are considered. The performance of several existing policies is analyzed, and new fundamental limitations on risk-averse learning is established. In particular, it is shown that although a logarithmic distribution-dependent regret in time $T$ is achievable (similar to the risk-neutral problem), the worst-case (i.e. minimax) regret is lower bounded by $\Omega(T)$ (in contrast to the $\Omega(\sqrt{T})$ lower bound in the risk-neutral problem). This sharp difference from the risk-neutral counterpart is caused by the the variance in the player's decisions, which, while absent in the regret under the expected reward criterion, contributes to excess mean-variance due to the non-linearity of this risk measure. The role of the decision variance in regret performance reflects a risk-averse player's desire for robust decisions and outcomes.

📄 PDF Abstract BibTeX arXiv:1807.09089

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scale-Free Algorithms for Online Linear Optimization

2015-02-19 · Francesco Orabona, David Pal

We design algorithms for online linear optimization that have optimal regret and at the same time do not need to know any upper or lower bounds on the norm of the loss vectors. We achieve adaptiveness to norms of loss ve…

Variance reduction combining pre-experiment and in-experiment data

2024-10-11 · Zhexiao Lin, Pablo Crespo

Online controlled experiments (A/B testing) are essential in data-driven decision-making for many companies. Increasing the sensitivity of these experiments, particularly with a fixed sample size, relies on reducing the …

Decision MakingSensitivity

Statistical Inference for Online Decision Making via Stochastic Gradient Descent

2020-10-14 · Haoyu Chen, Wenbin Lu, Rui Song

Online decision making aims to learn the optimal decision rule by making personalized decisions and updating the decision rule recursively. It has become easier than before with the help of big data, but new challenges a…

Decision Making

Online A-Optimal Design and Active Linear Regression

2019-06-20 · Xavier Fontaine, Pierre Perrault, Michal Valko, Vianney Perchet

We consider in this paper the problem of optimal experiment design where a decision maker can choose which points to sample to obtain an estimate $\hat{\beta}$ of the hidden parameter $\beta^{\star}$ of an underlying lin…

regression

Online learning in bandits with predicted context

2023-07-26 · Yongyi Guo, Ziping Xu, Susan Murphy

We consider the contextual bandit problem where at each time, the agent only has access to a noisy version of the context and the error variance (or an estimator of this variance). This setting is motivated by a wide ran…

Decision Making