Risk-averse estimation, an axiomatic approach to inference, and Wallace-Freeman without MML
We define a new class of Bayesian point estimators, which we refer to as risk averse. Using this definition, we formulate axioms that provide natural requirements for inference, e.g. in a scientific setting, and show that for well-behaved estimation problems the axioms uniquely characterise an estimator. Namely, for estimation problems in which some parameter values have a positive posterior probability (such as, e.g., problems with a discrete hypothesis space), the axioms characterise Maximum A Posteriori (MAP) estimation, whereas elsewhere (such as in continuous estimation) they characterise the Wallace-Freeman estimator. Our results provide a novel justification for the Wallace-Freeman estimator, which previously was derived only as an approximation to the information-theoretic Strict Minimum Message Length estimator. By contrast, our derivation requires neither approximations nor coding.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Risk averse non-stationary multi-armed bandits
This paper tackles the risk averse multi-armed bandits problem when incurred losses are non-stationary. The conditional value-at-risk (CVaR) is used as the objective function. Two estimation methods are proposed for this…
Multi-Armed BanditsRisk Averse Value Expansion for Sample Efficient and Robust Policy Learning
Model-based Reinforcement Learning(RL) has shown great advantage in sample-efficiency, but suffers from poor asymptotic performance and high inference cost. A promising direction is to combine model-based reinforcement l…
Model-based Reinforcement LearningMuJoCoreinforcement-learningReinforcement Learning+1Risk-Averse No-Regret Learning in Online Convex Games
We consider an online stochastic game with risk-averse agents whose goal is to learn optimal decisions that minimize the risk of incurring significantly high costs. Specifically, we use the Conditional Value at Risk (CVa…
Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning
We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variable. MVPI enjoys great flexibility in tha…
MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)Partial Uncertainty and Applications to Risk-Averse Valuation
This paper introduces an intermediary between conditional expectation and conditional sublinear expectation, called R-conditioning. The R-conditioning of a random-vector in $L^2$ is defined as the best $L^2$-estimate, gi…