Risk-Sensitive RL with Optimized Certainty Equivalents via Reduction to Standard RL
We study Risk-Sensitive Reinforcement Learning (RSRL) with the Optimized Certainty Equivalent (OCE) risk, which generalizes Conditional Value-at-risk (CVaR), entropic risk and Markowitz's mean-variance. Using an augmented Markov Decision Process (MDP), we propose two general meta-algorithms via reductions to standard RL: one based on optimistic algorithms and another based on policy optimization. Our optimistic meta-algorithm generalizes almost all prior RSRL theory with entropic risk or CVaR. Under discrete rewards, our optimistic theory also certifies the first RSRL regret bounds for MDPs with bounded coverability, e.g., exogenous block MDPs. Under discrete rewards, our policy optimization meta-algorithm enjoys both global convergence and local improvement guarantees in a novel metric that lower bounds the true OCE risk. Finally, we instantiate our framework with PPO, construct an MDP, and show that it learns the optimal risk-sensitive policy while prior algorithms provably fail.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Conditional divergence risk measures
Our paper contributes to the theory of conditional risk measures and conditional certainty equivalents. We adopt a random modular approach which proved to be effective in the study of modular convex analysis and conditio…
Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents
The optimized certainty equivalent (OCE) is a family of risk measures that cover important examples such as entropic risk, conditional value-at-risk and mean-variance models. In this paper, we propose a new episodic risk…
reinforcement-learningReinforcement Learning (RL)Robust optimized certainty equivalents and quantiles for loss positions with distribution uncertainty
The paper investigates the robust optimized certainty equivalents and analyzes the relevant properties of them as risk measures for loss positions with distribution uncertainty. On this basis, the robust generalized quan…
Learning Bounds for Risk-sensitive Learning
In risk-sensitive learning, one aims to find a hypothesis that minimizes a risk-averse (or risk-seeking) measure of loss, instead of the standard expected loss. In this paper, we propose to study the generalization prope…
Risk, utility and sensitivity to large losses
Risk and utility functionals are fundamental building blocks in economics and finance. In this paper we investigate under which conditions a risk or utility functional is sensitive to the accumulation of losses in the se…
Sensitivity