Learning Robust Options by Conditional Value at Risk Optimization
Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters, these methods only consider either the worst case or the average (ordinary) case for learning options. This limited consideration of the cases often produces options that do not work well in the unconsidered case. In this paper, we propose a conditional value at risk (CVaR)-based method to learn options that work well in both the average and worst cases. We extend the CVaR-based policy gradient method proposed by Chow and Ghavamzadeh (2014) to deal with robust Markov decision processes and then apply the extended method to learning robust options. We conduct experiments to evaluate our method in multi-joint robot control tasks (HopperIceBlock, Half-Cheetah, and Walker2D). Experimental results show that our method produces options that 1) give better worst-case performance than the options learned only to minimize the average-case loss, and 2) give better average-case performance than the options learned only to minimize the worst-case loss.
Code (1)
Similar Papers 제목 키워드 기반
Hedging Conditional Value at Risk with Options
We present a method of hedging Conditional Value at Risk of a position in stock using put options. The result leads to a linear programming problem that can be solved to optimise risk hedging.
PositionDistributional Reinforcement Learning on Path-dependent Options
We reinterpret and propose a framework for pricing path-dependent financial derivatives by estimating the full distribution of payoffs using Distributional Reinforcement Learning (DistRL). Unlike traditional methods that…
Distributional Reinforcement Learningreinforcement-learningReinforcement LearningUncertainty QuantificationRisk-averse Stochastic Optimization for Farm Management Practices and Cultivar Selection Under Uncertainty
Optimizing management practices and selecting the best cultivar for planting play a significant role in increasing agricultural food production and decreasing environmental footprint. In this study, we develop optimizati…
Bayesian OptimizationManagementStochastic OptimizationSystemic Risk of Optioned Portfolios: Controllability and Optimization
We investigate the portfolio selection problem against the systemic risk which is measured by CoVaR. We first demonstrate that the systemic risk of pure stock portfolios is essentially uncontrollable due to the contagion…
Portfolio OptimizationValuation of Exotic Options and Counterparty Games Based on Conditional Diffusion
This paper addresses the challenges of pricing exotic options and structured products, which traditional models often fail to handle due to their inability to capture real-world market phenomena like fat-tailed distribut…