QUOTA: The Quantile Option Architecture for Reinforcement Learning
In this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL). In QUOTA, decision making is based on quantiles of a value distribution, not only the mean. QUOTA provides a new dimension for exploration via making use of both optimism and pessimism of a value distribution. We demonstrate the performance advantage of QUOTA in both challenging video games and physical robot simulators.
Code (3)
Tasks
Decision MakingDistributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
EX-DRL: Hedging Against Heavy Losses with EXtreme Distributional Reinforcement Learning
Recent advancements in Distributional Reinforcement Learning (DRL) for modeling loss distributions have shown promise in developing hedging strategies in derivatives markets. A common approach in DRL involves learning th…
Distributional Reinforcement Learningquantile regressionDistributional Reinforcement Learning on Path-dependent Options
We reinterpret and propose a framework for pricing path-dependent financial derivatives by estimating the full distribution of payoffs using Distributional Reinforcement Learning (DistRL). Unlike traditional methods that…
Distributional Reinforcement Learningreinforcement-learningReinforcement LearningUncertainty QuantificationPricing methods for $α$-quantile and perpetual early exercise options based on Spitzer identities
We present new numerical schemes for pricing perpetual Bermudan and American options as well as $\alpha$-quantile options. This includes a new direct calculation of the optimal exercise barrier for early-exercise options…
Safety-Aware Reinforcement Learning for Control via Risk-Sensitive Action-Value Iteration and Quantile Regression
Mainstream approximate action-value iteration reinforcement learning (RL) algorithms suffer from overestimation bias, leading to suboptimal policies in high-variance stochastic environments. Quantile-based action-value i…
quantile regressionReinforcement Learning (RL)Gamma and Vega Hedging Using Deep Distributional Reinforcement Learning
We show how D4PG can be used in conjunction with quantile regression to develop a hedging strategy for a trader responsible for derivatives that arrive stochastically and depend on a single underlying asset. We assume th…
Distributional Reinforcement LearningPositionquantile regressionreinforcement-learning+2