Risk-Sensitive Reinforcement Learning: a Martingale Approach to Reward Uncertainty
We introduce a novel framework to account for sensitivity to rewards uncertainty in sequential decision-making problems. While risk-sensitive formulations for Markov decision processes studied so far focus on the distribution of the cumulative reward as a whole, we aim at learning policies sensitive to the uncertain/stochastic nature of the rewards, which has the advantage of being conceptually more meaningful in some cases. To this end, we present a new decomposition of the randomness contained in the cumulative reward based on the Doob decomposition of a stochastic process, and introduce a new conceptual tool - the \textit{chaotic variation} - which can rigorously be interpreted as the risk measure of the martingale component associated to the cumulative reward process. We innovate on the reinforcement learning side by incorporating this new risk-sensitive approach into model-free algorithms, both policy gradient and value function based, and illustrate its relevance on grid world and portfolio optimization problems.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingPortfolio Optimizationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Sequential Decision MakingSimilar Papers 제목 키워드 기반
Continuous-time Risk-sensitive Reinforcement Learning via Quadratic Variation Penalty
This paper studies continuous-time risk-sensitive reinforcement learning (RL) under the entropy-regularized, exploratory diffusion process formulation with the exponential-form objective. The risk-sensitive objective ari…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Risk-Sensitive Q-Learning in Continuous Time with Application to Dynamic Portfolio Selection
This paper studies the problem of risk-sensitive reinforcement learning (RSRL) in continuous time, where the environment is characterized by a controllable stochastic differential equation (SDE) and the objective is a po…
Reinforcement LearningRisk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret
We study risk-sensitive reinforcement learning in episodic Markov decision processes with unknown transition kernels, where the goal is to optimize the total reward under the risk measure of exponential utility. We propo…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Risk Sensitive Model-Based Reinforcement Learning using Uncertainty Guided Planning
Identifying uncertainty and taking mitigating actions is crucial for safe and trustworthy reinforcement learning agents, especially when deployed in high-risk environments. In this paper, risk sensitivity is promoted in …
Model-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Bayesian Robust Optimization for Imitation Learning
One of the main challenges in imitation learning is determining what action an agent should take when outside the state distribution of the demonstrations. Inverse reinforcement learning (IRL) can enable generalization t…
Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)