Risk-Averse Reinforcement Learning: An Optimal Transport Perspective on Temporal Difference Learning
The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance, frequently without considering risk or safety. In contrast, safe reinforcement learning seeks to reduce or avoid unsafe states. This letter introduces a risk-averse temporal difference algorithm that uses optimal transport theory to direct the agent toward predictable behavior. By incorporating a risk indicator, the agent learns to favor actions with predictable consequences. We evaluate the proposed algorithm in several case studies and show its effectiveness in the presence of uncertainty. The results demonstrate that our method reduces the frequency of visits to risky states while preserving performance. A Python implementation of the algorithm is available at https:// github.com/SAILRIT/Risk-averse-TD-Learning.
Code (1)
Tasks
Decision Makingreinforcement-learningReinforcement LearningSafe Reinforcement LearningSimilar Papers 제목 키워드 기반
Risk-Averse Learning by Temporal Difference Methods
We consider reinforcement learning with performance evaluated by a dynamic risk measure. We construct a projected risk-averse dynamic programming equation and study its properties. Then we propose risk-averse counterpart…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Optimal Transport and Risk Aversion in Kyle's Model of Informed Trading
We establish connections between optimal transport theory and the dynamic version of the Kyle model, including new characterizations of informed trading profits via conjugate duality and Monge-Kantorovich duality. We use…
Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control
We present an historical overview about the connections between the analysis of risk and the control of autonomous systems. We offer two main contributions. Our first contribution is to propose three overlapping paradigm…
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
Prior work on safe Reinforcement Learning (RL) has studied risk-aversion to randomness in dynamics (aleatory) and to model uncertainty (epistemic) in isolation. We propose and analyze a new framework to jointly model the…
Reinforcement Learning (RL)Safe Reinforcement LearningFinding Risk-Averse Shortest Path with Time-dependent Stochastic Costs
In this paper, we tackle the problem of risk-averse route planning in a transportation network with time-dependent and stochastic costs. To solve this problem, we propose an adaptation of the A* algorithm that accommodat…