paper-with-me

홈 › Papers

Robustness and risk management via distributional dynamic programming

2021-12-28 · Mastane Achab, Gergely Neu

In dynamic programming (DP) and reinforcement learning (RL), an agent learns to act optimally in terms of expected long-term return by sequentially interacting with its environment modeled by a Markov decision process (MDP). More generally in distributional reinforcement learning (DRL), the focus is on the whole distribution of the return, not just its expectation. Although DRL-based methods produced state-of-the-art performance in RL with function approximation, they involve additional quantities (compared to the non-distributional setting) that are still not well understood. As a first contribution, we introduce a new class of distributional operators, together with a practical DP algorithm for policy evaluation, that come with a robust MDP interpretation. Indeed, our approach reformulates through an augmented state space where each state is split into a worst-case substate and a best-case substate, whose values are maximized by safe and risky policies respectively. Finally, we derive distributional operators and DP algorithms solving a new control task: How to distinguish safe from risky optimal actions in order to break ties in the space of optimal policies?

📄 PDF Abstract BibTeX arXiv:2112.15430

Code (0)

등록된 구현이 없습니다.

Tasks

Distributional Reinforcement LearningManagementreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Optimizing Return Distributions with Distributional Dynamic Programming

2025-01-22 · Bernardo Ávila Pires, Mark Rowland, Diana Borsa, Zhaohan Daniel Guo 외

We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforcement learning as a special case. Previous distributional DP methods c…

Toward a Scalable Upper Bound for a CVaR-LQ Problem

2021-03-03 · Margaret P. Chapman, Laurent Lessard

We study a linear-quadratic, optimal control problem on a discrete, finite time horizon with distributional ambiguity, in which the cost is assessed via Conditional Value-at-Risk (CVaR). We take steps toward deriving a s…

Distributional Method for Risk Averse Reinforcement Learning

2023-02-27 · Ziteng Cheng, Sebastian Jaimungal, Nick Martin

We introduce a distributional method for learning the optimal policy in risk averse Markov decision process with finite state action spaces, latent costs, and stationary dynamics. We assume sequential observations of sta…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Nesterov Accelerated Distributional Deep Hedging for Efficient Volatility Risk Management

2025-02-25 · Lei Zhao, Lin Cai, Wu-Sheng Lu

In the field of financial derivatives trading, managing volatility risk is crucial for protecting investment portfolios from market changes. Traditional Vega hedging strategies, which often rely on basic and rule-based m…

Distributional Reinforcement LearningManagementreinforcement-learningReinforcement Learning

Lipschitz Networks and Distributional Robustness

2018-09-04 · Zac Cranko, Simon Kornblith, Zhan Shi, Richard Nock

Robust risk minimisation has several advantages: it has been studied with regards to improving the generalisation properties of models and robustness to adversarial perturbation. We bound the distributionally robust risk…