paper-with-me

홈 › Papers

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

2026-02-03 · Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon, Erick Delage arxiv

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depends on entire trajectories without admitting a recursive Bellman decomposition in the underlying Markov decision process. A classical resolution relies on state augmentation with a continuous variable. However, unless restricted to a specialized class of admissible value functions, this formulation induces sparse rewards and degenerate fixed points. In this work, we propose a novel formulation of the static CVaR objective based on augmentation. Our alternative approach leads to a Bellman operator with: (1) dense per-step rewards; (2) contracting properties on the full space of bounded value functions. Building on this theoretical foundation, we develop risk-averse value iteration and model-free Q-learning algorithms that rely on discretized augmented states. We further provide convergence guarantees and approximation error bounds due to discretization. Empirical results demonstrate that our algorithms successfully learn CVaR-sensitive policies and achieve effective performance-safety trade-offs.

📄 PDF Abstract BibTeX arXiv:2602.03778

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Risk-Averse Reinforcement Learning via Dynamic Time-Consistent Risk Measures

2023-01-14 · Xian Yu, Siqian Shen

Traditional reinforcement learning (RL) aims to maximize the expected total reward, while the risk of uncertain outcomes needs to be controlled to ensure reliable performance in a risk-averse setting. In this paper, we c…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Partial Policy Iteration for L1-Robust Markov Decision Processes

2020-06-16 · Chin Pang Ho, Marek Petrik, Wolfram Wiesemann

Robust Markov decision processes (MDPs) allow to compute reliable solutions for dynamic decision problems whose evolution is modeled by rewards and partially-known transition probabilities. Unfortunately, accounting for …

Improving PAC Exploration Using the Median Of Means

2016-12-01 · NeurIPS 2016 12 · Jason Pazis, Ronald E. Parr, Jonathan P. How

We present the first application of the median of means in a PAC exploration algorithm for MDPs. Using the median of means allows us to significantly reduce the dependence of our bounds on the range of values that the va…

Twice regularized MDPs and the equivalence between robustness and regularization

2021-10-12 · NeurIPS 2021 12 · Esther Derman, Matthieu Geist, Shie Mannor

Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, this significantly increases computational …

Being Optimistic to Be Conservative: Quickly Learning a CVaR Policy

2019-11-05 · Ramtin Keramati, Christoph Dann, Alex Tamkin, Emma Brunskill

While maximizing expected return is the goal in most reinforcement learning approaches, risk-sensitive objectives such as conditional value at risk (CVaR) are more suitable for many high-stakes applications. However, rel…

Reinforcement Learning