paper-with-me

홈 › Papers

On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes

2025-07-18 · Mathieu Godbout, Audrey Durand arxiv

It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail when based on the dual formulation, yet the root cause of this failure remains unclear. We expand on these findings by shifting focus from policy optimization to the seemingly simpler task of policy evaluation. We show that evaluating the static CVaR of a given policy can be framed as two distinct minimization problems. We introduce a set of ``risk-assignment consistency constraints'' that must be satisfied for their solutions to match and we demonstrate that an empty intersection of these constraints is the source of previously observed evaluation errors. Quantifying the evaluation error as the \emph{CVaR evaluation gap}, we demonstrate that the issues observed when optimizing over the dual-based CVaR DP are explained by the returned policy having a non-zero CVaR evaluation gap. Finally, we leverage our proposed risk-assignment constraints perspective to prove that the search for a single, uniformly optimal policy on the dual CVaR decomposition is fundamentally limited, identifying an MDP where no single policy can be optimal across all initial risk levels.

📄 PDF Abstract BibTeX arXiv:2507.14005

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes

2023-04-24 · NeurIPS 2023 11 · Jia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek Petrik

Optimizing static risk-averse objectives in Markov decision processes is difficult because they do not admit standard dynamic programming equations common in Reinforcement Learning (RL) algorithms. Dynamic programming de…

Reinforcement Learning (RL)

RMIX: Risk-Sensitive Multi-Agent Reinforcement Learning

2021-01-01 · Wei Qiu, Xinrun Wang, Runsheng Yu, Xu He 외

Centralized training with decentralized execution (CTDE) has become an important paradigm in multi-agent reinforcement learning (MARL). Current CTDE-based methods rely on restrictive decompositions of the centralized val…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement Learning

2025-01-03 · Mehrdad Moghimi, Hyejin Ku

In domains such as finance, healthcare, and robotics, managing worst-case scenarios is critical, as failure to do so can lead to catastrophic outcomes. Distributional Reinforcement Learning (DRL) provides a natural frame…

Decision MakingDistributional Reinforcement Learning

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

2026-02-03 · Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon 외 arxiv

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depe…

On the Generalization and Robustness in Conditional Value-at-Risk

2026-02-20 · Dinesh Karthik Mulumudi, Piyushi Manupriya, Gholamali Aminian, Anant Raj arxiv

Conditional Value-at-Risk (CVaR) is a widely used risk-sensitive objective for learning under rare but high-impact losses, yet its statistical behavior under heavy-tailed data remains poorly understood. Unlike expectatio…