paper-with-me

Papers

Risk-Averse Approximate Dynamic Programming with Quantile-Based Risk Measures

2015-09-07 · Daniel R. Jiang, Warren B. Powell

In this paper, we consider a finite-horizon Markov decision process (MDP) for which the objective at each stage is to minimize a quantile-based risk measure (QBRM) of the sequence of future costs; we call the overall objective a dynamic quantile-based risk measure (DQBRM). In particular, we consider optimizing dynamic risk measures where the one-step risk measures are QBRMs, a class of risk measures that includes the popular value at risk (VaR) and the conditional value at risk (CVaR). Although there is considerable theoretical development of risk-averse MDPs in the literature, the computational challenges have not been explored as thoroughly. We propose data-driven and simulation-based approximate dynamic programming (ADP) algorithms to solve the risk-averse sequential decision problem. We address the issue of inefficient sampling for risk applications in simulated settings and present a procedure, based on importance sampling, to direct samples toward the "risky region" as the ADP algorithm progresses. Finally, we show numerical results of our algorithms in the context of an application involving risk-averse bidding for energy storage.

📄 PDF Abstract BibTeX arXiv:1509.01920

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Boosting CVaR Policy Optimization with Quantile Gradients

2026-01-29 · Yudong Luo, Erick Delage arxiv

Optimizing Conditional Value-at-risk (CVaR) using policy gradient (a.k.a CVaR-PG) faces significant challenges of sample inefficiency. This inefficiency stems from the fact that it focuses on tail-end performance and ove…

Risk-Averse Learning by Temporal Difference Methods

2020-03-02 · Umit Kose, Andrzej Ruszczynski

We consider reinforcement learning with performance evaluated by a dynamic risk measure. We construct a projected risk-averse dynamic programming equation and study its properties. Then we propose risk-averse counterpart…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Distributional Method for Risk Averse Reinforcement Learning

2023-02-27 · Ziteng Cheng, Sebastian Jaimungal, Nick Martin

We introduce a distributional method for learning the optimal policy in risk averse Markov decision process with finite state action spaces, latent costs, and stationary dynamics. We assume sequential observations of sta…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Long-Term Sequential Decision Making under Risk

2026-07-22 · Irmaan, Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib arxiv

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distributi…

Decision Making

Risk-averse Total-reward MDPs with ERM and EVaR

2024-08-30 · Xihong Su, Julien Grand-Clément, Marek Petrik

Optimizing risk-averse objectives in discounted MDPs is challenging because most models do not admit direct dynamic programming equations and require complex history-dependent policies. In this paper, we show that the ri…