paper-with-me

홈 › Papers

Toward a Scalable Upper Bound for a CVaR-LQ Problem

2021-03-03 · Margaret P. Chapman, Laurent Lessard

We study a linear-quadratic, optimal control problem on a discrete, finite time horizon with distributional ambiguity, in which the cost is assessed via Conditional Value-at-Risk (CVaR). We take steps toward deriving a scalable dynamic programming approach to upper-bound the optimal value function for this problem. This dynamic program yields a novel, tunable risk-averse control policy, which we compare to existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2103.02136

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst Path

2022-06-06 · Yihan Du, Siwei Wang, Longbo Huang

In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each step, and focuses on tightly controlling th…

Autonomous Drivingreinforcement-learningReinforcement Learning (RL)

Concentration bounds for CVaR estimation: The cases of light-tailed and heavy-tailed distributions

2019-01-04 · ICML 2020 1 · Prashanth L. A., Krishna Jagannathan, Ravi Kumar Kolla

Conditional Value-at-Risk (CVaR) is a widely used risk metric in applications such as finance. We derive concentration bounds for CVaR estimates, considering separately the cases of light-tailed and heavy-tailed distribu…

Multi-Armed Bandits

Near-Optimal Sample Complexity for Iterated CVaR Reinforcement Learning with a Generative Model

2025-03-11 · Zilong Deng, Simon Khan, Shaofeng Zou

In this work, we study the sample complexity problem of risk-sensitive Reinforcement Learning (RL) with a generative model, where we aim to maximize the Conditional Value at Risk (CVaR) with risk tolerance level $\tau$ a…

Reinforcement Learning (RL)

Optimal Thompson Sampling strategies for support-aware CVaR bandits

2020-12-10 · Dorian Baudry, Romain Gautron, Emilie Kaufmann, Odalric-Ambryn Maillard

In this paper we study a multi-arm bandit problem in which the quality of each arm is measured by the Conditional Value at Risk (CVaR) at some level alpha of the reward distribution. While existing works in this setting …

Thompson Sampling

Optimizing Conditional Value-At-Risk of Black-Box Functions

2021-12-01 · NeurIPS 2021 12 · Quoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick Jaillet

This paper presents two Bayesian optimization (BO) algorithms with theoretical performance guarantee to maximize the conditional value-at-risk (CVaR) of a black-box function: CV-UCB and CV-TS which are based on the well-…

Bayesian OptimizationThompson Sampling