paper-with-me

Papers

Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback

2023-07-06 · Yu Chen, Yihan Du, Pihe Hu, Siwei Wang, Desheng Wu, Longbo Huang

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR) objective under both linear and general function approximations, enriched by human feedback. These new formulations provide a principled way to guarantee safety in each decision making step throughout the control process. Moreover, integrating human feedback into risk-sensitive RL framework bridges the gap between algorithmic decision-making and human participation, allowing us to also guarantee safety for human-in-the-loop systems. We propose provably sample-efficient algorithms for this Iterated CVaR RL and provide rigorous theoretical analysis. Furthermore, we establish a matching lower bound to corroborate the optimality of our algorithms in a linear context.

📄 PDF Abstract BibTeX arXiv:2307.02842

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingLEMMAreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst Path

2022-06-06 · Yihan Du, Siwei Wang, Longbo Huang

In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each step, and focuses on tightly controlling th…

Autonomous Drivingreinforcement-learningReinforcement Learning (RL)

Provably Efficient CVaR RL in Low-rank MDPs

2023-11-20 · Yulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho-fung Leung 외

We study risk-sensitive Reinforcement Learning (RL), where we aim to maximize the Conditional Value at Risk (CVaR) with a fixed risk tolerance $\tau$. Prior theoretical work studying risk-sensitive RL focuses on the tabu…

Reinforcement Learning (RL)Representation Learning

Online Risk-Averse Planning in POMDPs Using Iterated CVaR Value Function

2026-01-28 · Yaacov Pariente, Vadim Indelman arxiv

We study risk-sensitive planning under partial observability using the dynamic risk measure Iterated Conditional Value-at-Risk (ICVaR). A policy evaluation algorithm for ICVaR is developed with finite-time performance gu…

Near-Optimal Sample Complexity for Iterated CVaR Reinforcement Learning with a Generative Model

2025-03-11 · Zilong Deng, Simon Khan, Shaofeng Zou

In this work, we study the sample complexity problem of risk-sensitive Reinforcement Learning (RL) with a generative model, where we aim to maximize the Conditional Value at Risk (CVaR) with risk tolerance level $\tau$ a…

Reinforcement Learning (RL)

Policy Gradients for CVaR-Constrained MDPs

2014-05-12 · Prashanth L. A

We study a risk-constrained version of the stochastic shortest path (SSP) problem, where the risk measure considered is Conditional Value-at-Risk (CVaR). We propose two algorithms that obtain a locally risk-optimal polic…