paper-with-me

홈 › Papers

Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation

2026-05-01 · Andrzej Ruszczynski, Tiangang Zhang arxiv

For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem.

📄 PDF Abstract BibTeX arXiv:2605.00654

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Regret Bounds for Risk-sensitive Reinforcement Learning with Lipschitz Dynamic Risk Measures

2023-06-04 · Hao Liang, Zhi-Quan Luo

We study finite episodic Markov decision processes incorporating dynamic risk measures to capture risk sensitivity. To this end, we present two model-based algorithms applied to \emph{Lipschitz} dynamic risk measures, a …

reinforcement-learningSensitivity

A policy gradient approach for optimization of smooth risk measures

2022-02-22 · Nithia Vijayan, Prashanth L. A

We propose policy gradient algorithms for solving a risk-sensitive reinforcement learning (RL) problem in on-policy as well as off-policy settings. We consider episodic Markov decision processes, and model the risk using…

reinforcement-learningReinforcement Learning (RL)

Risk-sensitive reinforcement learning using expectiles, shortfall risk and optimized certainty equivalent risk

2026-02-10 · Sumedh Gupte, Shrey Rakeshkumar Patel, Soumen Pachal, Prashanth L. A. 외 arxiv

We propose risk-sensitive reinforcement learning algorithms catering to three families of risk measures, namely expectiles, utility-based shortfall risk and optimized certainty equivalent risk. For each risk measure, in …

Reinforcement Learning

Risk-Averse Reinforcement Learning via Dynamic Time-Consistent Risk Measures

2023-01-14 · Xian Yu, Siqian Shen

Traditional reinforcement learning (RL) aims to maximize the expected total reward, while the risk of uncertain outcomes needs to be controlled to ensure reliable performance in a risk-averse setting. In this paper, we c…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Variance-Based Risk Estimations in Markov Processes via Transformation with State Lumping

2019-07-09 · Shuai Ma, Jia Yuan Yu

Variance plays a crucial role in risk-sensitive reinforcement learning, and most risk measures can be analyzed via variance. In this paper, we consider two law-invariant risks as examples: mean-variance risk and exponent…

Reinforcement LearningReinforcement Learning (RL)