paper-with-me

Papers

Non-stationary Risk-sensitive Reinforcement Learning: Near-optimal Dynamic Regret, Adaptive Detection, and Separation Design

2022-11-19 · Yuhao Ding, Ming Jin, Javad Lavaei

We study risk-sensitive reinforcement learning (RL) based on an entropic risk measure in episodic non-stationary Markov decision processes (MDPs). Both the reward functions and the state transition kernels are unknown and allowed to vary arbitrarily over time with a budget on their cumulative variations. When this variation budget is known a prior, we propose two restart-based algorithms, namely Restart-RSMB and Restart-RSQ, and establish their dynamic regrets. Based on these results, we further present a meta-algorithm that does not require any prior knowledge of the variation budget and can adaptively detect the non-stationarity on the exponential value functions. A dynamic regret lower bound is then established for non-stationary risk-sensitive RL to certify the near-optimality of the proposed algorithms. Our results also show that the risk control and the handling of the non-stationarity can be separately designed in the algorithm if the variation budget is known a prior, while the non-stationary detection mechanism in the adaptive algorithm depends on the risk parameter. This work offers the first non-asymptotic theoretical analyses for the non-stationary risk-sensitive RL in the literature.

📄 PDF Abstract BibTeX arXiv:2211.10815

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Optimal Transport-Assisted Risk-Sensitive Q-Learning

2024-06-17 · Zahra Shahrooei, Ali Baheri

The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance without considering risk or safety. In contrast, safe reinforcement learning aims to mitigate or avoid…

Decision MakingQ-Learningreinforcement-learningReinforcement Learning+1

Policy Newton methods for Distortion Riskmetrics

2025-08-10 · Soumen Pachal, Mizhaan Prajit Maniyar, Prashanth L. A arxiv

We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskmetric (DRM) of the discounted reward in …

Reinforcement Learning

On the Convergence and Optimality of Policy Gradient for Markov Coherent Risk

2021-03-04 · Audrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar Azizzadenesheli

In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes conditional value-at-risk (CVaR). Because op…

Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret

2020-06-22 · NeurIPS 2020 12 · Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 외

We study risk-sensitive reinforcement learning in episodic Markov decision processes with unknown transition kernels, where the goal is to optimize the total reward under the risk measure of exponential utility. We propo…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Risk-sensitive reinforcement learning using expectiles, shortfall risk and optimized certainty equivalent risk

2026-02-10 · Sumedh Gupte, Shrey Rakeshkumar Patel, Soumen Pachal, Prashanth L. A. 외 arxiv

We propose risk-sensitive reinforcement learning algorithms catering to three families of risk measures, namely expectiles, utility-based shortfall risk and optimized certainty equivalent risk. For each risk measure, in …

Reinforcement Learning