paper-with-me

Papers

Online Bayesian Risk-Averse Reinforcement Learning

2025-09-17 · Yuhao Wang, Enlu Zhou arxiv

In this paper, we study the Bayesian risk-averse formulation in reinforcement learning (RL). To address the epistemic uncertainty due to a lack of data, we adopt the Bayesian Risk Markov Decision Process (BRMDP) to account for the parameter uncertainty of the unknown underlying model. We derive the asymptotic normality that characterizes the difference between the Bayesian risk value function and the original value function under the true unknown distribution. The results indicate that the Bayesian risk-averse approach tends to pessimistically underestimate the original value function. This discrepancy increases with stronger risk aversion and decreases as more data become available. We then utilize this adaptive property in the setting of online RL as well as online contextual multi-arm bandits (CMAB), a special case of online RL. We provide two procedures using posterior sampling for both the general RL problem and the CMAB problem. We establish a sub-linear regret bound, with the regret defined as the conventional regret for both the RL and CMAB settings. Additionally, we establish a sub-linear regret bound for the CMAB setting with the regret defined as the Bayesian risk regret. Finally, we conduct numerical experiments to demonstrate the effectiveness of the proposed algorithm in addressing epistemic uncertainty and verifying the theoretical properties.

📄 PDF Abstract BibTeX arXiv:2509.14077

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Risk-Averse Bayes-Adaptive Reinforcement Learning

2021-02-10 · NeurIPS 2021 12 · Marc Rigter, Bruno Lacerda, Nick Hawes

In this work, we address risk-averse Bayes-adaptive reinforcement learning. We pose the problem of optimising the conditional value at risk (CVaR) of the total return in Bayes-adaptive Markov decision processes (MDPs). W…

Bayesian Optimisationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

2026-07-29 · Mingxuan Che, Tsung-Yuan Tseng, Theresa Eimer, Marius Lindauer 외 arxiv

Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (…

Reinforcement Learning

Bayesian Risk-Averse Q-Learning with Streaming Observations

2023-05-18 · NeurIPS 2023 11

We consider a robust reinforcement learning problem, where a learning agent learns from a simulated training environment. To account for the model mis-specification between this training environment and the real environm…

Q-Learning

Risk-Averse Finetuning of Large Language Models

2025-01-12 · Sapana Chaudhary, Ujwal Dinesha, Dileep Kalathil, Srinivas Shakkottai

We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning t…

Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning

2020-04-22 · Shangtong Zhang, Bo Liu, Shimon Whiteson

We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variable. MVPI enjoys great flexibility in tha…

MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)