paper-with-me

Papers

Risk-Averse Bayes-Adaptive Reinforcement Learning

2021-02-10 · NeurIPS 2021 12 · Marc Rigter, Bruno Lacerda, Nick Hawes

In this work, we address risk-averse Bayes-adaptive reinforcement learning. We pose the problem of optimising the conditional value at risk (CVaR) of the total return in Bayes-adaptive Markov decision processes (MDPs). We show that a policy optimising CVaR in this setting is risk-averse to both the parametric uncertainty due to the prior distribution over MDPs, and the internal uncertainty due to the inherent stochasticity of MDPs. We reformulate the problem as a two-player stochastic game and propose an approximate algorithm based on Monte Carlo tree search and Bayesian optimisation. Our experiments demonstrate that our approach significantly outperforms baseline approaches for this problem.

📄 PDF Abstract BibTeX arXiv:2102.05762

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Optimisationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Online Bayesian Risk-Averse Reinforcement Learning

2025-09-17 · Yuhao Wang, Enlu Zhou arxiv

In this paper, we study the Bayesian risk-averse formulation in reinforcement learning (RL). To address the epistemic uncertainty due to a lack of data, we adopt the Bayesian Risk Markov Decision Process (BRMDP) to accou…

Reinforcement Learning

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

2026-07-29 · Mingxuan Che, Tsung-Yuan Tseng, Theresa Eimer, Marius Lindauer 외 arxiv

Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (…

Reinforcement Learning

Bayesian Risk-Averse Q-Learning with Streaming Observations

2023-05-18 · NeurIPS 2023 11

We consider a robust reinforcement learning problem, where a learning agent learns from a simulated training environment. To account for the model mis-specification between this training environment and the real environm…

Q-Learning

Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning

2020-04-22 · Shangtong Zhang, Bo Liu, Shimon Whiteson

We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variable. MVPI enjoys great flexibility in tha…

MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Risk-Tendency: Nano Drone Navigation in Cluttered Environments with Distributional Reinforcement Learning

2022-03-28 · Cheng Liu, Erik-Jan van Kampen, Guido C. H. E. de Croon

Enabling the capability of assessing risk and making risk-aware decisions is essential to applying reinforcement learning to safety-critical robots like drones. In this paper, we investigate a specific case where a nano …

Distributional Reinforcement LearningDrone navigationNavigatereinforcement-learning+2