paper-with-me

홈 › Papers

Risk Averse Value Expansion for Sample Efficient and Robust Policy Learning

2019-09-25 · Bo Zhou, Fan Wang, Hongsheng Zeng, Hao Tian

Model-based Reinforcement Learning(RL) has shown great advantage in sample-efficiency, but suffers from poor asymptotic performance and high inference cost. A promising direction is to combine model-based reinforcement learning with model-free reinforcement learning, such as model-based value expansion(MVE). However, the previous methods do not take into account the stochastic character of the environment, thus still suffers from higher function approximation errors. As a result, they tend to fall behind the best model-free algorithms in some challenging scenarios. We propose a novel Hybrid-RL method, which is developed from MVE, namely the Risk Averse Value Expansion(RAVE). In the proposed method, we use an ensemble of probabilistic models for environment modeling to generate imaginative rollouts, based on which we further introduce the aversion of risks by seeking the lower confidence bound of the estimation. Experiments on different environments including MuJoCo and robo-school show that RAVE yields state-of-the-art performance. Also we found that it greatly prevented some catastrophic consequences such as falling down and thus reduced the variance of the rewards.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Model-based Reinforcement LearningMuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization

2024-03-17 · Yudong Luo, Yangchen Pan, Han Wang, Philip Torr 외

Reinforcement learning algorithms utilizing policy gradients (PG) to optimize Conditional Value at Risk (CVaR) face significant challenges with sample inefficiency, hindering their practical applications. This inefficien…

MuJoCo

Risk-averse Total-reward MDPs with ERM and EVaR

2024-08-30 · Xihong Su, Julien Grand-Clément, Marek Petrik

Optimizing risk-averse objectives in discounted MDPs is challenging because most models do not admit direct dynamic programming equations and require complex history-dependent policies. In this paper, we show that the ri…

Actor-Critic Algorithm for Dynamic Expectile and CVaR

2026-05-08 · Yudong Luo, Erick Delage arxiv

Optimizing dynamic risk with stochastic policies is challenging in both policy updates and value learning. The former typically requires transition perturbation, while the latter may rely on model-based approaches. To ad…

Efficient and Robust Reinforcement Learning with Uncertainty-based Value Expansion

2019-12-10 · Bo Zhou, Hongsheng Zeng, Fan Wang, Yunxiang Li 외

By integrating dynamics models into model-free reinforcement learning (RL) methods, model-based value expansion (MVE) algorithms have shown a significant advantage in sample efficiency as well as value estimation. Howeve…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Efficient Risk-Averse Reinforcement Learning

2022-05-10 · Ido Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie Mannor

In risk-averse reinforcement learning (RL), the goal is to optimize some risk measure of the returns. A risk measure often focuses on the worst returns out of the agent's experience. As a result, standard methods for ris…

Autonomous Drivingreinforcement-learningReinforcement LearningReinforcement Learning (RL)