paper-with-me

홈 › Papers

Bayes-Adaptive Deep Model-Based Policy Optimisation

2020-10-29 · Tai Hoang, Ngo Anh Vien

We introduce a Bayesian (deep) model-based reinforcement learning method (RoMBRL) that can capture model uncertainty to achieve sample-efficient policy optimisation. We propose to formulate the model-based policy optimisation problem as a Bayes-adaptive Markov decision process (BAMDP). RoMBRL maintains model uncertainty via belief distributions through a deep Bayesian neural network whose samples are generated via stochastic gradient Hamiltonian Monte Carlo. Uncertainty is propagated through simulations controlled by sampled models and history-based policies. As beliefs are encoded in visited histories, we propose a history-based policy network that can be end-to-end trained to generalise across history space and will be trained using recurrent Trust-Region Policy Optimisation. We show that RoMBRL outperforms existing approaches on many challenging control benchmark tasks in terms of sample complexity and task performance. The source code of this paper is also publicly available on https://github.com/thobotics/RoMBRL.

📄 PDF Abstract BibTeX arXiv:2010.15948

Code (1)

thobotics/RoMBRL 공식 구현 tf

Tasks

modelModel-based Reinforcement Learning

Similar Papers 제목 키워드 기반

Risk-Averse Bayes-Adaptive Reinforcement Learning

2021-02-10 · NeurIPS 2021 12 · Marc Rigter, Bruno Lacerda, Nick Hawes

In this work, we address risk-averse Bayes-adaptive reinforcement learning. We pose the problem of optimising the conditional value at risk (CVaR) of the total return in Bayes-adaptive Markov decision processes (MDPs). W…

Bayesian Optimisationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Alternating Optimisation and Quadrature for Robust Control

2016-05-24 · Supratik Paul, Konstantinos Chatzilygeroudis, Kamil Ciosek, Jean-Baptiste Mouret 외

Bayesian optimisation has been successfully applied to a variety of reinforcement learning problems. However, the traditional approach for learning optimal policies in simulators does not utilise the opportunity to impro…

Bayesian OptimisationReinforcement Learning

Efficient Adaptive Data Acquisition via Pretrained Belief Representations

2026-06-23 · Daolang Huang, Zhuoyue Huang, Conor Hassan, Luigi Acerbi 외 arxiv

Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learnin…

Representation LearningActive Learning

MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning

2023-02-22 · Adam X. Yang, Laurence Aitchison, Henry B. Moss

In Bayesian optimisation, we often seek to minimise the black-box objective functions that arise in real-world physical systems. A primary contributor to the cost of evaluating such black-box objective functions is often…

Bayesian OptimisationMeta-Learning

Contextual Causal Bayesian Optimisation

2023-01-29 · Vahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar

Causal Bayesian optimisation (CaBO) combines causality with Bayesian optimisation (BO) and shows that there are situations where the optimal reward is not achievable if causal knowledge is ignored. While CaBO exploits ca…

Bayesian OptimisationMulti-Armed Bandits