paper-with-me

Papers

Augmented Bayesian Policy Search

2024-07-05 · Mahdi Kallel, Debabrota Basu, Riad Akrour, Carlo D'Eramo

Deterministic policies are often preferred over stochastic ones when implemented on physical systems. They can prevent erratic and harmful behaviors while being easier to implement and interpret. However, in practice, exploration is largely performed by stochastic policies. First-order Bayesian Optimization (BO) methods offer a principled way of performing exploration using deterministic policies. This is done through a learned probabilistic model of the objective function and its gradient. Nonetheless, such approaches treat policy search as a black-box problem, and thus, neglect the reinforcement learning nature of the problem. In this work, we leverage the performance difference lemma to introduce a novel mean function for the probabilistic model. This results in augmenting BO methods with the action-value function. Hence, we call our method Augmented Bayesian Search~(ABS). Interestingly, this new mean function enhances the posterior gradient with the deterministic policy gradient, effectively bridging the gap between BO and policy gradient methods. The resulting algorithm combines the convenience of the direct policy search with the scalability of reinforcement learning. We validate ABS on high-dimensional locomotion problems and demonstrate competitive performance compared to existing direct policy search schemes.

📄 PDF Abstract BibTeX arXiv:2407.04864

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian OptimizationLEMMAPolicy Gradient Methodsreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

Adversarial Augmentation Policy Search for Domain and Cross-Lingual Generalization in Reading Comprehension

2020-04-13 · Findings of the Association for Computational Linguistics 2020 · Adyasha Maharana, Mohit Bansal

Reading comprehension models often overfit to nuances of training datasets and fail at adversarial evaluation. Training with adversarially augmented dataset improves robustness against those adversarial attacks but hurts…

Data AugmentationReading Comprehension

The dynamic impact of monetary policy on regional housing prices in the US: Evidence based on factor-augmented vector autoregressions

2018-02-16

In this study interest centers on regional differences in the response of housing prices to monetary policy shocks in the US. We address this issue by analyzing monthly home price data for metropolitan regions using a fa…

Time-Varying Identification of Monetary Policy Shocks

2023-11-10 · Annika Camehl, Tomasz Woźniak

We propose a new Bayesian heteroskedastic Markov-switching structural vector autoregression with data-driven time-varying identification. The model selects alternative exclusion restrictions over time and, as a condition…

Stochastic Path Planning in Correlated Obstacle Fields

2025-09-23 · Li Zhou, Elvan Ceyhan arxiv

We introduce the Stochastic Correlated Obstacle Scene (SCOS) problem, a navigation setting with spatially correlated obstacles of uncertain blockage status, realistically constrained sensors that provide noisy readings a…

Reinforcement Learning

Bayesian Neural Networks for Macroeconomic Analysis

2022-11-09 · Niko Hauzenberger, Florian Huber, Karin Klieber, Massimiliano Marcellino

Macroeconomic data is characterized by a limited number of observations (small T), many time series (big K) but also by featuring temporal dependence. Neural networks, by contrast, are designed for datasets with millions…

Time SeriesTime Series Analysis