Deep Interactive Bayesian Reinforcement Learning via Meta-Learning
Agents that interact with other agents often do not know a priori what the other agents' strategies are, but have to maximise their own online return while interacting with and learning about others. The optimal adaptive behaviour under uncertainty over the other agents' strategies w.r.t. some prior can in principle be computed using the Interactive Bayesian Reinforcement Learning framework. Unfortunately, doing so is intractable in most settings, and existing approximation methods are restricted to small tasks. To overcome this, we propose to meta-learn approximate belief inference and Bayes-optimal behaviour for a given prior. To model beliefs over other agents, we combine sequential and hierarchical Variational Auto-Encoders, and meta-train this inference model alongside the policy. We show empirically that our approach outperforms existing methods that use a model-free approach, sample from the approximate posterior, maintain memory-free models of others, or do not fully utilise the known structure of the environment.
Code (0)
등록된 구현이 없습니다.
Tasks
Meta-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Reinforcement Learning Enhanced PicHunter for Interactive Search
With the tremendous increase in video data size, search performance could be impacted significantly. Specifically, in an interactive system, a real-time system allows a user to browse, search and refine a query. Without …
Bayesian Inferencereinforcement-learningReinforcement LearningBayesian Model-Agnostic Meta-Learning
Learning to infer Bayesian posterior from a few-shot dataset is an important step towards robust meta-learning due to the model uncertainty inherent in the problem. In this paper, we propose a novel Bayesian model-agnost…
Active Learningimage-classificationImage ClassificationMeta-Learning+3Bayesian Meta-reinforcement Learning for Traffic Signal Control
In recent years, there has been increasing amount of interest around meta reinforcement learning methods for traffic signal control, which have achieved better performance compared with traditional control methods. Howev…
Continual LearningMeta-LearningMeta Reinforcement Learningreinforcement-learning+3Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
Meta-reinforcement learning trains a single reinforcement learning agent on a distribution of tasks to quickly generalize to new tasks outside of the training set at test time. From a Bayesian perspective, one can interp…
Meta Reinforcement Learningreinforcement-learningReinforcement LearningVariational InferenceInteractive Text Ranking with Bayesian Optimisation: A Case Study on Community QA and Summarisation
For many NLP applications, such as question answering and summarisation, the goal is to select the best solution from a large space of candidates to meet a particular user's needs. To address the lack of user-specific tr…
Bayesian OptimisationCommunity Question AnsweringQuestion AnsweringReinforcement Learning+1