Variational Intrinsic Control Revisited
In this paper, we revisit variational intrinsic control (VIC), an unsupervised reinforcement learning method for finding the largest set of intrinsic options available to an agent. In the original work by Gregor et al. (2016), two VIC algorithms were proposed: one that represents the options explicitly, and the other that does it implicitly. We show that the intrinsic reward used in the latter is subject to bias in stochastic environments, causing convergence to suboptimal solutions. To correct this behavior and achieve the maximal empowerment, we propose two methods respectively based on the transitional probability model and Gaussian mixture model. We substantiate our claims through rigorous mathematical derivations and experimental analyses.
Code (0)
등록된 구현이 없습니다.
Tasks
reinforcement-learningReinforcement Learning (RL)Unsupervised Reinforcement LearningSimilar Papers 제목 키워드 기반
Intrinsic Control of Variational Beliefs in Dynamic Partially-Observed Visual Environments
Humans and animals explore their environment and acquire useful skills even in the absence of clear goals, exhibiting intrinsic motivation. The study of intrinsic motivation in artificial agents is concerned with the fol…
Deep Reinforcement LearningSystematic and multifactor risk models revisited
Systematic and multifactor risk models are revisited via methods which were already successfully developed in signal processing and in automatic control. The results, which bypass the usual criticisms on those risk model…
Universal AI maximizes Variational Empowerment
This paper presents a theoretical framework unifying AIXI -- a model of universal AI -- with variational empowerment as an intrinsic drive for exploration. We build on the existing framework of Self-AIXI -- a universal l…
VASE: Variational Assorted Surprise Exploration for Reinforcement Learning
Exploration in environments with continuous control and sparse rewards remains a key challenge in reinforcement learning (RL). Recently, surprise has been used as an intrinsic reward that encourages systematic and effici…
continuous-controlContinuous ControlEfficient Explorationreinforcement-learning+3Provably Safe Stein Variational Clarity-Aware Informative Planning
Autonomous robots are increasingly deployed for information-gathering tasks in environments that vary across space and time. Planning informative and safe trajectories in such settings is challenging because information …
Bayesian Inference