paper-with-me

홈 › Papers

Information Maximizing Exploration with a Latent Dynamics Model

2018-04-04 · Trevor Barron, Oliver Obst, Heni Ben Amor

All reinforcement learning algorithms must handle the trade-off between exploration and exploitation. Many state-of-the-art deep reinforcement learning methods use noise in the action selection, such as Gaussian noise in policy gradient methods or $\epsilon$-greedy in Q-learning. While these methods are appealing due to their simplicity, they do not explore the state space in a methodical manner. We present an approach that uses a model to derive reward bonuses as a means of intrinsic motivation to improve model-free reinforcement learning. A key insight of our approach is that this dynamics model can be learned in the latent feature space of a value function, representing the dynamics of the agent and the environment. This method is both theoretically grounded and computationally advantageous, permitting the efficient use of Bayesian information-theoretic methods in high-dimensional state spaces. We evaluate our method on several continuous control tasks, focusing on improving exploration.

📄 PDF Abstract BibTeX arXiv:1804.01238

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous ControlDeep Reinforcement LearningmodelPolicy Gradient MethodsQ-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

State-Wise Safe Reinforcement Learning With Pixel Observations

2023-11-03 · Simon Sinong Zhan, YiXuan Wang, Qingyuan Wu, Ruochen Jiao 외

In the context of safe exploration, Reinforcement Learning (RL) has long grappled with the challenges of balancing the tradeoff between maximizing rewards and minimizing safety violations, particularly in complex environ…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Exploration+1

VIME: Variational Information Maximizing Exploration

2016-05-31 · NeurIPS 2016 12 · Rein Houthooft, Xi Chen, Yan Duan, John Schulman 외

Scalable and effective exploration remains a key challenge in reinforcement learning (RL). While there are methods with optimality guarantees in the setting of discrete state and action spaces, these methods cannot be ap…

continuous-controlContinuous ControlReinforcement LearningReinforcement Learning (RL)+1

Efficiently Learning Nonstationary Gaussian Processes for Real World Impact

2018-04-27 · Sahil Garg

Most real world phenomena such as sunlight distribution under a forest canopy, minerals concentration, stock valuation, exhibit nonstationary dynamics i.e. phenomenon variation changes depending on the locality. Nonstati…

Gaussian ProcessesInformativeness

Entropic Desired Dynamics for Intrinsic Control

2021-12-01 · NeurIPS 2021 12 · Steven Hansen, Guillaume Desjardins, Kate Baumli, David Warde-Farley 외

An agent might be said, informally, to have mastery of its environment when it has maximised the effective number of states it can reliably reach. In practice, this often means maximizing the number of latent codes that …

Montezuma's Revenge

Latent-Predictive Empowerment: Measuring Empowerment without a Simulator

2024-10-15 · Andrew Levy, Alessandro Allievi, George Konidaris

Empowerment has the potential to help agents learn large skillsets, but is not yet a scalable solution for training general-purpose agents. Recent empowerment methods learn diverse skillsets by maximizing the mutual info…