paper-with-me

홈 › Papers

Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning

2018-10-15 · NeurIPS 2019 12 · David Janz, Jiri Hron, Przemysław Mazur, Katja Hofmann, José Miguel Hernández-Lobato, Sebastian Tschiatschek

Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network function approximation do not possess the properties which make PSRL effective, and provably fail in sparse reward problems. Moreover, we find that propagation of uncertainty, a property of PSRL previously thought important for exploration, does not preclude this failure. We use these insights to design Successor Uncertainties (SU), a cheap and easy to implement RVF algorithm that retains key properties of PSRL. SU is highly effective on hard tabular exploration benchmarks. Furthermore, on the Atari 2600 domain, it surpasses human performance on 38 of 49 games tested (achieving a median human normalised score of 2.09), and outperforms its closest RVF competitor, Bootstrapped DQN, on 36 of those.

📄 PDF Abstract BibTeX arXiv:1810.06530

Code (2)

DavidJanz/successor_uncertainties_atari pytorch
DavidJanz/successor_uncertainties_tabular pytorch

Tasks

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…
DQN A DQN, or Deep Q-Network, approximates a state-value function in a Q-Learning framework with a neural network. In the Atari…

Similar Papers 제목 키워드 기반

Temporal Difference Uncertainties as a Signal for Exploration

2020-10-05 · Sebastian Flennerhag, Jane X. Wang, Pablo Sprechmann, Francesco Visin 외

An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in tabular settings. However, in non-tabula…

AKF-SR: Adaptive Kalman Filtering-based Successor Representation

2022-03-31 · Parvin Malekzadeh, Mohammad Salimibeni, Ming Hou, Arash Mohammadi 외

Recent studies in neuroscience suggest that Successor Representation (SR)-based models provide adaptation to changes in the goal locations or reward function faster than model-free algorithms, together with lower computa…

Active LearningDecision Making

Probabilistic Successor Representations with Kalman Temporal Differences

2019-10-06 · Jesse P. Geerts, Kimberly L. Stachenfeld, Neil Burgess

The effectiveness of Reinforcement Learning (RL) depends on an animal's ability to assign credit for rewards to the appropriate preceding stimuli. One aspect of understanding the neural underpinnings of this process invo…

Reinforcement LearningReinforcement Learning (RL)

Exploration by Learning Diverse Skills through Successor State Measures

2024-06-14 · Paul-Antoine Le Tolguenec, Yann Besse, Florent Teichteil-Konigsbuch, Dennis G. Wilson 외

The ability to perform different skills can encourage agents to explore. In this work, we aim to construct a set of diverse skills which uniformly cover the state space. We propose a formalization of this search for dive…

Efficient Exploration

Multi-Agent Reinforcement Learning via Adaptive Kalman Temporal Difference and Successor Representation

2021-12-30 · Mohammad Salimibeni, Arash Mohammadi, Parvin Malekzadeh, Konstantinos N. Plataniotis

Distributed Multi-Agent Reinforcement Learning (MARL) algorithms has attracted a surge of interest lately mainly due to the recent advancements of Deep Neural Networks (DNNs). Conventional Model-Based (MB) or Model-Free …

Multi-agent Reinforcement LearningOpenAI Gymreinforcement-learningReinforcement Learning (RL)