paper-with-me

Papers

Successor Features Combine Elements of Model-Free and Model-based Reinforcement Learning

2019-01-31 · Lucas Lehnert, Michael L. Littman

A key question in reinforcement learning is how an intelligent agent can generalize knowledge across different inputs. By generalizing across different inputs, information learned for one input can be immediately reused for improving predictions for another input. Reusing information allows an agent to compute an optimal decision-making strategy using less data. State representation is a key element of the generalization process, compressing a high-dimensional input space into a low-dimensional latent state space. This article analyzes properties of different latent state spaces, leading to new connections between model-based and model-free reinforcement learning. Successor features, which predict frequencies of future observations, form a link between model-based and model-free learning: Learning to predict future expected reward outcomes, a key characteristic of model-based agents, is equivalent to learning successor features. Learning successor features is a form of temporal difference learning and is equivalent to learning to predict a single policy's utility, which is a characteristic of model-free agents. Drawing on the connection between model-based reinforcement learning and successor features, we demonstrate that representations that are predictive of future reward outcomes generalize across variations in both transitions and rewards. This result extends previous work on successor features, which is constrained to fixed transitions and assumes re-learning of the transferred state representation.

📄 PDF Abstract BibTeX arXiv:1901.11437

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingmodelModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

APS: Active Pretraining with Successor Features

2021-08-31 · Hao liu, Pieter Abbeel

We introduce a new unsupervised pretraining objective for reinforcement learning. During the unsupervised reward-free pretraining phase, the agent maximizes mutual information between tasks and states induced by the poli…

Unsupervised Reinforcement Learning

A neurally plausible model learns successor representations in partially observable environments

2019-06-22 · NeurIPS 2019 12 · Eszter Vertes, Maneesh Sahani

Animals need to devise strategies to maximize returns while interacting with their environment based on incoming noisy sensory observations. Task-relevant states, such as the agent's location within an environment or the…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Full-Gradient Successor Feature Representations

2026-04-01 · Ritish Shrirao, Aditya Priyadarshi, Raghuram Bharadwaj Diddigi arxiv

Successor Features (SF) combined with Generalized Policy Improvement (GPI) provide a robust framework for transfer learning in Reinforcement Learning (RL) by decoupling environment dynamics from reward functions. However…

Reinforcement LearningTransfer Learning

Successor Representation Active Inference

2022-07-20 · Beren Millidge, Christopher L Buckley

Recent work has uncovered close links between between classical reinforcement learning algorithms, Bayesian filtering, and Active Inference which lets us understand value functions in terms of Bayesian posteriors. An alt…

Reinforcement Learning (RL)

Deep Successor Reinforcement Learning

2016-06-08 · Tejas D. Kulkarni, Ardavan Saeedi, Simanta Gautam, Samuel J. Gershman

Learning robust value functions given raw observations and rewards is now possible with model-free and model-based deep reinforcement learning algorithms. There is a third alternative, called Successor Representations (S…

Deep Reinforcement LearningFPS GamesGame of Doomreinforcement-learning+2