A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation
Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. However, current state-of-the-art MIS methods rely on complex optimization tricks and succeed mostly on simple toy problems. We bridge the gap between MIS and deep reinforcement learning by observing that the density ratio can be computed from the successor representation of the target policy. The successor representation can be trained through deep reinforcement learning methodology and decouples the reward optimization from the dynamics of the environment, making the resulting algorithm stable and applicable to high-dimensional domains. We evaluate the empirical performance of our approach on a variety of challenging Atari and MuJoCo environments.
Code (1)
Tasks
Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Practical Marginalized Importance Sampling with the Successor Representation
Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…
Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2Marginalized Operators for Off-policy Reinforcement Learning
In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as spe…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling
Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Free Energy Evaluation Using Marginalized Annealed Importance Sampling
The evaluation of the free energy of a stochastic model is considered a significant issue in various fields of physics and machine learning. However, the exact free energy evaluation is computationally infeasible because…
State2vec: Off-Policy Successor Features Approximators
A major challenge in reinforcement learning (RL) is the design of agents that are able to generalize across tasks that share common dynamics. A viable solution is meta-reinforcement learning, which identifies common stru…
Meta Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)