paper-with-me

홈 › Papers

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

2021-06-12 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. However, current state-of-the-art MIS methods rely on complex optimization tricks and succeed mostly on simple toy problems. We bridge the gap between MIS and deep reinforcement learning by observing that the density ratio can be computed from the successor representation of the target policy. The successor representation can be trained through deep reinforcement learning methodology and decouples the reward optimization from the dynamics of the environment, making the resulting algorithm stable and applicable to high-dimensional domains. We evaluate the empirical performance of our approach on a variety of challenging Atari and MuJoCo environments.

📄 PDF Abstract BibTeX arXiv:2106.06854

Code (1)

sfujim/SR-DICE 공식 구현 pytorch

Tasks

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Practical Marginalized Importance Sampling with the Successor Representation

2021-01-01 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Marginalized Operators for Off-policy Reinforcement Learning

2022-03-30 · Yunhao Tang, Mark Rowland, Rémi Munos, Michal Valko

In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as spe…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling

2019-06-08 · NeurIPS 2019 12 · Tengyang Xie, Yifei Ma, Yu-Xiang Wang

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Free Energy Evaluation Using Marginalized Annealed Importance Sampling

2022-04-08 · Muneki Yasuda, Chako Takahashi

The evaluation of the free energy of a stochastic model is considered a significant issue in various fields of physics and machine learning. However, the exact free energy evaluation is computationally infeasible because…

State2vec: Off-Policy Successor Features Approximators

2019-10-22 · Sephora Madjiheurem, Laura Toni

A major challenge in reinforcement learning (RL) is the design of agents that are able to generalize across tasks that share common dynamics. A viable solution is meta-reinforcement learning, which identifies common stru…

Meta Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)