paper-with-me

홈 › Papers

Practical Marginalized Importance Sampling with the Successor Representation

2021-01-01 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. However, current state-of-the-art MIS methods rely on complex optimization tricks and succeed mostly on simple toy problems. We bridge the gap between MIS and deep reinforcement learning by observing that the density ratio can be computed from the successor representation of the target policy. The successor representation can be trained through deep reinforcement learning methodology and decouples the reward optimization from the dynamics of the environment, making the resulting algorithm stable and applicable to high-dimensional domains. We evaluate the empirical performance of our approach on a variety of challenging Atari and MuJoCo environments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

2021-06-12 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Marginalized Operators for Off-policy Reinforcement Learning

2022-03-30 · Yunhao Tang, Mark Rowland, Rémi Munos, Michal Valko

In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as spe…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Free Energy Evaluation Using Marginalized Annealed Importance Sampling

2022-04-08 · Muneki Yasuda, Chako Takahashi

The evaluation of the free energy of a stochastic model is considered a significant issue in various fields of physics and machine learning. However, the exact free energy evaluation is computationally infeasible because…

Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling

2019-06-08 · NeurIPS 2019 12 · Tengyang Xie, Yifei Ma, Yu-Xiang Wang

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Importance Weighted Actor-Critic for Optimal Conservative Offline Reinforcement Learning

2023-01-30 · NeurIPS 2023 11 · Hanlin Zhu, Paria Rashidinejad, Jiantao Jiao

We propose A-Crab (Actor-Critic Regularized by Average Bellman error), a new practical algorithm for offline reinforcement learning (RL) in complex environments with insufficient data coverage. Our algorithm combines the…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)