paper-with-me

홈 › Papers

Maximum a Posteriori Policy Optimisation

2018-06-14 · ICLR 2018 1 · Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, Martin Riedmiller

We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are competitive with the state-of-the-art in deep reinforcement learning. In particular, for continuous control, our method outperforms existing methods with respect to sample efficiency, premature convergence and robustness to hyperparameter settings while achieving similar or better final performance.

📄 PDF Abstract BibTeX arXiv:1806.06920

Code (3)

MotorCityCobra/C_plusplus_mpo
acyclics/MPO pytorch
deepmind/rgb_stacking

Tasks

continuous-controlContinuous ControlDeep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Probabilistic inverse reinforcement learning in unknown environments

2014-08-09 · Aristide Tossou, Christos Dimitrakakis

We consider the problem of learning by demonstration from agents acting in unknown stochastic Markov environments or games. Our aim is to estimate agent preferences in order to construct improved policies for the same ta…

Bayesian Inferencereinforcement-learningReinforcement LearningReinforcement Learning (RL)

Probabilistic inverse reinforcement learning in unknown environments

2013-07-14 · Aristide C. Y. Tossou, Christos Dimitrakakis

We consider the problem of learning by demonstration from agents acting in unknown stochastic Markov environments or games. Our aim is to estimate agent preferences in order to construct improved policies for the same ta…

Bayesian Inferencereinforcement-learningReinforcement LearningReinforcement Learning (RL)

Relative Entropy Regularized Policy Iteration

2018-12-05 · Abbas Abdolmaleki, Jost Tobias Springenberg, Jonas Degrave, Steven Bohez 외

We present an off-policy actor-critic algorithm for Reinforcement Learning (RL) that combines ideas from gradient-free optimization via stochastic search with learned action-value function. The result is a simple procedu…

continuous-controlContinuous ControlOpenAI GymReinforcement Learning+1

Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation

2024-07-25 · Jean Seong Bjorn Choe, Jong-Kook Kim

Entropy Regularisation is a widely adopted technique that enhances policy optimisation performance and stability. A notable form of entropy regularisation is augmenting the objective with an entropy term, thereby simulta…

MuJoCo

An end-to-end data-driven optimisation framework for constrained trajectories

2020-11-24 · Florent Dewez, Benjamin Guedj, Arthur Talpaert, Vincent Vandewalle

Many real-world problems require to optimise trajectories under constraints. Classical approaches are based on optimal control methods but require an exact knowledge of the underlying dynamics, which could be challenging…