paper-with-me

Papers

On Pathologies in KL-Regularized Reinforcement Learning from Expert Demonstrations

2022-12-28 · NeurIPS 2021 12 · Tim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal, Yee Whye Teh

KL-regularized reinforcement learning from expert demonstrations has proved successful in improving the sample efficiency of deep reinforcement learning algorithms, allowing them to be applied to challenging physical real-world tasks. However, we show that KL-regularized reinforcement learning with behavioral reference policies derived from expert demonstrations can suffer from pathological training dynamics that can lead to slow, unstable, and suboptimal online learning. We show empirically that the pathology occurs for commonly chosen behavioral policy classes and demonstrate its impact on sample efficiency and online policy performance. Finally, we show that the pathology can be remedied by non-parametric behavioral reference policies and that this allows KL-regularized reinforcement learning to significantly outperform state-of-the-art approaches on a variety of challenging locomotion and dexterous hand manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2212.13936

Code (1)

conglu1997/nppac 공식 구현 pytorch

Tasks

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Difference of Convex Functions Programming Applied to Control with Expert Data

2016-06-03 · Bilal Piot, Matthieu Geist, Olivier Pietquin

This paper reports applications of Difference of Convex functions (DC) programming to Learning from Demonstrations (LfD) and Reinforcement Learning (RL) with expert data. This is made possible because the norm of the Opt…

General Classificationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Demonstration-Regularized RL

2023-10-26 · Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 외

Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this extra information reduces RL's sample co…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Continuous Mean-Zero Disagreement-Regularized Imitation Learning (CMZ-DRIL)

2024-03-02 · Noah Ford, Ryan W. Gardner, Austin Juhl, Nathan Larson

Machine-learning paradigms such as imitation learning and reinforcement learning can generate highly performant agents in a variety of complex environments. However, commonly used methods require large quantities of data…

Imitation LearningMuJoCoreinforcement-learningReinforcement Learning

Convergence of a model-free entropy-regularized inverse reinforcement learning algorithm

2024-03-25 · Titouan Renard, Andreas Schlaginhaufen, Tingting Ni, Maryam Kamgarpour

Given a dataset of expert demonstrations, inverse reinforcement learning (IRL) aims to recover a reward for which the expert is optimal. This work proposes a model-free algorithm to solve entropy-regularized IRL problem.…

Wasserstein Adversarial Imitation Learning

2019-06-19 · Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche 외

Imitation Learning describes the problem of recovering an expert policy from demonstrations. While inverse reinforcement learning approaches are known to be very sample-efficient in terms of expert demonstrations, they u…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)