paper-with-me

Papers

Rethinking ValueDice: Does It Really Improve Performance?

2022-02-05 · Ziniu Li, Tian Xu, Yang Yu, Zhi-Quan Luo

Since the introduction of GAIL, adversarial imitation learning (AIL) methods attract lots of research interests. Among these methods, ValueDice has achieved significant improvements: it beats the classical approach Behavioral Cloning (BC) under the offline setting, and it requires fewer interactions than GAIL under the online setting. Are these improvements benefited from more advanced algorithm designs? We answer this question by the following conclusions. First, we show that ValueDice could reduce to BC under the offline setting. Second, we verify that overfitting exists and regularization matters in the low-data regime. Specifically, we demonstrate that with weight decay, BC also nearly matches the expert performance as ValueDice does. The first two claims explain the superior offline performance of ValueDice. Third, we establish that ValueDice does not work when the expert trajectory is subsampled. Instead, the mentioned success of ValueDice holds when the expert trajectory is complete, in which ValueDice is closely related to BC that performs well as mentioned. Finally, we discuss the implications of our research for imitation learning studies beyond ValueDice.

📄 PDF Abstract BibTeX arXiv:2202.02468

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Methods 이 논문이 사용한 방법론

GAIL Generative Adversarial Imitation Learning presents a new general framework for directly extracting a policy from data, as if it were obtained by reinforcement learning…

Similar Papers 제목 키워드 기반

Rethinking ValueDice: Does It Really Improve Performance?

2022-01-17 · ICLR Track Blog 2022 5 · Anonymous

Since the introduction of GAIL, adversarial imitation learning (AIL) methods attract lots of research interests. Among these methods, ValueDice has achieved significant improvements: it beats the classical approach Behav…

Imitation Learning

SoftDICE for Imitation Learning: Rethinking Off-policy Distribution Matching

2021-06-06 · Mingfei Sun, Anuj Mahajan, Katja Hofmann, Shimon Whiteson

We present SoftDICE, which achieves state-of-the-art performance for imitation learning. SoftDICE fixes several key problems in ValueDICE, an off-policy distribution matching approach for sample-efficient imitation learn…

Imitation LearningMuJoCo

Non-Adversarial Imitation Learning and its Connections to Adversarial Methods

2020-08-08 · Oleg Arenz, Gerhard Neumann

Many modern methods for imitation learning and inverse reinforcement learning, such as GAIL or AIRL, are based on an adversarial formulation. These methods apply GANs to match the expert's distribution over states and ac…

Imitation Learning

Adversarial Imitation Learning via Boosting

2024-04-12 · Jonathan D. Chang, Dhruv Sreenivas, Yingbing Huang, Kianté Brantley 외

Adversarial imitation learning (AIL) has stood out as a dominant framework across various imitation learning (IL) applications, with Discriminator Actor Critic (DAC) (Kostrikov et al.,, 2019) demonstrating the effectiven…

Imitation Learning

MT Quality Estimation for Computer-assisted Translation: Does it Really Help?

2015-07-01 · IJCNLP 2015 7 · Marco Turchi, Matteo Negri, Marcello Federico
Machine TranslationTranslation