paper-with-me

홈 › Papers

Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution Matching

2023-03-05 · Lantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger, Stefano Ermon

Offline imitation learning (IL) promises the ability to learn performant policies from pre-collected demonstrations without interactions with the environment. However, imitating behaviors fully offline typically requires numerous expert data. To tackle this issue, we study the setting where we have limited expert data and supplementary suboptimal data. In this case, a well-known issue is the distribution shift between the learned policy and the behavior policy that collects the offline data. Prior works mitigate this issue by regularizing the KL divergence between the stationary state-action distributions of the learned policy and the behavior policy. We argue that such constraints based on exact distribution matching can be overly conservative and hamper policy learning, especially when the imperfect offline data is highly suboptimal. To resolve this issue, we present RelaxDICE, which employs an asymmetrically-relaxed f-divergence for explicit support regularization. Specifically, instead of driving the learned policy to exactly match the behavior policy, we impose little penalty whenever the density ratio between their stationary state-action distributions is upper bounded by a constant. Note that such formulation leads to a nested min-max optimization problem, which causes instability in practice. RelaxDICE addresses this challenge by supporting a closed-form solution for the inner maximization problem. Extensive empirical study shows that our method significantly outperforms the best prior offline IL method in six standard continuous control environments with over 30% performance gain on average, across 22 settings where the imperfect dataset is highly suboptimal.

📄 PDF Abstract BibTeX arXiv:2303.02569

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous ControlImitation Learning

Similar Papers 제목 키워드 기반

Discriminator-Guided Model-Based Offline Imitation Learning

2022-07-01 · Wenjia Zhang, Haoran Xu, Haoyi Niu, Peng Cheng 외

Offline imitation learning (IL) is a powerful method to solve decision-making problems from expert demonstrations without reward labels. Existing offline IL methods suffer from severe performance degeneration under limit…

Decision MakingImitation Learningmodel

Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning

2025-10-22 · Kevin Huang, Rosario Scalise, Cleah Winston, Ayush Agrawal 외 arxiv

Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptabi…

Reinforcement LearningOffline RL

Discriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations

2022-07-20 · Haoran Xu, Xianyuan Zhan, Honglei Yin, Huiling Qin

We study the problem of offline Imitation Learning (IL) where an agent aims to learn an optimal expert behavior policy without additional online environment interactions. Instead, the agent is provided with a supplementa…

Imitation LearningOffline RLReinforcement Learning (RL)

Imitation from Observations with Trajectory-Level Generative Embeddings

2026-01-01 · Yongtao Qu, Shangzhe Li, Weitong Zhang arxiv

We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert behavior. Many existing distribution-matchin…

Language-Critique Imitation Learning from Suboptimal Demonstrations

2026-07-01 · Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang, Ping-Chun Hsieh 외 arxiv

Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inher…

Reinforcement LearningContinuous Control