Reward-free World Models for Online Imitation Learning
Imitation learning (IL) enables agents to acquire skills directly from expert demonstrations, providing a compelling alternative to reinforcement learning. However, prior online IL approaches struggle with complex tasks characterized by high-dimensional inputs and complex dynamics. In this work, we propose a novel approach to online imitation learning that leverages reward-free world models. Our method learns environmental dynamics entirely in latent spaces without reconstruction, enabling efficient and accurate modeling. We adopt the inverse soft-Q learning objective, reformulating the optimization process in the Q-policy space to mitigate the instability associated with traditional optimization in the reward-policy space. By employing a learned latent dynamics model and planning for control, our approach consistently achieves stable, expert-level performance in tasks with high-dimensional observation or action spaces and intricate dynamics. We evaluate our method on a diverse set of benchmarks, including DMControl, MyoSuite, and ManiSkill2, demonstrating superior empirical performance compared to existing approaches.
Code (1)
Tasks
Imitation LearningQ-LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance
Online Imitation Learning struggles with the gap between extensive online exploration space and limited expert trajectories, hindering efficient exploration due to inaccurate reward estimation. Inspired by the findings f…
Efficient ExplorationImitation LearningOptical Flow EstimationCoupled Distributional Random Expert Distillation for World Model Online Imitation Learning
Imitation Learning (IL) has achieved remarkable success across various domains, including robotics, autonomous driving, and healthcare, by enabling agents to learn complex behaviors from expert demonstrations. However, e…
Autonomous DrivingDensity EstimationImitation LearningReward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization
Sparse rewards pose a central challenge in reinforcement learning, since agents receive no informative signal until they reach their goal. Intrinsic-reward methods address this issue by optimizing non-stationary objectiv…
Reinforcement LearningReward-Free Continual Adaptation for Resilient Space Robots
Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation,…
Reinforcement LearningContinual LearningDITTO: Offline Imitation Learning with World Models
We propose DITTO, an offline imitation learning algorithm which uses world models and on-policy reinforcement learning to addresses the problem of covariate shift, without access to an oracle or any additional online int…
Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)