RIDM: Reinforced Inverse Dynamics Modeling for Learning from a Single Observed Demonstration
Augmenting reinforcement learning with imitation learning is often hailed as a method by which to improve upon learning from scratch. However, most existing methods for integrating these two techniques are subject to several strong assumptions---chief among them that information about demonstrator actions is available. In this paper, we investigate the extent to which this assumption is necessary by introducing and evaluating reinforced inverse dynamics modeling (RIDM), a novel paradigm for combining imitation from observation (IfO) and reinforcement learning with no dependence on demonstrator action information. Moreover, RIDM requires only a single demonstration trajectory and is able to operate directly on raw (unaugmented) state features. We find experimentally that RIDM performs favorably compared to a baseline approach for several tasks in simulation as well as for tasks on a real UR5 robot arm. Experiment videos can be found at https://sites.google.com/view/ridm-reinforced-inverse-dynami.
Code (0)
등록된 구현이 없습니다.
Tasks
Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
HybridMIM: A Hybrid Masked Image Modeling Framework for 3D Medical Image Segmentation
Masked image modeling (MIM) with transformer backbones has recently been exploited as a powerful self-supervised pre-training technique. The existing MIM methods adopt the strategy to mask random patches of the image and…
Contrastive LearningImage SegmentationMedical Image SegmentationSelf-Supervised Learning+1HybridMamba: A Dual-domain Mamba for 3D Medical Image Segmentation
In the domain of 3D biomedical image segmentation, Mamba exhibits the superior performance for it addresses the limitations in modeling long-range dependencies inherent to CNNs and mitigates the abundant computational ov…
Medical Image SegmentationHybridMimic: Hybrid RL-Centroidal Control for Humanoid Motion Mimicking
Motion mimicking, i.e., encouraging the control policy to mimic human motion, facilitates the learning of complex tasks via reinforcement learning (RL) for humanoid robots. Although standard RL frameworks demonstrate imp…
Reinforcement LearningCrash Time Matters: HybridMamba for Fine-Grained Temporal Localization in Traffic Surveillance Footage
Traffic crash detection in long-form surveillance videos is critical for emergency response and infrastructure planning but remains difficult due to the brief and rare nature of crash events. We introduce HybridMamba, a …
Temporal LocalizationSuperpixelGridCut, SuperpixelGridMean and SuperpixelGridMix Data Augmentation
A novel approach of data augmentation based on irregular superpixel decomposition is proposed. This approach called SuperpixelGridMasks permits to extend original image datasets that are required by training stages of ma…
Data Augmentationimage-classificationImage Classification