paper-with-me

홈 › Papers

RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes

2026-06-23 · Arsh Chawla, Rahul Shome arxiv

Object interaction tasks have been a focus of advances in imitation learning. End-to-end methods, dominated by diffusion and flow-based variants have shown leaps in performance while sacrificing interpretability. Object-centric and pose-informed variants have had a role in learning from demonstration in manipulation tasks. In this paper, we revisit a few modern imitation learning benchmarks for object interactions, with the aim of composing a framework that repurposes principled theories of manipulation, preserving both performance and interpretability. For image observations, lightweight training is proposed for model-free pose estimation of the target object, using self-supervision over the demonstration data available for imitation learning. This information is then used to inform a manipulation mode-aware retrieval of a demonstration, a mode-aware transformation, a replan step that connects to the retrieval point while preserving mode constraints, and finally rolling out the transformed demonstration. These compose four key steps of the proposed RE4 framework, evaluated over state-based and image-based benchmarks in Push-T and Robomimic. An adversarial benchmark that evaluates sparse data regions of image-based Push-T showcases the robustness, further bolstered by indications from low-data regime experiments. The current work shows promise in using simple interpretable building blocks to learn manipulation skills.

📄 PDF Abstract BibTeX arXiv:2606.24403

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

2026-04-09 · Jingjing Wang, Zhengdong Hong, Chong Bao, Yuke Zhu 외 arxiv

Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforcement learning, imitation learning, and vision-language-action-models (…

Zero-shot GeneralizationReinforcement Learning

Learning Spatial-Aware Manipulation Ordering

2025-10-29 · Yuxiang Yan, Zhiyuan Zhou, Xin Gao, Guanghao Li 외 arxiv

Manipulation in cluttered environments is challenging due to spatial dependencies among objects, where an improper manipulation order can cause collisions or blocked access. Existing approaches often overlook these spati…

Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation

2026-06-02 · Yilong Wang, Cheng Qian, Edward Johns arxiv

Deformable object manipulation (DOM) is challenging due to high-dimensional, partially observable states that evolve through long-horizon, topology-changing interactions with multiple valid manipulation modes. We introdu…

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy

2025-08-20 · Yijin Chen, Wenqiang Xu, Zhenjun Yu, Tutian Tang 외 arxiv

Dexterous in-hand manipulation is a long-standing challenge in robotics due to complex contact dynamics and partial observability. While humans synergize vision and touch for such tasks, robotic approaches often prioriti…

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine

2025-11-17 · Xincheng Shuai, Zhenyuan Qin, Henghui Ding, Dacheng Tao arxiv

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-awar…

3D ReconstructionImage Editing