paper-with-me

홈 › Papers

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

2026-03-27 · Wenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng, Xin Jin, Li Zhang arxiv

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video data. To address these issues, we propose DeFI, a novel framework that Decouples visual Forward and Inverse dynamics pretraining to exploit respective data sources, wherein video generation and action prediction are disentangled. We introduce the General Forward Dynamics Model (GFDM), pretrained on diverse human and robot videos for future prediction, and the General Inverse Dynamics Model (GIDM), trained via self-supervised learning to infer latent actions from unlabeled video transitions. These models are then integrated into a unified architecture for end-to-end finetuning on downstream tasks. In this manner, GFDM and GIDM first shine separately and then cooperate for mutual benefit. Extensive experiments on CALVIN ABC-D and SimplerEnv demonstrate state-of-the-art performance, with DeFI achieving an average task length of 4.51 for CALVIN, 51.2% success rate on SimplerEnv-Fractal benchmark and 81.3% success rate in real-world deployment, significantly outperforming prior methods.

📄 PDF Abstract BibTeX arXiv:2604.16391

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningVideo Generation

Similar Papers 제목 키워드 기반

Distributed Differentiable Dynamic Game for Multi-robot Coordination

2022-07-18 · Yizhi Zhou, Wanxin Jin, Xuan Wang

This paper develops a Distributed Differentiable Dynamic Game (D3G) framework, which can efficiently solve the forward and inverse problems in multi-robot coordination. We formulate multi-robot coordination as a dynamic …

A Flow Matching Framework for Soft-Robot Inverse Dynamics

2026-04-03 · Hang Yang, Fangju Yang, Yangming Zhang, Ibrahim Alsarraj 외 arxiv

Learning the inverse dynamics of soft continuum robots remains challenging due to high-dimensional nonlinearities and complex actuation coupling. Conventional feedback-based controllers often suffer from control chatteri…

Learning to Poke by Poking: Experiential Learning of Intuitive Physics

2016-06-23 · NeurIPS 2016 12 · Pulkit Agrawal, Ashvin Nair, Pieter Abbeel, Jitendra Malik 외

We investigate an experiential learning paradigm for acquiring an internal model of intuitive physics. Our model is evaluated on a real-world robotic manipulation task that requires displacing objects to target locations…

Decision Making

MoRight: Motion Control Done Right

2026-04-08 · Shaowei Liu, Xuanchi Ren, Tianchang Shen, Huan Ling 외 arxiv

Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to sep…

Action-Conditional Recurrent Kalman Networks For Forward and Inverse Dynamics Learning

2020-10-20 · Vaisakh Shaj, Philipp Becker, Dieter Buchler, Harit Pandya 외

Estimating accurate forward and inverse dynamics models is a crucial component of model-based control for sophisticated robots such as robots driven by hydraulics, artificial muscles, or robots dealing with different con…

FrictionState Estimation