paper-with-me

홈 › Papers

Unified Motion-Action Modeling for Heterogeneous Robot Learning

2026-06-15 · Yunhao Cao, Shitong Liu, Chao Feng, Meryl Zhang, Xuanchen Lu, Andrew Owens, Kuan Fang arxiv

We present Unified Motion-Action (UMA) Model, an approach that uses 3D object motion trajectories as a shared interface to bridge visuomotor control and dynamics modeling. UMA treats object motion and robot actions as co-evolving variables under a masked generative objective, in which the mask pattern determines both the supervision regime during pretraining and the inference mode at deployment. Using hindsight-relabeled motion contexts and a contrastive objective that disentangles task intent from scene geometry, UMA enables multi-task pretraining across heterogeneous data sources without requiring manually annotated task instructions. At deployment, the same pretrained parameters support motion-conditioned visuomotor control, motion-based dynamics modeling, and task adaptation from few-shot demonstrations. Pretrained on a mixture of robot demonstrations, human videos, and simulated data, UMA consistently outperforms state-of-the-art baselines specialized for each inference mode.

📄 PDF Abstract BibTeX arXiv:2606.16917

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

2026-08-26 · Xiaomi Embodied Intelligence Team, University of Macau, :, Shaoqing Xu 외 arxiv

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces…

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

2025-11-04 · Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang 외 arxiv

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (i) prod…

Motus: A Unified Latent Action World Model

2025-12-15 · Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 외 arxiv

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative ca…

Video Generation

Hydra-0: Action Flow for Generalist World Modeling and Control

2026-08-18 · Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang 외 arxiv

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action con…

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

2026-07-02 · Guoyang Xia, Fengfa Li, Hongjin Ji, Lei Ren 외 arxiv

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in archite…