paper-with-me

홈 › Papers

XPACE: Joint World and Action Modeling from Heterogeneous Experience

2026-09-15 · Jiacheng Wei, Jerry Bai, Xiaoyu Yue, Zidong Wang, Xiaoyang Guo, Cheng Chen, Fanqi Pu, Fan Wu, Zhixu Yue, Yizhuo Li, Feng Qiu, Bo Liu, Yuying Ge, Hui Zhou, Chenyi Chen, Yixiao Ge arxiv

A general-purpose robot needs to draw on diverse experience, choose actions, and anticipate how those actions will change the world. We introduce XPACE, a unified embodied world model that serves as both a world action model, jointly predicting executable robot actions and future video, and a world simulator, predicting the visual consequences of prescribed actions. Our key insight is that video prediction can both connect heterogeneous experience to action learning and generate new experience for policy improvement. With a shared video backbone between the policy and simulator, we use action-unlabeled video to learn visual dynamics and action-labeled human and robot demonstrations to jointly learn video and action prediction. Building on this architecture, a coarse-to-fine training curriculum progressively emphasizes robot control while retaining human experience, allowing the policy to learn behaviors beyond those covered by robot demonstrations. Beyond learning from recorded experience, XPACE uses its simulator to create additional recovery supervision for the policy. Specifically, we adapt the simulator to its own generated context, synthesize deviation-recovery trajectories around expert demonstrations, and fine-tune the policy on filtered recovery examples. Experiments on XPENG's IRON humanoid robot show that heterogeneous training improves robustness and enables transfer of human-observed skills to tasks absent from robot demonstrations, while recovery data generated by the model's own simulator further improves real-world task completion. Together, these results demonstrate how joint world and action modeling connects learning from heterogeneous experience with simulation-driven policy self-improvement.

📄 PDF Abstract BibTeX arXiv:2609.17372

Code (0)

등록된 구현이 없습니다.

Tasks

Video Prediction

Similar Papers 제목 키워드 기반

Riemann-1.0: An Embodied World Action Model for Physical AI

2026-08-27 · Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu 외 arxiv

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unif…

Robot Manipulation

ProgD: Progressive Multi-scale Decoding with Dynamic Graphs for Joint Multi-agent Motion Forecasting

2025-09-11 · Xing Gao, Zherui Huang, Weiyao Lin, Xiao Sun arxiv

Accurate motion prediction of surrounding agents is crucial for the safe planning of autonomous vehicles. Recent advancements have extended prediction techniques from individual agents to joint predictions of multiple in…

Autonomous VehiclesMotion Forecasting

Motubrain: An Advanced World Action Model for Robot Control

2026-04-30 · MotuBrain Team, Chendong Xiang, Fan Bao, Haitian Liu 외 arxiv

Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present Motubrain, a unified World Action Model that jointly models video and action under a Uni…

Video Generation

All-in-One: Heterogeneous Interaction Modeling for Cold-Start Rating Prediction

2024-03-26 · Shuheng Fang, Kangfei Zhao, Yu Rong, ZHIXUN LI 외

Cold-start rating prediction is a fundamental problem in recommender systems that has been extensively studied. Many methods have been proposed that exploit explicit relations among existing data, such as collaborative f…

AllCollaborative FilteringRecommendation Systems

Robust Learning on Heterogeneous Graphs with Heterophily: A Graph Structure Learning Approach

2026-04-30 · Yihan Zhang, Ercan E. Kuruoglu arxiv

Heterogeneous graphs with heterophily have emerged as a powerful abstraction for modeling complex real-world systems, where nodes of different types and labels interact in diverse and often non-homophilous ways. Despite …

Graph structure learningRepresentation LearningGraph Learning