paper-with-me

Papers

Learning Transferable Dynamics Priors from Action to World Modeling

2026-06-28 · Ze Huang, Jiahui Zhang, Hairuo Liu, Chenxi Zhang, Ran Cheng, Li Zhang arxiv

We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model captures reusable interaction dynamics beyond appearance-level video generation. Concretely, we pretrain a multi-view interactive base diffusion world model, A2World, on large-scale robot manipulation data with real action annotations. We validate the learned dynamics priors from two complementary perspectives. First, we adapt A2World into a task- or scene-specialized real-world simulator, A2World-sim, whose long-horizon rollouts support simulator-based policy evaluation and scalable what-if analysis by replacing real-robot rollouts with world model rollouts. Second, starting from the same pretrained weights, we adapt A2World into a video-action joint prediction model, A2World-policy, that predicts actions under visual and instruction conditioning. Experiments across simulation benchmarks and real-robot settings demonstrate that action-conditioned world model pretraining yields transferable dynamics priors that benefit both simulator-centric and policy-centric robot learning.

📄 PDF Abstract BibTeX arXiv:2606.29501

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationVideo Generation

Similar Papers 제목 키워드 기반

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

2026-02-10 · Zhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo 외 arxiv

Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the fi…

Video Generation

SCAR: Self-Supervised Continuous Action Representation Learning

2026-05-13 · Hongjia Liu, Fan Feng, Minghao Fu, Xinyue Wang 외 arxiv

Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across emb…

Representation Learning

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

2026-07-02 · Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji 외 arxiv

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this b…

Advancing Opinion Dynamics Modeling with Neural Diffusion-Convection-Reaction Equation

2026-02-05 · Chenghua Gong, Yihang Jiang, Hao Li, Rui Sun 외 arxiv

Advanced opinion dynamics modeling is vital for deciphering social behavior, emphasizing its role in mitigating polarization and securing cyberspace. To synergize mechanistic interpretability with data-driven flexibility…

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

2026-08-13 · Bing Zhan, Shuyao Shang, Jiahao Gu, Shuo Lu 외 arxiv

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action…

Autonomous Driving