paper-with-me

홈 › Papers

Latent Action Pretraining Through World Modeling

2025-09-22 · Bahey Tharwat, Yara Nasser, Ali Abouzeid, Ian Reid arxiv

Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $π_{0}$, were trained on large-scale, manually labeled action datasets collected through teleoperation. More recent approaches, including LAPA and villa-X, introduce latent action representations that enable unsupervised pretraining on unlabeled datasets by modeling abstract visual changes between frames. Although these methods have shown strong results, their large model sizes make deployment in real-world settings challenging. In this work, we propose LAWM, a model-agnostic framework to pretrain imitation learning models in a self-supervised way, by learning latent action representations from unlabeled video data through world modeling. These videos can be sourced from robot recordings or videos of humans performing actions with everyday objects. Our framework is able to transfer learned knowledge across tasks, environments, and embodiments. It outperforms models pretrained with ground-truth robot actions and other similar pretraining methods on the LIBERO benchmark and real-world setup, while being efficient and practical for real-world settings.

📄 PDF Abstract BibTeX arXiv:2509.18428

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

2026-02-23 · Manish Kumar Govind, Dominick Reilly, Pu Wang, Srijan Das arxiv

Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot action supervision. However, latent act…

Chain of World: World Model Thinking in Latent Motion

2026-03-03 · Fuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang 외 arxiv

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by pre…

Computational Efficiency

Motus: A Unified Latent Action World Model

2025-12-15 · Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 외 arxiv

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative ca…

Video Generation

A Generalization Theory for JEPA-Based World Models

2026-06-25 · Jingyi Cui, Qi Zhang, Hongwei Wen, Yisen Wang arxiv

Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input …

Graph Learning

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

2026-06-05 · Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung 외 arxiv

Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question thro…