paper-with-me

홈 › Papers

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

2026-08-04 · Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You, Fugui Fan, Yuting Wu, Minghui Wu, Chenxu Zhao, JiaHong Ning, Peiguang Jing arxiv

World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.

📄 PDF Abstract BibTeX arXiv:2608.03701

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LILA: Language-Informed Latent Actions

2021-11-05 · Siddharth Karamcheti, Megha Srivastava, Percy Liang, Dorsa Sadigh

We introduce Language-Informed Latent Actions (LILA), a framework for learning natural language interfaces in the context of human-robot collaboration. LILA falls under the shared autonomy paradigm: in addition to provid…

Imitation Learning

Causal-JEPA: Learning World Models through Object-Level Latent Masking

2026-02-11 · Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun 외 arxiv

World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-depend…

Visual Question Answering

LiLa-Net: Lightweight Latent LiDAR Autoencoder for 3D Point Cloud Reconstruction

2025-10-02 · Mario Resino, Borja Pérez, Jaime Godoy, Abdulla Al-Kaff 외 arxiv

This work proposed a 3D autoencoder architecture, named LiLa-Net, which encodes efficient features from real traffic environments, employing only the LiDAR's point clouds. For this purpose, we have real semi-autonomous v…

Point Clouds

LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training

2025-09-25 · Abhishek Moturu, Muhammad Muzammil, Anna Goldenberg, Babak Taati arxiv

Training deep neural networks with noise and data heterogeneity is a major challenge. We introduce Lightweight Learnable Adaptive Weighting (LiLAW), a method that dynamically adjusts the loss weight of each training samp…

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

2026-04-29 · Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki, David Joseph Tan 외 arxiv

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of …

Video Object SegmentationSemantic Segmentation