paper-with-me

홈 › Papers

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

2026-02-12 · Jiangran Lyu, Kai Liu, Xuheng Zhang, Haoran Liao, Yusen Feng, Wenxuan Zhu, Tingrui Shen, Jiayi Chen, Jiazhao Zhang, Yifei Dong, Wenbo Cui, Senmao Qi, Shuo Wang, Yixin Zheng, Mi Yan, Xuesong Shi, Haoran Li, Dongbin Zhao, Ming-Yu Liu, Zhizheng Zhang, Li Yi, Yizhou Wang, He Wang arxiv

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing instantiations struggle to scale to foundation-level due to coarse data usage and fragmented datasets. We introduce LDA-1B, a robot foundation model that scales through universal embodied data ingestion by jointly learning dynamics, policy, and visual forecasting, assigning distinct roles to data of varying quality. To support this regime at scale, we assemble and standardize EI-30k, an embodied interaction dataset comprising over 30k hours of human and robot trajectories in a unified format. Scalable dynamics learning over such heterogeneous data is enabled by prediction in a structured DINO latent space, which avoids redundant pixel-space appearance modeling. Complementing this representation, LDA-1B employs a multi-modal diffusion transformer to handle asynchronous vision and action streams, enabling stable training at the 1B-parameter scale. Experiments in simulation and the real world show LDA-1B outperforms prior methods (e.g., $π_{0.5}$) by up to 21\%, 48\%, and 23\% on contact-rich, dexterous, and long-horizon tasks, respectively. Notably, LDA-1B enables data-efficient fine-tuning, gaining 10\% by leveraging 30\% low-quality trajectories typically harmful and discarded.

📄 PDF Abstract BibTeX arXiv:2602.12215

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

2026-02-01 · Shuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li 외 arxiv

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception …

Robot Manipulation

Universal Actions for Enhanced Embodied Foundation Models

2025-01-17 · CVPR 2025 1 · Jinliang Zheng, Jianxiong Li, Dongxiu Liu, Yinan Zheng 외

Training on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availabili…

$ω$-EVA: Envision, Verify, and Act with Latent Interactive World Models

2026-06-08 · Zhenguo Sun, Yu Sun, Hande Huang, Alois Knoll arxiv

Embodied policies typically map current observations directly to actions, leaving candidate-action consequences implicit. World models provide predictive supervision, representations, or external simulation, but rarely l…

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

2026-08-31 · Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun 외 hf

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Ego…

Consistent Attack: Universal Adversarial Perturbation on Embodied Vision Navigation

2022-06-12 · Chengyang Ying, You Qiaoben, Xinning Zhou, Hang Su 외

Embodied agents in vision navigation coupled with deep neural networks have attracted increasing attention. However, deep neural networks have been shown vulnerable to malicious adversarial noises, which may potentially …