paper-with-me

Papers

X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

2025-10-11 · Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, Ya-Qin Zhang, Jiangmiao Pang, Jingjing Liu, Tai Wang, Xianyuan Zhan arxiv

Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/

📄 PDF Abstract BibTeX arXiv:2510.10274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DexFormer: Cross-Embodied Dexterous Manipulation via History-Conditioned Transformer

2026-02-09 · Ke Zhang, Lixin Xu, Chengyi Song, Junzhe Xu 외 arxiv

Dexterous manipulation remains one of the most challenging problems in robotics, requiring coherent control of high-DoF hands and arms under complex, contact-rich dynamics. A major barrier is embodiment variability: diff…

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

2026-05-12 · Yiren Song, Xiyao Deng, Pei Yang, Yihan Wang 외 arxiv

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge …

Video Generation

Embedding Morphology into Transformers for Cross-Robot Policy Learning

2026-02-26 · Kei Suzuki, Jing Liu, Ye Wang, Chiori Hori 외 arxiv

Cross-robot policy learning -- training a single policy to perform well across multiple embodiments -- remains a central challenge in robot learning. Transformer-based policies, such as vision-language-action (VLA) model…

Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation

2024-08-21 · Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari 외

Modern machine learning systems rely on large datasets to attain broad generalization, and this often poses a challenge in robot learning, where each robotic platform and task might have only a small dataset. By training…

Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

2026-07-28 · Huy Ha, C. Karen Liu, Shuran Song arxiv

An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot des…

Robot Manipulation