paper-with-me

홈 › Papers

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

2025-10-20 · Haochen Su, Cristian Meo, Francesco Stella, Andrea Peirone, Kai Junge, Josie Hughes arxiv

Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a language guided generalized control framework for real robots. However, their deployment has been limited to conventional serial link manipulators. Coupled by their rigidity and unpredictability of learning based control, the ability to safely interact with the environment is missing yet critical. In this work, we present the deployment of a VLA model on a soft continuum manipulator to demonstrate autonomous safe human-robot interaction. We present a structured finetuning and deployment pipeline evaluating two state-of-the-art VLA models (OpenVLA-OFT and $π_0$) across representative manipulation tasks, and show while out-of-the-box policies fail due to embodiment mismatch, through targeted finetuning the soft robot performs equally to the rigid counterpart. Our findings highlight the necessity of finetuning for bridging embodiment gaps, and demonstrate that coupling VLA models with soft robots enables safe and flexible embodied AI in human-shared environments.

📄 PDF Abstract BibTeX arXiv:2510.17369

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

2025-11-04 · Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang 외 arxiv

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (i) prod…

RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents

2026-07-30 · Sihyung Yoon, Minjong Yoo, Sanghyun Ahn, Seojeong Choi 외 arxiv

Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical ga…

Scene Understanding

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

2025-10-28 · Wentao Tan, Bowen Wang, Heng Zhi, Chenyu Liu 외 arxiv

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical …

Instruction Following

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

2026-06-26 · Sijin Chen, Kaixuan Jiang, Haixin Shi, Yanhui Wang 외 arxiv

We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it one of the most promising resources for…

On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation

2026-07-11 · Finn Ferchau, Daniel Pommer, Cristian Axenie arxiv

Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade …