paper-with-me

홈 › Papers

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

2026-05-25 · Perry Dong, Kuo-Han Hung, Tian Gao, Dorsa Sadigh, Chelsea Finn arxiv

The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse manipulation tasks, yet pretrained policies consistently fall short of the reliability required for real-world deployment. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches either train from scratch without fully leveraging pretrained priors, or fine-tune VLAs without achieving the sample efficiency and success rates that practical deployment demands. We present EXPO-FT, a system for stable, sample-efficient RL finetuning of pretrained VLA policies that closes this gap. Our system solves a suite of challenging manipulation tasks, including routing string lights and inserting the plug to light it up, striking a pool ball into a pocket, and inserting a flower into a wine bottle, each requiring combinations of high precision, dynamic actions, and robustness to varied initial states. Our system achieves perfect task performance (30/30 successes) across all evaluated tasks within an average of 19.1 minutes of online robot data, outperforming both prior RL-from-scratch and VLA finetuning approaches. We release an open-source codebase with the aim of facilitating broader adoption of RL finetuning of VLA models in robotics.

📄 PDF Abstract BibTeX arXiv:2605.25477

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

2026-03-21 · Zhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng 외 arxiv

Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety risks of real-world interaction. Trainin…

Reinforcement Learning

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

2026-05-26 · Heng Qu, Yike Liu, Renren Jin, Wenzong Zhang 외 arxiv

Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based agents in this domain. To facilitate rig…

Reinforcement Learning

VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness

2024-01-15 · Rongyu Zhang, Zefan Cai, Huanrui Yang, Zidong Liu 외

Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data points results in diminished training efficie…

Computational Efficiencyimage-classificationImage Classification

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

2025-05-26 · Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin 외

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language…

document understandingMultimodal ReasoningVisual Reasoning

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

2026-07-15 · Dwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim 외 hf

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwri…