paper-with-me

홈 › Papers

Improving Vision-Language-Action Model with Online Reinforcement Learning

2025-01-28 · Yanjiang Guo, Jianke Zhang, Xiaoyu Chen, Xiang Ji, Yen-Jen Wang, Yucheng Hu, Jianyu Chen

Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models are powerful, how to improve these large models during interaction with environments remains an open question. In this paper, we explore how to further improve these VLA models via Reinforcement Learning (RL), a commonly used fine-tuning technique for large models. However, we find that directly applying online RL to large VLA models presents significant challenges, including training instability that severely impacts the performance of large models, and computing burdens that exceed the capabilities of most local machines. To address these challenges, we propose iRe-VLA framework, which iterates between Reinforcement Learning and Supervised Learning to effectively improve VLA models, leveraging the exploratory benefits of RL while maintaining the stability of supervised learning. Experiments in two simulated benchmarks and a real-world manipulation suite validate the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2501.16664

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Vision-Language-Action

Similar Papers 제목 키워드 기반

Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning

2025-09-25 · Xiefeng Wu, Jing Zhao, Shu Zhang, Mingyu Hu arxiv

Online reinforcement learning in complex tasks is time-consuming, as massive interaction steps are needed to learn the optimal Q-function.Vision-language action (VLA) policies represent a promising direction for solving …

Reinforcement Learning

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning

2025-12-15 · Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui 외 arxiv

Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement L…

Reinforcement LearningAutonomous Driving

GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

2026-08-27 · Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang 외 arxiv

Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…

Robot ManipulationVisual Grounding

NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

2026-03-10 · Ziyue Zhu, Shangyang Wu, Shuai Zhao, Zhiqiu Zhao 외 arxiv

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning re…

Reinforcement Learning

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking

2026-05-11 · Andrew Choi, Wei Xu arxiv

Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key challenge, however, is learning an accurate critic in large state--action s…

Reinforcement Learning