paper-with-me

홈 › Papers

Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning

2025-09-25 · Xiefeng Wu, Jing Zhao, Shu Zhang, Mingyu Hu arxiv

Online reinforcement learning in complex tasks is time-consuming, as massive interaction steps are needed to learn the optimal Q-function.Vision-language action (VLA) policies represent a promising direction for solving diverse tasks; however, their performance on low-level control remains limited, and effective deployment often requires task-specific expert demonstrations for fine-tuning. In this paper, we propose \textbf{VARL} (\textbf{V}LM as \textbf{A}ction advisor for online \textbf{R}einforcement \textbf{L}earning), a framework that leverages the domain knowledge of vision-language models (VLMs) to provide action suggestions for reinforcement learning agents. Unlike previous methods, VARL provides action suggestions rather than designing heuristic rewards, thereby guaranteeing unchanged optimality and convergence. The suggested actions increase sample diversity and ultimately improve sample efficiency, especially in sparse-reward tasks. To validate the effectiveness of VARL, we evaluate it across diverse environments and agent settings. Results show that VARL greatly improves sample efficiency without introducing significant computational overhead. These advantages make VARL a general framework for online reinforcement learning and make it feasible to directly apply reinforcement learning from scratch in real-world environments.

📄 PDF Abstract BibTeX arXiv:2509.21126

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Online Transfer Learning in Reinforcement Learning Domains

2015-07-02 · Yusen Zhan, Matthew E. Taylor

This paper proposes an online transfer framework to capture the interaction among agents and shows that current transfer learning in reinforcement learning is a special case of online transfer. Furthermore, this paper re…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Bridging the Imitation Gap by Adaptive Insubordination

2020-07-23 · NeurIPS 2021 12 · Luca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador 외

In practice, imitation learning is preferred over pure reinforcement learning whenever it is possible to design a teaching agent to provide expert supervision. However, we show that when the teaching agent makes decision…

Imitation LearningMemorizationreinforcement-learningReinforcement Learning+1

Multi-Agent Advisor Q-Learning

2021-10-26 · Sriram Ganapathi Subramanian, Matthew E. Taylor, Kate Larson, Mark Crowley

In the last decade, there have been significant advances in multi-agent reinforcement learning (MARL) but there are still numerous challenges, such as high sample complexity and slow convergence to stable policies, that …

Decision MakingMulti-agent Reinforcement LearningQ-Learningreinforcement-learning+1

Are Generative AI Agents Effective Personalized Financial Advisors?

2025-04-08 · Takehiro Takayanagi, Kiyoshi Izumi, Javier Sanz-Cruzado, Richard McCreadie 외

Large language model-based agents are becoming increasingly popular as a low-cost mechanism to provide personalized, conversational advice, and have demonstrated impressive capabilities in relatively simple scenarios, su…

Large Language Model

Improving interactive reinforcement learning: What makes a good teacher?

2019-04-15 · Francisco Cruz, Sven Magg, Yukie Nagai, Stefan Wermter

Interactive reinforcement learning has become an important apprenticeship approach to speed up convergence in classic reinforcement learning problems. In this regard, a variant of interactive reinforcement learning is po…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)