paper-with-me

홈 › Papers

RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

2024-02-06 · YuFei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian, Erdem Biyik, David Held, Zackory Erickson

Reward engineering has long been a challenge in Reinforcement Learning (RL) research, as it often requires extensive human effort and iterative processes of trial-and-error to design effective reward functions. In this paper, we propose RL-VLM-F, a method that automatically generates reward functions for agents to learn new tasks, using only a text description of the task goal and the agent's visual observations, by leveraging feedbacks from vision language foundation models (VLMs). The key to our approach is to query these models to give preferences over pairs of the agent's image observations based on the text description of the task goal, and then learn a reward function from the preference labels, rather than directly prompting these models to output a raw reward score, which can be noisy and inconsistent. We demonstrate that RL-VLM-F successfully produces effective rewards and policies across various domains - including classic control, as well as manipulation of rigid, articulated, and deformable objects - without the need for human supervision, outperforming prior methods that use large pretrained models for reward generation under the same assumptions. Videos can be found on our project website: https://rlvlmf2024.github.io/

📄 PDF Abstract BibTeX arXiv:2402.03681

Code (1)

yufeiwang63/rl-vlm-f pytorch

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers

2023-12-14 · Taewook Nam, Juyong Lee, Jesse Zhang, Sung Ju Hwang 외

We propose a framework that leverages foundation models as teachers, guiding a reinforcement learning agent to acquire semantically meaningful behavior without human feedback. In our framework, the agent receives task in…

Language ModelingLanguage Modellingreinforcement-learningReinforcement Learning+1

Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models

2025-06-15 · Tung Minh Luu, Younghwan Lee, Donghoon Lee, Sunho Kim 외

Designing effective reward functions remains a fundamental challenge in reinforcement learning (RL), as it often requires extensive human effort and domain expertise. While RL from human feedback has been successful in a…

Reinforcement Learning (RL)

PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models

2025-09-19 · Ruiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao 외 arxiv

Preference-based reinforcement learning (PbRL) has emerged as a promising paradigm for teaching robots complex behaviors without reward engineering. However, its effectiveness is often limited by two critical challenges:…

Reinforcement Learning

CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks

2025-09-17 · Seoyeon Choi, Kanghyun Ryu, Jonghoon Ock, Negar Mehr arxiv

Multi-Agent Reinforcement Learning (MARL) provides a powerful framework for learning coordination in multi-agent systems. However, applying MARL to robotics remains challenging due to their high-dimensional continuous jo…

Multi-agent Reinforcement Learning

Reinforcement Learning from Denoising Feedback

2026-05-25 · Qi He, Huan Chen, Ya Guo, Huijia Zhu 외 arxiv

Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (DLMs). We introduce Reinforcement Learning from Denoising Feedback (RLDF), a novel tr…

Computational EfficiencyReinforcement Learning