VLRM: Vision-Language Models act as Reward Models for Image Captioning
In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and BLIP2-ITM as reward models. The RL-tuned model is able to generate longer and more comprehensive descriptions. Our model reaches impressive 0.90 R@1 CLIP Recall score on MS-COCO Carpathy Test Split. Weights are available at https://huggingface.co/sashakunitsyn/vlrm-blip2-opt-2.7b.
Code (0)
등록된 구현이 없습니다.
Tasks
Image Captioningreinforcement-learningReinforcement LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become …
Binary ClassificationHallucinationMathematical ReasoningTransferring Textual Preferences to Vision-Language Understanding through Model Merging
Large vision-language models (LVLMs) perform outstandingly across various multimodal tasks. However, their ability to evaluate generated content remains limited, and training vision-language reward models (VLRMs) with pr…
Multi-Level Policy and Reward Reinforcement Learning for Image Captioning
Image captioning is one of the most challenging hallmarks of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning …
Image CaptioningNatural Language Understandingreinforcement-learningReinforcement Learning+2Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
Inference time scaling drives extended reasoning to enhance the performance of Vision-Language Models (VLMs), thus forming powerful Vision-Language Reasoning Models (VLRMs). However, long reasoning dilutes visual tokens,…
HallucinationThe Devil is in the EOS: Sequence Training for Detailed Image Captioning
Despite significant advances in vision-language models (VLMs), image captioning often suffers from a lack of detail, with base models producing short, generic captions. This limitation persists even though VLMs are equip…
Image Captioning