paper-with-me

홈 › Papers

VLRM: Vision-Language Models act as Reward Models for Image Captioning

2024-04-02 · Maksim Dzabraev, Alexander Kunitsyn, Andrei Ivaniuta

In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and BLIP2-ITM as reward models. The RL-tuned model is able to generate longer and more comprehensive descriptions. Our model reaches impressive 0.90 R@1 CLIP Recall score on MS-COCO Carpathy Test Split. Weights are available at https://huggingface.co/sashakunitsyn/vlrm-blip2-opt-2.7b.

📄 PDF Abstract BibTeX arXiv:2404.01911

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

2025-03-10 · Jiacheng Ruan, Wenzhen Yuan, Xian Gao, Ye Guo 외

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become …

Binary ClassificationHallucinationMathematical Reasoning

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

2025-02-19 · Chen-An Li, Tzu-Han Lin, Yun-Nung Chen, Hung-Yi Lee

Large vision-language models (LVLMs) perform outstandingly across various multimodal tasks. However, their ability to evaluate generated content remains limited, and training vision-language reward models (VLRMs) with pr…

Multi-Level Policy and Reward Reinforcement Learning for Image Captioning

2018-06-15 · IJCAI 2018 6 · An-An Liu1, Ning Xu1, Hanwang Zhang2, Weizhi Nie1 외

Image captioning is one of the most challenging hallmarks of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning …

Image CaptioningNatural Language Understandingreinforcement-learningReinforcement Learning+2

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information

2025-05-29 · Xu Chu, Xinrong Chen, Guanyu Wang, Zhijie Tan 외

Inference time scaling drives extended reasoning to enhance the performance of Vision-Language Models (VLMs), thus forming powerful Vision-Language Reasoning Models (VLRMs). However, long reasoning dilutes visual tokens,…

Hallucination

The Devil is in the EOS: Sequence Training for Detailed Image Captioning

2025-07-26 · Abdelrahman Mohamed, Yova Kementchedjhieva arxiv

Despite significant advances in vision-language models (VLMs), image captioning often suffers from a lack of detail, with base models producing short, generic captions. This limitation persists even though VLMs are equip…

Image Captioning