paper-with-me

Papers

MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

2025-10-14 · Zhenxin Lei, Zhangwei Gao, Changyao Tian, Erfei Cui, Guanzhou Chen, Danni Yang, Yuchen Duan, Zhaokai Wang, Wenhao Li, Weiyun Wang, Xiangyu Zhao, Jiayi Ji, Yu Qiao, Wenhai Wang, Gen Luo arxiv

Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains. In this task, current open-source models present a large performance gap with commercial ones, which limits various applications such as data synthesis. To bridge the gap, this paper proposes CapFlow, a novel multi-agent collaboration workflow. CapFlow demonstrates for the first time that, by capitalizing on open-source models, it is possible to achieve caption quality on par with GPT-4.1 in various domains with an 89.5% reduction in costs. By leveraging CapFlow as the data synthesizer, we produce high-quality visual captions from image and video domains at scale, and obtain a generalist visual captioner via fine-tuning, namely MetaCaptioner. Through extensive experiments, we show that MetaCaptioner not only achieves comparable captioning capabilities with commercial models but also reaches top-tier multimodal performance in the open-source community. We hope CapFlow and MetaCaptioner can benefit future multimodal research by providing a strong and cost-effective visual captioning solution.

📄 PDF Abstract BibTeX arXiv:2510.12126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

2026-06-02 · Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong 외 arxiv

Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, and summarization. However, their practical use is rapidly expanding …

Visual Question AnsweringReferring ExpressionObject Detection

nocaps: novel object captioning at scale

2018-12-20 · ICCV 2019 10 · Harsh Agrawal, Karan Desai, YuFei Wang, Xinlei Chen 외

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, …

Image CaptioningObjectobject-detectionObject Detection

Vision-Language Models as a Source of Rewards

2023-12-14 · Kate Baumli, Satinder Baveja, Feryal Behbahani, Harris Chan 외

Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been…

reinforcement-learningReinforcement Learning

OneDiff: A Generalist Model for Image Difference Captioning

2024-07-08 · Erdong Hu, Longteng Guo, Tongtian Yue, Zijia Zhao 외

In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on specialist models, which restrict their applicab…

Language ModellingmodelMulti-Task Learning

Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation

2024-04-30 · CVPR 2024 1 · Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman, Tsung-Yi Lin 외

Existing automatic captioning methods for visual content face challenges such as lack of detail, content hallucination, and poor instruction following. In this work, we propose VisualFactChecker (VFC), a flexible trainin…

Caption GenerationHallucinationImage to textInstruction Following+6