paper-with-me

Papers

What Is a Good Caption? A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

2025-02-19 · Zhihang Liu, Chen-Wei Xie, Bin Wen, Feiwu Yu, Jixuan Chen, Boqiang Zhang, Nianzu Yang, Pandeng Li, Yinglu Li, Zuan Gao, Yun Zheng, Hongtao Xie

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions using F1-score. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides the first holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of capabilities.

📄 PDF Abstract BibTeX arXiv:2502.14914

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningKeyword Extraction

Similar Papers 제목 키워드 기반

What Makes for Good Image Captions?

2024-05-01 · Delong Chen, Samuel Cahyawijaya, Etsuko Ishii, Ho Shu Chan 외

This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode semantic units in images. Our framework posi…

HallucinationImage CaptioningRepresentation Learning

simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions

2018-08-27 · EMNLP 2018 10 · Fenglin Liu, Xuancheng Ren, Yuanxin Liu, Houfeng Wang 외

The encode-decoder framework has shown recent success in image captioning. Visual attention, which is good at detailedness, and semantic attention, which is good at comprehensiveness, have been separately proposed to gro…

DecoderImage Captioning

FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation

2021-06-19 · CVPR 2021 1 · Sijin Wang, Ziwei Yao, Ruiping Wang, Zhongqin Wu 외

Image caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. Wh…

Image Captioning

Explore and Tell: Embodied Visual Captioning in 3D Environments

2023-08-21 · ICCV 2023 1 · Anwen Hu, ShiZhe Chen, Liang Zhang, Qin Jin

While current visual captioning models have achieved impressive performance, they often assume that the image is well-captured and provides a complete view of the scene. In real-world scenarios, however, a single image m…

Image CaptioningNavigateScene Understanding

CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

2025-09-26 · Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 외 arxiv

Image captioning is a fundamental task that bridges the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Current state-of-the-art captioning models are typicall…

Reinforcement LearningImage Captioning