paper-with-me

홈 › Papers

Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

2025-03-18 · Sara Sarto, Marcella Cornia, Rita Cucchiara

The evaluation of machine-generated image captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of advancements in image captioning evaluation, analyzing the evolution, strengths, and limitations of existing metrics. We assess these metrics across multiple dimensions, including correlation with human judgment, ranking accuracy, and sensitivity to hallucinations. Additionally, we explore the challenges posed by the longer and more detailed captions generated by MLLMs and examine the adaptability of current metrics to these stylistic variations. Our analysis highlights some limitations of standard evaluation approaches and suggests promising directions for future research in image captioning assessment.

📄 PDF Abstract BibTeX arXiv:2503.14604

Code (1)

aimagelab/awesome-captioning-evaluation 공식 구현 pytorch

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

2024-12-04 · Davide Bucciarelli, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 외

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Lan…

Image CaptioningImage DescriptionPrompt Learning

Multimodal Claim Extraction for Fact-Checking

2026-02-01 · Joycelyn Teo, Rui Cao, Zhenyun Deng, Zifeng Ding 외 arxiv

Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with …

Visual Question AnsweringImage Captioning

ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs

2026-04-04 · Zitong Xu, Huiyu Duan, Shengyao Qin, Guangyu Yang 외 arxiv

Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in …

Zero-shot GeneralizationImage Captioning

GUI Action Narrator: Where and When Did That Action Take Place?

2024-06-19 · Qinchen Wu, Difei Gao, Kevin Qinghong Lin, Zhuoyu Wu 외

The advent of Multimodal LLMs has significantly enhanced image OCR recognition capabilities, making GUI automation a viable reality for increasing efficiency in digital tasks. One fundamental aspect of developing a GUI a…

Optical Character Recognition (OCR)Video Captioning

MileBench: Benchmarking MLLMs in Long Context

2024-04-29 · Dingjie Song, Shunian Chen, Guiming Hardy Chen, Fei Yu 외

Despite the advancements and impressive performance of Multimodal Large Language Models (MLLMs) on benchmarks, their effectiveness in real-world, long-context, and multi-image tasks is unclear due to the benchmarks' limi…

BenchmarkingDiagnostic