paper-with-me

홈 › Papers

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

2025-12-25 · Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki, Komei Sugiura arxiv

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics, that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Our project page is available at https://pearl.kinsta.page/.

📄 PDF Abstract BibTeX arXiv:2512.21582

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

2024-06-10 · Yebin Lee, Imseong Park, Myungjoo Kang

Most existing image captioning evaluation metrics focus on assigning a single numerical score to a caption by comparing it with reference captions. However, these methods do not provide an explanation for the assigned sc…

Image Captioning

HICEScore: A Hierarchical Metric for Image Captioning Evaluation

2024-07-26 · Zequn Zeng, JianQiao Sun, Hao Zhang, Tiansheng Wen 외

Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant …

DescriptiveImage Captioning

MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation

2026-02-10 · Nalin Srun, Parisa Rastin, Guénaël Cabanes, Lydia Boudjeloud Assala arxiv

We introduce MILE-RefHumEval, a reference-free framework for evaluating Large Language Models (LLMs) without ground-truth annotations or evaluator coordination. It leverages an ensemble of independently prompted evaluato…

Image Captioning

G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o

2024-12-18 · Tony Cheng Tong, Sirui He, Zhiwen Shao, Dit-yan Yeung

Evaluation metric of visual captioning is important yet not thoroughly explored. Traditional metrics like BLEU, METEOR, CIDEr, and ROUGE often miss semantic depth, while trained metrics such as CLIP-Score, PAC-S, and Pol…

Image CaptioningVideo Captioning

CLIPScore: A Reference-free Evaluation Metric for Image Captioning

2021-04-18 · EMNLP 2021 11 · Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 외

Image captioning has conventionally relied on reference-based automatic evaluations, where machine captions are compared against captions written by humans. This is in contrast to the reference-free manner in which human…

Hallucination Pair-wise Detection (1-ref)Hallucination Pair-wise Detection (4-ref)Human Judgment ClassificationHuman Judgment Correlation+1