Re-evaluating Automatic Metrics for Image Captioning
The task of generating natural language descriptions from images has received a lot of attention in recent years. Consequently, it is becoming increasingly important to evaluate such image captioning approaches in an automatic manner. In this paper, we provide an in-depth evaluation of the existing image captioning metrics through a series of carefully designed experiments. Moreover, we explore the utilization of the recently proposed Word Mover's Distance (WMD) document metric for the purpose of image captioning. Our findings outline the differences and/or similarities between metrics and their relative robustness by means of extensive correlation, accuracy and distraction based evaluations. Our results also demonstrate that WMD provides strong advantages over other metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningSimilar Papers 제목 키워드 기반
StyleM: Stylized Metrics for Image Captioning Built with Contrastive N-grams
In this paper, we build two automatic evaluation metrics for evaluating the association between a machine-generated caption and a ground truth stylized caption: OnlyStyle and StyleCIDEr.
Image CaptioningREO-Relevance, Extraness, Omission: A Fine-grained Evaluation for Image Captioning
Popular metrics used for evaluating image captioning systems, such as BLEU and CIDEr, provide a single score to gauge the system's overall effectiveness. This score is often not informative enough to indicate what specif…
Image CaptioningContrastive Semantic Similarity Learning for Image Captioning Evaluation with Intrinsic Auto-encoder
Automatically evaluating the quality of image captions can be very challenging since human language is quite flexible that there can be various expressions for the same meaning. Most of the current captioning metrics rel…
Image CaptioningRepresentation LearningSemantic SimilaritySemantic Textual Similarity+1VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily…
Image CaptioningText-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption Similarity
Automatic Audio Captioning (AAC) refers to the task of translating an audio sample into a natural language (NL) text that describes the audio events, source of the events and their relationships. Unlike NL text generatio…
Audio captioningImage CaptioningTAGText Generation